Skip to content

Beyond Simple Prompts: Architecting Self-Correcting LLM Agents for Robust Structured Data Extraction from Dynamic Feeds

Beyond Simple Prompts: Architecting Self-Correcting LLM Agents for Robust Structured Data Extraction from Dynamic Feeds
Learn to architect self-correcting prompt strategies that leverage schema validation and iterative refinement to reliably extract structured data from diverse, dynamic content streams in production.

Building LLM applications that consistently produce structured, valid output from varied, dynamic input is a significant challenge. If you've ever tried to coax a large language model into returning clean JSON from messy web content, you've likely encountered the frustration of malformed syntax, missing fields, or hallucinated values. This isn't just an inconvenience; in production, it leads to cascading failures, invalid data, and a constant need for manual intervention. This post is for engineers grappling with these issues, aiming to move beyond static, fire-and-forget prompts. I'll show you how to implement an adaptive, self-correction mechanism that ensures robust, schema-compliant data extraction, significantly improving system reliability. We'll build a system that can intelligently parse the dynamic content of the Stripe Blog RSS feed, validate its own output, and iteratively refine it until it meets our strict data requirements.

Key Takeaways

  • Schema validation, powered by tools like Pydantic, is indispensable as a feedback mechanism for LLM outputs.
  • Iterative prompting, where validation errors inform subsequent LLM calls, enables robust self-correction.
  • Explicitly providing the desired schema (e.g., JSON Schema) and past errors in the prompt significantly guides the LLM towards correct output.
  • Normalizing dynamic content before LLM processing is crucial for reducing ambiguity and improving extraction accuracy.
  • A well-defined retry strategy with clear limits prevents infinite loops and manages API costs effectively.

The Problem: Inconsistent LLM Output in Production

Imagine you're building a content aggregation service that pulls articles from various engineering blogs, including the Stripe Tech Blog, to populate a knowledge base. Your goal is to extract specific structured information from each article entry: its title, a concise summary, the publication date, and an estimated reading time. The challenge? RSS feeds are dynamic, and article summaries can vary wildly in length and content. A simple, one-shot prompt to an LLM might work most of the time, but inevitably, you'll encounter cases where the LLM omits a field, provides a date in an unexpected format, or, worst of all, generates malformed JSON. This isn't just a minor annoyance; it halts your downstream processing pipelines, requiring manual intervention to fix or re-run. My experience has shown that relying solely on a single prompt for structured output in a dynamic environment is a recipe for operational headaches. We need a system that can detect its own errors and fix them.

Data and Sources

For this demonstration, we'll be extracting information from the official Stripe Blog RSS feed. This feed provides a diverse set of article summaries, ranging from technical deep dives to product announcements, making it an excellent real-world test case for our self-correcting agent.

Data accessed on 2024-07-29.

Step 1 — Ingesting and Normalizing Dynamic Content

The first hurdle is getting the raw content into a usable format. RSS feeds, while structured, can be inconsistent. Summaries might contain HTML, dates can be in various formats, and some fields might be missing. Our goal here is to fetch the feed, parse it, and extract the most relevant text content in a normalized way, preparing it for the LLM.

What this step addresses:

This step tackles the variability of external data sources. We need to reliably fetch the RSS feed, handle potential network issues, and then extract the core textual content in a clean, consistent manner. The `feedparser` library is excellent for this, abstracting away much of the RSS parsing complexity.

How it works:

We use `feedparser.parse()` to fetch and parse the XML content from the RSS URL. It returns a dictionary-like object that's easy to navigate. For each entry, I focus on `title`, `link`, `summary`, and `published`. The `summary` often contains HTML, so I'll strip that out to present clean text to the LLM. I'm also ensuring we only grab a few recent entries to keep the example focused.

import feedparser
import re

def fetch_and_normalize_feed(rss_url: str, num_entries: int = 3) -> list[dict]:
    """
    Fetches an RSS feed, parses it, and normalizes the content for LLM processing.
    Strips HTML from summaries and extracts key fields.
    """
    try:
        feed = feedparser.parse(rss_url)
        if feed.bozo:
            print(f"Warning: RSS feed parsing issues for {rss_url}: {feed.bozo_exception}")

        normalized_entries = []
        for entry in feed.entries[:num_entries]:
            summary_clean = re.sub(r'<.*?>', '', entry.get('summary', '')).strip()
            # Fallback for description if summary is empty
            if not summary_clean and 'description' in entry:
                summary_clean = re.sub(r'<.*?>', '', entry.get('description', '')).strip()

            normalized_entries.append({
                "title": entry.get('title', 'No Title'),
                "link": entry.get('link', 'No Link'),
                "summary": summary_clean,
                "published": entry.get('published', 'No Date'),
                "raw_text": f"Title: {entry.get('title', 'No Title')}\nSummary: {summary_clean}\nPublished: {entry.get('published', 'No Date')}"
            })
        return normalized_entries
    except Exception as e:
        print(f"Error fetching or parsing RSS feed: {e}")
        return []

# Example usage (will be part of complete script later)
# stripe_feed_url = "https://stripe.com/blog/feed.rss"
# normalized_data = fetch_and_normalize_feed(stripe_feed_url)
# for item in normalized_data:
#     print(f"--- Article ---\nTitle: {item['title']}\nSummary: {item['summary'][:100]}...\n")

Step 2 — Defining the Target Structured Output Schema

For our LLM to produce reliable structured data, it needs a clear "contract." This contract is our schema. Pydantic is an absolute game-changer here, allowing us to define this schema using standard Python type hints. It not only provides strong validation but can also generate a JSON Schema representation that's perfect for instructing LLMs.

What this step addresses:

This defines the exact structure and types we expect from the LLM. Without a precise schema, the LLM has too much freedom, leading to inconsistent outputs. Pydantic ensures that whatever comes out of the LLM can be programmatically checked against our expectations.

How it works:

I define a `BlogArticle` Pydantic model with fields like `title`, `summary`, `published_date` (as a `datetime` for strict parsing), and `reading_time_minutes` (an `int`). Crucially, I include a `model_json_schema()` method call to get a JSON Schema representation. This schema will be embedded directly into our LLM prompt, telling the model exactly what output format we expect.

from pydantic import BaseModel, Field, ValidationError
from datetime import datetime

class BlogArticle(BaseModel):
    """
    Schema for a blog article extracted from an RSS feed.
    """
    title: str = Field(..., description="The title of the blog post.")
    summary: str = Field(..., description="A concise summary of the blog post, up to 200 words.")
    published_date: datetime = Field(..., description="The publication date of the article in YYYY-MM-DD format.")
    reading_time_minutes: int = Field(..., description="Estimated reading time of the article in minutes (integer).")

    class Config:
        json_schema_extra = {
            "example": {
                "title": "My Awesome Blog Post",
                "summary": "This is a brief summary of the blog post content.",
                "published_date": "2023-10-26",
                "reading_time_minutes": 5
            }
        }

# Get the JSON schema for use in prompts
BLOG_ARTICLE_SCHEMA = BlogArticle.model_json_schema()

# Example of how the schema looks (for explanation, not in final script output)
# import json
# print(json.dumps(BLOG_ARTICLE_SCHEMA, indent=2))

Step 3 — Initial Prompting for Structured Extraction

Now that we have our content and our schema, it's time to instruct the LLM. The initial prompt is critical for guiding the LLM towards the desired JSON output. I've found that being extremely explicit about the output format and including the JSON schema directly in the prompt yields the best results.

What this step addresses:

This is where we translate our data extraction task into instructions for the LLM. The sub-problem is getting the LLM to understand our requirements for structured output, specifically JSON that adheres to our defined Pydantic schema.

How it works:

I construct a system message that sets the context (you are a data extraction assistant) and provides the JSON schema. The user message then contains the specific article content to be processed. I'm using a placeholder `LLM_API_CALL` function for demonstration, as the actual API integration (e.g., OpenAI, Anthropic, or local `llama.cpp`) is beyond the scope of this particular post's focus but would be implemented here. The key is to tell the LLM to "respond ONLY with a JSON object" and "strictly adhere to the following JSON schema."

import json
import os
import time

# Placeholder for actual LLM API call
# In a real scenario, this would integrate with OpenAI, Anthropic, etc.
# For simplicity, we'll simulate an LLM response here.
# For a more advanced local setup, you might refer to:
# http://blogs.mausamadhikari.com.np/2026/10/from-cloud-to-core-architecting-local.html
def LLM_API_CALL(messages: list[dict], model: str = "gpt

إرسال تعليق

Hi! How can we help you? Send us a message and we'll get back to you.