When I first started building LLM agents for production, a common trap I fell into was designing them with static, hardcoded toolkits. It felt intuitive: define the tools, give the agent access, and let it choose. But as external information sources, like critical third-party service updates or market news feeds, began to evolve rapidly, my agents became brittle. They couldn't adapt to new types of information or unexpected operational needs because their tool selection logic was tied to a fixed understanding of the world. This post is for you if you're an engineer looking to move beyond this limitation, to empower your agents to intelligently 'self-tool' by analyzing real-time external feeds. We'll build a system that dynamically discovers and applies the most relevant tools, ensuring your agents remain agile and effective in truly dynamic environments.
Key Takeaways
- Implement a robust RSS parsing and data extraction pipeline using
feedparserand Pydantic for resilient data ingestion and schema enforcement. - Leverage semantic embeddings (e.g.,
sentence-transformers) to classify incoming feed entries into action-oriented categories like "Product Update" or "Security Notice," moving beyond brittle keyword matching. - Architect a dynamic tool dispatcher that maps classified content types to specific agent tools, enabling flexible agent behavior without requiring code changes for new content types.
- Integrate comprehensive error handling and fallback mechanisms to ensure agent resilience against network issues, malformed data, and unclassifiable content.
The Problem
In a world where critical information constantly streams from diverse external sources, relying on an agent with a predefined, static set of tools is a significant vulnerability. Imagine an agent tasked with monitoring security advisories or product updates from a crucial vendor. If a new type of advisory emerges – perhaps related to a novel attack vector – and your agent's tool selection logic doesn't anticipate it, it will fail to invoke the correct response tool. Traditional approaches, often based on simple keyword matching or rigid conditional logic, lead to agents that are not only difficult to maintain but also inherently reactive, constantly requiring manual updates to their tool manifests. My goal was to build an agent that could proactively understand the *intent* behind new content and select the right tool, even if that content structure was unforeseen.
Data and Sources
For this exploration, we're going to tap into a real-world, dynamic data stream: the Discord Engineering Blog RSS Feed. This provides a rich, varied source of content, from patch notes to feature announcements and security updates, perfect for testing our agent's adaptability.
- Discord Engineering Blog RSS Feed:
https://discord.com/blog/rss.xml - Feed Parsing Library:
feedparser - Data Validation Library: Pydantic V2
- Semantic Embedding Library:
sentence-transformers
Data accessed on 2026-10-27.
Step 1: Ingesting and Structuring Dynamic Feed Data
The first hurdle with any external feed is its inherent messiness. RSS feeds, while structured, can still vary in their content and reliability. Titles might be missing, summaries could be truncated, or dates might be in unexpected formats. To build a resilient agent, we need a robust ingestion pipeline that not only fetches the data but also enforces a clean, predictable schema. This is where Pydantic becomes indispensable for data validation and transformation.
My approach was to combine feedparser for fetching and initial parsing with Pydantic for strict schema definition and validation. This ensures that every piece of data our agent processes conforms to an expected structure, preventing downstream errors.
import feedparser
from pydantic import BaseModel, Field, HttpUrl, ValidationError
from typing import Optional, List
from datetime import datetime
class FeedEntry(BaseModel):
title: str = Field(..., description="Title of the blog post")
link: HttpUrl = Field(..., description="URL of the blog post")
published: datetime = Field(..., description="Publication date of the blog post")
summary: Optional[str] = Field(None, description="Summary or description of the blog post")
@classmethod
def from_feedparser_entry(cls, entry):
# Handle potential missing fields and convert types
published_parsed = entry.get('published_parsed')
published_dt = datetime(*published_parsed[:6]) if published_parsed else datetime.min
return cls(
title=entry.get('title', 'Untitled'),
link=entry.get('link', 'http://example.com'), # Provide a fallback for HttpUrl
published=published_dt,
summary=entry.get('summary')
)
def fetch_and_parse_feed(url: str) -> List[FeedEntry]:
try:
feed = feedparser.parse(url)
if feed.bozo:
print(f"Warning: Feed may be malformed for {url}. Reason: {feed.bozo_exception}")
parsed_entries = []
for entry_data in feed.entries:
try:
parsed_entries.append(FeedEntry.from_feedparser_entry(entry_data))
except ValidationError as e:
print(f"Skipping malformed entry: {e.errors()} in {entry_data.get('link')}")
return parsed_entries
except Exception as e:
print(f"Error fetching or parsing feed from {url}: {e}")
return []
# Example usage:
# discord_feed_url = "https://discord.com/blog/rss.xml"
# recent_entries = fetch_and_parse_feed(discord_feed_url)
# for entry in recent_entries[:2]:
# print(f"Title: {entry.title}\nLink: {entry.link}\nPublished: {entry.published}\n")
In this snippet, the FeedEntry Pydantic model defines our expected data structure. The from_feedparser_entry class method handles the conversion from feedparser's raw dictionary output, including type conversions and sensible fallbacks for missing data. The fetch_and_parse_feed function then orchestrates the fetching and validation, gracefully skipping malformed entries rather than crashing the entire pipeline. This robust ingestion is the bedrock for an adaptive agent.
Step 2: Semantic Classification for Contextual Understanding
Once we have clean, structured data, the next challenge is to understand its *meaning* at a deeper level than simple keyword searches. An agent needs to know if an entry is a "Product Update" or a "Security Notice" to choose the right tool. My previous experiences taught me that keyword-based rules are brittle and fail when content phrasing changes. This is where semantic embeddings shine for real-time feature extraction.
We'll use a pre-trained sentence-transformers model to convert the titles and summaries of our feed entries into dense vector representations. Then, we'll define a few "anchor" embeddings for our desired categories (e.g., "Product Update", "Security Alert", "Patch Notes") and classify each new entry based on its cosine similarity to these anchors. This allows for flexible, context-aware classification.
from sentence_transformers import SentenceTransformer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np
# Load a pre-trained sentence transformer model
# Using a lightweight model suitable for production inference
model = SentenceTransformer('all-MiniLM-L6-v2')
# Define our target categories and their "anchor" descriptions
category_definitions = {
"Product Update": "New feature announcement, product improvements, new release, upcoming changes.",
"Security Notice": "Security vulnerability, exploit, patch, data breach, incident response.",
"Patch Notes": "Bug fixes, performance improvements, technical updates, changelog.",
"Company News": "Blog post, company announcement, event, general news."
}
# Pre-compute embeddings for our category definitions
category_embeddings = {
cat: model.encode(desc, convert_to_tensor=True)
for cat, desc in category_definitions.items()
}
def classify_entry_semantically(entry: FeedEntry) -> str:
combined_text = f"{entry.title}. {entry.summary if entry.summary else ''}"
entry_embedding = model.encode(combined_text, convert_to_tensor=True)
similarities = {}
for category, cat_emb in category_embeddings.items():
# Ensure both tensors are on the same device (CPU in this case, as no .cuda() is called)
similarities[category] = cosine_similarity(
entry_embedding.cpu().reshape(1, -1),
cat_emb.cpu().reshape(1, -1)
)[0][0]
# Select the category with the highest similarity
best_category = max(similarities, key=similarities.get)
# Thresholding for "Unknown" category
if similarities[best_category] < 0.5: # A configurable threshold
return "Unknown"
return best_category
# Example usage:
# sample_entry = FeedEntry(
# title="Discord Update: September 25, 2026",
# link="https://discord.com/blog/update-september-25-2026",
# published=datetime(2026, 9, 25),
# summary="Here's the Discord