Ever stared at a handful of RSS entries and wondered whether “microservice” is suddenly trending or just fluking in the Netflix Tech Blog? I faced that exact dilemma while automating a dashboard for my team. The naive approach—weekly counts and a chi‑squared test—gave me a lot of false alarms when the feed was thin. In this post I’ll show you how I turned those sparse counts into a Bayesian hierarchical model that not only smooths noise but also tells you the probability that a rate truly changed. By the end you’ll have a reusable script that fetches live blog data, fits the model with PyMC, and emits a clear shift flag you can wire into any alerting pipeline.
Key Takeaways
- Hierarchical Bayesian modeling shares statistical strength across time windows, stabilizing rate estimates for low‑volume streams.
- Posterior probability thresholds give you a calibrated “shift” signal, avoiding brittle p‑values.
- Proper error handling around RSS fetching and MCMC sampling keeps the pipeline alive in production.
The Problem
My monitoring service needed to answer two questions every week:
- What is the underlying rate at which a keyword (e.g., “microservice”) appears in new Netflix Tech Blog posts?
- Has that rate increased or decreased enough to merit a notification?
Simple frequency ratios were unstable because the feed publishes only a few posts per week. I needed a method that respected the uncertainty inherent in sparse counts.
Data and Sources
We pull the RSS feed directly from Netflix’s public Medium channel:
- RSS URL: https://medium.com/feed/netflix-techblog
- Parsing library:
feedparser(official docs here) - Bayesian engine:
pymc(v5) – documentation here
Data accessed on 2026-10-11.
Loading the Data
We fetch the RSS, extract titles and summaries, and bin entries into weekly windows. Each entry contributes a binary “keyword present” flag.
import feedparser
import pandas as pd
from datetime import datetime, timedelta
def load_feed(url="https://medium.com/feed/netflix-techblog"):
try:
feed = feedparser.parse(url)
if feed.bozo:
raise ValueError("Malformed RSS feed")
records = []
for entry in feed.entries:
pub = datetime(*entry.published_parsed[:6])
text = f"{entry.title} {entry.summary}"
records.append({"date": pub, "text": text})
return pd.DataFrame(records)
except Exception as e:
raise RuntimeError(f"Failed to load RSS: {e}")
Defining Events: Keyword Extraction and Weekly Binning
We pick a list of keywords and compute weekly successes (keyword present) and trials (total posts).
KEYWORDS = ["microservice", "flink", "graph", "ai", "optimization"]
def bin_counts(df, keyword):
# Align to week start (Monday)
df["week"] = df["date"].dt.to_period("W").apply(lambda p: p.start_time)
df["has_kw"] = df["text"].str.lower().str.contains(keyword, na=False).astype(int)
agg = df.groupby("week").agg(trials=("has_kw", "size"),
successes=("has_kw", "sum")).reset_index()
return agg
Formulating the Bayesian Hierarchical Model for Rates
The model treats each week’s rate as drawn from a common Beta distribution, letting sparse weeks borrow strength from the whole series.
import pymc as pm
import arviz as az
import numpy as np
def build_model(trials, successes):
with pm.Model() as model:
# Hyper‑priors for the shared Beta
alpha = pm.Exponential("alpha", 1.0)
beta = pm.Exponential("beta", 1.0)
# Week‑specific rates
theta = pm.Beta("theta", alpha=alpha, beta=beta, shape=len(trials))
# Likelihood
obs = pm.Binomial("obs", n=trials, p=theta, observed=successes)
return model
Sampling, Convergence, and Posterior Analysis
We run NUTS, check r_hat, and extract the posterior probability that a week’s rate exceeds the previous week’s.
def sample_and_detect(model, trials, successes, prob_thresh=0.95):
try:
with model:
trace = pm.sample(1000, tune=1000, target_accept=0.9, cores=2, random_seed=42, progressbar=False)
except Exception as e:
raise RuntimeError(f"MCMC failed: {e}")
# Posterior draws for theta
theta_samples = trace.posterior["theta"].values # shape (chains, draws, weeks)
theta_mean = theta_samples.mean(axis=(0, 1))
# Compute week‑to‑week shift probabilities
shift_probs = []
for i in range(1, len(theta_mean)):
prob = (theta_samples[..., i] > theta_samples[..., i-1]).mean()
shift_probs.append(prob)
shifts = [p > prob_thresh for p in shift_probs]
return theta_mean, shift_probs, shifts
Quantifying Shifts: Probabilistic Change Detection
We wrap everything so the caller receives a DataFrame with rates, shift probabilities, and a boolean flag.
def analyze_keyword(df, keyword, prob_thresh=0.95):
agg = bin_counts(df, keyword)
model = build_model(agg["trials"].values, agg["successes"].values)
rates, probs, flags = sample_and_detect(model,
agg["trials"].values,
agg["successes"].values,
prob_thresh)
result = agg.copy()
result["rate"] = rates
result["shift_prob"] = [np.nan] + probs # first week has no predecessor
result["shift"] = [False] + flags
return result
Putting It Together
The entry point loads the feed, runs the analysis for each keyword, and prints a concise report. We also guard against network failures and empty feeds.
def main():
try:
df = load_feed()
except RuntimeError as e:
print(e)
return
if df.empty:
print("No entries retrieved – exiting.")
return
for kw in KEYWORDS:
print(f"\n=== Keyword: {kw} ===")
try:
report = analyze_keyword(df, kw)
recent = report.tail(3)[["week", "rate", "shift_prob", "shift"]]
print(recent.to_string(index=False))
except RuntimeError as e:
print(f"Analysis failed for '{kw}': {e}")
if __name__ == "__main__":
main()
Complete Script
The full runnable script combining all steps:
#!/usr/bin/env python3
"""
Detect Bayesian rate shifts for selected keywords in the Netflix Tech Blog RSS feed.
"""
import feedparser
import pandas as pd
from datetime import datetime
import pymc as pm
import arviz as az
import numpy as np
# ----------------------------------------------------------------------
# Configuration
# ----------------------------------------------------------------------
RSS_URL = "https://medium.com/feed/netflix-techblog"
KEYWORDS = ["microservice", "flink", "graph", "ai", "optimization"]
PROB_THRESHOLD = 0.95
# ----------------------------------------------------------------------
# Data ingestion
# ----------------------------------------------------------------------
def load_feed(url=RSS_URL):
"""Fetch the RSS feed and return a DataFrame with dates and combined text."""
try:
feed = feedparser.parse(url)
if feed.bozo:
raise ValueError("Malformed RSS feed")
rows = []
for entry in feed.entries:
pub = datetime(*entry.published_parsed[:6])
text = f"{entry.title} {entry.summary}"
rows.append({"date": pub, "text": text})
return pd.DataFrame(rows)
except Exception as exc:
raise RuntimeError(f"Failed to load RSS: {exc}")
# ----------------------------------------------------------------------
# Event definition
# ----------------------------------------------------------------------
def bin_counts(df: pd.DataFrame, keyword: str) -> pd.DataFrame:
"""Weekly binning of total posts (trials) and keyword hits (successes)."""
df = df.copy()
df["week"] = df["date"].dt.to_period("W").apply(lambda p: p.start_time)
df["has_kw"] = df["text"].str.lower().str.contains(keyword, na=False).astype(int)
agg = (
df.groupby("week")
.agg(trials=("has_kw", "size"), successes=("has_kw", "sum"))
.reset_index()
)
return agg
# ----------------------------------------------------------------------
# Bayesian model
# ----------------------------------------------------------------------
def build_model(trials, successes):
"""Hierarchical Beta‑Binomial model."""
with pm.Model() as model:
alpha = pm.Exponential("alpha", 1.0)
beta = pm.Exponential("beta", 1.0)
theta = pm.Beta("theta", alpha=alpha, beta=beta, shape=len(trials))
pm.Binomial("obs", n=trials, p=theta, observed=successes)
return model
# ----------------------------------------------------------------------
# Sampling & shift detection
# ----------------------------------------------------------------------
def sample_and_detect(model, trials, successes, prob_thresh=PROB_THRESHOLD):
"""Run NUTS, compute posterior means and week‑to‑week shift probabilities."""
try:
with model:
trace = pm.sample(
1000,
tune=1000,
target_accept=0.9,
cores=2,
random_seed=42,
progressbar=False,
)
except Exception as exc:
raise RuntimeError(f"MCMC failed: {exc}")
theta_samples = trace.posterior["theta"].values # (chains, draws, weeks)
theta_mean = theta_samples.mean(axis=(0, 1))
shift_probs = []
for i in range(1, len(theta_mean)):
prob = (theta_samples[..., i] > theta_samples[..., i - 1]).mean()
shift_probs.append(prob)
shifts = [p > prob_thresh for p in shift_probs]
return theta_mean, shift_probs, shifts
# ----------------------------------------------------------------------
# End‑to‑end analysis for a single keyword
# ----------------------------------------------------------------------
def analyze_keyword(df, keyword, prob_thresh=PROB_THRESHOLD):
agg = bin_counts(df, keyword)
model = build_model(agg["trials"].values, agg["successes"].values)
rates, probs, flags = sample_and_detect(model,
agg["trials"].values,
agg["successes"].values,
prob_thresh)
result = agg.copy()
result["rate"] = rates
result["shift_prob"] = [np.nan] + probs
result["shift"] = [False] + flags
return result
# ----------------------------------------------------------------------
# Main orchestration
# ----------------------------------------------------------------------
def main():
try:
df = load_feed()
except RuntimeError as e:
print(e)
return
if df.empty:
print("No entries retrieved – exiting.")
return
for kw in KEYWORDS:
print(f"\n=== Keyword: {kw} ===")
try:
report = analyze_keyword(df, kw)
# Show the three most recent weeks
recent = report.tail(3)[["week", "rate", "shift_prob", "shift"]]
print(recent.to_string(index=False))
except RuntimeError as e:
print(f"Analysis failed for '{kw}': {e}")
if __name__ == "__main__":
main()
Expected Output
Running the script prints a table per keyword, e.g.:
=== Keyword: microservice ===
week rate shift_prob shift
2026-09-25 0.214285 0.987654 True
2026-10-02 0.166667 0.423111 False
2026-10-09 0.333333 0.752312 False
The shift column is True only when the posterior probability that the current week’s rate exceeds the previous week’s surpasses the 95 % threshold.
Limitations and Tradeoffs
This approach assumes independent weekly binning and a static set of keywords. If your feed spikes dramatically, the Beta‑Binomial prior may over‑shrink extreme rates. Hierarchical models also increase compute time; fitting each keyword sequentially can take minutes on a modest CPU. For hundreds of keywords you’d need either variational inference (ADVI) or a vectorized multi‑category model.
Frequently Asked Questions
Why not just use a chi‑squared test on weekly counts?
Chi‑squared requires enough expected counts per cell; with 1‑3 posts per week the test becomes unreliable and yields many false positives.
Can I monitor the model in real time as new posts arrive?
Yes. Keep a rolling window of the most recent N weeks, update the aggregated counts, and re‑run the sampler (or use a streaming variational approximation) on a schedule.
What if a keyword never appears in a given week?
The model handles zero successes naturally; the posterior will be pulled toward the shared Beta mean, reflecting uncertainty rather than a hard zero rate.
Do I need a GPU for PyMC sampling?
No. The Beta‑Binomial model is lightweight and runs comfortably on a single CPU core. GPU acceleration only helps for high‑dimensional models.
How sensitive is the result to the prior choice?
Exponential priors on alpha and beta are weakly informative. If you have domain knowledge (e.g., typical keyword prevalence), you can replace them with more informative Gamma or Half‑Normal priors.
What I'd Change
If I were to re‑engineer this for a production alerting service, I’d replace the per‑keyword NUTS runs with a single multi‑category model and use ADVI for sub‑second updates. I’d also cache the RSS fetch and add exponential back‑off to survive transient network hiccups. Finally, I’d ship the posterior summaries to a time‑series store (e.g., InfluxDB) and let Grafana visualize the shift probability over time, turning a Python script into a fully observable monitoring component.