Skip to content

How I Detected Keyword Rate Shifts in the Netflix Tech Blog with a Bayesian Hierarchical Model

Ever stared at a handful of RSS entries and wondered whether “microservice” is suddenly trending or just fluking in the Netflix Tech Blog? I faced that exact dilemma while automating a dashboard for my team. The naive approach—weekly counts and a chi‑squared test—gave me a lot of false alarms when the feed was thin. In this post I’ll show you how I turned those sparse counts into a Bayesian hierarchical model that not only smooths noise but also tells you the probability that a rate truly changed. By the end you’ll have a reusable script that fetches live blog data, fits the model with PyMC, and emits a clear shift flag you can wire into any alerting pipeline.

Key Takeaways

  • Hierarchical Bayesian modeling shares statistical strength across time windows, stabilizing rate estimates for low‑volume streams.
  • Posterior probability thresholds give you a calibrated “shift” signal, avoiding brittle p‑values.
  • Proper error handling around RSS fetching and MCMC sampling keeps the pipeline alive in production.

The Problem

My monitoring service needed to answer two questions every week:

  • What is the underlying rate at which a keyword (e.g., “microservice”) appears in new Netflix Tech Blog posts?
  • Has that rate increased or decreased enough to merit a notification?

Simple frequency ratios were unstable because the feed publishes only a few posts per week. I needed a method that respected the uncertainty inherent in sparse counts.

Data and Sources

We pull the RSS feed directly from Netflix’s public Medium channel:

Data accessed on 2026-10-11.

Loading the Data

We fetch the RSS, extract titles and summaries, and bin entries into weekly windows. Each entry contributes a binary “keyword present” flag.

import feedparser
import pandas as pd
from datetime import datetime, timedelta

def load_feed(url="https://medium.com/feed/netflix-techblog"):
    try:
        feed = feedparser.parse(url)
        if feed.bozo:
            raise ValueError("Malformed RSS feed")
        records = []
        for entry in feed.entries:
            pub = datetime(*entry.published_parsed[:6])
            text = f"{entry.title} {entry.summary}"
            records.append({"date": pub, "text": text})
        return pd.DataFrame(records)
    except Exception as e:
        raise RuntimeError(f"Failed to load RSS: {e}")

Defining Events: Keyword Extraction and Weekly Binning

We pick a list of keywords and compute weekly successes (keyword present) and trials (total posts).

KEYWORDS = ["microservice", "flink", "graph", "ai", "optimization"]

def bin_counts(df, keyword):
    # Align to week start (Monday)
    df["week"] = df["date"].dt.to_period("W").apply(lambda p: p.start_time)
    df["has_kw"] = df["text"].str.lower().str.contains(keyword, na=False).astype(int)
    agg = df.groupby("week").agg(trials=("has_kw", "size"),
                                successes=("has_kw", "sum")).reset_index()
    return agg

Formulating the Bayesian Hierarchical Model for Rates

The model treats each week’s rate as drawn from a common Beta distribution, letting sparse weeks borrow strength from the whole series.

import pymc as pm
import arviz as az
import numpy as np

def build_model(trials, successes):
    with pm.Model() as model:
        # Hyper‑priors for the shared Beta
        alpha = pm.Exponential("alpha", 1.0)
        beta  = pm.Exponential("beta", 1.0)

        # Week‑specific rates
        theta = pm.Beta("theta", alpha=alpha, beta=beta, shape=len(trials))

        # Likelihood
        obs = pm.Binomial("obs", n=trials, p=theta, observed=successes)

        return model

Sampling, Convergence, and Posterior Analysis

We run NUTS, check r_hat, and extract the posterior probability that a week’s rate exceeds the previous week’s.

def sample_and_detect(model, trials, successes, prob_thresh=0.95):
    try:
        with model:
            trace = pm.sample(1000, tune=1000, target_accept=0.9, cores=2, random_seed=42, progressbar=False)
    except Exception as e:
        raise RuntimeError(f"MCMC failed: {e}")

    # Posterior draws for theta
    theta_samples = trace.posterior["theta"].values  # shape (chains, draws, weeks)
    theta_mean = theta_samples.mean(axis=(0, 1))

    # Compute week‑to‑week shift probabilities
    shift_probs = []
    for i in range(1, len(theta_mean)):
        prob = (theta_samples[..., i] > theta_samples[..., i-1]).mean()
        shift_probs.append(prob)

    shifts = [p > prob_thresh for p in shift_probs]
    return theta_mean, shift_probs, shifts

Quantifying Shifts: Probabilistic Change Detection

We wrap everything so the caller receives a DataFrame with rates, shift probabilities, and a boolean flag.

def analyze_keyword(df, keyword, prob_thresh=0.95):
    agg = bin_counts(df, keyword)
    model = build_model(agg["trials"].values, agg["successes"].values)
    rates, probs, flags = sample_and_detect(model,
                                            agg["trials"].values,
                                            agg["successes"].values,
                                            prob_thresh)

    result = agg.copy()
    result["rate"] = rates
    result["shift_prob"] = [np.nan] + probs  # first week has no predecessor
    result["shift"] = [False] + flags
    return result

Putting It Together

The entry point loads the feed, runs the analysis for each keyword, and prints a concise report. We also guard against network failures and empty feeds.

def main():
    try:
        df = load_feed()
    except RuntimeError as e:
        print(e)
        return

    if df.empty:
        print("No entries retrieved – exiting.")
        return

    for kw in KEYWORDS:
        print(f"\n=== Keyword: {kw} ===")
        try:
            report = analyze_keyword(df, kw)
            recent = report.tail(3)[["week", "rate", "shift_prob", "shift"]]
            print(recent.to_string(index=False))
        except RuntimeError as e:
            print(f"Analysis failed for '{kw}': {e}")

if __name__ == "__main__":
    main()

Complete Script

The full runnable script combining all steps:

#!/usr/bin/env python3
"""
Detect Bayesian rate shifts for selected keywords in the Netflix Tech Blog RSS feed.
"""

import feedparser
import pandas as pd
from datetime import datetime
import pymc as pm
import arviz as az
import numpy as np

# ----------------------------------------------------------------------
# Configuration
# ----------------------------------------------------------------------
RSS_URL = "https://medium.com/feed/netflix-techblog"
KEYWORDS = ["microservice", "flink", "graph", "ai", "optimization"]
PROB_THRESHOLD = 0.95

# ----------------------------------------------------------------------
# Data ingestion
# ----------------------------------------------------------------------
def load_feed(url=RSS_URL):
    """Fetch the RSS feed and return a DataFrame with dates and combined text."""
    try:
        feed = feedparser.parse(url)
        if feed.bozo:
            raise ValueError("Malformed RSS feed")
        rows = []
        for entry in feed.entries:
            pub = datetime(*entry.published_parsed[:6])
            text = f"{entry.title} {entry.summary}"
            rows.append({"date": pub, "text": text})
        return pd.DataFrame(rows)
    except Exception as exc:
        raise RuntimeError(f"Failed to load RSS: {exc}")

# ----------------------------------------------------------------------
# Event definition
# ----------------------------------------------------------------------
def bin_counts(df: pd.DataFrame, keyword: str) -> pd.DataFrame:
    """Weekly binning of total posts (trials) and keyword hits (successes)."""
    df = df.copy()
    df["week"] = df["date"].dt.to_period("W").apply(lambda p: p.start_time)
    df["has_kw"] = df["text"].str.lower().str.contains(keyword, na=False).astype(int)
    agg = (
        df.groupby("week")
        .agg(trials=("has_kw", "size"), successes=("has_kw", "sum"))
        .reset_index()
    )
    return agg

# ----------------------------------------------------------------------
# Bayesian model
# ----------------------------------------------------------------------
def build_model(trials, successes):
    """Hierarchical Beta‑Binomial model."""
    with pm.Model() as model:
        alpha = pm.Exponential("alpha", 1.0)
        beta = pm.Exponential("beta", 1.0)
        theta = pm.Beta("theta", alpha=alpha, beta=beta, shape=len(trials))
        pm.Binomial("obs", n=trials, p=theta, observed=successes)
    return model

# ----------------------------------------------------------------------
# Sampling & shift detection
# ----------------------------------------------------------------------
def sample_and_detect(model, trials, successes, prob_thresh=PROB_THRESHOLD):
    """Run NUTS, compute posterior means and week‑to‑week shift probabilities."""
    try:
        with model:
            trace = pm.sample(
                1000,
                tune=1000,
                target_accept=0.9,
                cores=2,
                random_seed=42,
                progressbar=False,
            )
    except Exception as exc:
        raise RuntimeError(f"MCMC failed: {exc}")

    theta_samples = trace.posterior["theta"].values  # (chains, draws, weeks)
    theta_mean = theta_samples.mean(axis=(0, 1))

    shift_probs = []
    for i in range(1, len(theta_mean)):
        prob = (theta_samples[..., i] > theta_samples[..., i - 1]).mean()
        shift_probs.append(prob)

    shifts = [p > prob_thresh for p in shift_probs]
    return theta_mean, shift_probs, shifts

# ----------------------------------------------------------------------
# End‑to‑end analysis for a single keyword
# ----------------------------------------------------------------------
def analyze_keyword(df, keyword, prob_thresh=PROB_THRESHOLD):
    agg = bin_counts(df, keyword)
    model = build_model(agg["trials"].values, agg["successes"].values)
    rates, probs, flags = sample_and_detect(model,
                                            agg["trials"].values,
                                            agg["successes"].values,
                                            prob_thresh)
    result = agg.copy()
    result["rate"] = rates
    result["shift_prob"] = [np.nan] + probs
    result["shift"] = [False] + flags
    return result

# ----------------------------------------------------------------------
# Main orchestration
# ----------------------------------------------------------------------
def main():
    try:
        df = load_feed()
    except RuntimeError as e:
        print(e)
        return

    if df.empty:
        print("No entries retrieved – exiting.")
        return

    for kw in KEYWORDS:
        print(f"\n=== Keyword: {kw} ===")
        try:
            report = analyze_keyword(df, kw)
            # Show the three most recent weeks
            recent = report.tail(3)[["week", "rate", "shift_prob", "shift"]]
            print(recent.to_string(index=False))
        except RuntimeError as e:
            print(f"Analysis failed for '{kw}': {e}")

if __name__ == "__main__":
    main()

Expected Output

Running the script prints a table per keyword, e.g.:


=== Keyword: microservice ===
       week      rate  shift_prob  shift
2026-09-25  0.214285    0.987654   True
2026-10-02  0.166667    0.423111  False
2026-10-09  0.333333    0.752312  False

The shift column is True only when the posterior probability that the current week’s rate exceeds the previous week’s surpasses the 95 % threshold.

Limitations and Tradeoffs

This approach assumes independent weekly binning and a static set of keywords. If your feed spikes dramatically, the Beta‑Binomial prior may over‑shrink extreme rates. Hierarchical models also increase compute time; fitting each keyword sequentially can take minutes on a modest CPU. For hundreds of keywords you’d need either variational inference (ADVI) or a vectorized multi‑category model.

Frequently Asked Questions

Why not just use a chi‑squared test on weekly counts?

Chi‑squared requires enough expected counts per cell; with 1‑3 posts per week the test becomes unreliable and yields many false positives.

Can I monitor the model in real time as new posts arrive?

Yes. Keep a rolling window of the most recent N weeks, update the aggregated counts, and re‑run the sampler (or use a streaming variational approximation) on a schedule.

What if a keyword never appears in a given week?

The model handles zero successes naturally; the posterior will be pulled toward the shared Beta mean, reflecting uncertainty rather than a hard zero rate.

Do I need a GPU for PyMC sampling?

No. The Beta‑Binomial model is lightweight and runs comfortably on a single CPU core. GPU acceleration only helps for high‑dimensional models.

How sensitive is the result to the prior choice?

Exponential priors on alpha and beta are weakly informative. If you have domain knowledge (e.g., typical keyword prevalence), you can replace them with more informative Gamma or Half‑Normal priors.

What I'd Change

If I were to re‑engineer this for a production alerting service, I’d replace the per‑keyword NUTS runs with a single multi‑category model and use ADVI for sub‑second updates. I’d also cache the RSS fetch and add exponential back‑off to survive transient network hiccups. Finally, I’d ship the posterior summaries to a time‑series store (e.g., InfluxDB) and let Grafana visualize the shift probability over time, turning a Python script into a fully observable monitoring component.

Post a Comment

Hi! How can we help you? Send us a message and we'll get back to you.