Have you ever found yourself grappling with a time series that clearly exhibits strong yearly cycles, yet also hides subtle, intricate daily or weekly patterns, all while being unexpectedly punctuated by external, non-periodic events? I certainly have. When I first attempted to predict the frequency of major events, like Formula 1 race meetings, relying on a single model often felt like trying to capture mist with a sieve. Facebook Prophet would effortlessly handle the annual seasonality and holiday effects, but sometimes smoothed over the finer, autoregressive dynamics. Conversely, a pure SARIMAX model could dissect intricate short-term correlations, but demanded meticulous feature engineering to account for complex, recurring calendar-based phenomena. This recurring dilemma led me to a powerful insight: why force a single model to do everything when you can combine their complementary strengths? In this post, I’ll take you through my process of architecting a resilient, hybrid forecasting pipeline, leveraging both Facebook Prophet and statsmodels' SARIMAX, dynamically enriched with external regressors derived directly from raw API data. You’ll learn how to build a system that not only predicts event frequency with enhanced accuracy but also provides the actionable interpretability crucial for production deployments.
Key Takeaways
- Complex time series often benefit from hybrid modeling, where a model strong in seasonality (like Prophet) handles the macro patterns, and another strong in autoregression (like SARIMAX) models the residuals.
- Dynamic external regressors, even simple counts or indicators derived from raw API metadata, can significantly improve forecast accuracy and model interpretability by providing crucial contextual information.
- Careful data structuring, including creating a consistent daily time series for event counts and handling potential API inconsistencies, is foundational for robust time series modeling.
- Prophet's strengths lie in its additive components for trend, seasonality, and holidays, making it excellent for baseline forecasts, while SARIMAX excels at capturing the unexplained autocorrelation in the residuals.
- A production-ready forecasting pipeline requires robust error handling, clear output visualization, and a well-defined evaluation strategy beyond just fitting the model.
The Problem
Imagine you're managing resources for a broadcast network or logistics company that supports global sporting events. Your operational efficiency hinges on accurately predicting the frequency of these events months in advance. Specifically, for Formula 1, while the overall season structure is somewhat predictable, the exact number of unique meetings per month, or the impact of new circuits and varying pre-season tests, introduces significant variability. A simple moving average is too crude. A purely seasonal model might miss the subtle year-over-year growth or decline, or the impact of a specific country hosting more races. How do we build a forecasting system that can robustly predict the number of F1 meetings in a given period, accounting for both its inherent seasonality and external factors, while being adaptable enough for production use cases?
Data and Sources
For this project, we'll be using the Open F1 API, which provides detailed data on Formula 1 meetings, sessions, drivers, and more. We'll specifically query the /v1/meetings endpoint to gather historical and current season meeting data.
- Open F1 API Documentation: https://api.openf1.org/
- Specific endpoint used: https://api.openf1.org/v1/meetings
Data accessed on 2024-07-28.
Step 1 — Ingesting and Structuring Dynamic F1 Event Data
Our first challenge is to reliably fetch the F1 meeting data from the API and transform it into a clean, consistent time series format. The API returns a list of dictionaries, each representing a race meeting. For forecasting event frequency, we need a daily count of meetings. This involves iterating through a range of years, making API calls, and then collapsing the start dates of these meetings into a daily frequency.
I start by defining a utility function to fetch data for a specific year, including basic error handling for network issues. Then, I iterate through several years to build a historical dataset. Once fetched, I parse the date_start field, convert it to a datetime object, and then use it to create a daily count of events. The key here is to ensure our time series is continuous, even on days with no events, by re-indexing to a daily frequency and filling missing values with zero.
import requests
import pandas as pd
from datetime import datetime, timedelta
def fetch_f1_meetings(year: int) -> list:
"""Fetches F1 meeting data for a given year."""
url = f"https://api.openf1.org/v1/meetings?year={year}"
try:
response = requests.get(url, timeout=10)
response.raise_for_status() # Raise an exception for HTTP errors
return response.json()
except requests.exceptions.RequestException as e:
print(f"Error fetching data for {year}: {e}")
return []
# Fetch data for multiple years (example: 2020-2024)
all_meetings = []
for year in range(2020, 2025):
all_meetings.extend(fetch_f1_meetings(year))
# Process into a DataFrame and create daily event counts
meeting_dates = []
for meeting in all_meetings:
try:
# Extract only the date part for daily counting
start_date = datetime.fromisoformat(meeting['date_start'].replace('Z', '+00:00')).date()
meeting_dates.append({'ds': start_date, 'meeting_key': meeting['meeting_key']})
except (KeyError, ValueError) as e:
print(f"Skipping malformed meeting entry: {meeting}. Error: {e}")
continue
df_meetings = pd.DataFrame(meeting_dates)
df_meetings['ds'] = pd.to_datetime(df_meetings['ds'])
# Aggregate to daily event counts
daily_counts = df_meetings.groupby('ds').size().reset_index(name='y')
# Create a complete date range to ensure continuous time series
min_date = daily_counts['ds'].min()
max_date = daily_counts['ds'].max()
full_date_range = pd.date_range(start=min_date, end=max_date, freq='D')
df_ts = pd.DataFrame({'ds': full_date_range})
df_ts = pd.merge(df_ts, daily_counts, on='ds', how='left').fillna(0)
The `df_ts` DataFrame now contains our target variable `y` (daily meeting count) and the `ds` (date) column, formatted perfectly for our forecasting models. This step is critical; a clean and continuous time series is the bedrock for any robust forecast.
Step 2 — Engineering Contextual Regressors from Metadata
While event counts provide our primary signal, the raw F1 meeting metadata contains valuable contextual information that can significantly enhance our models. Instead of just relying on the time series itself, we can engineer external regressors. For instance, the number of unique countries hosting races in a given year might indicate the scale or global reach of the season, which could influence overall event frequency. Another regressor could be a binary flag for "major" events versus testing sessions.
I decided to create a regressor that captures the diversity of host countries. This involves aggregating the unique countries per year and then mapping this back to our daily time series. This type of regressor helps Prophet understand broader yearly trends not purely driven by day-of-week or month-of-year effects.
# Extract unique countries per year as a regressor