Skip to content

Guardian Gates: Multi-Layered Data Quality with Great Expectations and Soda for Dynamic APIs

Guardian Gates: Multi-Layered Data Quality with Great Expectations and Soda for Dynamic APIs
Do you ever lie awake at night, wondering if the external API you rely on for critical financial data suddenly decided to change a field name, or worse, started sending `null` where you expected a number? I certainly have. While we've discussed architecting resilient data pipelines that gracefully handle API interactions and complex ETL, like the ones in our recent deep dive into Pandas vs. Polars for dynamic API feeds, there's a more insidious threat: silent data quality issues. This isn't about network retries or rate limits; it's about schema drift, missing fields, or outright nonsensical values silently corrupting your downstream analytics, machine learning models, and crucial financial reports. You need a proactive defense. This post will guide you through building a dual-layered data quality system using Great Expectations for rigid schema validation on raw API data and Soda Core for flexible business logic checks on your transformed datasets, ensuring your data assets remain trustworthy even when external APIs evolve without warning.

Key Takeaways

  • Raw API data demands proactive schema validation with tools like Great Expectations to prevent schema drift from propagating downstream.
  • Business logic and semantic checks on transformed data require a complementary tool like Soda Core, which offers flexible, human-readable data quality agreements.
  • A multi-layered approach, validating raw data structure and transformed data content, provides comprehensive protection against dynamic API changes.
  • Orchestrating both Great Expectations checkpoints and Soda scans within your data pipeline creates a unified data observability workflow.
  • Explicit error handling for API interactions and robust data quality checks are non-negotiable for production-grade data pipelines.

The Problem

My team recently faced a situation where a third-party API, providing crucial financial instrument data, unexpectedly changed the casing of a field from `dailyVolume` to `DailyVolume`. Our ETL process, built with Polars for performance, gracefully ingested the data, but our downstream analytics, which expected `dailyVolume`, started reporting zero values. This silent failure went unnoticed for a short period, leading to incorrect reporting. This experience highlighted a critical gap: even with robust ETL, the *quality* and *consistency* of the incoming data itself needed rigorous, automated checks. We needed a system that could not only alert us to structural changes (like schema drift) but also validate the semantic integrity of the data *after* our transformations.

Data and Sources

For this tutorial, we'll use the public GitHub API for repository information, specifically the CPython repository. This API provides dynamic JSON data that can represent a real-world scenario where data structures might evolve.

Data accessed on 2026-09-29.

Step 1 — Ingesting Dynamic API Data: The Unseen Vulnerability

The first challenge is simply getting the data. While fetching data seems straightforward, dynamic APIs are a common point of failure. A network issue, a malformed response, or an unexpected status code can halt your pipeline before data quality checks even begin. We need robust error handling right at the ingestion point.

Here's how I'd typically fetch the data, ensuring basic resilience:

import requests
import json
import os
import pandas as pd # For Great Expectations and Soda, we'll use pandas DataFrames
from great_expectations.data_context import DataContext
from great_expectations.checkpoint import Checkpoint
from soda.sampler.sample_schema import SampleSchema
from soda.scan import Scan

def fetch_github_repo_data(url: str) -> dict:
    """Fetches data from the GitHub API with basic error handling."""
    try:
        response = requests.get(url, timeout=10)
        response.raise_for_status() # Raises an HTTPError for bad responses (4xx or 5xx)
        return response.json()
    except requests.exceptions.HTTPError as e:
        print(f"HTTP error occurred: {e}")
        raise
    except requests.exceptions.ConnectionError as e:
        print(f"Connection error occurred: {e}")
        raise
    except requests.exceptions.Timeout as e:
        print(f"Timeout error occurred: {e}")
        raise
    except json.JSONDecodeError as e:
        print(f"Failed to decode JSON from response: {e}")
        raise
    except Exception as e:
        print(f"An unexpected error occurred during API fetch: {e}")
        raise

# Example usage (not part of complete script, just for illustration)
# try:
#     repo_data = fetch_github_repo_data("https://api.github.com/repos/python/cpython")
#     print(f"Fetched {len(repo_data)} fields.")
# except Exception as e:
#     print(f"Failed to fetch data: {e}")
This snippet demonstrates how I wrap the `requests` call with multiple `try/except` blocks. This ensures that transient network issues, timeouts, or unexpected API responses (like a 500 server error or non-JSON content) are caught and explicitly handled, preventing an ungraceful pipeline crash. The `raise_for_status()` call is particularly useful for catching HTTP errors.

Step 2 — Proactive Schema Guardrails with Great Expectations

Once we have the raw data, the immediate concern is its structure. Does it conform to what we expect? Are all the critical fields present? This is where Great Expectations shines. It allows us to define "Expectations" – assertions about our data – and then run these against our raw API output. My strategy is to use Great Expectations for strict schema validation on the *raw* data immediately after ingestion. This acts as the first "guardian gate."

I'll configure Great Expectations to validate the structure of the incoming JSON. For simplicity in a runnable script, I'll save the API response to a temporary JSON file, which Great Expectations can then read. In a production setting, this would typically read directly from a landing zone or an in-memory DataFrame.

# Assume repo_data is fetched from Step 1
# Save raw data to a temporary file for GE to read
raw_data_path = "raw_github_repo_data.json"
with open(raw_data_path, "w") as f:
    json.dump(repo_data, f, indent=2)

# Great Expectations setup (simplified for script)
# Initialize a DataContext (if not already present)
# In production, this would be pre-configured
if not os.path.exists("great_expectations"):
    os.system("great_expectations init --uncommitted") # Only if not initialized

# Add a simple expectation suite
# This is usually done via `great_expectations suite new`
# For this example, we'll generate a minimal suite programmatically
# or assume a pre-existing one. Let's create a temporary one.

# This is a simplified programmatic way to define expectations.
# In a real scenario, you'd use `great_expectations suite new` and `great_expectations suite edit`
# to create and refine your expectation suites.
def define_ge_expectations(context: DataContext, data_asset_name: str):
    suite_name = f"{data_asset_name}_raw_schema_suite"
    if suite_name not in context.list_expectation_suite_names():
        suite = context.create_expectation_suite(suite_name)
        # Add expectations for critical fields

إرسال تعليق

Hi! How can we help you? Send us a message and we'll get back to you.