The Freshness Alert Fired 40 Times a Day Until the Channel Muted It: Data Alert Design With Burn-Rate Budgets Instead of Static Thresholds

Somewhere in your Slack workspace there is a channel called #data-alerts. It fired forty times yesterday. Nobody read it. The one message that mattered — a source table that stopped receiving CDC events at 02:14 and stayed stale for six hours — scrolled past under thirty-nine “source X is 61 minutes stale” notifications that resolved themselves within two minutes.

That channel is not a monitoring failure. It is a design failure. The alert was written to answer the wrong question.

A static freshness threshold answers: is this one file late right now? The operational question is: are we on pace to miss the SLA? Those are different questions, and only one of them deserves a page.

What a static threshold actually measures

Take a dbt source with warn_after: {count: 1, period: hour} and error_after: {count: 2, period: hour}. The check compares the max load timestamp in the source against wall-clock time. If the upstream extractor runs at 02:00 and finishes at 02:03, the check at 02:05 passes. If it finishes at 03:01, the check fails. The alert fires on a single late file.

The Google SRE Workbook makes the cost of this pattern explicit. In its first alerting iteration — alert when the recent error rate equals the SLO over a short window — the authors note you could receive up to 144 alerts per day every day, not act upon any alerts, and still meet the SLO (SRE Workbook, Chapter 5). That is the arithmetic of a short window: excellent detection time, terrible precision. Every late file crosses the line. Almost none of them threaten the SLA.

The obvious fix — widen the window — has its own failure mode. The Workbook’s second iteration uses a 36-hour window to ensure only sustained problems alert. Precision improves. Reset time collapses: in the case of 100% outage, an alert will fire shortly after 2 minutes, and continue to fire for the next 36 hours. You have traded a noisy channel for a channel that lies to you for a day and a half after the incident is over.

The third iteration — add a for: 1h duration — is worse. The Workbook is blunt: because the duration does not scale with the severity of the incident, a 100% outage alerts after one hour, the same detection time as a 0.2% outage. A service that spikes to 100% errors for five minutes every ten minutes never triggers the alert at all, despite consuming 35% of the monthly budget. Duration parameters do not measure severity. They measure persistence.

Reframe freshness as an SLO with a budget

If your freshness SLA is “data is at most one hour old,” you have already defined an SLO. The good event is a load that arrives within the SLA. The bad event is a load that arrives late. The error budget is the fraction of loads you are allowed to deliver late over the measurement window — typically 30 days.

Once you frame it that way, a single late file is not an incident. It is budget spend. The pager should fire when the rate of budget spend threatens to exhaust the budget before the window closes. That rate is the burn rate.

Burn rate 1 means you are consuming budget at exactly the pace that exhausts it at the end of the window. Burn rate 2 exhausts it in half the time. The Workbook’s table for a 99.9% SLO over 30 days: burn rate 1 corresponds to a 0.1% error rate and 30 days to exhaustion; burn rate 10 corresponds to 1% and 3 days; burn rate 1,000 corresponds to 100% and 43 minutes.

For freshness, the mapping is direct. If your SLA is one hour and your measurement window is 30 days, then a source that is late 0.1% of the time is burning at rate 1. A source that is late 1% of the time is burning at rate 10 and will exhaust the budget in three days. A source that stops entirely is burning at rate 1,000 and will exhaust it in 43 minutes.

The rule shape that replaces the threshold

The Workbook recommends multiwindow, multi-burn-rate alerts as the most viable option. The recommended starting numbers: 2% budget consumption in one hour and 5% budget consumption in six hours as reasonable starting numbers for paging, and 10% budget consumption in three days as a good baseline for ticket alerts.

Translated into burn rates for a 30-day window:

  • Page: 14.4x burn rate over 1 hour, confirmed by 14.4x over 5 minutes. This is 2% of the budget in one hour.
  • Page: 6x burn rate over 6 hours, confirmed by 6x over 30 minutes. This is 5% of the budget in six hours.
  • Ticket: 1x burn rate over 3 days. This is 10% of the budget in three days.

The short confirmation window is the key mechanism. The Workbook’s guideline: make the short window 1/12 the duration of the long window. The long window establishes that a significant amount of budget has been spent. The short window establishes that the budget is still being spent. If the short window has recovered, the alert does not fire — which is exactly the class of alert that trains people to mute the channel.

In Prometheus, this maps onto the alerting rule primitives directly. The for clause waits a duration before firing; the keep_firing_for clause keeps the alert firing after the condition was last met, which the documentation describes as useful to prevent situations such as flapping alerts, false resolutions due to lack of data loss, etc. (Prometheus alerting rules). The same page is explicit that Prometheus alerting rules are not a notification solution: another layer is needed to add summarization, notification rate limiting, silencing and alert dependencies. That layer is Alertmanager. If you skip it, you will rebuild rate limiting and silencing by hand, badly.

Wiring it to the actual stack

dbt source freshness and Prometheus burn-rate rules are complementary, not competing. dbt produces the SLI — how stale is the source. Prometheus produces the alerting logic — how fast is the budget burning.

Two dbt behaviors matter for this design. First, dbt build does not include source freshness checks; you either select the “Run source freshness” checkbox in the job, which runs it as the first step and won’t break subsequent steps if it fails, or you add dbt source freshness as a run step, in which case if your source data is out of date — this step will “fail”, and subsequent steps will not run (dbt source freshness docs). The checkbox is the right choice if you want freshness to feed an alerting pipeline without blocking the models. The run step is the right choice if you want stale data to halt the DAG. Pick one deliberately; the default behavior of the run step has surprised more than one on-call engineer at 03:00.

Second, check frequency is not optional. dbt’s own guidance: you should run your source freshness jobs with at least double the frequency of your lowest SLA. If your SLA is one hour, run the check every 30 minutes. A daily freshness check against a one-hour SLA measures nothing useful — it tells you whether the source was fresh at the moment you happened to look.

The recording rule that turns freshness into a ratio is straightforward. Emit a gauge per source: 1 if fresh, 0 if stale. Then:

record: source:freshness_ratio_rate1h
expr: sum(rate(source_fresh[1h])) by (source)
      / sum(rate(source_expected[1h])) by (source)

The alert then compares that ratio against the burn-rate threshold. For a 99.9% freshness SLO, the 14.4x page rule is source:freshness_ratio_rate1h < (1 - 14.4 * 0.001) combined with the 5-minute confirmation. The exact numbers depend on your SLA; the shape does not.

Where the freshness SLI lies to you

A freshness check that only reads the max timestamp in the target table can miss a permanently skipped file. Snowpipe is the clearest example. Its file-loading metadata is maintained for 14 days. Files that failed to load — because of invalid content or stage access failures — are still registered in the pipe’s metadata, and the registered file names are ignored by subsequent pipe activity, including ALTER PIPE … REFRESH (Snowpipe troubleshooting). The target table’s max timestamp can look perfectly fresh while a specific partition is permanently missing rows.

The same page documents the inverse trap: files modified and staged again after 14 days are loaded again, potentially duplicating records. A freshness SLI that only measures recency will not catch either case. Pair it with a load-history check — COPY_HISTORY for status, SYSTEM$PIPE_STATUS for lastReceivedMessageTimestamp versus lastForwardedMessageTimestamp. The gap between those two timestamps distinguishes a service configuration problem from a path mismatch between the stage and pipe definitions. A freshness alert alone cannot make that distinction.

What the fix costs

Define the SLO per source: one hour of work per source if the SLA is already documented, half a day if it is not. Write the recording rule and the three alert rules: two to four hours for the first source, thirty minutes for each subsequent one once the pattern is templated. Wire Alertmanager routing and suppression: half a day, once.

The ongoing maintenance tax is real and should be priced honestly. You now have more windows, more thresholds, and more numbers to reason about. The Workbook names this directly as a disadvantage of multi-burn-rate alerting. The three-day ticket window also produces a longer reset time than a short-window alert would. And you need alert suppression, because a 10% budget spend in five minutes also means 5% was spent in six hours and 2% in one hour — three conditions true, three notifications, unless the monitoring system prevents it.

The trade is fewer pages and a ticket queue that catches slow burns. For an on-call engineer who was about to mute the channel, that is usually the right trade. It is not free, and it is not a platform migration — it is a few hours per source plus a routing layer you probably already have.

The operational rule

The pager should fire on budget spend, not on a single late file. The ticket queue catches the slow burns that would otherwise exhaust the budget unnoticed. The muted channel is the symptom; the static threshold is the cause.

If you want to test this on one source before committing: pick the noisiest freshness alert you have, count how many times it fired last week, and count how many of those firings corresponded to a real SLA miss. If the ratio is worse than 10:1, the threshold is measuring the wrong thing. Replace it with a burn-rate rule and watch the channel go quiet — not because you muted it, but because it stopped lying to you.

FAQ

Do I need Prometheus to do this? No. The Workbook’s examples use Prometheus syntax, but the page states the approach applies in any alerting framework. What you need is a way to compute a ratio over a window and compare it to a threshold, plus a notification layer that can suppress and route. Most observability platforms have both.

What if my SLA is not a clean number like 99.9%? The burn rate is derived from the SLA, not the other way around. If your SLA is 99% over 30 days, the error budget is 1%, and burn rate 1 corresponds to a 1% error rate. The recommended budget-consumption percentages (2% in 1h, 5% in 6h, 10% in 3d) stay the same; the burn-rate multipliers change.

Should I keep the old static threshold as a backstop? Only if you route it to a ticket queue, not a pager. A static threshold that pages is the problem you are trying to solve. A static threshold that opens a ticket is a cheap safety net for the case where your SLI pipeline itself breaks.

How do I handle sources with no fixed SLA? You cannot burn-rate alert on an undefined budget. Either define the SLA — even a loose one — or route the source to a dashboard and a weekly review. Paging on an undefined SLA is how you get forty alerts a day.

Row Counts Matched but the Money Didn’t: A Source-to-Warehouse Reconciliation Query That Catches Silent Decimal Coercion

Your reconciliation job reports green. Row counts match on every table. Finance closes the month, and a controller finds a variance that traces back to a column you have been loading for two years.

The failure class is silent decimal coercion. A source column declared numeric with no scale feeds a warehouse column declared NUMBER(18,2). Every value with more than two fractional digits is rounded on write. The row count is unchanged. The money is not.

This is not a pipeline bug in the usual sense. Nothing errored. Nothing was dropped. The type contract between source and warehouse was never tested, because count(*) cannot see it.

What precision and scale actually mean

In Snowflake, precision is the total number of digits allowed and scale is the number of digits allowed to the right of the decimal point. NUMBER defaults to precision 38 and scale 0, i.e. NUMBER(38,0) (Snowflake numeric data types).

In PostgreSQL, the same terms apply. The numeric type is recommended for storing monetary amounts and other quantities where exactness is required (PostgreSQL numeric types).

The asymmetry that matters: an unconstrained numeric column does not coerce input values to any particular scale, whereas a numeric column with a declared scale does coerce input values to that scale. If the scale of a value to be stored is greater than the declared scale of the column, the system rounds the value to the specified number of fractional digits. If the number of digits to the left of the decimal point exceeds the declared precision minus the declared scale, an error is raised.

So a source column with no declared scale feeding a warehouse column with a declared scale is a rounding operation, not a copy. The rounding is deterministic and silent. It does not raise. It does not warn. It does not change the row count.

Why row-count reconciliation is structurally blind

A row-count check verifies cardinality. It answers one question: did the same number of rows arrive? It says nothing about value fidelity.

Consider a payments table. The source column is numeric with no scale. The warehouse column is NUMBER(18,2). A refund of 0.005 rounds to 0.01 in the warehouse and stays 0.005 in the source. Row counts match. The sum differs by a fraction of a cent per row. This is illustrative, not a reported incident.

The same blindness applies to floating-point paths. Snowflake’s FLOAT type uses double-precision (64 bit) IEEE 754 floating-point numbers with precision of approximately 15 digits. Snowflake recommends comparing two floating-point numbers for approximate equality rather than exact equality. PostgreSQL’s real and double precision types are inexact, variable-precision numeric types, and comparing two floating-point values for equality might not always work as expected.

If your reconciliation compares count(*) and nothing else, you are testing the one property that decimal coercion does not change.

The reconciliation query pattern

Compare three things per money column per key range and time window: row count, SUM of the column, and a scale probe. The scale probe is the part most teams skip.

SUM catches aggregate drift. It does not catch offsetting errors. One row rounded up and another rounded down can net to the same total. The scale probe catches the case where rounding is uniform enough that SUM happens to match.

A scale probe asks: what is the maximum number of digits actually present to the right of the decimal point in this column, over this key range? If the source returns 4 and the warehouse returns 2, the warehouse has coerced. The exact expression depends on your engine; the point is to measure the stored scale, not the declared scale.

For per-row fidelity, add a hash or checksum of the money column to the same test. SUM plus scale probe plus per-row hash is one test, not three. Splitting them across separate jobs means a failure in one does not block the others, and the signal gets lost in the noise.

In dbt, this is a singular data test. Data tests are assertions about models and other resources, and a test passes when it returns zero failing rows. Data tests return one row for each failure, and the columns in the test’s SQL select statement are the columns visible when you debug failures, including when you store test failures (dbt data tests).

Write the test so it returns the key range, the source sum, the warehouse sum, the source scale, and the warehouse scale. When it fails, you want the numbers in the row, not a boolean.

The type contract is the artifact under test

The reconciliation query is testing a contract: the source column’s precision and scale must be compatible with the warehouse column’s precision and scale. If the source is unconstrained and the warehouse is declared, the contract is lossy by construction.

Pair the reconciliation query with a schema-diff check that fails when the two sides’ precision and scale declarations diverge. The schema diff catches the drift at deploy time. The reconciliation query catches it at data time. You need both because a schema diff cannot see values that were already rounded before the schema changed.

This is where the decision gets expensive. Snowflake’s DECFLOAT type stores numbers exactly, with up to 38 significant digits of precision, and uses a dynamic base-10 exponent to represent very large or small values. Snowflake lists ledgers, taxes, or compliance as use cases requiring exact numeric values, and notes that use of the DECFLOAT type might cause storage consumption to increase. The NUMBER and FLOAT types might provide better performance than the DECFLOAT type.

So the choice is not free. DECFLOAT buys exactness and costs storage and possibly performance. NUMBER with an explicit scale buys predictability and costs the fractional digits you did not declare. FLOAT buys range and costs exactness. There is no option that is free on all three axes.

CDC and backfill make it worse in a specific way

Debezium’s PostgreSQL connector relies on logical decoding, which does not support DDL changes. The connector is unable to report DDL change events back to consumers (Debezium PostgreSQL connector).

This means a column type change on the source is invisible to the connector. The streaming path continues. If the connector stops for any reason, upon restart it continues reading the WAL where it last left off. If it stops during a snapshot, it begins a new snapshot when it restarts.

The failure mode: a DDL change lands between the streaming path and the backfill path. The streaming path read the column under the old type. The backfill reads it under the new type. Both paths produce rows. Both paths pass row-count reconciliation. The values differ.

This is why reconciliation must run after backfills, not only after initial load. A backfill that re-reads a column after a type change can produce a different value than the streaming path for the same logical row. The reconciliation query is the only signal, because the connector will not tell you the type changed.

Iceberg’s schema evolution goals state that schema evolution supports safe column add, drop, reorder and rename, including in nested structures (Iceberg table spec). Safe in this context means the metadata is consistent. It does not mean your source and warehouse type declarations agree. The reconciliation query is still yours to write.

Pricing the fix

The query is cheap. Adding a singular dbt test or an Airflow task that compares SUM, scale probe, and per-row hash across source and warehouse is hours of work, not weeks. It runs on the same schedule as your existing reconciliation.

The expensive part is the type audit. You need to enumerate every money column on both sides, compare precision and scale declarations, and decide for each one whether to pin an explicit scale, migrate to DECFLOAT, or accept the rounding and document it. That audit is days to weeks depending on column count and how many teams own the schemas.

The migration decision is the expensive part. The query is the cheap part. Do the query first, because it tells you which columns actually need the migration decision.

Airflow’s TaskFlow API uses XComs to move inputs and outputs between tasks and requires that variables used as arguments be serializable (Airflow TaskFlow). If you pass reconciliation results between tasks, keep them to scalars and small dicts. Do not pass result sets through XCom.

The maintenance tax

Every new money column added to the warehouse is a new place the source/warehouse type contract can drift. The reconciliation query is a standing test, not a migration checklist item, because the drift is introduced by future schema changes, not by the original load.

The tax is not the query. The tax is the review step that asks, for every new money column, what the source scale is and what the warehouse scale is, and whether the difference is intentional. That review is minutes per column if it is part of the schema change process, and hours per column if it is discovered after the fact.

Row counts will keep matching. The money will keep not matching. The only thing that changes is whether you find out before or after the controller does.

FAQ

Does unique or not_null catch this? No. dbt ships with four generic data tests: unique, not_null, accepted_values, and relationships. None of them compares values across two systems. You need a singular test or a custom generic test that queries both sides.

Can I just compare SUM? No. SUM can hide offsetting errors. One row rounded up and another rounded down can net to the same total. Compare SUM, a scale probe, and a per-row hash in the same test.

Should I migrate everything to DECFLOAT? Not without pricing it. Snowflake notes that DECFLOAT may increase storage consumption and that NUMBER and FLOAT might provide better performance. Run the reconciliation query first to find which columns actually need exactness, then decide per column.

Why not just declare a scale on the source column? That is one valid fix. It makes the source coerce to the same scale as the warehouse, so the rounding happens on both sides. The tradeoff is that you are now rounding at the source, which may not be acceptable for the business logic that reads the source directly.

How often should the reconciliation run? After every backfill, after every schema change, and on the same schedule as your existing reconciliation. The backfill case is the one teams miss, because the backfill passes row-count reconciliation and looks healthy.

Nobody Owns the Pipeline at 3 a.m.: What a Working Data On-Call Rotation Requires Beyond a PagerDuty Schedule

A PagerDuty schedule is an assignment mechanism. It says who gets the page. It does not say who owns the pipeline, what the page means, or what the on-call engineer is supposed to do when the page fires at 3 a.m. and the only person awake is the one holding the phone.

Data teams running Kafka, Airflow, dbt, Postgres/CDC, and Snowflake or Iceberg stacks tend to discover this gap the hard way. The schedule exists. The rotation exists. The ownership does not. When a consumer group rebalances, a source freshness check fails, or a Postgres statistics counter resets after a crash, the on-call engineer is left to reconstruct intent from dashboards and Slack threads.

This article is about what the rotation needs beyond the schedule. It is not a best-practices list. It is a set of load limits, escalation paths, and recovery procedures, each priced in incidents per shift, hours of follow-up, and the specific configuration keys that generate recurring operational work.

Start with a load budget, not a coverage chart

Google’s SRE book caps the amount of time SREs spend on purely operational work at 50%, with at least 50% allocated to engineering projects. The same chapter states that dealing with an on-call incident — root-cause analysis, remediation, and follow-up like writing a postmortem and fixing bugs — takes 6 hours on average, and that the maximum number of incidents per day is 2 per 12-hour on-call shift.

Those numbers are not universal. They are a load budget. The useful move for a data team is to derive its own budget from the same arithmetic: how many hours does a real incident consume, from the first page to the closed postmortem? If a Kafka consumer lag alert takes 90 minutes to triage and a backfill takes four hours to verify, then a shift that absorbs three of those has already spent the engineer’s follow-up capacity. The fourth page lands on someone who is already behind.

The SRE workbook restates the target: a maximum of two incidents per on-call shift, to ensure adequate time for follow-up. The workbook also notes that on-call engineers should be fully supported by procedures and escalation paths because being on-call can be daunting and highly stressful. That is not a wellness slogan. It is a statement about cognitive load. An engineer who is stressed and underslept makes worse decisions during an incident, and worse decisions during a data incident tend to mean a longer backfill or a wider blast radius.

For a data rotation, the load budget has to be measured in the units the team actually experiences: pages per shift, incidents per week, and hours of follow-up per incident. If the team cannot state those numbers, it does not have a rotation. It has a schedule.

Not every alert is a page

The single most common failure in a data on-call rotation is treating every alert as a page. A dbt source freshness failure, a Kafka consumer lag spike, and a Postgres statistics reset are not the same kind of event. They have different urgency, different recovery paths, and different follow-up work.

The dbt documentation makes the distinction concrete. dbt build does not include source freshness checks when building and testing resources in the DAG. If you select the Run source freshness checkbox in a job’s execution settings, dbt runs dbt source freshness as the first step and does not break subsequent steps if it fails. If you instead add dbt source freshness as a run step, and the source data is out of date, that step fails and subsequent steps do not run.

That is a configuration decision with an on-call consequence. The checkbox produces a non-breaking signal: the job continues, the freshness state is visible, and the on-call engineer can decide whether to act. The run step produces a hard stop: the pipeline halts, downstream models do not run, and the page fires. Neither is wrong. But the team has to decide which one it wants before the alert fires, not after.

The dbt documentation also recommends running source freshness jobs with at least double the frequency of the lowest SLA. A one-hour SLA implies a check every 30 minutes. A daily SLA implies a check every 12 hours. If the check frequency is wrong, the alert either fires too late to be useful or fires so often that the on-call engineer learns to ignore it.

There is a further limitation worth knowing. dbt source freshness for Snowflake is calculated using the LAST_ALTERED column. That column reflects metadata changes, not necessarily data changes. A table can be altered without new rows arriving, and a table can receive new rows without the metadata changing in the way the check expects. The check is a signal, not a proof.

Kafka consumer lag has its own timing semantics. The Confluent consumer configuration reference documents heartbeat.interval.ms as the expected time between heartbeats to the consumer coordinator when using group management, with a default of 3000 ms and a note that it should typically be no higher than one-third of session.timeout.ms. It documents session.timeout.ms as the timeout used to detect client failures, with a default of 45000 ms. It documents max.poll.interval.ms as the maximum delay between invocations of poll(), with a default of 300000 ms, after which the consumer is considered failed and the group rebalances.

Those three values interact. A consumer that processes a batch slowly can exceed max.poll.interval.ms and be evicted from the group even though it is healthy. A consumer that is paused for a deploy can exceed session.timeout.ms and trigger a rebalance. The alert that fires is “consumer lag,” but the cause may be a configuration value, not a data volume problem. The on-call engineer needs to know which one before touching anything.

Escalation paths and named owners

Google’s SRE book lists clear escalation paths, well-defined incident-management procedures, and a blameless postmortem culture as the most important on-call resources. It also describes primary and secondary on-call rotations, with duties varying by team: the secondary may be a fall-through for pages the primary misses, or may handle non-urgent production activities while the primary handles pages.

Data teams need the same structure, but the ownership map is harder. A pipeline may be owned by a data engineer, an analytics engineer, or a platform engineer. The Kafka cluster may be owned by a platform team. The warehouse may be owned by a separate group. When a page fires, the on-call engineer needs to know which of those owners to escalate to, and what response time to expect.

That information has to be written down before the rotation starts. A useful artifact is a one-page ownership table per pipeline: pipeline name, primary owner, secondary owner, escalation contact, expected response time, and the specific failure modes that justify a page. The table is not documentation for its own sake. It is the difference between a 10-minute escalation and a 40-minute Slack search at 3 a.m.

The SRE workbook describes playbooks as high-level instructions on how to respond to automated alerts, explaining severity and impact, and including debugging suggestions and possible actions. It also recommends implementing automation if playbooks are a deterministic list of commands the on-call engineer runs every time a particular alert fires. That recommendation is directly applicable to data pipelines. If the playbook for a freshness failure is “run this query, check this table, restart this task,” the playbook should be a script, not a document.

The maintenance tax in configuration

The recurring operational work in a data stack does not come from the big architectural decisions. It comes from configuration drift. Three examples, each anchored to a documented behavior.

Kafka consumer timeouts. The defaults for heartbeat.interval.ms, session.timeout.ms, and max.poll.interval.ms are tuned for general-purpose consumers. A consumer that does heavy per-record processing, or that pauses during a deploy, may need different values. Every change to those values is a change to the failure mode. The on-call engineer needs to know which consumers have non-default values and why.

Postgres statistics collection. PostgreSQL’s cumulative statistics system supports collection and reporting of information about server activity, including accesses to tables and indexes in disk-block and individual-row terms. Collection is controlled by parameters such as track_activities, track_counts, track_functions, and track_io_timing. The statistics views do not update instantaneously: each server process flushes accumulated statistics to shared memory just before going idle, but not more frequently than once per PGSTAT_MIN_INTERVAL milliseconds, so the displayed information lags behind actual activity. And when a server starts from an unclean shutdown — after an immediate shutdown, a server crash, a base backup, or point-in-time recovery — all statistics counters are reset.

That last point matters for on-call. A dashboard that depends on cumulative counters will show a discontinuity after a crash. An alert threshold based on a counter that just reset will either fire spuriously or fail to fire. The on-call engineer needs to know which dashboards are counter-based and which are gauge-based, and what happens to each after a restart.

dbt freshness check placement. As noted above, the checkbox and the run step produce different failure behavior. The choice is a maintenance decision. If the check is a run step, every freshness failure is a pipeline halt and a page. If the check is a checkbox, every freshness failure is a signal that someone has to review. The first option creates more pages. The second creates more silent failures. The team has to pick which tax it wants to pay.

A rotation checklist that fits on one page

Before anyone goes on-call for a data pipeline, the following should exist and be current. This is not a maturity model. It is the minimum set of artifacts that makes the rotation functional.

  • Ownership table. Pipeline name, primary owner, secondary owner, escalation contact, expected response time.
  • Alert classification. Which alerts page, which alerts create a ticket, and which alerts are informational. For each paging alert, the specific failure mode it represents.
  • Load budget. The team’s target for incidents per shift and hours of follow-up per incident, derived from its own incident history.
  • Playbook per paging alert. Severity, impact, debugging steps, and the actions that mitigate or resolve the alert. If the steps are deterministic, they should be a script.
  • Recovery procedures. How to restart a consumer group, how to re-run a failed dbt model, how to verify a backfill, and how to confirm that a schema change did not break a downstream consumer.
  • Handoff template. What the outgoing on-call engineer writes down: open incidents, in-progress backfills, known configuration changes, and anything that is likely to page in the next shift.

The SRE workbook describes a training approach that is worth adapting: a checklist of focus areas, lab sessions for common debugging and mitigation tasks, and “Wheel of Misfortune” exercises where the team role-plays recent incidents. For a data team, the equivalent is a game day that exercises the actual failure modes: a backfill that runs long, a schema change that breaks a consumer, a freshness check that fails silently, and an escalation that reaches the wrong person.

The game day is not a drill for its own sake. It is how the team discovers which parts of the rotation are missing before the pager discovers it for them.

What the schedule cannot do

A PagerDuty schedule answers one question: who is holding the phone. It does not answer who owns the pipeline, what the page means, how long the recovery should take, or who to call when the first three steps do not work.

Those answers come from a load budget measured in incidents per shift, an alert classification that separates pages from signals, an ownership table with named escalation contacts, and recovery procedures that have been tested. The schedule is necessary. It is not sufficient.

The test is simple. Ask the on-call engineer to describe, without looking anything up, what happens when a Kafka consumer group rebalances at 3 a.m. If the answer is a specific sequence of checks, a named escalation contact, and a known recovery procedure, the rotation is working. If the answer is “I would look at the dashboard and figure it out,” the schedule is doing all the work, and the pipeline has no owner.

FAQ

How many incidents per shift should a data on-call rotation target?

Google’s SRE workbook targets a maximum of two incidents per on-call shift to allow adequate time for follow-up. That number is a starting point, not a universal rule. A data team should derive its own target from the average time a real incident consumes, including triage, remediation, and postmortem. If a typical incident takes four hours, two incidents per shift already exceeds a normal working day.

Should dbt source freshness checks break the pipeline or just warn?

It depends on the SLA and the downstream dependency. The dbt documentation describes two behaviors: the Run source freshness checkbox runs the check as a non-breaking first step, while adding dbt source freshness as a run step causes the step to fail and subsequent steps not to run. The first produces a signal; the second produces a halt. The team should choose based on whether downstream models can tolerate stale source data.

Why does a Kafka consumer get evicted from its group even when it is healthy?

The Confluent consumer configuration reference documents max.poll.interval.ms as the maximum delay between invocations of poll(), with a default of 300000 ms. If the consumer’s processing loop takes longer than that between polls, the consumer is considered failed and the group rebalances. A consumer that processes large batches or pauses during a deploy can hit this limit without any underlying data problem.

What happens to Postgres monitoring after a crash?

The PostgreSQL documentation states that when a server starts from an unclean shutdown — after an immediate shutdown, a server crash, a base backup, or point-in-time recovery — all statistics counters are reset. Dashboards and alerts that depend on cumulative counters will show a discontinuity. The on-call engineer should know which monitoring depends on those counters and how the alert thresholds behave after a reset.

Do we need a secondary on-call rotation for data pipelines?

Google’s SRE book describes primary and secondary rotations with duties that vary by team. For a data team, the secondary serves two purposes: fall-through for pages the primary misses, and a second person who knows the recovery procedures. The second purpose is the more important one. If only one person knows how to recover a pipeline, the rotation has a single point of failure that the schedule does not address.

The Producer Team Renamed a Field and Called It Non-Breaking: What a Cross-Team Schema Review Actually Has to Check

A producer team renames user_id to account_id in an Avro record. They run the schema through Schema Registry. It is accepted. They post in the shared channel: “Non-breaking change, no consumer action needed.”

Two days later, a consumer job fails to deserialize. A dbt incremental model silently stops populating a column. A CDC pipeline replays a batch it already processed.

The rename was not non-breaking. It was non-breaking relative to the compatibility mode that was actually in effect, the schema format in use, and the set of consumers that actually read the topic. The producer team checked one of those three things.

This is not a story about a careless team. It is a story about a review process that treats “the registry accepted it” as equivalent to “nothing downstream will break.” Those are different claims, and the gap between them is where on-call hours go to die.

What the registry actually checks

Schema Registry enforces compatibility by comparing a new schema version against previous versions using a configurable compatibility type. The default is BACKWARD, not BACKWARD_TRANSITIVE. That distinction matters more than most teams realize.

Under BACKWARD, a consumer using the new schema can process data written by producers using schema X or X-1, but not necessarily X-2. Under BACKWARD_TRANSITIVE, that same consumer can process data written by X, X-1, or X-2. The Confluent documentation is explicit: the default is BACKWARD, and the main reason is so that you can rewind consumers to the beginning of the topic.

Here is the failure mode. A team has three schema versions in production. They add a fourth. The registry checks version 4 against version 3 under BACKWARD. It passes. But a consumer that rewinds to the beginning of the topic will encounter version 1 and version 2 messages. If the change is not transitive-safe, that consumer breaks on old data.

The registry did its job. The review did not.

What each compatibility type actually permits

The Confluent compatibility tables for Avro and Protobuf show which operations are allowed under each mode. Adding an optional field is compatible under BACKWARD, FORWARD, and FULL. Removing an optional field is also compatible under all three. Adding a required field is compatible only under FORWARD. Removing a required field is compatible only under BACKWARD.

Renames are not listed as a distinct operation. In Avro, a rename is typically expressed as a field removal plus a field addition. Whether that passes depends on whether the removed field was optional or had a default value, and whether the added field has a default value. The documentation states: “the ability to delete a field and keep the schema compatible requires that the field was either specified as optional or provided a default value in the original version.”

So a rename can pass BACKWARD if the old field had a default and the new field has a default. It can pass FORWARD under similar conditions. It can pass FULL if both conditions hold. But passing the registry check does not mean consumers will find the data they expect. A consumer looking for user_id will not find it in a message that only contains account_id. The registry does not know what field names your consumer code references.

Schema format changes the rules

Avro, Protobuf, and JSON Schema have different compatibility rules. The Confluent documentation notes that Avro was developed with schema evolution in mind and its specification clearly states the rules for backward compatibility, whereas the rules for JSON Schema and Protobuf can be more nuanced.

For JSON Schema, compatibility behavior depends on both the compatibility policy (lenient or strict) and the content model (additionalProperties: true for open, false for closed). A change that passes under a lenient policy with an open content model may fail under a strict policy with a closed content model. The review must confirm which policy and which content model are actually in effect.

For Protobuf, the documentation notes that best practice is to use BACKWARD_TRANSITIVE, because adding new message types is not forward compatible. A team using BACKWARD with Protobuf may accept a change that breaks forward compatibility in ways the registry does not flag.

The effective compatibility mode may not be what you think

A REST API call to compatibility mode is global and overrides any compatibility parameters set in schema registry properties files. This means the effective mode for a subject may differ from what the properties file says. A team that set BACKWARD_TRANSITIVE in their properties file may find that a global API call reset it to BACKWARD.

The review must verify the effective mode for the specific subject, not the mode someone believes is configured. The diagnostic is straightforward:

curl -s http://schema-registry:8081/config
curl -s http://schema-registry:8081/config/<subject-name>

The first call returns the global compatibility level. The second returns the subject-level override, if any. If the subject-level value is absent, the global value applies. If a global API call was made, it overrides the properties file.

What the registry does not check

The registry checks schema compatibility. It does not check:

  • Whether consumers have been rewound to the beginning of the topic
  • Whether dbt incremental models will pick up the change
  • Whether downstream CDC pipelines handle the renamed field
  • Whether any consumer code references the old field name
  • Whether the change is transitive-safe across all schema versions in the topic

Each of these is a separate failure mode. Each requires a separate check.

The dbt layer: silent column drops

dbt incremental models have an on_schema_change configuration. The default is ignore. Under ignore, if you add a column to your incremental model and execute a dbt run, the column will not appear in the target table. If you remove a column and execute a dbt run, dbt will fail.

This means a renamed field can produce two different failure modes depending on which side of the rename the dbt model sees. If the model references the old field name and the source no longer provides it, the run fails. If the model references the new field name and the source provides it, but the target table was built with the old schema, the new column silently does not appear.

The documentation is explicit: “None of the on_schema_change behaviors backfill values in old records for newly added columns.” If you need to populate those values, you must run manual updates or trigger a --full-refresh.

There is another constraint: on_schema_change only tracks top-level column changes. It does not track nested column changes. A rename inside a nested structure will not trigger a schema change, even if on_schema_change is set appropriately.

The diagnostic is to check which models use incremental materialization and what their on_schema_change setting is:

dbt ls -s config.materialized:incremental --output json | jq '.[].config.on_schema_change'

If the output is null or "ignore", the model will not pick up new columns automatically.

The CDC layer: replay and idempotency

PostgreSQL logical decoding slots emit each change once in normal operation. But the current position of each slot is persisted only at checkpoint. In the case of a crash, the slot might return to an earlier LSN, which will cause recent changes to be sent again when the server restarts.

The documentation states: “Logical decoding clients are responsible for avoiding ill effects from handling the same message more than once.”

This means a CDC pipeline that consumes a renamed field must be idempotent against replay. If the pipeline processes a message with user_id, then a message with account_id, then a replayed message with user_id, it must not produce duplicate or inconsistent rows.

Replication slots persist across crashes and know nothing about the state of their consumers. They will prevent removal of required resources even when there is no connection using them. A slot that is no longer required should be dropped, but dropping it requires knowing which consumers depend on it.

The diagnostic is to check which slots exist and how far behind they are:

SELECT slot_name, plugin, slot_type, active, restart_lsn,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS lag
FROM pg_replication_slots;

A slot with a large lag and no active connection is a candidate for investigation. It may be holding WAL for a consumer that no longer exists, or it may be a consumer that is about to replay a large batch.

The Iceberg layer: column IDs and rename safety

The Iceberg spec states that schema evolution supports safe column add, drop, reorder and rename, including in nested structures. This is a stronger guarantee than what Schema Registry provides for Kafka topics, because Iceberg tracks column IDs rather than column names.

But the spec also notes that the format version number is incremented when new features are added that will break forward-compatibility. This means a table written with format version 3 may not be readable by a reader that only supports version 2. The review must confirm which format version the table uses and whether all readers support it.

The diagnostic is to check the table metadata:

SELECT * FROM catalog.db.table_snapshots LIMIT 1;
-- or
SELECT * FROM "table$snapshots" LIMIT 1;

The format-version field in the table metadata indicates which version the table uses. If the table was upgraded from version 2 to version 3, readers that only support version 2 will fail.

What a cross-team schema review must actually check

The review is not a single check. It is a sequence of checks, each of which can fail independently.

  1. Confirm the effective compatibility mode for the specific subject. Use the Schema Registry API to check both the global and subject-level settings. Do not rely on the properties file.
  2. Confirm the schema format and its specific rules. Avro, Protobuf, and JSON Schema have different compatibility rules. JSON Schema compatibility depends on both the policy and the content model.
  3. Confirm whether the change is transitive-safe. If any consumer rewinds to the beginning of the topic, the change must be compatible with all schema versions, not just the last one.
  4. Confirm whether any consumer rewinds to the beginning of the topic. This is a property of the consumer configuration, not the schema. Check consumer group offsets and retention settings.
  5. Confirm whether downstream dbt models use on_schema_change and what setting. The default is ignore, which means new columns silently do not appear.
  6. Confirm whether CDC pipelines are idempotent against replay. Logical decoding slots can replay changes after a crash. The pipeline must handle duplicate messages.
  7. Confirm whether the change is compatible with the Iceberg table format version in use. A table upgraded to a newer format version may not be readable by older readers.

Pricing the maintenance tax

The cost of a schema change is not paid at the moment of the change. It is paid when a consumer fails to deserialize, when a dbt incremental model silently drops a column, or when a CDC pipeline replays a change it already processed.

Each of these failure modes has a different cost profile:

  • Consumer deserialization failure: The consumer stops processing. If it is a real-time pipeline, the lag grows. If it is a batch pipeline, the batch fails. The fix is to update the consumer code and redeploy. The cost is the time to diagnose, fix, and redeploy, plus the cost of any data that was not processed during the outage.
  • dbt silent column drop: The model runs successfully but produces incomplete data. The failure is not detected until someone notices that a column is null or missing. The fix is to run a full refresh, which may be expensive if the model processes a large volume of data. The cost is the compute cost of the full refresh plus the time to diagnose why the column disappeared.
  • CDC replay: The pipeline processes a message it already processed. If the pipeline is not idempotent, it produces duplicate rows. The fix is to deduplicate the data and make the pipeline idempotent. The cost is the time to diagnose the duplication plus the cost of the deduplication job.

The review should price these failure modes in on-call hours, not just in registry API calls. A review that takes 30 minutes and catches a non-transitive change is cheaper than a review that takes 5 minutes and misses it.

Frequently asked questions

Does Schema Registry check for field renames?

Schema Registry checks compatibility based on the rules for the schema format and compatibility type. In Avro, a rename is typically expressed as a field removal plus a field addition. Whether that passes depends on whether the removed field was optional or had a default value, and whether the added field has a default value. The registry does not know what field names your consumer code references.

What is the difference between BACKWARD and BACKWARD_TRANSITIVE?

Under BACKWARD, a consumer using the new schema can process data written by producers using schema X or X-1, but not necessarily X-2. Under BACKWARD_TRANSITIVE, that same consumer can process data written by X, X-1, or X-2. The default is BACKWARD.

Why does my dbt incremental model not pick up a new column?

The default on_schema_change setting is ignore. Under ignore, if you add a column to your incremental model and execute a dbt run, the column will not appear in the target table. You must set on_schema_change to append_new_columns or sync_all_columns, or run a full refresh.

Can a CDC pipeline process the same message twice?

Yes. PostgreSQL logical decoding slots persist their position only at checkpoint. In the case of a crash, the slot might return to an earlier LSN, which will cause recent changes to be sent again when the server restarts. Logical decoding clients are responsible for avoiding ill effects from handling the same message more than once.

Does Iceberg handle column renames safely?

The Iceberg spec states that schema evolution supports safe column add, drop, reorder and rename, including in nested structures. This is because Iceberg tracks column IDs rather than column names. However, the format version number is incremented when new features are added that will break forward-compatibility, so readers must support the table’s format version.

Sources