Nobody Owns the Pipeline at 3 a.m.: What a Working Data On-Call Rotation Requires Beyond a PagerDuty Schedule

A PagerDuty schedule is an assignment mechanism. It says who gets the page. It does not say who owns the pipeline, what the page means, or what the on-call engineer is supposed to do when the page fires at 3 a.m. and the only person awake is the one holding the phone.

Data teams running Kafka, Airflow, dbt, Postgres/CDC, and Snowflake or Iceberg stacks tend to discover this gap the hard way. The schedule exists. The rotation exists. The ownership does not. When a consumer group rebalances, a source freshness check fails, or a Postgres statistics counter resets after a crash, the on-call engineer is left to reconstruct intent from dashboards and Slack threads.

This article is about what the rotation needs beyond the schedule. It is not a best-practices list. It is a set of load limits, escalation paths, and recovery procedures, each priced in incidents per shift, hours of follow-up, and the specific configuration keys that generate recurring operational work.

Start with a load budget, not a coverage chart

Google’s SRE book caps the amount of time SREs spend on purely operational work at 50%, with at least 50% allocated to engineering projects. The same chapter states that dealing with an on-call incident — root-cause analysis, remediation, and follow-up like writing a postmortem and fixing bugs — takes 6 hours on average, and that the maximum number of incidents per day is 2 per 12-hour on-call shift.

Those numbers are not universal. They are a load budget. The useful move for a data team is to derive its own budget from the same arithmetic: how many hours does a real incident consume, from the first page to the closed postmortem? If a Kafka consumer lag alert takes 90 minutes to triage and a backfill takes four hours to verify, then a shift that absorbs three of those has already spent the engineer’s follow-up capacity. The fourth page lands on someone who is already behind.

The SRE workbook restates the target: a maximum of two incidents per on-call shift, to ensure adequate time for follow-up. The workbook also notes that on-call engineers should be fully supported by procedures and escalation paths because being on-call can be daunting and highly stressful. That is not a wellness slogan. It is a statement about cognitive load. An engineer who is stressed and underslept makes worse decisions during an incident, and worse decisions during a data incident tend to mean a longer backfill or a wider blast radius.

For a data rotation, the load budget has to be measured in the units the team actually experiences: pages per shift, incidents per week, and hours of follow-up per incident. If the team cannot state those numbers, it does not have a rotation. It has a schedule.

Not every alert is a page

The single most common failure in a data on-call rotation is treating every alert as a page. A dbt source freshness failure, a Kafka consumer lag spike, and a Postgres statistics reset are not the same kind of event. They have different urgency, different recovery paths, and different follow-up work.

The dbt documentation makes the distinction concrete. dbt build does not include source freshness checks when building and testing resources in the DAG. If you select the Run source freshness checkbox in a job’s execution settings, dbt runs dbt source freshness as the first step and does not break subsequent steps if it fails. If you instead add dbt source freshness as a run step, and the source data is out of date, that step fails and subsequent steps do not run.

That is a configuration decision with an on-call consequence. The checkbox produces a non-breaking signal: the job continues, the freshness state is visible, and the on-call engineer can decide whether to act. The run step produces a hard stop: the pipeline halts, downstream models do not run, and the page fires. Neither is wrong. But the team has to decide which one it wants before the alert fires, not after.

The dbt documentation also recommends running source freshness jobs with at least double the frequency of the lowest SLA. A one-hour SLA implies a check every 30 minutes. A daily SLA implies a check every 12 hours. If the check frequency is wrong, the alert either fires too late to be useful or fires so often that the on-call engineer learns to ignore it.

There is a further limitation worth knowing. dbt source freshness for Snowflake is calculated using the LAST_ALTERED column. That column reflects metadata changes, not necessarily data changes. A table can be altered without new rows arriving, and a table can receive new rows without the metadata changing in the way the check expects. The check is a signal, not a proof.

Kafka consumer lag has its own timing semantics. The Confluent consumer configuration reference documents heartbeat.interval.ms as the expected time between heartbeats to the consumer coordinator when using group management, with a default of 3000 ms and a note that it should typically be no higher than one-third of session.timeout.ms. It documents session.timeout.ms as the timeout used to detect client failures, with a default of 45000 ms. It documents max.poll.interval.ms as the maximum delay between invocations of poll(), with a default of 300000 ms, after which the consumer is considered failed and the group rebalances.

Those three values interact. A consumer that processes a batch slowly can exceed max.poll.interval.ms and be evicted from the group even though it is healthy. A consumer that is paused for a deploy can exceed session.timeout.ms and trigger a rebalance. The alert that fires is “consumer lag,” but the cause may be a configuration value, not a data volume problem. The on-call engineer needs to know which one before touching anything.

Escalation paths and named owners

Google’s SRE book lists clear escalation paths, well-defined incident-management procedures, and a blameless postmortem culture as the most important on-call resources. It also describes primary and secondary on-call rotations, with duties varying by team: the secondary may be a fall-through for pages the primary misses, or may handle non-urgent production activities while the primary handles pages.

Data teams need the same structure, but the ownership map is harder. A pipeline may be owned by a data engineer, an analytics engineer, or a platform engineer. The Kafka cluster may be owned by a platform team. The warehouse may be owned by a separate group. When a page fires, the on-call engineer needs to know which of those owners to escalate to, and what response time to expect.

That information has to be written down before the rotation starts. A useful artifact is a one-page ownership table per pipeline: pipeline name, primary owner, secondary owner, escalation contact, expected response time, and the specific failure modes that justify a page. The table is not documentation for its own sake. It is the difference between a 10-minute escalation and a 40-minute Slack search at 3 a.m.

The SRE workbook describes playbooks as high-level instructions on how to respond to automated alerts, explaining severity and impact, and including debugging suggestions and possible actions. It also recommends implementing automation if playbooks are a deterministic list of commands the on-call engineer runs every time a particular alert fires. That recommendation is directly applicable to data pipelines. If the playbook for a freshness failure is “run this query, check this table, restart this task,” the playbook should be a script, not a document.

The maintenance tax in configuration

The recurring operational work in a data stack does not come from the big architectural decisions. It comes from configuration drift. Three examples, each anchored to a documented behavior.

Kafka consumer timeouts. The defaults for heartbeat.interval.ms, session.timeout.ms, and max.poll.interval.ms are tuned for general-purpose consumers. A consumer that does heavy per-record processing, or that pauses during a deploy, may need different values. Every change to those values is a change to the failure mode. The on-call engineer needs to know which consumers have non-default values and why.

Postgres statistics collection. PostgreSQL’s cumulative statistics system supports collection and reporting of information about server activity, including accesses to tables and indexes in disk-block and individual-row terms. Collection is controlled by parameters such as track_activities, track_counts, track_functions, and track_io_timing. The statistics views do not update instantaneously: each server process flushes accumulated statistics to shared memory just before going idle, but not more frequently than once per PGSTAT_MIN_INTERVAL milliseconds, so the displayed information lags behind actual activity. And when a server starts from an unclean shutdown — after an immediate shutdown, a server crash, a base backup, or point-in-time recovery — all statistics counters are reset.

That last point matters for on-call. A dashboard that depends on cumulative counters will show a discontinuity after a crash. An alert threshold based on a counter that just reset will either fire spuriously or fail to fire. The on-call engineer needs to know which dashboards are counter-based and which are gauge-based, and what happens to each after a restart.

dbt freshness check placement. As noted above, the checkbox and the run step produce different failure behavior. The choice is a maintenance decision. If the check is a run step, every freshness failure is a pipeline halt and a page. If the check is a checkbox, every freshness failure is a signal that someone has to review. The first option creates more pages. The second creates more silent failures. The team has to pick which tax it wants to pay.

A rotation checklist that fits on one page

Before anyone goes on-call for a data pipeline, the following should exist and be current. This is not a maturity model. It is the minimum set of artifacts that makes the rotation functional.

  • Ownership table. Pipeline name, primary owner, secondary owner, escalation contact, expected response time.
  • Alert classification. Which alerts page, which alerts create a ticket, and which alerts are informational. For each paging alert, the specific failure mode it represents.
  • Load budget. The team’s target for incidents per shift and hours of follow-up per incident, derived from its own incident history.
  • Playbook per paging alert. Severity, impact, debugging steps, and the actions that mitigate or resolve the alert. If the steps are deterministic, they should be a script.
  • Recovery procedures. How to restart a consumer group, how to re-run a failed dbt model, how to verify a backfill, and how to confirm that a schema change did not break a downstream consumer.
  • Handoff template. What the outgoing on-call engineer writes down: open incidents, in-progress backfills, known configuration changes, and anything that is likely to page in the next shift.

The SRE workbook describes a training approach that is worth adapting: a checklist of focus areas, lab sessions for common debugging and mitigation tasks, and “Wheel of Misfortune” exercises where the team role-plays recent incidents. For a data team, the equivalent is a game day that exercises the actual failure modes: a backfill that runs long, a schema change that breaks a consumer, a freshness check that fails silently, and an escalation that reaches the wrong person.

The game day is not a drill for its own sake. It is how the team discovers which parts of the rotation are missing before the pager discovers it for them.

What the schedule cannot do

A PagerDuty schedule answers one question: who is holding the phone. It does not answer who owns the pipeline, what the page means, how long the recovery should take, or who to call when the first three steps do not work.

Those answers come from a load budget measured in incidents per shift, an alert classification that separates pages from signals, an ownership table with named escalation contacts, and recovery procedures that have been tested. The schedule is necessary. It is not sufficient.

The test is simple. Ask the on-call engineer to describe, without looking anything up, what happens when a Kafka consumer group rebalances at 3 a.m. If the answer is a specific sequence of checks, a named escalation contact, and a known recovery procedure, the rotation is working. If the answer is “I would look at the dashboard and figure it out,” the schedule is doing all the work, and the pipeline has no owner.

FAQ

How many incidents per shift should a data on-call rotation target?

Google’s SRE workbook targets a maximum of two incidents per on-call shift to allow adequate time for follow-up. That number is a starting point, not a universal rule. A data team should derive its own target from the average time a real incident consumes, including triage, remediation, and postmortem. If a typical incident takes four hours, two incidents per shift already exceeds a normal working day.

Should dbt source freshness checks break the pipeline or just warn?

It depends on the SLA and the downstream dependency. The dbt documentation describes two behaviors: the Run source freshness checkbox runs the check as a non-breaking first step, while adding dbt source freshness as a run step causes the step to fail and subsequent steps not to run. The first produces a signal; the second produces a halt. The team should choose based on whether downstream models can tolerate stale source data.

Why does a Kafka consumer get evicted from its group even when it is healthy?

The Confluent consumer configuration reference documents max.poll.interval.ms as the maximum delay between invocations of poll(), with a default of 300000 ms. If the consumer’s processing loop takes longer than that between polls, the consumer is considered failed and the group rebalances. A consumer that processes large batches or pauses during a deploy can hit this limit without any underlying data problem.

What happens to Postgres monitoring after a crash?

The PostgreSQL documentation states that when a server starts from an unclean shutdown — after an immediate shutdown, a server crash, a base backup, or point-in-time recovery — all statistics counters are reset. Dashboards and alerts that depend on cumulative counters will show a discontinuity. The on-call engineer should know which monitoring depends on those counters and how the alert thresholds behave after a reset.

Do we need a secondary on-call rotation for data pipelines?

Google’s SRE book describes primary and secondary rotations with duties that vary by team. For a data team, the secondary serves two purposes: fall-through for pages the primary misses, and a second person who knows the recovery procedures. The second purpose is the more important one. If only one person knows how to recover a pipeline, the rotation has a single point of failure that the schedule does not address.