The Freshness Alert Fired 40 Times a Day Until the Channel Muted It: Data Alert Design With Burn-Rate Budgets Instead of Static Thresholds

Somewhere in your Slack workspace there is a channel called #data-alerts. It fired forty times yesterday. Nobody read it. The one message that mattered — a source table that stopped receiving CDC events at 02:14 and stayed stale for six hours — scrolled past under thirty-nine “source X is 61 minutes stale” notifications that resolved themselves within two minutes.

That channel is not a monitoring failure. It is a design failure. The alert was written to answer the wrong question.

A static freshness threshold answers: is this one file late right now? The operational question is: are we on pace to miss the SLA? Those are different questions, and only one of them deserves a page.

What a static threshold actually measures

Take a dbt source with warn_after: {count: 1, period: hour} and error_after: {count: 2, period: hour}. The check compares the max load timestamp in the source against wall-clock time. If the upstream extractor runs at 02:00 and finishes at 02:03, the check at 02:05 passes. If it finishes at 03:01, the check fails. The alert fires on a single late file.

The Google SRE Workbook makes the cost of this pattern explicit. In its first alerting iteration — alert when the recent error rate equals the SLO over a short window — the authors note you could receive up to 144 alerts per day every day, not act upon any alerts, and still meet the SLO (SRE Workbook, Chapter 5). That is the arithmetic of a short window: excellent detection time, terrible precision. Every late file crosses the line. Almost none of them threaten the SLA.

The obvious fix — widen the window — has its own failure mode. The Workbook’s second iteration uses a 36-hour window to ensure only sustained problems alert. Precision improves. Reset time collapses: in the case of 100% outage, an alert will fire shortly after 2 minutes, and continue to fire for the next 36 hours. You have traded a noisy channel for a channel that lies to you for a day and a half after the incident is over.

The third iteration — add a for: 1h duration — is worse. The Workbook is blunt: because the duration does not scale with the severity of the incident, a 100% outage alerts after one hour, the same detection time as a 0.2% outage. A service that spikes to 100% errors for five minutes every ten minutes never triggers the alert at all, despite consuming 35% of the monthly budget. Duration parameters do not measure severity. They measure persistence.

Reframe freshness as an SLO with a budget

If your freshness SLA is “data is at most one hour old,” you have already defined an SLO. The good event is a load that arrives within the SLA. The bad event is a load that arrives late. The error budget is the fraction of loads you are allowed to deliver late over the measurement window — typically 30 days.

Once you frame it that way, a single late file is not an incident. It is budget spend. The pager should fire when the rate of budget spend threatens to exhaust the budget before the window closes. That rate is the burn rate.

Burn rate 1 means you are consuming budget at exactly the pace that exhausts it at the end of the window. Burn rate 2 exhausts it in half the time. The Workbook’s table for a 99.9% SLO over 30 days: burn rate 1 corresponds to a 0.1% error rate and 30 days to exhaustion; burn rate 10 corresponds to 1% and 3 days; burn rate 1,000 corresponds to 100% and 43 minutes.

For freshness, the mapping is direct. If your SLA is one hour and your measurement window is 30 days, then a source that is late 0.1% of the time is burning at rate 1. A source that is late 1% of the time is burning at rate 10 and will exhaust the budget in three days. A source that stops entirely is burning at rate 1,000 and will exhaust it in 43 minutes.

The rule shape that replaces the threshold

The Workbook recommends multiwindow, multi-burn-rate alerts as the most viable option. The recommended starting numbers: 2% budget consumption in one hour and 5% budget consumption in six hours as reasonable starting numbers for paging, and 10% budget consumption in three days as a good baseline for ticket alerts.

Translated into burn rates for a 30-day window:

  • Page: 14.4x burn rate over 1 hour, confirmed by 14.4x over 5 minutes. This is 2% of the budget in one hour.
  • Page: 6x burn rate over 6 hours, confirmed by 6x over 30 minutes. This is 5% of the budget in six hours.
  • Ticket: 1x burn rate over 3 days. This is 10% of the budget in three days.

The short confirmation window is the key mechanism. The Workbook’s guideline: make the short window 1/12 the duration of the long window. The long window establishes that a significant amount of budget has been spent. The short window establishes that the budget is still being spent. If the short window has recovered, the alert does not fire — which is exactly the class of alert that trains people to mute the channel.

In Prometheus, this maps onto the alerting rule primitives directly. The for clause waits a duration before firing; the keep_firing_for clause keeps the alert firing after the condition was last met, which the documentation describes as useful to prevent situations such as flapping alerts, false resolutions due to lack of data loss, etc. (Prometheus alerting rules). The same page is explicit that Prometheus alerting rules are not a notification solution: another layer is needed to add summarization, notification rate limiting, silencing and alert dependencies. That layer is Alertmanager. If you skip it, you will rebuild rate limiting and silencing by hand, badly.

Wiring it to the actual stack

dbt source freshness and Prometheus burn-rate rules are complementary, not competing. dbt produces the SLI — how stale is the source. Prometheus produces the alerting logic — how fast is the budget burning.

Two dbt behaviors matter for this design. First, dbt build does not include source freshness checks; you either select the “Run source freshness” checkbox in the job, which runs it as the first step and won’t break subsequent steps if it fails, or you add dbt source freshness as a run step, in which case if your source data is out of date — this step will “fail”, and subsequent steps will not run (dbt source freshness docs). The checkbox is the right choice if you want freshness to feed an alerting pipeline without blocking the models. The run step is the right choice if you want stale data to halt the DAG. Pick one deliberately; the default behavior of the run step has surprised more than one on-call engineer at 03:00.

Second, check frequency is not optional. dbt’s own guidance: you should run your source freshness jobs with at least double the frequency of your lowest SLA. If your SLA is one hour, run the check every 30 minutes. A daily freshness check against a one-hour SLA measures nothing useful — it tells you whether the source was fresh at the moment you happened to look.

The recording rule that turns freshness into a ratio is straightforward. Emit a gauge per source: 1 if fresh, 0 if stale. Then:

record: source:freshness_ratio_rate1h
expr: sum(rate(source_fresh[1h])) by (source)
      / sum(rate(source_expected[1h])) by (source)

The alert then compares that ratio against the burn-rate threshold. For a 99.9% freshness SLO, the 14.4x page rule is source:freshness_ratio_rate1h < (1 - 14.4 * 0.001) combined with the 5-minute confirmation. The exact numbers depend on your SLA; the shape does not.

Where the freshness SLI lies to you

A freshness check that only reads the max timestamp in the target table can miss a permanently skipped file. Snowpipe is the clearest example. Its file-loading metadata is maintained for 14 days. Files that failed to load — because of invalid content or stage access failures — are still registered in the pipe’s metadata, and the registered file names are ignored by subsequent pipe activity, including ALTER PIPE … REFRESH (Snowpipe troubleshooting). The target table’s max timestamp can look perfectly fresh while a specific partition is permanently missing rows.

The same page documents the inverse trap: files modified and staged again after 14 days are loaded again, potentially duplicating records. A freshness SLI that only measures recency will not catch either case. Pair it with a load-history check — COPY_HISTORY for status, SYSTEM$PIPE_STATUS for lastReceivedMessageTimestamp versus lastForwardedMessageTimestamp. The gap between those two timestamps distinguishes a service configuration problem from a path mismatch between the stage and pipe definitions. A freshness alert alone cannot make that distinction.

What the fix costs

Define the SLO per source: one hour of work per source if the SLA is already documented, half a day if it is not. Write the recording rule and the three alert rules: two to four hours for the first source, thirty minutes for each subsequent one once the pattern is templated. Wire Alertmanager routing and suppression: half a day, once.

The ongoing maintenance tax is real and should be priced honestly. You now have more windows, more thresholds, and more numbers to reason about. The Workbook names this directly as a disadvantage of multi-burn-rate alerting. The three-day ticket window also produces a longer reset time than a short-window alert would. And you need alert suppression, because a 10% budget spend in five minutes also means 5% was spent in six hours and 2% in one hour — three conditions true, three notifications, unless the monitoring system prevents it.

The trade is fewer pages and a ticket queue that catches slow burns. For an on-call engineer who was about to mute the channel, that is usually the right trade. It is not free, and it is not a platform migration — it is a few hours per source plus a routing layer you probably already have.

The operational rule

The pager should fire on budget spend, not on a single late file. The ticket queue catches the slow burns that would otherwise exhaust the budget unnoticed. The muted channel is the symptom; the static threshold is the cause.

If you want to test this on one source before committing: pick the noisiest freshness alert you have, count how many times it fired last week, and count how many of those firings corresponded to a real SLA miss. If the ratio is worse than 10:1, the threshold is measuring the wrong thing. Replace it with a burn-rate rule and watch the channel go quiet — not because you muted it, but because it stopped lying to you.

FAQ

Do I need Prometheus to do this? No. The Workbook’s examples use Prometheus syntax, but the page states the approach applies in any alerting framework. What you need is a way to compute a ratio over a window and compare it to a threshold, plus a notification layer that can suppress and route. Most observability platforms have both.

What if my SLA is not a clean number like 99.9%? The burn rate is derived from the SLA, not the other way around. If your SLA is 99% over 30 days, the error budget is 1%, and burn rate 1 corresponds to a 1% error rate. The recommended budget-consumption percentages (2% in 1h, 5% in 6h, 10% in 3d) stay the same; the burn-rate multipliers change.

Should I keep the old static threshold as a backstop? Only if you route it to a ticket queue, not a pager. A static threshold that pages is the problem you are trying to solve. A static threshold that opens a ticket is a cheap safety net for the case where your SLI pipeline itself breaks.

How do I handle sources with no fixed SLA? You cannot burn-rate alert on an undefined budget. Either define the SLA — even a loose one — or route the source to a dashboard and a weekly review. Paging on an undefined SLA is how you get forty alerts a day.