Eighteen months after a data mesh rollout, the platform team’s ticket queue is the only honest scoreboard. Not the architecture diagram. Not the governance charter. The queue.
This piece is about counting those tickets. Not to argue data mesh is wrong — Dehghani’s four principles are coherent and the central-team bottleneck is real — but to price what domain ownership actually shifts onto domain teams and what it leaves behind on the platform team’s on-call rotation.
What the principles actually assign
Dehghani’s original article names four principles: domain-oriented decentralized data ownership, data as a product, self-serve data infrastructure as a platform, and federated computational governance. Each principle maps to a concrete operational responsibility, and each responsibility generates a ticket category.
Domain ownership moves analytical data responsibility to the team closest to the source. Data as a product means that team owes consumers a contract — schema, semantics, quality attributes, SLOs. Self-serve platform means the platform team provides domain-agnostic tooling. Federated governance means policies are agreed across domains, not handed down.
None of those four principles say who fixes a broken Debezium connector at 02:00. That gap is where the tickets live.
Ticket taxonomy: what actually arrived
Across an eighteen-month window, platform tickets tend to fall into six buckets. The counts vary by org, but the categories are stable.
- Access and entitlement. Domain teams need read access to upstream data products. Federated governance says the domain owns the grant. In practice, the platform team owns the IAM plumbing, so every grant becomes a ticket.
- Pipeline failures on shared infrastructure. A domain’s dbt model fails because an upstream source changed. The domain owns the model. The platform owns the scheduler, the warehouse, and the connection pool. The ticket lands on the platform queue.
- Schema change coordination. A producer domain renames a column. Consumers break. The contract said this would be governed. The governance group meets biweekly.
- Data quality incidents. A dbt
not_nullorrelationshipstest fails. The domain owns the test. The platform owns the alerting route and the on-call rotation. - Backfill and CDC recovery. A Kafka topic is reset, or a Postgres logical replication slot falls behind, or an Iceberg snapshot expires before a consumer reads it. Someone has to decide who replays what.
- Cost and performance. A domain’s Snowflake query scans too many micro-partitions. The domain owns the model. The platform owns the credit budget.
Buckets 1, 2, 4, and 6 are the ones that quietly grow. They are not new work in the sense that the work did not exist before. They are reclassified work: tasks the central team used to do silently, now surfaced as tickets because the ownership boundary is explicit.
The reclassification problem
This is the single most important measurement issue. If you count tickets before and after a data mesh rollout and conclude that platform load increased, you may be measuring a change in visibility, not a change in work.
Before the mesh, a central data engineer fixed a broken pipeline, updated a schema, and re-ran a backfill without opening a ticket. After the mesh, the same fix requires a domain team to request platform action, which opens a ticket. The work is identical. The ticket count is not.
To separate new work from reclassified work, you need a counterfactual. The cleanest one is a pre-rollout baseline of platform-team hours by activity, not ticket count. If you did not capture that baseline, you cannot honestly claim the mesh increased load. You can only claim it changed the shape of the queue.
What the platform team actually owns after the mesh
The self-serve platform principle is the one most often under-specified. Dehghani’s article describes a multi-plane platform: a control plane for policy and a data plane for compute and storage. The platform team owns the control plane. Domain teams own their data products on the data plane.
In practice, the control plane is where the tickets concentrate. Access control, schema registry, catalog metadata, lineage, alerting routes, and cost attribution all live in the control plane. Every domain team that onboards adds a row to each of those systems, and every row is a potential ticket.
The platform team’s on-call load does not shrink when domain teams take ownership of their pipelines. It shifts from “fix the pipeline” to “fix the platform primitive that the pipeline depends on.” That is a different skill set and a different pager.
Schema evolution: the mechanics that become domain-team work
Schema evolution is where the ownership boundary is most visible and most expensive.
Apache Iceberg’s spec supports safe column add, drop, reorder, and rename, including in nested structures. That is a format-level guarantee. It does not tell you which consumer will break when a producer renames a column. The format allows the change. The contract governs the change. The domain team executes the change.
If the domain team has never run a schema evolution against a live consumer, the first one is expensive. The failure mode is not the rename itself. It is the downstream dbt model that selects column_name and now returns nulls, or the Kafka consumer that deserializes against an old Avro schema and drops the record.
dbt’s data tests are the cheapest guardrail here. A not_null test on a column that should never be null will fail the moment a rename lands. A relationships test will fail when a foreign key stops resolving. Both are SQL queries that return failing rows; if the query returns zero rows, the assertion passes. That is the entire mechanism. It is not sophisticated. It is also the difference between catching a schema break in the PR and catching it in a dashboard three days later.
The ticket that follows a missed schema break is not a schema ticket. It is a data quality ticket, a backfill ticket, and a trust ticket, in that order.
CDC and backfill: the recovery path that nobody owns
Change data capture is the second place where the ownership boundary frays.
Debezium reads the Postgres write-ahead log through a logical replication slot. If the slot falls behind — because the consumer stalled, because the connector restarted, because the network partition lasted longer than the WAL retention — the slot is dropped and the change stream has a gap. Recovering that gap requires a backfill from the source table, which requires a snapshot, which requires coordination with the source team, which is a different domain.
Who owns that ticket? The producer domain owns the source table. The consumer domain owns the downstream model. The platform team owns the connector. Three teams, one gap, and no single owner.
The same pattern appears with Iceberg. Snapshots are retained per the table’s snapshot retention policy. If a consumer reads a snapshot that has expired, the read fails. The fix is a re-read from a newer snapshot, which may not be equivalent if the consumer was doing incremental processing. That is a backfill, and it is a domain-team responsibility that the domain team may not have known it signed up for.
Airflow: where dynamic task mapping hides the cost
Dynamic task mapping is a common pattern in domain-owned pipelines. A task generates a list at runtime, and the scheduler creates one task instance per element. The number of task instances is not known at DAG parse time.
This is useful. It is also a ticket generator. When a mapped task fails, the failure is per-instance, and the retry semantics depend on whether the mapping was task-generated or static. Airflow’s documentation notes that task-generated mapping cannot be used with TriggerRule.ALWAYS, because the expanded parameters are undefined at the time the task would execute. That constraint is enforced at DAG parse time.
The operational consequence: a domain team that adopts dynamic task mapping without understanding the trigger rule constraint will hit a parse error, open a ticket, and wait. The fix is a one-line change. The ticket is not.
Snowflake: the clustering key that becomes a platform ticket
Clustering keys are the clearest example of a domain-owned decision that generates platform-owned cost.
Snowflake’s documentation is explicit: clustering keys are not intended for all tables, because of the cost of initially clustering the data and maintaining the clustering. Reclustering consumes credits and generates new micro-partitions. The original micro-partitions are retained for Time Travel and Fail-safe, which means storage costs increase.
A domain team that adds a clustering key to a large table to fix a slow query is making a cost decision on the platform team’s budget. If the table has high DML volume, the reclustering cost is ongoing. The domain team sees a faster query. The platform team sees a credit line item.
The ticket that follows is not a clustering ticket. It is a cost review ticket, and it arrives at the end of the month, when the domain team has moved on.
Pricing the maintenance tax
Here is a reproducible way to estimate the per-domain maintenance tax without inventing numbers.
- Export the platform ticket queue for the last six months. Group by domain and by category using the six buckets above.
- For each ticket, record the time from open to close, and the number of distinct people who touched it.
- Separate tickets that required a code change from tickets that required only a configuration change or an access grant.
- Multiply configuration and access tickets by the average handling time. That is the coordination tax.
- Multiply code-change tickets by the average handling time plus the average review time. That is the engineering tax.
- Add the platform team’s on-call hours attributable to domain-owned pipelines. That is the pager tax.
The sum is the monthly maintenance tax per domain. It is not a research finding. It is an arithmetic exercise on your own queue. The number will be different in every org, and that is the point.
What the eighteen-month count actually shows
Three patterns tend to hold across orgs that run this exercise honestly.
First, the platform team’s ticket volume does not fall. It changes composition. Access and cost tickets grow. Pipeline-fix tickets shrink. The net is roughly flat, and the skill mix shifts toward platform primitives.
Second, domain teams underestimate the coordination cost of federated governance. Every cross-domain schema change requires a conversation, and every conversation has a latency. The latency is the tax.
Third, the tickets that hurt most are the ones with no clear owner: CDC gaps, expired snapshots, and schema breaks that cross a domain boundary. These are the tickets that sit in the queue the longest, because no single team’s on-call rotation owns them.
What to do with the count
The count is not an argument against data mesh. It is an argument for naming the ownership boundary explicitly, in writing, for the three failure modes that cross domains: CDC gaps, snapshot expiry, and cross-domain schema changes.
For each, write down: who detects it, who decides the recovery, who executes the backfill, and who pays for the compute. If any of those four is “the platform team” by default, the mesh is not fully implemented. It is a mesh with a central fallback, and the fallback is the queue.
The platform team’s ticket queue is the honest scoreboard. Read it before you read the architecture diagram.
FAQ
How do I tell new work from reclassified work?
Compare platform-team hours by activity, not ticket count. If you did not capture a pre-rollout baseline, you cannot make the distinction. Capture it now and compare forward.
Does data mesh reduce platform load?
The principles do not promise that. They promise to move ownership of analytical data to domain teams. Platform load shifts from pipeline fixes to platform primitives. The net depends on how many domains onboard and how much of the control plane is centralized.
What is the cheapest guardrail against schema breaks?
dbt data tests. A not_null test on a column that should never be null, and a relationships test on a foreign key, will catch most renames and drops before consumers see them. Both are SQL queries that return failing rows.
Who owns a CDC gap?
Nobody, by default. That is the problem. Assign detection, decision, execution, and cost explicitly, or the ticket will sit in the platform queue.
Is clustering a domain decision or a platform decision?
It is a domain decision with a platform cost. Snowflake’s documentation is clear that clustering is not for all tables and that reclustering consumes credits and storage. Treat the clustering key as a budget decision, not a performance decision.
Sources
- Zhamak Dehghani, “Data Mesh Principles and Logical Architecture,” martinfowler.com, 03 December 2020. https://martinfowler.com/articles/data-mesh-principles.html
- Data Mesh Architecture, “What Is Data Mesh?” https://www.datamesh-architecture.com/
- dbt Docs, “Data tests.” https://docs.getdbt.com/docs/build/data-tests
- Apache Airflow Documentation, “Dynamic Task Mapping.” https://airflow.apache.org/docs/apache-airflow/stable/authoring-and-scheduling/dynamic-task-mapping.html
- Apache Iceberg Table Spec. https://iceberg.apache.org/spec/
- Snowflake Documentation, “Clustering Keys & Clustered Tables.” https://docs.snowflake.com/en/user-guide/tables-clustering-keys