A producer team renames user_id to account_id in an Avro record. They run the schema through Schema Registry. It is accepted. They post in the shared channel: “Non-breaking change, no consumer action needed.”
Two days later, a consumer job fails to deserialize. A dbt incremental model silently stops populating a column. A CDC pipeline replays a batch it already processed.
The rename was not non-breaking. It was non-breaking relative to the compatibility mode that was actually in effect, the schema format in use, and the set of consumers that actually read the topic. The producer team checked one of those three things.
This is not a story about a careless team. It is a story about a review process that treats “the registry accepted it” as equivalent to “nothing downstream will break.” Those are different claims, and the gap between them is where on-call hours go to die.
What the registry actually checks
Schema Registry enforces compatibility by comparing a new schema version against previous versions using a configurable compatibility type. The default is BACKWARD, not BACKWARD_TRANSITIVE. That distinction matters more than most teams realize.
Under BACKWARD, a consumer using the new schema can process data written by producers using schema X or X-1, but not necessarily X-2. Under BACKWARD_TRANSITIVE, that same consumer can process data written by X, X-1, or X-2. The Confluent documentation is explicit: the default is BACKWARD, and the main reason is so that you can rewind consumers to the beginning of the topic.
Here is the failure mode. A team has three schema versions in production. They add a fourth. The registry checks version 4 against version 3 under BACKWARD. It passes. But a consumer that rewinds to the beginning of the topic will encounter version 1 and version 2 messages. If the change is not transitive-safe, that consumer breaks on old data.
The registry did its job. The review did not.
What each compatibility type actually permits
The Confluent compatibility tables for Avro and Protobuf show which operations are allowed under each mode. Adding an optional field is compatible under BACKWARD, FORWARD, and FULL. Removing an optional field is also compatible under all three. Adding a required field is compatible only under FORWARD. Removing a required field is compatible only under BACKWARD.
Renames are not listed as a distinct operation. In Avro, a rename is typically expressed as a field removal plus a field addition. Whether that passes depends on whether the removed field was optional or had a default value, and whether the added field has a default value. The documentation states: “the ability to delete a field and keep the schema compatible requires that the field was either specified as optional or provided a default value in the original version.”
So a rename can pass BACKWARD if the old field had a default and the new field has a default. It can pass FORWARD under similar conditions. It can pass FULL if both conditions hold. But passing the registry check does not mean consumers will find the data they expect. A consumer looking for user_id will not find it in a message that only contains account_id. The registry does not know what field names your consumer code references.
Schema format changes the rules
Avro, Protobuf, and JSON Schema have different compatibility rules. The Confluent documentation notes that Avro was developed with schema evolution in mind and its specification clearly states the rules for backward compatibility, whereas the rules for JSON Schema and Protobuf can be more nuanced.
For JSON Schema, compatibility behavior depends on both the compatibility policy (lenient or strict) and the content model (additionalProperties: true for open, false for closed). A change that passes under a lenient policy with an open content model may fail under a strict policy with a closed content model. The review must confirm which policy and which content model are actually in effect.
For Protobuf, the documentation notes that best practice is to use BACKWARD_TRANSITIVE, because adding new message types is not forward compatible. A team using BACKWARD with Protobuf may accept a change that breaks forward compatibility in ways the registry does not flag.
The effective compatibility mode may not be what you think
A REST API call to compatibility mode is global and overrides any compatibility parameters set in schema registry properties files. This means the effective mode for a subject may differ from what the properties file says. A team that set BACKWARD_TRANSITIVE in their properties file may find that a global API call reset it to BACKWARD.
The review must verify the effective mode for the specific subject, not the mode someone believes is configured. The diagnostic is straightforward:
curl -s http://schema-registry:8081/config
curl -s http://schema-registry:8081/config/<subject-name>
The first call returns the global compatibility level. The second returns the subject-level override, if any. If the subject-level value is absent, the global value applies. If a global API call was made, it overrides the properties file.
What the registry does not check
The registry checks schema compatibility. It does not check:
- Whether consumers have been rewound to the beginning of the topic
- Whether dbt incremental models will pick up the change
- Whether downstream CDC pipelines handle the renamed field
- Whether any consumer code references the old field name
- Whether the change is transitive-safe across all schema versions in the topic
Each of these is a separate failure mode. Each requires a separate check.
The dbt layer: silent column drops
dbt incremental models have an on_schema_change configuration. The default is ignore. Under ignore, if you add a column to your incremental model and execute a dbt run, the column will not appear in the target table. If you remove a column and execute a dbt run, dbt will fail.
This means a renamed field can produce two different failure modes depending on which side of the rename the dbt model sees. If the model references the old field name and the source no longer provides it, the run fails. If the model references the new field name and the source provides it, but the target table was built with the old schema, the new column silently does not appear.
The documentation is explicit: “None of the on_schema_change behaviors backfill values in old records for newly added columns.” If you need to populate those values, you must run manual updates or trigger a --full-refresh.
There is another constraint: on_schema_change only tracks top-level column changes. It does not track nested column changes. A rename inside a nested structure will not trigger a schema change, even if on_schema_change is set appropriately.
The diagnostic is to check which models use incremental materialization and what their on_schema_change setting is:
dbt ls -s config.materialized:incremental --output json | jq '.[].config.on_schema_change'
If the output is null or "ignore", the model will not pick up new columns automatically.
The CDC layer: replay and idempotency
PostgreSQL logical decoding slots emit each change once in normal operation. But the current position of each slot is persisted only at checkpoint. In the case of a crash, the slot might return to an earlier LSN, which will cause recent changes to be sent again when the server restarts.
The documentation states: “Logical decoding clients are responsible for avoiding ill effects from handling the same message more than once.”
This means a CDC pipeline that consumes a renamed field must be idempotent against replay. If the pipeline processes a message with user_id, then a message with account_id, then a replayed message with user_id, it must not produce duplicate or inconsistent rows.
Replication slots persist across crashes and know nothing about the state of their consumers. They will prevent removal of required resources even when there is no connection using them. A slot that is no longer required should be dropped, but dropping it requires knowing which consumers depend on it.
The diagnostic is to check which slots exist and how far behind they are:
SELECT slot_name, plugin, slot_type, active, restart_lsn,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS lag
FROM pg_replication_slots;
A slot with a large lag and no active connection is a candidate for investigation. It may be holding WAL for a consumer that no longer exists, or it may be a consumer that is about to replay a large batch.
The Iceberg layer: column IDs and rename safety
The Iceberg spec states that schema evolution supports safe column add, drop, reorder and rename, including in nested structures. This is a stronger guarantee than what Schema Registry provides for Kafka topics, because Iceberg tracks column IDs rather than column names.
But the spec also notes that the format version number is incremented when new features are added that will break forward-compatibility. This means a table written with format version 3 may not be readable by a reader that only supports version 2. The review must confirm which format version the table uses and whether all readers support it.
The diagnostic is to check the table metadata:
SELECT * FROM catalog.db.table_snapshots LIMIT 1;
-- or
SELECT * FROM "table$snapshots" LIMIT 1;
The format-version field in the table metadata indicates which version the table uses. If the table was upgraded from version 2 to version 3, readers that only support version 2 will fail.
What a cross-team schema review must actually check
The review is not a single check. It is a sequence of checks, each of which can fail independently.
- Confirm the effective compatibility mode for the specific subject. Use the Schema Registry API to check both the global and subject-level settings. Do not rely on the properties file.
- Confirm the schema format and its specific rules. Avro, Protobuf, and JSON Schema have different compatibility rules. JSON Schema compatibility depends on both the policy and the content model.
- Confirm whether the change is transitive-safe. If any consumer rewinds to the beginning of the topic, the change must be compatible with all schema versions, not just the last one.
- Confirm whether any consumer rewinds to the beginning of the topic. This is a property of the consumer configuration, not the schema. Check consumer group offsets and retention settings.
- Confirm whether downstream dbt models use
on_schema_changeand what setting. The default isignore, which means new columns silently do not appear. - Confirm whether CDC pipelines are idempotent against replay. Logical decoding slots can replay changes after a crash. The pipeline must handle duplicate messages.
- Confirm whether the change is compatible with the Iceberg table format version in use. A table upgraded to a newer format version may not be readable by older readers.
Pricing the maintenance tax
The cost of a schema change is not paid at the moment of the change. It is paid when a consumer fails to deserialize, when a dbt incremental model silently drops a column, or when a CDC pipeline replays a change it already processed.
Each of these failure modes has a different cost profile:
- Consumer deserialization failure: The consumer stops processing. If it is a real-time pipeline, the lag grows. If it is a batch pipeline, the batch fails. The fix is to update the consumer code and redeploy. The cost is the time to diagnose, fix, and redeploy, plus the cost of any data that was not processed during the outage.
- dbt silent column drop: The model runs successfully but produces incomplete data. The failure is not detected until someone notices that a column is null or missing. The fix is to run a full refresh, which may be expensive if the model processes a large volume of data. The cost is the compute cost of the full refresh plus the time to diagnose why the column disappeared.
- CDC replay: The pipeline processes a message it already processed. If the pipeline is not idempotent, it produces duplicate rows. The fix is to deduplicate the data and make the pipeline idempotent. The cost is the time to diagnose the duplication plus the cost of the deduplication job.
The review should price these failure modes in on-call hours, not just in registry API calls. A review that takes 30 minutes and catches a non-transitive change is cheaper than a review that takes 5 minutes and misses it.
Frequently asked questions
Does Schema Registry check for field renames?
Schema Registry checks compatibility based on the rules for the schema format and compatibility type. In Avro, a rename is typically expressed as a field removal plus a field addition. Whether that passes depends on whether the removed field was optional or had a default value, and whether the added field has a default value. The registry does not know what field names your consumer code references.
What is the difference between BACKWARD and BACKWARD_TRANSITIVE?
Under BACKWARD, a consumer using the new schema can process data written by producers using schema X or X-1, but not necessarily X-2. Under BACKWARD_TRANSITIVE, that same consumer can process data written by X, X-1, or X-2. The default is BACKWARD.
Why does my dbt incremental model not pick up a new column?
The default on_schema_change setting is ignore. Under ignore, if you add a column to your incremental model and execute a dbt run, the column will not appear in the target table. You must set on_schema_change to append_new_columns or sync_all_columns, or run a full refresh.
Can a CDC pipeline process the same message twice?
Yes. PostgreSQL logical decoding slots persist their position only at checkpoint. In the case of a crash, the slot might return to an earlier LSN, which will cause recent changes to be sent again when the server restarts. Logical decoding clients are responsible for avoiding ill effects from handling the same message more than once.
Does Iceberg handle column renames safely?
The Iceberg spec states that schema evolution supports safe column add, drop, reorder and rename, including in nested structures. This is because Iceberg tracks column IDs rather than column names. However, the format version number is incremented when new features are added that will break forward-compatibility, so readers must support the table’s format version.
Sources
- Confluent Schema Registry documentation, “Schema Evolution and Compatibility”: https://docs.confluent.io/platform/current/schema-registry/fundamentals/schema-evolution.html
- dbt documentation, “Incremental models”: https://docs.getdbt.com/docs/build/incremental-models
- PostgreSQL documentation, “Logical Decoding Concepts”: https://www.postgresql.org/docs/current/logicaldecoding-explanation.html
- Apache Iceberg documentation, “Iceberg Table Spec”: https://iceberg.apache.org/spec/
- Apache Airflow documentation, “Dynamic Task Mapping”: https://airflow.apache.org/docs/apache-airflow/stable/authoring-and-scheduling/dynamic-task-mapping.html