Kafka resilience and schema evolution
Skill vaquarkhan/data-engineering-agent-skills/skills/kafka-resilience-and-schema-evolution
Production-grade Agent Skills for data engineering AI agents: 73 workflows, platform presets, safe backfill/replay, Kafka & Spark reliability, MCP observability, and VS Code/JetBrains installers.
npx -y skills add vaquarkhan/data-engineering-agent-skills --skill kafka-resilience-and-schema-evolutionAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 21 stars21 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Enforces production Kafka guardrails including non-breaking schema evolution, dead-letter queues for poison messages, and acks=all producer durability. Use when designing or changing Kafka topics, producers, consumers, schema registry policies, or streaming recovery paths.
SKILL.md
4.7 KB, as published. Nobody here has run it
Kafka Resilience And Schema Evolution
Overview
Generic streaming guidance is not enough for production Kafka. Agents routinely introduce breaking schema changes, under-provisioned durability settings, and missing poison-message isolation. This skill mandates enforceable broker, producer, consumer, and registry guardrails before any production change ships.
When to Use
- creating or modifying
Kafkatopics, producers, or consumers - setting or changing schema registry compatibility policies
- designing dead-letter queue (DLQ) routing for poison pill messages
- hardening producer durability (
acks, retries, idempotence) - reviewing consumer lag, replay, or failover behavior on Kafka-backed pipelines
Pair with streaming-and-messaging-systems for broader event design. Pair with avro-protobuf-json-schema-registry when registry subjects and compatibility CI are in scope.
Workflow
-
Define the production contract before broker changes. Document:
- topic key strategy and partition count rationale
- retention, compaction, and replay policy
- schema format and registry subject naming
- consumer groups and downstream sinks
- delivery semantics target (at-least-once with idempotent sinks, or stricter)
-
Enforce producer durability defaults. Require unless explicitly waived with owner approval:
acks=all(oracks=-1)enable.idempotence=truewhen ordering and deduplication matter- bounded
retrieswithdelivery.timeout.msaligned to SLA max.in.flight.requests.per.connection=1when strict ordering is required- TLS/SASL configuration documented for non-development clusters
-
Block breaking schema evolution. Before any schema change:
- set compatibility policy per subject (
BACKWARD,FORWARD, orFULL— notNONEin production) - run compatibility checks in CI against registered schemas
- document producer-then-consumer or consumer-then-producer rollout order
- reject field removals, renames, or type changes without migration plan
- load
references/kafka-production-guardrails.mdfor DLQ and evolution patterns
- set compatibility policy per subject (
-
Mandate dead-letter and poison pill isolation. Every production consumer that parses external payloads must define:
- DLQ topic or sink with retention and access controls
- classification rules (deserialization failure, schema mismatch, business rule violation)
- alert routing when DLQ rate exceeds threshold
- replay procedure with deduplication keys
- no silent drop of unparseable records
-
Make lag and recovery observable. Plan for:
- consumer group lag alerts with owner routing
- offset reset policy documented and restricted
- replay runbook that does not bypass DLQ classification
- broker disk and retention monitoring for high-throughput topics
-
Load MCP observability when diagnosing live lag. Use
mcp-data-observability-integrationwithmcp/kafka.mcp.jsonto inspect consumer group lag and topic metadata before changing consumer code or partition counts.
Common Rationalizations
| Rationalization | Reality |
|---|---|
| "acks=1 is fine because Kafka is durable." | Leader acknowledgment without full ISR acknowledgment loses events under failure scenarios. |
| "We can fix schema breaks by redeploying consumers quickly." | Breaking changes propagate to many consumers and batch sinks before redeploy completes. |
| "DLQs add too much operational overhead." | Poison pills without DLQs stall partitions, inflate lag, and hide data loss as consumer retries. |
| "Schema compatibility NONE is okay for internal topics." | Internal topics still feed warehouses, stream processors, and audit systems. |
Red Flags
- production subjects use
NONEcompatibility - producers use
acks=0oracks=1without documented waiver - consumers have no DLQ path for deserialization failures
- schema changes ship without CI compatibility validation
- consumer group lag has no alert owner
- replay procedures reset offsets without reconciliation or publish pause
Verification
- Producer durability settings meet
acks=alland idempotence requirements - Schema compatibility policy is set and CI-validated for production subjects
- DLQ routing, alerts, and replay procedure are documented
- Consumer lag and retention monitoring exist with named owners
- Rollout order for schema changes is explicit and tested in non-production