agentsclimarketplace

Data resiliency testing and failure injection

Skill vaquarkhan/data-engineering-agent-skills/skills/data-resiliency-testing-and-failure-injection

Production-grade Agent Skills for data engineering AI agents: 73 workflows, platform presets, safe backfill/replay, Kafka & Spark reliability, MCP observability, and VS Code/JetBrains installers.

Install
npx -y skills add vaquarkhan/data-engineering-agent-skills --skill data-resiliency-testing-and-failure-injection

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 21 stars21 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Guides agents through resiliency testing for data platforms. Use when designing or running failure drills, recovery validation, failover tests, replay-safety checks, dependency outage exercises, or fault injection for pipelines and publishes.

SKILL.md

4.5 KB, as published. Nobody here has run it

Data Resiliency Testing And Failure Injection

Overview

Use this skill when the goal is to prove that a data system recovers safely under failure, not only when everything goes right. It helps agents design controlled drills for retries, restarts, dependency outages, state recovery, duplicate prevention, backlog catch-up, and publish protection.

When to Use

  • hardening a production pipeline before broad rollout
  • testing failover, restart, replay, or checkpoint recovery behavior
  • validating that retries do not duplicate or corrupt data
  • proving recovery objectives for orchestrators, jobs, streams, or warehouse publishes
  • converting a past incident into a repeatable resilience drill

Do not treat resilience testing as random breakage. The point is to validate recovery behavior with explicit safety limits and evidence.

Workflow

  1. Define the failure modes that matter. Prioritize:

    • source outage or delayed upstream delivery
    • worker or task restart
    • orchestrator retry and timeout behavior
    • duplicate event or duplicate file delivery
    • checkpoint or incremental-state recovery
    • credential, secret, or network dependency failure
    • partial publish or downstream unavailability
  2. Define the resilience objectives. Include:

    • acceptable data loss behavior
    • recovery time objective
    • replay or backlog catch-up expectation
    • duplicate-prevention requirement
    • publish block or quarantine behavior
    • alert and escalation expectation
  3. Choose the safest drill environment. Prefer:

    • staging or isolated non-production
    • canary datasets or partitions
    • synthetic or masked test data
    • bounded windows and rollback-ready test scope
  4. Inject one failure mode at a time. Use controlled exercises such as:

    • killing a task or worker
    • pausing an upstream dependency
    • delaying input arrival
    • replaying a duplicate input
    • forcing an expired secret or denied permission in a safe environment
    • simulating partial output and validating publish closure
  5. Validate the recovery path. Check:

    • whether the system resumes or fails safely
    • whether alerts fire with useful context
    • whether duplicates are prevented
    • whether backlog catch-up stays bounded
    • whether publish remains blocked until validation passes
  6. Record guardrails and automate the highest-value drills. The best resilience test is one the team can rerun after changes, not a one-time exercise that gets forgotten.

  7. Load companion skills by failure mode.

    • replay or backfill drills: safe-backfill-and-replay-orchestration
    • Kafka lag, DLQ, or schema drift: kafka-resilience-and-schema-evolution
    • serverless Spark checkpoint recovery: spark-serverless-reliability-and-state-management
    • live diagnosis before drills: mcp-data-observability-integration
    • drill patterns: references/data-resiliency-testing-patterns.md

Common Rationalizations

RationalizationReality
"If the job retries, we are resilient enough."Retry alone does not prove replay safety, duplicate prevention, or publish protection.
"We can test recovery during a real incident."Real incidents are the worst time to discover the recovery path is unclear or unsafe.
"Failure injection is too risky for data systems."Uncontrolled failure is riskier than bounded, reviewable drills in safe environments.
"The scheduler health page already proves resilience."Scheduler status does not prove data correctness, backlog catch-up, or downstream safety.

Red Flags

  • no list of prioritized failure modes exists
  • retries are enabled without idempotency proof
  • resilience drills have no rollback or blast-radius limits
  • checkpoint or incremental-state recovery has never been tested
  • alerts fire but recovery ownership is unclear
  • a past incident has no corresponding regression drill

Verification

  • High-impact failure modes are named and prioritized
  • Recovery objectives and acceptable failure behavior are explicit
  • The drill scope is bounded and safe to run
  • Recovery evidence covers alerts, replay safety, duplicate prevention, and publish protection
  • At least one incident-derived failure mode is turned into a repeatable drill

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.