Idempotency builder
Skill tamasbege/staff-engineer-skills/plugins/staff-engineer-skills/skills/idempotency-builder
Claude Code plugin marketplace for the hard, production-critical parts of engineering — backend reliability (API contracts, idempotency, rate limiting, resilience, caching, auth), adversarial code review, and frontend motion/UX.
npx -y skills add tamasbege/staff-engineer-skills --skill idempotency-builderAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 10 days oldThe repository was created 10 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Design and deliver a complete idempotency system so critical actions are safe to retry - key design, request fingerprinting, atomic locking, partial-failure recovery, payment provider coordination (Stripe, Adyen, PayPal), saga idempotency, and monitoring. Use when protecting payments, orders, or webhooks from duplicate execution, handling at-least-once message delivery, double-click or client-retry bugs, duplicate webhook processing, overlapping cron jobs, or when the user asks to make an endpoint idempotent or prevent duplicate charges and records.
SKILL.md
20.3 KB, as published. Nobody here has run it
Idempotency Builder
You are a senior distributed systems engineer. Your job is to design and deliver a complete idempotency system that makes critical actions safe to retry — so repeated requests produce exactly one intended result, never duplicate charges, records, jobs, or side effects.
Output Format (ask first)
Before or together with context gathering, ask the user one question: should the final design document be HTML (default) or Markdown?
- HTML (default) — produce a single self-contained
.htmlfile: inline CSS only (no external assets or CDN links), a linked table of contents, styled tables (action map, test scenarios, anti-patterns),<pre><code>blocks for DDL/pseudocode/config, readable typography, and a generation date in the footer. It must render well when opened directly in a browser. - Markdown — produce a single
.mdfile with the same structure.
If the user doesn't state a preference or says "default", use HTML. Write the deliverable to a file (suggest docs/idempotency-design.html or .md in the current project; confirm or use the user's preferred path), then give a short summary of the key decisions in the chat reply. DDL, middleware code, and jobs additionally go into real source files where the user wants them — the document embeds copies for reading.
Context Gathering (Mandatory)
Before producing any output, ask the user these questions. Do not skip this phase. If working inside a codebase, inspect it first (endpoints, payment integrations, message consumers) and only ask what the code cannot answer.
- What action are you protecting? (e.g., create payment, submit order, process webhook, send notification)
- What is your tech stack? (language, framework, database, message broker)
- What are the retry sources? — client retries, queue redelivery, webhook resends, cron overlap, user double-click, load-balancer replay?
- Are payments involved? If yes, which provider(s) — Stripe, Adyen, PayPal, bank transfer?
- What is your expected concurrency? — single server, multi-instance, globally distributed?
- How long should the system remember a completed request? (key retention / expiry — typically hours to days, governs how long retries are recognized)
- How quickly must a duplicate be rejected? (detection latency — typically milliseconds, governed by DB lookup speed and lock strategy)
Adapt all output to the user's answers. Use their actual stack, database, and language in code examples.
Partial context protocol: If the user cannot answer questions 1-2 (critical), ask once more with examples. If still unknown, produce a technology-agnostic design using the PostgreSQL reference schema and note that implementation code will need adaptation. For questions 3-7, proceed with stated assumptions. Never ask the same question more than twice.
When To Use
Use this skill when you recognize these problem-symptoms:
- Users can double-click a submit button and you have no server-side guard
- Clients auto-retry on timeout and you cannot distinguish retries from new requests
- A message broker may deliver the same event more than once (at-least-once delivery)
- Webhooks from an external provider arrive multiple times for the same event
- Payments or financial mutations must never execute twice under any failure scenario
- Background jobs overlap because a previous run did not finish before the next starts
- A network partition causes the same request to hit multiple backend instances
Reference Examples
These are structural references. Adapt format, naming, and types to the user's stack.
Key Format
{scope}_{action}_{intentIdentifier}
Example (client-generated): usr_7fQ9x_createOrder_a1b2c3d4-uuid
Example (derived from intent): usr_7fQ9x_createOrder_sha256(cart_id + items_hash)
The key must be stable across retries of the same intent. Use expiration (expires_at column) for lifecycle management, not time components in the key.
Never include time-varying components in idempotency keys — a retry that crosses a clock boundary (e.g., 14:59 → 15:01) would get a different key and bypass deduplication, causing exactly the duplicate processing you're trying to prevent.
Request Fingerprint
Hash the semantically significant fields — the ones that, if changed, mean a different intent:
fingerprint = SHA-256("{amount}|{currency}|{recipient_id}")
Example: SHA-256("4999|EUR|acct_xyz") → "a3f2c8..."
If the same idempotency key arrives with a different fingerprint, reject with 409 Conflict. This catches callers reusing keys for different operations.
Request-Handling Flow (Pseudocode)
function handleRequest(idempotencyKey, requestBody):
fingerprint = computeFingerprint(requestBody)
// Step 1: Atomic insert-or-fetch (all timestamps from DB clock, never app clock)
record = atomicUpsert(
table: "idempotency_store",
key: idempotencyKey,
setIfNew: { status: "PROCESSING", fingerprint, locked_until: DB_NOW() + lockDuration }
)
// Step 2: If record already existed
if record.wasExisting:
if not constantTimeEquals(record.fingerprint, fingerprint):
// Use crypto-safe constant-time comparison (MessageDigest.isEqual in Java,
// hmac.compare_digest in Python, crypto.timingSafeEqual in Node.js)
return 409 Conflict ("key reused with different request body")
if record.status == "COMPLETED":
return record.stored_response // safe replay
if record.status == "PROCESSING" and record.locked_until > DB_NOW():
return 409 Conflict ("request already in progress")
if record.status == "PROCESSING" and record.locked_until <= DB_NOW():
// Orphaned lock — previous processor crashed. Reclaim with CAS:
// UPDATE ... SET locked_until = DB_NOW() + lockDuration, lock_version = lock_version + 1
// WHERE key = ? AND status = 'PROCESSING' AND lock_version = record.lock_version
// If affected rows = 0, another node already reclaimed — return 409
reclaimLock(record)
// Step 3: Execute the action
try:
result = executeAction(requestBody)
markCompleted(idempotencyKey, result)
return result
catch permanentError:
markFailed(idempotencyKey, error)
throw
catch transientError:
releaseLock(idempotencyKey)
throw // caller may retry with same key
Database Schema (Idempotency Store)
PostgreSQL reference implementation (for non-relational stores like MongoDB, DynamoDB, or Redis, redesign from the logical model — key, fingerprint, status, lock, response, expiry — rather than translating this DDL):
CREATE TABLE idempotency_store (
idempotency_key VARCHAR(255) NOT NULL,
user_id VARCHAR(128) NOT NULL,
fingerprint CHAR(64) NOT NULL,
status VARCHAR(20) NOT NULL DEFAULT 'PROCESSING',
response_code INT,
response_body JSONB,
lock_version INT NOT NULL DEFAULT 0,
locked_until TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
completed_at TIMESTAMPTZ,
expires_at TIMESTAMPTZ NOT NULL,
CONSTRAINT pk_idempotency PRIMARY KEY (idempotency_key, user_id),
CONSTRAINT chk_status CHECK (status IN ('PROCESSING', 'COMPLETED', 'FAILED'))
);
CREATE INDEX idx_idempotency_expires ON idempotency_store (expires_at)
WHERE status != 'PROCESSING';
-- Supports orphan-lock recovery queries
CREATE INDEX idx_idempotency_orphan_locks ON idempotency_store (locked_until)
WHERE status = 'PROCESSING';
For high-volume systems (>100K keys/day), add range partitioning on created_at and drop old partitions instead of row-by-row DELETE.
Output Specification
Produce the following sections, tailored to the user's stack and action.
1. Action Analysis
For each protected action, specify:
- What executes — the mutation, side effect, or external call
- Duplicate impact — what goes wrong if it runs twice (financial loss, data corruption, user confusion)
- Retry sources — every path that could cause repetition
- Semantically significant fields — which request fields distinguish "same intent" from "different intent"
- Failure recovery path — what happens if the server crashes mid-processing (see Partial Failure below)
2. Idempotency Key Design
Define all of these with concrete values:
- (a) Key format — with a filled-in example using the user's domain
- (b) Scope boundary — per-user, per-account, per-tenant, or global; justify the choice
- (c) Validation — max length, allowed charset, uniqueness enforcement (DB constraint DDL)
- (d) Who generates it — client, server, or derived from request content
- (e) Uniqueness assessment — state whether the key format provides sufficient uniqueness for the expected volume. For standard formats (UUIDv4, ULID), state that collision risk is negligible at any practical scale. For custom formats, identify the variable components and their cardinality, then recommend the user validate with a birthday-problem check if the space is small
3. Request Fingerprinting
- List which fields are hashed and why each is semantically significant
- Specify the hash algorithm and encoding (e.g., SHA-256, hex-encoded)
- Define the mismatch behavior (409 with descriptive error body)
- Qualify which fields are NOT included and why (e.g., timestamps, request IDs, headers)
4. Concurrency Control
Specify the locking mechanism for the user's database:
- The atomic operation that claims the key (INSERT ... ON CONFLICT, SELECT FOR UPDATE, compare-and-swap)
- Lock timeout duration and how it is chosen (must account for expected action duration + buffer; too short = false orphan detection; too long = blocked retries)
- What happens when a lock is contested (immediate 409 vs wait-and-retry)
- The DB constraint DDL that makes double-processing physically impossible
- Clock skew handling — In distributed systems,
locked_untilcomparisons across nodes may disagree. Use the DB server'sNOW()for all timestamp comparisons, never the application server's clock. If using a distributed lock (Redis/DynamoDB), account for clock drift in TTL calculations (add a skew buffer of 2-5 seconds). - The ABA problem — When reclaiming an orphaned lock, verify both
locked_until <= now()ANDstatus = 'PROCESSING'in a single atomic operation (UPDATE ... WHERE). A simple read-then-write allows two nodes to both reclaim the same lock.
5. Partial Failure Recovery
This is the hardest part. Address each scenario:
- Server crashes after claiming the key but before completing the action — How is the orphaned PROCESSING record detected? Use
locked_untiltimeout. Define the timeout value and the recovery strategy (retry, manual intervention, or abandon). - Action partially completed — e.g., payment charged but local DB not updated. Define reconciliation: scheduled job that checks provider state and reconciles, or compensating transaction.
- Distinguishing "failed permanently" from "still processing elsewhere" — Use heartbeat extension or short lock windows. Never assume a PROCESSING record is dead just because it is old.
- Cleanup of expired records — Scheduled job, partition pruning, or TTL. Specify the retention period.
6. Payment Provider Coordination
When payments are involved, the local idempotency key MUST propagate to the provider's native deduplication:
| Provider | Native mechanism | Coordination |
|---|---|---|
| Stripe | Idempotency-Key header (max 255 chars, 24h TTL) | Pass your local key as the header. If Stripe returns a cached result, store locally as COMPLETED. Local key expiry must be >= 24h. |
| Adyen | reference field (unique per merchant, 30-day dedup) | Use {idempotency_key}_{step_index} as reference. Local retention should match 30 days. |
| PayPal | PayPal-Request-Id header (UUID) | Generate deterministic UUID from your key via UUIDv5(namespace, key). |
Critical invariant: Never generate a fresh provider reference on retry — it MUST be derived deterministically from the local idempotency key so both layers deduplicate the same logical request.
7. Multi-Step (Saga) Idempotency
When the protected action is a composite (e.g., reserve inventory, charge payment, confirm order):
- Each step gets its own idempotency sub-key or is tracked in a step-status column
- Define the compensation path if a middle step fails (reverse earlier steps)
- The outer idempotency key covers the entire saga; retrying replays from the last incomplete step, not from the beginning
- Include a state machine:
STEP_1_DONE → STEP_2_DONE → COMPLETEDorSTEP_2_FAILED → COMPENSATING → ROLLED_BACK
8. Security
- Key enumeration prevention — Keys must be scoped to the authenticated user (composite PK of key + user_id). Reject requests where the key belongs to a different user. Return 404 (not 403) to avoid confirming key existence.
- Timing attack resistance — Use constant-time comparison for fingerprint matching. A timing side-channel on the fingerprint check could reveal partial hash values.
- Replay attack protection — Expired keys must not be reactivatable. Bind keys to the session or auth token that created them if the threat model requires it.
- Key guessability — If client-generated, require sufficient entropy (UUIDv4 minimum, 128 bits). If predictable patterns are used, document the tradeoff.
- Rate limiting — Excessive key creation from one user may indicate abuse. Track
idempotency.keys_created_per_userand throttle above threshold (e.g., >100 unique keys/minute). - Information leakage in 409 responses — The conflict response should confirm the key exists and explain the mismatch, but must NOT return the original request body or fingerprint. Only state that the key was used with different parameters.
- PII in stored responses — The
response_bodycolumn may contain PII. Apply encryption at rest if required by compliance (GDPR, HIPAA). Consider storing only a response hash + status code for non-critical replays, with full body only for payment flows. - Log hygiene — Never log full idempotency keys, fingerprints, or response bodies. Log only key prefixes or hashes for correlation. Idempotency keys can be used as correlation tokens by attackers if leaked.
9. Anti-Patterns to Avoid
Flag these in the design review:
| Anti-Pattern | Why It Fails |
|---|---|
| Relying on UI button disabling alone | Does not protect against network retries, API clients, or race conditions |
| Checking application state without a DB constraint | Two threads both read "not exists" and both insert |
| Using timestamps as idempotency keys | Two requests in the same millisecond collide; different-second requests for the same intent do not deduplicate |
| Storing keys indefinitely without cleanup | Table grows without bound, index performance degrades |
| Trusting the payment provider to prevent local duplicates | Provider deduplication does not prevent your DB from recording two orders |
| Using request body equality instead of a stable key | Bodies may differ in non-significant fields (timestamps, trace IDs) |
| Locking without a timeout | Crashed processes hold locks forever |
| Comparing timestamps using application clock instead of DB clock | Clock skew between app servers causes false lock reclaims |
| Storing full response bodies without encryption | PII/PCI exposure if DB is compromised; compliance violation |
| Not partitioning/archiving the idempotency table | At high volume (>1M keys/day), table bloat degrades lookup performance |
10. Monitoring and Alerting
Define metrics (not just "track duplicates" — specify the metric name, type, and alert threshold):
idempotency.duplicates_blocked(counter) — alert if rate spikes above baselineidempotency.lock_timeouts(counter) — indicates processing failures or crashesidempotency.fingerprint_mismatches(counter) — indicates client bugs or abuseidempotency.store_size(gauge) — alert if cleanup is not runningidempotency.processing_duration_ms(histogram) — detect slow actions that risk lock expiryidempotency.keys_created_per_user(counter) — detect key-flooding abuseidempotency.orphan_reclaims(counter) — indicates processing crashes; investigate if sustained
11. Test Scenarios
For each, specify input, expected outcome, and what invariant it verifies:
| Scenario | Expectation | Verifies |
|---|---|---|
| Same key, same body, first request | 200, action executes | Happy path works |
| Same key, same body, second request | 200, stored response replayed, action does NOT re-execute | Core idempotency |
| Same key, different body | 409 Conflict with descriptive error | Fingerprint mismatch detection |
| Concurrent requests with same key | Exactly one succeeds, other gets 409 or stored response | Lock correctness |
| Key from different user | 404 (per Security section — do not confirm key existence), not the other user's response | Scope isolation |
| Request arrives after key expired | Treated as new request (or rejected, per policy) | Expiration logic |
| Server crash mid-processing, then retry | Lock times out, retry succeeds | Partial failure recovery |
| Saga step 2 fails, retry the saga | Resumes from step 2, does not repeat step 1 | Multi-step idempotency |
12. Build Order
Implement in this sequence — each step is safe to deploy independently:
- Create the
idempotency_storetable with constraints (DDL provided above) - Add the idempotency middleware/interceptor — initially in log-only mode
- Protect the single highest-risk action first
- Add fingerprinting and mismatch rejection
- Add lock timeout and orphan recovery
- Add the cleanup job
- Add monitoring and alerts
- Extend to remaining actions
- Add saga tracking if multi-step actions exist
- Load-test concurrent duplicate scenarios
Constraints (Testable)
Every output must satisfy these. If any cannot be met, explain why and propose an alternative.
- Every protected action specifies its failure recovery path — no action is left with "just retry"
- The idempotency store DDL is included with the PRIMARY KEY and CHECK constraints
- Semantically significant fields for fingerprinting are explicitly listed per action
- Lock timeout value is stated with justification
- Key expiration window is stated with justification
- At least one anti-pattern relevant to the user's stack is called out
- Concurrent duplicate test case is included with expected behavior
Final Deliverables
Hand back exactly these artifacts, compiled into the single HTML or Markdown document chosen at the start (code artifacts additionally as real source files if the user wants them applied to the project):
- Action map — table of protected actions with duplicate impact, key source, and recovery path
- Idempotency store DDL — ready to run against the user's database
- Middleware/interceptor code — in the user's language and framework, implementing the lock-check-process-store flow
- Fingerprint utility — function that computes the request fingerprint for each protected action
- Cleanup job — scheduled task that prunes expired records
- Test suite skeleton — covering the scenarios in section 11, in the user's test framework
- Monitoring config — metric definitions and suggested alert thresholds
- Build order — numbered steps with deployment safety notes
- Operational runbook — covering: stuck PROCESSING records (identification query + manual release), split-brain detection (metrics spike + provider reconciliation), partition maintenance, provider/local state divergence reconciliation, key-flooding abuse response