Forensic trail for fire and forget sends
Skill Ed3Design/ed3design-skill-bundles/async-forensik/skills/forensic-trail-for-fire-and-forget-sends
Claude Code skill bundles for software engineering: 56 skills + 5 Python tools + 6 hooks + 4 sub-agents across 6 thematic plugins (token-savers, code-quality, planning-disciplines, async-forensik, schema-discipline, skill-system-meta). Empirically TDD-validated patterns, MIT licensed.
npx -y skills add Ed3Design/ed3design-skill-bundles --skill forensic-trail-for-fire-and-forget-sendsAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Use when designing or hardening any "send + log + swallow" pattern (Telegram bot sends, email notifications, webhook dispatches, Slack pings, push notifications, SMS sends) where DB-state persistence is separated from external API send, send-errors fail-open via log-only, and future forensics may need to distinguish "send happened, user missed" from "send never happened" without relying on log-aggregator availability. Trigger on phrases like "log.warning on send-fail", "fire-and-forget Telegram", "why didn't the notification arrive", "DB save ran but Telegram didn't", "container logs are gone", "*_sent_at column". Produces (1) DB migration for `*_sent_at TIMESTAMPTZ`, (2) update-after-send-OK persistence, (3) forensic query template. Do NOT load for synchronous send-and-wait patterns, sends with guaranteed-persistent external log aggregation, or one-off ad-hoc scripts.
SKILL.md
6.4 KB, as published. Nobody here has run it
Forensic trail for fire-and-forget sends
Overview
The pattern "DB save → external send → log.warning on send-fail" is nice for resilience, but on container restart or log rotation the send-fail forensics are gone. When the user asks the next morning "why didn't my notification arrive?", you have to say "I don't know, logs are toast".
Core principle: every external-send action needs a DB trail (*_sent_at) as a log-independent audit layer.
When to use
- Telegram bot sends (action trigger or notification)
- Email dispatch (transactional or marketing)
- Webhook POSTs to external services
- Slack/Discord channel pings
- SMS sends via Twilio / etc.
- Push notifications (FCM/APN)
- Alert dispatcher (any external integration)
- Async
asyncio.create_task(send_x())(mitigates the silent-error class)
When NOT to use
- Sync send-and-wait where the caller has the send result directly
- External log aggregator (Datadog/Sentry/etc.) persistently covers ALL services
- Ad-hoc scripts with short lifetime
- Sends with their own external receipt tracking (e.g. Twilio status callback into your own DB)
The 4-step procedure
Step 1 — DB migration: *_sent_at TIMESTAMPTZ
ALTER TABLE <state_table>
ADD COLUMN IF NOT EXISTS <channel>_sent_at TIMESTAMPTZ;
Idempotent + additive. Conventions:
telegram_sent_atfor Telegram sendsemail_sent_atfor mailswebhook_posted_atfor webhooksslack_sent_atetc.
For multiple channels: separate columns. For ONE send per row: dedicated column directly. For 1:n (multiple sends per row over time): a separate *_dispatches table.
Step 2 — Update logic after send-OK
async def _send_x(record_id: int, ..., conn=None):
try:
await external_api.send(...)
except Exception as exc:
log.error("Send-Fail (record_id=%s): %s", record_id, exc)
return False
# Forensic trail
if conn is not None:
try:
await conn.execute(
"UPDATE <state_table> SET <channel>_sent_at = NOW() "
"WHERE id = $1",
record_id,
)
except Exception as exc:
log.warning("sent_at update failed: %s", exc)
return True
Important: update INSIDE the send function, NOT from the caller. Otherwise the caller can forget and the trail is gone again.
Step 3 — Forensic query as a template snippet
In a code comment or a dedicated docs/forensics/send-discrepancy.md:
-- Which records had DB save but no successful send (= send-fail)?
SELECT id, <key-columns>, created_at
FROM <state_table>
WHERE <channel>_sent_at IS NULL
AND created_at < NOW() - INTERVAL '1 minute' -- send-completion margin
ORDER BY created_at DESC
LIMIT 50;
Margin: large enough for send latency, small enough not to show current in-flight sends.
Step 4 — Tests + docs
async def test_sent_at_set_on_send_ok():
# mock external send → success
# call _send_x with conn + record_id
# assert SELECT sent_at FROM table WHERE id=record_id IS NOT NULL
async def test_sent_at_null_on_send_fail():
# mock external send → raise
# call _send_x with conn + record_id
# assert returns False, sent_at IS NULL
Anti-patterns
- ❌ Update before send —
UPDATE sent_atBEFORE the actual send → mark as sent when actually failed - ❌ Update from caller — caller forgets update or race condition
- ❌ Status column instead of timestamp (
status='sent') — loses the "when" information for latency forensics - ❌ Index on
*_sent_at IS NULLwithout partial index — can become expensive on large tables. UseCREATE INDEX ... WHERE sent_at IS NULLfor the most common forensic query - ❌ Send + update in one transaction — if send takes 30s, this holds the DB connection longer than needed. Separate transactions.
Worked example
User-driven forensics: in the morning at 06:02 UTC, two advisor outputs were persisted in claude_assessments (id=140 SI=F, id=141 ZC=F), but no Telegram arrived. A container restart at 09:03 UTC erased the logs — the send-fail cause was unprovable.
Fix:
ALTER TABLE claude_assessments
ADD COLUMN IF NOT EXISTS telegram_sent_at TIMESTAMPTZ;
Plus update-after-send-OK in _send_advisor_telegram.
Cycle-2 forensic query (in code comment):
SELECT id, symbol, direction, created_at FROM claude_assessments
WHERE telegram_sent_at IS NULL
AND model='per_signal_advisor'
AND created_at < NOW() - INTERVAL '1 minute';
That way, every future send discrepancy is readable from the DB in <1s, log-independent.
Skill-Composition
superpowers:test-driven-development— for the send testslibrary-subclass-explicit-type-classification— when send errors need to be classified as retry/fail-fast (see the twin skill)- The project CLAUDE.md your-app section — document the pattern there as a backlog hint
Why this is overlooked (common cause of regret)
During implementation, one believes "log.warning is robust enough — logs are persistent". But:
- Container restart (deploy, OOM kill, health-check failure) → in-memory logs gone
- Log rotation (logrotate, k8s fluentd quotas) → older logs gone
- Log-aggregator outage → nothing gets forwarded anymore
- Production stress spike → log buffer overruns
Precisely in those cases you need forensics MORE than usual. The DB trail is 1× migration + 5 lines of code, against years of future sessions of forensic value.