Forensic trail for fire and forget sends
Skill Ed3Design/ed3design-skill-bundles/async-forensik/skills/forensic-trail-for-fire-and-forget-sends
Use when designing or hardening any "send + log + swallow" pattern (Telegram bot sends, email notifications, webhook dispatches, Slack pings, push notifications, SMS sends) where DB-state persistence is separated from external API send, send-errors fail-open via log-only, and future forensics may need to distinguish "send happened, user missed" from "send never happened" without relying on log-aggregator availability. Trigger on phrases like "log.warning on send-fail", "fire-and-forget Telegram", "why didn't the notification arrive", "DB save ran but Telegram didn't", "container logs are gone", "*_sent_at column". Produces (1) DB migration for `*_sent_at TIMESTAMPTZ`, (2) update-after-send-OK persistence, (3) forensic query template. Do NOT load for synchronous send-and-wait patterns, sends with guaranteed-persistent external log aggregation, or one-off ad-hoc scripts.From its SKILL.md
npx -y skills add Ed3Design/ed3design-skill-bundles --skill forensic-trail-for-fire-and-forget-sendsAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
6.4 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it
Forensic trail for fire-and-forget sends
Overview
The pattern "DB save → external send → log.warning on send-fail" is nice for resilience, but on container restart or log rotation the send-fail forensics are gone. When the user asks the next morning "why didn't my notification arrive?", you have to say "I don't know, logs are toast".
Core principle: every external-send action needs a DB trail (*_sent_at) as a log-independent audit layer.
When to use
- Telegram bot sends (action trigger or notification)
- Email dispatch (transactional or marketing)
- Webhook POSTs to external services
- Slack/Discord channel pings
- SMS sends via Twilio / etc.
- Push notifications (FCM/APN)
- Alert dispatcher (any external integration)
- Async
asyncio.create_task(send_x())(mitigates the silent-error class)
When NOT to use
- Sync send-and-wait where the caller has the send result directly
- External log aggregator (Datadog/Sentry/etc.) persistently covers ALL services
- Ad-hoc scripts with short lifetime
- Sends with their own external receipt tracking (e.g. Twilio status callback into your own DB)
The 4-step procedure
Step 1 — DB migration: *_sent_at TIMESTAMPTZ
ALTER TABLE <state_table>
ADD COLUMN IF NOT EXISTS <channel>_sent_at TIMESTAMPTZ;
Idempotent + additive. Conventions:
telegram_sent_atfor Telegram sendsemail_sent_atfor mailswebhook_posted_atfor webhooksslack_sent_atetc.
For multiple channels: separate columns. For ONE send per row: dedicated column directly. For 1:n (multiple sends per row over time): a separate *_dispatches table.
Step 2 — Update logic after send-OK
async def _send_x(record_id: int, ..., conn=None):
try:
await external_api.send(...)
except Exception as exc:
log.error("Send-Fail (record_id=%s): %s", record_id, exc)
return False
# Forensic trail
if conn is not None:
try:
await conn.execute(
"UPDATE <state_table> SET <channel>_sent_at = NOW() "
"WHERE id = $1",
record_id,
)
except Exception as exc:
log.warning("sent_at update failed: %s", exc)
return True
Important: update INSIDE the send function, NOT from the caller. Otherwise the caller can forget and the trail is gone again.
Step 3 — Forensic query as a template snippet
In a code comment or a dedicated docs/forensics/send-discrepancy.md:
-- Which records had DB save but no successful send (= send-fail)?
SELECT id, <key-columns>, created_at
FROM <state_table>
WHERE <channel>_sent_at IS NULL
AND created_at < NOW() - INTERVAL '1 minute' -- send-completion margin
ORDER BY created_at DESC
LIMIT 50;
Margin: large enough for send latency, small enough not to show current in-flight sends.
Step 4 — Tests + docs
async def test_sent_at_set_on_send_ok():
# mock external send → success
# call _send_x with conn + record_id
# assert SELECT sent_at FROM table WHERE id=record_id IS NOT NULL
async def test_sent_at_null_on_send_fail():
# mock external send → raise
# call _send_x with conn + record_id
# assert returns False, sent_at IS NULL
Anti-patterns
- ❌ Update before send —
UPDATE sent_atBEFORE the actual send → mark as sent when actually failed - ❌ Update from caller — caller forgets update or race condition
- ❌ Status column instead of timestamp (
status='sent') — loses the "when" information for latency forensics - ❌ Index on
*_sent_at IS NULLwithout partial index — can become expensive on large tables. UseCREATE INDEX ... WHERE sent_at IS NULLfor the most common forensic query - ❌ Send + update in one transaction — if send takes 30s, this holds the DB connection longer than needed. Separate transactions.
Worked example
User-driven forensics: in the morning at 06:02 UTC, two advisor outputs were persisted in claude_assessments (id=140 SI=F, id=141 ZC=F), but no Telegram arrived. A container restart at 09:03 UTC erased the logs — the send-fail cause was unprovable.
Fix:
ALTER TABLE claude_assessments
ADD COLUMN IF NOT EXISTS telegram_sent_at TIMESTAMPTZ;
Plus update-after-send-OK in _send_advisor_telegram.
Cycle-2 forensic query (in code comment):
SELECT id, symbol, direction, created_at FROM claude_assessments
WHERE telegram_sent_at IS NULL
AND model='per_signal_advisor'
AND created_at < NOW() - INTERVAL '1 minute';
That way, every future send discrepancy is readable from the DB in <1s, log-independent.
Skill-Composition
superpowers:test-driven-development— for the send testslibrary-subclass-explicit-type-classification— when send errors need to be classified as retry/fail-fast (see the twin skill)- The project CLAUDE.md your-app section — document the pattern there as a backlog hint
Why this is overlooked (common cause of regret)
During implementation, one believes "log.warning is robust enough — logs are persistent". But:
- Container restart (deploy, OOM kill, health-check failure) → in-memory logs gone
- Log rotation (logrotate, k8s fluentd quotas) → older logs gone
- Log-aggregator outage → nothing gets forwarded anymore
- Production stress spike → log buffer overruns
Precisely in those cases you need forensics MORE than usual. The DB trail is 1× migration + 5 lines of code, against years of future sessions of forensic value.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.