agentsclimarketplace

Forensic trail for fire and forget sends

Skill Ed3Design/ed3design-skill-bundles/async-forensik/skills/forensic-trail-for-fire-and-forget-sends

Claude Code skill bundles for software engineering: 56 skills + 5 Python tools + 6 hooks + 4 sub-agents across 6 thematic plugins (token-savers, code-quality, planning-disciplines, async-forensik, schema-discipline, skill-system-meta). Empirically TDD-validated patterns, MIT licensed.

Install
npx -y skills add Ed3Design/ed3design-skill-bundles --skill forensic-trail-for-fire-and-forget-sends

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when designing or hardening any "send + log + swallow" pattern (Telegram bot sends, email notifications, webhook dispatches, Slack pings, push notifications, SMS sends) where DB-state persistence is separated from external API send, send-errors fail-open via log-only, and future forensics may need to distinguish "send happened, user missed" from "send never happened" without relying on log-aggregator availability. Trigger on phrases like "log.warning on send-fail", "fire-and-forget Telegram", "why didn't the notification arrive", "DB save ran but Telegram didn't", "container logs are gone", "*_sent_at column". Produces (1) DB migration for `*_sent_at TIMESTAMPTZ`, (2) update-after-send-OK persistence, (3) forensic query template. Do NOT load for synchronous send-and-wait patterns, sends with guaranteed-persistent external log aggregation, or one-off ad-hoc scripts.

SKILL.md

6.4 KB, as published. Nobody here has run it

Forensic trail for fire-and-forget sends

Overview

The pattern "DB save → external send → log.warning on send-fail" is nice for resilience, but on container restart or log rotation the send-fail forensics are gone. When the user asks the next morning "why didn't my notification arrive?", you have to say "I don't know, logs are toast".

Core principle: every external-send action needs a DB trail (*_sent_at) as a log-independent audit layer.

When to use

  • Telegram bot sends (action trigger or notification)
  • Email dispatch (transactional or marketing)
  • Webhook POSTs to external services
  • Slack/Discord channel pings
  • SMS sends via Twilio / etc.
  • Push notifications (FCM/APN)
  • Alert dispatcher (any external integration)
  • Async asyncio.create_task(send_x()) (mitigates the silent-error class)

When NOT to use

  • Sync send-and-wait where the caller has the send result directly
  • External log aggregator (Datadog/Sentry/etc.) persistently covers ALL services
  • Ad-hoc scripts with short lifetime
  • Sends with their own external receipt tracking (e.g. Twilio status callback into your own DB)

The 4-step procedure

Step 1 — DB migration: *_sent_at TIMESTAMPTZ

ALTER TABLE <state_table>
ADD COLUMN IF NOT EXISTS <channel>_sent_at TIMESTAMPTZ;

Idempotent + additive. Conventions:

  • telegram_sent_at for Telegram sends
  • email_sent_at for mails
  • webhook_posted_at for webhooks
  • slack_sent_at etc.

For multiple channels: separate columns. For ONE send per row: dedicated column directly. For 1:n (multiple sends per row over time): a separate *_dispatches table.

Step 2 — Update logic after send-OK

async def _send_x(record_id: int, ..., conn=None):
    try:
        await external_api.send(...)
    except Exception as exc:
        log.error("Send-Fail (record_id=%s): %s", record_id, exc)
        return False

    # Forensic trail
    if conn is not None:
        try:
            await conn.execute(
                "UPDATE <state_table> SET <channel>_sent_at = NOW() "
                "WHERE id = $1",
                record_id,
            )
        except Exception as exc:
            log.warning("sent_at update failed: %s", exc)

    return True

Important: update INSIDE the send function, NOT from the caller. Otherwise the caller can forget and the trail is gone again.

Step 3 — Forensic query as a template snippet

In a code comment or a dedicated docs/forensics/send-discrepancy.md:

-- Which records had DB save but no successful send (= send-fail)?
SELECT id, <key-columns>, created_at
FROM <state_table>
WHERE <channel>_sent_at IS NULL
  AND created_at < NOW() - INTERVAL '1 minute'  -- send-completion margin
ORDER BY created_at DESC
LIMIT 50;

Margin: large enough for send latency, small enough not to show current in-flight sends.

Step 4 — Tests + docs

async def test_sent_at_set_on_send_ok():
    # mock external send → success
    # call _send_x with conn + record_id
    # assert SELECT sent_at FROM table WHERE id=record_id IS NOT NULL

async def test_sent_at_null_on_send_fail():
    # mock external send → raise
    # call _send_x with conn + record_id
    # assert returns False, sent_at IS NULL

Anti-patterns

  • Update before sendUPDATE sent_at BEFORE the actual send → mark as sent when actually failed
  • Update from caller — caller forgets update or race condition
  • Status column instead of timestamp (status='sent') — loses the "when" information for latency forensics
  • Index on *_sent_at IS NULL without partial index — can become expensive on large tables. Use CREATE INDEX ... WHERE sent_at IS NULL for the most common forensic query
  • Send + update in one transaction — if send takes 30s, this holds the DB connection longer than needed. Separate transactions.

Worked example

User-driven forensics: in the morning at 06:02 UTC, two advisor outputs were persisted in claude_assessments (id=140 SI=F, id=141 ZC=F), but no Telegram arrived. A container restart at 09:03 UTC erased the logs — the send-fail cause was unprovable.

Fix:

ALTER TABLE claude_assessments
ADD COLUMN IF NOT EXISTS telegram_sent_at TIMESTAMPTZ;

Plus update-after-send-OK in _send_advisor_telegram.

Cycle-2 forensic query (in code comment):

SELECT id, symbol, direction, created_at FROM claude_assessments
WHERE telegram_sent_at IS NULL
  AND model='per_signal_advisor'
  AND created_at < NOW() - INTERVAL '1 minute';

That way, every future send discrepancy is readable from the DB in <1s, log-independent.

Skill-Composition

  • superpowers:test-driven-development — for the send tests
  • library-subclass-explicit-type-classification — when send errors need to be classified as retry/fail-fast (see the twin skill)
  • The project CLAUDE.md your-app section — document the pattern there as a backlog hint

Why this is overlooked (common cause of regret)

During implementation, one believes "log.warning is robust enough — logs are persistent". But:

  • Container restart (deploy, OOM kill, health-check failure) → in-memory logs gone
  • Log rotation (logrotate, k8s fluentd quotas) → older logs gone
  • Log-aggregator outage → nothing gets forwarded anymore
  • Production stress spike → log buffer overruns

Precisely in those cases you need forensics MORE than usual. The DB trail is 1× migration + 5 lines of code, against years of future sessions of forensic value.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.