Asyncio fire and forget loop exit await
Skill Ed3Design/ed3design-skill-bundles/async-forensik/skills/asyncio-fire-and-forget-loop-exit-await
Use when writing or reviewing Python code that uses `asyncio.run(...)` as entry-point AND inside the entry-coroutine uses `asyncio.create_task(...)` for fire-and-forget background work (notifications, telemetry, advisor-calls, fan-out HTTP/DB-writes). Without an explicit `await asyncio.gather(*pending_tasks)` before the entry-coroutine returns, `asyncio.run()` closes the event-loop and CANCELS still-pending tasks — silent data-loss of 0-50% of background work. Trigger on phrases like "asyncio.run + create_task", "fire-and-forget cancelled at loop-exit", "telemetry silent dropping", "Loop closed before task finished". Do NOT load for asyncio.Runner / gather-as-entry patterns, long-running daemons (uvicorn, asyncio.Event.wait), sync code with threading.Thread, or non-Python async (JS/Go).From its SKILL.md
npx -y skills add Ed3Design/ed3design-skill-bundles --skill asyncio-fire-and-forget-loop-exit-awaitAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
9.4 KB, ~2.1k tokens by cl100k_base, as published. Nobody here has run it
asyncio fire-and-forget loop-exit-await pattern
The Iron Law
If you use asyncio.run(main()) as entry-point AND asyncio.create_task(...) for background work inside main(), main() MUST await all pending tasks before returning — otherwise asyncio.run() closes the loop and silently cancels them.
When the Bug Bites
The bug is silent and partial:
- Tasks that complete BEFORE
main()returns → fine - Tasks still pending when
main()returns → cancelled, no error, no log (unless you handle CancelledError)
Typical victim pattern: main() does N iterations, each iteration dispatches a quick fire-and-forget task. The Nth iteration's task starts a 10s Claude-API-call, then main() returns immediately → that last task is killed mid-flight. Earlier tasks may also be partially incomplete.
In the real-world case: a Per-Signal-Advisor calling the Claude API (5-30s) was dispatched after each signal in an orchestrator loop that completed in ~2-5s. Estimated 0-50% silent task-cancellation depending on signal-count and API-latency.
The 3-Step Fix
async def main():
tasks: list[asyncio.Task] = []
for item in items:
# ... existing work ...
# Step 1: TRACK the task in a local list (not just create_task and forget)
tasks.append(asyncio.create_task(background_work(item)))
# Step 2: Before returning, await ALL pending tasks with a sensible timeout
if tasks:
n_pending = sum(1 for t in tasks if not t.done())
if n_pending:
try:
results = await asyncio.wait_for(
asyncio.gather(*tasks, return_exceptions=True),
timeout=60.0, # tune to your slowest expected task
)
# Step 3: Inspect results, log errors, NEVER re-raise
n_errors = sum(1 for r in results if isinstance(r, BaseException))
if n_errors:
logging.warning("%d/%d background tasks raised", n_errors, len(tasks))
except asyncio.TimeoutError:
n_still_pending = sum(1 for t in tasks if not t.done())
logging.warning("%d tasks exceeded %ds — cancelled", n_still_pending, 60)
return 0 # NOW it's safe to return; asyncio.run() can close loop cleanly
if __name__ == "__main__":
sys.exit(asyncio.run(main()))
Why return_exceptions=True: without it, the first exception in any task cancels all sibling tasks and propagates up — usually NOT what fire-and-forget callers want. With it, every task either completes-result or completes-exception, and you log/count without surprise re-raises.
Why a timeout: protects against one hung task blocking shutdown indefinitely. Pick a value greater than your slowest expected background-call (Claude API: 30s sane upper bound, so use 60s).
Detection Checklist (for code-review)
When reviewing a Python file with asyncio.run(...) as entry-point:
- Grep for
asyncio.create_taskinside the entry-coroutine or its callees. - For each
create_task: trace whether the task-handle is awaited before the outermost coroutine returns. If the handle is discarded (anonymous expression) → 🔴 Critical. - Check: is the dispatched coroutine longer-running than the dispatcher? (Common offender: API-calls, DB-writes-with-network, slow file-IO.) If yes → 🔴 Critical for silent data-loss.
- Check: is there an existing
asyncio.gather(...)or similar at the bottom of the entry-coroutine that catches pending tasks? If yes → ✓.
Anti-Patterns
- ❌
asyncio.create_task(f())with no name, no list, no await — pure fire-and-forget. Bug-guaranteed inasyncio.run()contexts. - ❌ Trying to await INSIDE the loop (
await asyncio.create_task(...)) — that defeats the parallelism purpose. - ❌ Using
asyncio.ensure_futureinstead ofcreate_task— same problem, slightly less obvious. - ❌ Catching
CancelledErrorinside the background-coroutine and treating it as "normal" — masks the bug, doesn't fix it. - ❌ "Add a
time.sleep(5)at end of main to give tasks time" — fragile, doesn't scale, looks bad in code-review.
When You Don't Need This
- Long-running daemons (FastAPI/Uvicorn, asyncio-based servers,
asyncio.Event().wait()loops): the loop runs until externally stopped — pending tasks have time to complete, or you have explicit shutdown-hooks. No bug. - Coroutines using
asyncio.gatheras their main fan-out:gather()itself awaits, so no orphan tasks. - Sync-only code: no event-loop, no issue.
Related
- Python docs: asyncio.run — note the "Closes the loop and finalizes asynchronous generators" line, which is exactly what cancels pending tasks.
- Sibling skill:
code-review-chunk-dispatch(for catching this class of bug via subagent-review). - The maxim "Code-Review as standard": this bug was caught Day 2 in a row by the subagent-review pattern.
Edge-Cases & Aggravators
- Sync-IO in async context: if the symbol loop makes blocking calls (yfinance without await, blocking psycopg2 cursor), it blocks the event loop and prevents parallel Advisor-tasks from making progress — the silent-cancel rate rises. Check during review: are all IO calls async? Otherwise the bug escalates further.
- Extended-Thinking-Timeouts: the default
timeout=60sis sized for Claude Sonnet without Extended-Thinking. Withthinking_enabled=Trueand a high token budget, calls can take 60-180s — raise the timeout or rethink the pattern architecture (background worker instead of fire-and-forget). - uvicorn reload mode: in dev with
--reload, uvicorn periodically restarts — pending tasks are cancelled on each reload. In prod (no reload) this is not the problem, but dev tests can mask the bug if they straddle a reload cycle.
Background: Real-World Case
A Per-Signal-Advisor in scripts/v3_orchestrator.py of your-app:
_orchestrate(args)loops over symbols, each iteration: scan_signal + persist + send_telegram +asyncio.create_task(advisor_consult(signal))- After the loop: stage-monitor + PnL-update (~2s) →
return 0 asyncio.run(_orchestrate(args))closes the loop → all create_task'd Advisor-calls cancelled- Advisor-calls take 5-30s (Claude Sonnet API)
- Silent data-loss estimated 0-50% of Advisor verdicts (depending on signal-count and order)
Fix: tasks tracked in advisor_tasks: list[asyncio.Task], await asyncio.wait_for(asyncio.gather(*advisor_tasks, return_exceptions=True), timeout=60.0) before return 0. Caught by code-review subagent in a Spec→Plan→Implementation→Review cycle before any deploy.
Background: TDD Log (Bulletproofing)
Cycle 1 — PASS via Subagent-Pair-Dispatch (Tempo-Booster class)
-
RED subagent (without skill, prompt: review a v3_orchestrator snippet with
create_task(advisor_consult)+ asyncio.run): found the critical bug on its own! Wrote a precise fix withawait asyncio.gather(*advisor_tasks, return_exceptions=True)— nearly identical to the skill's 3-step fix. RED was surprisingly smart; the skill here is not a "bug preventer" but a "tempo booster + consistency guarantee". -
GREEN subagent (with skill, identical prompt): structured output, impact quote 0-50% explicitly computed,
wait_for + timeoutadded,return_exceptions=Truerationale clean. Additionally discovered an important issue "sync-IO in async context as an aggravator" — RED without the skill would not have included that in the review. -
Refactor applied: section "Edge-Cases & Aggravators" added (sync-IO, Extended-Thinking-Timeouts, uvicorn reload mode) — from GREEN self-reflection. Improves the depth of the detection checklist.
Polish-vs-Promote verdict
Comparable to pytest-venv-first-triage: smart-RED finds the bug, the skill provides tempo + consistency + avoidance of alternative anti-patterns (time.sleep, ensure_future, CancelledError-catch). Promote argument: in reviews with cognitive load (multi-bug file) consistency is more valuable than per-bug hit rate.
Cycle-2 Backlog (Polish, non-blocking)
- AnyIO variant: same pattern for
anyio.run(...)+anyio.create_task_group()— anyio enforces structured concurrency but pattern drift in codebases is possible - TaskGroup PEP-654: Python 3.11+
asyncio.TaskGroupas the idiomatic replacement for manualgather— once widely adopted, add a dedicated "Modern Alternative" section - Telemetry integration: logging hook for the
n_errors/n_pendingcounters into a central telemetry system (instead of only log.warning) — application-specific
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.