Ci fix with memory
Automatically invoke when CI checks fail on a PR. Triggers on: any mention of CI failure, red checks, failed GitHub Actions, or "CI is broken/failing". Reads session handoffs and a known-issues log to avoid repeating failed fix attempts. Do NOT wait for user to ask — invoke as soon as a CI failure is confirmed. Not for local-only test failures (use test-fix instead) or lint-only failures with an obvious single-line fix.From its SKILL.md
npx -y skills add jerseycheese/agent-skills --skill ci-fix-with-memoryAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
3 things to look at
- skips confirmationTells the agent to proceed without asking first, 2 times: "Do NOT wait for user to ask — invoke as soon as a CI failure is confirmed" and 1 more.
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 8 commands, including `cat .handoff/latest.md` and 7 more.
SKILL.md
9.5 KB, ~2.4k tokens by cl100k_base, as published. Nobody here has run it
When to invoke (auto-trigger)
Invoke automatically whenever:
- The user mentions CI is failing, red, or blocked
gh pr checksshows a failure- A push triggers a failed workflow run
Do NOT invoke for: local test failures (use test-fix instead), lint-only failures with an obvious single-line fix.
Self-Healing CI Pipeline with Handoff Memory
When to Use This Skill
Use this skill when:
- CI checks are failing on a PR
- You want to avoid repeating previously failed fix attempts
- You need systematic CI debugging with context preservation
- Multiple CI check types are failing (lint, tests, build, etc.)
Critical Setup
This skill requires a .handoff/ directory structure:
.handoff/
├── latest.md # Most recent session handoff
└── known-issues.md # Persistent log of failed approaches
If these don't exist, the skill will create them on first run.
Pipeline Steps
Phase 1: Ingest Context (Before ANY Action)
-
Read the handoff file
cat .handoff/latest.mdThis contains:
- What was tried in the previous session
- What worked and what didn't
- Current branch state
- Open questions or blockers
-
Read the known issues log
cat .handoff/known-issues.mdThis contains:
- Approaches that have been tried and failed
- Specific commands or fixes that don't work
- Patterns to avoid
- Environmental quirks
-
Acknowledge context Present a brief summary:
## Context Loaded ### From Previous Session [Key points from latest.md] ### Known Failed Approaches [List from known-issues.md] ### Starting Fresh Investigation
Phase 2: Get Current CI Status
-
Fetch current CI check status
gh pr checks [PR_NUMBER] -
For each failing check, get full logs
gh pr checks [PR_NUMBER] --watch # OR gh api repos/{owner}/{repo}/commits/{sha}/check-runs -
Create TodoWrite checklist
- [ ] Fix lint failures - [ ] Fix TypeScript errors - [ ] Fix test failures - [ ] Fix build errors
Phase 3: Cross-Reference Against Known Issues
For each failure:
-
Check if this exact error was seen before
- Search known-issues.md for error message
- Search latest.md for previous attempts
-
If found in known-issues.md
⚠️ Known Issue Detected Failure: [error message] Previously Tried: [approach from known-issues.md] Result: Failed Skipping this approach. Trying alternative: [different approach] -
If NOT found
New failure detected: [error message] No prior attempts recorded. Proceeding with standard fix approach.
Phase 4: Fix Failures (Priority Order)
Fix in this priority order:
- Build failures first (they may cascade to other failures)
- Lint failures second (fix all in one pass)
- Test failures third (individual tests, one at a time)
- Type errors fourth (often caused by other fixes)
For Each Failure Type:
Lint Failures
# Get all lint errors
npm run lint 2>&1 | tee lint-errors.txt
# Apply fixes
npm run lint:fix
npm run lint:css:fix
# Verify
npm run lint && npm run lint:css
Update checklist: Mark lint item as done if passed
Test Failures
# Get all failing tests
npm test -- --listTests --json 2>&1 | tee test-failures.txt
# Fix one test at a time
npm test -- [specific-test-file]
# Apply 3-attempt limit per test (use test-fix principles)
Update checklist: Mark each test as done when fixed
Build Failures
# Get build errors
npm run build 2>&1 | tee build-errors.txt
# Fix errors
[apply fixes]
# Verify
npm run build
Update checklist: Mark build as done if passed
Phase 5: Dev Server Protection
Critical checks before running tests:
# Check if dev server is running
DEV_SERVER_PID=$(lsof -ti:3000)
if [ -n "$DEV_SERVER_PID" ]; then
echo "✅ Dev server running (PID: $DEV_SERVER_PID)"
echo "Tests will use existing server"
else
echo "⚠️ No dev server found"
echo "Ask user: Should I start dev server?"
fi
Never:
- Kill the dev server
- Restart the dev server
- Run commands that interfere with port 3000
Phase 6: Commit Safety
Before committing, verify:
# Check for merge conflict markers
git diff --check
# Check for common mistakes
grep -r "<<<<<<< HEAD" src/ && echo "❌ Merge conflicts found" || echo "✅ No merge conflicts"
grep -r "console.log" src/ && echo "⚠️ Debug code found" || echo "✅ No debug code"
grep -r "debugger" src/ && echo "⚠️ Debugger statements found" || echo "✅ No debugger"
If all checks pass:
git add -u
git commit -m "fix: resolve CI failures
[List of what was fixed]
Applied via ci-fix-with-memory skill"
git push origin [BRANCH]
Phase 7: Update Memory Files
-
Update .handoff/latest.md
# Session Handoff: [DATE] ## PR Context - PR #[NUMBER]: [TITLE] - Branch: [BRANCH_NAME] - Status: [CI status after fixes] ## What Was Done - [List of fixes applied] - [Commands run] - [Tests that passed/failed] ## What Worked - [Successful approaches] ## What Didn't Work - [Failed approaches - add to known-issues.md] ## Current State - Lint: [PASS/FAIL] - Tests: [PASS/FAIL] - Build: [PASS/FAIL] - Type Check: [PASS/FAIL] ## Open Questions - [Any blockers or unclear issues] ## Next Steps - [What should happen next] -
Update .handoff/known-issues.md
# Known Issues Log ## [DATE] - [Issue Category] **Error**: [exact error message] **Attempted Fix**: [what was tried] **Result**: Failed **Why It Failed**: [explanation] **Alternative Approach**: [what to try instead] --- [Previous entries...]
Phase 8: Report Status
Present final summary:
## CI Fix Pipeline Results
### Context Used
- Loaded previous session handoff
- Cross-referenced [N] known failed approaches
- Avoided repeating [specific approaches]
### Fixes Applied
✅ [List of successful fixes]
❌ [List of failed fixes with reasons]
⏭️ [List of skipped issues]
### Final CI Status
- **Lint**: [PASS/FAIL]
- **Type Check**: [PASS/FAIL]
- **Tests**: [PASS/FAIL - with count]
- **Build**: [PASS/FAIL]
### Memory Updated
- Updated .handoff/latest.md with session details
- Added [N] new entries to .handoff/known-issues.md
### Next Steps
[What user should do next]
### Handoff for Next Session
[Key context to carry forward]
Example Known Issues Log Entry
## 2025-01-15 - E2E Test Timeout
**Error**: `Timeout waiting for element: [data-testid="submit-button"]`
**Attempted Fix**: Added `page.waitForSelector('[data-testid="submit-button"]', { timeout: 10000 })`
**Result**: Failed - timeout still occurred
**Why It Failed**: Button exists but is covered by a loading overlay during test
**Alternative Approach**: Wait for overlay to disappear before clicking
**Reference**: See `.handoff/latest.md` from 2025-01-15 session
---
## 2025-01-12 - Lint Rule Conflict
**Error**: `Unexpected var, use let or const instead (no-var)`
**Attempted Fix**: Ran `eslint --fix` to auto-convert var to let
**Result**: Failed - introduced scope bugs in loop
**Why It Failed**: var has function scope, let has block scope - not always safe to auto-convert
**Alternative Approach**: Manually review each var and decide let vs const based on reassignment
---
Example Handoff Entry
# Session Handoff: 2025-01-15 14:30
## PR Context
- PR #89: Fix async form submission race condition
- Branch: `fix/form-submit-timing`
- Status: CI failing (2 of 4 checks)
## What Was Done
- Fixed lint errors in `SubmitButton.tsx` and `FormContext.tsx`
- Attempted to fix E2E test `form-submit.spec.ts` (3 attempts, failed)
- Updated TypeScript types for form state
## What Worked
- Lint fixes: Changed `var` to `const` manually (not auto-fix)
- Type fixes: Added proper typing for `FormState` interface
## What Didn't Work
- E2E test still timing out on submit button click
- Root cause: Loading overlay covers button before it becomes interactive
- Tried: Waiting for selector (failed), adding explicit delay (failed), clicking with force (failed)
## Current State
- Lint: PASS
- Tests: FAIL (1 test: form-submit.spec.ts)
- Build: PASS
- Type Check: PASS
## Open Questions
- Should we dismiss the overlay programmatically in test setup?
- Is this a timing issue in the component or the test?
## Next Steps
- May need browser DevTools inspection to confirm overlay z-index
- Consider checking if overlay has a data-testid we can wait on
Success Metrics
Track these outcomes:
- Number of repeat fix attempts avoided
- Time saved by checking known-issues log first
- How many handoffs actually helped next session
- Patterns discovered through known-issues tracking
Notes
- Addresses context loss between sessions (cached state, unknown file locations)
- Prevents "merge-conflict-markers-left-in-files" class of errors
- Known-issues log becomes useful project documentation over time
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most memory context skills give in ~2.4k tokens
Counted across 754 of the 1,056 authors here whose files we hold, read 2026-09-06
- Preserve existing content structurein 15 of 754, across 9 files
- Front-load the leading wordin 14 of 754, across 10 files
- Update existing entries instead of duplicatingin 14 of 754, across 7 files
- Keep CLAUDE.md under one hundred linesin 14 of 754, across 12 files
- Read CLAUDE.md at the project rootin 14 of 754
- Keep each meaning in a single source of truthin 12 of 754, across 8 files
- Redact sensitive information before committingin 11 of 754, across 4 files
- Scan for all CLAUDE.md filesin 11 of 754, across 7 files
- Use frontmatter for metadata on filesin 10 of 754, across 3 files
- Repeat user interactions 10 timesin 10 of 754, across 4 files
- Write the CLAUDE.md file into the target folderin 10 of 754, across 8 files
- Use memlab to process snapshotsin 9 of 754, across 3 files
Said here and by no other author read
- Read the handoff file
- Read the known issues log
- Fetch current CI check status
- Cross-reference against known issues
- Fix build failures first
- Fix lint failures second
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.