Agent tester
Test agent: dry-run, unit, integration, compatibilityFrom its SKILL.md
npx -y skills add aAAaqwq/AGI-Super-Team --skill agent-testerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
3.2 KB, 828 tokens by cl100k_base, as published. Nobody here has run it
Agent Tester
Tests a built agent: dry-run, unit tests, integration, compatibility with other agents.
When to use
- After Agent Builder has finished
- "test agent X"
- "check agent compatibility"
Input
- Agent from
$AGENTS_PATH/[name]/ - Spec from
$AGENTS_PATH/specs/[name].spec.md
How to execute
Step 1: Static analysis
Check the agent code:
- File exists and runs without syntax errors
- All imports resolve
- Config file is valid
- Paths in config exist
- Credentials are accessible
- Dry-run mode is implemented
Step 2: Dry-run test
Run the agent with --dry-run:
python3 $AGENTS_PATH/[name]/[name]_agent.py --dry-run
Check:
- Agent starts without errors
- Logs are clear
- Shows what it WOULD do (without real side effects)
- Execution time is reasonable
Step 3: Unit tests
Run tests:
python3 -m pytest $AGENTS_PATH/[name]/test_[name].py -v
Minimum tests:
- Input parsing works
- Business logic is correct on test data
- Error handling works (bad input, missing files, API timeout)
- Output format is correct
Step 4: Integration test (one run on real data)
WARNING: only with human approval!
- Back up data that the agent modifies:
cp [target.csv] [target.csv.backup]
-
Run the agent once on real data
-
Check output:
- Data was written correctly
- Format matches schema.yaml
- Nothing broke
- Git commit was created (if needed)
-
If something is wrong -- rollback:
cp [target.csv.backup] [target.csv]
Step 5: Compatibility test
Check that the new agent does not conflict with existing ones:
## Compatibility Matrix
| Agent | Shared Files | Potential Conflict | Status |
|-------|-------------|-------------------|--------|
| Email Pipeline | activities.csv | Write conflict | ? |
| [other agents] | ... | ... | ? |
Specific checks:
- File locks: can two agents write to the same CSV simultaneously
- Data consistency: does the agent overwrite another agent's data
- ID generation: do IDs conflict (person_id, activity_id, etc.)
- Schedule overlap: do agents run at the same time
- Git conflicts: does auto-commit create merge conflicts
Step 6: Report
Create a test report file:
$AGENTS_PATH/specs/[name].test-report.md
Report structure:
# Test Report: [Agent Name]
## Date: YYYY-MM-DD
## Tester: Process Analyst Agent
## Results
| Test | Status | Notes |
|------|--------|-------|
| Static analysis | PASS/FAIL | |
| Dry-run | PASS/FAIL | |
| Unit tests | PASS/FAIL | X/Y passed |
| Integration | PASS/FAIL | |
| Compatibility | PASS/FAIL | |
## Issues Found
1. [Issue description + severity]
## Recommendation
- [ ] READY for production
- [ ] NEEDS FIXES (list what)
- [ ] BLOCKED (list why)
Output
- Test report in
$AGENTS_PATH/specs/[name].test-report.md - PASS/FAIL verdict
- List of issues if any
Related skills
process-analyst— creates the specagent-builder— builds the agentchange-review— validates CRM/PM changes
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most context ai engineering skills give in 828 tokens
Counted across 1,328 of the 2,349 authors here whose files we hold, read 2026-09-06
- Dispatch a fresh subagent for each taskin 76 of 1328, across 59 files
- Perform spec compliance review before code quality reviewin 44 of 1328, across 34 files
- Dispatch a final code reviewer after all tasksin 38 of 1328, across 26 files
- Answer subagent questions before allowing implementationin 36 of 1328, across 26 files
- Use the least powerful model capable of the taskin 33 of 1328, across 26 files
- Create a TodoWrite list for all tasksin 32 of 1328, across 22 files
- Perform a task review after each implementationin 31 of 1328, across 24 files
- Extract all tasks and context from the planin 29 of 1328, across 20 files
- Provide full task text to subagentsin 28 of 1328, across 20 files
- Use git worktrees for isolated workspacesin 25 of 1328, across 20 files
- Specify the model explicitly when dispatching a subagentin 23 of 1328, across 18 files
- Execute all tasks from the plan without stoppingin 21 of 1328, across 16 files
Said here and by no other author read
- Run agent with dry-run flag
- Execute unit tests using pytest
- Backup data before integration testing
- Run integration test only with human approval
- Check for compatibility with existing agents
- Create a test report file
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.