Benchmark suite manager
Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration
npx -y skills add a5c-ai/babysitter --skill benchmark-suite-managerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
Manage benchmarks for algorithm engineering experiments and evaluations
SKILL.md
1.4 KB, as published. Nobody here has run it
Benchmark Suite Manager
Purpose
Provides expert guidance on managing benchmark suites for algorithm engineering and experimental evaluation.
Capabilities
- Standard benchmark suite access (DIMACS, TSPLIB, etc.)
- Instance generation for specific problem classes
- Statistical analysis of results
- Performance comparison tables
- Visualization of scaling behavior
- Reproducibility support
Usage Guidelines
- Suite Selection: Choose appropriate benchmark suite
- Instance Selection: Select representative instances
- Execution: Run experiments systematically
- Analysis: Perform statistical analysis
- Reporting: Generate comparison tables and plots
Tools/Libraries
- DIMACS
- TSPLIB
- SuiteSparse Matrix Collection
- Statistical tools