Simulation validator
Skill HeshamFS/materials-simulation-skills/skills/simulation-workflow/simulation-validator
Validate simulations across three stages — run pre-flight checks on configuration files (parameter ranges, required fields, disk space), monitor runtime logs for residual growth, NaN/Inf, and adaptive dt collapse, and perform post-flight validation of results (physical bounds, mass/energy conservation, convergence). Diagnose failed simulations with probable-cause analysis and recommended fixes. Use when preparing to launch a simulation, checking whether a running job is healthy, verifying that finished results are trustworthy, or debugging a crash or blow-up, even if the user only says "my simulation crashed" or "can I trust these results."From its SKILL.md
npx -y skills add HeshamFS/materials-simulation-skills --skill simulation-validatorAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- runs commandsInstructs the agent to run 4 commands, including `scripts/preflight_checker.py --config simulation.json` and 3 more.
SKILL.md
15.9 KB, ~3.6k tokens by cl100k_base, as published. Nobody here has run it
Simulation Validator
Goal
Provide a three-stage validation protocol: pre-flight checks, runtime monitoring, and post-flight validation for materials simulations.
Requirements
- Python 3.10+
- No external dependencies (uses Python standard library only)
- Works on Linux, macOS, and Windows
Inputs to Gather
Before running validation scripts, collect from the user:
| Input | Description | Example |
|---|---|---|
| Config file | Simulation configuration (JSON/YAML) | simulation.json |
| Log file | Runtime output log | simulation.log |
| Metrics file | Post-run metrics (JSON) | results.json |
| Required params | Parameters that must exist | dt,dx,kappa |
| Valid ranges | Parameter bounds | dt:1e-6:1e-2 |
Decision Guidance
When to Run Each Stage
Is simulation about to start?
├── YES → Run Stage 1: preflight_checker.py
│ └── BLOCK status? → Fix issues, do NOT run simulation
│ └── WARN status? → Review warnings, document if accepted
│ └── PASS status? → Proceed to run simulation
│
Is simulation running?
├── YES → Run Stage 2: runtime_monitor.py (periodically)
│ └── Alerts? → Consider stopping, check parameters
│
Has simulation finished?
├── YES → Run Stage 3: result_validator.py
│ └── Failed checks? → Do NOT use results
│ → Run failure_diagnoser.py
│ └── All passed? → Results are valid
Choosing Validation Thresholds
| Metric | Conservative | Standard | Relaxed |
|---|---|---|---|
| Mass tolerance | 1e-6 | 1e-3 | 1e-2 |
| Residual growth | 2x | 10x | 100x |
| dt reduction | 10x | 100x | 1000x |
Script Outputs (JSON Fields)
| Script | Output Fields |
|---|---|
scripts/preflight_checker.py | report.status, report.blockers, report.warnings |
scripts/runtime_monitor.py | alerts, residual_stats, dt_stats (alerts include NaN/Inf/overflow detection, residual growth, and dt collapse) |
scripts/result_validator.py | checks, confidence_score, failed_checks, status (PASS / FAIL / INSUFFICIENT_DATA); confidence_score is null when no check ran |
scripts/failure_diagnoser.py | probable_causes, recommended_fixes |
Three-Stage Validation Protocol
Stage 1: Pre-flight (Before Simulation)
- Run
scripts/preflight_checker.py --config simulation.json - BLOCK status: Stop immediately, fix all blocker issues
- WARN status: Review warnings, document accepted risks
- PASS status: Proceed to simulation
Note:
preflight_checker.pyvalidates required keys, numeric ranges, output-directory access, and disk space. It does not evaluate numerical stability (CFL / diffusion-Fourier). For explicit stability gating useskills/core-numerical/numerical-stability/scripts/cfl_checker.py.
python3 scripts/preflight_checker.py \
--config simulation.json \
--required dt,dx,kappa \
--ranges "dt:1e-6:1e-2,dx:1e-4:1e-1" \
--min-free-gb 1.0 \
--json
Stage 2: Runtime (During Simulation)
- Run
scripts/runtime_monitor.py --log simulation.logperiodically - Configure alert thresholds based on problem type
- Stop simulation if critical alerts appear
python3 scripts/runtime_monitor.py \
--log simulation.log \
--residual-growth 10.0 \
--dt-drop 100.0 \
--json
Stage 3: Post-flight (After Simulation)
- Run
scripts/result_validator.py --metrics results.json - All checks PASS: Results are valid for analysis
- Any check FAIL: Do NOT use results, diagnose failure
python3 scripts/result_validator.py \
--metrics results.json \
--bound-min 0.0 \
--bound-max 1.0 \
--mass-tol 1e-3 \
--json
For variational / gradient-flow models (Allen-Cahn, Cahn-Hilliard), add
--variational to enforce a strict monotone non-increasing energy check.
Failure Diagnosis
When validation fails:
python3 scripts/failure_diagnoser.py --log simulation.log --json
Conversational Workflow Example
User: My phase field simulation crashed after 1000 steps. Can you help me figure out why?
Agent workflow:
- First, check the log for obvious errors:
python3 scripts/failure_diagnoser.py --log simulation.log --json - If diagnosis suggests numerical blow-up, check runtime stats:
python3 scripts/runtime_monitor.py --log simulation.log --json - Recommend fixes based on findings:
- If residual grew rapidly → reduce time step
- If dt collapsed → check stability conditions
- If NaN detected → check initial conditions
Error Handling
| Error | Cause | Resolution |
|---|---|---|
Config not found | File path invalid | Verify config path exists |
Non-numeric value | Parameter is not a number | Fix config file format |
out of range | Parameter outside bounds | Adjust parameter or bounds |
Output directory not writable | Permission issue | Check directory permissions |
Insufficient disk space at <path> | Disk nearly full on the output volume | Free up space or reduce output |
Invalid parameter name | --required name has disallowed characters | Use only letters, digits, _, ., - |
range max ... must be greater than min | Inverted/degenerate --ranges or bounds | Ensure max > min |
must be a finite positive number | nan/inf/negative threshold supplied | Pass a finite positive value |
Log file too large | Log exceeds the 500 MB parse cap | Truncate or pre-filter the log |
Interpretation Guidance
Status Meanings
| Status | Meaning | Action |
|---|---|---|
| PASS | All checks passed | Proceed with confidence |
| WARN | Non-critical issues found | Review and document |
| BLOCK | Critical issues found | Must fix before proceeding |
Confidence Score Interpretation
| Score | Meaning |
|---|---|
| 1.0 | All validation checks passed → proceed with confidence |
| 0.75+ | Most checks passed, minor issues |
| 0.5-0.75 | Significant issues, review carefully |
| < 0.5 | Major problems, do not trust results |
null (status INSUFFICIENT_DATA) | No recognized metrics fields; no check ran — NOT a pass. Inspect the metrics file. |
A requested bound (--bound-min/--bound-max) with no matching field_min/field_max
in the metrics is reported as a failed bounds_unverifiable check, never a vacuous pass.
For variational/gradient-flow runs, pass --variational (or set "energy_variational": true
in the metrics) to enforce a strict monotone non-increasing energy check (energy_monotone);
otherwise a weaker energy_net_decrease check is used, which does not detect mid-run spikes.
Common Failure Patterns
| Pattern in Log | Likely Cause | Recommended Fix |
|---|---|---|
| NaN, Inf, overflow | Numerical instability | Reduce dt, increase damping |
| max iterations, did not converge | Solver failure | Tune preconditioner, tolerances |
| out of memory | Memory exhaustion | Reduce mesh, enable out-of-core |
| dt reduced | Adaptive stepping triggered | May be okay if controlled |
Verification checklist
Do not trust a validation verdict until each applicable item below is satisfied with the concrete artifact named. Record these in your summary to the user.
- Ran
result_validator.py --jsonand confirmedresults.statusisPASS(notINSUFFICIENT_DATA) ANDresults.confidence_score == 1.0; anullscore orINSUFFICIENT_DATAmeans no check ran — treat as unverified, not as a pass. - Listed
results.checksand confirmed every requested check actually appears (e.g.mass_conserved,bounds_satisfied,no_nan, andenergy_monotone/energy_net_decrease); confirmedresults.failed_checksis empty and contains nobounds_unverifiableentry (which means a requested bound had nofield_min/field_maxto compare against). - For variational/gradient-flow models (Allen-Cahn, Cahn-Hilliard), passed
--variational(or set"energy_variational": true) soenergy_monotoneis enforced; recorded that the weakerenergy_net_decreasewas NOT relied on, since it cannot detect mid-run energy spikes. - Recorded the mass drift tolerance used (
--mass-tol, default1e-3) and confirmed it matches the Conservative/Standard/Relaxed column appropriate to the run; did not silently accept the default for a tight-conservation problem. - Ran
runtime_monitor.py --jsonand recordedresidual_stats(min/max/last) anddt_stats; confirmed there are noalertsfor NaN/Inf/overflow, residual growth above--residual-growth, or dt collapse below--dt-drop. - Confirmed numerical stability was gated separately via
core-numerical/numerical-stability/scripts/cfl_checker.py(CFL/Fourier limit) —preflight_checker.pydoes NOT evaluate CFL/Fourier and a PASS preflight says nothing about temporal/spatial stability. - On any
FAILor alert, ranfailure_diagnoser.py --jsonand recorded theprobable_causes/recommended_fixes, rather than reusing the results.
Common pitfalls & rationalizations
| Tempting shortcut | Why it's wrong / what to do |
|---|---|
| "Preflight passed, so the run is numerically stable." | preflight_checker.py checks required keys, ranges, output-dir writability, and disk space only. It does NOT compute CFL/Fourier. Gate stability with cfl_checker.py separately. |
"result_validator printed a confidence score, so results are good." | An empty or unrecognized metrics file returns confidence_score: null and status INSUFFICIENT_DATA — that is "no check ran", not a pass. Verify recognized fields are present and status == PASS. |
| "Energy ends lower than it started, so the dissipative run is fine." | The default energy_net_decrease only compares first vs last and misses mid-run spikes. For gradient-flow models use --variational to enforce the strict monotone energy_monotone check. |
"I asked for bounds and didn't get a bounds_satisfied: false, so bounds hold." | If field_min/field_max are absent the validator emits bounds_unverifiable (a FAILED check), never a vacuous pass. Ensure the metrics file actually carries the field extrema. |
| "The simulation finished without crashing, so the results are trustworthy." | Run completion is not correctness. Verify mass conservation, energy behavior, physical bounds, and a clean runtime_monitor alert list before using results. |
| "dt got smaller during the run, so the solver is failing." | runtime_monitor dt-collapse is direction-aware (running-max vs current) and only alerts past --dt-drop; a controlled adaptive ramp is expected. Check the actual dt_stats and whether an alert fired. |
| "I'll just use the default thresholds." | Defaults (--mass-tol 1e-3, --residual-growth 10, --dt-drop 100) are the Standard column; a conservation-critical problem needs the Conservative tolerances. Pick thresholds for the physics, then record them. |
Security
Input Validation
- Config file paths are validated for existence before parsing; non-existent paths produce clear errors (exit code 2)
--requiredparameter names are validated against a safe-character allowlist (^[A-Za-z0-9_.-]+$); names with shell metacharacters are rejected--rangesentries are parsed asname:min:maxwith finite numeric bounds enforced andmax > minrequired--min-free-gbis validated as a finite positive number (negatives, zero,nan,infrejected)--residual-growthand--dt-dropthresholds are validated as finite positive numbers--bound-minand--bound-maxare validated as finite numbers (nan/infrejected), and--bound-max > --bound-minis enforced;--mass-tolis validated as a finite positive number- Invalid input exits with code 2 and an explanatory message
File Access
preflight_checker.pyreads a single user-specified config file (JSON/YAML) and checks disk space on the volume hosting the resolved output directoryruntime_monitor.pyreads a single log file specified by--log; log files are size-limited (500 MB max) and rejected before parsing if largerresult_validator.pyreads a single metrics file (JSON) specified by--metricsfailure_diagnoser.pyreads a single log file specified by--log; log files are size-limited (500 MB max) before parsing- No scripts write to the filesystem; all output goes to stdout
Tool Restrictions
- Read: Used to inspect script source, references, config files, and simulation logs
- Bash: Used to execute the four Python validation scripts (
preflight_checker.py,runtime_monitor.py,result_validator.py,failure_diagnoser.py) with explicit argument lists - Write: Used to save validation reports; writes are scoped to the user's working directory
- Grep/Glob: Used to locate log files, config files, and search references
Safety Measures
- No
eval(),exec(), or dynamic code generation - All subprocess calls use explicit argument lists (no
shell=True) failure_diagnoser.pyuses hardcoded, pre-compiled diagnostic regex patterns;runtime_monitor.pyaccepts optional--residual-pattern/--dt-patternoverrides that are compiled withre.compile(noeval) and applied only to the user's own log- Diagnostic strings emitted in output are drawn from the skill's fixed cause/fix table, not interpolated from raw log content
Limitations
- Not a real-time monitor: Scripts analyze logs after-the-fact
- Regex-based: Log parsing depends on pattern matching; may miss unusual formats
- No automatic fixes: Scripts diagnose but don't modify simulations
References
references/validation_protocol.md- Detailed checklist and criteriareferences/log_patterns.md- Common failure signatures and regex patterns
Version History
- v1.2.2 (2026-06-24): Added a Verification checklist (evidence-based, tied to the four scripts' JSON outputs) and a Common pitfalls & rationalizations table to harden agent interpretation of validation verdicts.
- v1.2.0 (2026-06-23): Corrected diagnostic regexes (no false convergence/blow-up on healthy logs), direction-aware dt-collapse detection, NaN/Inf scan in runtime monitor, strict variational energy check, non-vacuous bounds/confidence, config-relative output-dir + correct-volume disk check, and implemented the documented input-validation/file-size safeguards
- v1.1.0 (2024-12-24): Enhanced documentation, decision guidance, Windows compatibility
- v1.0.0: Initial release with 4 validation scripts
What ships with it: 14 files
56.7 KB alongside SKILL.md, 4 of them executable
evals/
- evals.json9.1 KB
- files/config.json116 B
- files/crash.log223 B
- files/results.json158 B
- files/run.log162 B
- files/simulation.json146 B
- files/simulation.log231 B
references/
- log_patterns.md7.4 KB
- validation_protocol.md7.7 KB
scripts/
- failure_diagnoser.pyruns3.1 KB
- preflight_checker.pyruns9.4 KB
- result_validator.pyruns7.9 KB
- runtime_monitor.pyruns6.3 KB
- CHANGELOG.md4.7 KB
Gives 0 of the 12 instructions most quality gates skills give in ~3.6k tokens
Counted across 1,524 of the 2,830 authors here whose files we hold, read 2026-09-06
- Read full output and check exit codein 45 of 1524, across 40 files
- Verify output confirms the claimin 44 of 1524, across 39 files
- Identify the command that proves the claimin 43 of 1524, across 39 files
- Execute the full verification commandin 36 of 1524, across 30 files
- Produce a verification reportin 34 of 1524, across 18 files
- Review git diff changesin 30 of 1524, across 16 files
- Fix build failures immediatelyin 29 of 1524, across 9 files
- Group findings by severityin 28 of 1524
- State claim only with evidencein 27 of 1524, across 22 files
- Verify regression tests with red-green cyclein 26 of 1524, across 22 files
- Run the full test suitein 26 of 1524, across 25 files
- Run test suite with coveragein 25 of 1524, across 10 files
Said here and by no other author read
- run preflight_checker.py before starting a simulation
- run runtime_monitor.py periodically during simulation
- run result_validator.py after simulation finishes
- run failure_diagnoser.py when validation fails
- fix all blocker issues before running simulation
- document accepted risks for warnings
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.