Hcod result reporting
Skill jiangkaiqi2005/jiangkaiqi-Skills/hcod-result-reporting
Personal AI skills by Jiang Kaiqi for study tutoring, guided learning, exam review, and VMamba environment troubleshooting.
npx -y skills add jiangkaiqi2005/jiangkaiqi-Skills --skill hcod-result-reportingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Use when HyperCOD/VMamba experiment results are ready and Codex needs to archive train logs, choose or confirm an output root, generate train-log HTML visualizations, identify the best iteration, or collect paper-style metrics.
SKILL.md
5.5 KB, ~1.2k tokens by cl100k_base, as published. Nobody here has run it
HCOD Result Reporting
Overview
Archive a finished HyperCOD experiment into a clear outputs folder, produce a compact report, and verify which checkpoint should be treated as best. The core rule is to classify the experiment family before naming folders; never dump method-family results into the baseline bucket just because paths look similar.
Workflow
-
Locate the source run directory and artifacts.
- Find
train.nohup.log, config, checkpoints, existingpaper_metrics*, and any previous HTML reports. - Read project notes when version semantics are unclear:
CHANGELOG.md,version_modules.md, run-local notes, and nearby configs.
- Find
-
Choose the output family name before creating files.
- Inspect existing directories under the user-specified output root when one is provided.
E:\ViM\outputsis only an example of a possible root, not a default. - Use the method family, not the implementation shortcut. Examples:
- Basic/baseline experiments:
hcod_hsi_basic_cache_160k/<run-name> - HSC-SAM imitation experiments:
hcod_hsc_sam_cache_160k/<run-name> - RGB/spectral experiments: follow existing
hcod_rgb_*naming.
- Basic/baseline experiments:
- If unsure whether a run is baseline, HSC-SAM imitation, ablation, or another family, infer from run notes/configs first; ask the user only if local evidence is ambiguous.
- Keep child folders versioned and descriptive, for example
hcod_hsc_sam_v0.1_cache_160k.
- Inspect existing directories under the user-specified output root when one is provided.
-
Create or use the destination.
- If the user did not specify an output root, ask where to put the results before creating or copying files. Do not assume
E:\ViM\outputs. - If the user specifies
E:\ViM\outputsor another existing convention, prefer that real path when mounted or accessible. - If the requested drive is unavailable, create a same-structure local mirror under the repo only as a fallback, such as
outputs/<family>/<run>, and clearly say it is not the requested root. - Copy
train.nohup.loginto the destination. Do not copy unrelated run notes whose labels conflict with the method family.
- If the user did not specify an output root, ask where to put the results before creating or copying files. Do not assume
-
Parse the training log.
- Extract train points from
Iter(train)lines: iter, loss, decode/aux losses, accuracy, lr, time. - Extract validation points from
Iter(val)summary lines: iter,aAcc,mIoU,mAcc,mDice, timing. - Associate validation rows with the immediately preceding checkpoint save when the val line does not include the iter.
- Extract train points from
-
Pick the best iteration.
- Default rule: highest validation
mIoU; break ties bymDice, thenaAcc. - If the user names a different primary metric, use that and state the rule.
- Compare the best point with the latest point. Mention when training loss improves but validation quality regresses.
- Default rule: highest validation
-
Generate the HTML visualization.
- Match the style of the closest existing report when the user provides one; otherwise create a self-contained HTML file that opens directly.
- Include summary cards, training loss curves, validation metric curves, validation table, best-iter conclusion, paper metrics, and the reproduce command.
- Avoid external CDN dependencies unless the existing report already requires them.
-
Establish paper-style metrics for the selected iter.
- Prefer running the project tool:
CUDA_VISIBLE_DEVICES=<gpu> python segmentation/tools/evaluate_hcod_paper_metrics.py \
<config> <checkpoint> \
--split test \
--with-flops \
--out-dir <destination>/paper_metrics_iter<iter>
- Save
summary_paper_metrics.json,summary_paper_metrics.csv,per_image_paper_metrics.csv, andmetrics_command.txt. - If the command fails, debug the root cause. If valid metrics already exist from the same config/checkpoint, copying them is acceptable only with an explicit note in
metrics_command.txtand the final response. - Remove failed-run debris such as timestamped work dirs that only contain failed configs/logs.
- Verify before reporting completion.
- Confirm expected files exist and are non-empty.
- Re-read the validation CSV/JSON and recompute best iter.
- Confirm the HTML contains embedded data, the best iter, and chart elements.
- Scan the destination for stale/wrong family names.
- State anything that could not be verified, such as unavailable
E:mount or missing browser tooling.
Required Outputs
train.nohup.logtrain_log_visualization_<version>.htmlvalidation_summary_<version>.csvbest_iter_analysis_<version>.mdpaper_metrics_iter<best_iter>/summary_paper_metrics.jsonpaper_metrics_iter<best_iter>/summary_paper_metrics.csvpaper_metrics_iter<best_iter>/per_image_paper_metrics.csvpaper_metrics_iter<best_iter>/metrics_command.txt
Common Mistakes
- Do not assume
hcod_hsi_basic_cache_160kis correct just because the source work dir contains HSI/cache/160k. That folder is for baseline/basic runs. - Do not assume the output root. If the user did not name one, ask first.
- Do not say files were downloaded to the requested root when only a local mirror was created.
- Do not pick the final checkpoint as best without checking validation metrics.
- Do not hide paper-metrics failures. Record the exact failure and whether metrics were copied from an existing successful run.
- Do not leave wrong-name temporary outputs after the user confirms they downloaded the mirror; remove the local mirror when asked.
What ships with it: 1 file
317 B alongside SKILL.md
agents/
- openai.yaml317 B