agentsclimarketplace

Iclr experiments

Skill brycewang-stanford/Awesome-Journal-Skills/ICLR-Skills/skills/iclr-experiments

Use when designing or auditing ICLR experiments, including baselines, ablations, scaling laws, robustness, statistics, benchmarks, human evaluation, and compute reporting. Use when a reviewer questions whether a representation-learning or model gain is real, when you must isolate one mechanism with an ablation, or when preparing a small compute-matched control that can be posted inline during the public discussion period.From its SKILL.md

Install
npx -y skills add brycewang-stanford/Awesome-Journal-Skills --skill iclr-experiments

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

3.6 KB, 678 tokens by cl100k_base, as published. Nobody here has run it

ICLR Experiments

Use this before submission or during a revision pass to stress-test empirical claims. ICLR experiments should answer the scientific question, not merely assemble a leaderboard.

Experiment audit

  • Match each experiment to a claim in the introduction.
  • Compare against current strong baselines, open-source systems, and the most relevant recent OpenReview/arXiv papers.
  • Add ablations that isolate one mechanism at a time.
  • Report variance across seeds or runs when randomness can change conclusions.
  • Include robustness checks for dataset shift, prompt changes, architecture variants, hyperparameter sensitivity, or compute scale when those affect the claim.
  • State compute budget, hardware, training time, inference cost, and environmental or access limits where relevant.
  • For human evaluation, document task, annotator instructions, aggregation, quality control, and IRB or ethics status when needed.

Reviewer questions to pre-answer

  • Is the baseline tuned fairly?
  • Does the method win because of more compute, data, parameters, or prompt search?
  • Does the effect persist outside the easiest benchmark?
  • Are negative results hidden?
  • Can a reviewer reproduce the headline table from the supplement or artifact?

What ICLR reviewers reward in evidence

ICLR's empirical culture prizes honest ablations and mechanism over leaderboard position. A clean ablation that explains why a representation works often outscores a larger raw number.

Claim typeEvidence that convinces ICLR reviewersCommon reject trigger
New objective helpsAblate the objective with everything else fixedGains confounded with extra tuning
Method scalesSeveral model sizes/tasks with a trendOne large run, no scaling curve
Robust representationTests across shifts, seeds, promptsSingle-seed peak on one benchmark
Beats prior methodTuned, current, open-source baselineStale or under-tuned baseline

Worked vignette

A paper claims a new self-supervised pretext task yields better linear-probe accuracy. Reviewers ask whether the gain is the pretext task or simply longer pretraining. The author audit: hold total pretraining compute fixed, swap only the pretext objective, and report linear-probe accuracy with error bars over five seeds. The compute-matched ablation isolates the mechanism and is small enough to post inline during discussion, where the table becomes part of the permanent public record.

Reviewer-pushback patterns

  • "You win because of more compute." Add a compute-matched control; report FLOPs, not just wall time.
  • "Only one seed." Report mean and spread across seeds; an unstable benchmark needs variance.
  • "Baseline is weak." Cite the baseline's own recommended settings and show you matched them.
  • "Ablation removes two things at once." Split into single-mechanism ablations a reviewer can read.

Output format

[Claim] <paper claim>
[Experiment evidence] sufficient / needs baseline / needs ablation / needs robustness
[Fairness issue] <compute, tuning, data, prompt, metric>
[Fast fix] <experiment or analysis feasible before deadline>
[Appendix placement] <what can move out of main text>

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.