agentsclimarketplace

Aamas artifact evaluation

Skill brycewang-stanford/Awesome-Journal-Skills/AAMAS-Skills/skills/aamas-artifact-evaluation

Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的 Claude Code/Codex 期刊技能包,从选题、识别策略到表格规范与审稿回复全流程,助你快速发论文。

Install
npx -y skills add brycewang-stanford/Awesome-Journal-Skills --skill aamas-artifact-evaluation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Use when packaging AAMAS multiagent code, environments, opponent and population sets, random seeds, game definitions, and logs as anonymous supplementary evidence or a public post-acceptance release, even without a separate artifact badge, so that game-theory and MARL reviewers can inspect and re-run the interaction claims.

SKILL.md

3.7 KB, as published. Nobody here has run it

AAMAS Artifact Evaluation

Use this for evidence packaging around AAMAS. Because the venue is about interaction, an artifact must make a multiagent claim inspectable: the game, the other agents, and the protocol, not just a single trained model.

Artifact plan

  • Decide what a reviewer needs to believe the interaction claim: game or environment code, opponent/population definitions, the training regime, seeds, payoff logs, proofs, or qualitative episode traces.
  • Keep decision-critical evidence in the main paper or appendix; optional bulk runs can live in the supplementary zip.
  • Anonymize repository history, paths, environment names, license headers, cluster paths, and commit authors for the review version.
  • Include a minimal reproduction map: environment build, dependencies, hardware, commands, expected outputs, per-run wall-clock, seeds, and known nondeterminism (especially in self-play).
  • For a deployed or human-subject setting, give enough provenance for credible reproduction without violating data-use terms.
  • After acceptance, replace anonymous archives with a public, licensed, citable artifact.

What AAMAS evidence reviewers open first

The single fact that shapes packaging: a reviewer will re-run a small game far sooner than they will retrain a large policy, so make the strategic core turnkey before polishing anything.

Claim typeFirst artifact inspectedCommon failure caught
Convergence to an equilibriumThe game definition and the learning-rule codeSolution concept named in the paper but not encoded in the evaluation
Emergent cooperation/defectionThe environment and reward specificationResult depends on an undocumented reward-shaping constant
Beats other agentsThe opponent/population set and match protocolOnly self-play reported; no held-out opponents
Mechanism is truthfulThe payment rule plus a strategic-deviation testNo script that lets an agent try to game the mechanism

Worked vignette: packaging a self-play study

A hypothetical submission claims a learning rule that converges to a correlated equilibrium in a repeated congestion game, shown by self-play.

  • Ship the game as one parameterized generator (number of agents, capacity, payoff scale) rather than constants buried in a notebook, so reviewers can vary the interaction.
  • Record the exact seed sequence and replication count behind every convergence plot; an equilibrium-convergence claim without seeds is unfalsifiable.
  • Emit payoff and regret tables directly from logged results so PDF and artifact numbers cannot drift.
  • Include a strategic-deviation harness: a script that drops in a non-conforming agent and measures whether it profits, because that is exactly what a game-theory reviewer will try.

Calibration anchors

  • Supplement inspection at AAMAS is at reviewer discretion; assume only the README and one entry script get opened, and design the top level accordingly.
  • Supplement size and format caps vary by cycle (25 MB single zip in 2026); verify against the current OpenReview form rather than a past year.

Output format

[Artifact role] anonymous supplement / camera-ready release / public archive
[Contents] <game/env/opponents/seeds/proofs/logs>
[Anonymity risks] <paths/metadata/licenses/URLs>
[Reproduction level] turnkey / scripted / descriptive / weak
[Fixes before upload] <ordered list>

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.