agentsclimarketplace

Cross harness

Skill tmuskal/arc-agi-benchmarker/plugins/longmemeval-benchmarker/skills/cross-harness

Prove you achieved AGI at home by testing your claude-code setup against arc-agi-3 benchmarks

Install
npx -y skills add tmuskal/arc-agi-benchmarker --skill cross-harness

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Cross-harness benchmarking for LongMemEval - emit instruction packs for Codex/Gemini/OpenCode and ingest their results into a comparable scorecard

SKILL.md

2.9 KB, as published. Nobody here has run it

LongMemEval Cross-Harness

Step 1: Parse args

ArgumentValues
<action>emit / import
--harnesscodex / gemini / opencode
--run-idrequired for import
--results-pathpath to foreign item-results.jsonl (for import)

Step 2: emit — write instruction pack

Produce .longmemeval-benchmarks/cross-harness/<harness>-instructions.md containing:

  1. Link to the LongMemEval repo.
  2. The exact dataset file to use (cfg.datasetPath).
  3. The exact schema for item-results.jsonl (copy from SPEC.md).
  4. Instructions: for each item, produce a hypothesis. Emit one jsonl line per item with the schema fields. Do NOT judge — return the raw hypotheses. The Claude harness will re-judge for fairness.
  5. The output jsonl file they should produce, and how to hand it back.

Step 3: import — judge foreign hypotheses

Re-judge the foreign hypotheses with our judge (so judge model is constant across harnesses):

$VENV_PYTHON -c "
import json, sys, uuid
from pathlib import Path
from datetime import datetime, timezone

sys.path.insert(0, 'plugins/longmemeval-benchmarker/scripts')
from checkpoint_io import append_item_result, mark_completed, write_atomic_json
from judge_shim import judge

cfg = json.load(open('.longmemeval-benchmarks/config.json'))
harness = '<HARNESS>'
run_id = '<RUN_ID>'
foreign = Path('<RESULTS_PATH>')

run_dir = Path(cfg['runs_dir']) / run_id
run_dir.mkdir(parents=True, exist_ok=True)
meta = {
    'schemaVersion': '1.0.0', 'runId': run_id, 'harness': harness,
    'timestamp': datetime.now(timezone.utc).isoformat(),
    'datasetVariant': cfg['datasetVariant'], 'datasetPath': cfg['datasetPath'],
    'maxEvals': cfg['maxEvals'], 'seed': cfg['seed'],
    'targetModel': 'foreign', 'judgeModel': cfg['judgeModel'], 'judgeProvider': cfg['judgeProvider'],
    'harness_config': {'model': 'foreign', 'plugins': [], 'skills': [], 'mcp_servers': []},
    'status': 'running', 'plugin_version': '1.0.0',
}
write_atomic_json(run_dir / 'run-meta.json', meta)

for line in foreign.read_text().splitlines():
    if not line.strip(): continue
    r = json.loads(line)
    j = judge(r['question_type'], r['question'], r['answer'], r['hypothesis'],
              Path(cfg['longmemeval_root']),
              provider=cfg['judgeProvider'], model=cfg['judgeModel'])
    r['judgment'] = {'model': j['model'], 'label': j['label'], 'raw': j['raw']}
    r.setdefault('schemaVersion', '1.0.0')
    append_item_result(run_dir, r)
    mark_completed(run_dir, r['question_id'])

meta['status'] = 'completed'
write_atomic_json(run_dir / 'run-meta.json', meta)
print('imported run', run_id)
"

Step 4: Report + compare

Run /longmemeval-report <run_id> on the imported run, then /longmemeval-compare-runs <claude_run> <imported_run>.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.