Paper review lite codex
Skill scdenney/open-science-skills/codex/paper-review-lite-codex
Agentic skills for Claude Code and Codex, built from published social-science methods sources. Covers experimental design, computational text analysis, manuscript QA, and transparent reporting.
npx -y skills add scdenney/open-science-skills --skill paper-review-lite-codexAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
What its author says it does
Copied from the file, not written here
Run a cross-model adversarial pre-submission audit from Codex. Use when a user explicitly requests the heavier paper-review-lite-codex workflow, independent Codex and Claude review passes, or cross-model verification of a manuscript. Apply the paper-review-lite protocol, cross-check findings against the manuscript, and retain only evidence-grounded issues.
SKILL.md
7.0 KB, as published. Nobody here has run it
Paper review lite: Codex lead, cross-model audit
Run Codex as the lead and use Claude Code's non-interactive CLI as the independent model family. This is the Codex-side equivalent of the Claude skill with the same name; the direction of delegation is reversed.
Read ../paper-review-lite/SKILL.md completely before starting. It defines the review dimensions, severity rubric, evidence requirements, and final report format. If the sibling skill is unavailable, stop and tell the user to install it; do not reconstruct a partial protocol from memory.
Sandbox constraint — read before the first claude -p call
claude -p is a different binary than codex exec, so it does not hit the in-process IPC failure that breaks nested codex exec calls under sandbox. But under workspace-write sandbox, an outbound claude -p network call was observed (July 2026) to hang rather than complete or fail cleanly — Codex's sandbox restricts network access, and claude -p needs it to reach Anthropic's API. That observation is less rigorously isolated than the nested-codex exec failure (no distinct error message, just an unresponsive process that had to be killed), so treat it as a strong warning rather than a certainty. If a claude -p call hangs rather than returning, do not assume it will eventually resolve — request escalation (sandbox_permissions: require_escalated) for that call, or run from an unsandboxed session, before retrying.
Preflight
- Confirm that the user explicitly requested
$paper-review-lite-codexor a cross-model audit. Otherwise use$paper-review-lite. - Locate the manuscript, supplement, bibliography, figures, preregistration, and replication archive.
- Check
command -v claudeand run a harmless authentication/status check supported by the installed CLI. Do not print credentials. - Before the first call, confirm the user accepts an external model call that may consume separate credits — unless they already authorized Claude or cross-model execution.
- If Claude is unavailable or authorization is declined, offer the fallback in “Reduced-diversity mode.”
Claude Code documents claude -p as its non-interactive interface. Use --output-format text, --no-session-persistence, and read-only tools for review calls. See Run Claude Code programmatically.
Orient once
Follow $paper-review-lite Phase 1. Read the manuscript before spawning reviewers. Record actual absolute paths, design family, source format, target journal, and missing inputs. Create:
.review-tmp/
├── codex/
├── claude/
└── cross-check/
Phase 1: independent Red Teams
Codex cohort
Run the nine $paper-review-lite review dimensions in parallel batches that respect the runtime's agent limit. Each subagent receives only:
- the relevant dimension block from
$paper-review-lite; - the manuscript and related absolute paths;
- the common severity and quote requirements; and
- its assigned output path under
.review-tmp/codex/.
Keep agents blind to one another. Require direct manuscript evidence for every critical or recommended issue.
Claude cohort
Run three independent Claude CLI calls concurrently when safe:
- argument, claims, and numerical consistency;
- references, DOI status, writing, figures, and tables;
- methods, CONSORT/preregistration where applicable, and replication readiness.
Construct each prompt from the matching dimension blocks in $paper-review-lite; do not paraphrase away requirements. Tell Claude to read the named files, print structured Markdown only, and never edit them. Redirect stdout to the assigned file under .review-tmp/claude/.
Use this command shape, substituting absolute paths and a complete prompt file:
claude -p \
--output-format text \
--no-session-persistence \
--allowedTools "Read" \
"$(< /absolute/path/to/prompt.txt)" \
> /absolute/path/to/.review-tmp/claude/review-N.md
Do not use --bare unless API-key authentication is configured explicitly; bare mode skips normal OAuth and keychain discovery.
Validate every output file. A non-zero exit, empty file, or refusal is a failed reviewer, not evidence that the manuscript passed.
Phase 2: blind cross-check
After both cohorts finish, launch four verification passes:
- two Codex subagents verify Claude's content/methods and technical findings;
- two Claude CLI calls verify Codex's corresponding findings.
For each critical or recommended issue, require this sequence:
- Find the quoted span at the cited location. If absent, mark
QUOTE FAILEDand drop it. - Check whether the manuscript already answers the concern elsewhere. If yes, mark
REFUTEDorDOWNGRADEwith the contradicting location. - Check the issue on its merits. Mark
CONFIRMED,REFUTED, orDOWNGRADE. - Do not add new issues during cross-checking.
Cross-checkers read the manuscript plus the other cohort's files, never their own cohort's findings.
Phase 3: adjudicate
Apply these rules in order:
- Mutual and confirmed: retain at the stronger justified severity; label
Mutual (Codex + Claude). - Single-cohort and cross-confirmed: retain; label the originating cohort and confirmation.
- Single-cohort and cross-refuted: drop unless the lead rereads the source and records specific contrary evidence. If retained, demote one tier.
- Quote failed: drop.
- Reviewer failed: record reduced coverage; never convert missing output into a pass.
Deduplicate by underlying issue, not wording. Then follow $paper-review-lite Phase 4 to write the recommendation, editor's note, severity-ranked issues, strengths, readiness checklist, and “What Still Needs Your Input.” Add a Confidence column to critical and recommended findings.
Reduced-diversity mode
If Claude cannot run, ask whether to continue with two blind Codex cohorts. If approved:
- give each cohort fresh context and identical specifications;
- keep cohorts blind until cross-checking;
- label the report
same-model independent review, nevercross-model; and - state that agreement is weaker evidence because model-family blind spots are shared.
If the user declines, run ordinary $paper-review-lite or stop as requested.
Completion checks
- Every retained critical or recommended issue has a verified quote and location.
- Both directions of cross-check completed, or reduced coverage is stated explicitly alongside any failed or empty reviewer output.
- No agent edited the manuscript.
- The final report is self-contained and does not expose scratch prompts or agent chatter.
.review-tmp/is removed after delivery unless the user asked to keep it.