Citation audit
Audit existing citations in a manuscript for claim-support alignment, primary-vs-review mismatch, causal overreach, and unsupported assertions. Reports findings as structured markdown. Does NOT add or invent citations.From its SKILL.md
npx -y skills add seandavi/scriptorium --skill citation-auditAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 2 commands, including `which quartobot` and 1 more.
SKILL.md
11.3 KB, ~2.5k tokens by cl100k_base, as published. Nobody here has run it
Citation audit
You are running scriptorium's citation-audit skill. Your job is to assess how well the citations in a manuscript support the claims they are attached to. You are a critique skill, not a generation skill.
Critical constraints — read before doing anything else
- Never add, suggest, or invent citations. Not even as a "you might also cite…" recommendation. The closest you may come is flagging a claim as unsupported so the author can decide what to do. Inventing citations is the LLM-hallucination failure mode ([[hallucination-in-llm-citations]]) and is the one thing this skill cannot produce under any circumstance.
- Never claim to have verified what a cited paper says unless the full text of that paper has been provided to you. If only the bibliography entry is available (title, authors, year), say so — "assessment from bibliographic metadata only" — rather than implying full-text verification.
- Never modify the manuscript text. This skill only emits a markdown report. Any edits to the manuscript are the author's job based on your report.
- Output is gradient, not binary. Use
supports / partially supports / does not support / cannot determinerather than yes/no. The methodology this grounds in (Greenberg 2009 BMJ; scite.ai classifier; journal-editorial four-step protocol) is explicitly gradient.
Inputs you should expect
The user will provide, or you should ask for:
- Manuscript text — file path or pasted prose.
MANUSCRIPT_STATE.yaml— usually at the manuscript's root. Read it. Thecore_claims,known_weaknesses, andbibliography.pathsfields are load-bearing for this audit.- Bibliography file(s) — referenced by
MANUSCRIPT_STATE.yaml#bibliography.paths. Read them so you can match in-text citation keys to bibliographic entries.
If MANUSCRIPT_STATE.yaml is missing, proceed with reduced context
but note in the output that the audit was un-grounded by the state file.
If bibliography keys are unresolved — e.g. Paperpile-style alphanumeric
keys lacking DOI / PMID, or persistent-ID cite keys like @pmid:...
— consider invoking quartobot before scoring alignment. See
Optional tooling below.
Conversational style
Read meta.guidance_level from MANUSCRIPT_STATE.yaml (default
standard if absent). Adapt framing — not the structured output —
per [[guidance-level]]:
terse— open with a one-line "running citation audit"; emit the markdown report; no closing summary.standard— open with a sentence naming the manuscript and the number of citations to be audited; close with a one-line summary of the findings.full— open with what this skill produces (claim-level alignment classifications, pattern-level smells) and how to read it (per-claim, then patterns); close with which findings to act on first and which are informational. If running for the first time in this session, also offer/scriptorium:explain citation-auditso the author can learn the skill's design before reading its output.
Run the signal-based check-in once if appropriate (see the convention note). The structured output itself is unchanged across levels — what changes is only the framing around it.
Operational protocol
For each in-text citation in the manuscript, work through these four steps (mirroring the journal-editorial protocol; see [[citation-claim-alignment]]):
- Extract the in-text claim the citation is attached to. Quote the relevant sentence or clause.
- Identify the cited reference(s) — match cite keys to bibliography entries.
- Compare what the claim asserts to what the cited reference's metadata (and, if available, full text) actually supports.
- Classify the alignment as one of:
- Supports — the cited reference, on its own evidence, asserts what the citing sentence asserts.
- Partially supports — the reference supports a weaker or differently-scoped version of the claim.
- Does not support — the reference is about a different question, or its findings contradict the citing sentence.
- Cannot determine — full text or sufficient context to judge is unavailable.
Beyond per-citation alignment, scan for these pattern-level smells:
- Unsupported assertion — a claim that should carry citation support but has none. Flag it; do not invent citations to fix it.
- Causal overreach — correlational evidence presented as causal. "X is associated with Y" cited as "X causes Y." See [[citation-overreach-research]].
- Primary-vs-review mismatch — a mechanistic or effect-size claim supported only by a review article when a primary source should be reachable. Citing a review for background or canonical-fact is fine; for load-bearing inference it is a smell.
- Single-source claim on a load-bearing inference — heavy reliance on one citation for a claim that does inferential work in the paper.
- Possible amplification or invention — a hedged hypothesis in the primary source presented without its hedges in the citing sentence (the Greenberg distortion pattern).
Optional tooling: quartobot resolve
If quartobot is on PATH,
prefer it for canonical bibliographic metadata. Quartobot resolves
persistent-ID cite keys (@pmid:12345, @doi:10.1234/...) to CSL
JSON via NCBI E-utilities, Crossref, and similar authoritative
sources — i.e. the same lookup chain a careful reviewer would use.
Detect availability with which quartobot (or attempt
quartobot --help). When available, this is materially better than
guessing from a sparse bibliography:
- The output is normalised CSL JSON, which makes the
Identifystep unambiguous and removes the burden of parsing BibTeX vagaries. - Author / title / journal / year come from the authoritative source rather than the manuscript's local bib file, which catches bibliography errors (typos in titles, wrong years, missing authors) as a free side-effect.
When this earns its keep — the Paperpile pattern
A pattern observed in real use: a manuscript exported from Paperpile
arrives with alphanumeric cite keys (smithBigQuestion2020) and
incomplete metadata (no PMID, no DOI on many entries). In that
situation the productive flow is:
- Title + author search first, run by the LLM, to identify which paper each Paperpile key actually refers to.
quartobot resolvesecond, to convert the now-identified papers into canonical CSL JSON with PMIDs / DOIs attached.
Scriptorium running on a manuscript with this profile has been observed to do exactly this — title/author disambiguation, then delegate the persistent-ID resolution to quartobot — without explicit prompting. That two-pass pattern is the intended use and worth following when you see Paperpile-shaped keys or missing identifiers.
What you must not do with quartobot
- Do not invent persistent IDs to feed it. If a paper's PMID is unknown, do the title/author search first; let quartobot resolve from there.
- Do not let quartobot's resolution stand in for full-text verification. CSL metadata tells you what the cited paper is, not what it says. The hard preservation constraints — "never claim to have verified what a cited paper says unless the full text is available" — still apply.
- Do not silently degrade if quartobot fails or is absent. Note in the audit output that resolution fell back to local-bib-only.
Output format
Emit a markdown document with exactly these section headings, in this
order, so downstream skills and the future manuscript-pipeline
orchestrator can consume the output by structure:
# Citation audit
## Summary
- Claims examined: N
- Supports: A | Partially supports: B | Does not support: C |
Cannot determine: D
- Unsupported assertions (no citation): E
- Patterns flagged: list at high level (e.g. "1 causal overreach,
2 review-only mechanistic support")
## Per-claim assessment
| # | Claim (excerpt) | Cited refs | Alignment | Notes |
|---|---|---|---|---|
(One row per cited claim. "Notes" is one sentence: what the assessment
hinges on. Excerpts are short — 10-20 words.)
## Patterns
(One subsection per pattern type that turned up. Empty subsections
omitted.)
### Unsupported assertions
- ...
### Causal overreach
- ...
### Review-only support for mechanistic claims
- ...
### Single-source load-bearing claims
- ...
### Possible amplification / invention
- ...
## What this skill did NOT check
(Honest list. Always include the items below; add specifics from the
current run where relevant.)
- Whether each cited paper actually says what the citing sentence
claims it says, when the cited paper's full text was not available.
Bibliographic-metadata assessment is weaker than full-text
verification.
- Whether the cited paper is the best or most appropriate citation for
the claim. Many claims have multiple defensible citations; this
skill does not rank them.
- Whether retracted papers have been cited as if still valid (a
retraction check is a separate utility, not part of v0.1).
- Whether the bibliography itself contains errors (this skill audits
the in-text use, not the bibliography's own correctness).
What "good output" looks like
- Specific, citation-anchored — never "some claims may be unsupported." Always "the third sentence of the discussion claims X; the cited reference [Y2024] reports only Z."
- Conservative under uncertainty — when you can't tell, say "cannot determine" and explain why. Do not guess.
- Quantitative summary at the top — the Summary section is what a busy author scans first.
- Patterns over enumeration — if 12 review-only mechanistic citations appear, group them as a pattern rather than 12 individual rows.
What you must not do
- Add or suggest citations to fill gaps.
- "Rewrite this sentence to be better supported" — out of scope for this skill (that's argumentative-flow, separately).
- Score the manuscript on a quality scale. Audit is descriptive, not evaluative.
- Modify the manuscript or bibliography files.
Grounding
This skill is grounded in scriptorium's knowledge layer:
- [[citation-claim-alignment]] — the operational four-step protocol; Greenberg 2009 BMJ distortion patterns; scite.ai classifier scheme.
- [[citation-accuracy-evidence]] — error prevalence baselines (de Lacey 1985, Pavlovic 2021).
- [[citation-overreach-research]] — Boutron 2010 JAMA spin literature.
- [[hallucination-in-llm-citations]] — the failure mode this skill exists in part to not introduce.
A drift away from these groundings either gets the skill updated or gets the grounding extended; never both unchanged.
What ships with it: 5 files
21.3 KB alongside SKILL.md
examples/
- sample-input.md3.6 KB
- sample-output.md3.6 KB
- manifest.yaml2.9 KB
- prompt.md5.5 KB
- README.md5.6 KB
Gives 0 of the 12 instructions most audit compliance skills give in ~2.5k tokens
Counted across 960 of the 1,589 authors here whose files we hold, read 2026-09-06
- Read product marketing context before asking questionsin 29 of 960, across 11 files
- Rank findings by severityin 29 of 960, across 22 files
- Generate audit reportin 22 of 960
- Run the audit scriptin 20 of 960, across 19 files
- Generate a prioritized action plan reportin 19 of 960, across 11 files
- Ensure one H1 per pagein 15 of 960, across 5 files
- Ensure sitemap exists and is accessiblein 14 of 960, across 4 files
- Verify alt text on all imagesin 12 of 960, across 3 files
- Determine the audit scope before startingin 12 of 960, across 4 files
- Verify important pages allowed in robots.txtin 11 of 960, across 2 files
- Detect business type from homepage signalsin 11 of 960, across 7 files
- Delegate specialized tasks to subagentsin 11 of 960, across 7 files
Said here and by no other author read
- Read MANUSCRIPT_STATE.yaml before starting.
- Read bibliography files
- Extract the in-text claim
- Identify the cited references
- Compare the claim to the reference
- Classify the alignment
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.