Research critique
Evaluate scientific claims, review manuscripts, assess evidence quality. Bias detection, GRADE assessment, logical fallacy identification, structured peer review reports. Use when user says 'review this paper', 'evaluate this claim', 'peer review', 'is this study valid', 'check the methodology', 'evidence quality', 'critical analysis'. Do NOT use for literature search (use research-search). Do NOT use for code review (use code-review).From its SKILL.md
npx -y skills add GRIDLOCK-NYC/claude-skills --skill research-critiqueAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
15.2 KB, ~3.0k tokens by cl100k_base, as published. Nobody here has run it
Research Review
Overview
Two operating modes depending on input:
- Claim/Evidence Evaluation — assess a claim, study finding, or piece of evidence using critical thinking frameworks
- Manuscript Review — formally review a paper, preprint, or presentation using structured peer review workflow
Both modes draw on the same shared frameworks: bias detection, statistical evaluation, and evidence quality. Mode is detected from context. If ambiguous, ask.
Mode 1: Claim/Evidence Evaluation
Use when given a claim, result, or study to assess rather than a full manuscript to formally review.
Claim Evaluation Process
- Identify the claim — What exactly is asserted? Causal, associational, or descriptive? How strong is the claim (proven / likely / suggested / possible)?
- Assess the evidence — What evidence is provided? Direct or indirect? Sufficient for the strength of claim? Are alternative explanations ruled out?
- Check logical connection — Do conclusions follow from data? Logical leaps? Correlational data supporting causal claims? Limitations acknowledged?
- Evaluate proportionality — Is confidence proportional to evidence? Are hedging words used appropriately? Is speculation labeled?
- Check for overgeneralization — Do claims extend beyond the sample? Population restrictions acknowledged? Context-dependence recognized?
- Red flags — Causal language from correlational data; "proves" or absolute certainty; cherry-picked citations; ignored contradictions; extrapolation beyond data.
Output format for claim evaluation:
- Summary — what was evaluated
- Strengths — what is well-supported
- Concerns — organized by severity (critical / important / minor)
- Recommendations — actionable suggestions
- Overall Assessment — what conclusions the evidence actually supports
Mode 2: Manuscript Review
Use when given a full paper, manuscript, preprint, or presentation to formally review.
Stage 1: Initial Assessment
- What is the central research question or hypothesis?
- What are the main findings and conclusions?
- Is the work scientifically sound and significant?
- Are there immediate major flaws that preclude publication?
Output: 2-3 sentence summary capturing the manuscript's essence and initial impression.
Stage 2: Section-by-Section Review
Abstract and Title
- Accuracy: Does the abstract reflect actual content and conclusions?
- Clarity: Is the title specific, accurate, informative?
- Completeness: Are key findings and methods summarized?
Introduction
- Context: Is background adequate and current?
- Rationale: Is the research question clearly motivated?
- Novelty: Is originality and significance articulated?
- Objectives: Are aims/hypotheses clearly stated?
Methods
- Reproducibility: Can another researcher replicate from this description?
- Rigor: Are methods appropriate for the research question?
- Detail: Are protocols, reagents, parameters sufficiently described?
- Ethics: Are approvals, consent, and data handling documented?
- Statistics: Are statistical methods appropriate, described, justified?
- Validation: Are controls, replicates, and validation adequate?
Critical elements: sample sizes, power calculations, randomization, blinding, inclusion/exclusion criteria, software versions, multiple comparison correction.
Results
- Presentation: Logical, clear, without over-interpretation?
- Statistics: Effect sizes, CIs, p-values properly reported?
- Completeness: Are negative results included?
- Common issues: selective reporting, missing error bars, overfitting, batch effects, missing controls.
Discussion
- Interpretation: Conclusions supported by data?
- Limitations: Acknowledged and discussed?
- Context: Findings placed within existing literature?
- Red flags: Overstated conclusions; ignored contradictions; causal claims from correlational data; mechanistic claims without mechanistic evidence.
References
- Completeness, currency, balance (contrary viewpoints cited), accuracy, appropriate self-citation.
Stage 3: Reproducibility and Transparency
Data Availability
- Raw data deposited in appropriate repositories?
- Accession numbers provided?
- Data sharing restrictions justified?
Code and Materials
- Analysis code available (GitHub, Zenodo)?
- Unique materials available or described for recreation?
Reporting Standards
- Does the manuscript follow discipline-specific guidelines: CONSORT (RCTs), PRISMA (systematic reviews), ARRIVE (animal research), STROBE (observational studies), MIAME/MINSEQE (omics)?
- Are all elements of the relevant checklist addressed?
Stage 4: Figure and Data Presentation
Quality checks
- High resolution, clearly labeled axes with units?
- Error bars defined (SD, SEM, CI)?
- Color schemes colorblind-accessible?
- Scale bars included for images?
- Figures stand alone with legends?
Integrity checks
- Signs of image manipulation (duplication, splicing)?
- Western blots/gels appropriately presented?
- Representative images truly representative?
- All conditions shown (no selective presentation)?
Stage 5: Ethical Considerations
Human subjects — IRB approval documented; informed consent described; vulnerable populations protected; conflicts of interest disclosed.
Animal research — IACUC approval documented; 3Rs (replacement, reduction, refinement) considered; procedures humane and justified.
Research integrity — Data fabrication/falsification concerns; authorship appropriate; competing interests disclosed; funding disclosed; plagiarism or duplicate publication concerns.
Stage 6: Special Considerations by Manuscript Type
Original research — Emphasize rigor, reproducibility, novelty; verify conclusions are data-driven; check complete methods and appropriate controls.
Reviews and meta-analyses — Comprehensiveness of literature coverage; search strategy and inclusion/exclusion criteria; systematic approach; critical analysis vs. mere summarization; for meta-analyses, assess statistical approach and heterogeneity.
Methods papers — Validation and comparison to existing methods; reproducibility; sufficient detail for implementation.
Preprints — Recognize pre-peer-review status; still apply rigorous standards; constructive feedback to help before journal submission.
Presentations and slide decks — NEVER read the PDF directly. Convert to images first using any PDF-to-image tool available (e.g., pdftoppm, ImageMagick's convert, or a Python script using pdf2image). Inspect each slide image sequentially. Document issues with specific slide numbers.
Presentation-specific criteria: text ≥18pt; contrast ≥4.5:1; no overflow or element overlaps; one main idea per slide; ≤6 bullets per slide; results dominate (40-50%); ~1 slide per minute; narrative arc (intro → methods → results → discussion); citations in intro and discussion.
Shared Frameworks
These apply in both modes.
Bias Detection
Cognitive biases (researcher)
- Confirmation bias: Only supporting findings highlighted?
- HARKing: Hypotheses stated a priori or formed after seeing results?
- Publication bias: Negative results missing from literature?
- Cherry-picking: Evidence selectively reported?
- Check for preregistration and analysis plan transparency.
Selection biases
- Sampling bias: Is sample representative of target population?
- Volunteer bias: Do participants self-select systematically?
- Attrition bias: Is dropout differential between groups?
- Survivorship bias: Only "survivors" visible in sample?
Measurement biases
- Observer bias: Could expectations influence observations?
- Recall bias: Are retrospective reports systematically inaccurate?
- Social desirability: Responses biased toward acceptability?
- Instrument bias: Do measurement tools systematically err?
Analysis biases
- P-hacking: Multiple analyses until significance emerged?
- Outcome switching: Non-significant outcomes replaced with significant ones?
- Selective reporting: All planned analyses reported?
- Subgroup fishing: Subgroup analyses without correction?
- Check study registration and compare to published outcomes.
Confounding
- What variables could affect both exposure and outcome?
- Were confounders measured and controlled?
- Could unmeasured confounding explain findings?
- Plausible alternative explanations?
Statistical Analysis Evaluation
- Sample size and power — A priori power analysis conducted? Adequate for detecting meaningful effects? Underpowered studies (common) raise flags for inflated effect sizes in significant results.
- Statistical tests — Appropriate for data type and distribution? Assumptions checked? Parametric vs. non-parametric choice justified? Analysis matched to design (paired vs. independent)?
- Multiple comparisons — Multiple hypotheses tested? Correction applied (Bonferroni, FDR)? Primary outcomes distinguished from exploratory? Could findings be false positives?
- P-value interpretation — Interpreted correctly? Non-significance incorrectly treated as "no effect"? Statistical significance conflated with practical importance? Suspicious clustering just below .05?
- Effect sizes and CIs — Effect sizes reported alongside significance? CIs show precision? Effect size meaningful in practical terms?
- Missing data — How much missing? Mechanism considered (MCAR, MAR, MNAR)? Handling method (deletion, imputation, maximum likelihood) appropriate?
- Regression and modeling — Overfitting? Extrapolation outside data range? Multicollinearity addressed? Model assumptions checked?
- Common pitfalls — Correlation as causation; regression to the mean ignored; base rate neglect; Texas sharpshooter; Simpson's paradox.
Evidence Quality Assessment
Study design hierarchy (highest to lowest for intervention effects):
- Systematic reviews / meta-analyses
- Randomized controlled trials
- Cohort studies
- Case-control studies
- Cross-sectional studies
- Case series / reports
- Expert opinion
Higher-level designs are not always better quality. A well-designed observational study can be stronger than a poorly-conducted RCT.
GRADE framework
- Start: RCT = high; observational = low
- Downgrade for: risk of bias; inconsistency across studies; indirectness (wrong population/intervention/outcome); imprecision (wide CIs, small samples); publication bias
- Upgrade for: large effect sizes; dose-response relationships; confounders would reduce (not increase) the effect
Convergence of evidence
- Stronger: multiple independent replications; different research groups and settings; different methodologies converging; mechanistic and empirical evidence aligned
- Weaker: single study or group; contradictory findings; publication bias evident; no replication attempts
Contextual factors — Biological/theoretical plausibility; consistency with established knowledge; temporality (cause precedes effect); specificity; strength of association.
Logical Fallacy Identification
Causation fallacies — Post hoc ergo propter hoc; correlation = causation; reverse causation; single cause fallacy.
Generalization fallacies — Hasty generalization; anecdotal fallacy; cherry-picking; ecological fallacy (group patterns applied to individuals).
Authority and source fallacies — Appeal to authority (without evidence); ad hominem; genetic fallacy; appeal to nature.
Statistical fallacies — Base rate neglect; Texas sharpshooter; multiple comparisons not corrected; prosecutor's fallacy (confusing P(E|H) with P(H|E)).
Structural fallacies — False dichotomy; moving goalposts; begging the question; straw man.
Science-specific fallacies — Galileo gambit; argument from ignorance; nirvana fallacy; unfalsifiability.
When identifying fallacies: name the specific fallacy; explain why the reasoning is flawed; identify what evidence would be needed for valid inference; note that fallacious reasoning doesn't prove the conclusion false — just that this argument doesn't support it.
Manuscript Review Report Format
Summary Statement
1-2 paragraphs: brief synopsis; overall recommendation (accept / minor revisions / major revisions / reject); 2-3 key strengths; 2-3 key weaknesses; bottom-line assessment.
Major Comments
Numbered sequentially. Each entry: (1) state the issue, (2) explain why it's problematic, (3) suggest specific solutions, (4) indicate if essential for publication.
Typical major comments: fundamental methodological flaws; inappropriate statistical analyses; unsupported conclusions; missing critical controls; serious reproducibility concerns; ethical concerns.
Minor Comments
Numbered sequentially with specific location (section, paragraph, figure). Typical: unclear figure labels, missing methodological details, typos, suggestions for improved data presentation, requests for clarification.
Questions for Authors
Methodological details that are unclear; seemingly contradictory results; missing information needed to evaluate; requests for additional data or analyses.
Tone and Approach
- Constructive: frame criticism as opportunities for improvement
- Specific: concrete examples and actionable suggestions
- Balanced: acknowledge strengths as well as weaknesses
- Respectful: authors have invested significant effort
- Objective: focus on the science, not the scientists
Avoid: personal attacks; sarcasm; vague criticism without examples; demanding unnecessary experiments beyond scope; applying stricter standards to results you dislike.
Research Design Guidance
When helping plan or improve a study:
- Research question — Specific, answerable, falsifiable? Addresses a gap? Feasible? Variables operationally defined?
- Design selection — Matched to question (causal → experimental; associational → observational)? Feasibility and ethics considered?
- Bias minimization — Randomization where possible; blinding at all feasible levels; confounders identified and controlled; procedures standardized.
- Sample planning — A priori power analysis; account for attrition; clear inclusion/exclusion criteria; representativeness.
- Measurement strategy — Validated, reliable instruments; objective measures preferred; multiple measures of key constructs; inter-rater reliability procedures.
- Analysis planning — Prespecify all hypotheses and analyses; designate primary outcome; plan assumption checks; specify missing data handling; plan to report effect sizes and CIs; consider multiple comparison corrections.
- Transparency — Preregister study and analysis plan; use reporting guidelines (CONSORT, STROBE, PRISMA); plan to report all outcomes; distinguish confirmatory from exploratory; commit to data/code sharing.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most research analysis skills give in ~3.0k tokens
Counted across 1,063 of the 1,754 authors here whose files we hold, read 2026-08-07
- Generate a markdown reportin 32 of 1063, across 23 files
- Cite each claim's sourcein 30 of 1063, across 15 files
- Define the ideal customer profilein 20 of 1063, across 2 files
- Search for companies matching the criteriain 20 of 1063, across 2 files
- Assign a fit score from one to tenin 20 of 1063, across 2 files
- Analyze the codebase to understand the productin 19 of 1063, across 1 file
- Ask clarifying questions about the value propositionin 19 of 1063, across 1 file
- Look for signals of immediate needin 19 of 1063, across 1 file
- Identify the target decision maker rolein 19 of 1063, across 1 file
- Suggest a personalized contact strategyin 19 of 1063, across 1 file
- Provide conversation starters for outreachin 19 of 1063, across 1 file
- Format results in a scannable markdown templatein 19 of 1063, across 1 file
Said here and by no other author read
- identify the claim type and strength
- evaluate statistical methods and assumptions
- apply GRADE framework to assess evidence
- check for researcher and selection biases
- verify reproducibility and data availability
- inspect images for manipulation
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.