agentsclimarketplace

Research init

Skill brycewang-stanford/Auto-Empirical-Research-Skills/skills/61-phdemotions-research-methods/skills/research-init

Scaffold a new research project with full reproducibility infrastructure in R and/or Python. Creates directory structure, pipeline stubs (targets/Snakemake), environment lockfiles (renv/uv), documentation templates (codebook, decision log, pre-registration, Cornell README), Quarto manuscript template, and proper .gitignore. Can wrap existing data in gold-standard structure. Use when the user says "new project," "scaffold," "start a study," "set up a project," "I have data and need to organize it," or when /research-intake recommends scaffolding.From its SKILL.md

Install
npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill research-init

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

SKILL.md

3.3 KB, 643 tokens by cl100k_base, as published. Nobody here has run it

/research-init — Project Scaffolding

You create the structure that makes everything else possible. A well-scaffolded project is halfway to reproducibility before a single line of analysis is written.

How to scaffold a project

Step 1 — Gather requirements

Ask the researcher (or infer from context):

  1. Project name — will become the directory name
  2. Language — R, Python, or both? (Default: both)
  3. Existing data? — If yes, where? What format? This changes the workflow.
  4. Research question — one sentence, for the README and pre-registration skeleton
  5. Target journal — if known, for formatting defaults

If the researcher provides a project name and says "scaffold it," don't over-ask. Use sensible defaults and get them started.

Step 2 — Create directory structure

Create the full structure documented in references/criteria.md. Use the templates in references/templates/ for each file.

Step 3 — Initialize environments

For R:

  1. Create _targets.R from template
  2. Initialize renv (if R is available on the system)
  3. Create R/00_setup.R from template

For Python:

  1. Create Snakefile from template
  2. Create pyproject.toml with research stack dependencies
  3. Create python/00_setup.py from template

Step 4 — Handle existing data

If the researcher has existing data:

  1. Copy (not move) files to data/raw/
  2. Set data/raw/ as conceptually read-only (the raw-data-guard hook enforces this)
  3. Note the original file locations in the README provenance section
  4. Suggest running /data-validate next

Step 5 — Initialize git

If not already in a git repo:

  1. Create .gitignore from template
  2. git init
  3. Create initial commit with structure (but NOT data files — those go in .gitignore or are tracked separately)

Step 6 — Print summary and next steps

Show the researcher what was created and suggest next steps per _shared/next-steps.md.

Principles

Read references/principles.md for the foundational principles behind every scaffolding decision.

Voice

Efficient and organized. You're setting up a workspace, not giving a lecture. Create the structure, explain what each piece is for briefly, and get the researcher moving. Show the directory tree at the end so they can see what was built.

Argument handling

  • research-init my-study → creates ./my-study/
  • research-init my-study --lang r → R only
  • research-init my-study --existing-data ~/data/survey.csv → copies data to data/raw/
  • research-init (no args) → asks for project name

What ships with it: 13 files

28.6 KB alongside SKILL.md, 1 of them executable

Gives 1 of the 12 instructions most project setup skills give in 643 tokens

Counted across 1,553 of the 3,091 authors here whose files we hold, read 2026-09-06

  • Write the configuration filein 36 of 1553
  • Create the directory structurehere, and in 35 of 1553, across 33 files
  • Verify the setupin 31 of 1553, across 28 files
  • Run the setup scriptin 30 of 1553, across 29 files
  • Pre-determine the required sample sizein 29 of 1553, across 12 files
  • Check if the configuration already existsin 29 of 1553
  • Document every testin 26 of 1553, across 10 files
  • Start with a hypothesisin 26 of 1553, across 11 files
  • Ask one question at a timein 22 of 1553
  • Test a single variable per testin 21 of 1553, across 9 files
  • Read product marketing context before asking questionsin 19 of 1553, across 8 files
  • Do not peek and stop earlyin 18 of 1553, across 7 files

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.