Batch analyze
Skill Baikodis/claude-knowledge-base-skills/skills/batch-analyze
Claude Code skills that turn research and raw data into an LLM-ready knowledge base: source-discovery → corpus-ingest → wiki-build → wiki-lint (+ batch-analyze). Installable plugin marketplace.
npx -y skills add Baikodis/claude-knowledge-base-skills --skill batch-analyzeAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Multi-agent data analysis — auto-chunking, parallel agents, assembly with validation
SKILL.md
8.5 KB, ~2.4k tokens by cl100k_base, as published. Nobody here has run it
/batch-analyze — Multi-agent data analysis
When to use
- Processing a file with 50+ records where each needs LLM logic
- Extraction, verification, classification, enrichment, summarization
- Any "for each of N records, do X" task
Modes
- extract — from raw data → structured output
- verify — check an existing table against a source
PHASE 0: Understand the task (BEFORE any calculations)
1. Determine the INPUT (automatically):
- What file is it? Auto-detect the format (JSON/CSV/PDF/txt/dir)
- PDF → how many pages? Scanned or text?
- JSON/CSV → how many records?
- Directory → how many files, what formats?
2. Sample from DIFFERENT places (silently):
- 3 records: start, middle, end of the file
- If PDF: 3 pages (1, N/2, N)
- Estimate the size spread (uniform or not?)
3. Show the user and ask:
📂 File: [path] ([format], [N records/pages])
Here's what the data looks like:
[record 1 — brief, 3-5 lines]
[record 2 — brief, 3-5 lines]
Questions:
- What do you want to extract/verify/classify?
- What format should the output be saved in? (CSV/JSON/...)
- Are there records that should be skipped?
- Show one example of an ideal result row
STOP. Wait for an answer. Without one — don't calculate, don't launch.
If the task is already obvious from the invocation (the user provided the file + task + format) — skip the questions and move to Phase 1.
PHASE 1: Recon and calculation
Sampling (from 3 different places):
# Take records from the start, middle, and end
samples = [data[0], data[len(data)//2], data[-1]]
sample_tokens = [len(json.dumps(s, ensure_ascii=False)) / 3.5 for s in samples]
avg_record_tokens = mean(sample_tokens)
max_record_tokens = max(sample_tokens)
# If the spread is > 3x — warn
if max_record_tokens / avg_record_tokens > 3:
print("⚠️ Records differ significantly in size")
Calculation:
CONST:
agent_usable = 80_000 tokens
prompt_overhead = 2_000
output_per_record = 50
CALCULATION:
# Capacity by context
capacity = (agent_usable - prompt_overhead) / (avg_record_tokens + output_per_record)
# Safe chunking: divide by 2
safe_per_agent = capacity / 2
# Bounds
if safe_per_agent > 300: safe_per_agent = 300 # cap: too many = loss of attention
if safe_per_agent < 10: warn about a large number of rounds
total_agents = ceil(total_records / safe_per_agent)
rounds = ceil(total_agents / 12)
For PDF (separate branch):
# Determine the type: text or scan
# Sample 3 pages → tokens_per_page
# pages_per_agent = (agent_usable - prompt) / (avg_page_tokens + output) / 2
# For scans: first OCR → text, then the standard pipeline
Show the user:
📊 Task analysis
- Type: [extract/verify]
- Format: [JSON/CSV/PDF/dir]
- Records: N (~X tokens total)
- Average record: ~Y tokens (spread: min Z — max W)
⚙️ Resources
- Agents: K (M records each)
- Rounds: R
- Estimate: ~3-5 min
📝 Agent prompt:
[FULL template]
📤 Output example:
[2-3 rows]
🔍 Validation: spot-check 5 records after assembly
Proceed? ✔ or ✖
STOP. Wait for confirmation.
PHASE 2: Preparation
mkdir -p /tmp/batch_chunks /tmp/batch_results
rm -f /tmp/batch_chunks/* /tmp/batch_results/*
Chunking by data type:
JSON:
import json, os
data = json.load(open(SOURCE_FILE))
per_chunk = len(data) // K
for i in range(K):
start = i * per_chunk
end = start + per_chunk + (len(data) % K if i == K-1 else 0)
with open(f'/tmp/batch_chunks/chunk_{i+1:02d}.json', 'w') as f:
json.dump(data[start:end], f, ensure_ascii=False)
CSV:
import csv
rows = list(csv.DictReader(open(SOURCE)))
# similarly split by per_chunk
PDF:
# Split by pages: pages [0:M], [M:2M], ...
# Each chunk = a separate PDF or a text file with the extracted text
Directory of files:
# Distribute files across folders: chunk_01/, chunk_02/, ...
ls /tmp/batch_chunks/ # confirm creation
PHASE 3: Launch — CRITICAL
Rules (absolute, no exceptions):
- ALL agents in a round go in ONE tool-call message — not one at a time
- If there is more than 1 round — start the next one ONLY after confirming the artifacts of the previous one
- Agent model:
sonnet(faster, sufficient for extraction) - Every agent MUST save a file to
/tmp/batch_results/ - If fewer than K agents were launched — IMMEDIATELY tell the user
Agent prompt template (extract):
INPUT: /tmp/batch_chunks/chunk_NN.json
TASK: {task from the user — full prompt with rules}
OUTPUT: Save to /tmp/batch_results/result_NN.csv
- Header: {columns}
- Encoding: UTF-8
- Quote all text fields
RULES:
- Process EVERY record in the file
- If a record doesn't fit the task — skip it
- At the end print: "Done: processed X, extracted Y, skipped Z"
- Save file BEFORE printing summary
Agent prompt template (verify):
INPUT: /tmp/batch_chunks/verify_NN.json
Format: array of {csv_row: {...}, source: {...}}
TASK: For each pair, compare csv_row fields against source text.
Report ONLY errors.
OUTPUT: Save to /tmp/batch_results/errors_NN.csv
Header: link,field,was,should_be,reason
Only write rows with actual errors.
At the end: "Checked: X, errors: Y"
PHASE 4: Verification
4a. Proof of work (MANDATORY):
ls /tmp/batch_results/result_*.csv | wc -l # must = K
wc -l /tmp/batch_results/result_*.csv # lines per file
Show the user:
- File | Lines | Status
- If files < K — name the missing ones, relaunch
4b. Spot-check (5 random, MANDATORY):
import random, json, csv
data = json.load(open(SOURCE))
results = {}
for i in range(1, K+1):
for row in csv.DictReader(open(f'/tmp/batch_results/result_{i:02d}.csv')):
results[row[KEY_FIELD]] = row
samples = random.sample(data, 5)
for s in samples:
key = s[KEY_FIELD]
print(f"SOURCE: {s}")
print(f"RESULT: {results.get(key, 'NOT FOUND')}")
Show the 5 pairs to the user.
Criteria:
- 5/5 correct → assembly
- 4/5 → continue, mark as ~80%
- ≤3/5 → STOP. Show the errors, adjust the prompt
PHASE 5: Assembly
head -1 /tmp/batch_results/result_01.csv > FINAL_OUTPUT.csv
for f in /tmp/batch_results/result_*.csv; do
tail -n +2 "$f" >> FINAL_OUTPUT.csv
done
wc -l FINAL_OUTPUT.csv
Statistics:
import csv
rows = list(csv.DictReader(open('FINAL_OUTPUT.csv')))
print(f"Total: {len(rows)}")
for col in fieldnames:
filled = sum(1 for r in rows if r[col])
print(f" {col}: {filled}/{len(rows)} ({filled*100//len(rows)}%)")
dupes = len(rows) - len(set(r[KEY] for r in rows))
print(f"Duplicates: {dupes}")
Encoding (CSV for Excel/Sheets):
printf '\xEF\xBB\xBF' > FINAL_bom.csv
cat FINAL_OUTPUT.csv >> FINAL_bom.csv
Delivery:
Deliver the final file to the user via your usual file-delivery mechanism (e.g. an upload endpoint or a shared directory).
GUARANTEES (anti-fabrication)
Absolute, no exceptions:
- Results come ONLY from tool results. Not in the output = doesn't exist
- After the agents —
lsandwc -lthe real files - Spot-check to the user BEFORE saying "done"
- Files < K — say so directly, relaunch
- Numbers come only from files (wc -l, len()), not from generation
- Don't append "expected" results
SPECIAL CASES
Data not in a file (API, DB)
Export → /tmp/batch_source.json → standard pipeline
PDF with scans
OCR (Tesseract/Vision) → text → standard pipeline. Batch by pages, not by records.
Task needs context BETWEEN records
Two-pass: pass 1 = extract keys (in parallel), pass 2 = compare (1 agent)
Vague prompt
Show 3 records → "show an example of a result row" → do NOT launch without specifics
Spot-check > 20% errors
STOP → show the errors → refine the prompt / add few-shot examples / reduce batch size
Agent crashed / empty file
Relaunch ONLY that one. Crashed again → show the error to the user
EXAMPLES
/batch-analyze
File: ./data/messages.json
Task: extract name, position, company, city
Output: CSV
/batch-analyze verify
Source: ./data/raw.json
Table: ./data/contacts.csv
Join on: link
Verify: name, position, company, city
/batch-analyze
File: ./docs/report.pdf (120 pages, scans)
Task: extract all tables with numeric data
Output: CSV
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.