agentsclimarketplace

Research aggregator

Skill diamitani/rostr-agent/skills/research/research-aggregator

Autonomous research aggregator: web search → find articles → catalog to spreadsheet with download links → sync to storage → maintain long.md with transcribed content. Handles thousands of pages of accumulated research.From its SKILL.md

Install
npx -y skills add diamitani/rostr-agent --skill research-aggregator

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

6.6 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it

Research Aggregator

Autonomous research collection pipeline: search → extract → catalog → store → transcribe.

Use this skill when you need to:

  • Search the web for articles on a topic
  • Catalog results in a spreadsheet with metadata
  • Download/sync content to local storage
  • Maintain a comprehensive long.md file with transcribed content (can grow to thousands of pages)

1. Pipeline Overview

Query → Web Search → Extract Articles → Catalog (CSV/JSON) → Storage Sync → long.md Update

Output Artifacts

ArtifactPurposeFormat
catalog.csvSearchable index of all sourcesCSV spreadsheet
catalog.jsonMachine-readable catalogJSON
storage/Downloaded article contentMarkdown files
long.mdComplete transcribed contentSingle markdown file

2. Catalog Schema (Spreadsheet)

id,title,url,domain,tier,date_published,date_added,topics,summary,storage_path,word_count,status
uuid,Article Title,https://...,arxiv.org,1,2025-07-01,2025-07-23,"ai,llm,research",3-sentence summary,storage/uuid.md,2500,transcribed

Fields

FieldTypeDescription
idUUIDUnique identifier
titleStringArticle title
urlURLSource URL
domainStringSource domain
tier1/2/3Credibility tier (RAG DAL)
date_publishedISO DateWhen published
date_addedISO DateWhen added to catalog
topicsComma-sepTopic tags
summaryString3-5 sentence summary
storage_pathPathLocal file path
word_countIntContent length
statusEnum`pending

3. Execution Workflow

Step 1: Initialize Research Session

mkdir -p research/{storage,exports}
touch research/catalog.csv research/catalog.json research/long.md

Step 2: Web Search Phase

from hermes_tools import web_search

queries = [
    "topic name",
    "topic site:arxiv.org",
    "topic research paper 2025",
    "topic case study",
]

all_results = []
for q in queries:
    r = web_search(q, limit=10)
    all_results.extend(r.get('data', {}).get('web', []))

# Dedupe by URL
seen = set()
unique = [r for r in all_results if r['url'] not in seen and not seen.add(r['url'])]

Step 3: Extract and Catalog

from hermes_tools import web_extract
import uuid
from datetime import date

TIER_1 = ['arxiv.org', 'pubmed', '.gov', '.edu', 'acm.org', 'ieee.org', 'nature.com']
TIER_2 = ['techcrunch.com', 'wired.com', 'reuters.com', 'bbc.com', 'nytimes.com']

def classify_tier(url):
    url_lower = url.lower()
    if any(d in url_lower for d in TIER_1): return 1
    if any(d in url_lower for d in TIER_2): return 2
    return 3

catalog = []
for r in unique:
    catalog.append({
        'id': str(uuid.uuid4())[:8],
        'title': r.get('title', 'Untitled'),
        'url': r['url'],
        'domain': r['url'].split('/')[2],
        'tier': classify_tier(r['url']),
        'date_added': str(date.today()),
        'status': 'pending'
    })

Step 4: Extract Content to Storage

for i in range(0, len(catalog), 5):
    batch = catalog[i:i+5]
    urls = [e['url'] for e in batch]
    extracted = web_extract(urls)
    
    for entry, result in zip(batch, extracted.get('results', [])):
        if result.get('error'):
            entry['status'] = 'failed'
            continue
        
        content = result.get('content', '')
        path = f"research/storage/{entry['id']}.md"
        write_file(path, f"# {entry['title']}\n\nSource: {entry['url']}\n\n---\n\n{content}")
        
        entry['storage_path'] = path
        entry['word_count'] = len(content.split())
        entry['status'] = 'extracted'

Step 5: Build long.md

extracted = [e for e in catalog if e['status'] == 'extracted']
total_words = sum(e['word_count'] for e in extracted)

long_md = f"""# Research: {TOPIC}

Generated: {date.today()}
Sources: {len(extracted)} | Words: {total_words:,}

---

## Table of Contents

"""
for e in extracted:
    long_md += f"- [{e['title']}](#{e['id']})\n"

long_md += "\n---\n\n"

for e in extracted:
    content = read_file(e['storage_path'])['content']
    long_md += f"## {e['title']} {{#{e['id']}}}\n\nTier {e['tier']} | {e['word_count']} words\n\n{content}\n\n---\n\n"

write_file('research/long.md', long_md)

4. Incremental Updates

# Load existing
with open('research/catalog.json') as f:
    existing = json.load(f)
existing_urls = {e['url'] for e in existing}

# Search for new
new_results = [r for r in search_results if r['url'] not in existing_urls]

# Append to catalog and long.md
existing.extend(new_entries)

5. Storage Sync Options

Local: research/storage/*.md

S3/R2: aws s3 sync research/ s3://bucket/research/

Git: cd research && git add . && git commit -m "Update"


6. Handling Thousands of Pages

When long.md grows large, split by chapters:

MAX_CHARS = 500_000  # ~100 pages per file
chapters = []
current = ""
for entry in catalog:
    content = read_file(entry['storage_path'])['content']
    if len(current) + len(content) > MAX_CHARS:
        chapters.append(current)
        current = ""
    current += f"\n\n## {entry['title']}\n\n{content}"
chapters.append(current)

for i, chapter in enumerate(chapters):
    write_file(f"research/long_part_{i+1}.md", chapter)

7. Quick Reference

Directory:     research/
Catalog:       catalog.csv, catalog.json
Content:       storage/*.md
Compiled:      long.md (or long_part_*.md for large)

Workflow:      Search → Extract → Catalog → Store → long.md
Incremental:   Load existing → Filter new → Append

8. Integration with RAG DAL

  1. RAG DAL — Multi-pass retrieval with credibility scoring (in-session)
  2. Research Aggregator — Persistent catalog + storage (cross-session)

Export RAG DAL results to catalog for future reference.

What ships with it: 1 file

8.0 KB alongside SKILL.md, 1 of them executable

scripts/

Keep looking

Skills are one crate of 326,736. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.