Github trending analyzer
Crawl GitHub trending repositories, analyze with LLM for Chinese insights, categorize by themes, compute diffs against history, and generate Markdown reports. Default brief mode stops at trend analysis; optional detailed mode appends per-project analysis. Supports incremental gap-filling and selective re-analysis with caching.From its SKILL.md
npx -y skills add Dianel555/DSkills --skill github-trending-analyzerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- fetches URLsInstructs the agent to fetch 1 URL, including https://github.com/trending[/{language}]?since={daily|weekly|monthly}.
SKILL.md
8.7 KB, ~2.1k tokens by cl100k_base, as published. Nobody here has run it
GitHub Trending Analyzer
A workflow protocol for tracking GitHub trending repositories with LLM-powered analysis. Fetches trending projects, enriches each with structured Chinese insights (what/analogy/help/who), classifies by themes, compares against historical snapshots, and generates reports in two modes — a compact brief (default) or a detailed report with per-project analysis (opt-in).
Trigger Signals
- GitHub trending analysis
- Weekly tech trend report
- Repository discovery automation
- Incremental analysis refresh
- Theme-based repo categorization
Preconditions
- HTTP access to github.com/trending (no auth required for public trending)
- LLM backend capable of JSON-structured output (for the 4-field analysis schema)
- File system access for memory cache and report output
- HTML parsing capability (regex or DOM parser)
Strategy
Run the five-step pipeline in order.
Step 1: Fetch trending HTML
Construct the URL with time range and optional language filter:
https://github.com/trending[/{language}]?since={daily|weekly|monthly}
Fetch with a browser User-Agent to avoid bot detection. Parse the HTML to extract:
name(org/repo)url(full GitHub link)desc(one-line description from the page)lang(primary language)stars(total stargazers count)today_stars(increment for this period)
Regex patterns (reference from source):
- Project name:
<h2[^>]*>.*?<a href="/([^"]+)" - Description:
<p class="[^"]*col-9[^"]*"[^>]*>\s*(.*?)\s*</p> - Language:
<span itemprop="programmingLanguage">([^<]+)</span> - Stars: parse from
/stargazerslink text after stripping HTML tags - Today increment:
([\d,]+)\s*stars?\s*(?:this|today)(case-insensitive)
Step 2: Batch LLM analysis
For each batch of 5 projects (to avoid token limits), send this prompt to your LLM:
Analyze the following {N} GitHub Trending projects. Output strict JSON array.
Each project needs 4 fields:
- what: What it is (≤30 Chinese characters)
- analogy: Life analogy (one sentence)
- help: What it helps you do (2 items, each ≤40 chars, array)
- who: Who needs it (one sentence, ≤30 chars)
Project list:
1. org/repo (Language) — description...
2. ...
Output ONLY the JSON array, no other text. Example:
[{"name":"org/repo","what":"...","analogy":"...","help":["...","..."],"who":"..."}]
Parse the response:
- Strip markdown code fences (
```json/```) - Clean trailing commas:
,\s*([\]}])→\1 - Extract the JSON array via regex:
\[.*\](DOTALL) - Decode with
json.loads()or equivalent - Match results back to projects by name suffix (case-insensitive)
Fallback: If array parsing fails, extract individual objects via bracket-counting and parse one by one.
Deep mode (optional): Use longer limits (what ≤50 chars, help 3 items) for richer analysis.
Step 3: Theme classification
Load the bundled theme_rules.json. For each project:
- Concatenate
name + " " + descand lowercase - Iterate themes by priority order
- Check if any keyword from the theme appears in the text
- Assign to first matching theme
- Default to "🌐 其他" if no match
Result: {theme_name: [projects...]} dictionary.
Step 4: Compute diff (optional)
If you maintain a memory cache (JSON file storing past runs):
[
{
"date": "2026-06-19",
"since": "weekly",
"lang": "python",
"repos": [{"name":"...", "url":"...", "desc":"...", "lang":"...", "stars":..., "today_stars":..., "analysis":{...}}]
}
]
Compare current repos against the latest entry with the same since value:
- new: projects in current but not in last
- hot: projects in both
- dropped: projects in last but not in current
- last_date: baseline timestamp
Step 5: Generate reports
Two report modes, driven by the bundled templates:
- Brief (default):
report_template_brief.md— stops at "💡 Trend Analysis". Always emitted. - Detailed (opt-in):
report_template_detailed.md— the brief content plus a per-project "📋 Project Details" section with the 4-field analysis. Emitted only when the user asks for detail (or whendeepanalysis was run).
Trend insight prompt (used in the "Trend Analysis" section of both modes):
基于以下GitHub Trending项目摘要,用3-5句话分析当前最强技术趋势和驱动力:
{list of "name: what" for all projects}
Save to date-stamped files (e.g. trending_briefing_2026-06-19.md, trending_detailed_2026-06-19.md). Overwrite on same-day re-runs.
Empty tables: when a section (new/hot/dropped) has no rows, render the table header followed by a single *none* row; keep "Theme Breakdown" and "Trend Analysis" only if there are classified projects. On a first run (no memory baseline), omit the "Dropped Off" section rather than showing it empty.
Constraints
Core rules
- Batch size = 5 for LLM calls to avoid truncation. For 20 repos, make 4 separate calls.
- JSON-only LLM output. The prompt explicitly forbids explanatory text. Parse defensively (strip fences, clean commas).
- Name matching is fuzzy. Match by suffix (
org/repovsrepo) and case-insensitive substring. - Theme priority matters. A project matching both "AI" and "Dev Tools" gets classified as "AI" (priority 1 < 4).
- Memory is append-only list. Each run appends one entry. Keep last 30 to prevent unbounded growth.
Incremental modes (optional)
- Gap-fill mode: Load the latest memory entry → detect repos without
analysisfield → re-run LLM only for those → merge back → regenerate reports. - Selective re-analysis: User specifies project names (comma-separated, partial match) → find matching repos in memory → re-run LLM with optional deep mode → update memory → regenerate reports.
Implementation hint: detect_gaps(repos) returns [r for r in repos if not r.get('analysis')].
Error handling
- HTML fetch fails: Retry once with 5s delay, then abort with clear error message.
- LLM returns non-JSON: Log warning, continue with raw description as fallback for that batch.
- Memory file missing: Treat as first run (no diff section in reports).
Output Protocol
Emit Markdown files to a reports directory, based on mode:
- Brief (default) (
trending_briefing_{date}.md): sections new/hot/dropped/themes + trend insight. Stops at "Trend Analysis" — no per-project blocks. - Detailed (opt-in) (
trending_detailed_{date}.md): brief content followed by one "📋 Project Details" block per project with the 4-field analysis. Only when the user requests detail.
Overwrite if file exists (same-day re-runs replace prior reports).
Console output during execution:
- "Fetching {since} trending..." → "Got {N} projects"
- "LLM batch {i}/{total}..." → "✅ Batch complete: {n} items"
- "📄 Brief saved: {path}"
- "📄 Detailed saved: {path}" (only when detailed mode runs)
- (Gap-fill) "Coverage: {covered}/{total} ({pct}%)"
Validation
Before emitting reports, confirm:
- All repos have
name,url,desc,lang,stars,today_starsfields. - At least one theme contains projects (not all "其他").
- LLM analysis covers ≥50% of projects (log warning if lower).
- Both report files are valid UTF-8 Markdown.
- Memory JSON is valid (can be reloaded without error).
Adapting and Extending
Custom themes
Edit the bundled theme_rules.json:
- Add new themes with emoji prefix and priority
- Extend keyword lists for existing themes
- Adjust priority order to prefer certain classifications
Alternative LLM schemas
The 4-field schema (what/analogy/help/who) is optimized for Chinese tech audiences. Adapt for other contexts:
- English reports: Change field names and prompt language
- Different insights: Replace "analogy" with "use cases" or "risks"
- Richer detail: Increase char limits in deep mode
Different trending sources
The HTML parsing patterns are GitHub-specific. To adapt for other platforms (Hacker News, Product Hunt):
- Replace Step 1 fetch logic
- Adjust regex patterns for that site's DOM structure
- Keep Steps 2-5 unchanged (LLM + themes + diff + reports)
Memory backends
The reference uses local JSON. For multi-agent or cloud deployments:
- Swap
load_memory()/save_memory()with a DB or object storage client - Maintain the same list-of-dicts schema
- Add concurrency locks if multiple agents run in parallel
What ships with it: 3 files
5.3 KB alongside SKILL.md
- report_template_brief.md1.4 KB
- report_template_detailed.md1.7 KB
- theme_rules.json2.1 KB
Gives 0 of the 12 instructions most research analysis skills give in ~2.1k tokens
Counted across 1,213 of the 2,113 authors here whose files we hold, read 2026-09-06
- Cite sources for every important claimin 47 of 1213, across 38 files
- Separate facts from inferences and recommendationsin 21 of 1213, across 12 files
- Write findings to a markdown filein 19 of 1213
- Label every insight with a confidence levelin 18 of 1213, across 8 files
- Read product marketing context before asking questionsin 18 of 1213, across 8 files
- Rank themes by frequency and intensityin 16 of 1213, across 6 files
- Establish research mode before proceedingin 16 of 1213, across 6 files
- Segment survey responses by customer tier or tenurein 16 of 1213, across 6 files
- Categorize support tickets before analyzingin 16 of 1213, across 6 files
- Weight research sources from the last twelve monthsin 16 of 1213, across 6 files
- Use at least five data points per segmentin 15 of 1213, across 5 files
- Extract verbatim quotes for all research findingsin 15 of 1213, across 5 files
Said here and by no other author read
- Fetch trending repositories using a browser User-Agent
- Process projects in batches of five for LLM analysis
- Strip markdown fences and clean commas from LLM output
- Classify projects by theme using keyword matching
- Compare current repositories against historical snapshots
- Generate reports in brief or detailed mode
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.