agentsclimarketplace

Sn search academic

Skill OpenSenseNova/SenseNova-Skills/skills/sn-search-academic

用于学术调研、论文精读、相关工作梳理、百科知识查询和引用链追溯。From its SKILL.md

Install
npx -y skills add OpenSenseNova/SenseNova-Skills --skill sn-search-academic

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

14.1 KB, ~4.9k tokens by cl100k_base, as published. Nobody here has run it

sn-search-academic - 学术搜索

凭证配置

API key、token 与 cookie 统一建议写在仓库根目录 .env(参考 .env.example),并由 runtime 或用户在执行前加载为同名环境变量。脚本仍只从环境变量或显式 CLI 参数读取凭证;不要把真实密钥写入 skill payload、报告、日志或提交。

使用三个统一入口完成学术调研:

  • search.py:搜索论文和百科条目
  • paper.py:列出论文章节,读取论文全文或指定章节
  • refTree.py:查询论文的 references 和 citations

不要直接调用历史 provider 脚本;它们只是统一入口的内部实现细节。 需要 provider 回退链、参数分发或完整输出字段时,按需读取 references/search.mdreferences/paper.mdreferences/refTree.md

可用脚本

脚本用途主要输入主要输出
scripts/search.py搜索论文/百科query,可选 --source--limit--category--lang按 source 分组的论文/百科条目,位于 source_results[*].items
scripts/paper.py列出章节,读取论文全文或章节论文 ID,可选 --source--list_section--section章节列表位于 sections;全文或章节正文位于 content
scripts/refTree.py查询引用树--paper_id--title,可选 --direction参考文献与被引论文,位于 source_results[*].references / source_results[*].citations

执行约定

本技能的 scripts/...requirements.txtreferences/... 路径均相对本 skill 目录;若当前工作目录不同,先解析为绝对路径,不要依赖 ${SKILL_DIR} 运行时变量。

调用约定:

  • 不要并行启动多个本技能脚本;search.pyrefTree.py 内部已经处理并发、超时和 provider 回退链。
  • 长结果优先加 --output <path> 写入文件,再读取必要字段,避免终端输出过长。
  • --provider-timeout 表示单个 provider 超时;默认使用脚本内置超时。

依赖

首次运行或脚本提示缺库时,使用本技能的依赖清单安装到当前 Python 环境:

python3 -m pip install -r requirements.txt

不要在脚本内部自动安装依赖。若安装失败、网络不可用或包不可用,停止使用对应脚本并改用 WebSearch/browser-use,说明缺少依赖。

Crawler 回退还需要额外运行时环境:

python3 -m playwright install firefox

arxiv_crawler_search.pysemantic_scholar_crawler_refTree.py 还需要 Node.js,以及某个当前目录或祖先目录中已安装 camoufox-jsnode_modules。缺少这些环境时,不要尝试绕过;改用非 crawler provider 或网页搜索。

参数说明

search.py

统一搜索入口。默认搜索所有支持的 source,并按 source 分组返回结果。

python3 scripts/search.py <query> [选项]
参数说明默认值
query搜索关键词,必填位置参数-
--source, --sources, -s搜索源;支持重复传参或逗号分隔all
--limit, -n每个 source 返回数量10
--category, -cArXiv 分类过滤,只传给支持分类的 source-
--lang, -l语言提示,只传给支持语言参数的 source-
--output, -o将最终 JSON 写入文件-
--provider-timeout每个 provider 的超时时间,单位秒;0 表示不限制60

支持的 --source

  • all
  • arxiv
  • semantic
  • google_scholar
  • pubmed
  • wikipedia

示例:

python3 scripts/search.py "retrieval augmented generation" --limit 5
python3 scripts/search.py "diffusion model" --source arxiv,semantic --category cs.CV --limit 5
python3 scripts/search.py "阿尔茨海默病 多模态诊断" --source pubmed,wikipedia --lang zh --limit 5
python3 scripts/search.py "agentic memory" --source all --limit 8 --output results/search.json

paper.py

统一论文阅读入口。默认按 arXiv 论文读取;读取 PMC 论文时显式传 --source pmc。不确定章节名时先用 --list_section 列出可用章节,再用 --section 精读。

python3 scripts/paper.py <id> [选项]
参数说明默认值
id论文 ID。arXiv 支持原始 ID、arXiv: 前缀、abs/pdf URL;PMC 支持 PMC1111914311119143、PMC URL-
--source论文来源:arxivpmcarxiv
--section, -s读取指定章节;不填则读取全文-
--list_section, --list-section列出论文可用章节,不返回正文;不能和 --section 同时使用false
--output, -o将最终 JSON 写入文件-

示例:

python3 scripts/paper.py 2603.00729
python3 scripts/paper.py 2603.00729 --list_section
python3 scripts/paper.py arXiv:2603.00729 --section introduction
python3 scripts/paper.py 2603.00729 --section method --output results/paper-method.json
python3 scripts/paper.py PMC11119143 --source pmc
python3 scripts/paper.py PMC11119143 --source pmc --list-section
python3 scripts/paper.py PMC11119143 --source pmc --section results

refTree.py

统一引用树入口。--paper_id--title 都必填;标题用于回退时精确匹配。

python3 scripts/refTree.py --paper_id <paper_id> --title <title> [选项]
参数说明默认值
--paper_id论文 ID:Semantic Scholar ID、DOI、ArXiv ID、PMID 等-
--title论文标题,必填-
--direction查询方向:referencescitations;不填则两者都查-
--source, --sources, -s引用树 source;当前支持 allsemanticall
--limit, -n每个 source、每个 direction 返回数量10
--api-keySemantic Scholar API 密钥,可选-
--provider-timeout每个 provider 的超时时间,单位秒;0 表示不限制60
--output, -o将最终 JSON 写入文件-

注意:参数名是 --paper_id,不是 --paper-idpaper_id 不支持位置参数。

示例:

python3 scripts/refTree.py --paper_id "2309.16609" --title "Qwen Technical Report"
python3 scripts/refTree.py --paper_id "2309.16609" --title "Qwen Technical Report" --direction references --limit 20
python3 scripts/refTree.py --paper_id "10.1038/s41586-024-07487-w" --title "AlphaFold 3" --direction citations
python3 scripts/refTree.py --paper_id "2309.16609" --title "Qwen Technical Report" --output results/refTree.json

输出格式

所有脚本都输出 JSON。先看顶层 success;失败时读取 errorerrorsattempts 判断是无结果、超时还是 provider 失败。

search.py 输出

CLI 输出的顶层不包含 items,论文条目在 source_results[*].items 中:

{
  "success": true,
  "query": "retrieval augmented generation",
  "provider": "search.py",
  "sources": ["arxiv", "semantic"],
  "source_results": [
    {
      "source": "arxiv",
      "success": true,
      "provider": "arxiv_official",
      "items": [
        {
          "source": "arxiv",
          "provider": "arxiv_official",
          "title": "Example title",
          "abstract": "Example abstract",
          "citation_count": null,
          "arxiv_id": "2301.00001",
          "url": "https://arxiv.org/abs/2301.00001"
        }
      ],
      "attempts": [],
      "error": null
    }
  ],
  "errors": [],
  "error": null
}

常用 item 字段:

  • 通用:titleabstractsnippeturlcitation_countdoi
  • arXiv:arxiv_idpdf_urlcategories
  • Semantic Scholar:paper_idvenueyear
  • PubMed:pmidpmc_idjournalpub_date
  • Wikipedia:page_idword_countsection_title

paper.py 输出

默认读取全文;指定 --section 时读取章节。正文在顶层 content

{
  "success": true,
  "source": "arxiv",
  "provider": "arxiv_html",
  "arxiv_id": "2603.00729",
  "section": "introduction",
  "content": "<全文或章节正文>",
  "char_count": 12345,
  "attempts": [],
  "error": null
}

指定 --list_section 时只返回章节结构,不返回 content

{
  "success": true,
  "source": "arxiv",
  "provider": "arxiv_html",
  "arxiv_id": "2603.00729",
  "section_count": 2,
  "sections": [
    {"name": "Abstract", "level": 0},
    {"name": "1 Introduction", "level": 1}
  ],
  "attempts": [],
  "error": null
}

常用字段:

  • arXiv:arxiv_idtitleabs_urlhtml_urlpdf_urlsection_countsections
  • PMC:pmc_idpmidtitlepmc_urlsection_countsections
  • 指定 --list_section 时返回 sectionssection_count,不包含 content
  • 指定 --section 时会包含 section;不指定 --list_section / --section 时读取全文

refTree.py 输出

引用树结果在 source_results[*].referencessource_results[*].citations

{
  "success": true,
  "id": "2309.16609",
  "title": "Qwen Technical Report",
  "provider": "refTree.py",
  "direction": "all",
  "source_results": [
    {
      "source": "semantic",
      "success": true,
      "provider": "semantic_official",
      "references": [
        {
          "title": "Example reference",
          "abstract": "Example abstract",
          "citation_count": 128,
          "paper_id": "example-reference-id",
          "arxiv_id": "2301.00001"
        }
      ],
      "citations": [
        {
          "title": "Example citing paper",
          "abstract": "Example abstract",
          "citation_count": 42,
          "paper_id": "example-citing-id",
          "doi": "10.1234/example"
        }
      ],
      "attempts": [],
      "error": null
    }
  ],
  "errors": [],
  "error": null
}

如果使用 --output,三个脚本都会在 JSON 中额外加入 output_path

并发与限流约定

这些脚本会访问外部学术服务,必须控制请求频率。 执行本技能脚本时:

  • 不要并发运行多个搜索脚本。
  • 不要使用并行工具同时调用多个 python3 scripts/... 命令。
  • 一次只运行一个脚本命令,等待结果返回后再运行下一个。
  • 批量查询时,优先使用脚本自带的 --limit--id-list 等参数,而不是启动多个进程。
  • 如果需要连续调用,按顺序执行,并在必要时等待数秒。

全文阅读工作流

搜索结果只有摘要时,用 paper.py 先列章节,再补充全文或关键章节。

  1. 先用 search.py 搜索,优先从 source_results[*].items 里记录 titlearxiv_idpmc_idpaper_iddoicitation_count
  2. 如果条目有 arxiv_id,先用 python3 scripts/paper.py <arxiv_id> --source arxiv --list_section 查看章节;再用 --section <section> 精读。
  3. 如果条目有 pmc_id,先用 python3 scripts/paper.py <pmc_id> --source pmc --list_section 查看章节;再按需读 --section <section>
  4. 如果需要整体理解,再不带 --section / --list_section 读取全文。
  5. 全文很长时使用 --output results/paper.json,再读取 contentsectionschar_count 等字段。

引用追溯工作流

通过论文的引用关系发现关键词搜索覆盖不到的相关工作。

通过 references 找奠基工作,通过 citations 找后续进展。refTree.py 需要同时传论文 ID 和标题。

后向追溯(找奠基工作)

  1. 关键词搜索找到高相关论文 → 取其 paper_idarxiv_idtitle
  2. refTree.py --paper_id "<id>" --title "<title>" --direction references --limit 20 → 找到高引参考文献
  3. 筛选与研究问题相关的条目 → 用 paper.py深入阅读

前向追踪(找后续进展)

  1. 找到领域奠基论文或关键论文 → 取其 ID
  2. refTree.py --paper_id "<id>" --title "<title>" --direction citations --limit 20 → 找到近期高引跟进工作
  3. 筛选与研究问题相关的条目 → 用 paper.py深入阅读

引用链:构建演化路径

  1. 从种子论文 A 出发 → backward 找到 A 的关键参考文献 B
  2. 从 B 出发 → forward 找到引用 B 的后续工作(可能发现 A 没引用的相关论文 C)
  3. 形成 B → A → ... 和 B → C → ... 的知识脉络

主工作流

严格遵循本工作流去执行学术搜索的全流程

  1. 在提供的学术平台选择所有可能的平台搜索学术文献
  2. 如果摘要不足或论文高度相关时,列出章节,尝试读取论文章节或全文,判断论文和搜索需求的相关性
  3. 选择相关性高的论文,搜索它的参考文献和被引(使用引用追溯工作流)。
  4. 选择引用树中 高引用的文献,执行步骤 2、步骤 3。
  5. 重复以上步骤,进行多轮搜索,尽可能多的进行搜索。
  6. 当文献数量、引用链和全文证据足够支撑回答时停止搜索,并在结论中说明主要依据。

ArXiv 分类速查

顶层领域可直接用(如 --category cs),子分类更精确(如 --category cs.AI)。

领域分类代码说明
计算机科学cs.AI人工智能
cs.LG机器学习
cs.CL计算语言学 / NLP
cs.CV计算机视觉
cs.IR信息检索
cs.RO机器人
cs.SE软件工程
cs.DC分布式/并行计算
cs.NI网络与互联网
cs.CR密码学与安全
cs.DB数据库
cs.HC人机交互
统计stat.ML统计机器学习
stat.AP应用统计
stat.ME统计方法论
数学math.OC优化与控制
math.ST统计理论
math.CO组合数学
物理physics物理(全类)
cond-mat凝聚态物理
quant-ph量子物理
hep-th高能理论物理
经济/金融econ.GN经济学综合
q-fin.CP计算金融
q-fin.ST统计金融
生物/医学q-bio.NC神经科学
q-bio.GN基因组学
q-bio.QM定量方法

What ships with it: 26 files

249.2 KB alongside SKILL.md, 22 of them executable

references/

Keep looking

Skills are one crate of 326,144. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.