Methodology sample flow extractor zh
Skill gy910210/hermes-research-skills/skills/methodology-sample-flow-extractor-zh
方法与样本流抽取器:专门从论文、附录、官方页面中抽取训练方式、样本流、任务阶段、prompt/template、loss 栈、披露比例和工业结果,生成可比较的方法卡与样本流对照表。适合深挖生成式搜推、训练数据构造、methodology 例子集和实验设计模板时调用。From its SKILL.md
npx -y skills add gy910210/hermes-research-skills --skill methodology-sample-flow-extractor-zhAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
3.7 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it
Hermes 适配说明
- 本 skill 现面向 Hermes 使用,优先依赖 Hermes 原生工具:
search_files、read_file、write_file、patch、session_search、delegate_task、cronjob、browser、web/search、vision。 - 若正文提到
references/...或scripts/...,优先读取当前 skill 目录下对应文件,不再依赖 Claude 专属目录结构。 - 原始 Claude
agents/openai.yaml不作为执行前提;需要并行研究、分工精读或角色评审时,改用 Hermes 的delegate_task。 - 保留原有研究方法论与产物契约,但执行层统一按 Hermes 工具体系落地。
Methodology Sample Flow Extractor(中文)
这个 skill 负责抽取“方法怎么做”,尤其是:
- 训练样本怎么构造;
- 分几个阶段;
- 哪些是公开 prompt / template;
- 哪些比例和规则公开了,哪些没有。
它不是普通 paper-reader-zh 的替代,而是一个更偏“实验与实现结构”的专用抽取器。
何时使用
- 你要做“训练方式 / 样本流”专题。
- 你想把多篇论文的 methodology 放到统一模板里比较。
- 你想知道论文有没有公开 prompt、appendix template、task mixture、sample ratio。
- 你想产出方法例子集、实验设计模板、样本构造路线图。
输入
sources[]- 论文、PDF、appendix、官方技术页
focussample_flowpublic_prompt_evidenceloss_stackstagewise_trainingindustrial_result_mapping
granularitypaper_cardcomparison_tableexecution_template
输出
methodology_packettask_and_sceneraw_data_objectstoken_objectsstagewise_sample_flowtraining_targetsloss_stackpublic_prompt_evidencedisclosed_ratiosundisclosed_gapsindustrial_results
method_comparison_rows[]illustrative_examples[]
工作流
- 先判断论文属于哪一类:
- tokenizer / SID
- generative retrieval/search
- recommendation backbone
- reranking / control
- training system / serving
- 优先抽“对象定义”,再抽“怎么训练”:
- raw tables / raw logs
- sample object
- token object
- list object
- 再按阶段抽取:
Stage Atokenizer / SIDStage Bbackbone supervisionStage Calignment / preferenceStage Donline refresh / serving coupling
- 识别是否存在公开 prompt / template:
- 正文里的 instruction 示例
- appendix 表格里的 prompt template
- figure / pseudo-template
- 把没有公开的地方显式记成:
未披露摘要级可知appendix 可知
- 如果用户要求例子集,可在论文支持范围外补
illustrative example,但必须明确标注为示意,不得冒充论文原文。
守护
- 不把“方法直觉”写成“论文披露的实现细节”。
- 对 prompt/template 尤其保守:只有公开正文/附录稳定可见时,才记为公开证据。
- 对 ratio、filtering threshold、sampling policy,只要没看到原文,就写
未披露。 - 对工业收益,不做跨论文强行横比;只记录论文公开写出的结果。
推荐产出形式
- 单篇 methodology 卡
- 样本流对照表
- prompt 证据表
- 可执行实验设计模板
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.