deerflow-code/batch-distill/README.md
2026-09-07 18:24:55 +08:00

34 lines
1.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 批量人物蒸馏 → llmwiki
一个**完全独立的单文件脚本 + 内置定时器**:读 Excel 名单,串行调用线上 DeerFlow 服务问答接口,
复用「女娲 / huashu-nuwa」技能逐个蒸馏人物,**由脚本抓取回复并保存**成 llmwiki 风格 markdown。
- **不联网**:只用已配置的知识库查询技能(`knowledge-search-v2`)检索资料。
- **定时执行**:每天 **23:00 ~ 次日 08:00** 自动跑,到点暂停、次日继续。
- **md 由 py 保存**:agent 只产出内容,脚本抓最终回复写成 `.md`(不让项目自己写文件)。
## 你只需要做两件事
1. **改一行地址**:打开 `distill_people.py`,把顶部 `BASE_URL = "..."` 改成你的线上服务地址。
2. **放名单**:同目录放 `people.xlsx`(或 `people.csv`),**第一列 = 人名**,其余列=可选提示。
其余(用户名 `distiller`、口令 `123ewq`、技能名、时间窗)都已内置,不用填。
## 运行
```bash
python distill_people.py # 常驻,按 23:00~08:00 窗口自动执行(推荐挂后台)
python distill_people.py --now # 忽略时间窗,立刻把名单跑完(测试用)
python distill_people.py --redo # 忽略已完成,全部重做
```
读 `.xlsx` 需要 `openpyxl`;用 `.csv` 则零额外依赖。
## 输出(都在 output/ 一个文件夹)
- `<人名>.md` —— 每个人的 llmwiki 档案
- `manifest.json` —— 执行清单(状态/耗时/run_id)
- `index.md` —— 可读清单表(谁完成/失败,点开看 md)
**断点续跑**:已完成的人自动跳过;中途往名单里加人,下个执行窗会带上。