deerflow-code/offline-backend-20260512/backend/app/gateway/roundtable_run_policy.py
2026-09-07 18:24:55 +08:00

263 lines
14 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""圆桌各角色「单轮 run 参数」的**单一数据源**。
历史问题:前端手动路径(``multi_agent.py`` 的 ``_leader_run`` / ``_special_run``)
和后台进程内网关(``roundtable_inprocess_gateway.py``)各自硬编码了一份「每角色
工具裁剪 + thinking + reasoning」,结果**漂移**出差异(后台席位 / 报告少禁了
``web_search``,能联网搜索而前端不能)。
本模块把这套策略收敛成唯一定义,两条路径都从这里取,保证「后台挂起」与「前端
手动调用」传参**逐字一致**,且以后改一处即可,不再漂移。
角色(``role``):
- ``leader`` —— 总控派活:短决策,关 thinking + low reasoning(首字节快);
**只保留 ``agent_orchestration`` / ``ask_clarification`` / ``present_files``**,
其余「干活类」工具(``write_file`` / ``bash`` / ``str_replace`` / ``read_file`` /
``ls`` / ``view_image`` / ``task`` / ``write_todos`` / ``memory`` / ``hindsight_*``)
一律禁用——否则弱模型(如 deepseek-chat)会概率性地**自己写文件、跑 bash 干活**
而不是派活(已在生产日志里观测到总控 ``wc -l outputs/*.md`` 自检自己写的报告)。
总控的唯一职责是「派活 / 澄清 / 收口呈现」,没有任何理由动文件系统或执行命令。
并配合 ``SkillStopMiddleware`` 在 ``agent_orchestration`` 调用后停图,便于网关从
AIMessage.tool_calls 解析派活列表。
- ``seat`` —— 子智能体交付:开 thinking + medium reasoning;**可写文件**
(``write_file`` / ``str_replace`` 产出 md 等成品文档)+ **可跑 ``bash``**
(2026-06 放开:业务需要跑知识库检索 / Python 辅助分析)。禁派活 / 澄清 /
展示文件 / 看图 / 联网搜索。``report`` 通过 ``[*_SEAT_EXCLUDED, ...]`` 继承本角色,
所以 Step3 报告也跟着拿到 bash;只有 ``leader`` 和 ``seat_parallel`` 仍各自显式禁。
历史警示(保留):弱模型用 ``bash`` 跑 ``python xxx.py > out.md`` 生成产物会
**绕过前端虚拟预览**(预览只拦截 ``write_file.args.content``)→ 沙箱里有文件但
Step2 右侧预览空白。席位 SOUL 必须明确「bash 只用于辅助查询拿数据,**最终交付
必须用 write_file 写 .md 落盘**,不要让脚本生成最终产物」。
- ``seat_parallel`` —— **DAG 快速(fast)模式的并行席位(方案 A,2026-06)**:**禁写文件**
(``write_file`` / ``str_replace``,产物走正文)+ **禁 ``bash``** + 禁「打断 / 越权」类,
**放开 ``web_search`` 与技能**(搜索不碰 sandbox;技能靠 ``read_file`` SKILL.md)——
否则并行席位几乎调不动工具、产生不了步骤(步骤条不显示)。
并行席位同时跑 bash 会在共享 sandbox 里抢路径 / 占进程数;且 fast 模式整体不
写文件,bash 跑出来的产物既不能预览也没意义。
「写文件产出 md」归 **沙箱(sandbox)模式**(走 ``seat`` policy,单列 + 沙箱跟随预览);
fast 模式快、双列、产物走正文 —— 用户明确要求 fast 不生成文件。详见
``docs/roundtable-dag-orchestration.md`` §4 / §12.1 / §12.8 / §12.10。
与 ``seat`` 的差异:``seat`` 禁 ``web_search``(旧「圆桌一律禁搜索」),
``seat_parallel`` **放开搜索**(本期诉求);``seat`` 已放开 ``bash``,本角色仍禁。
- ``summary`` —— Step3「总结报告」:写 Markdown 总结报告 + 支持问答/增量改。
- ``action_plan`` —— 岗位会商收口:写 Markdown 行动规划与 JSON 子任务清单。
工具面与 ``report`` 完全一致(继承 ``seat`` 的 bash 放开 + 额外禁 ``str_replace`` /
读前序席位),单列一个角色名便于以后独立调优。
- ``report`` —— Step3 报告绘制:在 ``seat`` 基础上**额外禁 ``str_replace``**:
* 继承 ``seat`` 的 bash 放开(2026-06)—— Step3 也可跑 bash 做辅助查询,同
``seat`` 警示:**结果必须 write_file 落盘**,否则前端虚拟预览会空白
(预览只抓 ``write_file.args.content``,抓不到脚本产物 → 预览停在半截
``write_file``,正是「没写完就去生成 py 文件」的现象);
* 禁 ``str_replace`` → 强制返工用整文件 ``write_file`` 重写
(``str_replace`` 没有 ``content`` 字段,会让前端虚拟预览读到 undefined
而清空沙箱)。
"""
from __future__ import annotations
from dataclasses import dataclass
# ── 每角色基础工具裁剪(唯一定义处)─────────────────────────────────────────────
#
# 注意:``seat`` / ``report`` 都包含 ``web_search`` —— 圆桌研讨里联网搜索一律禁用
# (与前端 ``_special_run`` 的 excluded_tools 对齐)。
#
# ``leader`` 用「白名单」思路反推黑名单:总控**只允许** agent_orchestration(派活)、
# ask_clarification(澄清)、present_files(收口呈现子智能体产出的文件)三件事,
# 因此把其余所有工具全部禁掉。关键是 write_file / bash / str_replace —— 留着它们,
# 弱模型会绕过派活直接自己写报告(生产日志已实锤)。memory / hindsight_* 一并禁,
# 配合 ``memory_injection_enabled=False`` 让单例总控不被上轮会话的记忆带偏席位。
_LEADER_EXCLUDED: list[str] = [
"web_search",
"view_image",
"bash",
"ls",
"read_file",
"write_file",
"str_replace",
"task",
"write_todos",
"memory",
"hindsight_recall",
"hindsight_reflect",
"hindsight_retain",
"read_peer_delivery", # 总控靠 _collect_delivery_status 直读各席位,不需要这个席位侧工具
]
# seat:席位**可以写文件**(write_file / str_replace)产出 md 等成品文档,**bash 放开**
# (2026-06 决策,业务需要 bash 跑知识库检索 / Python 辅助分析)。禁派活 / 澄清 /
# 展示文件 / 看图 / 联网搜索。``report`` 通过 ``[*_SEAT_EXCLUDED, ...]`` 继承本角色,
# 所以 Step3 报告跟着拿到 bash;只有 leader / seat_parallel 仍各自显式禁。
#
# **历史警示(保留)**:弱模型用 bash 跑 ``python xxx.py > out.md`` 生成最终产物会
# 绕过前端虚拟预览(只拦 write_file.args.content) → 沙箱里有文件但 Step2 右侧预览
# 空白。必须靠席位 SOUL 强约束「bash 只用于辅助查询拿数据,最终交付必须用 write_file
# 写 .md 落盘」。
_SEAT_EXCLUDED: list[str] = [
"web_search",
"ask_clarification",
"present_files",
"view_image",
"agent_orchestration",
]
# 并行席位 = **快速(fast)模式**(双列):**禁写文件**(产物走正文)+ 禁跑命令 +「打断/越权」类。
# **禁 web_search** —— 离线内网部署无外网,席位检索一律走「技能(skill)」(靠 read_file 读
# SKILL.md,不联网),与 leader / seat / report 三角色保持一致(2026-06 离线部署决策)。
# 「写文件」归 **沙箱(sandbox)模式**(seat policy,单列 + 沙箱跟随预览);fast 模式快、双列、
# 产物走正文 —— 用户明确要求 fast 不生成文件(2026-06 §12.10)。
_SEAT_PARALLEL_EXCLUDED: list[str] = [
"web_search", # 离线部署无外网,检索走技能(skill),不联网
"ask_clarification", # 并行批次不打断问澄清
"present_files", # 产物走正文,不呈现文件
"view_image",
"agent_orchestration", # 席位不派活
"bash", # 禁跑脚本(防 .py 脚本生成产物)
"write_file", # fast 模式产物走正文,不写文件(写文件归 sandbox 模式)
"str_replace",
]
# report:在 seat 基础上额外禁 str_replace(bash 已在 seat 禁)——强制整文件 write_file 重写,
# 让前端虚拟预览能拿到完整 args.content。report 是 Step3 单独绘制,不读前序席位 → 禁 read_peer_delivery。
_REPORT_EXCLUDED: list[str] = [*_SEAT_EXCLUDED, "str_replace", "read_peer_delivery"]
# summary:Step3「总结报告」智能体(写 Markdown 报告 + 支持问答/增量改)。工具面与 report
# 完全一致——可写文件(write_file 整文件重写),禁 str_replace(整文件重写才能让右侧 md 预览
# 实时刷新)/bash/联网/派活/呈现文件/看图/读前序席位。单列出来便于以后独立调优。
_SUMMARY_EXCLUDED: list[str] = [*_SEAT_EXCLUDED, "str_replace", "read_peer_delivery"]
# 岗位会商行动规划与总结报告一样:只能归纳已给出的材料,最终交付必须直接 write_file
# 到 outputs,不能联网、派活或改动其他线程的文件。
_ACTION_PLAN_EXCLUDED: list[str] = [*_SUMMARY_EXCLUDED]
# ingest:Step3「入库」——让 roundtable-structure 真正**调用入库技能**把方案成果(总结报告 +
# 结构化流程图 + taskId)沉淀进知识库。和 dashboard 相反:dashboard 禁一切工具(只输出 JSON),
# ingest 必须**放开技能执行链路**——技能靠 ``read_file`` 读 SKILL.md、按其指令用 ``bash`` /
# ``write_file`` 等工具落地入库(如 curl 入库接口、写中转文件)。在 ``seat`` 基础上额外禁:
# ``str_replace``(整文件重写即可)、``read_peer_delivery``(单轮入库不读前序席位)、``task``
# (不派子代理)、``write_todos`` / ``memory`` / ``hindsight_*``(单例 agent 不吃跨会话记忆,
# 避免污染)。保留 ``read_file`` / ``write_file`` / ``bash`` / ``ls`` 让技能能真正执行。
# thinking + medium:需要先判断该用哪个入库技能、再按其契约调用。
_INGEST_EXCLUDED: list[str] = [
*_SEAT_EXCLUDED,
"str_replace",
"read_peer_delivery",
"task",
"write_todos",
"memory",
"hindsight_recall",
"hindsight_reflect",
"hindsight_retain",
]
# dashboard:Step3「大屏多页」数据智能体(roundtable-dashboard)——纯「研讨成果 → report-json」
# 提炼器,**禁用一切工具**(同 leader 的黑名单全集)。它只需读 prompt、分析各席位、在对话里
# 输出一个 report-json 代码块;留任何干活类工具(尤其 write_file / bash)只会让弱模型跑偏去
# 写文件而不输出 JSON(report agent 上已实锤这类行为)。thinking + medium:需要逐席位分析归纳。
_DASHBOARD_EXCLUDED: list[str] = [
"web_search",
"view_image",
"bash",
"ls",
"read_file",
"write_file",
"str_replace",
"task",
"write_todos",
"memory",
"hindsight_recall",
"hindsight_reflect",
"hindsight_retain",
"read_peer_delivery",
"agent_orchestration",
"ask_clarification",
"present_files",
]
# leader 内禀的 skill_stop:``agent_orchestration`` 调用后停图。调用方可在此基础上
# 再叠加自己的 skill_stop(见 ``merge_skill_stop_names``)。
_LEADER_SKILL_STOP: list[str] = ["agent_orchestration"]
@dataclass(frozen=True)
class RoundtableRunPolicy:
"""一个角色单轮 run 的归一化参数。
与 ``multi_agent._run_payload`` 的入参一一对应:``excluded_tools`` /
``skill_stop_names`` / ``thinking_enabled`` / ``reasoning_effort``。
"""
excluded_tools: list[str]
skill_stop_names: list[str]
thinking_enabled: bool
reasoning_effort: str
def roundtable_run_policy(role: str) -> RoundtableRunPolicy:
"""返回某角色(``leader`` / ``seat`` / ``seat_parallel`` / ``report``)的单轮 run 参数。
返回值里的列表是**新副本**,调用方可安全地再 append/merge 而不污染本模块常量。
未知角色抛 ``ValueError``(早失败,避免悄悄退化到错误的工具面)。
"""
if role == "leader":
return RoundtableRunPolicy(
excluded_tools=list(_LEADER_EXCLUDED),
skill_stop_names=list(_LEADER_SKILL_STOP),
thinking_enabled=False,
reasoning_effort="low",
)
if role == "seat":
return RoundtableRunPolicy(
excluded_tools=list(_SEAT_EXCLUDED),
skill_stop_names=[],
thinking_enabled=True,
reasoning_effort="medium",
)
if role == "seat_parallel":
# 同 seat(thinking + medium),但额外禁写文件 / 跑命令(防并行沙箱打架)。
return RoundtableRunPolicy(
excluded_tools=list(_SEAT_PARALLEL_EXCLUDED),
skill_stop_names=[],
thinking_enabled=True,
reasoning_effort="medium",
)
if role == "report":
return RoundtableRunPolicy(
excluded_tools=list(_REPORT_EXCLUDED),
skill_stop_names=[],
thinking_enabled=True,
reasoning_effort="medium",
)
if role == "summary":
return RoundtableRunPolicy(
excluded_tools=list(_SUMMARY_EXCLUDED),
skill_stop_names=[],
thinking_enabled=True,
reasoning_effort="medium",
)
if role == "action_plan":
return RoundtableRunPolicy(
excluded_tools=list(_ACTION_PLAN_EXCLUDED),
skill_stop_names=[],
thinking_enabled=True,
reasoning_effort="medium",
)
if role == "dashboard":
return RoundtableRunPolicy(
excluded_tools=list(_DASHBOARD_EXCLUDED),
skill_stop_names=[],
thinking_enabled=True,
reasoning_effort="medium",
)
if role == "ingest":
return RoundtableRunPolicy(
excluded_tools=list(_INGEST_EXCLUDED),
skill_stop_names=[],
thinking_enabled=True,
reasoning_effort="medium",
)
raise ValueError(
f"unknown roundtable role: {role!r} (expected leader/seat/seat_parallel/report/summary/action_plan/dashboard/ingest)"
)
def merge_skill_stop_names(policy_names: list[str], caller_names: list[str] | None) -> list[str]:
"""把角色内禀 skill_stop 与调用方追加的 skill_stop 合并去重(保序)。
取代历史上 ``list({"agent_orchestration", *skill_stop_names})`` 的 set 写法
(set 顺序不确定);这里用 dict.fromkeys 保序去重,结果确定、便于断言。
"""
return list(dict.fromkeys([*policy_names, *(caller_names or [])]))