# CLAUDE.md ?? This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## Project Overview DeerFlow is a LangGraph-based AI super agent system with a full-stack architecture. The backend provides a "super agent" with sandbox execution, persistent memory, subagent delegation, and extensible tool integration - all operating in per-thread isolated environments. **Architecture**: - **Gateway API** (port 8001): REST API plus embedded LangGraph-compatible agent runtime - **Frontend** (port 3000): Next.js web interface - **Nginx** (port 2026): Unified reverse proxy entry point - **Provisioner** (port 8002, optional in Docker dev): Started only when sandbox is configured for provisioner/Kubernetes mode **Runtime**: - `make dev`, Docker dev, and production all run the agent runtime in Gateway via `RunManager` + `run_agent()` + `StreamBridge` (`packages/harness/deerflow/runtime/`). Nginx exposes that runtime at `/api/langgraph/*` and rewrites it to Gateway's native `/api/*` routers. **Project Structure**: ``` deer-flow/ ??? Makefile # Root commands (check, install, dev, stop) ??? config.yaml # Main application configuration ??? extensions_config.json # MCP servers and skills configuration ??? backend/ # Backend application (this directory) ? ??? Makefile # Backend-only commands (dev, gateway, lint) ? ??? langgraph.json # LangGraph Studio graph configuration ? ??? packages/ ? ? ??? harness/ # deerflow-harness package (import: deerflow.*) ? ? ??? pyproject.toml ? ? ??? deerflow/ ? ? ??? agents/ # LangGraph agent system ? ? ? ??? lead_agent/ # Main agent (factory + system prompt) ? ? ? ??? middlewares/ # 10 middleware components ? ? ? ??? memory/ # Memory V2: provider/manager/builtin/hindsight ? ? ? ??? thread_state.py # ThreadState schema ? ? ??? sandbox/ # Sandbox execution system ? ? ? ??? local/ # Local filesystem provider ? ? ? ??? sandbox.py # Abstract Sandbox interface ? ? ? ??? tools.py # bash, ls, read/write/str_replace ? ? ? ??? middleware.py # Sandbox lifecycle management ? ? ??? subagents/ # Subagent delegation system ? ? ? ??? builtins/ # general-purpose, bash agents ? ? ? ??? executor.py # Background execution engine ? ? ? ??? registry.py # Agent registry ? ? ??? tools/builtins/ # Built-in tools (present_files, ask_clarification, view_image) ? ? ??? mcp/ # MCP integration (tools, cache, client) ? ? ??? models/ # Model factory with thinking/vision support ? ? ??? skills/ # Skills discovery, loading, parsing ? ? ??? config/ # Configuration system (app, model, sandbox, tool, etc.) ? ? ??? community/ # Community tools (tavily, jina_ai, firecrawl, image_search, aio_sandbox) ? ? ??? reflection/ # Dynamic module loading (resolve_variable, resolve_class) ? ? ??? utils/ # Utilities (network, readability) ? ? ??? client.py # Embedded Python client (DeerFlowClient) ? ??? app/ # Application layer (import: app.*) ? ? ??? gateway/ # FastAPI Gateway API ? ? ? ??? app.py # FastAPI application ? ? ? ??? routers/ # FastAPI route modules (models, mcp, memory, skills, uploads, threads, artifacts, agents, suggestions, channels) ? ? ??? channels/ # IM platform integrations ? ??? tests/ # Test suite ? ??? docs/ # Documentation ??? frontend/ # Next.js frontend application ??? skills/ # Agent skills directory ??? public/ # Public skills (committed) ??? custom/ # Custom skills (gitignored) ``` ## Important Development Guidelines ### Documentation Update Policy **CRITICAL: Always update README.md and CLAUDE.md after every code change** When making code changes, you MUST update the relevant documentation: - Update `README.md` for user-facing changes (features, setup, usage instructions) - Update `CLAUDE.md` for development changes (architecture, commands, workflows, internal systems) - Keep documentation synchronized with the codebase at all times - Ensure accuracy and timeliness of all documentation ### Workflow Studio stream and canvas compatibility - `app/gateway/workflow_agent_runner.py` is the boundary between LangGraph token chunks and durable workflow events. Provider reasoning in `AIMessage.additional_kwargs.reasoning_content`/`reasoning` is emitted to the UI as balanced inline `` blocks; only visible answer content becomes the node's `text` output for downstream nodes. It must also emit every original `[AIMessage|ToolMessage, metadata]` tuple as `node.message.data.messages` with no whitelist, preview or length crop. Tool results and `additional_kwargs.clarification` are needed by the Studio's DeerFlow-compatible message UI and must survive durable SSE replay. - An agent `ask_clarification` ToolMessage pauses the workflow at that agent node (`awaiting_input`), not as a terminal run. The pause records the source `nodeRunId`/`toolCallId`; resuming turns the submitted reply into the next user message on the same `wf-{run}-{node}` LangGraph thread. Keep the raw ToolMessage authoritative; do not invent a second card payload. - `app/gateway/workflow_proposal_planner.py` converts the workflow controller's selected catalog roles into per-agent `mission`/`deliverable`/`scope`/`handoff` contracts. Only a selected visible `agentId` may contribute a contract; every omitted or malformed field falls back to a deterministic role-aware contract. Preserve this prompt-level worker boundary when adding strategies: researchers must not silently become final writers, and sequential roles must receive their upstream handoff binding plus an explicit next handoff. - `app/gateway/routers/workflows_coze_compat.py` must return both `creator.self: true` and editable VCS metadata for an owned standalone workflow. Coze Playground derives its preview/read-only mode from those fields, so omitting `creator.self` disables node dragging even when the standalone page passes `readonly={false}`. ## Commands **Root directory** (for full application): ```bash make check # Check system requirements make install # Install all dependencies (frontend + backend) make dev # Start all services (Gateway + Frontend + Nginx), with config.yaml preflight make start # Start production services locally make stop # Stop all services ``` **Backend directory** (for backend development only): ```bash make install # Install backend dependencies make dev # Run Gateway API with reload (port 8001) make gateway # Run Gateway API only (port 8001) make test # Run all backend tests make lint # Lint with ruff make format # Format code with ruff ``` Regression tests related to Docker/provisioner behavior: - `tests/test_docker_sandbox_mode_detection.py` (mode detection from `config.yaml`) - `tests/test_provisioner_kubeconfig.py` (kubeconfig file/directory handling) Boundary check (harness ? app import firewall): - `tests/test_harness_boundary.py` ? ensures `packages/harness/deerflow/` never imports from `app.*` CI runs these regression tests for every pull request via [.github/workflows/backend-unit-tests.yml](../.github/workflows/backend-unit-tests.yml). ## Architecture ### Position roundtable recovery persistence `/api/position-roundtable` uses the shared `/api/intent/*` LangGraph flow as the live Step-1 execution source and stores an additional SQL recovery copy in `position_roundtable_sessions`/`position_roundtable_nodes`. The recovery copy contains complete intent/node/summary/action-plan conversations, unsent composer drafts, workflow snapshots, and bounded full text for generated Markdown/JSON artifacts. This lets history and downstream closing agents keep working when container-local checkpoint or sandbox volumes are replaced. Revision `20260723_05` adds the conversation columns. Session deletion explicitly removes child nodes so SQLite deployments do not depend on FK cascade pragmas. The router is mounted canonically below `/api/position-roundtable` and, for intranet gateway allowlists, below the schema-hidden compatibility prefix `/api/multi-agent/position-roundtable`. Clients may retry the compatibility route only for a framework/proxy-level `404 Not Found`; domain-level missing-session responses must remain real 404s. The position-roundtable state machine is multi-worker safe: node commands (`complete`, `bind`, conversation save, reject) and session commands (`activate`, archive, delete) lock the parent session and affected nodes in one transaction, then record a unique command id for retry replay. The task and intent snapshot writes use the same parent-row lock; they are permitted only while `status=intent_pending`, so a late “confirm intent” write cannot change a chain that another user has already activated. Activation itself rechecks the stored confirmed intent after taking the lock. `tests/test_position_roundtable_ commands.py` and `tests/test_position_roundtable_concurrency.py` cover the state transitions; `tests/test_roundtable_postgres_concurrency.py` is the real PostgreSQL acceptance suite (set `DEERFLOW_POSTGRES_CONCURRENCY_TEST_URL`). Roundtable backend jobs use a database unique active-dedupe key plus a leased dispatcher, never process-local start deduplication. Cancellation is two phase (`cancel_requested` then `cancelled`): the lease owner’s watcher normally finalizes it, while the dispatcher reaps a cancellation whose lease has expired after a worker crash. A worker that loses lease stops its local model run immediately. Do not reintroduce direct `executor.cancel()` from the HTTP cancel route, because it can cancel the watcher before the database state is finalized. The boolean defaults on these new rows must use SQLAlchemy `false()` rather than numeric `0`, so fresh PostgreSQL schemas can be created. Position roundtable selects the separately seeded built-in agent `position-roundtable-intent` by passing `intent_agent_id` to the shared `/api/intent/init` and `/api/intent/stream` endpoints. Its request lifecycle, files, and SSE protocol match `roundtable-intent`, but its cap is five **user-assistance rounds**. A single round may emit several independent `ask_clarification` tool calls; those calls are separate cards, not one card containing several numbered questions. Count that as one round by the tool-calling AI message, never by card count. Topic-only requests must request all material missing decisions as separate cards rather than being treated as a resolved intent; after the fifth set is answered it must emit completed final JSON without unresolved placeholders. The required arrays are `coreGoals`, `riskWarnings`, `keyPoints`, and `strategicSignificance`. Do not treat a fresh dedicated-agent `ask_clarification` as stale: it must win over a same-run premature `[INTENT_READY]` and leave the UI waiting for the user. Conversely, the prior turn’s historical clarification must not block the user reply from completing. Do not change `roundtable-intent`'s `{objective,constraints,assumptions}` contract for this feature. The position router accepts the new schema for activation and injects all four sections into downstream position-node and closing-agent context; legacy stored summaries are mapped read-only for compatibility. `/api/intent/init` must create the empty Step-1 thread through `multi_agent._create_thread_direct`, exactly like multi-agent initialization; do not restore the old HTTP loopback to `/api/threads`, because deployed Gateway ports/proxies may make the automatic page initialization fail while later manual turns still work. If checkpoint creation fails or the parallel seat task is cancelled, the helper compensates by deleting any partial checkpoint and metadata row (the cancellation branch runs under `asyncio.shield` so a second cancellation cannot abort it half-way). Streaming stays on the existing loopback path. `ThreadMetaStore.create` never issues a post-commit `session.refresh()`. Every `ThreadMetaRow` column is assigned before commit and the session factory uses `expire_on_commit=False`, so the returned dict is built from the in-memory row. The removed read-back ran on a *different* pooled connection than the INSERT, which a MySQL read/write-splitting endpoint can route to a lagging replica — the row comes back empty and SQLAlchemy raises `Could not refresh instance`, surfacing as `500 "Failed to create thread"`. Concurrency makes that far more likely, so a single-user deployment never sees it. No caller uses the return value, and all seven call sites (`multi_agent._create_thread_direct`, `services.start_run`, `routers/threads.create_thread`, `roundtable_inprocess_gateway`, `scheduler/service`, `thread_shares`, `authz`) benefit. Keep new `create` implementations read-back-free; `tests/test_position_roundtable_sessions.py:: test_thread_meta_create_never_reads_back_after_commit` guards this. ### Harness / App Split The backend is split into two layers with a strict dependency direction: - **Harness** (`packages/harness/deerflow/`): Publishable agent framework package (`deerflow-harness`). Import prefix: `deerflow.*`. Contains agent orchestration, tools, sandbox, models, MCP, skills, config ? everything needed to build and run agents. - **App** (`app/`): Unpublished application code. Import prefix: `app.*`. Contains the FastAPI Gateway API and IM channel integrations (Feishu, Slack, Telegram, DingTalk). **Dependency rule**: App imports deerflow, but deerflow never imports app. This boundary is enforced by `tests/test_harness_boundary.py` which runs in CI. ### Report Collaboration (AgentScope, in progress) Independent of Workflow Studio and the existing roundtable. Plan: repo-root `docs/agentscope-多智能体报告协作工作台-后端实施计划.md`(`RC-BE-000`~`018` 已落地,下一票 `RC-BE-019`)。Spike baseline: `docs/AGENTSCOPE_BASELINE.md`. - **Config**: `config.yaml → report_collaboration` (`deerflow.config.report_collaboration_config.ReportCollaborationConfig`). Default **off** (`enabled` / `worker_enabled`). - **Contracts**: `app/report_collaboration/contracts/` — snake_case Pydantic models matching `frontend-web/src/report-collaboration/api/types.ts`. SSE envelope uses camelCase (`eventId` / `sessionId` / `runId`). `node_run_id` is always `${node_id}-attempt{attempt}`. - **Persistence**: `deerflow.persistence.report_collaboration` (`report_collaboration_*` tables, migrations `20260905_01` / `20260906_01` / `20260907_01`). SQL + memory stores; one active run per session via `active_dedupe_key`; command/message/run/version idempotency keys refuse cross-session reuse; event `seq` taken from the run row. Session delete removes children explicitly. - **Gateway**: `/api/report-collaboration` (`app/gateway/routers/report_collaboration.py`). `enabled=false` → 503. Plan/run writes enqueue durable jobs and return immediately. Pre-run `POST /sessions/{id}/messages` runs a rule-based `IntentRouterAgent` + `RequirementResolver` (`app/report_collaboration/requirement/`): writes a requirement revision and clarification messages, never auto-starts planning. `POST /sessions/{id}/plan-requests` then runs `PlanProposalService` (`app/report_collaboration/planning/`): catalog projection (no SOUL/keys) → selector → 2–3 strategy RoleDemand graphs → assembler binds catalog agents/tools/quality gates → validator (acyclic DAG, artifact types, allowlisted tools). Select and start-run stay independent commands. `POST /runs` freezes role models, report-structure template and run budget. Rewrite / apply / restore are live (`reporting/`). `GET /runs/{id}/stream` is a read-only durable SSE tail (`execution/live_hub.py`): persist then publish, replay + `after_seq` / `Last-Event-ID`, heartbeat comments, token-delta coalesce before seq assignment. Disconnect does **not** cancel the run. Lifespan mounts `ReportCollaborationDispatcher` only when `enabled` **and** `worker_enabled` are both true; default both false. - **Runtime**: AgentScope adapters live in `app/report_collaboration/agentscope_runtime/` (not the publishable harness). `ConversationReplica` maps AgentScope 2.0.5 events 1:1 onto collaboration SSE types; TeamSay is `HINT_BLOCK`. `DeerFlowChatModelAdapter` wraps `create_chat_model` (no second Provider config); role models are frozen onto node `agent_snapshot` at `POST /runs` and missing/incapable models fail with 422 before the run is created. Fallback is recorded on the snapshot plus a durable `heartbeat` audit event. Nodes complete only after a validated artifact (`app/report_collaboration/execution/completion.py`). `TaskLedger` (`execution/ledger.py`) owns ready waves, attempts and terminal node status; AgentScope / coordinator suggestions never pick the next node. A writer-bearing run is `completed` only after `ReportFinalizer` publishes a formal version. `WaveExecutor` + `QualityRoleKernel` (`execution/quality_roles.py`) remain for fixtures: Verifier marks stale/contested claims, Analyst/Writer only cite supported IDs, Reviewer `revise` replans only the target write node plus review. Production worker injects `AgentScopeMemberKernel` (`agentscope_runtime/member_kernel.py`) with DeerFlow tools and per-run budget. `QualityGate` (`execution/quality_gate.py`) is the claim-level gate: conflicts/missing units, unverified facts, coverage gaps, and unmapped review targets fail before `validated`; `ArtifactBoard.supersede_previous` hides old attempts from downstream. `ReportCollaborationLiveHub` (`execution/live_hub.py`) fans out persisted envelopes to SSE viewers; `PublishingReportCollaborationStore` sanitizes secrets/privacy/hidden reasoning without truncating visible tool results, and writes redaction + event audits. Dispatcher (`execution/dispatcher.py`) atomically claims queued/recovering/expired-lease runs, heartbeats via `renew_lease`, two-phase cancels (`cancel_requested` → `finalize_cancel`), and recovers orphans (`execution/recovery.py`): keep completed work, mark leftover in-flight `interrupted`/`superseded`, open a new `planned` attempt. Late artifacts after cancel/reclaim are `superseded` and must not change terminal run status. Tests drive `dispatch_once()`; do not rely on the background loop. Per-run budgets (`security/budget.py`) fail-closed; per-user active-run caps are enforced at `create_run`. - **Quality eval**: `app/report_collaboration/eval/` (`RC-BE-018`) — 20 annotated tasks (policy/market/enterprise/event/comparison ×4) with required angles, key facts, forbidden fabrications, templates and review rules. Deterministic scorer (no LLM judge) gates citation/angle/template/directed-mod; gold reports must pass, canned failures (including a roundtable-style essay) must fail and stay reproducible. Recommended runtime knobs live in `eval/thresholds.py` and match current `ReportCollaborationConfig` defaults (`worker_enabled` remains false). Tests: `tests/test_report_collaboration_eval.py`. - **Runtime intervention**: `app/report_collaboration/interventions/` owns run-time natural-language commands. `RuntimeIntentRouterAgent` uses explicit rules first and a tool-free coordinator-model fallback only for ambiguity; its Pydantic decision cannot dispatch or mutate anything. `ImpactAnalyzer` validates server-side node/agent ownership, resolves angle-specific lineage, and enforces confidence/cost/parallel-branch confirmation (`intent_confirmation_threshold`, default `0.8`). `RuntimeCommandService` mirrors the human message, emits frozen command events, CAS-claims accepted commands, and applies them only through `TaskLedger` safe points. Rework persists command overlays, supersedes affected attempts/artifacts, creates monotonic attempts, and prevents late output from entering downstream. Waiting-node answers resume only those nodes. `rerun_node` without a legal target does not expand to the whole graph; requirement patches apply only after rework succeeds; stuck `executing` commands are reclaimed on takeover or lease timeout; `command.cancelled` is a durable SSE event. In-run `POST /sessions/{id}/messages` returns 409 (`RUN_ACTIVE`). Formal rewrite candidates and version apply/restore live in `reporting/` (`RC-BE-015`). - **Dependency**: optional extra `report-collaboration` → `agentscope==2.0.5`. Do not add AgentScope to the harness `pyproject.toml`. Tests: `tests/test_report_collaboration_*.py`. ### Workflow Studio Coze Studio is canvas-only; ChatDev is reference-only. Plan + status per phase: `docs/WORKFLOW_STUDIO_BACKEND_DEV_ZH.md`; decisions: `docs/adr/0001-workflow-studio-runtime.md`; contracts: `docs/WORKFLOW_SCHEMA_V1_ZH.md` + `docs/WORKFLOW_SSE_V1_ZH.md` + `docs/WORKFLOW_OPENAPI_DRAFT_V1.yaml`. Master switch + all limits: `config.yaml` → `workflows` (`deerflow.config.workflow_config.WorkflowConfig`). `workflows.agent_recursion_limit` defaults to 250: LangGraph counts each model→tool research exchange as about two super-steps, so workflow agent/skill nodes must not fall back to a chat-sized 60-step limit. In the local Studio proxy, keep Rsbuild `server.compress: false`; dev gzip buffers small SSE frames and makes live planning/run output arrive only at completion despite the backend's `no-transform` / `X-Accel-Buffering: no` headers. The composer’s optional `modelName` must be validated against `AppConfig.models`: it selects the controller model and is copied only into newly generated candidate agent configs (or the projected direct-agent task), never injected into an already-confirmed proposal/run. **Execution kernel is a custom DAG scheduler, not a compiled LangGraph.** LangGraph stays the kernel *inside* an `agent`/`skill` node. The reason is single-source-of-truth recovery: a re-leased or resumed run replays completed node outputs from `workflow_node_runs` plus the `workflow_run_events` log, so a second (checkpoint) copy of workflow-level state would have no defensible tie-breaker. See the ADR's phase-2 amendment. **Harness** (`packages/harness/deerflow/workflows/`, never imports `app.*`): - `schemas` / `events` / `errors` / `ports` / `validator` / `samples` — frozen v1 contracts. - `expressions.py` — `{{ inputs.x }}` / `{{ nodes.n.data.y }}` templates and conditions resolved by path walking + a small comparator set. **No `eval`.** Roots are restricted to `inputs` / `nodes` / `run` / `loop` / `env`. - `runtime/engine.py` — the scheduler. Every edge is `pending` → `active` (source chose this port) or `pruned`; a node is ready when its non-back incoming edges are all resolved and at least one is `active`, and is skipped when all are `pruned` (skips propagate). That one rule covers fan-out, condition branching, and merge joins with no topological pre-pass. The `loop` node's `continue` back edge is the only permitted cycle; re-entry clears the body's results and invalidates its node runs. - `runtime/{sink,sql_runner,code_runner,context}.py`, `nodes/*` (15 executors, including deterministic `evidence_normalizer` and durable `deep_research_write`), `security/{http_policy,sql_policy,secret_box}.py`. **Persistence**: `workflow_definitions` / `workflow_versions` (migration `20260828_01`), `workflow_runs` / `workflow_node_runs` / `workflow_run_events` / `workflow_run_artifacts` (`20260828_02`), `workflow_data_sources` (`20260828_03`), and short-lived-but-durable `workflow_planning_sessions` / `workflow_planning_proposals` (`20260831_01`). Stores under `deerflow.persistence.workflow_{runs,events,data_sources,planning}` with SQL + in-memory implementations. Because planning proposal rows have a scalar FK rather than an ORM relationship, `SqlWorkflowPlanningStore.create_session` must flush the parent session before adding its children; otherwise SQLite can insert the child first and return a 500. A planning proposal is distinct from a run: its graph can be edited with optimistic `graph_revision` checks, but it becomes executable only once selected and atomically bound to one run. **App layer**: `workflow_dispatcher.py` (lease scan → atomic claim; an expired lease is how a dead worker's run is reclaimed, so every worker can run the loop), `workflow_executor.py` (one lease's lifecycle: replay → engine under a run timeout with heartbeat + cancel watcher → exactly one terminal transition; also runs subworkflows inline, depth ≤ 3), `workflow_retention.py` (hourly sweep of `workflows.retention.event_retention_days` / `run_retention_days`), `workflow_live_hub.py` (in-process SSE fan-out), `workflow_agent_runner.py` (agent/skill nodes → lead agent; `disable_tools=True` means an empty bound-tool set for control-plane callers), `workflow_deep_research_adapter.py` (Evidence Pack → selected Deep Research sources/job/event projection; no internal HTTP call), and `workflow_proposal_planner.py` (calls the dedicated `workflow-planner` controller over visible-agent metadata, then constructs safe candidate graphs server-side). `workflow-planner` is a separate Workflow Studio control-plane agent; never reuse or alias `roundtable-coordinator` for workflow planning, and never let it dispatch a run, call research tools, or return arbitrary graph JSON. It always runs with tools and model thinking disabled. **Gateway routes** (register static prefixes before `workflows.router`'s `/{workflow_id}` catch-alls): `/api/workflows` CRUD/draft/validate/publish/copy, `/api/workflows/node-types` + `/resources/*`, `/api/workflows/{id}/planning-sessions{/stream}` + `/api/workflows/planning-sessions/{session_id}{,/proposals/{proposal_id},/proposals/{proposal_id}/graph}`, `/api/workflows/{id}/runs` + `/api/workflows/runs/{run_id}{,/events,/stream,/cancel,/resume,/retry,/feedback,/artifacts}`, `/api/workflows/data-sources` (+ `/{id}/introspect`, `/{id}/validate-query`), `/api/workflows/embed/{tickets,verify}`, `/api/workflow_api/*` Coze compat. The planning `/stream` endpoint sends product-safe real controller milestones (`catalog`, `controller`, `selection`, `assembly`, `validation`, `complete`) then a final `result` session; never stream agent catalog descriptions or raw planner JSON. Formal runs accept the selected `planningSessionId` and `proposalId`; creating the same pair again returns the already-bound run rather than launching another execution. **Guarantees worth not breaking**: events are persisted *then* published, and the DB is authoritative for `seq` ordering (live frames at or below the replay cursor are dropped and re-read on the next poll, so a reconnect sees no gap; `Last-Event-ID` wins over `?after=`); a run reaching a terminal status always gets a terminal event, synthesised if necessary, or SSE clients hang; `resumeToken` appears only in the `run.awaiting_input` event, never in the run view; `POST /retry` on a `failed` run copies completed node outputs onto a new run marked `retry_of_run_id` so the engine resumes from the failed node; `POST /feedback` never mutates its source run—on a terminal acyclic run it may target only a completed `agent`/`skill`, reuses completed nodes outside that node's downstream closure, records those `reusedNodeIds` alongside the feedback summary, and puts the bounded feedback only in `RunContext.env.nodeFeedback[target]` before creating a revision run; data-source secrets are write-only (encrypted via `WORKFLOW_SECRET_KEY`, decrypted only inside the executor, reads expose `masked_target`); the `http` node denies private/loopback/link-local targets by default (SSRF), `sql_read` allows one read-only statement with bound parameters, and the `code` node is **off by default** because its runner is a hardened subprocess rather than a container. HTTP interface resources persist a non-secret base_url plus an allowed_methods allowlist (migration 20260829_01); request headers stay only in the encrypted secret blob. The resource catalog may return the address and methods to the owning user, but must never return headers or any decrypted secret. The executor receives HTTP metadata only through its data-source resolver. The Workflow Studio agent/skill resource catalog is a read-only display projection over the existing agent and skill ownership stores. It returns stable `agentId` / `skillId`, an agent's cached skill names, and a `createdBy` display label. Resolve human labels from the immutable owner id at request time and fall back to that id if lookup fails; built-in records must return `系统内置`. Do not create a parallel creator table or accept creator data from the browser. `evidence_normalizer` is a deterministic data node, not an LLM: it consumes only configured/upstream node results and emits a bounded `evidencePack` of claims with node provenance, role labels and gaps. The planner inserts it between parallel research roles and the synthesis Agent. It must not treat source text as verified fact or fetch any external resource, and its explicit `sources[].nodeId` values must stay upstream-only under `validate_workflow_graph`. `deep_research_write` is a durable report node, not an Agent prompt shortcut. It accepts only an Evidence Pack, derives a session identity from the workflow run/node plus the topic, evidence snapshot and writing config, writes the pack as selected `other` source rows with upstream-node provenance, and queues/reuses the existing Deep Research `regenerate` job. Its live `report_delta` and persisted `report_chunk` frames are deduplicated before becoming workflow `node.output.delta` events; only report prose is forwarded (never planning/reasoning deltas). Workflow cancellation requests cancellation on that same job. Do not make it self-call `/api/deep-research/*`, use Deep Research `multi_agent` inside this node, or treat Evidence Pack claims as verified facts. The parallel-research candidate grants the node an explicit 1800-second timeout; deployments that override `workflows.node_timeout_seconds` must keep it at least that high. Tests: `tests/test_workflow_schemas.py`, `tests/test_workflow_definitions.py`, `tests/test_workflow_engine.py`, `tests/test_workflow_runs.py`. **Import conventions**: ```python # Harness internal from deerflow.agents import make_lead_agent from deerflow.models import create_chat_model # App internal from app.gateway.app import app from app.channels.service import start_channel_service # App ? Harness (allowed) from deerflow.config import get_app_config # Harness ? App (FORBIDDEN ? enforced by test_harness_boundary.py) # from app.gateway.routers.uploads import ... # ? will fail CI ``` ### Agent System **Lead Agent** (`packages/harness/deerflow/agents/lead_agent/agent.py`): - Entry point: `make_lead_agent(config: RunnableConfig)` registered in `langgraph.json` - Dynamic model selection via `create_chat_model()` with thinking/vision support - Tools loaded via `get_available_tools()` - combines sandbox, built-in, MCP, community, and subagent tools - System prompt generated by `apply_prompt_template()` with skills, memory, and subagent instructions **agentfx built-in analyst + report convert skills.** `agentfx-analyst` (「任务研判报告助手」) is seeded from `app/gateway/routers/_agentfx_seed_assets/` before `_sync_legacy_agents`. After mandatory `web_search`, the agent writes each tool JSON to `/mnt/user-data/workspace/hits/qN.json` and runs **one** command — `task-report-build` (group `all`) — to fill the 14-category `task-reports.json` for the dashboard slots; the four per-page skills `task-report-{enemy,our,env,judge}` remain for re-running a single dashboard page. All five are thin CLIs over the shared stdlib engine `skills/public/task-report-lib/convert_lib.py` (no SKILL.md on the lib, so it never enters the skill list). The model must not `write_file` that JSON. Because the sandbox maps `/mnt/user-data/...` onto a host dir, a virtual absolute hits path can be `isdir`-true yet list empty on a Windows host — `resolve_hit_files` therefore tries the path as given, then the cwd-mapped sandbox path (`discover_user_data_root` + `resolve_sandbox_path` — bash `cd`s into the thread `workspace`, so `/mnt/user-data/outputs/task-reports.json` must land in that thread's `outputs/`, never `F:\mnt\...` on a Windows host), then the virtual-prefix-stripped relative form, then the bare dir name, and picks the first candidate that actually contains `.json`; a miss returns a `warnings` entry listing every tried form so the agent never has to write a debug script. Assemble of the 14 categories is the same function (`assemble_records`); the model must not copy or merge files. If the model is interrupted after convert, `task-reports.json` is already in outputs and the frontend import card also scans `GET /api/threads/{id}/artifacts` so入库 still works without `present_files`. Import remains `task-report-import`. Existing agent/skill directories are never overwritten. Tests: `tests/test_task_report_convert.py`, `tests/test_agentfx_seed.py`. **Forced-research built-in agent.** `forced-research-responder` (「强制检索输出助手」) is seeded from `app/gateway/routers/_forced_research_seed_assets/` before `_sync_legacy_agents`. `ForcedResearchMiddleware` is mounted only for this stable id. On every visible user turn it requires a successful information-gathering tool result before the model may finish; `tool_choice=required` is backed by an `after_model -> model` retry for gateways that ignore tool choice, with an eight-attempt failure cap that permits only an honest failure response. Skill discovery, reading `SKILL.md`, and writing a helper script do not satisfy the gate; web/knowledge/MCP results, non-skill uploaded-file reads, or successful Python skill-script execution do. Tool filtering removes `present_files`, `skill_manage`, task/schedule/memory/clarification surfaces. The middleware permits `write_file`/`str_replace` only for `/mnt/user-data/workspace/*.py` and permits `bash` only for `.py` execution without redirection/document paths, so skills can run scripts but the agent cannot create Markdown/document artifacts or write `outputs`. Tests: `tests/test_forced_research_agent.py`. **ThreadState** (`packages/harness/deerflow/agents/thread_state.py`): - Extends `AgentState` with: `sandbox`, `thread_data`, `title`, `artifacts`, `todos`, `uploaded_files`, `viewed_images` - Uses custom reducers: `merge_artifacts` (deduplicate), `merge_viewed_images` (merge/clear) **Runtime Configuration** (via `config.configurable`): - `thinking_enabled` - Enable model's extended thinking - `model_name` - Select specific LLM model - `is_plan_mode` - Enable TodoList middleware - `subagent_enabled` - Enable task delegation tool ### Middleware Chain Lead-agent middlewares are assembled in strict append order across `packages/harness/deerflow/agents/middlewares/tool_error_handling_middleware.py` (`build_lead_runtime_middlewares`) and `packages/harness/deerflow/agents/lead_agent/agent.py` (`_build_middlewares`): 1. **ThreadDataMiddleware** - Creates per-thread directories under the creator's stable isolation scope (`backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/{workspace,uploads,outputs}`); resolves persisted threads through `resolve_path_user_id(thread_id)` (falls back to `get_effective_user_id()` / `"default"` when metadata is unavailable); Web UI thread deletion now follows LangGraph thread removal with Gateway cleanup of the local thread directory 2. **UploadsMiddleware** - Tracks and injects newly uploaded files into conversation 3. **SandboxMiddleware** - Acquires sandbox, stores `sandbox_id` in state 4. **DanglingToolCallMiddleware** - Injects placeholder ToolMessages for AIMessage tool_calls that lack responses (e.g., due to user interruption), including raw provider tool-call payloads preserved only in `additional_kwargs["tool_calls"]` 5. **LLMErrorHandlingMiddleware** - Normalizes provider/model invocation failures into recoverable assistant-facing errors before later middleware/tool stages run 6. **GuardrailMiddleware** - Pre-tool-call authorization via pluggable `GuardrailProvider` protocol (optional, if `guardrails.enabled` in config). Evaluates each tool call and returns error ToolMessage on deny. Three provider options: built-in `AllowlistProvider` (zero deps), OAP policy providers (e.g. `aport-agent-guardrails`), or custom providers. See [docs/GUARDRAILS.md](docs/GUARDRAILS.md) for setup, usage, and how to implement a provider. 7. **SandboxAuditMiddleware** - Audits sandboxed shell/file operations for security logging before tool execution continues 8. **ToolErrorHandlingMiddleware** - Converts tool exceptions into error `ToolMessage`s so the run can continue instead of aborting 9. **SummarizationMiddleware** (`DeerFlowSummarizationMiddleware`) - Context reduction when approaching token limits (optional, if enabled). **Non-destructive**: overrides the stock `before_model` (which persists a `RemoveMessage(REMOVE_ALL_MESSAGES)`) to a no-op and instead compresses **transiently** in `wrap_model_call` ? the model request is replaced with `[summary, *recent]` while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). The `summary` message is `hide_from_ui` and never persisted; summaries are cached **per-thread in a process-wide, thread-safe LRU** (`_GLOBAL_SUMMARY_CACHE`, bounded to 512 threads, sticky boundary). The cache is module-level — **not** on the middleware instance — because `make_lead_agent` rebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed by `thread_id` (resolved like `ThreadDataMiddleware`), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. **Threshold-first gating**: `_plan_compression` returns passthrough (no compression at all) whenever the *full* current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. **Stickiness headroom** (`_apply_stickiness_headroom`, `_KEEP_HEADROOM_FRACTION=0.75`): after the `keep` policy picks a cutoff, the kept suffix is forced below ~75% of *every* trigger before summarizing. Without it a `keep` in *messages* (e.g. 20) can itself exceed a *token* trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a `{"type":"context_compacting","message":"正在压缩上下文…"}` event on LangGraph's `custom` stream channel (the cheap reuse path stays silent); the chat frontend's `onCustomEvent` renders it as a transient toast. Tests: `tests/test_summarization_middleware.py` 10. **TodoListMiddleware** - Task tracking with `write_todos` tool (optional, if plan_mode) 11. **TokenUsageMiddleware** - Records token usage metrics when token tracking is enabled (optional) 12. **TitleMiddleware** - Two-phase auto-title (never shows "untitled"): `before_model` sets an instant question-derived *provisional* title the moment the first user message arrives (no LLM, zero latency), then `after_model` asynchronously generates the final LLM title and replaces it (tracked via the `title_provisional` state flag). Normalizes structured message content before prompting the title model 13. **MemoryMiddleware** - Memory V2: `wrap_model_call` injects Hindsight recall transiently; `after_agent` auto-retains the turn to Hindsight 14. **ViewImageMiddleware** - Injects base64 image data before LLM call (conditional on vision support) 15. **DeferredToolFilterMiddleware** - Hides deferred tool schemas from the bound model until tool search is enabled (optional) 16. **SubagentLimitMiddleware** - Truncates excess `task` tool calls from model response to enforce `MAX_CONCURRENT_SUBAGENTS` limit (optional, if `subagent_enabled`) 17. **LoopDetectionMiddleware** - Detects repeated tool-call loops; hard-stop responses clear both structured `tool_calls` and raw provider tool-call metadata before forcing a final text answer 18. **ClarificationMiddleware** - Intercepts `ask_clarification` tool calls, interrupts via `Command(goto=END)` (must be last) 8a. **ToolMetricsMiddleware** - Appended immediately after `ToolErrorHandlingMiddleware` in `_build_runtime_middlewares`, so it sits *deepest* in the `wrap_tool_call` chain. Records one `tool_call_metrics` row per tool/skill/MCP invocation ? timing duration and bucketing failures into a coarse `error_category` (rate_limited / ip_blocked / auth_error / timeout / network_error / not_found / empty_result / tool_error). It catches both raised exceptions (before #8 converts them) and *self-handled* errors returned as error `ToolMessage`s (search tools that swallow HTTP 429/403). It also captures the targeted Agent Skill into `tool_call_metrics.skill_name`: the lead agent invokes a skill by `read_file`-ing its `/mnt/skills///SKILL.md` path (per the `` prompt block ? there is no dedicated "use skill" tool), so `skill_name` is recovered by scanning every tool call's string arguments for the `skills/{public,custom}/{name}/` path shape (`_SKILL_PATH_RE`), plus the explicit `name` argument of `skill_view`/`skill_manage`/`skill_archive`. `skill_view`-specific success detection handles its swallowed failures (`Skill 'x' not found.` returned as a plain string rather than raised). Regression coverage: `tests/test_tool_metrics_middleware.py`. Writes are best-effort fire-and-forget. **Per-run skill merging**: `SqlToolMetricsStore.record` upserts skill rows by `(run_id, skill_name)` instead of inserting one row per call ? a skill touched by several tool calls / retries inside one run collapses to a single row. Any success in the run wins (a skill that failed then finally succeeded is recorded as one `success`); a skill that only ever failed yields one `error` row whose `error_message` accumulates every attempt's message; durations are summed. Non-skill tool calls, and skill calls with no `run_id`, keep one row each. The read?merge?write is guarded by a per-event-loop `asyncio.Lock`. Surfaced by the admin `GET /api/admin/tool-metrics` router (which has a `kind=skill` filter and a per-skill `/summary` breakdown). 9. **PromptPrefixMiddleware** - Injects the admin-configured prompt prefix into the model request only. The prefix is tagged onto the first human message's `additional_kwargs['prompt_prefix']` by the Gateway (`app.gateway.services`) and prepended to content here at call time, so the persisted/displayed message stays prefix-free 10. **SummarizationMiddleware** (`DeerFlowSummarizationMiddleware`) - Context reduction when approaching token limits (optional, if enabled). **Non-destructive**: overrides the stock `before_model` (which persists a `RemoveMessage(REMOVE_ALL_MESSAGES)`) to a no-op and instead compresses **transiently** in `wrap_model_call` ? the model request is replaced with `[summary, *recent]` while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). The `summary` message is `hide_from_ui` and never persisted; summaries are cached **per-thread in a process-wide, thread-safe LRU** (`_GLOBAL_SUMMARY_CACHE`, bounded to 512 threads, sticky boundary). The cache is module-level — **not** on the middleware instance — because `make_lead_agent` rebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed by `thread_id` (resolved like `ThreadDataMiddleware`), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. **Threshold-first gating**: `_plan_compression` returns passthrough (no compression at all) whenever the *full* current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. **Stickiness headroom** (`_apply_stickiness_headroom`, `_KEEP_HEADROOM_FRACTION=0.75`): after the `keep` policy picks a cutoff, the kept suffix is forced below ~75% of *every* trigger before summarizing. Without it a `keep` in *messages* (e.g. 20) can itself exceed a *token* trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a `{"type":"context_compacting","message":"正在压缩上下文…"}` event on LangGraph's `custom` stream channel (the cheap reuse path stays silent); the chat frontend's `onCustomEvent` renders it as a transient toast. Tests: `tests/test_summarization_middleware.py` 11. **TodoListMiddleware** - Task tracking with `write_todos` tool (optional, if plan_mode) 12. **TokenUsageMiddleware** - Records token usage metrics when token tracking is enabled (optional) 13. **TitleMiddleware** - Two-phase auto-title (never shows "untitled"): `before_model` sets an instant question-derived *provisional* title the moment the first user message arrives (no LLM, zero latency), then `after_model` asynchronously generates the final LLM title and replaces it (tracked via the `title_provisional` state flag). Normalizes structured message content before prompting the title model 14. **MemoryMiddleware** - Queues conversations for async memory update (filters to user + final AI responses) 15. **ViewImageMiddleware** - Injects base64 image data before LLM call (conditional on vision support) 16. **DeferredToolFilterMiddleware** - Hides deferred tool schemas from the bound model until tool search is enabled (optional) 17. **SubagentLimitMiddleware** - Truncates excess `task` tool calls from model response to enforce `MAX_CONCURRENT_SUBAGENTS` limit (optional, if `subagent_enabled`) 18. **LoopDetectionMiddleware** - Detects repeated tool-call loops; hard-stop responses clear both structured `tool_calls` and raw provider tool-call metadata before forcing a final text answer 19. **ClarificationMiddleware** - Intercepts `ask_clarification` tool calls, interrupts via `Command(goto=END)` (must be last) ### Configuration System **Main Configuration** (`config.yaml`): Setup: Copy `config.example.yaml` to `config.yaml` in the **project root** directory. **Config Versioning**: `config.example.yaml` has a `config_version` field. On startup, `AppConfig.from_file()` compares user version vs example version and emits a warning if outdated. Missing `config_version` = version 0. Run `make config-upgrade` to auto-merge missing fields. When changing the config schema, bump `config_version` in `config.example.yaml`. **Config Caching**: `get_app_config()` caches the parsed config, but automatically reloads it when the resolved config path changes or the file's mtime increases. This keeps Gateway and LangGraph reads aligned with `config.yaml` edits without requiring a manual process restart. Configuration priority: 1. Explicit `config_path` argument 2. `DEER_FLOW_CONFIG_PATH` environment variable 3. `config.yaml` in current directory (backend/) 4. `config.yaml` in parent directory (project root - **recommended location**) Config values starting with `$` are resolved as environment variables (e.g., `$OPENAI_API_KEY`). `ModelConfig` also declares `use_responses_api` and `output_version` so OpenAI `/v1/responses` can be enabled explicitly while still using `langchain_openai:ChatOpenAI`. **Extensions Configuration** (`extensions_config.json`): MCP servers and skills are configured together in `extensions_config.json` in project root: Configuration priority: 1. Explicit `config_path` argument 2. `DEER_FLOW_EXTENSIONS_CONFIG_PATH` environment variable 3. `extensions_config.json` in current directory (backend/) 4. `extensions_config.json` in parent directory (project root - **recommended location**) ### Gateway API (`app/gateway/`) FastAPI application on port 8001 with health checks at `GET|HEAD /health`, `/api/health`, and `/api/langgraph/health`. These exact paths are always public in `auth_middleware.PUBLIC_HEALTH_PATHS`, including when authentication is enabled; reverse-proxy prefixes and trailing slashes are normalized before the whitelist check. Set `GATEWAY_ENABLE_DOCS=false` to disable `/docs`, `/redoc`, and `/openapi.json` in production (default: enabled). Regression coverage: `tests/test_health_auth_bypass.py`. **Routers**: | Router | Endpoints | |--------|-----------| | **Models** (`/api/models`) | `GET /` - list models; `GET /{name}` - model details | | **MCP** (`/api/mcp`) | `GET /config` - get config; `PUT /config` - update config (saves to extensions_config.json) | | **Skills** (`/api/skills`) | `GET /` - list skills; `GET /{name}` - details; `PUT /{name}` - update enabled; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) | | **Memory V2** (`/api/memory/v2`) | `GET /` - builtin data (USER.md/MEMORY.md); `POST/PUT/DELETE /entries` - add/replace/remove entry; `GET/PATCH/DELETE /config` - effective config + per-user overrides; `POST /search` - Hindsight recall/reflect proxy | | **Agents** (`/api/agents`) | `GET /` - list built-in, owned, and published agents; `POST /` - create user-owned agent; `GET/PUT/DELETE /{id}` - read/update/delete by stable agent id; `GET /check` - validate non-empty display name (duplicates allowed) | | **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags`); `GET /{name}` - details; `PUT /{name}` - update enabled; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) | | **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) | | **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `GET /{name}/content` - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) | | **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + repeatable `tag_ids` multi-select filter, OR-matched; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `GET /{name}/content` - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) | | **Memory** (`/api/memory`) | `GET /` - memory data; `POST /reload` - force reload; `GET /config` - config; `GET /status` - config + data | | **Agents** (`/api/agents`) | `GET /` - list built-in, owned, and published agents (optional `search` + repeatable `tag_ids` multi-select filter, OR-matched; each item carries `tags` and `featured_order`); `POST /` - create user-owned agent; `GET/PUT/DELETE /{id}` - read/update/delete by stable agent id; `GET /check` - validate non-empty display name (duplicates allowed); `GET /selector` - homepage agent picker (single admin-curated list, visibility-filtered, sorted by `featured_order` asc); `PUT /{id}/featured` - **admin** sets/clears the homepage `featured_order` (must be `published=true` to feature; passing `order: null` clears unconditionally); `PUT /featured/order` - **admin** bulk-reorders the homepage list (body `{agent_ids: [...]}` ? indices become 0..N-1) | | **Tags** (`/api/tags`) | `GET /` - list tags (optional `tag_type` filter; open to all); `POST /` - create tag (admin); `PUT /{id}` / `DELETE /{id}` - update/delete tag (admin); `GET /assignments/{target_type}/{target_id}` - tags on a resource; `PUT /assignments/{target_type}/{target_id}` - replace a resource's tags (admin). Tags are typed `agent`/`skill`/`scheduled_task` and may only tag resources of their own type. | | **Uploads** (`/api/threads/{id}/uploads`) | `POST /` - upload files (auto-converts PDF/PPT/Excel/Word); `GET /list` - list; `DELETE /{filename}` - delete | | **Threads** (`/api/threads/{id}`) | `DELETE /` - remove DeerFlow-managed local thread data after LangGraph thread deletion; unexpected failures are logged server-side and return a generic 500 detail | | **Threads search** (`POST /threads/search`, LangGraph-compat) | Conversation listing. **System threads are excluded by default**: rows whose `metadata.system` is truthy (scheduler runs, roundtable workers) never crowd the user's list; a caller whose `metadata` filter explicitly references `thread_type`/`system` (e.g. studio notebooks) opts out of the exclusion. New `query` field = case-insensitive title (`display_name`) substring search, used by the chats page's server-side pagination. `POST /threads/count` takes the same request body (limit/offset ignored) and returns `{total}` under identical filtering/exclusion semantics ? powers the chats page's ??/??? display without fetching rows (`ThreadMetaStore.count`). `ThreadMetaStore.search` gained `exclude_system`/`query` params; JSON-blob filters batch-scan until the requested `limit`/`offset` window is full (so pagination over a heavily-filtered set never under-fills), with a `thread_id` tie-breaker on the `updated_at DESC` ordering for stable OFFSET paging. Tests: `tests/test_thread_meta_search.py` | | **Artifacts** (`/api/threads/{id}/artifacts`, `/api/artifact-library`) | `GET /{path}` - serve artifacts; active content types (`text/html`, `application/xhtml+xml`, `image/svg+xml`) are always forced as download attachments to reduce XSS risk; `?download=true` still forces download for other file types. `GET /api/artifact-library` scans every current-user normal thread's authoritative `outputs/` directory (including files omitted from `present_files`), excludes uploads and image/audio/video media, joins the owning conversation title/agent route metadata, and exposes search/filter/pagination with a bounded 15-second per-user cache plus `refresh=true` bypass. Frontend route `/page/workspace/library` links each item back to the source chat with `?artifact=` so the chat opens the existing artifact panel directly. | | **Word export** (`POST /api/writing/export/docx`) | Shared Markdown → DOCX generator used by conversation, artifact, scheduled-task, Deep Research, and AI-writing download surfaces. `app/gateway/word_export.py` builds OOXML with no Python conversion dependency: A4 portrait, 3.7/3.5/2.8/2.6 cm margins, real multilevel heading numbering (`一、` / `(一)` / `1.`), two-character first-line indent on heading levels 1–3, formal body/table styles, page numbers beginning at 1 with odd pages left-aligned and even pages right-aligned, and embedded `app/gateway/assets/fonts/方正小标宋简体.ttf`. Before numbering it removes complete manual heading prefixes (including `2.1` / `2.1.1`) and drops Markdown horizontal rules from Word output. The ordinary-Q&A Lead Agent can emit the matching Markdown heading forms and avoid unordered bullet lists when the admin enables `prompt_prefix.ordinary_qa_markdown_format_enabled` in `system_settings.json`; the switch defaults off, while writing setup, writing mode, notebook, scheduled runs, and SOUL-driven custom agents keep their independent prompts. The font loader validates its internal family, SHA-256, and OpenType embedding permission (`OS/2.fsType`) before export. Tests: `tests/test_word_export.py`, `tests/test_es_query_routing_prompt.py`. | | **Deep Research** (`/api/deep-research`) | Durable, user-owned research sessions/jobs with replayable SSE events and hidden-thread artifacts. `basic`/`quick` use the compact runner; `detailed` creates bounded subtopics, searches/writes sections concurrently, and assembles a report; `deep` recursively searches follow-up questions under a deterministic breadth/depth query budget. `multi_agent` is a LangGraph 1.x editor/researcher/writer/reviewer graph: its initial plan emits `awaiting_input`, the recorder stores the interrupt and releases the lease, and `POST /jobs/{id}/resume` rebuilds the graph from the persisted approved/revised plan. A bounded `custom_outline` is frozen into the session; headings/list entries deterministically supply the detailed and multi-agent plan, and the shared writer receives it only as a guarded content constraint. Completed reports support owner-scoped follow-up (`GET /sessions/{id}/messages`, `POST /sessions/{id}/chat`): context is only the current report, selected sources and recent messages; source-id citations are persisted, while `allow_new_research=true` is explicitly rejected. Owners may change that evidence set only after the job stops via `PATCH /sessions/{id}/sources/{source_id}`; this applies to future follow-ups rather than rewriting a completed report. Optional report illustrations are gated by `config.yaml -> deep_research.image_generation`: the Harness `OpenAICompatibleImageProvider` requires base64 output, the shared runner step writes only bounded `research-image-01..04.(png|jpg|webp)` artifacts and Markdown uses private `deep-research://` links. `GET /sessions/{id}/artifacts/{name}` checks session ownership before serving an image; an unavailable provider yields a warning but never fails the text report. Runners live in `packages/harness/deerflow/agents/deep_research/runners/`, may never import `app.*`, and use only the `AdapterBundle`. | | **Gateway auth** | Owner-scoped Deep Research routes await the platform's asynchronous `get_current_user` resolver before they read or write persistence, ensuring each session and job operation receives a concrete user ID. Deep Research artifacts are always resolved from a whitelisted filename via `/mnt/user-data/outputs/` before mapping to a host path, preserving the sandbox guard on native Windows and Linux hosts. | | **Suggestions** (`/api/threads/{id}/suggestions`) | `POST /` - generate follow-up questions; rich list/block model content is normalized before JSON parsing | | **Thread Runs** (`/api/threads/{id}/runs`) | `POST /` - create background run; `POST /stream` - create + SSE stream; `POST /wait` - create + block; `GET /` - list runs; `GET /{rid}` - run details; `POST /{rid}/cancel` - cancel; `GET /{rid}/join` - join SSE; `GET /{rid}/messages` - paginated messages `{data, has_more}`; `GET /{rid}/events` - full event stream; `GET /../messages` - thread messages with feedback; `GET /../token-usage` - aggregate tokens | | **Feedback** (`/api/threads/{id}/runs/{rid}/feedback`) | `PUT /` - upsert feedback; `DELETE /` - delete user feedback; `POST /` - create feedback; `GET /` - list feedback; `GET /stats` - aggregate stats; `DELETE /{fid}` - delete specific | | **Runs** (`/api/runs`) | `POST /stream` - stateless run + SSE; `POST /wait` - stateless run + block; `GET /{rid}/messages` - paginated messages by run_id `{data, has_more}` (cursor: `after_seq`/`before_seq`); `GET /{rid}/feedback` - list feedback by run_id | | **Open Chat** (`/api/open/chat`) | **Unauthenticated** open Q&A entry for external systems (the `/api/open/` prefix is whitelisted in `auth_middleware._PUBLIC_PATH_PREFIXES` + CSRF-exempt in `csrf_middleware.should_check_csrf`). `POST /chat` `{message \| messages[], model_name?, agent_name?, agent_id?, thinking_enabled=false, strip_think=true, thread_id?, recursion_limit=100}` → resolves the agent by **Chinese display name** (`agent_name`, matched against each `.deer-flow/agents/*/config.yaml` `name` via `_resolve_agent_id`; explicit `agent_id` wins if given), builds the lead agent via `make_lead_agent` (so the chosen **agent's** SOUL/tools/skills all apply), `ainvoke`s it once (stateless — no checkpointer, no thread persistence; runs under the `"default"` user bucket), and returns **only the final** `{answer, model_name, agent_name, agent_id, thinking_enabled, thread_id}` (no intermediate tool-call/process). **`POST /chat/stream`** is the **SSE streaming variant** (same request body) for **long tasks** — the blocking `/chat` produces no downstream bytes for the whole run, so a front nginx `proxy_read_timeout` (default 60s) cuts it off with a **504**; the stream keeps the connection warm (a `: keepalive` SSE comment every `_HEARTBEAT_SECONDS=15`s via a background producer-task + queue, plus `event: delta` AI-text increments and `event: progress`), and ends with **`event: result`** carrying the same final `{answer, …}` JSON (or `event: error`). Clients only need to read `event: result`; intermediate events are ignorable. Both endpoints share `_build_run_context` (config/agent/thread). `GET /agents` lists `{name, id}` so callers know what `agent_name` to send. **Thinking off is reliable**: when `thinking_enabled=false` the route also sets `thinking_force_disabled=true` so the lead agent calls `create_chat_model(force_disable_thinking=True)`, which injects **three** OpenAI-compatible "thinking off" shapes into `extra_body` — top-level `enable_thinking=false` (**SiliconFlow / DashScope hybrid models** like `deepseek-ai/DeepSeek-V4-*` on `api.siliconflow.cn` — the knob the nested `chat_template_kwargs` shape does NOT reach), `chat_template_kwargs.enable_thinking=false` (vLLM/Qwen native), and `thinking.type="disabled"` (Zhipu GLM) — even for models that only declare `supports_thinking: true`. For the **web-UI path** (which sends `thinking_enabled=false` but NOT `force_disable`), SiliconFlow models need `when_thinking_disabled: {extra_body: {enable_thinking: false}}` in their `config.yaml` entry (added to `deepseek-ai/DeepSeek-V4-Flash`) for thinking-off to take effect. Sets `is_scheduled_run=true` internally to skip the auto-title LLM call. **Generated-file return (`include_files=true`, default on)**: an agent that produces its report as a file (writes to `/mnt/user-data/outputs/` + `present_files`) returns only a short summary in `answer`; to recover the real deliverable both endpoints now also return `files: [{name, virtual_path, encoding(text|base64), content}]` — `_collect_output_files` scans the run's `state['thread_data']['outputs_path']` (the open thread is stateless so artifacts can't be fetched by thread, but the endpoint runs server-side and the sandbox already wrote the files to disk), text decoded as utf-8 else base64, capped by `_MAX_OUTPUT_FILES=20` / `_MAX_OUTPUT_FILE_BYTES=5MB`. `/chat` puts `files` on the response; `/chat/stream` puts it in the `event: result` payload. A run with no answer **and** no files now 502s ("…文本答案或文件"). Router `app/gateway/routers/open_chat.py`; standalone caller `test_open_chat_api.py` (repo root, streaming by default); tests `tests/test_open_chat.py`. | | **Open 3qfx Ask** (`/api/open/3qfx/ask`) | **Unauthenticated** open entry that, given just a **任务id**, directly kicks off the 3qfx(三情分析)「第一次问答」as a **real persisted conversation** (sibling of `/api/open/chat`; same `/api/open/` auth-whitelist + CSRF-exempt). `POST /ask` `{task_id, message?, model_name?, thinking_enabled=false}` → (1) resolves the `3Q` agent from `app.state.business_mapping_store.get_mapping("3Q")` (the **same** agent the frontend deep-link `goPath=3qfx` uses; 503 if unconfigured); (2) **fetches the task detail server-side** (consumer `cop-task-detail`) and builds the opening text **identical to the frontend** `buildOpeningText` 3qfx branch — `任务名称:…\n任务内容:…\n任务id(task_id)为{task_id}` (an explicit `message` overrides and **skips** the fetch); (3) creates a fresh thread tagged `metadata.{taskId, agent_id}` (the exact filter `useTaskThreads` uses, and `metadata.taskId` makes it task-deeplink-shared / creator-path-resolved) and **starts a background persisted run** via `start_run(..., enforce_access=False)` (writes the checkpointer, runs under the `"default"` bucket); (4) **returns immediately** `{thread_id, run_id, agent_id, status, opening_text}` — the caller just needs "started", then watches the run in the 3qfx 任务工作区 conversation list. `start_run` gained a keyword-only `enforce_access` (default `True`); the open entry passes `False` so the admin-configured agent isn't rejected by the per-(missing)-login visibility check. **Consumer fetch config** lives in `config.yaml → task_deeplink`: `cop_task_detail_url` (full URL, appends `?taskId=`), `cop_task_detail_authorization` (default `Bearer admin`), `cop_task_detail_test` + `mock.cop_task_detail` (offline fake when the upstream is unreachable — on in this deployment). Router `app/gateway/routers/open_3qfx.py`; tests `tests/test_open_3qfx.py`. | | **Thread Shares** (`/api/thread-shares`) | Per-conversation sharing with two modes (`thread_shares.mode`). **`import`**: `POST /` `{thread_id, mode}` - create (or reuse) a revocable share code; `GET /{code}/preview` + `POST /{code}/import` - copy that conversation into the caller's account (fresh `thread_id` via `app/gateway/thread_copy.py` + new `threads_meta` row + `user-data` copy; dedup on `metadata.shared_origin`). **`view`**: a public, read-only static page ? see Public Shares below. Both modes: `GET /` - list my codes; `DELETE /{code}` - revoke. | | **Public Shares** (`/api/public/shares`) | **Unauthenticated** (the `/api/public/` prefix is whitelisted in `auth_middleware._PUBLIC_PATH_PREFIXES`). `GET /{code}` - read-only conversation snapshot (title + serialized messages + artifacts) for a `view`-mode share; messages are read straight from the latest checkpoint. `GET /{code}/files/{path}` - serves a file referenced by that conversation (images, Markdown, text, PDF, ?) so the shared page's links are openable; path resolution is confined to the owner's thread directory. Active content (`text/html`, `application/xhtml+xml`, `image/svg+xml`) is forced to download (`_public_share_disposition`); everything else is served inline. Missing/revoked/non-`view` codes all return 404. | | **Roundtable Drafts** (`/api/roundtable-drafts`) | **Strictly per-user** persistence for the roundtable-planning page (replaces the old browser `localStorage`). `GET /` - list drafts (lightweight meta — incl. `chain_title`, the business-chain name used this session, cheaply regex-extracted from the in-memory `step2` blob in `_draft_to_meta`/`_extract_chain_title`; null = no chain / 自由模式; surfaced on the 第一步「分析记录」history card + sidebar list instead of the old「第 X 步」); `POST /` - create; `GET/PUT/DELETE /{id}` - full read / partial update / delete (all owner-scoped). A draft = one full session (Step 1 intent + Step 2 roundtable), step snapshots stored as JSON in `PortableLongText`. Recommendation history bound to a draft: `GET /{id}/recommendations` - list newest-first; `POST /{id}/recommendations` - append one run (objective/status/rationale/picks/candidates). taskId 深链会商记录 are **not** here — they live in the separate `/api/roundtable-task-drafts` store (see below) so personal history never mixes with task records. Backed by `deerflow.persistence.roundtable_drafts` (tables `roundtable_drafts`, `roundtable_recommend_history`); store wired in `deps.py` as `app.state.roundtable_draft_store`. Tests: `tests/test_roundtable_task_drafts.py`. | | **Roundtable Task Drafts** (`/api/roundtable-task-drafts`) | **Task-scoped, unpartitioned** store for taskId-deeplink roundtable 聊天记录 (无界嵌入抽屉). Keyed **only by `task_id`** (one task → many records), **never partitioned by user** — any authenticated caller may read/write/delete a task's records (`created_by` is audit-only). Physically separate from `roundtable_drafts` so the two never mix. `GET /by-task/{task_id}` - all drafts for a task, newest-first (history dropdown), enriched with the latest background-job status per draft via `RoundtableJobRepository.status_map_by_drafts` (each meta also carries `chain_title`, same `_extract_chain_title` regex as the personal store); `GET /by-task/{task_id}/latest` - newest draft (回显最新一条); `POST /` - create (`task_id` required); `GET/PUT/DELETE /{id}` - full read / partial update / delete (all user-agnostic). Backed by `deerflow.persistence.roundtable_task_drafts` (table `roundtable_task_drafts`, auto-created by `create_all`; migration `20260618_01`); store wired in `deps.py` as `app.state.roundtable_task_draft_store`. Existing `task_id`-bearing `roundtable_drafts` rows are migrated over by `scripts/migrate_roundtable_task_drafts.py`. Tests: `tests/test_roundtable_task_drafts.py`, `tests/test_roundtable_drafts_by_task.py`. | | **Roundtable Diagnostics** (`/api/roundtable-diagnostics`) | 会商全流程诊断日志 (报错 + 关键事件). One row per event in `roundtable_diagnostics`: `{scope (background = 后台挂起 / foreground = 前台服务端观测 / client = 前端上报), stage (step1_intent/step2_init/step2_leader/step2_seat/step2_recommend/step2_orchestration/step3_summary/step3_report/step3_dashboard/step3_structure/upstream/job), level (warning/error — **`info` 级运行日志不再收集**:`record_diagnostic` 在写库前丢弃 `level=="info"`,只留报错;begin/consensus/gather/job_started 等纯进度事件即使仍调 `.diag(level="info")` 也成 no-op,故诊断页 = 报错日志页), event, message, detail (JSON), job_id/draft_id/task_id/user_id/agent_id/agent_name/cycle, created_at}`. **Reads are admin-only**: `GET /` - filter (scope/stage/level/event/job_id/draft_id/task_id/user_id/q/since/until; `scope` accepts a comma-separated set so 前端「前台调用」= `foreground,client`) + paginate (newest-first), each item enriched with `userName` (user_id → 邮箱 `@`-prefix via `get_local_provider().get_user`, best-effort batched cache); `GET /facets` - distinct values for filter dropdowns; `POST /admin/cleanup` - retention purge. **`POST /report` is NOT admin-gated** — any authenticated user's frontend posts the errors *the user actually hit* (HTTP 502 before any SSE frame / network drop / SSE error frame) so they're collected too; stored as `scope="client"`. Backed by `deerflow.persistence.roundtable_diagnostics` (table `roundtable_diagnostics`, auto-created by `create_all`; migration `20260619_01`); store wired in `deps.py` as `app.state.roundtable_diagnostic_store` + a process-level default store (`register_default_store`) so the harness-side engine/gateway can write across the app/harness boundary. **Writers** (all best-effort — diagnostics never break a run): background path uses `DiagnosticsRecorder` (bound to job ctx) threaded through `RoundtableJobExecutor` → `JobProgressWriter.diag(...)` (engine transitions/step3 failures/empty-delivery) + `InProcessRoundtableGateway` (per-turn timeout/report·summary·dashboard truncation+missing); foreground path uses `app/gateway/roundtable_diag.py::record_foreground_diag` in `multi_agent.py`/`intent.py`/`recommend.py` error frames. **Bug-fix shipped alongside**: `InProcessRoundtableGateway._run_agent_turn` guards each agent run with an **inactivity (idle) timeout** (`_await_run_with_idle_guard`, `ROUNDTABLE_SEAT_IDLE_TIMEOUT_SECONDS` default 1200s = 20min, tuned high for unstable intranet models) — **not** a total-duration cap, so a slow-but-alive run that keeps streaming tokens runs as long as it needs; only a run that emits **zero StreamBridge events for `idle_timeout` seconds straight** (a true freeze: 内网模型连接池挂死 / sqlite 锁在挂载卷上挂死) is cancelled + raises (顶层把 job 标 error,而非永久冻结) + records a `seat_turn_timeout` diagnostic. Activity is observed via `app.state.stream_bridge` (the worker publishes every chunk there). Optional absolute cap `ROUNDTABLE_SEAT_MAX_TOTAL_SECONDS` (default 0 = off). Frontend: admin page `pages/RoundtableDiagnosticsPage.tsx` (route `/page/strategy/admin/roundtable-diagnostics`), opened from the **会商页右上角设置按钮** (admin only — `handleSettingsClick` in `roundtable-planning/pages/RoundtablePlanningPage.tsx`); admin read client `strategy-components/api/roundtable-diagnostics.ts`. **Client error reporter** `roundtable-planning/api/diagnostics-report.ts` (`reportRoundtableError` fire-and-forget + `keepalive`, `setRoundtableErrorContext` for draft/task attribution set by the page, `stageForRoundtableAgent`) is called at every client-observable failure: `multi-agent.ts` (`streamMultiAgent` — covers Step2 leader/seat AND Step3 report/summary/dashboard/structure — reports HTTP error + SSE `{status:"error"}` frame + worker `event:error` frame [模型调用没成功] + fetch-reject [网络不可达] + mid-stream drop [连接重置/超时], deduped via a `reported` flag; `initMultiAgent` HTTP/error-frame/incomplete/mid-stream), `intent.ts` (Step1 HTTP+error-frame), `recommend.ts` (Step2 recommend HTTP+error-frame), `roundtable-jobs.ts` (job api/stream). The admin page's first filter is a **后台挂起 / 前台调用** toggle (frontend = `foreground,client`); the stage dropdown options adapt to it (background → step2+step3+job; frontend → step1+step2+step3) and each row shows the reporting user's name + 👤. Tests: `tests/test_roundtable_diagnostics.py`, `tests/test_roundtable_inprocess_timeout.py`. **Seat skill-recognition (弱模型补偿)**: a seat = one custom agent run whose configured skills (`agent config.yaml → skills`) are injected **only** via the system prompt's `` block (model must `read_file` the SKILL.md itself) — weak models (deepseek-chat 等) ignore it and never use their skills. `app/gateway/roundtable_seat_skills.py::append_seat_skill_directive(agent_id, task, *, chain_enabled=False)` appends an explicit Chinese directive (each configured-and-enabled skill's name + description + SKILL.md container path + "先 read_file 再按其流程取信息/执行") to the **seat's per-turn task (human turn)** — far stronger pull on weak models than the system prompt. Wired into **both** seat paths: foreground `multi_agent._special_run` (only `role in {seat, seat_parallel}`; Step3 singletons report/summary/dashboard/ingest excluded) + background `InProcessRoundtableGateway.run_seat`. No-op (returns task unchanged) when the seat has no configured skills. **2026-06-20: default-off + two opt-in switches (OR semantics)** — was global-on, now injects only when **either** the seat agent's own switch (`AgentConfig.seat_skill_directive`, edited in the agent edit dialog `agent-card.tsx` + `agent-detail-panel.tsx`) **or** the **per-seat** switch on that agent **within the selected business chain** is on. The per-seat chain switch lives **in the chain's `seats` JSON** (`ChainSeatSchema.seat_skill_directive`, no DB column — rides the existing JSON like the per-seat `model` override), edited via a small ✨ icon → dropdown (Switch + description) on each agent card in `BusinessChainEditorPage`. It is threaded **per-seat** to each seat run: foreground `multi_agent` per-call field `seat_skill_directive` (orchestration looks it up per agentId from a `{agentId: bool}` session map) → `_special_run` → `chain_enabled`; background as the job field `seat_skill_directives: dict[str,bool]` (StartJobRequest → StartParams → deps factory → `InProcessRoundtableGateway._seat_skill_directives`) → `run_seat` looks up `agent_id` → `chain_enabled`. `ROUNDTABLE_SEAT_SKILL_DIRECTIVE` is only a **master kill-switch** (default allow; set `0/false` to hard-disable regardless of the switches — never force-enables). Tests: `tests/test_roundtable_seat_skills.py`, `tests/test_agents_backend.py`. **2026-06-20 Step2 提速三件套**:(1) **每席位推理深度覆盖** —— 业务链条 `seats` JSON 里逐席位可配 `seat_mode`(flash/thinking/pro/ultra,覆盖链条级 `seatMode`;编辑器席位卡右侧 ⏱Gauge 图标下拉,null=跟随链条默认)。前台经 `useStep2Orchestration` 三处 dispatch 的 `seatModeToRequestFields(seatModes[id] ?? selectedMode)` 透传;后台经 job `seat_modes: dict[str,str]` → `StartParams` → `InProcessRoundtableGateway._seat_modes`,`run_seat` 用 `_SEAT_MODE_PARAMS`(与前端 `seatModeToRunParams` 逐字对齐)解析覆盖作业级 seat_*(未命中→跟随作业级→policy 默认)。无 DB 列(ride `seats` JSON,但 `ChainSeatSchema.seat_mode` 须显式声明否则 Pydantic 剥掉)。Tests: `tests/test_roundtable_seat_modes.py`。(2) **recommend 批内并行** —— 后端 `engine._run_recommend` 由 for-await 串行改 `asyncio.gather`(对齐前端 2026-06-13 已并行;总控一轮派多席位=判定相互独立→并发;批内用派活前 deliveries 快照、批后统一并入,语义不变)。(3) **研讨前并行『取数』两段式**(业务链条 `gather_first` 列,逐条 opt-in,migration `20260620_02`,editor「研讨前并行取数」开关)—— `gather_first=True` 时 `run_orchestration` 在分发前先 `engine._run_gather_phase`(全席位 `asyncio.gather` 跑 `_GATHER_TASK`「只取数不研讨」),产出 `gathered: dict` 经 `_seat_message(gathered=...)` 作为「已收集资料」块注入 recommend/chain/dag 各研讨轮席位任务(与「前序交付」块分开),把最慢的技能延迟前置并行、给(尤其 chain 串行的)研讨段提速;研讨阶段不限制再调技能。前台对称实现:`useStep2Orchestration.runGatherPhase` 在 `runActiveOrchestration` 分发前跑一次(`gatherPhaseDoneRef` 闸门,resetInitGate 复位),产出存 `gatheredDeliveriesRef` 经 `peerDeliveries`/初始 `priorDeliveries` 注入研讨席位。链条级 `gatherFirst`(camel)↔`gather_first`(snake) 经 `api/chains.ts` 映射;job 经 `StartJobRequest.gather_first`→`StartParams`→`run_orchestration(gather_first=…)`。Tests: `tests/test_roundtable_gather_phase.py`。(4) **削减重复广播**(前端 only,多智能体 foreground 路径)—— `suppress_pre_broadcast`(仅 leader:抑制「总控本轮输入→各席位 thread」预广播 + 「总控给X派活」→兄弟席位的 post 广播,二者一个 flag)此前只在 DAG 派活轮开;现扩到 recommend(`safetyCounter > 1`,即第 2 轮起的续轮/综合)+ chain 逐席位(`i > 0`,第 2 个席位起)+ chain 收口轮(恒 true)。**保留每种模式的首轮预广播**(它为席位打底任务背景),只 suppress 重复的续轮广播(N 次读写巨大席位 thread 的纯浪费 + 弱模型模仿派活);席位的前序交付另经「席位交付广播」`broadcastSeatDeliveryToOthers` 送达,不受影响。后台引擎路径无此广播(`_seat_message` 直接拼上下文),故 #4 仅前端。**2026-06-20 深链接业务链抽取规范**:深链接会商(taskId+goPath,businessCode 如 rwfx→6BF)下,结构化抽取(flow-json)与方案总结报告都按所选业务链**完整层级**强制校正。结构化抽取实为 `roundtable-structure` 智能体(`flow_extract.py` 的 `_system_prompt()=load_agent_soul(STRUCTURE_AGENT_ID)`,抽取→校验两段式 `model.astream`);其 SOUL 全局一份,旧实现把某业务层级写死进 SOUL→所有业务都按「目的→行为体」抽(bug)。修法:**单份领域知识、两角色用法**的 `deerflow.agents.roundtable_orchestrator.extraction_spec`(放 harness 层,app+harness 都可 import;`build_structure_directive`/`build_summary_directive` 按 businessCode 取;rwfx=6BF 写实「任务→目的→行为体→预期行为→驱动因素→叙事,预期行为↔驱动因素 1:1、余 1:N、label 不含步骤/阶段名、优先报告回退席位」,其余业务通用规范=强约束完整性+命名+取数且明令不套固定模板;business_code 空→""零副作用)。注入:(1) 结构化抽取 `flow_extract._hierarchy_hint`+`_validate_messages` 用 `build_structure_directive` 权威覆盖默认层级(前端 `streamFlowExtract` 传 `business_code`,源 `ingestBusinessCode`);(2) 方案总结报告 `build_summary_directive`——前台 `multi_agent._special_run`(role==summary 且 business_code) 追加、后台 `engine.build_summary_prompt(business_code=…)`(StartParams/StartJobRequest 透传)。**仅深链接生效**。种子 `roundtable-structure/SOUL.md` 层级顺序改 spec-driven、连接示例去 rwfx 写死(存量 live SOUL 不被 seed 覆盖,但注入规范权威可纠正)。Tests: `tests/test_roundtable_extraction_spec.py`。 | | **Roundtable Draft Shares** (`/api/roundtable-draft-shares`) | Per-draft sharing for the roundtable-planning page, mirroring Thread Shares but the shared resource is a `roundtable_drafts` row. Two modes (`roundtable_draft_shares.mode`). **`import`**: `POST /` `{draft_id, mode}` - create (or reuse) a revocable code; `GET /{code}/preview` + `POST /{code}/import` - copy that draft (step1/2/3) into the caller's account under a deterministic per-(code,recipient) id (`imp_` ? re-import is a no-op). **`view`**: a public, read-only result page ? see Public Roundtable Shares. Both: `GET /` - list my codes; `DELETE /{code}` - revoke. Backed by `deerflow.persistence.roundtable_draft_shares` (table `roundtable_draft_shares`, migration `20260606_01`); store wired in `deps.py` as `app.state.roundtable_draft_share_store`. Tests: `tests/test_roundtable_draft_shares.py`. | | **Public Roundtable Shares** (`/api/public/roundtable-shares`) | **Unauthenticated** (`/api/public/` whitelisted). `GET /{code}` - read-only snapshot for a `view`-mode draft share: returns the Step 3??????report HTML (`step3.html`, a self-contained document) + title/summary/`has_report`. The frontend renders it in a sandboxed `iframe` (no `allow-scripts`) at `#/roundtable-share/`. Missing/revoked/non-`view` codes all return 404. | | **Scheduled Tasks** (`/api/scheduled-tasks`) | Per-user cron-like agent tasks. `GET/POST /` - list/create owned tasks; `PATCH/DELETE /{id}` - update/delete (owner); `POST /{id}/{pause,resume,trigger}`; `GET /{id}/runs`, `GET /runs/{rid}`, run-file preview/edit/download; `GET /runs/{rid}/messages` - agent messages for the run's underlying agent run (used to inspect a stuck/running run; authorized via `_load_run`); `GET /runs/{rid}/html` - generated HTML document for an `html_page` run; `DELETE /runs/{rid}` - owner-scoped, with an admin fallback so admins can delete any user's run; publish/subscribe under `/{id}/{publish,unpublish,subscribe}` and `GET /published`. **Admin (cross-user):** `GET /admin/all` - every task; `PATCH/DELETE /admin/{id}` - edit/delete any task; `POST /admin/{id}/trigger`; `GET /admin/{id}/runs`. Admin routes require `system_role == "admin"`. | | **HTML Page Favorites** (`/api/html-page-favorites`) | Bookmarked generated HTML pages, owner-scoped. `GET /` - list; `POST /` `{page_id, title?, description?, tags?}` - bookmark a `scheduled_html_pages` row (extracts a reusable style/layout summary via GLM5 at favorite time); a page can only be favorited once per user (`DuplicateFavoriteError` ? HTTP 409). `GET/DELETE /{id}` - read/delete (the create-task reference picker deletes favorites here). A favorite can be selected as a style reference when creating a new `html_page` task. | | **Public Scheduled Tasks** (`/api/public/scheduled-tasks`) | **Unauthenticated** mirror for embedding. Adds `GET /runs/{rid}/html` and `GET /{task_id}/latest-html` for the `html_page` preview; the run result keeps `result.html` (only file paths are redacted). | | **Light Apps** (`/api/light-apps`) | Light-app registry (????? / ????), **global/shared** (no user scoping; `created_by`/`updated_by` audit-only). Two kinds: `window_open` (launched in a new tab from the ???? card grid) and `iframe` (embedded under a parent sidebar menu via `mount_location` and dynamically injected into the sidebar). Both carry `route_params` ? a list of `{key, paramType: login\|theme, source, customValue, themeValue}` descriptors the **frontend** resolves at launch time (login params read from `localStorage.userInfo`; theme params use a literal value) and appends to `url` as query string. `GET /` - list (filters `app_type`/`mount_location`/`enabled_only`/`search`); `GET /{app_id}` - read one (both open to any authenticated user); `POST /`, `PUT /{app_id}`, `DELETE /{app_id}` - **admin only**; iframe apps must carry a `mount_location` (400 otherwise); the unique `name` collides as HTTP 409. Backed by `deerflow.persistence.light_apps` (table `light_apps`, auto-created by `create_all`); store wired in `deps.py` as `app.state.light_app_store`. Tests: `tests/test_light_apps.py`. | | **Menu Overrides** (`/api/menu-overrides`) | Sidebar-menu overrides (菜单管理), **global/shared** (no user scoping; `updated_by` audit-only). The static menu tree (icons/routes/`adminOnly`) lives in the **frontend** code (`core/page-layout/sidebar-menu.ts`); iframe light-apps are injected at render time. This router persists a per-node **overlay** the frontend merges onto that tree: `parent_id` (re-parent, moves the whole subtree; NULL = top level), `sort_order` (reorder among siblings), `label` (rename; NULL keeps code label), `disabled` (停用 — hides the node **and its subtree**). A node with no row renders exactly as code defines it, so new code menus / newly-registered iframe apps appear automatically and are immediately configurable. `GET /` - list all overrides (open to any authenticated user); `PUT /` - **admin only**, atomically **replaces the whole set** (`{overrides: [...]}`); rejects re-parenting that would exceed **3 levels** or form a cycle (`_validate_depth`, HTTP 400). Backed by `deerflow.persistence.menu_overrides` (table `menu_overrides`, auto-created by `create_all`); store wired in `deps.py` as `app.state.menu_override_store`. Frontend: merge logic `core/page-layout/menu-overrides.ts` (`applyMenuOverrides` in `use-sidebar-menu.ts`), client `strategy-components/api/menu-overrides.ts`, admin page `pages/MenuManagementPage.tsx` (@dnd-kit drag-reorder/re-parent + 「移动到」parent selector + ↑↓ arrows + 改名/停用), entry in the bottom-left settings dropdown (`components/page-sidebar.tsx`, admin-only). Tests: `tests/test_menu_overrides.py`, `frontend-web/src/core/page-layout/menu-overrides.test.ts`. | | **Compact Menu Overrides** (`/api/compact-menu-overrides`) | Compact/sidebar-summary left-nav overlay (简介模式菜单), **global/shared**. Two independent sets keyed by `menu_type` (`A` default / `B`); saving one type does not wipe the other. Type B rows use a `B::` node_id prefix so they stay unique on databases that still use `node_id` as the sole PK. `GET ?menu_type=` lists one type (open); `PUT {menu_type, overrides}` admin-only replaces that type. Frontend: `MenuManagementPage` 简介模式 type switcher + `runtime-config.js` `VITE_COMPACT_MENU_TYPE`. Tests: `tests/test_compact_menu_overrides.py`. | | **Sensitive Words** (`/api/sensitive-words`) | Sensitive-word / desensitize map (敏感词管理 / 脱敏词), **global/shared** (no user scoping; `updated_by` audit-only). Each row maps an original `word` → `replacement`, with `enabled` (停用 parks a word without deleting) + `sort_order`. The **frontend** ships a seed map in code (`src/lib/desensitize.ts` `SEED_DESENSITIZE_MAP`) used as the default on a fresh deployment; once an admin saves, this table is the source of truth and the frontend loads it at startup (`DesensitizeWordsLoader` → `replaceDesensitizeMap`, **enabled** rows only). `GET /` - list all words (open to any authenticated user — every page desensitizes); `PUT /` - **admin OR the `lqq` account** (`_require_manager`: `system_role == "admin"` or email `@`-prefix `== "lqq"`), atomically **replaces the whole set** (`{words: [...]}`, de-duped by `word`, last wins). Backed by `deerflow.persistence.sensitive_words` (table `sensitive_words`, PK `word` `String(191)`, auto-created by `create_all`); store wired in `deps.py` as `app.state.sensitive_word_store`. Frontend: client `strategy-components/api/sensitive-words.ts`, page `pages/SensitiveWordsPage.tsx` (route `/page/strategy/admin/sensitive-words`, entry in the bottom-left settings dropdown — **gated on username `lqq`, not admin role**, present in both sidebar modes; 「导入种子词」 merges missing seed entries). Tests: `tests/test_sensitive_words.py`. | | **Sentiment Agent Proxy** (`/api/sentiment-agent`) | BFF for the frontend 舆情分析 virtual agent. `POST /stream` accepts the AG-UI body plus `apiUrl`/`authUsername`/`authPassword` from `runtime-config.js`, forwards to that URL with TLS verify off, and byte-pipes upstream SSE. Connection/HTTP failures return the full error in `detail`; mid-stream failures emit AG-UI `RUN_ERROR`. Distinct from `skill-packages/sentiment-analysis.skill`. Tests: `tests/test_sentiment_agent_proxy.py`. | | **Dashboard Sessions** (`/api/dashboard-sessions`) | 「大屏绘制智能体」独立页面会话,**按外部 `task_id` 存取,不按 user 分权**(open read + shared writes;`user_id` audit-only)。一个 task 一行(upsert):`transcript`(聊天记录 JSON)+ `report_json`(最新结构化数据,可解析 = 「有效结果」)。`GET /by-task/{task_id}` - 取该 task 会话(无则 404);`PUT /by-task/{task_id}` - upsert(`{transcript?, report_json?}`,只覆盖传入项)。Backed by `deerflow.persistence.dashboard_sessions`(table `dashboard_sessions`, auto-created by `create_all`,无迁移);store wired in `deps.py` as `app.state.dashboard_session_store`。前端独立页 `roundtable-planning/pages/DashboardAgentPage.tsx`(路由 `/page/strategy/dashboard-agent?taskId=`,侧边栏「问答管理 → 大屏绘制」)让智能体分析任务报告 JSON → report-json → 渲染 `DashboardApp`。Tests: `tests/test_dashboard_sessions.py`。 | | **Task Reports** (`/api/task-reports`) | 任务级研判情报报告详情,**按 `task_id` 共享、不按 user 分权**(写入者只留 `updated_by` 审计)。`GET /by-task/{task_id}` 始终返回原 Consumer 兼容外壳 `{state:"200", msg:"操作成功!", data:[...]}`(未保存为 `data: []`);`POST /by-task/{task_id}` 首次保存,已有报告返回 409;`PUT /by-task/{task_id}` 原子替换完整 `data` 数组。`POST /import/by-task/{task_id}` 面向 TaskCOP 导入页,接收 `sq-report-mock.json` 兼容的 `data`/`records` 数组并完整覆盖目标任务;路径中的选定任务 id 优先于文件内的 `taskId`,避免误写。每一项严格是 `{id, taskId, createTime, updateTime, sessionId, contentJson, categoryType}`;`contentJson` 以 JSON **字符串**原样落在 `records_json`,避免破坏大屏按类目解析的输入格式。保存或导入非空 `data` 后,router 会通过 `app.state.taskcop_task_store.mark_analysis_completed` 将同 id 的 TaskCOP 任务升级至状态 `25`,使旧任务列表显示并筛选为“已完成”;普通保存时不存在匹配任务的报告仍可独立保存,导入则要求目标任务存在。Backed by `deerflow.persistence.task_reports`(table `task_reports`, auto-created by `create_all`);store wired in `deps.py` as `app.state.task_report_store`。前端客户端 `strategy-components/api/task-reports.ts`,`core/auth/dashboard-input.ts` 与 rwfx 3q 完成检查均直接读此接口,不再依赖 Consumer 或本地 sq-report mock。Tests: `tests/test_task_reports.py`。 | | **Task Buttons** (`/api/task-buttons`) | Task-workspace configurable jump buttons (按钮管理), **global/shared** (no user scoping; `updated_by` audit-only). Each row = one jump button in a task deep-link workspace's bottom 快捷跳转 bar: `{id, business (rwfx/3qfx/xdfx), label, link_type (url/business/purpose/host), target, append_task_id, id_param (editable query-param name, e.g. id/task_id), id_kind (task/action — only xdfx differs, its deep-link id is an 行动id), open_mode (blank/self), login_params, append_auth, auth_token_param, auth_name_param, enabled, sort_order}`. **Two independent, coexisting login-append mechanisms for `url` buttons** — both resolved frontend-side: (1) `login_params` (JSON column `login_params_json`, `PortableJSON`, migration `20260625_01`) = a list of `{key, source, custom_value}` resolved at click time (`source` = a `userInfo` field or `other`→literal `custom_value`) and appended to the target URL — mirrors the light-app login params; empty = nothing appended. (2) `append_auth` + `auth_token_param` + `auth_name_param` (columns, migration `20260625_02`, default off) — appends `?{auth_token_param}={token1}&{auth_name_param}={username}` so the target system can identify the current user; gated on **token login** — the frontend only stashes `token1` (the `?authToken=`) + the raw `getTokenInfo` username (`login.token1`/`login.username`) when login went through `/login/token`, so a non-token-login session appends neither. The **frontend** holds code-seed defaults (rwfx's 4 historical buttons built from the live `task_deeplink` config); only when this table is empty do the workspace + admin page fall back to them, else this table is the sole source of truth. `GET /` - list all (open to any authenticated user — the workspace renders them); `PUT /` - **admin only** (`_require_admin`), atomically **replaces the whole set** (`{buttons: [...]}`, de-duped by `id`, last wins). Backed by `deerflow.persistence.task_buttons` (table `task_buttons`, auto-created by `create_all`); store wired in `deps.py` as `app.state.task_button_store`. Frontend: helpers `core/auth/task-buttons.ts`, client `strategy-components/api/task-buttons.ts`, admin page `pages/TaskButtonsPage.tsx` (route `/page/strategy/admin/task-buttons`, entry「按钮管理」in the bottom-left settings dropdown, admin-only, both sidebar modes; per-business tabs + 「导入默认」). Tests: `tests/test_task_buttons.py`. | | **Fixed Questions** (`/api/fixed-questions`) | Administrator-managed exact-match fixed Q&A (问题配置), **global/shared**. Each row is `{id, question, answer, enabled, tokens_per_second, sort_order}`; stream speed is configurable per row from `1–1000 token/s` and defaults to `100`. `GET /` and whole-set atomic `PUT /` are both admin-only. At the start of an external chat run, `services._find_fixed_answer` checks the latest eligible plain-text human message with exact, case/space/punctuation-sensitive equality. Enabled matches switch the run to `make_fixed_answer_agent`: a local `BaseChatModel` tokenizes with `cl100k_base`, replays the answer in rate-limited token chunks, and keeps the existing LangGraph SSE `messages`/`values` protocol, run journal, checkpoint history, deterministic local title, and refresh behavior intact while no provider/model request or model-concurrency slot is used. File/multimodal/context-prefixed turns, `canvas_agent`, `ai_writing`, writing mode, resume commands, and internal orchestration/background child runs bypass matching. Storage failures fail open to the normal model path. Backed by `deerflow.persistence.fixed_questions` (table `fixed_questions`, SHA-256 indexed lookup plus exact original-string collision check, migrations `20260821_01` + `20260821_02`), wired as `app.state.fixed_question_store`. Frontend: `strategy-components/api/fixed-questions.ts`, admin page `pages/FixedQuestionsPage.tsx`, route `/page/strategy/admin/fixed-questions`, entries in both full-layout and compact-workspace settings menus. Tests: `tests/test_fixed_questions.py`. | | **Knowledge** (`/api/knowledge`) | Product knowledge base, **global/shared** (no user isolation; writes record `created_by`/`updated_by` for audit). `POST /ingest/thread/{thread_id}` - sediment one thread into a note (reads messages from the checkpointer, `threads:read` owner-checked); `POST /ingest/threads/batch` - batch-sediment the caller's recent threads (`skip_existing` via `.manifest.json` dedup). **Defaults to `background=true`**: the slow per-thread extraction loop is deferred to a FastAPI background task and the endpoint returns `{queued:true, processed:N}` immediately (each thread's transcript is truncated to `max_input_chars` before the LLM call). Running the loop inline for a large batch outlives the nginx gateway timeout (504), which the frontend's `parseResponse` surfaces as a generic "Request failed"; `background=false` keeps the blocking path (full per-thread `created/skipped/failed` tally) for small batches/scripts. Shared loop: `_run_batch_ingest`. **Optional model selection**: both ingest endpoints accept `model_name` (overrides `knowledge.extract_model_name` for that capture; threaded through `KnowledgeService.capture_thread` ? `extract_thread_llm_multi`). **Model-error surfacing**: when the chosen model produces nothing usable (timeout / API error / non-JSON) capture still **falls back to rule-based extraction and lands a draft** (`llm_fallback:true`); the single-thread (non-background) response and the batch response (`llm_fallback` count) carry this so the UI tells the user "???????". The "??????" chat button stays background/fire-and-forget (passes `model_name`, no per-save error toast); the batch-import dialog is synchronous and shows the fallback count. `POST /ingest/file` - **import an uploaded file** (multipart: `file` + optional `title`/`tags`/`status`/`model_name`) into wiki note(s): the file is converted to Markdown (md/txt read directly, pdf/word/ppt/excel via `convert_file_to_markdown`), distilled by the same wiki-capture LLM pipeline as thread capture (`source_type="file"`), and falls back to a verbatim **draft** note when the model returns nothing usable (`llm_fallback:true`); `POST /ingest/search` - sediment search/tool results as a reference note; `POST /notes` - manual note; `GET /notes` (keyword/tag/source_type/status + pagination); `GET/PUT/DELETE /notes/{id}` (DELETE is soft = archive, `?hard=true` removes the vault file); `POST /search` - `mode=keyword\|vector\|hybrid` (vector?keyword fallback when no embedding endpoint); `GET /notes/{id}/versions` - edit history; `GET /duplicates` + `POST /merge` - dedup; `POST /reindex` - rebuild vector index; `GET /graph` - product graph data (notes + entities, graph-blacklist filtered); `GET/POST /graph/blacklist` + `DELETE /graph/blacklist/{id}` - graph node blacklist (labels = entity names / note titles, case-insensitive match; blacklisted nodes and their edges are dropped from `/graph` and the wiki-export artifacts; table `knowledge_graph_blacklist`, auto-created); `POST /export/graph` + `GET /export/files/{graph.json\|graph.html\|graph.graphml\|cypher.txt}` - obsidian-wiki `wiki-export` artifacts (self-contained `graph.html` for iframe); `POST /resolve` `{targets:[?]}` + `GET /notes/{id}/wikilinks` - resolve `[[wikilink]]` targets to a note id / vault page / missing (powers clickable wikilinks); `GET /vault-page?path=` - read a raw vault meta/hub Markdown page (no DB note). `GET/POST /folders` + `PUT/DELETE /folders/{id}` - user-managed directories (notes carry a `folder` path; `GET /notes?folder=` filters a dir + sub-dirs); `GET/POST /extract-templates` + `GET/PUT/DELETE /extract-templates/{id}` - editable extraction-prompt templates (thread/document ingest accept `template_id`; capture/PUT-note accept `folder`). Backed by the `deerflow.knowledge` module (see below). | Proxied through nginx: `/api/langgraph/*` ? LangGraph, all other `/api/*` ? Gateway. ### TaskCOP Situation-Overview Compatibility The standalone legacy Vue page calls the Gateway URL in its runtime `b1ConsumerUrl` configuration directly (no Vite `/api` reverse proxy). The Gateway CORS middleware merges standard local Vite origins (`localhost`/`127.0.0.1` on ports 5173, 5174, 3000, and 8080) even when `GATEWAY_CORS_ORIGINS` was set for another frontend. This avoids a restart regression when the existing Vite server is on 5173 but the launcher sets 5174. Deployments must set `GATEWAY_CORS_ORIGINS` to any additional exact frontend origins that are allowed to call the Gateway. `app/gateway/routers/taskcop_tasks.py` is a temporary compatibility layer for the separately deployed legacy Vue page at `/situation-overview/action/sentiment` and its task import page at `/situation-overview/action/task/import`. It preserves the old Consumer paths and response envelopes: `POST /cop/saveSuperiorTask` creates a row with a sequential four-digit string id (`0001`, `0002`, ...) and returns `{state: 201, msg, data}`; `POST /api/taskcop/import/tasks` accepts a JSON array or `tasks`/`records`/legacy list envelope, uses each supplied `id` or `taskId` as the primary key, and completely replaces duplicate IDs; `DELETE /cop/tasks/{task_id}` removes that task and its shared report document; `GET /taskAnalyseSearch/cop-task-three-list` (all), `GET /taskAnalyseSearch/cop-my-tasks` (creator-scoped), and `GET /taskAnalyseSearch/cop-task-list` return `{state: 200, data: {total, records}}`; `GET /taskAnalyseSearch/cop-task-detail?taskId=` returns one task. Existing UUID-style task rows are retained, and do not affect the new numeric sequence. An empty TaskCOP store remains empty—there are no built-in task or report seed records. It supports the page's existing `content`, `taskStatus`, `taskDirection`, `startTime`, `endTime`, `pageNum`, and `pageSize` query fields. Tasks are stored in `taskcop_tasks` through `deerflow.persistence.taskcop_tasks`, registered by `persistence.models`, and initialized as `app.state.taskcop_task_store` in `deps.py`. A non-empty save/update/import through `/api/task-reports/*` calls `mark_analysis_completed(task_id)`: it changes only pre-completion task states to `25`, so the list renders and filters the task as “已完成” without downgrading later workflow states. `trendSucc` is derived from the same status for the legacy sentiment list. Zero-padded task ids stay strings in the report API, so `0001` is never converted to `1`. The Vue router calls DeerFlow's `/api/v1/auth/login/username` directly with an incoming `username`/`userId`/`yUserId` parameter; it does not depend on the retired Consumer login service, and sends the returned DeerFlow JWT to every TaskCOP call. The TaskCOP routes are no longer public or CSRF-exempt. The Gateway fallback CORS allow-list includes TaskCOP Vite origins `http://localhost:8080` and `http://127.0.0.1:8080`; deployments must use explicit `GATEWAY_CORS_ORIGINS`. Tests: `tests/test_taskcop_tasks.py` and `tests/test_taskcop_cors.py`. ### AI Writing Intent Router (对话驱动写作台意图路由) Backs the frontend「AI 写作·对话版」page's unified bottom composer. `POST /api/ai-writing/sessions/{session_id}/intent` (owner-checked; in `routers/ai_writing.py`) reads the session's graph checkpoint via the shared `app.state.checkpointer` — the **pending interrupt is the authoritative pause point** (the frontend's claimed `pause_point` is only a fallback when the checkpoint can't be read) — then routes the user's free text through `deerflow/agents/ai_writing/intent_router.py::parse_user_intent`: an LLM call (session's `model_name` + `fallback_chain` model fault-tolerance) that returns `{intent, pause_point, action, payload, need_research, research_query, answer, confidence, clarify}`. Safety boundaries in `normalize_intent_result`: **per-pause action whitelist strictly aligned to graph.py's real resume actions** (`ALLOWED_ACTIONS`), no pause → no action, high-cost actions (`finalize` / `force_finalize` / query-less覆盖式 `re_search`) require confidence ≥ 0.85 else downgrade to `clarify`, payload key-whitelist cleaning (material-id filtering, review-issue index bounds, section-decision back-fill with the majority decision). QA intents return `answer` directly from the context snapshot (`build_context_snapshot` — stage-tailored: materials with ids / outline / draft body capped / review issues with indices). Model failures never 5xx — they degrade to clarify. Prompts in `agents/ai_writing/prompts/intent_router_prompts.py`; tests `tests/test_ai_writing_intent.py`. ### AI Writing Built-in Functional Agents (????????) The AI-writing graph (`packages/harness/deerflow/agents/ai_writing/`) has four stage "brains" ? ?????? / ????? / ?? / ?? ? that can each run in one of two modes, switched per-stage by `config.yaml ? ai_writing.use_builtin_agents.{researcher,outliner,writer,editor}` (default **false**, mtime hot-reloaded, new sessions pick up changes immediately): - **Legacy (flag off)**: hard-coded prompts in `prompts/*.py`, direct LLM call (unchanged behavior). - **Builtin-agent (flag on)**: the node loads the corresponding functional agent from `.deer-flow/agents/ai-writing-/` ? `SOUL.md` becomes the system prompt and `config.yaml`'s `model` overrides the stage model (priority: `agent.model > node explicit fallback (keyword_model/rank_model) > state.model_name`). Admins edit SOUL/model from the agent management page. Runner: `nodes/_agent_runner.py` (`run_writing_agent` / `stream_writing_agent` / `parse_marker_json`). Output contracts (marker + ```json``` blocks, intent.py-style): researcher `[KEYWORDS_READY]` / `[MATERIALS_RANKED]`, outliner `[OUTLINE_READY]`, editor (per-dimension) `[REVIEW_READY]`; writer emits raw section Markdown (keeps per-section streaming) with the legacy `[[SECTION_BLOCKED]]` sentinel in strict mode. **Orchestration stays in the nodes** (search fan-out, per-section loop, 3-dimension review loop, all pause points) ? the agent only does single structured generations, and a missing/empty SOUL silently falls back to the legacy path (`WritingAgentUnavailable`). Seeding mirrors the roundtable pattern: verified `config.yaml` + `SOUL.md` ship in `app/gateway/routers/_ai_writing_seed_assets//`; `_ai_writing_seed.py::ensure_ai_writing_functional_agents()` copies them to `.deer-flow/agents/` when missing (never overwrites local edits), called from the startup lifespan (before `_sync_legacy_agents`, so the agents appear in `/api/agents` on first boot) and self-heals from `GET /api/ai-writing/sessions` + `GET /api/ai-writing/article-types`. Tests: `tests/test_ai_writing_seed.py`, `tests/test_ai_writing_agent_runner.py`. **Sample imitation mode (样文仿写).** The initial writing form supports `writing_mode_type=imitate`: the user can paste sample text or upload Word/PDF/Markdown/TXT. `POST /api/ai-writing/sample/extract` converts supported document files to Markdown via `deerflow.utils.file_conversion` and returns capped plain text to the frontend. The LangGraph entry point now starts at `sample_analyzer` before `intent_parser`; non-imitation sessions return immediately, while imitation sessions analyze up to 12k chars into `sample_style_profile` (`SampleStyleProfile` in `state.py`) and emit `sample_profile_ready`. `writer_outline` and `writer_draft` call `render_sample_style_section(state)` so both outline planning and section drafting receive the same style profile, imitation strength (`light|medium|strong`), optional structure-preservation preference, and guardrails: learn structure/tone/paragraph rhythm/sentence style, but do not copy sample sentences, proprietary facts, people, organizations, cases, or data. Frontend state mirrors this with `analyzing_sample`, uploaded/pasted sample controls, history/timeline restoration of `sampleStyleProfile`, and the setup-chat direct-start path preserves the same fields. Tests live with the knowledge-base AI-writing tests (`tests/test_ai_writing_knowledge_base.py`). **Skill-driven retrieval (检索类型).** `material_source` has three values: `general` (通用检索, default), `knowledge_base` (知识库检索), and `notebook` (我的空间); legacy `skill` is normalized to `general` by the researcher node. **`knowledge_base` 直连内网 ES 接口(= 旧版「内网直连检索」)** — it bypasses skill orchestration entirely and calls `IntranetSearchProvider` over `ai_writing.intranet_search_url` (concurrent over the generated keywords; `intranet_knowledge_base`/`intranet_search_timeout` apply). **An empty/unconfigured `intranet_search_url` does NOT disable it** — the empty string is passed straight to `IntranetSearchProvider`, which falls back to its built-in default endpoint (`intranet_stub._DEFAULT_ENDPOINT`), exactly like the legacy `get_search_provider(search_provider=intranet)` path. So 内网直连检索 works out of the box without configuring an address (configure `intranet_search_url` only to point at a different endpoint). This is the "弱模型识别不了技能" escape hatch — the old no-skill retrieval path restored as an explicit option (weak models like deepseek-chat often fail to invoke configured skills). Tests: `tests/test_ai_writing_knowledge_base.py`. General retrieval is orchestrated from the **skills configured on the `ai-writing-researcher` agent** (`config.yaml → skills`, editable from the agent management page): the built-in `knowledge-base-search` skill (知识库检索) runs the native keyword fan-out web/intranet search (the old default path), while any other configured skill searches with a `description`-targeted query; skills run in configured order with early-stop at `max_materials`, and an empty/unresolvable skills config falls back to the built-in knowledge search so retrieval always works. The skill itself is seeded from `_ai_writing_seed_assets/skills/knowledge-base-search/SKILL.md` into `skills/public/` (never overwrites), and a pre-existing researcher `config.yaml` lacking the `skills` key is patched in place with the default (explicit `skills: []` is respected). **PAUSE-1 input append flow**: when the user types natural language into the material-confirm input (`re_search`/legacy `enable_skill_search` + `user_query`), the node matches the most relevant configured skill via `SkillPicker(candidates=…)` (skipped when only one skill is configured), searches with the user's words, and **appends** the new materials to the existing package (URL-dedup, original order preserved, no LLM re-rank/truncation, 100-item hard cap). **Skill-failure fault tolerance**: a configured skill that times out / raises / has all its keyword searches fail (vs. legitimately returning 0 results) no longer just silently moves on — the researcher node emits a `skill_search_step` of type `skill_error` (carrying the error `reason`), and if the configured skills collected **no** new materials at all it runs a `fallback` pass over other available skills (`_fallback_candidate_skills`: the built-in `knowledge-base-search` native path first, then other enabled catalog skills not yet tried) instead of erroring out. The frontend `SkillStepsCard` renders `skill_error` (rose) and `fallback` (amber) so the timeline shows "技能报错 → 改用其他技能搜集". Tests: `tests/test_ai_writing_material_source.py`. The intranet client uses `trust_env=False` and defaults `intranet_verify_ssl=false`, so internal HTTPS endpoints with self-signed/internal-CA certificates do not fail on certificate verification; set `ai_writing.intranet_verify_ssl: true` only for environments that require strict certificate checks. **Skill-call timeout** is hardcoded **300s (5 min)** — `researcher.py` `_SKILL_SEARCH_TIMEOUT_SECONDS` (per-skill wrapper) and `_skill_agent.py` `_REAL_AGENT_TIMEOUT_SECONDS` (whole real-agent skill run). The intranet ES fallback (`IntranetSearchProvider`) gets a longer **600s (10 min)** default (`intranet_stub.py`; overridable via `ai_writing.intranet_search_timeout`) — that interface is the slow one, so our own fallback timeout is generous to avoid being blamed for "检不出". **Parallel intranet fallback (技能 0 条 → 直接用并行预跑的内网结果).** When `ai_writing.intranet_search_url` is configured, the researcher node launches the intranet ES search (`IntranetSearchProvider` over the generated keywords) as a **concurrent `asyncio.Task` the moment the skill loop starts** (not sequentially after all skills fail). If the configured + fallback skills collect **nothing**, the node awaits that already-running task and uses its results (`_collect`) — no extra wait. If the skills **did** get materials, the task is cancelled (`_cancel_intranet_task`, suppresses CancelledError) so the intranet call isn't wasted. URL empty → no task, behavior unchanged. Tests: `tests/test_ai_writing_intranet_parallel.py`. **Partial-result harvest + query simplification (检不出/超时的容错).** When `researcher_skill_agent: true` runs a non-native skill via `run_skill_via_real_agent` (the deployment's real path — the skill, e.g. `knowledge-search-v2`, is invoked by the real agent through `bash python3 scripts/search.py ""`), one writing run issues many queries and **some time out** (the skill script prints `{"success":false,"error":"请求超时…","count":0}`) while others succeed (`{"success":true,"count":30,"results":[…]}`). Two fixes so successful data is never dropped: (1) `_drive()` captures every **ToolMessage** stdout and `_harvest_tool_results()` parses the skill's JSON **by reusing the lead-agent 参考文献 mode's proven extractor** (`deerflow.runtime.references`: `_parse_json_like` + `_first_list` + `_normalize_reference_item`) — the same logic that already extracts skill JSON into citations 100% reliably on ChatPage/AgentChatPage, with its huge field-name union (title/m_title/name, content/page_content/content_preview/snippet/…, recUuid/uuid/docId/…) and per-skill display-config support — plus skips `success:false` outputs — so even if the model's final `{materials:[]}` summary gives up (confused by the timeouts), the successful queries' real results are still collected; the function merges model materials + harvested results (dedup by url/recUuid/title, capped at `max_results`). **The harvest+merge (`_finalize`) runs on EVERY exit path — normal end, overall timeout, AND `await task` exceptions (e.g. `GraphRecursionError`)** — so when the run is cut off before the model writes its summary (observed: 10 queries × ~2 graph steps > the old `recursion_limit`, so the agent errored out with the real 30+ results already in `tool_outputs`), those results are still salvaged instead of returning `[]`. `recursion_limit` was also raised from `_REAL_AGENT_MAX_TURNS` to `_REAL_AGENT_MAX_TURNS*2+4` (it counts graph super-steps ≈ 2 per model→tool round, not model turns) so the model normally has room to finish its summary. **Early-stop on `max_results`:** `_drive` re-harvests after every tool output and, once the deduped harvested count reaches `max_results` (= `per_skill` = `max_materials`), **breaks the nested `astream`** and emits a 「已获取 N 条素材,达到上限,停止本技能检索」 step — without it a weak model keeps issuing new queries forever (observed: 100+ step-bar entries, materials far over the limit). Because the skill now returns exactly `max_materials`, the researcher loop's own `collect_limit` early-stop fires right after, so later configured skills don't run either. (2) The model is instructed (in `_MATERIAL_OUTPUT_CONTRACT`) that a 0-result/timeout query must be **retried with a simpler, core query** (drop topic suffixes like 竞选/诉讼/政策, keep the core entity: 「特朗普竞选」→「特朗普」) and that any data already retrieved must always be kept. The same deterministic simplifier `search/base.py::simplify_query` is wired into `IntranetSearchProvider.search`: an empty first attempt retries once with the simplified query. **(3) Root cause of the instability — bash output middle-truncation.** The skill prints its results as JSON to stdout, and the `bash` tool **middle-truncates** any output over `sandbox.bash_output_max_chars` (default **20000**, inserting a `... [middle truncated …] ...` marker). A `count:30` skill response (~30 records × ~600 chars) exceeds 20000 → the JSON is cut in the middle → **unparseable as a whole** by both the model and `_try_load_json` → all materials lost (while a small `count:9` response parses fine — hence the flapping). Fix: when whole-JSON parse yields no items, `_harvest_tool_results` falls back to `_scan_balanced_json_objects` — a brace-depth/string-aware scanner that recovers **every individually-complete `{…}` record object** from the truncated text (the records on each side of the truncation marker are still valid JSON), so ~16+16 records survive a 20000-char cut, far above `max_materials`. This is why the lead-agent chat pages (`/page/workspace/chats/new`, agent chat) call the same skill with 100% success — they only need the model to summarize prose from whatever it sees, never to parse a structured list, so truncation doesn't hurt them; ai-writing needs the structured records, so it must salvage them itself. Tests: `tests/test_ai_writing_skill_real_agent.py`, `tests/test_ai_writing_query_simplify.py`. **Agentic skill execution (`ai_writing.researcher_skill_agent`).** When this flag is on (the deployment `config.yaml` enables it), the directed-skill branch runs `nodes/_skill_agent.py::run_skill_via_real_agent`: it reuses `SubagentExecutor` with the full tool stack, sandbox middleware, and a one-skill whitelist so the skill can `read_file` its `SKILL.md`, call `bash`, MCP tools, or whatever the skill instructions require. The nested run now passes `agent_id=ai-writing-researcher` in both `config.configurable` and runtime `context`, matching normal AgentChat skill visibility/whitelist behavior. It also appends a human-turn hard directive listing the exact SKILL.md path and skill directory, requiring weak models to first `read_file` the SKILL.md and then `cd && python/python3 scripts/*.py ...` when the skill is script-backed; the model is explicitly forbidden from returning `materials` before executing the skill's required tools. If the real-agent path fails or yields no materials, `_execute_skill` still falls back to the legacy description→query search so materials are never empty. The built-in `knowledge-base-search` native concurrent path is unchanged. Tests: `tests/test_ai_writing_skill_real_agent.py`. **Targeted user revision (局部修改 — 只改命中章节).** When the user submits a free-text 修改意见 on the draft-confirm card (`action == "user_revise"`), `writer_draft_node` no longer always rewrites the whole article section-by-section. Before the full-rewrite loop it tries `_targeted_user_revision` (`nodes/writer_draft.py`): `_split_draft_sections` parses the previous `current_draft.full_markdown` into `(article_title, [(section_title, body)], references_block)` (level-1 `#` title, level-2 `##` sections, trailing `## 参考文献` pulled out and kept verbatim), then `_detect_revision_scope` runs a small `llm_json` classifier (prompts `DRAFT_REVISE_SCOPE_{SYSTEM,USER}`) returning **per-section targets** `[{index, new_title, rewrite_body}]`: it maps the note to the section indices it targets ("第一段/开头/引言"→0, "结尾/结论"→last, by title/topic) **and** distinguishes a **章节标题改名** (`new_title` set, `rewrite_body=false` → rename only, body untouched) from a **正文修改** (`rewrite_body=true`). Title-only renames apply directly to the assembled `## {new_title}` heading **and are written back into `current_outline.sections[i].section_title`** (so the renamed title persists through confirm/finalize/history — this fixes "改章节标题不生效"). Body rewrites run `_stream_revise_section` (prompt `DRAFT_REVISE_SECTION_USER`, streamed as `draft_chunk`). **Only hit sections are touched; unhit sections are preserved byte-for-byte and are NOT re-emitted to the stream** — the frontend clears the right canvas to its initial placeholder on entering `revising` and shows only the live rewrite of the changed section(s), so the user sees「清空 → 流式重写 → 终稿」rather than the whole article rebuilding. The classifier returns `whole_article=true` (整体性意见 like 「整体语气更正式」/「全文都改」, **structure changes** 加章/删章/调整顺序, or a **文章总标题** change) → `_detect_revision_scope` yields `None` → `_targeted_user_revision` returns `None` → the node **falls back to the whole-article rewrite** (which still runs `_revise_outline_with_notes` for title/structure changes). Also falls back when there is no prior draft, an empty note, or scope-detection errors. Editor-review revisions (`accept_review`) keep the holistic rewrite. Reassembly preserves the article title + structure and re-appends the original references block; word count is computed on the body (excluding references). Frontend: `isRevising` in `AIWritingState` (`types/ai-writing.ts`) is set in `useAIWritingStream.ts`'s `intervention_submitted` (`ns === 'revising'`) and cleared on `draft_ready`; `AIWritingDraftPanel.tsx` shows the initial `DraftWaitingPlaceholder` while `isRevising && streamingDraftSections.length === 0`, then the streaming preview once chunks arrive. Tests: `tests/test_ai_writing_targeted_revision.py`. **History draft editing (历史回看可编辑成稿).** The right-side draft in 历史回看 is no longer read-only — the user can edit it like the live 写作完成态 (Tiptap + 目录 + 格式工具栏 + AI 润色) and **persist** the change back to the session. `PUT /api/ai-writing/sessions/{id}/draft` `{draft_markdown, draft_title?}` (owner-checked → 404 otherwise; `draft_title` omitted ⇒ keep原标题) updates the row via the existing `AIWritingSessionRepository.update`. Frontend: `AIWritingDraftPanel` now passes `readOnly={false}` + an `onSaveDraft` handler (only in history mode) → `TextArtifactRenderer` → `WritingToolbar` 「保存」button (loading + toast); the handler is `AIWritingHistoryContext.saveHistoryDraft(markdown, title?)` → `updateAIWritingDraft` (`api/ai-writing-sessions.ts`), which write-backs the returned summary to `historySession`. Edits to the **live** completed draft stay ephemeral (export-only) as before — only history mode shows 保存. Tests: `tests/test_ai_writing_draft_update.py`. **Background suspend (后台挂起).** A writing session can be handed off to the server to auto-drive the remaining flow to finalize, surviving page leave / refresh / account switch. `POST /api/ai-writing/sessions/{id}/background` (owner-checked, idempotent) starts `AIWritingJobExecutor` (`app/gateway/ai_writing_job_executor.py`), wired in `deps.py` as `app.state.ai_writing_job_executor`. The executor **drives the same `ai_writing` graph in-process** — it builds `make_ai_writing_graph()`, attaches the shared `app.state.checkpointer` (the very saver the runtime worker uses, so it continues the real thread state the frontend already wrote), sets the owning user via `set_current_user`, and loops `aget_state` → `ainvoke(Command(resume=…))` until `status==done`. **Auto-action per interrupt**: material/outline confirm → `confirm`; section_help → every blocked section `loose` (continue with general knowledge, never re-loop to research); draft_confirm → `to_editor` until reviewed, then `finalize` (on pass, or when `revision_count >= max_revisions`); review_confirm → `accept_review` (revise & re-review until pass or budget exhausted, then forced finalize). To avoid racing a still-running runtime worker (suspend mid-node), the driver **waits for an actionable state** (`_wait_for_actionable`): it only acts when the graph is at an interrupt / done, or when the checkpoint stops advancing (no worker → abandoned/restart). No new table — the graph checkpoint + `ai_writing_sessions` already hold all state; `status="background"` marks 挂起中 (surfaced in the history dropdown), nodes write `draft`/`review_result`/`status=done` as usual, so the user reopens the finished article from history. Per-user concurrency cap `AI_WRITING_BACKGROUND_SLOTS_PER_USER` (default 2, same SQLite-lock rationale as the roundtable executor). Startup `reconcile` (in `deps.py`) relaunches drivers for sessions left `background` after a restart (`AIWritingSessionRepository.list_by_status`). **Progress flowchart popup**: `GET /api/ai-writing/sessions/{id}/progress` (owner-checked) reads the graph checkpoint via `executor.get_progress` and derives a canonical pipeline `stage` key (`intent/research/material_confirm/outline/outline_confirm/draft/draft_confirm/review/done`, from the pending interrupt's `pause_point` or `snapshot.next` node) + `note` (素材不足求助 / 编辑打回) + review progress + `isRunning`. Frontend: 「后台挂起」 button + auto-suspend on route unmount / `beforeunload` (`AIWritingPanel.tsx`, `suspendAIWritingToBackground[Keepalive]` in `api/ai-writing-sessions.ts`); 挂起中 badge in `AIWritingHistoryDropdown.tsx`, and **clicking a 挂起中 history record opens a flowchart dialog** (`AIWritingFlowDialog.tsx` polls `getAIWritingProgress` every 2.5s, renders the fixed vertical pipeline `AIWritingFlowChart.tsx` lit by the live `stage`) so the user sees how far the background run has gotten without reopening/taking over the session. **Transcript consistency**: the rich history timeline (`AIWritingHistoryView` → `AIWritingTimeline`) replays `transcript = {progressEvents, completedInterventions}`, normally assembled+saved client-side. A background run has no client, so the executor rebuilds it at the end (`_save_transcript`/`_build_transcript_from_checkpoint`) **entirely from the final checkpoint** — both `progressEvents` and `completedInterventions` are derived from the same set of `resumed` events (every pause-resume emits exactly one), so they are always 1:1 aligned. For each `resumed`: classify its message → `(pausePoint, action)` (`_classify_resumed`, matched against `graph.py`'s resume messages), insert an `await_user` (with the pause prompt) before it, and build a **rich** intervention card enriched from final state (material_confirm→`materials`, outline_confirm→`outline`, draft_confirm→`draftTitle`/`draftWordCount`/`reviewResult`, review_confirm→`reviewResult`); `materials_ready` is inserted before the first material-confirm. (Earlier code merged the frontend's pre-suspend `completedInterventions` with auto ones and aligned by index — but the frontend's last save can predate its most recent decision, so the counts drifted and cards rendered off-by-one with empty data; deriving everything from the checkpoint root-fixes both.) Raw (snake_case) events **and** intervention data are normalized on load by the **exported** `normalizeProgressDict` / `normalize{Material,Outline,ReviewResult}` (idempotent on already-camelCase live transcripts — one normalize path for both sources), so reopening a background-completed session shows the same cards as a frontend-driven one. Tests: `tests/test_ai_writing_background.py`. ### Admin Leaderboard Snapshots (`/api/admin/users/leaderboard*`) The admin ????? is served from **precomputed per-day snapshots** in `admin_leaderboard_daily_stats` (keyed by `(stat_date, settings_hash)`), not by aggregating raw rows on every request. A background scheduler inside the Gateway (`_leaderboard_snapshot_scheduler_loop` in `app/gateway/app.py`, singleton file-locked) ticks every `ADMIN_LEADERBOARD_BACKGROUND_INTERVAL_SECONDS` (default 300s): it prewarms recent days, then builds any `queued`/`error`/stale-`running` day. `GET /leaderboard` merges the per-day snapshots for the requested range. **Today** is rebuilt whenever its snapshot is older than `ADMIN_LEADERBOARD_TODAY_TTL_SECONDS` (default 43200s = half a day, so the leaderboard refreshes ~twice daily); a past day, once `ready`, is treated as fresh forever. Two on-demand/automatic refresh paths layer on top: - **Manual "????"** ? `POST /leaderboard/refresh-today` force-rebuilds *today's* snapshot **synchronously** (bounded by the same semaphore/timeout the scheduler uses), then returns the merged range. Lets an admin pull the latest real-time records into the stats without waiting for the TTL. Frontend: a "????" button on `AdminUserLeaderboardPage` ? `useRefreshTodayLeaderboard` (`core/admin/hooks.ts`), which primes + invalidates the `["admin-activity","leaderboard"]` query. - **Daily finalize of the previous day** ? once per day, on/after `ADMIN_LEADERBOARD_FINALIZE_HOUR` (local Beijing, default `1` = ?? 1 ?), the scheduler tick force-recomputes *yesterday* a single time so runs/tool-calls that only landed near midnight are fully captured. The decision is the pure `deerflow.persistence.admin_stats.finalize.should_finalize_previous_day` (status/generated_at/now/hour) ? the rebuild stamps `generated_at` to today, which self-debounces it for the rest of the day (no extra state file). Tests: `tests/test_leaderboard_finalize.py`. ### Admin Concurrency Monitor (实时并发监控 / 系统压力) A live admin monitor embedded **at the top of `AdminUserLeaderboardPage`** (`components/workspace/admin/concurrency-monitor.tsx`, `ConcurrencyMonitorPanel`) that answers "当前有几个并发的大模型调用、是谁、能不能直接停掉" and charts system pressure over time. Built on the existing in-memory `RunManager` (no new run plumbing). - **Live concurrency + 真正停止** — `RunManager.list_active()` returns every run whose status is `pending`/`running` (the registry keeps finished records around until `cleanup`, so it filters by status). `len(list_active())` is the current concurrent-call count. Router `app/gateway/routers/admin_active_runs.py` (`/api/admin/active-runs`, **admin only** via the shared `admin_users._require_admin`): `GET /` returns `{count, server_time, runs:[{run_id, thread_id, status, created_at, elapsed_seconds, user_id, user_name, agent_id, agent_name, thread_title, kind, multitask_strategy}]}` — each run best-effort enriched (batched DB) with the owning user's email (`threads_meta.user_id`→`users.email`) and the agent's display name (`agents.name`); `kind` is derived from thread metadata (圆桌会商 / 系统/后台 / 任务对话 / 对话). `POST /{run_id}/cancel` calls `RunManager.cancel(run_id, action="interrupt")`, which sets the abort event **and cancels the run's asyncio task** — i.e. it truly aborts the in-flight model stream, freeing capacity. `POST /cancel-all` stops every in-flight run. - **并发量折线图 (按分钟/小时)** — a lightweight background **sampler** (`_concurrency_sampler_loop` in `app/gateway/app.py`, singleton file-locked, started in the lifespan) records `len(list_active())` every `CONCURRENCY_SAMPLE_INTERVAL_SECONDS` (default 15s) into the new `concurrency_samples` table (`deerflow.persistence.concurrency`, `ConcurrencySampleRow{sampled_at, active_count}`, auto-created by `create_all`, **no migration**; store wired as `app.state.concurrency_sample_store`). ~Hourly it purges rows older than `CONCURRENCY_SAMPLE_RETENTION_HOURS` (default 168h). `GET /concurrency?granularity=minute|hour&hours=N` aggregates samples into per-bucket **peak + avg** (`ConcurrencySampleStore.aggregate`, bucketed in Python so it's identical across sqlite/mysql/postgres; fetch capped at 60k rows). Default lookback: minute→3h, hour→48h. - **大模型出字速度折线图 (tokens/sec, 按分钟/小时)** — `GET /token-speed?granularity=&hours=` aggregates the persisted `llm_call_metrics` (`status="success"`) into per-bucket avg/min/max tokens/sec (deriving tokens/sec from `output_tokens÷duration_ms` when the stored value is null). A dip pinpoints the model/provider as the bottleneck for that window — "知道具体原因出现在哪里". - **On/off + 不影响线上** — the whole feature is designed to never impact production: sampling is one tiny INSERT, retention-purged, all paths best-effort/guarded, queries bounded. A **runtime switch** lives in `system_settings.json → concurrency_monitor.enabled` (`ConcurrencyMonitorSettings`, **default false / opt-in**, mtime-hot-reloaded): the background loop always runs (just sleeps) but only samples when an admin turns it on, and it re-reads the flag **every tick** so toggling it on/off **takes effect immediately without a restart** (the on-demand live snapshot + 停止 stay available regardless; only the historical chart needs sampling on). `GET|PUT /api/admin/active-runs/settings` reads/writes it (admin only); the panel's 「后台采样」 Switch drives it. Env `CONCURRENCY_SAMPLE_ENABLED=0` is a hard master kill (never starts the loop). Frontend client/hooks: `core/admin/active-runs.ts` + `core/admin/hooks.ts` (`useActiveRuns` polls 4s while the panel is open, `useConcurrencySeries`/`useTokenSpeedSeries` poll 15s, `useCancelActiveRun`/`useCancelAllActiveRuns`, `useMonitorSettings`/`useUpdateMonitorSettings`). Tests: `tests/test_concurrency_monitor.py`. ### Scheduled Task Runtime `ScheduledTaskService` (`deerflow/runtime/scheduler/service.py`) polls due tasks every `poll_interval_seconds` and executes each in its own throwaway thread, serialized by a shared `_agent_run_lock`. **Stuck-run recovery**: each execution attempt is bounded by `_RUN_ATTEMPT_TIMEOUT_SECONDS` (30 min) via `asyncio.wait_for`. On timeout the run is retried up to `_MAX_RUN_RETRIES` (3) times, then marked `failed`. A watchdog (`_reconcile_stuck_runs`, run every poll) recovers runs left in `running` with no in-process owner ? e.g. orphaned by a crash or restart ? re-driving them through the same retry budget so a task can never stay perpetually "executing". The `scheduled_task_runs.retry_count` column tracks consumed retries. **Unique task names**: task names are unique per user (`uq_scheduled_tasks_user_name`). `create_task`/`update_task`/`admin_update_task` proactively check via `_assert_name_available()` and raise `DuplicateTaskNameError` on a collision; the router translates that into HTTP `409` instead of letting the raw DB `IntegrityError` surface as a `500`. **HTML page tasks (`task_kind == "html_page"`)**: a second task kind that renders a full HTML document each run instead of a Markdown report. The kind + its config (`html_config`: `page_prompt`, `negative_prompt`, `reference_favorite_id`, `model_name`) live inside `execution_context` ? no schema change to `scheduled_tasks`. Data gathering is delegated to the agent's own search skills (driven by `page_prompt`), so there is no explicit search-query/data-source field; the render model is user-selectable (`model_name`, default `zai-org/GLM-5-FP8`). `_run_task_attempt()` dispatches by kind; `_run_html_page_attempt()` runs the pipeline: gather real data via the lead agent (its search tools) ? optionally load a reference favorite's reusable style/layout ? render with the chosen model (thinking off) ? fault-tolerant JSON parse ? sanitize (BeautifulSoup, strips `