deerflow-code/offline-backend-20260512/backend/CLAUDE.md
2026-09-07 18:24:55 +08:00

1847 lines
269 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CLAUDE.md
??
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Project Overview
DeerFlow is a LangGraph-based AI super agent system with a full-stack architecture. The backend provides a "super agent" with sandbox execution, persistent memory, subagent delegation, and extensible tool integration - all operating in per-thread isolated environments.
**Architecture**:
- **Gateway API** (port 8001): REST API plus embedded LangGraph-compatible agent runtime
- **Frontend** (port 3000): Next.js web interface
- **Nginx** (port 2026): Unified reverse proxy entry point
- **Provisioner** (port 8002, optional in Docker dev): Started only when sandbox is configured for provisioner/Kubernetes mode
**Runtime**:
- `make dev`, Docker dev, and production all run the agent runtime in Gateway via `RunManager` + `run_agent()` + `StreamBridge` (`packages/harness/deerflow/runtime/`). Nginx exposes that runtime at `/api/langgraph/*` and rewrites it to Gateway's native `/api/*` routers.
**Project Structure**:
```
deer-flow/
??? Makefile # Root commands (check, install, dev, stop)
??? config.yaml # Main application configuration
??? extensions_config.json # MCP servers and skills configuration
??? backend/ # Backend application (this directory)
? ??? Makefile # Backend-only commands (dev, gateway, lint)
? ??? langgraph.json # LangGraph Studio graph configuration
? ??? packages/
? ? ??? harness/ # deerflow-harness package (import: deerflow.*)
? ? ??? pyproject.toml
? ? ??? deerflow/
? ? ??? agents/ # LangGraph agent system
? ? ? ??? lead_agent/ # Main agent (factory + system prompt)
? ? ? ??? middlewares/ # 10 middleware components
? ? ? ??? memory/ # Memory V2: provider/manager/builtin/hindsight
? ? ? ??? thread_state.py # ThreadState schema
? ? ??? sandbox/ # Sandbox execution system
? ? ? ??? local/ # Local filesystem provider
? ? ? ??? sandbox.py # Abstract Sandbox interface
? ? ? ??? tools.py # bash, ls, read/write/str_replace
? ? ? ??? middleware.py # Sandbox lifecycle management
? ? ??? subagents/ # Subagent delegation system
? ? ? ??? builtins/ # general-purpose, bash agents
? ? ? ??? executor.py # Background execution engine
? ? ? ??? registry.py # Agent registry
? ? ??? tools/builtins/ # Built-in tools (present_files, ask_clarification, view_image)
? ? ??? mcp/ # MCP integration (tools, cache, client)
? ? ??? models/ # Model factory with thinking/vision support
? ? ??? skills/ # Skills discovery, loading, parsing
? ? ??? config/ # Configuration system (app, model, sandbox, tool, etc.)
? ? ??? community/ # Community tools (tavily, jina_ai, firecrawl, image_search, aio_sandbox)
? ? ??? reflection/ # Dynamic module loading (resolve_variable, resolve_class)
? ? ??? utils/ # Utilities (network, readability)
? ? ??? client.py # Embedded Python client (DeerFlowClient)
? ??? app/ # Application layer (import: app.*)
? ? ??? gateway/ # FastAPI Gateway API
? ? ? ??? app.py # FastAPI application
? ? ? ??? routers/ # FastAPI route modules (models, mcp, memory, skills, uploads, threads, artifacts, agents, suggestions, channels)
? ? ??? channels/ # IM platform integrations
? ??? tests/ # Test suite
? ??? docs/ # Documentation
??? frontend/ # Next.js frontend application
??? skills/ # Agent skills directory
??? public/ # Public skills (committed)
??? custom/ # Custom skills (gitignored)
```
## Important Development Guidelines
### Documentation Update Policy
**CRITICAL: Always update README.md and CLAUDE.md after every code change**
When making code changes, you MUST update the relevant documentation:
- Update `README.md` for user-facing changes (features, setup, usage instructions)
- Update `CLAUDE.md` for development changes (architecture, commands, workflows, internal systems)
- Keep documentation synchronized with the codebase at all times
- Ensure accuracy and timeliness of all documentation
### Workflow Studio stream and canvas compatibility
- `app/gateway/workflow_agent_runner.py` is the boundary between LangGraph token chunks and durable workflow events. Provider reasoning in `AIMessage.additional_kwargs.reasoning_content`/`reasoning` is emitted to the UI as balanced inline `<think>` blocks; only visible answer content becomes the node's `text` output for downstream nodes. It must also emit every original `[AIMessage|ToolMessage, metadata]` tuple as `node.message.data.messages` with no whitelist, preview or length crop. Tool results and `additional_kwargs.clarification` are needed by the Studio's DeerFlow-compatible message UI and must survive durable SSE replay.
- An agent `ask_clarification` ToolMessage pauses the workflow at that agent node (`awaiting_input`), not as a terminal run. The pause records the source `nodeRunId`/`toolCallId`; resuming turns the submitted reply into the next user message on the same `wf-{run}-{node}` LangGraph thread. Keep the raw ToolMessage authoritative; do not invent a second card payload.
- `app/gateway/workflow_proposal_planner.py` converts the workflow controller's selected catalog roles into per-agent `mission`/`deliverable`/`scope`/`handoff` contracts. Only a selected visible `agentId` may contribute a contract; every omitted or malformed field falls back to a deterministic role-aware contract. Preserve this prompt-level worker boundary when adding strategies: researchers must not silently become final writers, and sequential roles must receive their upstream handoff binding plus an explicit next handoff.
- `app/gateway/routers/workflows_coze_compat.py` must return both `creator.self: true` and editable VCS metadata for an owned standalone workflow. Coze Playground derives its preview/read-only mode from those fields, so omitting `creator.self` disables node dragging even when the standalone page passes `readonly={false}`.
## Commands
**Root directory** (for full application):
```bash
make check # Check system requirements
make install # Install all dependencies (frontend + backend)
make dev # Start all services (Gateway + Frontend + Nginx), with config.yaml preflight
make start # Start production services locally
make stop # Stop all services
```
**Backend directory** (for backend development only):
```bash
make install # Install backend dependencies
make dev # Run Gateway API with reload (port 8001)
make gateway # Run Gateway API only (port 8001)
make test # Run all backend tests
make lint # Lint with ruff
make format # Format code with ruff
```
Regression tests related to Docker/provisioner behavior:
- `tests/test_docker_sandbox_mode_detection.py` (mode detection from `config.yaml`)
- `tests/test_provisioner_kubeconfig.py` (kubeconfig file/directory handling)
Boundary check (harness ? app import firewall):
- `tests/test_harness_boundary.py` ? ensures `packages/harness/deerflow/` never imports from `app.*`
CI runs these regression tests for every pull request via [.github/workflows/backend-unit-tests.yml](../.github/workflows/backend-unit-tests.yml).
## Architecture
### Position roundtable recovery persistence
`/api/position-roundtable` uses the shared `/api/intent/*` LangGraph flow as
the live Step-1 execution source and stores an additional SQL recovery copy in
`position_roundtable_sessions`/`position_roundtable_nodes`. The recovery copy
contains complete intent/node/summary/action-plan conversations, unsent
composer drafts, workflow snapshots, and bounded full text for generated
Markdown/JSON artifacts. This lets history and downstream closing agents keep
working when container-local checkpoint or sandbox volumes are replaced.
Revision `20260723_05` adds the conversation columns. Session deletion
explicitly removes child nodes so SQLite deployments do not depend on FK
cascade pragmas. The router is mounted canonically below
`/api/position-roundtable` and, for intranet gateway allowlists, below the
schema-hidden compatibility prefix `/api/multi-agent/position-roundtable`.
Clients may retry the compatibility route only for a framework/proxy-level
`404 Not Found`; domain-level missing-session responses must remain real 404s.
The position-roundtable state machine is multi-worker safe: node commands
(`complete`, `bind`, conversation save, reject) and session commands
(`activate`, archive, delete) lock the parent session and affected nodes in one
transaction, then record a unique command id for retry replay. The task and
intent snapshot writes use the same parent-row lock; they are permitted only
while `status=intent_pending`, so a late “confirm intent” write cannot change a
chain that another user has already activated. Activation itself rechecks the
stored confirmed intent after taking the lock. `tests/test_position_roundtable_
commands.py` and `tests/test_position_roundtable_concurrency.py` cover the
state transitions; `tests/test_roundtable_postgres_concurrency.py` is the real
PostgreSQL acceptance suite (set `DEERFLOW_POSTGRES_CONCURRENCY_TEST_URL`).
Roundtable backend jobs use a database unique active-dedupe key plus a leased
dispatcher, never process-local start deduplication. Cancellation is two
phase (`cancel_requested` then `cancelled`): the lease owner’s watcher normally
finalizes it, while the dispatcher reaps a cancellation whose lease has
expired after a worker crash. A worker that loses lease stops its local model
run immediately. Do not reintroduce direct `executor.cancel()` from the HTTP
cancel route, because it can cancel the watcher before the database state is
finalized. The boolean defaults on these new rows must use SQLAlchemy
`false()` rather than numeric `0`, so fresh PostgreSQL schemas can be created.
Position roundtable selects the separately seeded built-in agent
`position-roundtable-intent` by passing `intent_agent_id` to the shared
`/api/intent/init` and `/api/intent/stream` endpoints. Its request lifecycle,
files, and SSE protocol match `roundtable-intent`, but its cap is five
**user-assistance rounds**. A single round may emit several independent
`ask_clarification` tool calls; those calls are separate cards, not one card
containing several numbered questions. Count that as one round by the
tool-calling AI message, never by card count. Topic-only requests must request
all material missing decisions as separate cards rather than being treated as a
resolved intent; after the fifth set is answered it must emit completed final
JSON without unresolved placeholders. The required arrays are
`coreGoals`, `riskWarnings`, `keyPoints`, and `strategicSignificance`. Do not
treat a fresh dedicated-agent `ask_clarification` as stale: it must win over
a same-run premature `[INTENT_READY]` and leave the UI waiting for the user.
Conversely, the prior turn’s historical clarification must not block the user
reply from completing. Do not
change `roundtable-intent`'s `{objective,constraints,assumptions}` contract for
this feature. The position router accepts the new schema for activation and
injects all four sections into downstream position-node and closing-agent
context; legacy stored summaries are mapped read-only for compatibility.
`/api/intent/init` must create the empty Step-1 thread through
`multi_agent._create_thread_direct`, exactly like multi-agent initialization;
do not restore the old HTTP loopback to `/api/threads`, because deployed
Gateway ports/proxies may make the automatic page initialization fail while
later manual turns still work. If checkpoint creation fails or the parallel seat
task is cancelled, the helper compensates by deleting any partial checkpoint and
metadata row (the cancellation branch runs under `asyncio.shield` so a second
cancellation cannot abort it half-way). Streaming stays on the existing loopback
path.
`ThreadMetaStore.create` never issues a post-commit `session.refresh()`. Every
`ThreadMetaRow` column is assigned before commit and the session factory uses
`expire_on_commit=False`, so the returned dict is built from the in-memory row.
The removed read-back ran on a *different* pooled connection than the INSERT,
which a MySQL read/write-splitting endpoint can route to a lagging replica — the
row comes back empty and SQLAlchemy raises `Could not refresh instance`, surfacing
as `500 "Failed to create thread"`. Concurrency makes that far more likely, so a
single-user deployment never sees it. No caller uses the return value, and all
seven call sites (`multi_agent._create_thread_direct`, `services.start_run`,
`routers/threads.create_thread`, `roundtable_inprocess_gateway`,
`scheduler/service`, `thread_shares`, `authz`) benefit. Keep new `create`
implementations read-back-free; `tests/test_position_roundtable_sessions.py::
test_thread_meta_create_never_reads_back_after_commit` guards this.
### Harness / App Split
The backend is split into two layers with a strict dependency direction:
- **Harness** (`packages/harness/deerflow/`): Publishable agent framework package (`deerflow-harness`). Import prefix: `deerflow.*`. Contains agent orchestration, tools, sandbox, models, MCP, skills, config ? everything needed to build and run agents.
- **App** (`app/`): Unpublished application code. Import prefix: `app.*`. Contains the FastAPI Gateway API and IM channel integrations (Feishu, Slack, Telegram, DingTalk).
**Dependency rule**: App imports deerflow, but deerflow never imports app. This boundary is enforced by `tests/test_harness_boundary.py` which runs in CI.
### Report Collaboration (AgentScope, in progress)
Independent of Workflow Studio and the existing roundtable. Plan: repo-root `docs/agentscope-多智能体报告协作工作台-后端实施计划.md`(`RC-BE-000`~`018` 已落地,下一票 `RC-BE-019`)。Spike baseline: `docs/AGENTSCOPE_BASELINE.md`.
- **Config**: `config.yaml → report_collaboration` (`deerflow.config.report_collaboration_config.ReportCollaborationConfig`). Default **off** (`enabled` / `worker_enabled`).
- **Contracts**: `app/report_collaboration/contracts/` — snake_case Pydantic models matching `frontend-web/src/report-collaboration/api/types.ts`. SSE envelope uses camelCase (`eventId` / `sessionId` / `runId`). `node_run_id` is always `${node_id}-attempt{attempt}`.
- **Persistence**: `deerflow.persistence.report_collaboration` (`report_collaboration_*` tables, migrations `20260905_01` / `20260906_01` / `20260907_01`). SQL + memory stores; one active run per session via `active_dedupe_key`; command/message/run/version idempotency keys refuse cross-session reuse; event `seq` taken from the run row. Session delete removes children explicitly.
- **Gateway**: `/api/report-collaboration` (`app/gateway/routers/report_collaboration.py`). `enabled=false` → 503. Plan/run writes enqueue durable jobs and return immediately. Pre-run `POST /sessions/{id}/messages` runs a rule-based `IntentRouterAgent` + `RequirementResolver` (`app/report_collaboration/requirement/`): writes a requirement revision and clarification messages, never auto-starts planning. `POST /sessions/{id}/plan-requests` then runs `PlanProposalService` (`app/report_collaboration/planning/`): catalog projection (no SOUL/keys) → selector → 2–3 strategy RoleDemand graphs → assembler binds catalog agents/tools/quality gates → validator (acyclic DAG, artifact types, allowlisted tools). Select and start-run stay independent commands. `POST /runs` freezes role models, report-structure template and run budget. Rewrite / apply / restore are live (`reporting/`). `GET /runs/{id}/stream` is a read-only durable SSE tail (`execution/live_hub.py`): persist then publish, replay + `after_seq` / `Last-Event-ID`, heartbeat comments, token-delta coalesce before seq assignment. Disconnect does **not** cancel the run. Lifespan mounts `ReportCollaborationDispatcher` only when `enabled` **and** `worker_enabled` are both true; default both false.
- **Runtime**: AgentScope adapters live in `app/report_collaboration/agentscope_runtime/` (not the publishable harness). `ConversationReplica` maps AgentScope 2.0.5 events 1:1 onto collaboration SSE types; TeamSay is `HINT_BLOCK`. `DeerFlowChatModelAdapter` wraps `create_chat_model` (no second Provider config); role models are frozen onto node `agent_snapshot` at `POST /runs` and missing/incapable models fail with 422 before the run is created. Fallback is recorded on the snapshot plus a durable `heartbeat` audit event. Nodes complete only after a validated artifact (`app/report_collaboration/execution/completion.py`). `TaskLedger` (`execution/ledger.py`) owns ready waves, attempts and terminal node status; AgentScope / coordinator suggestions never pick the next node. A writer-bearing run is `completed` only after `ReportFinalizer` publishes a formal version. `WaveExecutor` + `QualityRoleKernel` (`execution/quality_roles.py`) remain for fixtures: Verifier marks stale/contested claims, Analyst/Writer only cite supported IDs, Reviewer `revise` replans only the target write node plus review. Production worker injects `AgentScopeMemberKernel` (`agentscope_runtime/member_kernel.py`) with DeerFlow tools and per-run budget. `QualityGate` (`execution/quality_gate.py`) is the claim-level gate: conflicts/missing units, unverified facts, coverage gaps, and unmapped review targets fail before `validated`; `ArtifactBoard.supersede_previous` hides old attempts from downstream. `ReportCollaborationLiveHub` (`execution/live_hub.py`) fans out persisted envelopes to SSE viewers; `PublishingReportCollaborationStore` sanitizes secrets/privacy/hidden reasoning without truncating visible tool results, and writes redaction + event audits. Dispatcher (`execution/dispatcher.py`) atomically claims queued/recovering/expired-lease runs, heartbeats via `renew_lease`, two-phase cancels (`cancel_requested` → `finalize_cancel`), and recovers orphans (`execution/recovery.py`): keep completed work, mark leftover in-flight `interrupted`/`superseded`, open a new `planned` attempt. Late artifacts after cancel/reclaim are `superseded` and must not change terminal run status. Tests drive `dispatch_once()`; do not rely on the background loop. Per-run budgets (`security/budget.py`) fail-closed; per-user active-run caps are enforced at `create_run`.
- **Quality eval**: `app/report_collaboration/eval/` (`RC-BE-018`) — 20 annotated tasks (policy/market/enterprise/event/comparison ×4) with required angles, key facts, forbidden fabrications, templates and review rules. Deterministic scorer (no LLM judge) gates citation/angle/template/directed-mod; gold reports must pass, canned failures (including a roundtable-style essay) must fail and stay reproducible. Recommended runtime knobs live in `eval/thresholds.py` and match current `ReportCollaborationConfig` defaults (`worker_enabled` remains false). Tests: `tests/test_report_collaboration_eval.py`.
- **Runtime intervention**: `app/report_collaboration/interventions/` owns run-time natural-language commands. `RuntimeIntentRouterAgent` uses explicit rules first and a tool-free coordinator-model fallback only for ambiguity; its Pydantic decision cannot dispatch or mutate anything. `ImpactAnalyzer` validates server-side node/agent ownership, resolves angle-specific lineage, and enforces confidence/cost/parallel-branch confirmation (`intent_confirmation_threshold`, default `0.8`). `RuntimeCommandService` mirrors the human message, emits frozen command events, CAS-claims accepted commands, and applies them only through `TaskLedger` safe points. Rework persists command overlays, supersedes affected attempts/artifacts, creates monotonic attempts, and prevents late output from entering downstream. Waiting-node answers resume only those nodes. `rerun_node` without a legal target does not expand to the whole graph; requirement patches apply only after rework succeeds; stuck `executing` commands are reclaimed on takeover or lease timeout; `command.cancelled` is a durable SSE event. In-run `POST /sessions/{id}/messages` returns 409 (`RUN_ACTIVE`). Formal rewrite candidates and version apply/restore live in `reporting/` (`RC-BE-015`).
- **Dependency**: optional extra `report-collaboration` → `agentscope==2.0.5`. Do not add AgentScope to the harness `pyproject.toml`.
Tests: `tests/test_report_collaboration_*.py`.
### Workflow Studio
Coze Studio is canvas-only; ChatDev is reference-only. Plan + status per phase: `docs/WORKFLOW_STUDIO_BACKEND_DEV_ZH.md`; decisions: `docs/adr/0001-workflow-studio-runtime.md`; contracts: `docs/WORKFLOW_SCHEMA_V1_ZH.md` + `docs/WORKFLOW_SSE_V1_ZH.md` + `docs/WORKFLOW_OPENAPI_DRAFT_V1.yaml`. Master switch + all limits: `config.yaml` → `workflows` (`deerflow.config.workflow_config.WorkflowConfig`). `workflows.agent_recursion_limit` defaults to 250: LangGraph counts each model→tool research exchange as about two super-steps, so workflow agent/skill nodes must not fall back to a chat-sized 60-step limit. In the local Studio proxy, keep Rsbuild `server.compress: false`; dev gzip buffers small SSE frames and makes live planning/run output arrive only at completion despite the backend's `no-transform` / `X-Accel-Buffering: no` headers. The composer’s optional `modelName` must be validated against `AppConfig.models`: it selects the controller model and is copied only into newly generated candidate agent configs (or the projected direct-agent task), never injected into an already-confirmed proposal/run.
**Execution kernel is a custom DAG scheduler, not a compiled LangGraph.** LangGraph stays the kernel *inside* an `agent`/`skill` node. The reason is single-source-of-truth recovery: a re-leased or resumed run replays completed node outputs from `workflow_node_runs` plus the `workflow_run_events` log, so a second (checkpoint) copy of workflow-level state would have no defensible tie-breaker. See the ADR's phase-2 amendment.
**Harness** (`packages/harness/deerflow/workflows/`, never imports `app.*`):
- `schemas` / `events` / `errors` / `ports` / `validator` / `samples` — frozen v1 contracts.
- `expressions.py` — `{{ inputs.x }}` / `{{ nodes.n.data.y }}` templates and conditions resolved by path walking + a small comparator set. **No `eval`.** Roots are restricted to `inputs` / `nodes` / `run` / `loop` / `env`.
- `runtime/engine.py` — the scheduler. Every edge is `pending` → `active` (source chose this port) or `pruned`; a node is ready when its non-back incoming edges are all resolved and at least one is `active`, and is skipped when all are `pruned` (skips propagate). That one rule covers fan-out, condition branching, and merge joins with no topological pre-pass. The `loop` node's `continue` back edge is the only permitted cycle; re-entry clears the body's results and invalidates its node runs.
- `runtime/{sink,sql_runner,code_runner,context}.py`, `nodes/*` (15 executors, including deterministic `evidence_normalizer` and durable `deep_research_write`), `security/{http_policy,sql_policy,secret_box}.py`.
**Persistence**: `workflow_definitions` / `workflow_versions` (migration `20260828_01`), `workflow_runs` / `workflow_node_runs` / `workflow_run_events` / `workflow_run_artifacts` (`20260828_02`), `workflow_data_sources` (`20260828_03`), and short-lived-but-durable `workflow_planning_sessions` / `workflow_planning_proposals` (`20260831_01`). Stores under `deerflow.persistence.workflow_{runs,events,data_sources,planning}` with SQL + in-memory implementations. Because planning proposal rows have a scalar FK rather than an ORM relationship, `SqlWorkflowPlanningStore.create_session` must flush the parent session before adding its children; otherwise SQLite can insert the child first and return a 500. A planning proposal is distinct from a run: its graph can be edited with optimistic `graph_revision` checks, but it becomes executable only once selected and atomically bound to one run.
**App layer**: `workflow_dispatcher.py` (lease scan → atomic claim; an expired lease is how a dead worker's run is reclaimed, so every worker can run the loop), `workflow_executor.py` (one lease's lifecycle: replay → engine under a run timeout with heartbeat + cancel watcher → exactly one terminal transition; also runs subworkflows inline, depth ≤ 3), `workflow_retention.py` (hourly sweep of `workflows.retention.event_retention_days` / `run_retention_days`), `workflow_live_hub.py` (in-process SSE fan-out), `workflow_agent_runner.py` (agent/skill nodes → lead agent; `disable_tools=True` means an empty bound-tool set for control-plane callers), `workflow_deep_research_adapter.py` (Evidence Pack → selected Deep Research sources/job/event projection; no internal HTTP call), and `workflow_proposal_planner.py` (calls the dedicated `workflow-planner` controller over visible-agent metadata, then constructs safe candidate graphs server-side). `workflow-planner` is a separate Workflow Studio control-plane agent; never reuse or alias `roundtable-coordinator` for workflow planning, and never let it dispatch a run, call research tools, or return arbitrary graph JSON. It always runs with tools and model thinking disabled.
**Gateway routes** (register static prefixes before `workflows.router`'s `/{workflow_id}` catch-alls): `/api/workflows` CRUD/draft/validate/publish/copy, `/api/workflows/node-types` + `/resources/*`, `/api/workflows/{id}/planning-sessions{/stream}` + `/api/workflows/planning-sessions/{session_id}{,/proposals/{proposal_id},/proposals/{proposal_id}/graph}`, `/api/workflows/{id}/runs` + `/api/workflows/runs/{run_id}{,/events,/stream,/cancel,/resume,/retry,/feedback,/artifacts}`, `/api/workflows/data-sources` (+ `/{id}/introspect`, `/{id}/validate-query`), `/api/workflows/embed/{tickets,verify}`, `/api/workflow_api/*` Coze compat. The planning `/stream` endpoint sends product-safe real controller milestones (`catalog`, `controller`, `selection`, `assembly`, `validation`, `complete`) then a final `result` session; never stream agent catalog descriptions or raw planner JSON. Formal runs accept the selected `planningSessionId` and `proposalId`; creating the same pair again returns the already-bound run rather than launching another execution.
**Guarantees worth not breaking**: events are persisted *then* published, and the DB is authoritative for `seq` ordering (live frames at or below the replay cursor are dropped and re-read on the next poll, so a reconnect sees no gap; `Last-Event-ID` wins over `?after=`); a run reaching a terminal status always gets a terminal event, synthesised if necessary, or SSE clients hang; `resumeToken` appears only in the `run.awaiting_input` event, never in the run view; `POST /retry` on a `failed` run copies completed node outputs onto a new run marked `retry_of_run_id` so the engine resumes from the failed node; `POST /feedback` never mutates its source run—on a terminal acyclic run it may target only a completed `agent`/`skill`, reuses completed nodes outside that node's downstream closure, records those `reusedNodeIds` alongside the feedback summary, and puts the bounded feedback only in `RunContext.env.nodeFeedback[target]` before creating a revision run; data-source secrets are write-only (encrypted via `WORKFLOW_SECRET_KEY`, decrypted only inside the executor, reads expose `masked_target`); the `http` node denies private/loopback/link-local targets by default (SSRF), `sql_read` allows one read-only statement with bound parameters, and the `code` node is **off by default** because its runner is a hardened subprocess rather than a container.
HTTP interface resources persist a non-secret base_url plus an allowed_methods allowlist (migration
20260829_01); request headers stay only in the encrypted secret blob. The resource catalog may return
the address and methods to the owning user, but must never return headers or any decrypted secret. The
executor receives HTTP metadata only through its data-source resolver.
The Workflow Studio agent/skill resource catalog is a read-only display projection over the existing
agent and skill ownership stores. It returns stable `agentId` / `skillId`, an agent's cached skill names,
and a `createdBy` display label. Resolve human labels from the immutable owner id at request time and
fall back to that id if lookup fails; built-in records must return `系统内置`. Do not create a parallel
creator table or accept creator data from the browser.
`evidence_normalizer` is a deterministic data node, not an LLM: it consumes only configured/upstream
node results and emits a bounded `evidencePack` of claims with node provenance, role labels and gaps.
The planner inserts it between parallel research roles and the synthesis Agent. It must not treat
source text as verified fact or fetch any external resource, and its explicit `sources[].nodeId` values
must stay upstream-only under `validate_workflow_graph`.
`deep_research_write` is a durable report node, not an Agent prompt shortcut. It accepts only an
Evidence Pack, derives a session identity from the workflow run/node plus the topic, evidence snapshot and writing config, writes
the pack as selected `other` source rows with upstream-node provenance, and queues/reuses the existing
Deep Research `regenerate` job. Its live `report_delta` and persisted `report_chunk` frames are
deduplicated before becoming workflow `node.output.delta` events; only report prose is forwarded (never
planning/reasoning deltas). Workflow cancellation requests cancellation on that same job. Do not make it
self-call `/api/deep-research/*`, use Deep Research `multi_agent` inside this node, or treat Evidence
Pack claims as verified facts. The parallel-research candidate grants the node an explicit 1800-second
timeout; deployments that override `workflows.node_timeout_seconds` must keep it at least that high.
Tests: `tests/test_workflow_schemas.py`, `tests/test_workflow_definitions.py`, `tests/test_workflow_engine.py`, `tests/test_workflow_runs.py`.
**Import conventions**:
```python
# Harness internal
from deerflow.agents import make_lead_agent
from deerflow.models import create_chat_model
# App internal
from app.gateway.app import app
from app.channels.service import start_channel_service
# App ? Harness (allowed)
from deerflow.config import get_app_config
# Harness ? App (FORBIDDEN ? enforced by test_harness_boundary.py)
# from app.gateway.routers.uploads import ... # ? will fail CI
```
### Agent System
**Lead Agent** (`packages/harness/deerflow/agents/lead_agent/agent.py`):
- Entry point: `make_lead_agent(config: RunnableConfig)` registered in `langgraph.json`
- Dynamic model selection via `create_chat_model()` with thinking/vision support
- Tools loaded via `get_available_tools()` - combines sandbox, built-in, MCP, community, and subagent tools
- System prompt generated by `apply_prompt_template()` with skills, memory, and subagent instructions
**agentfx built-in analyst + report convert skills.** `agentfx-analyst` (「任务研判报告助手」) is seeded from `app/gateway/routers/_agentfx_seed_assets/` before `_sync_legacy_agents`. After mandatory `web_search`, the agent writes each tool JSON to `/mnt/user-data/workspace/hits/qN.json` and runs **one** command — `task-report-build` (group `all`) — to fill the 14-category `task-reports.json` for the dashboard slots; the four per-page skills `task-report-{enemy,our,env,judge}` remain for re-running a single dashboard page. All five are thin CLIs over the shared stdlib engine `skills/public/task-report-lib/convert_lib.py` (no SKILL.md on the lib, so it never enters the skill list). The model must not `write_file` that JSON. Because the sandbox maps `/mnt/user-data/...` onto a host dir, a virtual absolute hits path can be `isdir`-true yet list empty on a Windows host — `resolve_hit_files` therefore tries the path as given, then the cwd-mapped sandbox path (`discover_user_data_root` + `resolve_sandbox_path` — bash `cd`s into the thread `workspace`, so `/mnt/user-data/outputs/task-reports.json` must land in that thread's `outputs/`, never `F:\mnt\...` on a Windows host), then the virtual-prefix-stripped relative form, then the bare dir name, and picks the first candidate that actually contains `.json`; a miss returns a `warnings` entry listing every tried form so the agent never has to write a debug script. Assemble of the 14 categories is the same function (`assemble_records`); the model must not copy or merge files. If the model is interrupted after convert, `task-reports.json` is already in outputs and the frontend import card also scans `GET /api/threads/{id}/artifacts` so入库 still works without `present_files`. Import remains `task-report-import`. Existing agent/skill directories are never overwritten. Tests: `tests/test_task_report_convert.py`, `tests/test_agentfx_seed.py`.
**Forced-research built-in agent.** `forced-research-responder` (「强制检索输出助手」) is seeded from `app/gateway/routers/_forced_research_seed_assets/` before `_sync_legacy_agents`. `ForcedResearchMiddleware` is mounted only for this stable id. On every visible user turn it requires a successful information-gathering tool result before the model may finish; `tool_choice=required` is backed by an `after_model -> model` retry for gateways that ignore tool choice, with an eight-attempt failure cap that permits only an honest failure response. Skill discovery, reading `SKILL.md`, and writing a helper script do not satisfy the gate; web/knowledge/MCP results, non-skill uploaded-file reads, or successful Python skill-script execution do. Tool filtering removes `present_files`, `skill_manage`, task/schedule/memory/clarification surfaces. The middleware permits `write_file`/`str_replace` only for `/mnt/user-data/workspace/*.py` and permits `bash` only for `.py` execution without redirection/document paths, so skills can run scripts but the agent cannot create Markdown/document artifacts or write `outputs`. Tests: `tests/test_forced_research_agent.py`.
**ThreadState** (`packages/harness/deerflow/agents/thread_state.py`):
- Extends `AgentState` with: `sandbox`, `thread_data`, `title`, `artifacts`, `todos`, `uploaded_files`, `viewed_images`
- Uses custom reducers: `merge_artifacts` (deduplicate), `merge_viewed_images` (merge/clear)
**Runtime Configuration** (via `config.configurable`):
- `thinking_enabled` - Enable model's extended thinking
- `model_name` - Select specific LLM model
- `is_plan_mode` - Enable TodoList middleware
- `subagent_enabled` - Enable task delegation tool
### Middleware Chain
Lead-agent middlewares are assembled in strict append order across `packages/harness/deerflow/agents/middlewares/tool_error_handling_middleware.py` (`build_lead_runtime_middlewares`) and `packages/harness/deerflow/agents/lead_agent/agent.py` (`_build_middlewares`):
1. **ThreadDataMiddleware** - Creates per-thread directories under the creator's stable isolation scope (`backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/{workspace,uploads,outputs}`); resolves persisted threads through `resolve_path_user_id(thread_id)` (falls back to `get_effective_user_id()` / `"default"` when metadata is unavailable); Web UI thread deletion now follows LangGraph thread removal with Gateway cleanup of the local thread directory
2. **UploadsMiddleware** - Tracks and injects newly uploaded files into conversation
3. **SandboxMiddleware** - Acquires sandbox, stores `sandbox_id` in state
4. **DanglingToolCallMiddleware** - Injects placeholder ToolMessages for AIMessage tool_calls that lack responses (e.g., due to user interruption), including raw provider tool-call payloads preserved only in `additional_kwargs["tool_calls"]`
5. **LLMErrorHandlingMiddleware** - Normalizes provider/model invocation failures into recoverable assistant-facing errors before later middleware/tool stages run
6. **GuardrailMiddleware** - Pre-tool-call authorization via pluggable `GuardrailProvider` protocol (optional, if `guardrails.enabled` in config). Evaluates each tool call and returns error ToolMessage on deny. Three provider options: built-in `AllowlistProvider` (zero deps), OAP policy providers (e.g. `aport-agent-guardrails`), or custom providers. See [docs/GUARDRAILS.md](docs/GUARDRAILS.md) for setup, usage, and how to implement a provider.
7. **SandboxAuditMiddleware** - Audits sandboxed shell/file operations for security logging before tool execution continues
8. **ToolErrorHandlingMiddleware** - Converts tool exceptions into error `ToolMessage`s so the run can continue instead of aborting
9. **SummarizationMiddleware** (`DeerFlowSummarizationMiddleware`) - Context reduction when approaching token limits (optional, if enabled). **Non-destructive**: overrides the stock `before_model` (which persists a `RemoveMessage(REMOVE_ALL_MESSAGES)`) to a no-op and instead compresses **transiently** in `wrap_model_call` ? the model request is replaced with `[summary, *recent]` while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). The `summary` message is `hide_from_ui` and never persisted; summaries are cached **per-thread in a process-wide, thread-safe LRU** (`_GLOBAL_SUMMARY_CACHE`, bounded to 512 threads, sticky boundary). The cache is module-level — **not** on the middleware instance — because `make_lead_agent` rebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed by `thread_id` (resolved like `ThreadDataMiddleware`), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. **Threshold-first gating**: `_plan_compression` returns passthrough (no compression at all) whenever the *full* current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. **Stickiness headroom** (`_apply_stickiness_headroom`, `_KEEP_HEADROOM_FRACTION=0.75`): after the `keep` policy picks a cutoff, the kept suffix is forced below ~75% of *every* trigger before summarizing. Without it a `keep` in *messages* (e.g. 20) can itself exceed a *token* trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a `{"type":"context_compacting","message":"正在压缩上下文…"}` event on LangGraph's `custom` stream channel (the cheap reuse path stays silent); the chat frontend's `onCustomEvent` renders it as a transient toast. Tests: `tests/test_summarization_middleware.py`
10. **TodoListMiddleware** - Task tracking with `write_todos` tool (optional, if plan_mode)
11. **TokenUsageMiddleware** - Records token usage metrics when token tracking is enabled (optional)
12. **TitleMiddleware** - Two-phase auto-title (never shows "untitled"): `before_model` sets an instant question-derived *provisional* title the moment the first user message arrives (no LLM, zero latency), then `after_model` asynchronously generates the final LLM title and replaces it (tracked via the `title_provisional` state flag). Normalizes structured message content before prompting the title model
13. **MemoryMiddleware** - Memory V2: `wrap_model_call` injects Hindsight recall transiently; `after_agent` auto-retains the turn to Hindsight
14. **ViewImageMiddleware** - Injects base64 image data before LLM call (conditional on vision support)
15. **DeferredToolFilterMiddleware** - Hides deferred tool schemas from the bound model until tool search is enabled (optional)
16. **SubagentLimitMiddleware** - Truncates excess `task` tool calls from model response to enforce `MAX_CONCURRENT_SUBAGENTS` limit (optional, if `subagent_enabled`)
17. **LoopDetectionMiddleware** - Detects repeated tool-call loops; hard-stop responses clear both structured `tool_calls` and raw provider tool-call metadata before forcing a final text answer
18. **ClarificationMiddleware** - Intercepts `ask_clarification` tool calls, interrupts via `Command(goto=END)` (must be last)
8a. **ToolMetricsMiddleware** - Appended immediately after `ToolErrorHandlingMiddleware` in `_build_runtime_middlewares`, so it sits *deepest* in the `wrap_tool_call` chain. Records one `tool_call_metrics` row per tool/skill/MCP invocation ? timing duration and bucketing failures into a coarse `error_category` (rate_limited / ip_blocked / auth_error / timeout / network_error / not_found / empty_result / tool_error). It catches both raised exceptions (before #8 converts them) and *self-handled* errors returned as error `ToolMessage`s (search tools that swallow HTTP 429/403). It also captures the targeted Agent Skill into `tool_call_metrics.skill_name`: the lead agent invokes a skill by `read_file`-ing its `/mnt/skills/<category>/<name>/SKILL.md` path (per the `<skill_system>` prompt block ? there is no dedicated "use skill" tool), so `skill_name` is recovered by scanning every tool call's string arguments for the `skills/{public,custom}/{name}/` path shape (`_SKILL_PATH_RE`), plus the explicit `name` argument of `skill_view`/`skill_manage`/`skill_archive`. `skill_view`-specific success detection handles its swallowed failures (`Skill 'x' not found.` returned as a plain string rather than raised). Regression coverage: `tests/test_tool_metrics_middleware.py`. Writes are best-effort fire-and-forget. **Per-run skill merging**: `SqlToolMetricsStore.record` upserts skill rows by `(run_id, skill_name)` instead of inserting one row per call ? a skill touched by several tool calls / retries inside one run collapses to a single row. Any success in the run wins (a skill that failed then finally succeeded is recorded as one `success`); a skill that only ever failed yields one `error` row whose `error_message` accumulates every attempt's message; durations are summed. Non-skill tool calls, and skill calls with no `run_id`, keep one row each. The read?merge?write is guarded by a per-event-loop `asyncio.Lock`. Surfaced by the admin `GET /api/admin/tool-metrics` router (which has a `kind=skill` filter and a per-skill `/summary` breakdown).
9. **PromptPrefixMiddleware** - Injects the admin-configured prompt prefix into the model request only. The prefix is tagged onto the first human message's `additional_kwargs['prompt_prefix']` by the Gateway (`app.gateway.services`) and prepended to content here at call time, so the persisted/displayed message stays prefix-free
10. **SummarizationMiddleware** (`DeerFlowSummarizationMiddleware`) - Context reduction when approaching token limits (optional, if enabled). **Non-destructive**: overrides the stock `before_model` (which persists a `RemoveMessage(REMOVE_ALL_MESSAGES)`) to a no-op and instead compresses **transiently** in `wrap_model_call` ? the model request is replaced with `[summary, *recent]` while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). The `summary` message is `hide_from_ui` and never persisted; summaries are cached **per-thread in a process-wide, thread-safe LRU** (`_GLOBAL_SUMMARY_CACHE`, bounded to 512 threads, sticky boundary). The cache is module-level — **not** on the middleware instance — because `make_lead_agent` rebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed by `thread_id` (resolved like `ThreadDataMiddleware`), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. **Threshold-first gating**: `_plan_compression` returns passthrough (no compression at all) whenever the *full* current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. **Stickiness headroom** (`_apply_stickiness_headroom`, `_KEEP_HEADROOM_FRACTION=0.75`): after the `keep` policy picks a cutoff, the kept suffix is forced below ~75% of *every* trigger before summarizing. Without it a `keep` in *messages* (e.g. 20) can itself exceed a *token* trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a `{"type":"context_compacting","message":"正在压缩上下文…"}` event on LangGraph's `custom` stream channel (the cheap reuse path stays silent); the chat frontend's `onCustomEvent` renders it as a transient toast. Tests: `tests/test_summarization_middleware.py`
11. **TodoListMiddleware** - Task tracking with `write_todos` tool (optional, if plan_mode)
12. **TokenUsageMiddleware** - Records token usage metrics when token tracking is enabled (optional)
13. **TitleMiddleware** - Two-phase auto-title (never shows "untitled"): `before_model` sets an instant question-derived *provisional* title the moment the first user message arrives (no LLM, zero latency), then `after_model` asynchronously generates the final LLM title and replaces it (tracked via the `title_provisional` state flag). Normalizes structured message content before prompting the title model
14. **MemoryMiddleware** - Queues conversations for async memory update (filters to user + final AI responses)
15. **ViewImageMiddleware** - Injects base64 image data before LLM call (conditional on vision support)
16. **DeferredToolFilterMiddleware** - Hides deferred tool schemas from the bound model until tool search is enabled (optional)
17. **SubagentLimitMiddleware** - Truncates excess `task` tool calls from model response to enforce `MAX_CONCURRENT_SUBAGENTS` limit (optional, if `subagent_enabled`)
18. **LoopDetectionMiddleware** - Detects repeated tool-call loops; hard-stop responses clear both structured `tool_calls` and raw provider tool-call metadata before forcing a final text answer
19. **ClarificationMiddleware** - Intercepts `ask_clarification` tool calls, interrupts via `Command(goto=END)` (must be last)
### Configuration System
**Main Configuration** (`config.yaml`):
Setup: Copy `config.example.yaml` to `config.yaml` in the **project root** directory.
**Config Versioning**: `config.example.yaml` has a `config_version` field. On startup, `AppConfig.from_file()` compares user version vs example version and emits a warning if outdated. Missing `config_version` = version 0. Run `make config-upgrade` to auto-merge missing fields. When changing the config schema, bump `config_version` in `config.example.yaml`.
**Config Caching**: `get_app_config()` caches the parsed config, but automatically reloads it when the resolved config path changes or the file's mtime increases. This keeps Gateway and LangGraph reads aligned with `config.yaml` edits without requiring a manual process restart.
Configuration priority:
1. Explicit `config_path` argument
2. `DEER_FLOW_CONFIG_PATH` environment variable
3. `config.yaml` in current directory (backend/)
4. `config.yaml` in parent directory (project root - **recommended location**)
Config values starting with `$` are resolved as environment variables (e.g., `$OPENAI_API_KEY`).
`ModelConfig` also declares `use_responses_api` and `output_version` so OpenAI `/v1/responses` can be enabled explicitly while still using `langchain_openai:ChatOpenAI`.
**Extensions Configuration** (`extensions_config.json`):
MCP servers and skills are configured together in `extensions_config.json` in project root:
Configuration priority:
1. Explicit `config_path` argument
2. `DEER_FLOW_EXTENSIONS_CONFIG_PATH` environment variable
3. `extensions_config.json` in current directory (backend/)
4. `extensions_config.json` in parent directory (project root - **recommended location**)
### Gateway API (`app/gateway/`)
FastAPI application on port 8001 with health checks at `GET|HEAD /health`, `/api/health`, and `/api/langgraph/health`. These exact paths are always public in `auth_middleware.PUBLIC_HEALTH_PATHS`, including when authentication is enabled; reverse-proxy prefixes and trailing slashes are normalized before the whitelist check. Set `GATEWAY_ENABLE_DOCS=false` to disable `/docs`, `/redoc`, and `/openapi.json` in production (default: enabled). Regression coverage: `tests/test_health_auth_bypass.py`.
**Routers**:
| Router | Endpoints |
|--------|-----------|
| **Models** (`/api/models`) | `GET /` - list models; `GET /{name}` - model details |
| **MCP** (`/api/mcp`) | `GET /config` - get config; `PUT /config` - update config (saves to extensions_config.json) |
| **Skills** (`/api/skills`) | `GET /` - list skills; `GET /{name}` - details; `PUT /{name}` - update enabled; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
| **Memory V2** (`/api/memory/v2`) | `GET /` - builtin data (USER.md/MEMORY.md); `POST/PUT/DELETE /entries` - add/replace/remove entry; `GET/PATCH/DELETE /config` - effective config + per-user overrides; `POST /search` - Hindsight recall/reflect proxy |
| **Agents** (`/api/agents`) | `GET /` - list built-in, owned, and published agents; `POST /` - create user-owned agent; `GET/PUT/DELETE /{id}` - read/update/delete by stable agent id; `GET /check` - validate non-empty display name (duplicates allowed) |
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags`); `GET /{name}` - details; `PUT /{name}` - update enabled; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `GET /{name}/content` - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + repeatable `tag_ids` multi-select filter, OR-matched; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `GET /{name}/content` - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
| **Memory** (`/api/memory`) | `GET /` - memory data; `POST /reload` - force reload; `GET /config` - config; `GET /status` - config + data |
| **Agents** (`/api/agents`) | `GET /` - list built-in, owned, and published agents (optional `search` + repeatable `tag_ids` multi-select filter, OR-matched; each item carries `tags` and `featured_order`); `POST /` - create user-owned agent; `GET/PUT/DELETE /{id}` - read/update/delete by stable agent id; `GET /check` - validate non-empty display name (duplicates allowed); `GET /selector` - homepage agent picker (single admin-curated list, visibility-filtered, sorted by `featured_order` asc); `PUT /{id}/featured` - **admin** sets/clears the homepage `featured_order` (must be `published=true` to feature; passing `order: null` clears unconditionally); `PUT /featured/order` - **admin** bulk-reorders the homepage list (body `{agent_ids: [...]}` ? indices become 0..N-1) |
| **Tags** (`/api/tags`) | `GET /` - list tags (optional `tag_type` filter; open to all); `POST /` - create tag (admin); `PUT /{id}` / `DELETE /{id}` - update/delete tag (admin); `GET /assignments/{target_type}/{target_id}` - tags on a resource; `PUT /assignments/{target_type}/{target_id}` - replace a resource's tags (admin). Tags are typed `agent`/`skill`/`scheduled_task` and may only tag resources of their own type. |
| **Uploads** (`/api/threads/{id}/uploads`) | `POST /` - upload files (auto-converts PDF/PPT/Excel/Word); `GET /list` - list; `DELETE /{filename}` - delete |
| **Threads** (`/api/threads/{id}`) | `DELETE /` - remove DeerFlow-managed local thread data after LangGraph thread deletion; unexpected failures are logged server-side and return a generic 500 detail |
| **Threads search** (`POST /threads/search`, LangGraph-compat) | Conversation listing. **System threads are excluded by default**: rows whose `metadata.system` is truthy (scheduler runs, roundtable workers) never crowd the user's list; a caller whose `metadata` filter explicitly references `thread_type`/`system` (e.g. studio notebooks) opts out of the exclusion. New `query` field = case-insensitive title (`display_name`) substring search, used by the chats page's server-side pagination. `POST /threads/count` takes the same request body (limit/offset ignored) and returns `{total}` under identical filtering/exclusion semantics ? powers the chats page's ??/??? display without fetching rows (`ThreadMetaStore.count`). `ThreadMetaStore.search` gained `exclude_system`/`query` params; JSON-blob filters batch-scan until the requested `limit`/`offset` window is full (so pagination over a heavily-filtered set never under-fills), with a `thread_id` tie-breaker on the `updated_at DESC` ordering for stable OFFSET paging. Tests: `tests/test_thread_meta_search.py` |
| **Artifacts** (`/api/threads/{id}/artifacts`, `/api/artifact-library`) | `GET /{path}` - serve artifacts; active content types (`text/html`, `application/xhtml+xml`, `image/svg+xml`) are always forced as download attachments to reduce XSS risk; `?download=true` still forces download for other file types. `GET /api/artifact-library` scans every current-user normal thread's authoritative `outputs/` directory (including files omitted from `present_files`), excludes uploads and image/audio/video media, joins the owning conversation title/agent route metadata, and exposes search/filter/pagination with a bounded 15-second per-user cache plus `refresh=true` bypass. Frontend route `/page/workspace/library` links each item back to the source chat with `?artifact=<path>` so the chat opens the existing artifact panel directly. |
| **Word export** (`POST /api/writing/export/docx`) | Shared Markdown → DOCX generator used by conversation, artifact, scheduled-task, Deep Research, and AI-writing download surfaces. `app/gateway/word_export.py` builds OOXML with no Python conversion dependency: A4 portrait, 3.7/3.5/2.8/2.6 cm margins, real multilevel heading numbering (`一、` / `(一)` / `1.`), two-character first-line indent on heading levels 1–3, formal body/table styles, page numbers beginning at 1 with odd pages left-aligned and even pages right-aligned, and embedded `app/gateway/assets/fonts/方正小标宋简体.ttf`. Before numbering it removes complete manual heading prefixes (including `2.1` / `2.1.1`) and drops Markdown horizontal rules from Word output. The ordinary-Q&A Lead Agent can emit the matching Markdown heading forms and avoid unordered bullet lists when the admin enables `prompt_prefix.ordinary_qa_markdown_format_enabled` in `system_settings.json`; the switch defaults off, while writing setup, writing mode, notebook, scheduled runs, and SOUL-driven custom agents keep their independent prompts. The font loader validates its internal family, SHA-256, and OpenType embedding permission (`OS/2.fsType`) before export. Tests: `tests/test_word_export.py`, `tests/test_es_query_routing_prompt.py`. |
| **Deep Research** (`/api/deep-research`) | Durable, user-owned research sessions/jobs with replayable SSE events and hidden-thread artifacts. `basic`/`quick` use the compact runner; `detailed` creates bounded subtopics, searches/writes sections concurrently, and assembles a report; `deep` recursively searches follow-up questions under a deterministic breadth/depth query budget. `multi_agent` is a LangGraph 1.x editor/researcher/writer/reviewer graph: its initial plan emits `awaiting_input`, the recorder stores the interrupt and releases the lease, and `POST /jobs/{id}/resume` rebuilds the graph from the persisted approved/revised plan. A bounded `custom_outline` is frozen into the session; headings/list entries deterministically supply the detailed and multi-agent plan, and the shared writer receives it only as a guarded content constraint. Completed reports support owner-scoped follow-up (`GET /sessions/{id}/messages`, `POST /sessions/{id}/chat`): context is only the current report, selected sources and recent messages; source-id citations are persisted, while `allow_new_research=true` is explicitly rejected. Owners may change that evidence set only after the job stops via `PATCH /sessions/{id}/sources/{source_id}`; this applies to future follow-ups rather than rewriting a completed report. Optional report illustrations are gated by `config.yaml -> deep_research.image_generation`: the Harness `OpenAICompatibleImageProvider` requires base64 output, the shared runner step writes only bounded `research-image-01..04.(png|jpg|webp)` artifacts and Markdown uses private `deep-research://` links. `GET /sessions/{id}/artifacts/{name}` checks session ownership before serving an image; an unavailable provider yields a warning but never fails the text report. Runners live in `packages/harness/deerflow/agents/deep_research/runners/`, may never import `app.*`, and use only the `AdapterBundle`. |
| **Gateway auth** | Owner-scoped Deep Research routes await the platform's asynchronous `get_current_user` resolver before they read or write persistence, ensuring each session and job operation receives a concrete user ID. Deep Research artifacts are always resolved from a whitelisted filename via `/mnt/user-data/outputs/<name>` before mapping to a host path, preserving the sandbox guard on native Windows and Linux hosts. |
| **Suggestions** (`/api/threads/{id}/suggestions`) | `POST /` - generate follow-up questions; rich list/block model content is normalized before JSON parsing |
| **Thread Runs** (`/api/threads/{id}/runs`) | `POST /` - create background run; `POST /stream` - create + SSE stream; `POST /wait` - create + block; `GET /` - list runs; `GET /{rid}` - run details; `POST /{rid}/cancel` - cancel; `GET /{rid}/join` - join SSE; `GET /{rid}/messages` - paginated messages `{data, has_more}`; `GET /{rid}/events` - full event stream; `GET /../messages` - thread messages with feedback; `GET /../token-usage` - aggregate tokens |
| **Feedback** (`/api/threads/{id}/runs/{rid}/feedback`) | `PUT /` - upsert feedback; `DELETE /` - delete user feedback; `POST /` - create feedback; `GET /` - list feedback; `GET /stats` - aggregate stats; `DELETE /{fid}` - delete specific |
| **Runs** (`/api/runs`) | `POST /stream` - stateless run + SSE; `POST /wait` - stateless run + block; `GET /{rid}/messages` - paginated messages by run_id `{data, has_more}` (cursor: `after_seq`/`before_seq`); `GET /{rid}/feedback` - list feedback by run_id |
| **Open Chat** (`/api/open/chat`) | **Unauthenticated** open Q&A entry for external systems (the `/api/open/` prefix is whitelisted in `auth_middleware._PUBLIC_PATH_PREFIXES` + CSRF-exempt in `csrf_middleware.should_check_csrf`). `POST /chat` `{message \| messages[], model_name?, agent_name?, agent_id?, thinking_enabled=false, strip_think=true, thread_id?, recursion_limit=100}` → resolves the agent by **Chinese display name** (`agent_name`, matched against each `.deer-flow/agents/*/config.yaml` `name` via `_resolve_agent_id`; explicit `agent_id` wins if given), builds the lead agent via `make_lead_agent` (so the chosen **agent's** SOUL/tools/skills all apply), `ainvoke`s it once (stateless — no checkpointer, no thread persistence; runs under the `"default"` user bucket), and returns **only the final** `{answer, model_name, agent_name, agent_id, thinking_enabled, thread_id}` (no intermediate tool-call/process). **`POST /chat/stream`** is the **SSE streaming variant** (same request body) for **long tasks** — the blocking `/chat` produces no downstream bytes for the whole run, so a front nginx `proxy_read_timeout` (default 60s) cuts it off with a **504**; the stream keeps the connection warm (a `: keepalive` SSE comment every `_HEARTBEAT_SECONDS=15`s via a background producer-task + queue, plus `event: delta` AI-text increments and `event: progress`), and ends with **`event: result`** carrying the same final `{answer, …}` JSON (or `event: error`). Clients only need to read `event: result`; intermediate events are ignorable. Both endpoints share `_build_run_context` (config/agent/thread). `GET /agents` lists `{name, id}` so callers know what `agent_name` to send. **Thinking off is reliable**: when `thinking_enabled=false` the route also sets `thinking_force_disabled=true` so the lead agent calls `create_chat_model(force_disable_thinking=True)`, which injects **three** OpenAI-compatible "thinking off" shapes into `extra_body` — top-level `enable_thinking=false` (**SiliconFlow / DashScope hybrid models** like `deepseek-ai/DeepSeek-V4-*` on `api.siliconflow.cn` — the knob the nested `chat_template_kwargs` shape does NOT reach), `chat_template_kwargs.enable_thinking=false` (vLLM/Qwen native), and `thinking.type="disabled"` (Zhipu GLM) — even for models that only declare `supports_thinking: true`. For the **web-UI path** (which sends `thinking_enabled=false` but NOT `force_disable`), SiliconFlow models need `when_thinking_disabled: {extra_body: {enable_thinking: false}}` in their `config.yaml` entry (added to `deepseek-ai/DeepSeek-V4-Flash`) for thinking-off to take effect. Sets `is_scheduled_run=true` internally to skip the auto-title LLM call. **Generated-file return (`include_files=true`, default on)**: an agent that produces its report as a file (writes to `/mnt/user-data/outputs/` + `present_files`) returns only a short summary in `answer`; to recover the real deliverable both endpoints now also return `files: [{name, virtual_path, encoding(text|base64), content}]` — `_collect_output_files` scans the run's `state['thread_data']['outputs_path']` (the open thread is stateless so artifacts can't be fetched by thread, but the endpoint runs server-side and the sandbox already wrote the files to disk), text decoded as utf-8 else base64, capped by `_MAX_OUTPUT_FILES=20` / `_MAX_OUTPUT_FILE_BYTES=5MB`. `/chat` puts `files` on the response; `/chat/stream` puts it in the `event: result` payload. A run with no answer **and** no files now 502s ("…文本答案或文件"). Router `app/gateway/routers/open_chat.py`; standalone caller `test_open_chat_api.py` (repo root, streaming by default); tests `tests/test_open_chat.py`. |
| **Open 3qfx Ask** (`/api/open/3qfx/ask`) | **Unauthenticated** open entry that, given just a **任务id**, directly kicks off the 3qfx(三情分析)「第一次问答」as a **real persisted conversation** (sibling of `/api/open/chat`; same `/api/open/` auth-whitelist + CSRF-exempt). `POST /ask` `{task_id, message?, model_name?, thinking_enabled=false}` → (1) resolves the `3Q` agent from `app.state.business_mapping_store.get_mapping("3Q")` (the **same** agent the frontend deep-link `goPath=3qfx` uses; 503 if unconfigured); (2) **fetches the task detail server-side** (consumer `cop-task-detail`) and builds the opening text **identical to the frontend** `buildOpeningText` 3qfx branch — `任务名称:…\n任务内容:…\n任务id(task_id)为{task_id}` (an explicit `message` overrides and **skips** the fetch); (3) creates a fresh thread tagged `metadata.{taskId, agent_id}` (the exact filter `useTaskThreads` uses, and `metadata.taskId` makes it task-deeplink-shared / creator-path-resolved) and **starts a background persisted run** via `start_run(..., enforce_access=False)` (writes the checkpointer, runs under the `"default"` bucket); (4) **returns immediately** `{thread_id, run_id, agent_id, status, opening_text}` — the caller just needs "started", then watches the run in the 3qfx 任务工作区 conversation list. `start_run` gained a keyword-only `enforce_access` (default `True`); the open entry passes `False` so the admin-configured agent isn't rejected by the per-(missing)-login visibility check. **Consumer fetch config** lives in `config.yaml → task_deeplink`: `cop_task_detail_url` (full URL, appends `?taskId=`), `cop_task_detail_authorization` (default `Bearer admin`), `cop_task_detail_test` + `mock.cop_task_detail` (offline fake when the upstream is unreachable — on in this deployment). Router `app/gateway/routers/open_3qfx.py`; tests `tests/test_open_3qfx.py`. |
| **Thread Shares** (`/api/thread-shares`) | Per-conversation sharing with two modes (`thread_shares.mode`). **`import`**: `POST /` `{thread_id, mode}` - create (or reuse) a revocable share code; `GET /{code}/preview` + `POST /{code}/import` - copy that conversation into the caller's account (fresh `thread_id` via `app/gateway/thread_copy.py` + new `threads_meta` row + `user-data` copy; dedup on `metadata.shared_origin`). **`view`**: a public, read-only static page ? see Public Shares below. Both modes: `GET /` - list my codes; `DELETE /{code}` - revoke. |
| **Public Shares** (`/api/public/shares`) | **Unauthenticated** (the `/api/public/` prefix is whitelisted in `auth_middleware._PUBLIC_PATH_PREFIXES`). `GET /{code}` - read-only conversation snapshot (title + serialized messages + artifacts) for a `view`-mode share; messages are read straight from the latest checkpoint. `GET /{code}/files/{path}` - serves a file referenced by that conversation (images, Markdown, text, PDF, ?) so the shared page's links are openable; path resolution is confined to the owner's thread directory. Active content (`text/html`, `application/xhtml+xml`, `image/svg+xml`) is forced to download (`_public_share_disposition`); everything else is served inline. Missing/revoked/non-`view` codes all return 404. |
| **Roundtable Drafts** (`/api/roundtable-drafts`) | **Strictly per-user** persistence for the roundtable-planning page (replaces the old browser `localStorage`). `GET /` - list drafts (lightweight meta — incl. `chain_title`, the business-chain name used this session, cheaply regex-extracted from the in-memory `step2` blob in `_draft_to_meta`/`_extract_chain_title`; null = no chain / 自由模式; surfaced on the 第一步「分析记录」history card + sidebar list instead of the old「第 X 步」); `POST /` - create; `GET/PUT/DELETE /{id}` - full read / partial update / delete (all owner-scoped). A draft = one full session (Step 1 intent + Step 2 roundtable), step snapshots stored as JSON in `PortableLongText`. Recommendation history bound to a draft: `GET /{id}/recommendations` - list newest-first; `POST /{id}/recommendations` - append one run (objective/status/rationale/picks/candidates). taskId 深链会商记录 are **not** here — they live in the separate `/api/roundtable-task-drafts` store (see below) so personal history never mixes with task records. Backed by `deerflow.persistence.roundtable_drafts` (tables `roundtable_drafts`, `roundtable_recommend_history`); store wired in `deps.py` as `app.state.roundtable_draft_store`. Tests: `tests/test_roundtable_task_drafts.py`. |
| **Roundtable Task Drafts** (`/api/roundtable-task-drafts`) | **Task-scoped, unpartitioned** store for taskId-deeplink roundtable 聊天记录 (无界嵌入抽屉). Keyed **only by `task_id`** (one task → many records), **never partitioned by user** — any authenticated caller may read/write/delete a task's records (`created_by` is audit-only). Physically separate from `roundtable_drafts` so the two never mix. `GET /by-task/{task_id}` - all drafts for a task, newest-first (history dropdown), enriched with the latest background-job status per draft via `RoundtableJobRepository.status_map_by_drafts` (each meta also carries `chain_title`, same `_extract_chain_title` regex as the personal store); `GET /by-task/{task_id}/latest` - newest draft (回显最新一条); `POST /` - create (`task_id` required); `GET/PUT/DELETE /{id}` - full read / partial update / delete (all user-agnostic). Backed by `deerflow.persistence.roundtable_task_drafts` (table `roundtable_task_drafts`, auto-created by `create_all`; migration `20260618_01`); store wired in `deps.py` as `app.state.roundtable_task_draft_store`. Existing `task_id`-bearing `roundtable_drafts` rows are migrated over by `scripts/migrate_roundtable_task_drafts.py`. Tests: `tests/test_roundtable_task_drafts.py`, `tests/test_roundtable_drafts_by_task.py`. |
| **Roundtable Diagnostics** (`/api/roundtable-diagnostics`) | 会商全流程诊断日志 (报错 + 关键事件). One row per event in `roundtable_diagnostics`: `{scope (background = 后台挂起 / foreground = 前台服务端观测 / client = 前端上报), stage (step1_intent/step2_init/step2_leader/step2_seat/step2_recommend/step2_orchestration/step3_summary/step3_report/step3_dashboard/step3_structure/upstream/job), level (warning/error — **`info` 级运行日志不再收集**:`record_diagnostic` 在写库前丢弃 `level=="info"`,只留报错;begin/consensus/gather/job_started 等纯进度事件即使仍调 `.diag(level="info")` 也成 no-op,故诊断页 = 报错日志页), event, message, detail (JSON), job_id/draft_id/task_id/user_id/agent_id/agent_name/cycle, created_at}`. **Reads are admin-only**: `GET /` - filter (scope/stage/level/event/job_id/draft_id/task_id/user_id/q/since/until; `scope` accepts a comma-separated set so 前端「前台调用」= `foreground,client`) + paginate (newest-first), each item enriched with `userName` (user_id → 邮箱 `@`-prefix via `get_local_provider().get_user`, best-effort batched cache); `GET /facets` - distinct values for filter dropdowns; `POST /admin/cleanup` - retention purge. **`POST /report` is NOT admin-gated** — any authenticated user's frontend posts the errors *the user actually hit* (HTTP 502 before any SSE frame / network drop / SSE error frame) so they're collected too; stored as `scope="client"`. Backed by `deerflow.persistence.roundtable_diagnostics` (table `roundtable_diagnostics`, auto-created by `create_all`; migration `20260619_01`); store wired in `deps.py` as `app.state.roundtable_diagnostic_store` + a process-level default store (`register_default_store`) so the harness-side engine/gateway can write across the app/harness boundary. **Writers** (all best-effort — diagnostics never break a run): background path uses `DiagnosticsRecorder` (bound to job ctx) threaded through `RoundtableJobExecutor` → `JobProgressWriter.diag(...)` (engine transitions/step3 failures/empty-delivery) + `InProcessRoundtableGateway` (per-turn timeout/report·summary·dashboard truncation+missing); foreground path uses `app/gateway/roundtable_diag.py::record_foreground_diag` in `multi_agent.py`/`intent.py`/`recommend.py` error frames. **Bug-fix shipped alongside**: `InProcessRoundtableGateway._run_agent_turn` guards each agent run with an **inactivity (idle) timeout** (`_await_run_with_idle_guard`, `ROUNDTABLE_SEAT_IDLE_TIMEOUT_SECONDS` default 1200s = 20min, tuned high for unstable intranet models) — **not** a total-duration cap, so a slow-but-alive run that keeps streaming tokens runs as long as it needs; only a run that emits **zero StreamBridge events for `idle_timeout` seconds straight** (a true freeze: 内网模型连接池挂死 / sqlite 锁在挂载卷上挂死) is cancelled + raises (顶层把 job 标 error,而非永久冻结) + records a `seat_turn_timeout` diagnostic. Activity is observed via `app.state.stream_bridge` (the worker publishes every chunk there). Optional absolute cap `ROUNDTABLE_SEAT_MAX_TOTAL_SECONDS` (default 0 = off). Frontend: admin page `pages/RoundtableDiagnosticsPage.tsx` (route `/page/strategy/admin/roundtable-diagnostics`), opened from the **会商页右上角设置按钮** (admin only — `handleSettingsClick` in `roundtable-planning/pages/RoundtablePlanningPage.tsx`); admin read client `strategy-components/api/roundtable-diagnostics.ts`. **Client error reporter** `roundtable-planning/api/diagnostics-report.ts` (`reportRoundtableError` fire-and-forget + `keepalive`, `setRoundtableErrorContext` for draft/task attribution set by the page, `stageForRoundtableAgent`) is called at every client-observable failure: `multi-agent.ts` (`streamMultiAgent` — covers Step2 leader/seat AND Step3 report/summary/dashboard/structure — reports HTTP error + SSE `{status:"error"}` frame + worker `event:error` frame [模型调用没成功] + fetch-reject [网络不可达] + mid-stream drop [连接重置/超时], deduped via a `reported` flag; `initMultiAgent` HTTP/error-frame/incomplete/mid-stream), `intent.ts` (Step1 HTTP+error-frame), `recommend.ts` (Step2 recommend HTTP+error-frame), `roundtable-jobs.ts` (job api/stream). The admin page's first filter is a **后台挂起 / 前台调用** toggle (frontend = `foreground,client`); the stage dropdown options adapt to it (background → step2+step3+job; frontend → step1+step2+step3) and each row shows the reporting user's name + 👤. Tests: `tests/test_roundtable_diagnostics.py`, `tests/test_roundtable_inprocess_timeout.py`. **Seat skill-recognition (弱模型补偿)**: a seat = one custom agent run whose configured skills (`agent config.yaml → skills`) are injected **only** via the system prompt's `<skill_system>` block (model must `read_file` the SKILL.md itself) — weak models (deepseek-chat 等) ignore it and never use their skills. `app/gateway/roundtable_seat_skills.py::append_seat_skill_directive(agent_id, task, *, chain_enabled=False)` appends an explicit Chinese directive (each configured-and-enabled skill's name + description + SKILL.md container path + "先 read_file 再按其流程取信息/执行") to the **seat's per-turn task (human turn)** — far stronger pull on weak models than the system prompt. Wired into **both** seat paths: foreground `multi_agent._special_run` (only `role in {seat, seat_parallel}`; Step3 singletons report/summary/dashboard/ingest excluded) + background `InProcessRoundtableGateway.run_seat`. No-op (returns task unchanged) when the seat has no configured skills. **2026-06-20: default-off + two opt-in switches (OR semantics)** — was global-on, now injects only when **either** the seat agent's own switch (`AgentConfig.seat_skill_directive`, edited in the agent edit dialog `agent-card.tsx` + `agent-detail-panel.tsx`) **or** the **per-seat** switch on that agent **within the selected business chain** is on. The per-seat chain switch lives **in the chain's `seats` JSON** (`ChainSeatSchema.seat_skill_directive`, no DB column — rides the existing JSON like the per-seat `model` override), edited via a small ✨ icon → dropdown (Switch + description) on each agent card in `BusinessChainEditorPage`. It is threaded **per-seat** to each seat run: foreground `multi_agent` per-call field `seat_skill_directive` (orchestration looks it up per agentId from a `{agentId: bool}` session map) → `_special_run` → `chain_enabled`; background as the job field `seat_skill_directives: dict[str,bool]` (StartJobRequest → StartParams → deps factory → `InProcessRoundtableGateway._seat_skill_directives`) → `run_seat` looks up `agent_id` → `chain_enabled`. `ROUNDTABLE_SEAT_SKILL_DIRECTIVE` is only a **master kill-switch** (default allow; set `0/false` to hard-disable regardless of the switches — never force-enables). Tests: `tests/test_roundtable_seat_skills.py`, `tests/test_agents_backend.py`. **2026-06-20 Step2 提速三件套**:(1) **每席位推理深度覆盖** —— 业务链条 `seats` JSON 里逐席位可配 `seat_mode`(flash/thinking/pro/ultra,覆盖链条级 `seatMode`;编辑器席位卡右侧 ⏱Gauge 图标下拉,null=跟随链条默认)。前台经 `useStep2Orchestration` 三处 dispatch 的 `seatModeToRequestFields(seatModes[id] ?? selectedMode)` 透传;后台经 job `seat_modes: dict[str,str]` → `StartParams` → `InProcessRoundtableGateway._seat_modes`,`run_seat` 用 `_SEAT_MODE_PARAMS`(与前端 `seatModeToRunParams` 逐字对齐)解析覆盖作业级 seat_*(未命中→跟随作业级→policy 默认)。无 DB 列(ride `seats` JSON,但 `ChainSeatSchema.seat_mode` 须显式声明否则 Pydantic 剥掉)。Tests: `tests/test_roundtable_seat_modes.py`。(2) **recommend 批内并行** —— 后端 `engine._run_recommend` 由 for-await 串行改 `asyncio.gather`(对齐前端 2026-06-13 已并行;总控一轮派多席位=判定相互独立→并发;批内用派活前 deliveries 快照、批后统一并入,语义不变)。(3) **研讨前并行『取数』两段式**(业务链条 `gather_first` 列,逐条 opt-in,migration `20260620_02`,editor「研讨前并行取数」开关)—— `gather_first=True` 时 `run_orchestration` 在分发前先 `engine._run_gather_phase`(全席位 `asyncio.gather` 跑 `_GATHER_TASK`「只取数不研讨」),产出 `gathered: dict` 经 `_seat_message(gathered=...)` 作为「已收集资料」块注入 recommend/chain/dag 各研讨轮席位任务(与「前序交付」块分开),把最慢的技能延迟前置并行、给(尤其 chain 串行的)研讨段提速;研讨阶段不限制再调技能。前台对称实现:`useStep2Orchestration.runGatherPhase` 在 `runActiveOrchestration` 分发前跑一次(`gatherPhaseDoneRef` 闸门,resetInitGate 复位),产出存 `gatheredDeliveriesRef` 经 `peerDeliveries`/初始 `priorDeliveries` 注入研讨席位。链条级 `gatherFirst`(camel)↔`gather_first`(snake) 经 `api/chains.ts` 映射;job 经 `StartJobRequest.gather_first`→`StartParams`→`run_orchestration(gather_first=…)`。Tests: `tests/test_roundtable_gather_phase.py`。(4) **削减重复广播**(前端 only,多智能体 foreground 路径)—— `suppress_pre_broadcast`(仅 leader:抑制「总控本轮输入→各席位 thread」预广播 + 「总控给X派活」→兄弟席位的 post 广播,二者一个 flag)此前只在 DAG 派活轮开;现扩到 recommend(`safetyCounter > 1`,即第 2 轮起的续轮/综合)+ chain 逐席位(`i > 0`,第 2 个席位起)+ chain 收口轮(恒 true)。**保留每种模式的首轮预广播**(它为席位打底任务背景),只 suppress 重复的续轮广播(N 次读写巨大席位 thread 的纯浪费 + 弱模型模仿派活);席位的前序交付另经「席位交付广播」`broadcastSeatDeliveryToOthers` 送达,不受影响。后台引擎路径无此广播(`_seat_message` 直接拼上下文),故 #4 仅前端。**2026-06-20 深链接业务链抽取规范**:深链接会商(taskId+goPath,businessCode 如 rwfx→6BF)下,结构化抽取(flow-json)与方案总结报告都按所选业务链**完整层级**强制校正。结构化抽取实为 `roundtable-structure` 智能体(`flow_extract.py` 的 `_system_prompt()=load_agent_soul(STRUCTURE_AGENT_ID)`,抽取→校验两段式 `model.astream`);其 SOUL 全局一份,旧实现把某业务层级写死进 SOUL→所有业务都按「目的→行为体」抽(bug)。修法:**单份领域知识、两角色用法**的 `deerflow.agents.roundtable_orchestrator.extraction_spec`(放 harness 层,app+harness 都可 import;`build_structure_directive`/`build_summary_directive` 按 businessCode 取;rwfx=6BF 写实「任务→目的→行为体→预期行为→驱动因素→叙事,预期行为↔驱动因素 1:1、余 1:N、label 不含步骤/阶段名、优先报告回退席位」,其余业务通用规范=强约束完整性+命名+取数且明令不套固定模板;business_code 空→""零副作用)。注入:(1) 结构化抽取 `flow_extract._hierarchy_hint`+`_validate_messages` 用 `build_structure_directive` 权威覆盖默认层级(前端 `streamFlowExtract` 传 `business_code`,源 `ingestBusinessCode`);(2) 方案总结报告 `build_summary_directive`——前台 `multi_agent._special_run`(role==summary 且 business_code) 追加、后台 `engine.build_summary_prompt(business_code=…)`(StartParams/StartJobRequest 透传)。**仅深链接生效**。种子 `roundtable-structure/SOUL.md` 层级顺序改 spec-driven、连接示例去 rwfx 写死(存量 live SOUL 不被 seed 覆盖,但注入规范权威可纠正)。Tests: `tests/test_roundtable_extraction_spec.py`。 |
| **Roundtable Draft Shares** (`/api/roundtable-draft-shares`) | Per-draft sharing for the roundtable-planning page, mirroring Thread Shares but the shared resource is a `roundtable_drafts` row. Two modes (`roundtable_draft_shares.mode`). **`import`**: `POST /` `{draft_id, mode}` - create (or reuse) a revocable code; `GET /{code}/preview` + `POST /{code}/import` - copy that draft (step1/2/3) into the caller's account under a deterministic per-(code,recipient) id (`imp_<sha1>` ? re-import is a no-op). **`view`**: a public, read-only result page ? see Public Roundtable Shares. Both: `GET /` - list my codes; `DELETE /{code}` - revoke. Backed by `deerflow.persistence.roundtable_draft_shares` (table `roundtable_draft_shares`, migration `20260606_01`); store wired in `deps.py` as `app.state.roundtable_draft_share_store`. Tests: `tests/test_roundtable_draft_shares.py`. |
| **Public Roundtable Shares** (`/api/public/roundtable-shares`) | **Unauthenticated** (`/api/public/` whitelisted). `GET /{code}` - read-only snapshot for a `view`-mode draft share: returns the Step 3??????report HTML (`step3.html`, a self-contained document) + title/summary/`has_report`. The frontend renders it in a sandboxed `iframe` (no `allow-scripts`) at `#/roundtable-share/<code>`. Missing/revoked/non-`view` codes all return 404. |
| **Scheduled Tasks** (`/api/scheduled-tasks`) | Per-user cron-like agent tasks. `GET/POST /` - list/create owned tasks; `PATCH/DELETE /{id}` - update/delete (owner); `POST /{id}/{pause,resume,trigger}`; `GET /{id}/runs`, `GET /runs/{rid}`, run-file preview/edit/download; `GET /runs/{rid}/messages` - agent messages for the run's underlying agent run (used to inspect a stuck/running run; authorized via `_load_run`); `GET /runs/{rid}/html` - generated HTML document for an `html_page` run; `DELETE /runs/{rid}` - owner-scoped, with an admin fallback so admins can delete any user's run; publish/subscribe under `/{id}/{publish,unpublish,subscribe}` and `GET /published`. **Admin (cross-user):** `GET /admin/all` - every task; `PATCH/DELETE /admin/{id}` - edit/delete any task; `POST /admin/{id}/trigger`; `GET /admin/{id}/runs`. Admin routes require `system_role == "admin"`. |
| **HTML Page Favorites** (`/api/html-page-favorites`) | Bookmarked generated HTML pages, owner-scoped. `GET /` - list; `POST /` `{page_id, title?, description?, tags?}` - bookmark a `scheduled_html_pages` row (extracts a reusable style/layout summary via GLM5 at favorite time); a page can only be favorited once per user (`DuplicateFavoriteError` ? HTTP 409). `GET/DELETE /{id}` - read/delete (the create-task reference picker deletes favorites here). A favorite can be selected as a style reference when creating a new `html_page` task. |
| **Public Scheduled Tasks** (`/api/public/scheduled-tasks`) | **Unauthenticated** mirror for embedding. Adds `GET /runs/{rid}/html` and `GET /{task_id}/latest-html` for the `html_page` preview; the run result keeps `result.html` (only file paths are redacted). |
| **Light Apps** (`/api/light-apps`) | Light-app registry (????? / ????), **global/shared** (no user scoping; `created_by`/`updated_by` audit-only). Two kinds: `window_open` (launched in a new tab from the ???? card grid) and `iframe` (embedded under a parent sidebar menu via `mount_location` and dynamically injected into the sidebar). Both carry `route_params` ? a list of `{key, paramType: login\|theme, source, customValue, themeValue}` descriptors the **frontend** resolves at launch time (login params read from `localStorage.userInfo`; theme params use a literal value) and appends to `url` as query string. `GET /` - list (filters `app_type`/`mount_location`/`enabled_only`/`search`); `GET /{app_id}` - read one (both open to any authenticated user); `POST /`, `PUT /{app_id}`, `DELETE /{app_id}` - **admin only**; iframe apps must carry a `mount_location` (400 otherwise); the unique `name` collides as HTTP 409. Backed by `deerflow.persistence.light_apps` (table `light_apps`, auto-created by `create_all`); store wired in `deps.py` as `app.state.light_app_store`. Tests: `tests/test_light_apps.py`. |
| **Menu Overrides** (`/api/menu-overrides`) | Sidebar-menu overrides (菜单管理), **global/shared** (no user scoping; `updated_by` audit-only). The static menu tree (icons/routes/`adminOnly`) lives in the **frontend** code (`core/page-layout/sidebar-menu.ts`); iframe light-apps are injected at render time. This router persists a per-node **overlay** the frontend merges onto that tree: `parent_id` (re-parent, moves the whole subtree; NULL = top level), `sort_order` (reorder among siblings), `label` (rename; NULL keeps code label), `disabled` (停用 — hides the node **and its subtree**). A node with no row renders exactly as code defines it, so new code menus / newly-registered iframe apps appear automatically and are immediately configurable. `GET /` - list all overrides (open to any authenticated user); `PUT /` - **admin only**, atomically **replaces the whole set** (`{overrides: [...]}`); rejects re-parenting that would exceed **3 levels** or form a cycle (`_validate_depth`, HTTP 400). Backed by `deerflow.persistence.menu_overrides` (table `menu_overrides`, auto-created by `create_all`); store wired in `deps.py` as `app.state.menu_override_store`. Frontend: merge logic `core/page-layout/menu-overrides.ts` (`applyMenuOverrides` in `use-sidebar-menu.ts`), client `strategy-components/api/menu-overrides.ts`, admin page `pages/MenuManagementPage.tsx` (@dnd-kit drag-reorder/re-parent + 「移动到」parent selector + ↑↓ arrows + 改名/停用), entry in the bottom-left settings dropdown (`components/page-sidebar.tsx`, admin-only). Tests: `tests/test_menu_overrides.py`, `frontend-web/src/core/page-layout/menu-overrides.test.ts`. |
| **Compact Menu Overrides** (`/api/compact-menu-overrides`) | Compact/sidebar-summary left-nav overlay (简介模式菜单), **global/shared**. Two independent sets keyed by `menu_type` (`A` default / `B`); saving one type does not wipe the other. Type B rows use a `B::` node_id prefix so they stay unique on databases that still use `node_id` as the sole PK. `GET ?menu_type=` lists one type (open); `PUT {menu_type, overrides}` admin-only replaces that type. Frontend: `MenuManagementPage` 简介模式 type switcher + `runtime-config.js` `VITE_COMPACT_MENU_TYPE`. Tests: `tests/test_compact_menu_overrides.py`. |
| **Sensitive Words** (`/api/sensitive-words`) | Sensitive-word / desensitize map (敏感词管理 / 脱敏词), **global/shared** (no user scoping; `updated_by` audit-only). Each row maps an original `word` → `replacement`, with `enabled` (停用 parks a word without deleting) + `sort_order`. The **frontend** ships a seed map in code (`src/lib/desensitize.ts` `SEED_DESENSITIZE_MAP`) used as the default on a fresh deployment; once an admin saves, this table is the source of truth and the frontend loads it at startup (`DesensitizeWordsLoader` → `replaceDesensitizeMap`, **enabled** rows only). `GET /` - list all words (open to any authenticated user — every page desensitizes); `PUT /` - **admin OR the `lqq` account** (`_require_manager`: `system_role == "admin"` or email `@`-prefix `== "lqq"`), atomically **replaces the whole set** (`{words: [...]}`, de-duped by `word`, last wins). Backed by `deerflow.persistence.sensitive_words` (table `sensitive_words`, PK `word` `String(191)`, auto-created by `create_all`); store wired in `deps.py` as `app.state.sensitive_word_store`. Frontend: client `strategy-components/api/sensitive-words.ts`, page `pages/SensitiveWordsPage.tsx` (route `/page/strategy/admin/sensitive-words`, entry in the bottom-left settings dropdown — **gated on username `lqq`, not admin role**, present in both sidebar modes; 「导入种子词」 merges missing seed entries). Tests: `tests/test_sensitive_words.py`. |
| **Sentiment Agent Proxy** (`/api/sentiment-agent`) | BFF for the frontend 舆情分析 virtual agent. `POST /stream` accepts the AG-UI body plus `apiUrl`/`authUsername`/`authPassword` from `runtime-config.js`, forwards to that URL with TLS verify off, and byte-pipes upstream SSE. Connection/HTTP failures return the full error in `detail`; mid-stream failures emit AG-UI `RUN_ERROR`. Distinct from `skill-packages/sentiment-analysis.skill`. Tests: `tests/test_sentiment_agent_proxy.py`. |
| **Dashboard Sessions** (`/api/dashboard-sessions`) | 「大屏绘制智能体」独立页面会话,**按外部 `task_id` 存取,不按 user 分权**(open read + shared writes;`user_id` audit-only)。一个 task 一行(upsert):`transcript`(聊天记录 JSON)+ `report_json`(最新结构化数据,可解析 = 「有效结果」)。`GET /by-task/{task_id}` - 取该 task 会话(无则 404);`PUT /by-task/{task_id}` - upsert(`{transcript?, report_json?}`,只覆盖传入项)。Backed by `deerflow.persistence.dashboard_sessions`(table `dashboard_sessions`, auto-created by `create_all`,无迁移);store wired in `deps.py` as `app.state.dashboard_session_store`。前端独立页 `roundtable-planning/pages/DashboardAgentPage.tsx`(路由 `/page/strategy/dashboard-agent?taskId=`,侧边栏「问答管理 → 大屏绘制」)让智能体分析任务报告 JSON → report-json → 渲染 `DashboardApp`。Tests: `tests/test_dashboard_sessions.py`。 |
| **Task Reports** (`/api/task-reports`) | 任务级研判情报报告详情,**按 `task_id` 共享、不按 user 分权**(写入者只留 `updated_by` 审计)。`GET /by-task/{task_id}` 始终返回原 Consumer 兼容外壳 `{state:"200", msg:"操作成功!", data:[...]}`(未保存为 `data: []`);`POST /by-task/{task_id}` 首次保存,已有报告返回 409;`PUT /by-task/{task_id}` 原子替换完整 `data` 数组。`POST /import/by-task/{task_id}` 面向 TaskCOP 导入页,接收 `sq-report-mock.json` 兼容的 `data`/`records` 数组并完整覆盖目标任务;路径中的选定任务 id 优先于文件内的 `taskId`,避免误写。每一项严格是 `{id, taskId, createTime, updateTime, sessionId, contentJson, categoryType}`;`contentJson` 以 JSON **字符串**原样落在 `records_json`,避免破坏大屏按类目解析的输入格式。保存或导入非空 `data` 后,router 会通过 `app.state.taskcop_task_store.mark_analysis_completed` 将同 id 的 TaskCOP 任务升级至状态 `25`,使旧任务列表显示并筛选为“已完成”;普通保存时不存在匹配任务的报告仍可独立保存,导入则要求目标任务存在。Backed by `deerflow.persistence.task_reports`(table `task_reports`, auto-created by `create_all`);store wired in `deps.py` as `app.state.task_report_store`。前端客户端 `strategy-components/api/task-reports.ts`,`core/auth/dashboard-input.ts` 与 rwfx 3q 完成检查均直接读此接口,不再依赖 Consumer 或本地 sq-report mock。Tests: `tests/test_task_reports.py`。 |
| **Task Buttons** (`/api/task-buttons`) | Task-workspace configurable jump buttons (按钮管理), **global/shared** (no user scoping; `updated_by` audit-only). Each row = one jump button in a task deep-link workspace's bottom 快捷跳转 bar: `{id, business (rwfx/3qfx/xdfx), label, link_type (url/business/purpose/host), target, append_task_id, id_param (editable query-param name, e.g. id/task_id), id_kind (task/action — only xdfx differs, its deep-link id is an 行动id), open_mode (blank/self), login_params, append_auth, auth_token_param, auth_name_param, enabled, sort_order}`. **Two independent, coexisting login-append mechanisms for `url` buttons** — both resolved frontend-side: (1) `login_params` (JSON column `login_params_json`, `PortableJSON`, migration `20260625_01`) = a list of `{key, source, custom_value}` resolved at click time (`source` = a `userInfo` field or `other`→literal `custom_value`) and appended to the target URL — mirrors the light-app login params; empty = nothing appended. (2) `append_auth` + `auth_token_param` + `auth_name_param` (columns, migration `20260625_02`, default off) — appends `?{auth_token_param}={token1}&{auth_name_param}={username}` so the target system can identify the current user; gated on **token login** — the frontend only stashes `token1` (the `?authToken=`) + the raw `getTokenInfo` username (`login.token1`/`login.username`) when login went through `/login/token`, so a non-token-login session appends neither. The **frontend** holds code-seed defaults (rwfx's 4 historical buttons built from the live `task_deeplink` config); only when this table is empty do the workspace + admin page fall back to them, else this table is the sole source of truth. `GET /` - list all (open to any authenticated user — the workspace renders them); `PUT /` - **admin only** (`_require_admin`), atomically **replaces the whole set** (`{buttons: [...]}`, de-duped by `id`, last wins). Backed by `deerflow.persistence.task_buttons` (table `task_buttons`, auto-created by `create_all`); store wired in `deps.py` as `app.state.task_button_store`. Frontend: helpers `core/auth/task-buttons.ts`, client `strategy-components/api/task-buttons.ts`, admin page `pages/TaskButtonsPage.tsx` (route `/page/strategy/admin/task-buttons`, entry「按钮管理」in the bottom-left settings dropdown, admin-only, both sidebar modes; per-business tabs + 「导入默认」). Tests: `tests/test_task_buttons.py`. |
| **Fixed Questions** (`/api/fixed-questions`) | Administrator-managed exact-match fixed Q&A (问题配置), **global/shared**. Each row is `{id, question, answer, enabled, tokens_per_second, sort_order}`; stream speed is configurable per row from `1–1000 token/s` and defaults to `100`. `GET /` and whole-set atomic `PUT /` are both admin-only. At the start of an external chat run, `services._find_fixed_answer` checks the latest eligible plain-text human message with exact, case/space/punctuation-sensitive equality. Enabled matches switch the run to `make_fixed_answer_agent`: a local `BaseChatModel` tokenizes with `cl100k_base`, replays the answer in rate-limited token chunks, and keeps the existing LangGraph SSE `messages`/`values` protocol, run journal, checkpoint history, deterministic local title, and refresh behavior intact while no provider/model request or model-concurrency slot is used. File/multimodal/context-prefixed turns, `canvas_agent`, `ai_writing`, writing mode, resume commands, and internal orchestration/background child runs bypass matching. Storage failures fail open to the normal model path. Backed by `deerflow.persistence.fixed_questions` (table `fixed_questions`, SHA-256 indexed lookup plus exact original-string collision check, migrations `20260821_01` + `20260821_02`), wired as `app.state.fixed_question_store`. Frontend: `strategy-components/api/fixed-questions.ts`, admin page `pages/FixedQuestionsPage.tsx`, route `/page/strategy/admin/fixed-questions`, entries in both full-layout and compact-workspace settings menus. Tests: `tests/test_fixed_questions.py`. |
| **Knowledge** (`/api/knowledge`) | Product knowledge base, **global/shared** (no user isolation; writes record `created_by`/`updated_by` for audit). `POST /ingest/thread/{thread_id}` - sediment one thread into a note (reads messages from the checkpointer, `threads:read` owner-checked); `POST /ingest/threads/batch` - batch-sediment the caller's recent threads (`skip_existing` via `.manifest.json` dedup). **Defaults to `background=true`**: the slow per-thread extraction loop is deferred to a FastAPI background task and the endpoint returns `{queued:true, processed:N}` immediately (each thread's transcript is truncated to `max_input_chars` before the LLM call). Running the loop inline for a large batch outlives the nginx gateway timeout (504), which the frontend's `parseResponse` surfaces as a generic "Request failed"; `background=false` keeps the blocking path (full per-thread `created/skipped/failed` tally) for small batches/scripts. Shared loop: `_run_batch_ingest`. **Optional model selection**: both ingest endpoints accept `model_name` (overrides `knowledge.extract_model_name` for that capture; threaded through `KnowledgeService.capture_thread` ? `extract_thread_llm_multi`). **Model-error surfacing**: when the chosen model produces nothing usable (timeout / API error / non-JSON) capture still **falls back to rule-based extraction and lands a draft** (`llm_fallback:true`); the single-thread (non-background) response and the batch response (`llm_fallback` count) carry this so the UI tells the user "???????". The "??????" chat button stays background/fire-and-forget (passes `model_name`, no per-save error toast); the batch-import dialog is synchronous and shows the fallback count. `POST /ingest/file` - **import an uploaded file** (multipart: `file` + optional `title`/`tags`/`status`/`model_name`) into wiki note(s): the file is converted to Markdown (md/txt read directly, pdf/word/ppt/excel via `convert_file_to_markdown`), distilled by the same wiki-capture LLM pipeline as thread capture (`source_type="file"`), and falls back to a verbatim **draft** note when the model returns nothing usable (`llm_fallback:true`); `POST /ingest/search` - sediment search/tool results as a reference note; `POST /notes` - manual note; `GET /notes` (keyword/tag/source_type/status + pagination); `GET/PUT/DELETE /notes/{id}` (DELETE is soft = archive, `?hard=true` removes the vault file); `POST /search` - `mode=keyword\|vector\|hybrid` (vector?keyword fallback when no embedding endpoint); `GET /notes/{id}/versions` - edit history; `GET /duplicates` + `POST /merge` - dedup; `POST /reindex` - rebuild vector index; `GET /graph` - product graph data (notes + entities, graph-blacklist filtered); `GET/POST /graph/blacklist` + `DELETE /graph/blacklist/{id}` - graph node blacklist (labels = entity names / note titles, case-insensitive match; blacklisted nodes and their edges are dropped from `/graph` and the wiki-export artifacts; table `knowledge_graph_blacklist`, auto-created); `POST /export/graph` + `GET /export/files/{graph.json\|graph.html\|graph.graphml\|cypher.txt}` - obsidian-wiki `wiki-export` artifacts (self-contained `graph.html` for iframe); `POST /resolve` `{targets:[?]}` + `GET /notes/{id}/wikilinks` - resolve `[[wikilink]]` targets to a note id / vault page / missing (powers clickable wikilinks); `GET /vault-page?path=` - read a raw vault meta/hub Markdown page (no DB note). `GET/POST /folders` + `PUT/DELETE /folders/{id}` - user-managed directories (notes carry a `folder` path; `GET /notes?folder=` filters a dir + sub-dirs); `GET/POST /extract-templates` + `GET/PUT/DELETE /extract-templates/{id}` - editable extraction-prompt templates (thread/document ingest accept `template_id`; capture/PUT-note accept `folder`). Backed by the `deerflow.knowledge` module (see below). |
Proxied through nginx: `/api/langgraph/*` ? LangGraph, all other `/api/*` ? Gateway.
### TaskCOP Situation-Overview Compatibility
The standalone legacy Vue page calls the Gateway URL in its runtime `b1ConsumerUrl` configuration directly (no Vite `/api` reverse proxy). The Gateway CORS middleware merges standard local Vite origins (`localhost`/`127.0.0.1` on ports 5173, 5174, 3000, and 8080) even when `GATEWAY_CORS_ORIGINS` was set for another frontend. This avoids a restart regression when the existing Vite server is on 5173 but the launcher sets 5174. Deployments must set `GATEWAY_CORS_ORIGINS` to any additional exact frontend origins that are allowed to call the Gateway.
`app/gateway/routers/taskcop_tasks.py` is a temporary compatibility layer for the separately deployed legacy Vue page at `/situation-overview/action/sentiment` and its task import page at `/situation-overview/action/task/import`. It preserves the old Consumer paths and response envelopes: `POST /cop/saveSuperiorTask` creates a row with a sequential four-digit string id (`0001`, `0002`, ...) and returns `{state: 201, msg, data}`; `POST /api/taskcop/import/tasks` accepts a JSON array or `tasks`/`records`/legacy list envelope, uses each supplied `id` or `taskId` as the primary key, and completely replaces duplicate IDs; `DELETE /cop/tasks/{task_id}` removes that task and its shared report document; `GET /taskAnalyseSearch/cop-task-three-list` (all), `GET /taskAnalyseSearch/cop-my-tasks` (creator-scoped), and `GET /taskAnalyseSearch/cop-task-list` return `{state: 200, data: {total, records}}`; `GET /taskAnalyseSearch/cop-task-detail?taskId=` returns one task. Existing UUID-style task rows are retained, and do not affect the new numeric sequence. An empty TaskCOP store remains empty—there are no built-in task or report seed records. It supports the page's existing `content`, `taskStatus`, `taskDirection`, `startTime`, `endTime`, `pageNum`, and `pageSize` query fields. Tasks are stored in `taskcop_tasks` through `deerflow.persistence.taskcop_tasks`, registered by `persistence.models`, and initialized as `app.state.taskcop_task_store` in `deps.py`. A non-empty save/update/import through `/api/task-reports/*` calls `mark_analysis_completed(task_id)`: it changes only pre-completion task states to `25`, so the list renders and filters the task as “已完成” without downgrading later workflow states. `trendSucc` is derived from the same status for the legacy sentiment list. Zero-padded task ids stay strings in the report API, so `0001` is never converted to `1`. The Vue router calls DeerFlow's `/api/v1/auth/login/username` directly with an incoming `username`/`userId`/`yUserId` parameter; it does not depend on the retired Consumer login service, and sends the returned DeerFlow JWT to every TaskCOP call. The TaskCOP routes are no longer public or CSRF-exempt. The Gateway fallback CORS allow-list includes TaskCOP Vite origins `http://localhost:8080` and `http://127.0.0.1:8080`; deployments must use explicit `GATEWAY_CORS_ORIGINS`. Tests: `tests/test_taskcop_tasks.py` and `tests/test_taskcop_cors.py`.
### AI Writing Intent Router (对话驱动写作台意图路由)
Backs the frontend「AI 写作·对话版」page's unified bottom composer. `POST /api/ai-writing/sessions/{session_id}/intent` (owner-checked; in `routers/ai_writing.py`) reads the session's graph checkpoint via the shared `app.state.checkpointer` — the **pending interrupt is the authoritative pause point** (the frontend's claimed `pause_point` is only a fallback when the checkpoint can't be read) — then routes the user's free text through `deerflow/agents/ai_writing/intent_router.py::parse_user_intent`: an LLM call (session's `model_name` + `fallback_chain` model fault-tolerance) that returns `{intent, pause_point, action, payload, need_research, research_query, answer, confidence, clarify}`. Safety boundaries in `normalize_intent_result`: **per-pause action whitelist strictly aligned to graph.py's real resume actions** (`ALLOWED_ACTIONS`), no pause → no action, high-cost actions (`finalize` / `force_finalize` / query-less覆盖式 `re_search`) require confidence ≥ 0.85 else downgrade to `clarify`, payload key-whitelist cleaning (material-id filtering, review-issue index bounds, section-decision back-fill with the majority decision). QA intents return `answer` directly from the context snapshot (`build_context_snapshot` — stage-tailored: materials with ids / outline / draft body capped / review issues with indices). Model failures never 5xx — they degrade to clarify. Prompts in `agents/ai_writing/prompts/intent_router_prompts.py`; tests `tests/test_ai_writing_intent.py`.
### AI Writing Built-in Functional Agents (????????)
The AI-writing graph (`packages/harness/deerflow/agents/ai_writing/`) has four stage "brains" ? ?????? / ????? / ?? / ?? ? that can each run in one of two modes, switched per-stage by `config.yaml ? ai_writing.use_builtin_agents.{researcher,outliner,writer,editor}` (default **false**, mtime hot-reloaded, new sessions pick up changes immediately):
- **Legacy (flag off)**: hard-coded prompts in `prompts/*.py`, direct LLM call (unchanged behavior).
- **Builtin-agent (flag on)**: the node loads the corresponding functional agent from `.deer-flow/agents/ai-writing-<stage>/` ? `SOUL.md` becomes the system prompt and `config.yaml`'s `model` overrides the stage model (priority: `agent.model > node explicit fallback (keyword_model/rank_model) > state.model_name`). Admins edit SOUL/model from the agent management page. Runner: `nodes/_agent_runner.py` (`run_writing_agent` / `stream_writing_agent` / `parse_marker_json`). Output contracts (marker + ```json``` blocks, intent.py-style): researcher `[KEYWORDS_READY]` / `[MATERIALS_RANKED]`, outliner `[OUTLINE_READY]`, editor (per-dimension) `[REVIEW_READY]`; writer emits raw section Markdown (keeps per-section streaming) with the legacy `[[SECTION_BLOCKED]]` sentinel in strict mode. **Orchestration stays in the nodes** (search fan-out, per-section loop, 3-dimension review loop, all pause points) ? the agent only does single structured generations, and a missing/empty SOUL silently falls back to the legacy path (`WritingAgentUnavailable`).
Seeding mirrors the roundtable pattern: verified `config.yaml` + `SOUL.md` ship in `app/gateway/routers/_ai_writing_seed_assets/<id>/`; `_ai_writing_seed.py::ensure_ai_writing_functional_agents()` copies them to `.deer-flow/agents/` when missing (never overwrites local edits), called from the startup lifespan (before `_sync_legacy_agents`, so the agents appear in `/api/agents` on first boot) and self-heals from `GET /api/ai-writing/sessions` + `GET /api/ai-writing/article-types`. Tests: `tests/test_ai_writing_seed.py`, `tests/test_ai_writing_agent_runner.py`.
**Sample imitation mode (样文仿写).** The initial writing form supports `writing_mode_type=imitate`: the user can paste sample text or upload Word/PDF/Markdown/TXT. `POST /api/ai-writing/sample/extract` converts supported document files to Markdown via `deerflow.utils.file_conversion` and returns capped plain text to the frontend. The LangGraph entry point now starts at `sample_analyzer` before `intent_parser`; non-imitation sessions return immediately, while imitation sessions analyze up to 12k chars into `sample_style_profile` (`SampleStyleProfile` in `state.py`) and emit `sample_profile_ready`. `writer_outline` and `writer_draft` call `render_sample_style_section(state)` so both outline planning and section drafting receive the same style profile, imitation strength (`light|medium|strong`), optional structure-preservation preference, and guardrails: learn structure/tone/paragraph rhythm/sentence style, but do not copy sample sentences, proprietary facts, people, organizations, cases, or data. Frontend state mirrors this with `analyzing_sample`, uploaded/pasted sample controls, history/timeline restoration of `sampleStyleProfile`, and the setup-chat direct-start path preserves the same fields. Tests live with the knowledge-base AI-writing tests (`tests/test_ai_writing_knowledge_base.py`).
**Skill-driven retrieval (检索类型).** `material_source` has three values: `general` (通用检索, default), `knowledge_base` (知识库检索), and `notebook` (我的空间); legacy `skill` is normalized to `general` by the researcher node. **`knowledge_base` 直连内网 ES 接口(= 旧版「内网直连检索」)** — it bypasses skill orchestration entirely and calls `IntranetSearchProvider` over `ai_writing.intranet_search_url` (concurrent over the generated keywords; `intranet_knowledge_base`/`intranet_search_timeout` apply). **An empty/unconfigured `intranet_search_url` does NOT disable it** — the empty string is passed straight to `IntranetSearchProvider`, which falls back to its built-in default endpoint (`intranet_stub._DEFAULT_ENDPOINT`), exactly like the legacy `get_search_provider(search_provider=intranet)` path. So 内网直连检索 works out of the box without configuring an address (configure `intranet_search_url` only to point at a different endpoint). This is the "弱模型识别不了技能" escape hatch — the old no-skill retrieval path restored as an explicit option (weak models like deepseek-chat often fail to invoke configured skills). Tests: `tests/test_ai_writing_knowledge_base.py`. General retrieval is orchestrated from the **skills configured on the `ai-writing-researcher` agent** (`config.yaml → skills`, editable from the agent management page): the built-in `knowledge-base-search` skill (知识库检索) runs the native keyword fan-out web/intranet search (the old default path), while any other configured skill searches with a `description`-targeted query; skills run in configured order with early-stop at `max_materials`, and an empty/unresolvable skills config falls back to the built-in knowledge search so retrieval always works. The skill itself is seeded from `_ai_writing_seed_assets/skills/knowledge-base-search/SKILL.md` into `skills/public/` (never overwrites), and a pre-existing researcher `config.yaml` lacking the `skills` key is patched in place with the default (explicit `skills: []` is respected). **PAUSE-1 input append flow**: when the user types natural language into the material-confirm input (`re_search`/legacy `enable_skill_search` + `user_query`), the node matches the most relevant configured skill via `SkillPicker(candidates=…)` (skipped when only one skill is configured), searches with the user's words, and **appends** the new materials to the existing package (URL-dedup, original order preserved, no LLM re-rank/truncation, 100-item hard cap). **Skill-failure fault tolerance**: a configured skill that times out / raises / has all its keyword searches fail (vs. legitimately returning 0 results) no longer just silently moves on — the researcher node emits a `skill_search_step` of type `skill_error` (carrying the error `reason`), and if the configured skills collected **no** new materials at all it runs a `fallback` pass over other available skills (`_fallback_candidate_skills`: the built-in `knowledge-base-search` native path first, then other enabled catalog skills not yet tried) instead of erroring out. The frontend `SkillStepsCard` renders `skill_error` (rose) and `fallback` (amber) so the timeline shows "技能报错 → 改用其他技能搜集". Tests: `tests/test_ai_writing_material_source.py`.
The intranet client uses `trust_env=False` and defaults `intranet_verify_ssl=false`, so internal HTTPS endpoints with self-signed/internal-CA certificates do not fail on certificate verification; set `ai_writing.intranet_verify_ssl: true` only for environments that require strict certificate checks.
**Skill-call timeout** is hardcoded **300s (5 min)** — `researcher.py` `_SKILL_SEARCH_TIMEOUT_SECONDS` (per-skill wrapper) and `_skill_agent.py` `_REAL_AGENT_TIMEOUT_SECONDS` (whole real-agent skill run). The intranet ES fallback (`IntranetSearchProvider`) gets a longer **600s (10 min)** default (`intranet_stub.py`; overridable via `ai_writing.intranet_search_timeout`) — that interface is the slow one, so our own fallback timeout is generous to avoid being blamed for "检不出".
**Parallel intranet fallback (技能 0 条 → 直接用并行预跑的内网结果).** When `ai_writing.intranet_search_url` is configured, the researcher node launches the intranet ES search (`IntranetSearchProvider` over the generated keywords) as a **concurrent `asyncio.Task` the moment the skill loop starts** (not sequentially after all skills fail). If the configured + fallback skills collect **nothing**, the node awaits that already-running task and uses its results (`_collect`) — no extra wait. If the skills **did** get materials, the task is cancelled (`_cancel_intranet_task`, suppresses CancelledError) so the intranet call isn't wasted. URL empty → no task, behavior unchanged. Tests: `tests/test_ai_writing_intranet_parallel.py`.
**Partial-result harvest + query simplification (检不出/超时的容错).** When `researcher_skill_agent: true` runs a non-native skill via `run_skill_via_real_agent` (the deployment's real path — the skill, e.g. `knowledge-search-v2`, is invoked by the real agent through `bash python3 scripts/search.py "<query>"`), one writing run issues many queries and **some time out** (the skill script prints `{"success":false,"error":"请求超时…","count":0}`) while others succeed (`{"success":true,"count":30,"results":[…]}`). Two fixes so successful data is never dropped: (1) `_drive()` captures every **ToolMessage** stdout and `_harvest_tool_results()` parses the skill's JSON **by reusing the lead-agent 参考文献 mode's proven extractor** (`deerflow.runtime.references`: `_parse_json_like` + `_first_list` + `_normalize_reference_item`) — the same logic that already extracts skill JSON into citations 100% reliably on ChatPage/AgentChatPage, with its huge field-name union (title/m_title/name, content/page_content/content_preview/snippet/…, recUuid/uuid/docId/…) and per-skill display-config support — plus skips `success:false` outputs — so even if the model's final `{materials:[]}` summary gives up (confused by the timeouts), the successful queries' real results are still collected; the function merges model materials + harvested results (dedup by url/recUuid/title, capped at `max_results`). **The harvest+merge (`_finalize`) runs on EVERY exit path — normal end, overall timeout, AND `await task` exceptions (e.g. `GraphRecursionError`)** — so when the run is cut off before the model writes its summary (observed: 10 queries × ~2 graph steps > the old `recursion_limit`, so the agent errored out with the real 30+ results already in `tool_outputs`), those results are still salvaged instead of returning `[]`. `recursion_limit` was also raised from `_REAL_AGENT_MAX_TURNS` to `_REAL_AGENT_MAX_TURNS*2+4` (it counts graph super-steps ≈ 2 per model→tool round, not model turns) so the model normally has room to finish its summary. **Early-stop on `max_results`:** `_drive` re-harvests after every tool output and, once the deduped harvested count reaches `max_results` (= `per_skill` = `max_materials`), **breaks the nested `astream`** and emits a 「已获取 N 条素材,达到上限,停止本技能检索」 step — without it a weak model keeps issuing new queries forever (observed: 100+ step-bar entries, materials far over the limit). Because the skill now returns exactly `max_materials`, the researcher loop's own `collect_limit` early-stop fires right after, so later configured skills don't run either. (2) The model is instructed (in `_MATERIAL_OUTPUT_CONTRACT`) that a 0-result/timeout query must be **retried with a simpler, core query** (drop topic suffixes like 竞选/诉讼/政策, keep the core entity: 「特朗普竞选」→「特朗普」) and that any data already retrieved must always be kept. The same deterministic simplifier `search/base.py::simplify_query` is wired into `IntranetSearchProvider.search`: an empty first attempt retries once with the simplified query. **(3) Root cause of the instability — bash output middle-truncation.** The skill prints its results as JSON to stdout, and the `bash` tool **middle-truncates** any output over `sandbox.bash_output_max_chars` (default **20000**, inserting a `... [middle truncated …] ...` marker). A `count:30` skill response (~30 records × ~600 chars) exceeds 20000 → the JSON is cut in the middle → **unparseable as a whole** by both the model and `_try_load_json` → all materials lost (while a small `count:9` response parses fine — hence the flapping). Fix: when whole-JSON parse yields no items, `_harvest_tool_results` falls back to `_scan_balanced_json_objects` — a brace-depth/string-aware scanner that recovers **every individually-complete `{…}` record object** from the truncated text (the records on each side of the truncation marker are still valid JSON), so ~16+16 records survive a 20000-char cut, far above `max_materials`. This is why the lead-agent chat pages (`/page/workspace/chats/new`, agent chat) call the same skill with 100% success — they only need the model to summarize prose from whatever it sees, never to parse a structured list, so truncation doesn't hurt them; ai-writing needs the structured records, so it must salvage them itself. Tests: `tests/test_ai_writing_skill_real_agent.py`, `tests/test_ai_writing_query_simplify.py`.
**Agentic skill execution (`ai_writing.researcher_skill_agent`).** When this flag is on (the deployment `config.yaml` enables it), the directed-skill branch runs `nodes/_skill_agent.py::run_skill_via_real_agent`: it reuses `SubagentExecutor` with the full tool stack, sandbox middleware, and a one-skill whitelist so the skill can `read_file` its `SKILL.md`, call `bash`, MCP tools, or whatever the skill instructions require. The nested run now passes `agent_id=ai-writing-researcher` in both `config.configurable` and runtime `context`, matching normal AgentChat skill visibility/whitelist behavior. It also appends a human-turn hard directive listing the exact SKILL.md path and skill directory, requiring weak models to first `read_file` the SKILL.md and then `cd <skill-dir> && python/python3 scripts/*.py ...` when the skill is script-backed; the model is explicitly forbidden from returning `materials` before executing the skill's required tools. If the real-agent path fails or yields no materials, `_execute_skill` still falls back to the legacy description→query search so materials are never empty. The built-in `knowledge-base-search` native concurrent path is unchanged. Tests: `tests/test_ai_writing_skill_real_agent.py`.
**Targeted user revision (局部修改 — 只改命中章节).** When the user submits a free-text 修改意见 on the draft-confirm card (`action == "user_revise"`), `writer_draft_node` no longer always rewrites the whole article section-by-section. Before the full-rewrite loop it tries `_targeted_user_revision` (`nodes/writer_draft.py`): `_split_draft_sections` parses the previous `current_draft.full_markdown` into `(article_title, [(section_title, body)], references_block)` (level-1 `#` title, level-2 `##` sections, trailing `## 参考文献` pulled out and kept verbatim), then `_detect_revision_scope` runs a small `llm_json` classifier (prompts `DRAFT_REVISE_SCOPE_{SYSTEM,USER}`) returning **per-section targets** `[{index, new_title, rewrite_body}]`: it maps the note to the section indices it targets ("第一段/开头/引言"→0, "结尾/结论"→last, by title/topic) **and** distinguishes a **章节标题改名** (`new_title` set, `rewrite_body=false` → rename only, body untouched) from a **正文修改** (`rewrite_body=true`). Title-only renames apply directly to the assembled `## {new_title}` heading **and are written back into `current_outline.sections[i].section_title`** (so the renamed title persists through confirm/finalize/history — this fixes "改章节标题不生效"). Body rewrites run `_stream_revise_section` (prompt `DRAFT_REVISE_SECTION_USER`, streamed as `draft_chunk`). **Only hit sections are touched; unhit sections are preserved byte-for-byte and are NOT re-emitted to the stream** — the frontend clears the right canvas to its initial placeholder on entering `revising` and shows only the live rewrite of the changed section(s), so the user sees「清空 → 流式重写 → 终稿」rather than the whole article rebuilding. The classifier returns `whole_article=true` (整体性意见 like 「整体语气更正式」/「全文都改」, **structure changes** 加章/删章/调整顺序, or a **文章总标题** change) → `_detect_revision_scope` yields `None` → `_targeted_user_revision` returns `None` → the node **falls back to the whole-article rewrite** (which still runs `_revise_outline_with_notes` for title/structure changes). Also falls back when there is no prior draft, an empty note, or scope-detection errors. Editor-review revisions (`accept_review`) keep the holistic rewrite. Reassembly preserves the article title + structure and re-appends the original references block; word count is computed on the body (excluding references). Frontend: `isRevising` in `AIWritingState` (`types/ai-writing.ts`) is set in `useAIWritingStream.ts`'s `intervention_submitted` (`ns === 'revising'`) and cleared on `draft_ready`; `AIWritingDraftPanel.tsx` shows the initial `DraftWaitingPlaceholder` while `isRevising && streamingDraftSections.length === 0`, then the streaming preview once chunks arrive. Tests: `tests/test_ai_writing_targeted_revision.py`.
**History draft editing (历史回看可编辑成稿).** The right-side draft in 历史回看 is no longer read-only — the user can edit it like the live 写作完成态 (Tiptap + 目录 + 格式工具栏 + AI 润色) and **persist** the change back to the session. `PUT /api/ai-writing/sessions/{id}/draft` `{draft_markdown, draft_title?}` (owner-checked → 404 otherwise; `draft_title` omitted ⇒ keep原标题) updates the row via the existing `AIWritingSessionRepository.update`. Frontend: `AIWritingDraftPanel` now passes `readOnly={false}` + an `onSaveDraft` handler (only in history mode) → `TextArtifactRenderer` → `WritingToolbar` 「保存」button (loading + toast); the handler is `AIWritingHistoryContext.saveHistoryDraft(markdown, title?)` → `updateAIWritingDraft` (`api/ai-writing-sessions.ts`), which write-backs the returned summary to `historySession`. Edits to the **live** completed draft stay ephemeral (export-only) as before — only history mode shows 保存. Tests: `tests/test_ai_writing_draft_update.py`.
**Background suspend (后台挂起).** A writing session can be handed off to the server to auto-drive the remaining flow to finalize, surviving page leave / refresh / account switch. `POST /api/ai-writing/sessions/{id}/background` (owner-checked, idempotent) starts `AIWritingJobExecutor` (`app/gateway/ai_writing_job_executor.py`), wired in `deps.py` as `app.state.ai_writing_job_executor`. The executor **drives the same `ai_writing` graph in-process** — it builds `make_ai_writing_graph()`, attaches the shared `app.state.checkpointer` (the very saver the runtime worker uses, so it continues the real thread state the frontend already wrote), sets the owning user via `set_current_user`, and loops `aget_state` → `ainvoke(Command(resume=…))` until `status==done`. **Auto-action per interrupt**: material/outline confirm → `confirm`; section_help → every blocked section `loose` (continue with general knowledge, never re-loop to research); draft_confirm → `to_editor` until reviewed, then `finalize` (on pass, or when `revision_count >= max_revisions`); review_confirm → `accept_review` (revise & re-review until pass or budget exhausted, then forced finalize). To avoid racing a still-running runtime worker (suspend mid-node), the driver **waits for an actionable state** (`_wait_for_actionable`): it only acts when the graph is at an interrupt / done, or when the checkpoint stops advancing (no worker → abandoned/restart). No new table — the graph checkpoint + `ai_writing_sessions` already hold all state; `status="background"` marks 挂起中 (surfaced in the history dropdown), nodes write `draft`/`review_result`/`status=done` as usual, so the user reopens the finished article from history. Per-user concurrency cap `AI_WRITING_BACKGROUND_SLOTS_PER_USER` (default 2, same SQLite-lock rationale as the roundtable executor). Startup `reconcile` (in `deps.py`) relaunches drivers for sessions left `background` after a restart (`AIWritingSessionRepository.list_by_status`). **Progress flowchart popup**: `GET /api/ai-writing/sessions/{id}/progress` (owner-checked) reads the graph checkpoint via `executor.get_progress` and derives a canonical pipeline `stage` key (`intent/research/material_confirm/outline/outline_confirm/draft/draft_confirm/review/done`, from the pending interrupt's `pause_point` or `snapshot.next` node) + `note` (素材不足求助 / 编辑打回) + review progress + `isRunning`. Frontend: 「后台挂起」 button + auto-suspend on route unmount / `beforeunload` (`AIWritingPanel.tsx`, `suspendAIWritingToBackground[Keepalive]` in `api/ai-writing-sessions.ts`); 挂起中 badge in `AIWritingHistoryDropdown.tsx`, and **clicking a 挂起中 history record opens a flowchart dialog** (`AIWritingFlowDialog.tsx` polls `getAIWritingProgress` every 2.5s, renders the fixed vertical pipeline `AIWritingFlowChart.tsx` lit by the live `stage`) so the user sees how far the background run has gotten without reopening/taking over the session. **Transcript consistency**: the rich history timeline (`AIWritingHistoryView` → `AIWritingTimeline`) replays `transcript = {progressEvents, completedInterventions}`, normally assembled+saved client-side. A background run has no client, so the executor rebuilds it at the end (`_save_transcript`/`_build_transcript_from_checkpoint`) **entirely from the final checkpoint** — both `progressEvents` and `completedInterventions` are derived from the same set of `resumed` events (every pause-resume emits exactly one), so they are always 1:1 aligned. For each `resumed`: classify its message → `(pausePoint, action)` (`_classify_resumed`, matched against `graph.py`'s resume messages), insert an `await_user` (with the pause prompt) before it, and build a **rich** intervention card enriched from final state (material_confirm→`materials`, outline_confirm→`outline`, draft_confirm→`draftTitle`/`draftWordCount`/`reviewResult`, review_confirm→`reviewResult`); `materials_ready` is inserted before the first material-confirm. (Earlier code merged the frontend's pre-suspend `completedInterventions` with auto ones and aligned by index — but the frontend's last save can predate its most recent decision, so the counts drifted and cards rendered off-by-one with empty data; deriving everything from the checkpoint root-fixes both.) Raw (snake_case) events **and** intervention data are normalized on load by the **exported** `normalizeProgressDict` / `normalize{Material,Outline,ReviewResult}` (idempotent on already-camelCase live transcripts — one normalize path for both sources), so reopening a background-completed session shows the same cards as a frontend-driven one. Tests: `tests/test_ai_writing_background.py`.
### Admin Leaderboard Snapshots (`/api/admin/users/leaderboard*`)
The admin ????? is served from **precomputed per-day snapshots** in `admin_leaderboard_daily_stats` (keyed by `(stat_date, settings_hash)`), not by aggregating raw rows on every request. A background scheduler inside the Gateway (`_leaderboard_snapshot_scheduler_loop` in `app/gateway/app.py`, singleton file-locked) ticks every `ADMIN_LEADERBOARD_BACKGROUND_INTERVAL_SECONDS` (default 300s): it prewarms recent days, then builds any `queued`/`error`/stale-`running` day. `GET /leaderboard` merges the per-day snapshots for the requested range. **Today** is rebuilt whenever its snapshot is older than `ADMIN_LEADERBOARD_TODAY_TTL_SECONDS` (default 43200s = half a day, so the leaderboard refreshes ~twice daily); a past day, once `ready`, is treated as fresh forever.
Two on-demand/automatic refresh paths layer on top:
- **Manual "????"** ? `POST /leaderboard/refresh-today` force-rebuilds *today's* snapshot **synchronously** (bounded by the same semaphore/timeout the scheduler uses), then returns the merged range. Lets an admin pull the latest real-time records into the stats without waiting for the TTL. Frontend: a "????" button on `AdminUserLeaderboardPage` ? `useRefreshTodayLeaderboard` (`core/admin/hooks.ts`), which primes + invalidates the `["admin-activity","leaderboard"]` query.
- **Daily finalize of the previous day** ? once per day, on/after `ADMIN_LEADERBOARD_FINALIZE_HOUR` (local Beijing, default `1` = ?? 1 ?), the scheduler tick force-recomputes *yesterday* a single time so runs/tool-calls that only landed near midnight are fully captured. The decision is the pure `deerflow.persistence.admin_stats.finalize.should_finalize_previous_day` (status/generated_at/now/hour) ? the rebuild stamps `generated_at` to today, which self-debounces it for the rest of the day (no extra state file). Tests: `tests/test_leaderboard_finalize.py`.
### Admin Concurrency Monitor (实时并发监控 / 系统压力)
A live admin monitor embedded **at the top of `AdminUserLeaderboardPage`** (`components/workspace/admin/concurrency-monitor.tsx`, `ConcurrencyMonitorPanel`) that answers "当前有几个并发的大模型调用、是谁、能不能直接停掉" and charts system pressure over time. Built on the existing in-memory `RunManager` (no new run plumbing).
- **Live concurrency + 真正停止** — `RunManager.list_active()` returns every run whose status is `pending`/`running` (the registry keeps finished records around until `cleanup`, so it filters by status). `len(list_active())` is the current concurrent-call count. Router `app/gateway/routers/admin_active_runs.py` (`/api/admin/active-runs`, **admin only** via the shared `admin_users._require_admin`): `GET /` returns `{count, server_time, runs:[{run_id, thread_id, status, created_at, elapsed_seconds, user_id, user_name, agent_id, agent_name, thread_title, kind, multitask_strategy}]}` — each run best-effort enriched (batched DB) with the owning user's email (`threads_meta.user_id`→`users.email`) and the agent's display name (`agents.name`); `kind` is derived from thread metadata (圆桌会商 / 系统/后台 / 任务对话 / 对话). `POST /{run_id}/cancel` calls `RunManager.cancel(run_id, action="interrupt")`, which sets the abort event **and cancels the run's asyncio task** — i.e. it truly aborts the in-flight model stream, freeing capacity. `POST /cancel-all` stops every in-flight run.
- **并发量折线图 (按分钟/小时)** — a lightweight background **sampler** (`_concurrency_sampler_loop` in `app/gateway/app.py`, singleton file-locked, started in the lifespan) records `len(list_active())` every `CONCURRENCY_SAMPLE_INTERVAL_SECONDS` (default 15s) into the new `concurrency_samples` table (`deerflow.persistence.concurrency`, `ConcurrencySampleRow{sampled_at, active_count}`, auto-created by `create_all`, **no migration**; store wired as `app.state.concurrency_sample_store`). ~Hourly it purges rows older than `CONCURRENCY_SAMPLE_RETENTION_HOURS` (default 168h). `GET /concurrency?granularity=minute|hour&hours=N` aggregates samples into per-bucket **peak + avg** (`ConcurrencySampleStore.aggregate`, bucketed in Python so it's identical across sqlite/mysql/postgres; fetch capped at 60k rows). Default lookback: minute→3h, hour→48h.
- **大模型出字速度折线图 (tokens/sec, 按分钟/小时)** — `GET /token-speed?granularity=&hours=` aggregates the persisted `llm_call_metrics` (`status="success"`) into per-bucket avg/min/max tokens/sec (deriving tokens/sec from `output_tokens÷duration_ms` when the stored value is null). A dip pinpoints the model/provider as the bottleneck for that window — "知道具体原因出现在哪里".
- **On/off + 不影响线上** — the whole feature is designed to never impact production: sampling is one tiny INSERT, retention-purged, all paths best-effort/guarded, queries bounded. A **runtime switch** lives in `system_settings.json → concurrency_monitor.enabled` (`ConcurrencyMonitorSettings`, **default false / opt-in**, mtime-hot-reloaded): the background loop always runs (just sleeps) but only samples when an admin turns it on, and it re-reads the flag **every tick** so toggling it on/off **takes effect immediately without a restart** (the on-demand live snapshot + 停止 stay available regardless; only the historical chart needs sampling on). `GET|PUT /api/admin/active-runs/settings` reads/writes it (admin only); the panel's 「后台采样」 Switch drives it. Env `CONCURRENCY_SAMPLE_ENABLED=0` is a hard master kill (never starts the loop). Frontend client/hooks: `core/admin/active-runs.ts` + `core/admin/hooks.ts` (`useActiveRuns` polls 4s while the panel is open, `useConcurrencySeries`/`useTokenSpeedSeries` poll 15s, `useCancelActiveRun`/`useCancelAllActiveRuns`, `useMonitorSettings`/`useUpdateMonitorSettings`). Tests: `tests/test_concurrency_monitor.py`.
### Scheduled Task Runtime
`ScheduledTaskService` (`deerflow/runtime/scheduler/service.py`) polls due tasks every `poll_interval_seconds` and executes each in its own throwaway thread, serialized by a shared `_agent_run_lock`.
**Stuck-run recovery**: each execution attempt is bounded by `_RUN_ATTEMPT_TIMEOUT_SECONDS` (30 min) via `asyncio.wait_for`. On timeout the run is retried up to `_MAX_RUN_RETRIES` (3) times, then marked `failed`. A watchdog (`_reconcile_stuck_runs`, run every poll) recovers runs left in `running` with no in-process owner ? e.g. orphaned by a crash or restart ? re-driving them through the same retry budget so a task can never stay perpetually "executing". The `scheduled_task_runs.retry_count` column tracks consumed retries.
**Unique task names**: task names are unique per user (`uq_scheduled_tasks_user_name`). `create_task`/`update_task`/`admin_update_task` proactively check via `_assert_name_available()` and raise `DuplicateTaskNameError` on a collision; the router translates that into HTTP `409` instead of letting the raw DB `IntegrityError` surface as a `500`.
**HTML page tasks (`task_kind == "html_page"`)**: a second task kind that renders a full HTML document each run instead of a Markdown report. The kind + its config (`html_config`: `page_prompt`, `negative_prompt`, `reference_favorite_id`, `model_name`) live inside `execution_context` ? no schema change to `scheduled_tasks`. Data gathering is delegated to the agent's own search skills (driven by `page_prompt`), so there is no explicit search-query/data-source field; the render model is user-selectable (`model_name`, default `zai-org/GLM-5-FP8`). `_run_task_attempt()` dispatches by kind; `_run_html_page_attempt()` runs the pipeline: gather real data via the lead agent (its search tools) ? optionally load a reference favorite's reusable style/layout ? render with the chosen model (thinking off) ? fault-tolerant JSON parse ? sanitize (BeautifulSoup, strips `<script>`/`<iframe>`/`on*`/`javascript:`) ? structural validate ? one GLM5 repair pass if invalid ? **completeness gate**: if the HTML is still not a full valid document (missing doctype/html/head/body, empty body, or residual scripts) after the repair pass, raise `HtmlGenerationError` so the run is marked **failed** (a page only counts as successful when it is a complete, previewable HTML document) ? best-effort Playwright screenshot (no-op when Playwright absent) ? archive `index.html` ? persist a `scheduled_html_pages` row ? return `result` with `format: "html"` and the full HTML inline. Pure helpers live in `deerflow/runtime/scheduler/html_page.py`. The frontend previews the HTML in a sandboxed `iframe` (no `allow-scripts`) on the run-detail and public pages; a user can bookmark a page (style summary extracted via GLM5) and later pick it as a style reference. New tables `scheduled_html_pages` + `html_page_favorites` (`deerflow/persistence/html_pages/`, wired as `app.state.html_page_store`) are created by `Base.metadata.create_all` with no migration. Tests: `tests/test_html_page_tasks.py`.
**Unified output switches (`require_html` + `require_markdown`).** The task form no longer splits "???? / HTML ??" into separate kinds. `execution_context` now carries two **independent** booleans that may both be on: `require_html` (???? HTML) and `require_markdown` (???? Markdown). `_task_requires_html()` returns true for `require_html` **or** legacy `task_kind == "html_page"` (so old tasks need no data migration), and `_run_task_attempt()` dispatches to `_run_html_page_attempt` when it is true, else the regular agent attempt. When **both** switches are on, `_gather_html_data()` returns `(text, files)` ? the gather agent run is additionally told to save a `.md` file ? and `_run_html_page_attempt` validates it (`MarkdownRequirementError` on failure) and surfaces it alongside the generated page, so the run delivers both artifacts. The single main `prompt` doubles as the HTML page requirement (`html_config.page_prompt`); there is no separate page-prompt field. Tests: `tests/test_scheduled_task_require_html.py` + the markdown-coexistence cases in `tests/test_html_page_tasks.py`.
**Template-fill mode (JSON → 固定 HTML 模板填充).** A third execution mode, distinct from the free-form `require_html` render pipeline: convert a user-supplied JSON into a **fixed** HTML template's data structure and fill it. Triggered when `execution_context.template_config.template_id` is set (chosen in the task form after selecting the built-in **「模板数据转换器」** agent `template-json-builder`). `_task_template_id()` reads it and `_run_task_attempt()` dispatches to `_run_template_fill_attempt` **before** the `require_html` check. Pipeline: take the source (the task `prompt`, pasted or uploaded .json) — which may be **pure JSON, plain text, or text-with-embedded-JSON** (`template_fill.extract_embedded_json` pre-extracts any embedded `{…}`/`[…]` with a string-aware balanced scan and surfaces it as a labeled block while keeping the full original text; plain text → no block, the model reads it semantically) → load the selected agent's `SOUL.md`+model and run a **direct** `_invoke_conversion_model` call (system=SOUL, human=conversion prompt carrying the template's compact **schema skeleton** + the recognized-JSON block + source) → `extract_json_object` → `template_fill.validate_data` (top-level keys + container types) → **on failure feed the structure errors + `field_coverage_gaps` back and retry up to `_MAX_TEMPLATE_FILL_RETRIES`=2** → `template_fill.fill_template` replaces the template's `const DATA={…}` line (and refreshes `MANIFEST.系统.更新时间`) → archive `index.html` + persist a `scheduled_html_pages` row → return `format:"html"` (reusing the existing HTML preview at `/runs/{id}/html`). Any terminal failure raises `TemplateFillError` → run marked failed (no retry, message shown). **Pure engine** `deerflow/runtime/scheduler/template_fill.py` (harness, no `app.*`): the template registry (`TEMPLATES`: id=`no1` name=`no.1` 台湾滨海防卫大屏, and id=`no2` name=`no.2` 台湾飞行部队平台 with **`convertible=False`** — a 「只展示」 template that renders its existing HTML **as-is**, no model/validation/fill, until conversion is wired in later; `_run_template_fill_attempt` short-circuits to `read_template_text` for it, and `list_templates`/`is_convertible` expose the flag so the form hides the JSON input), `list_templates`/`reference_data`/`schema_skeleton_json`/`validate_data`/`field_coverage_gaps`/`fill_template` — **the template's own embedded `DATA` is the schema source of truth** (parsed at load, 1-sample-per-array skeleton fed to the LLM, top-level structure as the validation basis); template HTML files live under `runtime/scheduler/templates/<id>/`, add a `TemplateSpec` to register a no.2. The agent is seeded like the AI-writing ones: `app/gateway/routers/_template_builder_seed.py` + `_template_builder_seed_assets/template-json-builder/{config.yaml,SOUL.md}`, called from the lifespan before `_sync_legacy_agents` (which upserts it into the `agents` table as built-in/published so it shows in `/api/agents` → the task form's 执行智能体 dropdown). Template list API: `GET /api/scheduled-tasks/html-templates` → `{templates:[{id,name,description,convertible}]}` (open). **Debug-render API** `POST /api/scheduled-tasks/template-fill/render` `{template_id, text?|data?}` → `{ok, display_only?, errors, gaps, html}` (open): no model call — `extract_embedded_json(text)` (or `data`) → `validate_data` → `fill_template`, returning the filled HTML **plus** the validation errors/field-gaps (fills even when invalid so you can see the effect). Powers the chat-page debug button below; non-convertible templates return their raw HTML. **Templates as `require_html` 参考收藏页面 (use the REAL template, all sub-pages)**: the 参考收藏页面 list in the task form also offers the templates (value `tpl:<id>`, e.g. `tpl:no1`). These templates are **JS-driven multi-page SPAs** (sub-pages/tabs are rendered by script from `DATA`), so a free-form model render would only ever produce the main page and **drop every sub-page**. Therefore, when `_run_html_page_attempt` sees a `tpl:<id>` reference (`template_fill.reference_template_id` parses it; distinguishes a template marker from a real favorite id), it short-circuits **before** the free-form pipeline to `_run_html_template_reference`, which renders the **actual template file** (all sub-pages intact): convertible templates gather real data (`_gather_html_data`) → convert to the template `DATA` (`_convert_to_template_data`, shared with template-fill mode) → **`merge_with_reference`** (per top-level key = per sub-page: any sub-page/field with no extracted content falls back to the template's **original** data, so sub-pages never render blank) → `fill_template`; non-convertible templates (no.2) output the raw template as-is; if the task prompt **and** description are both empty, or gather/conversion yields nothing, it uses the template's original data outright. **数据抽取开关** `html_config.template_extract` (default **off**, shown in the form only when a `tpl:` reference is picked): it **only** controls whether the HTML uses **extracted** data vs the template's **default** data — it does **not** affect anything else. Off → HTML fills the template's default data (no convert); on → gather+convert+merge. **Markdown is independent**: when `require_markdown` is also on, the gather agent run still runs and the `.md` is produced/validated/surfaced **regardless** of the extract switch (so turning extraction off never blocks the Markdown). Display-only templates (no.2) never extract; they still honor `require_markdown`. `_run_html_template_reference` runs gather once when `do_extract` (extract on + has prompt/desc + convertible) **or** `require_markdown`, and threads any `.md` through `_finalize_template_html(extra_files=…)`. `merge_with_reference` is also applied in template-fill mode (`_run_template_fill_attempt`) so empty sub-pages there fall back too. Archive/persist/result are shared via `_finalize_template_html`. (The earlier `style_reference` CSS-extraction approach was removed — it lost the sub-pages.) Frontend: when the agent dropdown = `template-json-builder` the form hides the HTML/Markdown switches and shows a 「HTML模板」 select (`useHtmlTemplates`) + a .json upload (reads file text into the prompt box). Tests: `tests/test_template_fill.py`, `tests/test_template_fill_attempt.py`, `tests/test_template_builder_seed.py`.
### Per-user Prompts & UI Preferences
**Custom prompts can be scoped to one agent.** `user_custom_prompts` gains a nullable `agent_id` column (migration `20260603_02`): `NULL` = global (the chip shows in every chat), otherwise the prompt only surfaces when that agent is selected. The visibility filter is applied in the frontend `CustomPromptBar` (it receives the current agent id from `ChatPage`'s selector or `AgentChatPage`'s `assistantId`); the backend (`/api/user-prompts` CRUD) just stores/returns/clears `agent_id` (update accepts an explicit `null` to clear). Tests: `tests/test_user_prompts_agent_id.py`.
**Per-user UI preferences** (`deerflow/persistence/user_preferences/`, wired as `app.state.user_preferences_store`; migration `20260603_03`, also auto-created by `create_all`): one JSON-blob row per user so new toggles need no schema change. Router `/api/user-preferences` (`GET`/`PUT`, account-scoped). Current key: `hide_admin_prompt_prefix` ? when true the chat input hides the admin-configured prompt-prefix switch and never applies the prefix (the user uses only their own prompts; enforced in `ChatPage` for both `shouldApplyPrefix` and `promptPrefixProp`). Toggled from the "???????" dialog. Tests: `tests/test_user_preferences.py`.
**HTML skills (tools the render step uses).** Two offline skills under `skills/public/` shape the generated HTML and are toggled in the skills settings page (enabled by default): `html-page-builder` (a self-contained inline-CSS design system ? colors/cards/tables/KPI/responsive) and `html-charts` (inline-SVG charts that render in the no-`allow-scripts` preview iframe; optional local ECharts documented in its `assets/README.md`). A skill opts into HTML injection via `metadata.kind: html_page` in its SKILL.md frontmatter; `_collect_html_skill_instructions()` reads the bodies of **enabled** such skills and `build_generation_prompt(skill_instructions=?)` injects them, so the (raw-model) render step honors the user's skill selection. The data-gathering step and the render/repair/style passes all use the task's selected `model_name` (no silent fallback to the default model ? and there is **no hardcoded GLM-5**; an unset selection falls back to config `models[0]`). To stop auxiliary LLM calls from leaking onto the default model, **`TitleMiddleware` is skipped when `is_scheduled_run` is set** (`_build_middlewares` in `lead_agent/agent.py`) ? scheduled/HTML runs use throwaway threads whose titles are never shown, and `config.title.model_name` is usually unset (? would call `models[0]`, not the chosen model). The completeness gate (`validate_html`) also requires a non-empty `<title>` and forbids external stylesheets, on top of the doctype/html/head/body + non-empty-body + no-script checks.
### Knowledge Base (`packages/harness/deerflow/knowledge/`)
Self-contained product knowledge base. **All knowledge capability lives in this one module** (harness layer, `deerflow.*` only ? no `app.*` imports). Wired as `app.state.knowledge_service` in `deps.py` via `make_knowledge_service(sf, config.knowledge)` (returns `None` on the memory backend or when `knowledge.enabled=false` ? router 503). Surfaced by `app/gateway/routers/knowledge.py` (`/api/knowledge`). Config: `config.yaml` ? `knowledge` (`deerflow/config/knowledge_config.py`).
**Two stores, kept in sync on every write:**
- **DB mirror** ? tables `knowledge_notes` / `knowledge_sources` / `knowledge_tags` / `knowledge_note_tags` (`knowledge/models.py`, registered in `persistence/models/__init__.py`, auto-created by `Base.metadata.create_all`; portable types `PortableLongText`/`BeijingDateTime`). Powers listing/filter/search/audit. Global/shared ? no `user_id` column; `created_by`/`updated_by` are audit-only (resolved via `get_effective_user_id()`).
- **Markdown vault** ? obsidian-wiki compatible, default `{$DEER_FLOW_HOME}/data/knowledge/obsidian-vault`. `ObsidianWikiAdapter` (`obsidian_adapter.py`) writes wiki-capture frontmatter+body, and after **every** write syncs `index.md` / `log.md` (CAPTURE/INGEST/UPDATE/ARCHIVE/EXPORT) / `hot.md` / `.manifest.json` (source-hash dedup for `skip_existing`). `markdown_vault.py` handles frontmatter build/parse, slugging, wikilink extraction and **path-traversal-proof** resolution.
**Module layout:** `models.py` (ORM: notes/sources/tags/note_tags + embeddings/versions/entities/relations) ? `repository.py` (DB only, no LLM) ? `schemas.py` (NoteDraft/SourceDraft DTOs) ? `skill_templates.py` ? `markdown_vault.py` ? `manifest.py` ? `extractor.py` (rule-based fallback) ? `llm_extractor.py` (**default**: distills declarative knowledge per wiki-capture, structured JSON incl. entities/relations; falls back to `extractor.py`) ? `sources.py` ? `exporter.py` (`graph.json`/`graph.graphml`/`cypher.txt` + self-contained vanilla-JS `graph.html`; entity nodes + relation edges. **`graph.html` is theme-aware** ? it reads `?theme=light|dark` from its own URL since the embedding iframe can't inherit the app's CSS; `KnowledgeGraphPage.tsx` appends the active day/night mode and re-skins on toggle. The layout cools down and **freezes** so dense graphs stay clickable, node radius scales with degree, and hovering/clicking a node focuses it + its direct neighbours while dimming the rest + a detail panel ? taming heavy node/edge clutter) ? `chunking.py` + `embeddings.py` (OpenAI-compatible embedding client, keyword fallback) + `search.py` (keyword/vector/hybrid) ? `sensitive.py` (secret/PII redaction) ? `dedup.py` (Jaccard + hash dedup) ? `auto_ingest.py` (background queue/worker) ? `obsidian_adapter.py` ? `service.py` (`KnowledgeService` orchestrator ? the only thing the router touches; capture entry points `capture_thread` / `capture_search` / `capture_file` / `create_manual_note`, all converging on `_persist_draft`). Tests: `tests/test_knowledge.py` (20 cases). Frontend: `frontend-web/src/core/knowledge/*` + `pages/KnowledgeBasePage.tsx` / `KnowledgeGraphPage.tsx` + chat `knowledge-save-trigger.tsx` / `knowledge-rag-toggle.tsx`; spec at `frontend-web/docs/product-knowledge-base-dev.md`.
**Wikilinks & multi-note capture** (`knowledge.{related_notes_enabled,related_notes_limit,multi_note_enabled}`, all default-on; `project_name` defaults to `zncm`):
- **Navigable `[[wikilinks]]`** ? `service.resolve_wikilink(s)` maps a `[[target]]` to a knowledge note (by id ? `vault_path` ? title) or a vault meta page (`read_vault_page`, path-traversal-proof), else `missing`. `markdown_vault.normalize_wikilink_target` canonicalizes targets (drops `|alias`, slashes, `.md`). The frontend `MarkdownRenderer` renders non-numeric `[[...]]` as clickable links from the resolver map (numeric `[[123]]` stay citation badges); note links navigate, vault links open a `VaultPageDialog`, missing links are greyed.
- **Real related notes only (no structural meta-links)** ? `extractor.related_links()` now returns `[]`: notes no longer auto-link `index` / `_meta/taxonomy` / a project hub (they were noise in every `## Related`). On capture, `service._find_related_notes` cross-links **only by shared *distinctive* tags** (project/`knowledge-base` defaults excluded; free-text title/summary matching was removed because CJK tokenization linked unrelated notes via generic words like ??/????) and `_inject_related_notes` writes real `[[Title]]` links into `## Related` (plus note?note `related` graph edges); no distinctive tags ? no related notes. The renderers omit the `## Related` section entirely when there are no related notes and no sources. The vault's own `index.md`/`_meta/taxonomy.md` still exist (obsidian scaffold) and remain resolvable for manually-written links. The router enriches `GET /notes/{id}` with `created_by_name`/`updated_by_name` (auth-provider email lookup) so the UI shows a person, not a UUID. The frontend `MarkdownRenderer` renders wiki-capture provenance markers `^[inferred]`/`^[ambiguous]` as small "??"/"??" badges.
- **LLM-extraction outcomes** ? capture distinguishes three LLM outcomes: (1) **success** ? declarative note(s); (2) **deliberate rejection** ? the model returns `should_ingest:false` with no notes (greeting / meta chatter) ? `extract_thread_llm_multi` raises `LLMRejectedIngest` and `capture_thread` raises `EmptyThreadError` (skip ? nothing persisted, so worthless chats never enter the KB); (3) **genuine failure** ? API error / non-JSON ? returns `[]`, capture falls back to the rule-based verbatim extractor and lands the note as a **`draft`** (not `approved`), logs a WARNING, and returns `llm_fallback:true` (surfaced by the router). `llm_extractor` logs an actionable WARNING/INFO for each non-success path so extraction issues are diagnosable.
- **Multi-note split** ? `llm_extractor.extract_thread_llm_multi` lets one conversation distill into 1?4 cross-linked notes (`notes:[?]` JSON array; legacy single-object output still accepted). `capture_thread` persists each, sibling-cross-links them by title, and returns `{note, created, notes:[?]}`. Gated by `multi_note_enabled` (false ? first note only). Tests: `tests/test_knowledge_wikilinks.py`.
**Phases 2?4 (all config-gated, defaults safe):**
- **Phase 2 ? vector RAG**: `knowledge.embedding.{enabled,base_url,api_key,model,...}` enables an OpenAI-compatible embedding endpoint; notes are chunked + embedded on write (`knowledge_embeddings`), `POST /search` supports `mode=keyword|vector|hybrid` (vector degrades to keyword when unconfigured). `KnowledgeRagMiddleware` (lead-agent chain) injects a transient `<knowledge-context>` recall before the model call, gated **per-conversation** by the `knowledge_rag_enabled` runtime flag (frontend chat toggle) falling back to global `knowledge.rag_enabled`. `POST /reindex` rebuilds the index.
- **Phase 3 ? governance**: `sensitive.py` redacts secrets/PII before persistence (`sensitive_redaction_enabled`); every update snapshots `knowledge_note_versions` (`GET /notes/{id}/versions`); `GET /duplicates` + `POST /merge` for dedup; background `AutoIngestQueue` (wired in `deps.py`, started/stopped in the gateway lifespan) auto-sediments finished conversations **off the chat path** when `auto_ingest_enabled` ? enqueued from `KnowledgeRagMiddleware.after_agent`, value-judged by min-chars + `skip_existing`, landing as `draft`.
- **Phase 4 ? graph**: LLM extraction also yields entities/relations ? `knowledge_entities`/`knowledge_relations`; `exporter.build_graph` adds entity nodes + note?entity/entity?entity edges (reflected in `graph.html`/`graph.json`); `GET /graph` returns product graph data.
**User directories + editable extraction templates + display polish (knowledge UX overhaul):**
- **Custom directories (`knowledge_notes.folder` + `knowledge_folders`)** ? notes carry a free-form, slash-joined `folder` path (e.g. `投研/行业`) independent of the wiki-capture `category` and of `vault_path`. `knowledge_folders` is the manageable directory list (supports empty folders + rename/move that re-points descendant folders **and** filed notes). Router CRUD under `/api/knowledge/folders` (`GET`/`POST`/`PUT/{id}`/`DELETE/{id}?reassign_to=`); `GET /notes?folder=` filters a directory + its sub-folders (`folder=""` ? unfiled). Capture (`capture_thread`/`capture_file`/`create_manual_note`) and `PUT /notes/{id}` accept `folder` (the note PUT uses `model_fields_set` so an absent key leaves the folder untouched, a present `null` clears it). Frontend `buildKnowledgeTree(notes, folderPaths)` builds the left tree from `folder` (no longer from `vault_path`); folder management dialog + `FolderSelect` in the create/import dialogs and the note detail header.
- **Editable extraction prompts (`knowledge_extract_templates`)** ? the hard-coded distillation prompts are seeded as editable templates on startup (`KnowledgeService.seed_default_templates`, idempotent — `默认 · 对话提炼` / `默认 · 文档提炼`, each `is_default` for its scope). `scope ∈ {thread, document, both}`; capture resolves an explicit `template_id` → the scope default → the built-in constant, and passes it to `llm_extractor.extract_thread_llm_multi` / `extract_document_llm_multi` as `system_prompt` (used verbatim). Router CRUD under `/api/knowledge/extract-templates` (409 on duplicate name). Frontend: extraction-template management dialog + `TemplateSelect` in the batch/file import dialogs. Editing a template is the supported way to fix "提炼太精简" (loosen the prompt to keep more detail) and to add alternate extraction logics.
- **Single source list** ? `extractor._related_section_md` no longer emits a `### Sources` sub-section into the body; the note detail shows sources once via the structured `来源` block (`note.sources`). `## Related` stays pure wikilinks so the frontend extracts it cleanly into the related-docs module.
- **Chinese section headings (display-only)** ? the frontend `localizeHeadings` (wikiUtils) maps `Context/Finding / Decision/Reasoning/Implications/Related/Sources` to 中文 at render time (applied to the body fed to both `extractToc` and the viewer, so anchors stay consistent). No data migration — historical English-heading notes render in Chinese too.
- **Historical-data cleanup** ? `scripts/cleanup_knowledge_related.py` (one-off, idempotent, `--dry-run`/`--note-id`) strips legacy `## Related` structural links (vault scaffold: `index`/`_meta/taxonomy`/`projects/…`, any slash-bearing target) and the old `### Sources` sub-section from existing notes (DB `content_md` + vault files, via `KnowledgeService.update_note`).
- Migration `20260616_02` + idempotent `scripts/sql/column_additions_mysql.sql` cover `knowledge_notes.folder` (+ index); the two new tables are auto-created by `create_all` / `_ensure_orm_columns_sync` on startup. Tests: `tests/test_knowledge.py`.
### Custom Agent Persistence
Custom agent metadata lives in the `agents` ORM table (`deerflow.persistence.agents`). `agents.id` is the stable runtime id and the filesystem directory name under `.deer-flow/agents/{id}/`; `agents.name` is display-only, can contain Chinese characters, and is not unique. `user_id IS NULL` marks built-in agents, user-owned agents are visible only to their owner unless `published = true`, and update/delete operations require ownership. The runtime accepts `agent_id` in context/configurable, and non-default `assistant_id` values are treated as agent ids after visibility checks.
**Listing cache + pagination** ? to keep `GET /api/agents` fast, the gallery's per-card extras (skills + tool_groups) are denormalized into two cache tables instead of reading every agent's `config.yaml` on each request: `agent_skills` (relational, one row per (agent, skill) with `position`) and `agent_extras` (one row per agent: `tool_groups` JSON + `model` + `skills_synced_at`). `config.yaml` remains authoritative for the agent runtime; these tables are a read cache, write-through on create/update and backfilled from disk by `_sync_legacy_agents` (mtime-vs-`skills_synced_at` drift check). The list endpoint enriches via `AgentStore.get_extras_for([ids])` (two batched `IN` queries). Both tables are auto-created by `Base.metadata.create_all` (no migration needed). `GET /api/agents` supports **opt-in server-side pagination**: `scope` (`mine` = owned + favorited / `square` = built-in/published public), `page`, `page_size`, returning `{agents, total, page, page_size}`; omitting `scope`/`page` preserves the legacy "return everything" behavior. Name `search` matches the raw OR desensitized name (the frontend's static term?abbrev map is mirrored server-side in `deerflow/config/agent_name_desensitize.py`); `tag_ids` is an SQL OR-filter; `owner` keeps only agents published by that user id (the card's ??? ? built-in agents never match), powering the gallery's click-???-to-filter. `POST /api/agents/private-square/list` takes the same pagination params (incl. `owner`). Repository entry point: `AgentStore.list_paginated(...)`. Tests: `tests/test_agent_listing_cache.py::test_list_paginated_owner_filter_keeps_one_publisher`.
**Homepage selector (agent picker on the new-chat page)** ? single admin-curated list backed by `agents.featured_order` (nullable BIGINT). NULL = not featured. Only `published = true` agents can be featured (router enforces on writes); any published agent is eligible regardless of ownership. Cleared with `order: null`. The list rendered to end users is `WHERE featured_order IS NOT NULL` + visibility-filter, sorted by `featured_order ASC`. Admins reorder by sending the desired ids to `PUT /api/agents/featured/order` (assigns indices 0..N-1). BIGINT is used rather than INT so the management UI can seed initial values from `Date.now()` (overflows INT on MySQL).
### Publish Squares (????)
Admin-configurable, **access-control-free** categories that published agents/skills are filed under (distinct from the password-gated ????). The square list lives in `system_settings.json` ? `publish_squares` (`PublishSquaresSettings`: ordered `squares: [{id,name,description}]` + `default_square_id`); helpers `list_publish_squares` / `resolve_default_square_id` / `normalize_square_id` always yield a valid square (built-in "????" fallback when none configured). Admin CRUD: `GET/PUT /api/system-settings/publish-squares` (GET readable by any user ? the gallery sub-tabs + publish picker need it; PUT admin-only, auto-slugs missing ids and re-anchors a dangling default).
Each agent/skill carries a single `square_id` column (`AgentRow` / `SkillRow`, `String(64)` NOT NULL default `''`, migration `20260611_01`, idempotent SQL in `column_additions_mysql.sql`). One column = at most one square, so an item can never be "published twice". `''` = unassigned ? resolved to the default square at read/filter time. Writes normalize via `normalize_square_id`: agents on create/update (`square_id` field), skills on publish (`PUT /api/skills/custom/{name}/published` accepts optional `square_id`). Listing filters: `GET /api/agents?square=` and `GET /api/skills?square=` keep only that square's items; the **default** square additionally sweeps in legacy `''` rows (`square_is_default`). Tests: `tests/test_publish_squares.py`. Frontend: `core/system-settings` (`usePublishSquares`/`useUpdatePublishSquares`), admin `PublishSquaresSettingsPage` (embedded in the appearance/settings hub), square picker in the agent edit dialog + per-skill card, and a ??-tab square sub-filter in both galleries.
### Reference-mode Entry Switch
The 参考文献/RAG citation-mode composer entry is system-controlled through
`system_settings.json` → `citation_display.reference_mode_enabled` (default
`false`). `GET /api/system-settings/citation-display` remains readable so the
frontend can hide or show the entry consistently; `PUT` is admin-only and
partially updates `reference_mode_enabled` and/or `cleanup_words`. The normal
workspace input box, iframe chat toolbar, send context, right reference panel,
and appearance settings all use this flag, while already locked citation
threads and selected LLMWiki knowledge bases still render their source
documents in the reference panel. Tests:
`tests/test_prompt_prefix.py` and
`frontend-web/src/pages/iframe-chat-mode-buttons.test.ts`.
### Agent Module Tag Config (zzJC / zzXD 按类型聚合)
Backs the zzJC / zzXD agent-chat left sidebar (「按类型聚合智能体 + 聊天记录」). An admin picks which agent **tags** belong to a module; the module's left sidebar then shows only agents under those tags (empty/unconfigured = show all visible agents). Stored in `system_settings.json` → `agent_modules` (`AgentModuleSettings`: `modules: {module_key → [tag_id, ...]}`); pure helpers `get_module_tag_ids` / `set_module_tag_ids` (de-dupe + strip). Endpoints (in `app/gateway/routers/system_settings.py`): `GET /api/system-settings/agent-modules/{key}` (readable by any user — the sidebar needs it to filter), `PUT /api/system-settings/agent-modules/{key}` (admin only, full replacement of the tag-id list). Module keys are frontend-defined (`AGENT_MODULES` in `components/workspace/agents/agent-module-sidebar.tsx`, keyed by the module's landing agent id: `6bf`→`zzjc`, `10d69…`→`zzxd`). Tests: `tests/test_agent_modules.py`. Frontend: `core/system-settings` (`useAgentModuleConfig`/`useUpdateAgentModuleConfig`), admin `AgentModuleSettingsDialog` (gear button in the agent-chat header).
### Business-mapping authorization
`/api/business-mapping` remains globally readable. Updates to 3Q, 6BF, and 7BF require `system_role == "admin"`; the 8BF public business-chain mapping is intentionally different and can be updated only by the trusted authenticated account whose email prefix is `lqq`. Do not rely on a username supplied in an external login URL. The frontend follows the same email-prefix rule when deciding whether to render the 8BF configuration button, while the router (`app/gateway/routers/business_mapping.py`) remains the enforcement point. Tests: `tests/test_business_mapping.py`.
### Position-collaboration foundation
### Roundtable business-chain enable/disable
`roundtable_chains.enabled` (BOOLEAN NOT NULL DEFAULT 1, migration `20260819_02`)
is the per-chain 启停 switch. New and existing chains default to **enabled**.
Owner or admin can flip it via `PUT /api/roundtable-chains/{id}` `{enabled}`.
Disabled chains stay visible in the management list (with an 启用/停用 filter)
and still show `updated_at`; the frontend hides them from the picker / public
entry rail so they cannot be selected into a roundtable. Tests:
`tests/test_roundtable_chain_stages.py` (`test_create_defaults_enabled_then_toggle`).
### Roundtable business-chain human validation
The existing `/api/roundtable-chains` `seats` JSON also accepts the optional
boolean `human_validation` field on each seat (`ChainSeatSchema`). It is a
frontend Step 2 checkpoint, not a server-side approval task: after the marked
seat completes (or, in a DAG, after its stage completes), the client pauses and
uses the existing resume button or intervention composer to continue. No new
database column, migration, or validation card is required.
The position-collaboration workspace is a separate frontend route and must not
use the original roundtable coordinator or jobs. Its shared configuration
integration point is the existing `/api/roundtable-chains` store: each `seats` JSON item
may carry an optional `position_id` (validated by `ChainSeatSchema`, no schema
migration). The original roundtable only reads its established seat fields and
therefore remains unchanged. The separate `/api/position-roundtable` router
(also mounted at `/api/multi-agent/position-roundtable` for intranet gateway
compatibility)
persists task/intent/frozen-chain snapshots in `position_roundtable_sessions`
and direct-agent node indexes in `position_roundtable_nodes` (migration
`20260721_01`). Message history remains in the established LangGraph thread
store; a node completion unlocks its next chain stage and a later upstream
revision marks downstream records `stale` for refresh.
`POST /api/position-roundtable/sessions/{id}/nodes/{node_key}/reject` uses the
same cascade: its reason is intentionally optional, the selected node becomes
`rejected`, and downstream material becomes `stale` (or `locked` when no
material exists) with `invalidated_by` set. Clients must replace their complete
session snapshot from this response so the flow, delivery list, and active
node state stay consistent. A rejected/stale node returns to `done` only after
a fresh Markdown delivery is recorded.
For direct-agent handoff, the router resolves artifact paths only inside the
completed upstream node's own thread, then includes bounded UTF-8 excerpts for
text outputs; binary artifacts stay as metadata-only references. Because each
seat runs on an isolated LangGraph thread, a downstream `read_file` on the
upstream artifact path would fail with "File not found"; the handoff therefore
also **mirrors** every upstream `outputs/` deliverable into the current seat's
own `workspace/upstream/` directory (`_mirror_upstream_artifacts`, bytes from
the upstream sandbox file with the SQL recovery copy as fallback, idempotent
per turn via size comparison, capped at 8 files / 2 MB each / 8 MB total) and
rewrites each handoff artifact ref's `path` to the mirror (`upstream_path`
keeps the original; `mirrored_to_sandbox: true`), with a matching note in the
hidden instructions telling the agent it may `read_file` the full text. The
mirror deliberately stays outside `outputs/` so it never appears in the seat's
own delivery manifest.
`POST /sessions/{id}/archive` turns the session into a read-only historical
record: task/intent updates, node binding, hidden-context construction, and
turn completion are rejected, while read endpoints remain available.
Built-in summary and action-plan completion also preserves a visible report
answer when the model misses its final `write_file` call: the router writes
that exact answer as the required Markdown artifact in the owning thread's
`outputs/` directory, so users receive both chat text and a downloadable file.
**Per-turn outputs write cap (弱模型防重复产物)**: ordinary position-node chats
and the summary builtin send run-context `position_artifact_cap: 1` (whitelisted
in `services.merge_run_context_overrides`). `PositionArtifactCapMiddleware`
then keeps at most one successful `/mnt/user-data/outputs/` `write_file` **and**
one `present_files` per user turn — truncates excess tool calls in the same
model response (and syncs raw provider `additional_kwargs.tool_calls`), rejects
further writes/presents after an `OK` / `Successfully presented files` or an
earlier sibling in the same batch, ends the turn when only excess calls remain,
and reminds the model to present once (or wrap up if already presented).
**Action-plan does not send the flag** (it must write report `.md` +
`action-plan-subtasks.json`). Tests: `tests/test_position_artifact_cap.py`.
### 岗位 (Position) Subscription System
An admin-managed aggregation layer that decides, per user, which agents / skills
/ scheduled tasks / recommended questions they are subscribed to or can see.
**A user has at most one position** (`users.position_id`, 1:1); unassigned users
inherit the single `is_default` position. Note: the admin-configured prompts
shown to users ARE the **recommended questions** (`recommended_questions` table)
— there is no separate prompt-template store; the per-user `user_custom_prompts`
(personal 常用提示词 bar) is untouched by this system.
**Persistence** (`deerflow/persistence/positions/`, wired as
`app.state.position_store`; tables auto-created by `create_all`, migration
`20260616_01`):
- `positions` (`PositionRow`): `id`, `name` (unique), `description`,
`is_default` (only one at a time — the store clears the flag on others).
- `position_tag_bindings` (`PositionTagBindingRow`): one row per
`(position, resource_type, tag)`. **No binding for a resource_type = "use the
default set".** `POSITION_RESOURCE_TYPES = (agent, skill, scheduled_task,
recommended_question)` — the same set drives the tag types (`TAG_TYPES` was
extended with `recommended_question`).
**Sync engine** (`deerflow/services/position_sync.py`, wired as
`app.state.position_sync`). `sync_user(user)` **materializes** three resource
types into the existing "我的" tables, idempotently:
- agents → `agent_favorites` rows with `origin='position'`
- **skills → `skill_favorites` rows with `origin='position'`** (table
`SkillFavoriteRow`, migration `20260616_04`, auto-created by `create_all`)
- scheduled tasks → `scheduled_task_subscriptions` rows with `origin='position'`
Skills behave **exactly like agents**: a position only *adds* its tagged skills
to the member's 我的技能 — the browse-all 广场 is never restricted. (This replaced
the earlier query-time visibility filter, which wrongly shrank the square to the
bound skill.) `PositionSyncService.resolve_visible(user, type)` returns
`set | None` (None = unrestricted) and is now used **only** for the one remaining
query-time type, `recommended_question`.
Sync only ever adds/removes `origin='position'` rows, so a user's own
favorites/subscriptions are never clobbered. Default sets (no binding for a
type): agents = built-ins (`user_id IS NULL`); skills = empty (square stays
full, nothing added to 我的); scheduled tasks = empty; questions = **untagged
only** (默认岗位 / 无标签绑定 只看到未打标的题目;打标的题目只在绑定了对应标签的
岗位展示 — the router interprets `resolve_visible → None` as "show only untagged",
not "show everything").
`origin`/`position_id` columns were added to `agent_favorites` and
`scheduled_task_subscriptions` — see migration `20260616_01` +
`scripts/sql/column_additions_mysql.sql`. The running app also auto-adds these
via `_ensure_orm_columns_sync` on startup.
**Triggers**: assigning/removing members, changing tag bindings (re-syncs all
members), becoming the default position (`sync_default_members`), and login
(`auth.py::_sync_user_position`, best-effort so a brand-new user inheriting the
default gets their grants materialized).
**API** (`app/gateway/routers/positions.py`, `/api/positions`, **admin only**):
position CRUD; `GET/PUT /{id}/tag-bindings`; `GET/POST/DELETE /{id}/members`
(batch assign/unassign → triggers sync); `POST /{id}/resync`;
`GET /users` (all users + their position); `PUT /users/{user_id}` (set one
user's position). **`GET /me` is the one non-admin route** — it returns the
caller's own effective position (explicit assignment, else the default position
unassigned users inherit, else `null`); powers the top-bar 岗位 label
(`strategy-components/Header.tsx` user cluster, via `useMyPosition`). Related
changes: `GET /api/recommended-questions` is now
filtered by the caller's position (`?all=1` bypasses for admin management) and
admin can tag questions via `/api/tags/assignments/recommended_question/{id}`.
**Default-position rule**: tags are bulk-attached to **all** questions first,
then the visibility filter applies — a position binding recommended_question 标签
sees only those tagged questions; a 默认岗位 / 无绑定 (`resolve_visible → None`)
sees only the **untagged** questions (打标的题目不再泄漏到默认岗位). `?all=1`
(admin management) skips this filter and returns every question with its tags.
The list also accepts a repeatable `tag_ids`
OR-filter — so the 推荐问题配置 page can display per-question 标签 and search by
them (frontend `RecommendedQuestionsPage` + `TagSearchBar`). Tests:
`tests/test_recommended_questions_tags.py`.
`GET /api/skills` no longer restricts by position; each skill carries a
`favorited` flag (true = in the caller's 我的技能, e.g. position-granted) so the
skills page's 我的 tab includes owned ∪ position-granted skills while the 广场
stays complete. Frontend: `core/positions`
+ admin `PositionManagementPage` (侧边栏「综合管理 → 岗位管理」 + 右下角设置菜单).
Recommended questions stay rendered **above** the chat input (their original
position), just filtered by the caller's position. Tests:
`tests/test_positions.py`.
### Portable Column Types (`deerflow/persistence/types.py`)
Two `TypeDecorator`s keep ORM columns correct across the `sqlite` / `postgres` / `mysql` backends. **All ORM timestamp/JSON columns must use these, not the raw SQLAlchemy types.**
- **`PortableJSON`** ? native `JSON` on SQLite/Postgres; `TEXT` (with manual `json.dumps`/`loads`) on MySQL, where older servers lack JSON support.
- **`PortableLongText`** ? a large text column: maps to MySQL `LONGTEXT` (4 GB) and plain `TEXT` on SQLite/Postgres (already unbounded). Use for big JSON/markdown blobs that would overflow MySQL's 64 KB `TEXT` cap ? e.g. `ai_writing_sessions.transcript`.
- **`BeijingDateTime`** ? stores timestamps as Beijing (UTC+8) wall-clock time. The app records time with `datetime.now(UTC)`, but MySQL `DATETIME` / SQLite have no timezone and persist whatever wall clock the driver serializes ? so a raw `DateTime` column would store UTC and read 8 hours behind Beijing. `BeijingDateTime` converts UTC?Beijing on bind and re-attaches `+08:00` on result, so the database holds Beijing time while Python still gets timezone-aware datetimes (instant comparisons stay correct, including in WHERE clauses). PostgreSQL `timestamptz` stores the instant unambiguously, so values pass through unchanged there. Uses a fixed UTC+8 offset (China has no DST), so no `tzdata` dependency.
### Schema Migrations / Manual Column-Add SQL (`scripts/sql/`)
Gateway startup runs Alembic through `app.gateway.app._run_db_migrations` after ORM table creation. The migration graph must always have unique revision ids and exactly one head; `tests/test_migration_graph.py` enforces both. Historical duplicate revision `20260821_01` is repaired by the `20260821_03` / `20260821_04` fixed-question chain and merge revision `20260825_01`. Startup DDL has a bounded timeout and failures are logged without preventing the app from serving.
**Convention (required):** whenever code adds a field to an existing table (a new `mapped_column` on an ORM model), you must, in the same change:
1. Keep the matching alembic version file under `persistence/migrations/versions/` (guarded with `_has_table`/`_has_column`), and
2. Append the idempotent `ADD COLUMN` statement to **`scripts/sql/column_additions_mysql.sql`**, with a definition matching the ORM column exactly (type, length, nullability).
Operators upgrading an existing DB may run the whole file (`USE <db>;` then execute) as a recovery path; it is idempotent (existing columns are skipped via `information_schema.columns`). Index-only changes go in the sibling `scripts/sql/leaderboard_indexes_mysql.sql` (same idempotent pattern) and stay out of startup migrations because online index creation on large metric tables can delay service boot.
Leaderboard scheduled-run exclusion uses correlated `NOT EXISTS` against the
indexed `scheduled_task_runs.agent_run_id`; do not restore `NOT IN`. Legacy
scheduler detection for tool metrics is also a correlated existence check and
must not outer-join the complete `runs` table into metric aggregations. Keep the
created-at + dimension compound indexes in ORM metadata and the idempotent
MySQL index script synchronized.
### Sandbox System (`packages/harness/deerflow/sandbox/`)
**Interface**: Abstract `Sandbox` with `execute_command`, `read_file`, `write_file`, `list_dir`
**Provider Pattern**: `SandboxProvider` with `acquire`, `get`, `release` lifecycle
**Implementations**:
- `LocalSandboxProvider` - Singleton local filesystem execution with path mappings
- `AioSandboxProvider` (`packages/harness/deerflow/community/`) - Docker-based isolation
**Virtual Path System**:
- Agent sees: `/mnt/user-data/{workspace,uploads,outputs}`, `/mnt/skills`
- Physical: `backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/...`, `deer-flow/skills/`
- Translation: `replace_virtual_path()` / `replace_virtual_paths_in_command()`
- Detection: `is_local_sandbox()` checks `sandbox_id == "local"`
- **`{user_id}` for path resolution = `resolve_path_user_id(thread_id)`, using the persisted creator for every tracked thread.** The route's regular owner check still authorizes access first; creator resolution only selects the stable physical directory (`threads_meta.user_id`) and avoids a lost request context falling back to `users/default`. This also preserves cross-account task-deeplink conversations (rwfx/3qfx/xdfx) and roundtable system threads. Positive creator mappings are cached process-wide and DB/memory/missing-row failures fall back to the current login. It is wired at the thread-path consumers including `ThreadDataMiddleware.before_agent`, `present_file_tool`, uploads, artifact serving, and embedded `client.get_artifact`. **Rewrite-history recovery is metadata-only and job-owner scoped** (`list_by_session(..., user_id=current_user)`), so its GET/list endpoints intentionally skip redundant thread-owner checks and return an empty history instead of a history-only 404; the start/commit endpoints retain strict thread write authorization. Tests: `tests/test_thread_paths_owner.py`, `tests/test_thread_paths_roundtable.py`.
- **Durable document-rewrite commit:** `document_rewrite_pipeline.atomic_write_text` writes through a deliberately short sibling temporary filename so long Windows artifact paths stay below the path-length limit, and guarantees the target `outputs` parent exists. If the sandbox recreates that directory in the narrow commit window, it repairs the directory and retries; the original file is otherwise only changed through `os.replace`.
- **Deep Research structural regeneration:** `POST /api/deep-research/sessions/{id}/report-variants` creates a distinct completed session from the source report and its selected evidence. It rebases both saved `[来源:source_id]` and legacy streamed `[[source:source_id]]` markers to cloned ids, and removes inherited markers with no selected clone. The frontend uses this endpoint only for “选择结构再生成”, then queues the durable writer with `operation=generate`; generation receives the confirmed outline and complete selected evidence set but never receives the inherited report as text to edit. The new report's completed card can subsequently start a separate in-place `AI 全文改写`. Requirement plans are deterministic and operation-specific, model reasoning is forwarded as live SSE before the first Markdown token, and generation prompts emit display-ready `[来源:id]` citations. Citation validation repairs only a unique near-complete `drs_src_*` prefix (provider-truncated suffix) and still rejects genuinely unknown ids. The report-only stream requests 8192 output tokens (except Codex Responses models, which reject token-limit kwargs) and rejects a malformed trailing citation instead of committing a provider-truncated report. Never restore the former second planning-model call unless its latency and streaming contract are redesigned together.
**Sandbox Tools** (in `packages/harness/deerflow/sandbox/tools.py`):
- `bash` - Execute commands with path translation and error handling
- `ls` - Directory listing (tree format, max 2 levels)
- `read_file` - Read file contents with optional line range
- `write_file` - Write/append to files, creates directories
- `str_replace` - Substring replacement (single or all occurrences); same-path serialization is scoped to `(sandbox.id, path)` so isolated sandboxes do not contend on identical virtual paths inside one process
### Subagent System (`packages/harness/deerflow/subagents/`)
**Built-in Agents**: `general-purpose` (all tools except `task`) and `bash` (command specialist)
**Execution**: Dual thread pool - `_scheduler_pool` (3 workers) + `_execution_pool` (3 workers)
**Concurrency**: `MAX_CONCURRENT_SUBAGENTS = 3` enforced by `SubagentLimitMiddleware` (truncates excess tool calls in `after_model`), 15-minute timeout
**Flow**: `task()` tool ? `SubagentExecutor` ? background thread ? poll 5s ? SSE events ? result
**Events**: `task_started`, `task_running`, `task_completed`/`task_failed`/`task_timed_out`
### Tool System (`packages/harness/deerflow/tools/`)
`get_available_tools(groups, include_mcp, model_name, subagent_enabled)` assembles:
1. **Config-defined tools** - Enabled entries are resolved from `config.yaml` via `resolve_variable()`; entries with `enabled: false` are omitted before import and are not registered with the agent
2. **MCP tools** - From enabled MCP servers (lazy initialized, cached with mtime invalidation)
3. **Built-in tools**:
- `present_files` - Make output files visible to user (only `/mnt/user-data/outputs`)
- `ask_clarification` - Request clarification (intercepted by ClarificationMiddleware ? interrupts)
- `view_image` - Read image as base64 (added only if model supports vision)
4. **Subagent tool** (if enabled):
- `task` - Delegate to subagent (description, prompt, subagent_type, max_turns)
**Community tools** (`packages/harness/deerflow/community/`):
- `tavily/` - Web search (5 results default) and web fetch (4KB limit)
- `jina_ai/` - Web fetch via Jina reader API with readability extraction
- `firecrawl/` - Web scraping via Firecrawl API
- `configurable_search/` - Config-driven `POST` JSON `web_search` replacement. It deep-copies the configured payload and overwrites only `query`, supports HTTP/HTTPS plus configurable `verify_ssl`, timeout and result limit, normalizes nullable `recUuid`/`content`/`title`/`url` fields, and rewrites a result URL through `result_url_template` only when `recUuid` is non-empty. The required response envelope is `{"results": [...]}`. Tests use `httpx.MockTransport` and never require the deployment endpoint.
**ACP agent tools**:
- `invoke_acp_agent` - Invokes external ACP-compatible agents from `config.yaml`
- ACP launchers must be real ACP adapters. The standard `codex` CLI is not ACP-compatible by itself; configure a wrapper such as `npx -y @zed-industries/codex-acp` or an installed `codex-acp` binary
- Missing ACP executables now return an actionable error message instead of a raw `[Errno 2]`
- Each ACP agent uses a per-thread workspace at `{base_dir}/users/{user_id}/threads/{thread_id}/acp-workspace/`. The workspace is accessible to the lead agent via the virtual path `/mnt/acp-workspace/` (read-only). In docker sandbox mode, the directory is volume-mounted into the container at `/mnt/acp-workspace` (read-only); in local sandbox mode, path translation is handled by `tools.py`
- `image_search/` - Image search via DuckDuckGo
### MCP System (`packages/harness/deerflow/mcp/`)
- Uses `langchain-mcp-adapters` `MultiServerMCPClient` for multi-server management
- **Lazy initialization**: Tools loaded on first use via `get_cached_mcp_tools()`
- **Cache invalidation**: Detects config file changes via mtime comparison
- **Transports**: stdio (command-based), SSE, HTTP
- **OAuth (HTTP/SSE)**: Supports token endpoint flows (`client_credentials`, `refresh_token`) with automatic token refresh + Authorization header injection
- **Runtime updates**: Gateway API saves to extensions_config.json; LangGraph detects via mtime
### Skills System (`packages/harness/deerflow/skills/`)
- **Location**: `deer-flow/skills/{public,custom}/`
- **Format**: Directory with `SKILL.md` (YAML frontmatter: name, description, license, allowed-tools)
- **Loading**: `load_skills()` recursively scans `skills/{public,custom}` for `SKILL.md`, parses metadata, and reads enabled state from extensions_config.json. **A directory containing `SKILL.md` is a self-contained leaf** — discovery does not descend into its subdirectories, so a skill *package* that bundles its own `examples/`/`references/` sub-skills (each with a SKILL.md) surfaces as **one** skill, not a swarm of phantom entries.
- **Discovery↔delete consistency**: a skill is keyed for discovery by its frontmatter `name`, but its on-disk directory name may differ (hand-dropped / mis-extracted package). `LocalSkillStorage._resolve_skill_dir(name, category)` resolves the *discovered* directory (canonical `<category>/<name>` fast-path, else scan-by-frontmatter-name); `custom_skill_exists`/`read_custom_skill`/`delete_custom_skill` all go through it, so **anything that gets listed can also be read and deleted** regardless of directory name (the old code assumed `custom/<name>/` and 404'd "Custom skill 'x' not found." on any non-canonical layout). Delete is path-guarded to never escape the `custom/` root. Tests: `tests/test_skill_storage_delete.py`.
- **Injection**: Enabled skills listed in agent system prompt with container paths
- **Configurable `es_query` routing prompt**: `config.yaml → skills.es_query_routing`
contains an `enabled` switch and editable `prompt`. When enabled,
`lead_agent.prompt._build_es_query_routing_section` injects the rule into the
default lead-agent system prompt. Custom-agent-only prompts receive it only
when their explicit skill allowlist contains `es_query`; disabling the switch
or leaving the prompt blank emits no block. Immediately before an ordinary Q&A
graph is created, `lead_agent.agent._log_es_query_routing_registration` checks
the exact configured block against the final system-prompt string and logs
`enabled`, allowlist eligibility, `registered`, character count, and a short
SHA-256 fingerprint (never the prompt body). Tests:
`tests/test_es_query_routing_prompt.py`.
- **Installation**: `POST /api/skills/install` extracts .skill ZIP archive to custom/ directory. Direct uploads use the ownership-aware `POST /api/skills/validate?target_category=...` preflight plus `POST /api/skills/install-upload`: a same-owner custom conflict may be uploaded again with `overwrite=true`, while a different owner's conflict is non-overwritable and returns that owner's username (falling back to user id). The install endpoint repeats the ownership check, including for admins targeting `custom`, so callers cannot bypass the preflight. Public-skill overwrite remains admin-only.
- **Installable sentiment analysis package**: `skill-packages/sentiment-analysis.skill` is deliberately separate from the frontend virtual-agent chat page. The chat page goes through Gateway `POST /api/sentiment-agent/stream` (runtime-config URL/Basic Auth in the body, TLS verify off, raw AG-UI SSE passthrough). The skill package's self-contained `config/sentiment-agent.json` owns the external AG-UI address for agent-skill callers; Basic Auth is resolved from `SENTIMENT_AGENT_AUTH_USERNAME` / `SENTIMENT_AGENT_AUTH_PASSWORD` at runtime (not committed into the package). `scripts/sentiment_stream.py` uses only the Python standard library and converts upstream SSE to stable NDJSON events (`thinking_delta`, `answer_delta`, `done`), with a Markdown convenience mode for normal agent calls. Installation/calling guide: `skill-packages/sentiment-analysis/docs/调用与安装说明.md`.
- **Chinese name (`name_zh`)**: a Chinese display name stored in the **`skills` DB table** (`SkillRow.name_zh`, `String(128)`, migration `20260609_01`) ? never written to the on-disk SKILL.md, so it works for built-in skills too. Threaded through `SkillStore` create/ensure_legacy/update/update_admin exactly like `detail`; surfaced on `SkillResponse.name_zh` from the ownership record in `_skill_to_response`. Authorization mirrors `detail`: custom skill ? owner or admin; built-in/public ? admin only (row materialized on demand via `ensure_legacy`). Endpoints (in `app/gateway/routers/skills.py`): `PUT /api/skills/{name}/name-zh` (set/clear), `POST /api/skills/{name}/name-zh/generate` (LLM generates one), `POST /api/skills/name-zh/generate-missing` (????: **admin-only** bulk-generate for every skill that lacks one, bounded-concurrency `asyncio.gather`). **Optional model selection**: both generate endpoints accept an optional JSON body `{model_name}` ? an explicit name wins, else the curator model setting, else the main model (`_resolve_skill_name_zh_model_name`). **Model-error surfacing**: `_generate_skill_name_zh` raises `SkillNameZhModelError` when the model itself fails (build/timeout/API error) ? distinct from the model returning an unusable name (? `None`). The single endpoint maps that to `502 ?????,?????????:?`; the bulk endpoint fails fast on a model-build error (502) and otherwise carries a representative `model_error` string in `SkillNameZhBulkResponse` so the UI can say "N ?????????". The frontend (`skill-settings-page.tsx` list + `skill-detail-dialog.tsx` editor, plus the agent skill pickers/badges in `NewAgentPage.tsx` + `agent-card.tsx`) shows `name_zh || name`, offers per-skill "AI ??" plus an **admin-only** header "???????" button, and a compact model dropdown (`components/workspace/model-name-select.tsx`, default = configured model) beside each action.
#### Skill Address Management (技能地址管理 — URL/IP 批量替换)
Admin-only maintenance tool to find every `http(s)://` URL and IPv4 (`ip[:port]`)
address in the **running** skills (`skills/{public,custom}`, all `.md` + `.py`,
recursing into `scripts/`/`references/`/`templates/`) and rewrite them in bulk —
e.g. swapping a UAT endpoint for production across all skills at once. The same
address commonly lives in both a skill's `SKILL.md` and its `scripts/*.py`, so
results are **aggregated by unique address** with every occurrence (`file:line` +
snippet) and replaced together.
- **Pure engine** — `deerflow/skills/address_scan.py` (harness, no `app.*`):
`find_spans` / `scan_occurrences` (URL regex with trailing-punctuation strip;
IPv4 with 0–255 octet check + optional port; an IP inside a URL is *not*
separately surfaced), `apply_replacements` (**substring-safe**: re-detects exact
token spans and only rewrites a span whose text equals a requested `from`, so
`1.2.3.4` is never rewritten inside `11.2.3.40`), `infer_kind`,
`is_valid_replacement`. Tested in `tests/test_skill_addresses.py`.
- **Router** — `app/gateway/routers/skill_addresses.py` (`/api/skill-addresses`,
**admin only**). **Availability at scale (150+ skills):** every filesystem
walk/read/write runs via `asyncio.to_thread` so the event loop is never blocked.
Routes: `GET /scan` (aggregated, one-shot) + `GET /scan/stream` (per-skill NDJSON
progress `{type:start|progress|result}`); `POST /preview` (dry-run diff +
validation warnings); `POST /apply` + `POST /apply/stream` (per-file NDJSON
progress); `GET /backups`; `POST /rollback/{backup_id}`. **Apply safety:**
optimistic lock via `expected_total` (409 if files changed since the scan),
path-guard (`is_relative_to` the skills root), auto disk backup to
`{skills_root}/.address-edits/{timestamp}/` (the `.`-prefixed dir is excluded
from scans), then `refresh_skills_system_prompt_cache_async()` so edits take
effect immediately. Registered in `app.py` after `sensitive_words.router`.
**Per-edit history:** each apply also writes a sibling manifest
`{skills_root}/.address-edits/{timestamp}.manifest.json` (best-effort) capturing
`{created_at, replaced_total, replacements:[{from,to,kind}], changed_files}`; the
sibling-not-child placement keeps it out of both the restore set and the file
count. `GET /backups` returns each record enriched with `replaced_total` +
`replacements` (from the manifest; falls back to dir mtime when absent), and
`POST /rollback/{backup_id}` restores a single record — so the UI shows a full
modification history and rolls back any one entry independently.
- **Frontend** — the panel is the reusable `SkillAddressesPanel` (exported from
`pages/SkillAddressesPage.tsx`; its default export keeps the standalone route
`/page/strategy/admin/skill-addresses` as a direct-link fallback). It is
**embedded as a「技能地址管理」tab inside the admin「技能管理」page**
`pages/AdminSkillsPage.tsx` (Tabs: 技能列表 + 技能地址管理; reached from the
「技能管理」item in the「设置和更多」dropdown `workspace-nav-menu.tsx` →
`/page/workspace/admin/skills`) — it is no longer a separate entry in the
`components/page-sidebar.tsx` settings dropdown. The panel:
auto-scan on open with a **progress bar** (扫描逐技能 / 应用逐文件 X/N), URL/IP
type filter, search, a「片段批量套用」helper (e.g. all `uat.4.cn` → `prod.4.cn`),
per-address「替换为」input with expandable occurrences, inline 预览 panel, 应用,
**multi-keyword search** (space-separated, all-must-hit), and a **修改记录**
(modification history) panel — each apply is one record (time / which addresses
changed from→to / file & replacement counts) with a per-record 回退此次
(single-record rollback, with inline confirm). Client + NDJSON stream parsing in
`strategy-components/api/skill-addresses.ts`. Scope is the running skills only;
the `deploy/minimal/skills/` copy is not touched (sync separately).
#### Skill Compression (技能压缩)
When the skill catalogue grows large, injecting every enabled skill full-text
into the lead agent's system prompt bloats it and wastes tokens. **Skill
compression** is an admin-gated mode that keeps only *常驻(免压缩)* skills in the
prompt body and surfaces the rest on demand via keyword retrieval.
- **Master switch** — `system_settings.json → skill_compression`
(`SkillCompressionSettings`: `enabled` default **false**, `search_top_n`
default 5, `keep_index_in_prompt` default true). Read/written by
`GET`/`PUT /api/system-settings/skill-compression` (GET open to any user — the
skill detail page reads it; PUT admin-only, refreshes the prompt cache).
- **Per-skill 常驻 flag** — stored in the **DB**: `SkillRow.always_on`
(`Boolean`, migration `20260616_03`, idempotent SQL in
`column_additions_mysql.sql`), **not** in `extensions_config.json` (the
source-tree config file is read-only / unsafe to rewrite in some deployments,
and built-in skills carry no file entry). Threaded through `SkillStore`
create/ensure_legacy/update/update_admin + `list_always_on_names()` exactly
like `name_zh`/`detail`; surfaced on `SkillResponse.always_on` from the
ownership record in `_skill_to_response`. Toggled via
`PUT /api/skills/{name}/always-on` (custom skill → owner or admin; built-in /
public → admin only, materializes a row on demand via `ensure_legacy`). The
prompt builder stamps `Skill.always_on` from `list_always_on_names()` via a
sync DB bridge (`_apply_always_on` in `lead_agent/prompt.py`).
- **Prompt split** — `get_skills_prompt_section` (`lead_agent/prompt.py`) reads
the compression settings (skipped for restricted custom agents — they already
have a curated allowlist). When **off**, behaviour is unchanged (all enabled
skills full-text). When **on**, `always_on` skills render as full `<skill>`
blocks; the rest are dropped from the body, replaced by a `search_skills`
usage note plus (when `keep_index_in_prompt`) a name-only `<compressed_skills>`
index. The `_get_cached_skills_prompt_section` cache key includes the
`always_on` flag + compression flags so changes invalidate correctly.
- **Retrieval tool** — `search_skills` (`tools/builtins/skill_tools.py`,
`search_skills_impl`) keyword-scores `name + description` over enabled,
visible, non-archived skills (same visibility filter as `skill_list`) and
returns top-N `{name, description, category, location}`. It is registered in
`get_available_tools` **only when compression is enabled** (offline,
deterministic, zero external deps). The agent then `read_file`s the returned
`location`.
Tests: `tests/test_skill_compression.py`.
#### Skill Curator (`packages/harness/deerflow/skills/curator.py`)
Background skill lifecycle manager. The Gateway `lifespan` starts a resident
scheduler task (`_curator_scheduler_loop` in `app/gateway/app.py`) that ? after a
5-minute startup delay ? calls `maybe_run_curator()` every hour. A cross-process
file lock (`.curator_scheduler.lock`) elects a single leader worker so the
scheduler runs once per deployment. Per run:
- **Master switch** ? `skill_evolution.enabled` (default **`false`**) gates the
whole feature; `curator.enabled` (default **`false`**) gates the curator
specifically. Both must be `true` for the curator to act. `should_run_now()`
and `_curator_thread_target()` both enforce this ? so manual `POST
/api/curator/run` also respects the master switch. Config is
`config.yaml.skill_evolution` deep-merged with per-user
`users/{user_id}/skill_evolution.json`.
- **Per-user run lock** ? `.curator_state` carries `running` / `started_at`;
`_try_acquire_run_lock()` prevents the same user's curator from running
concurrently (scheduler vs. manual trigger, or double manual clicks). A lock
whose `started_at` is older than `_CURATOR_STALE_LOCK_HOURS` (6h) is treated as
process-crash residue and can be taken over.
- **Concurrency cap** ? `maybe_run_curator()` collects all due users and hands
them to `_dispatch_curator_runs()`, which runs them through a
`ThreadPoolExecutor` capped at `_CURATOR_MAX_CONCURRENCY` (5) ? avoids a
thundering herd of LLM calls when many users come due at once (e.g. first
deploy).
- **Failure back-off** ? `_release_run_lock()` always stamps `last_attempt_at`.
When a first run keeps crashing (`last_run_at` never set), `should_run_now()`
retries on `_CURATOR_RETRY_HOURS` (6h) instead of every hourly tick.
- **`should_run_now()`** is the real gate: the hourly tick only checks "is the
per-user `interval_hours` (default 168h) elapsed". Timestamp parsing tolerates
naive (tz-less) strings ? a corrupt `last_run_at` triggers a re-run rather than
a permanent deadlock.
- **Published-skill filter** ? `SkillStore.list_published_names()` is a batch
query; the curator fetches the published-name set once per run instead of one
DB round-trip per candidate skill.
- **LLM review timeout** ? `_run_llm_review()` runs `agent.invoke()` in a worker
thread bounded by `curator.review_timeout_seconds` (default 600s). The model's
per-request `request_timeout` can't bound a multi-turn React agent loop, so on
timeout the run is abandoned (a warning is written to `curator_log`, visible in
the Web UI) and the curator finishes cleanly + releases the run lock. The
worker context is copied via `contextvars.copy_context()` so tool calls still
resolve the correct `user_id`.
- **Review model** ? `_run_llm_review()` uses `curator.model_name` (a
`config.yaml` `models[]` entry name) if set, else the primary model; it is
editable from the Web UI evolution panel. The skill security scanner
(`scan_skill_content` in `security_scanner.py`) reuses the same
`skill_evolution.curator.model_name` setting ? review and moderation share one
model. There is no separate moderation-model config field.
### Model Factory (`packages/harness/deerflow/models/factory.py`)
- `create_chat_model(name, thinking_enabled)` instantiates LLM from config via reflection
- Supports `thinking_enabled` flag with per-model `when_thinking_enabled` overrides
- `force_disable_thinking=True` (kwarg) hard-disables thinking for models that explicitly declare `supports_thinking: true` even when the model config declares no `when_thinking_*`/`thinking` shape — many internal vLLM/Qwen models only set that flag, so `thinking_enabled=False` alone injects nothing and they keep reasoning on. The flag injects the OpenAI-compatible disable shapes (`extra_body.enable_thinking=false`, `extra_body.chat_template_kwargs.enable_thinking=false`, and `extra_body.thinking.type="disabled"`) only for classes that accept `extra_body`. Models with `supports_thinking: false`, especially generic OpenAI-compatible proxies such as LiteLLM, receive no vendor-specific thinking fields because those gateways may reject unknown parameters. Mirrors `TitleMiddleware._force_disable_thinking`. The lead agent reads the runtime flag `thinking_force_disabled` and forwards it (used by `/api/open/chat`)
- Supports vLLM-style thinking toggles via `when_thinking_enabled.extra_body.chat_template_kwargs.enable_thinking` for Qwen reasoning models, while normalizing legacy `thinking` configs for backward compatibility
- Supports `supports_vision` flag for image understanding models
- Config values starting with `$` resolved as environment variables
- Missing provider modules surface actionable install hints from reflection resolvers (for example `uv add langchain-google-genai`)
### Roundtable Model Fault Tolerance (会商模型自动容错 / 换下一个模型)
会商第二步(leader/seat)+ 第三步(summary/report/dashboard + 结构化抽取)单轮执行时,**某模型出问题就自动切到下一个模型继续**,并告知用户「『X』模型异常 → 已切换『Y』」。前台实时(`multi_agent.py`)与后台挂起(`engine.py` + `InProcessRoundtableGateway`)两条路径都覆盖。
- **可检测的模型失败**:`LLMErrorHandlingMiddleware` 在同一模型重试 3 次后**不抛异常**,而是返回兜底 `AIMessage` —— 现在给它打 `additional_kwargs[LLM_ERROR_MARKER="deerflow_llm_error"]=reason`(model_dump 保留,经 checkpointer 序列化仍在)。`is_llm_error_message(msg)` 据标记(回退按已知兜底文案正文)识别,返回 reason。
- **候选模型链**:`app/gateway/roundtable_model_fallback.py::fallback_model_chain(current)` 按 `config.yaml models[]` 顺序给出 `[当前模型, …其余]`(去重保序)。`reason_text(reason)` 翻译成简短中文。总开关 env `ROUNDTABLE_MODEL_FALLBACK`(默认开,`0/false/off` 关 → 单模型旧行为);`ROUNDTABLE_MODEL_FALLBACK_MAX`(0=试完全部)。
- **触发换模型的情形**:① 单轮 run 抛异常(start_run 失败 / idle 卡死 → `timeout`,其余 `exception`);② 兜底错误文案(`is_llm_error_message`);③ 交付无效(仅 seat/leader;leader 还含「既无正文又无派活」)。席位**只认剥掉 `<think>` 后的可见正文**——思考-only / 空串 / 断流都不算已交付(`deerflow.agents.roundtable_orchestrator.delivery`:`visible_seat_text` / `classify_seat_delivery`)。前台 `_special_run` 对此 yield `status:"invalid_delivery"` 且不广播;`GET /last-ai-message` 回 `{delivered, content, length, invalid_reason}`(思考-only 标 `thinking_only`,总控席位清单标 ⚠️)。后台引擎只把可见正文写入 `deliveries`,续轮用 `build_continue_prompt`:有无效席位时强制重派,禁止声称「已完成交付」。Step3 单例(report/summary/dashboard/ingest/结构化)的产物是落盘文件,**不**把「正文为空」当失败,只对真模型错误换模型。
- **后台**:`InProcessRoundtableGateway._run_turn_with_fallback(...)` 包住 `_run_agent_turn`,逐候选模型试;换模型落 `model_switch` 诊断、全失败落 `model_fallback_exhausted`。`run_leader/run_seat/run_report/run_summary/run_dashboard_data` 全走它(report/summary/dashboard 的产物自愈重试保留,并记住已切到的可用模型用于后续续写轮)。各 Turn/Result 带回 `switches: list[ModelSwitch]`,引擎 `_emit_switches` → `JobProgressWriter.model_switch(...)` 往研讨时间线插一条 `role="system"` 系统提示气泡。
- **前台**:`_special_run`/`_leader_run` stream 收尾后检测失败 → yield `{status:"model_switch", role, agent_name, failed_model, next_model, reason}` 帧并用下一个模型重新流式;终态帧带 `model_used`。结构化抽取 `flow_extract.py` 的 round1 `model.astream` 也包了同款换模型循环(emit `event:"model_switch"` 帧)。
- **Step2 单席位隔离**:`engine._run_recommend/_run_chain/_run_dag` 的每个席位执行包 try/except —— 一个席位彻底失败(落 `seat_failed` 诊断 + 置空交付)不拖垮整批/整链/整个作业。
- **前端**:`multi-agent.ts` 解析 `model_switch` 帧 → `onModelSwitch` 回调(清空失败尝试的半截气泡 + 累计的 tool_call delta,phase 退回 connected);`useStep2Orchestration` 弹 toast + 时间线插「系统提示」气泡;`useStep3Summary/Report/Dashboard` + `flow-extract.ts` 重置对应气泡并在步骤卡留可见提示;后台路径经 `job-live.ts::toStep2Dialogue` 把 `role="system"` 渲染成 amber「系统提示」气泡。
- Tests: `tests/test_roundtable_model_fallback.py`。
### vLLM Provider (`packages/harness/deerflow/models/vllm_provider.py`)
- `VllmChatModel` subclasses `langchain_openai:ChatOpenAI` for vLLM 0.19.0 OpenAI-compatible endpoints
- Preserves vLLM's non-standard assistant `reasoning` field on full responses, streaming deltas, and follow-up tool-call turns
- Designed for configs that enable thinking through `extra_body.chat_template_kwargs.enable_thinking` on vLLM 0.19.0 Qwen reasoning models, while accepting the older `thinking` alias
### IM Channels System (`app/channels/`)
Bridges external messaging platforms (Feishu, Slack, Telegram, DingTalk) to the DeerFlow agent via the LangGraph Server.
**Architecture**: Channels communicate with Gateway through the `langgraph-sdk` HTTP client (same as the frontend), ensuring threads are created and managed server-side. The internal SDK client injects process-local internal auth plus a matching CSRF cookie/header pair so Gateway accepts state-changing thread/run requests from channel workers without relying on browser session cookies.
**Components**:
- `message_bus.py` - Async pub/sub hub (`InboundMessage` ? queue ? dispatcher; `OutboundMessage` ? callbacks ? channels)
- `store.py` - JSON-file persistence mapping `channel_name:chat_id[:topic_id]` ? `thread_id` (keys are `channel:chat` for root conversations and `channel:chat:topic` for threaded conversations)
- `manager.py` - Core dispatcher: creates threads via `client.threads.create()`, routes commands, keeps Slack/Telegram on `client.runs.wait()`, and uses `client.runs.stream(["messages-tuple", "values"])` for Feishu incremental outbound updates
- `base.py` - Abstract `Channel` base class (start/stop/send lifecycle)
- `service.py` - Manages lifecycle of all configured channels from `config.yaml`
- `slack.py` / `feishu.py` / `telegram.py` / `dingtalk.py` - Platform-specific implementations (`feishu.py` tracks the running card `message_id` in memory and patches the same card in place; `dingtalk.py` optionally uses AI Card streaming for in-place updates when `card_template_id` is configured)
**Message Flow**:
1. External platform -> Channel impl -> `MessageBus.publish_inbound()`
2. `ChannelManager._dispatch_loop()` consumes from queue
3. For chat: look up/create thread through Gateway's LangGraph-compatible API
4. Feishu chat: `runs.stream()` ? accumulate AI text ? publish multiple outbound updates (`is_final=False`) ? publish final outbound (`is_final=True`)
5. Slack/Telegram chat: `runs.wait()` ? extract final response ? publish outbound
6. Feishu channel sends one running reply card up front, then patches the same card for each outbound update (card JSON sets `config.update_multi=true` for Feishu's patch API requirement)
7. DingTalk AI Card mode (when `card_template_id` configured): `runs.stream()` ? create card with initial text ? stream updates via `PUT /v1.0/card/streaming` ? finalize on `is_final=True`. Falls back to `sampleMarkdown` if card creation or streaming fails
8. For commands (`/new`, `/status`, `/models`, `/memory`, `/help`): handle locally or query Gateway API
9. Outbound ? channel callbacks ? platform reply
**Configuration** (`config.yaml` -> `channels`):
- `langgraph_url` - LangGraph-compatible Gateway API base URL (default: `http://localhost:8001/api`)
- `gateway_url` - Gateway API URL for auxiliary commands (default: `http://localhost:8001`)
- In Docker Compose, IM channels run inside the `gateway` container, so `localhost` points back to that container. Use `http://gateway:8001/api` for `langgraph_url` and `http://gateway:8001` for `gateway_url`, or set `DEER_FLOW_CHANNELS_LANGGRAPH_URL` / `DEER_FLOW_CHANNELS_GATEWAY_URL`.
- Per-channel configs: `feishu` (app_id, app_secret), `slack` (bot_token, app_token), `telegram` (bot_token), `dingtalk` (client_id, client_secret, optional `card_template_id` for AI Card streaming)
### Memory System ? Memory V2 (`packages/harness/deerflow/agents/memory/`)
Pluggable memory architecture (Hermes-style). See `docs/MEMORY_V2_DESIGN_ZH.md`
and `docs/HINDSIGHT_SETUP_ZH.md` for the full design.
**Components**:
- `provider.py` - `MemoryProvider` abstract base (12 lifecycle/hook methods)
- `manager.py` - `MemoryManager` orchestrates builtin + ?1 external provider; `build_memory_manager()` is the construction entry point
- `providers/builtin.py` - `BuiltinFileProvider`: local USER.md / MEMORY.md, ? delimiter, char quotas, file locking, atomic writes, frozen snapshot
- `providers/hindsight.py` - `HindsightProvider`: HTTP client to a self-hosted Hindsight Docker instance, dual-bank routing
- `security.py` - injection/exfiltration scan + invisible-unicode block for memory content
- `message_processing.py` - `filter_messages_for_memory` (last-turn extraction for Hindsight retain)
**Two stores (binary classification)**:
- `USER.md` at `{base_dir}/users/{user_id}/USER.md` ? user profile, shared across agents
- `MEMORY.md` at `{base_dir}/users/{user_id}/agents/{agent_id}/MEMORY.md` ? agent-private working notes
- `user_id` resolved via `get_effective_user_id()`; `"default"` in no-auth mode
**Write paths**:
- LLM actively calls the `memory` tool (`tools/builtins/memory_tool.py`) ? writes USER.md/MEMORY.md, mirrors to Hindsight
- LLM calls `hindsight_recall`/`reflect`/`retain` tools (`tools/builtins/hindsight_tools.py`) ? only in `tools`/`hybrid` memory_mode
- Hindsight `auto_retain` ? every turn auto-stored to the work bank (via `MemoryMiddleware.after_agent`)
**Injection**:
- builtin frozen snapshot ? injected into the system prompt via `lead_agent/prompt.py` `_get_memory_context`
- Hindsight dynamic recall ? injected per model call via `MemoryMiddleware.wrap_model_call` (transient ? not persisted to state, so it never leaks to the frontend)
**Hindsight dual-bank isolation** (B-scheme):
- global user bank `deerflow-user-{user_id}` ? cross-agent
- agent work bank `deerflow-work-{user_id}-{agent_id}` ? per-agent
- `hindsight-client` is an optional dependency; when missing, `HindsightProvider.is_available()` returns False and the system falls back to builtin-only
**Migration**: `PYTHONPATH=. python scripts/migrate_memory_to_v2.py [--all|--user-id X] [--dry-run]` converts legacy `memory.json` into USER.md/MEMORY.md.
**Configuration** (`config.yaml` ? `memory`, nested): `enabled`, `provider` (builtin/hindsight),
`builtin.{memory_char_limit,user_char_limit,...}`, `hindsight.{mode,api_url,api_key,bank_id_template,work_bank_id_template,memory_mode,...}`,
`injection.{enabled,max_tokens,include_builtin,include_hindsight,context_tag}`, `security.{scan_content,block_invisible_unicode}`.
Per-user runtime overrides (whitelisted fields) live at `{base_dir}/users/{user_id}/memory_config.json` ? see `get_effective_memory_config()`.
**Gateway API** ? see the Memory router below (`/api/memory/v2*`).
### Reflection System (`packages/harness/deerflow/reflection/`)
- `resolve_variable(path)` - Import module and return variable (e.g., `module.path:variable_name`)
- `resolve_class(path, base_class)` - Import and validate class against base class
### Config Schema
**`config.yaml`** key sections:
- `models[]` - LLM configs with `use` class path, `supports_thinking`, `supports_vision`, provider-specific fields
- vLLM reasoning models should use `deerflow.models.vllm_provider:VllmChatModel`; for Qwen-style parsers prefer `when_thinking_enabled.extra_body.chat_template_kwargs.enable_thinking`, and DeerFlow will also normalize the older `thinking` alias
- `tools[]` - Tool configs with `use` variable path and `group`
- `tool_groups[]` - Logical groupings for tools
- `sandbox.use` - Sandbox provider class path
- `skills.path` / `skills.container_path` - Host and container paths to skills directory
- `title` - Auto-title generation (enabled, max_words, max_chars, prompt_template)
- `summarization` - Context summarization (enabled, trigger conditions, keep policy)
- `subagents.enabled` - Master switch for subagent delegation
- `memory` - Memory system (enabled, storage_path, debounce_seconds, model_name, max_facts, fact_confidence_threshold, injection_enabled, max_injection_tokens)
- `auth_login` - Custom login modes layered on the base auth (`deerflow/config/login_config.py`):
- `username_login.{require_password, password}` - **Shared** password gate for `POST /api/v1/auth/login/username`. When `require_password=true`, the (otherwise passwordless) username login additionally requires `?password=<shared-secret>`; empty `password` falls back to `123ewq`. Gate only ? it does not change which account logs in. **Per-account override (账号级登录口令)**: independent of this shared gate, an admin can assign **one specific account** an individual **plaintext** password via the 白名单管理 page (`users.login_password` column, migration `20260625_04`, nullable). When the target account carries a `login_password`, `/login/username` **requires** `?password=` to match it **exactly** and that personal password **takes precedence over** the shared gate (the shared secret will not unlock that account); accounts with no personal password keep the shared-gate / passwordless behavior. The check happens before auto-register via the lookup-only helper `_lookup_user_by_username` (so a brand-new username is unaffected). Surfaced/managed admin-only: `GET /api/v1/auth/users` returns `UserListItem.login_password` (plaintext — the list is already admin-gated), and `PUT /api/v1/auth/users/{user_id}/login-password` `{password: str|null}` sets (non-empty) / clears (null·empty) it without touching `token_version`. Plaintext storage is intentional (the password travels in the login URL query anyway, so it stays re-copyable). The user logs in via the existing `#/login/<username>?password=xxxx`. Tests: `tests/test_login_personal_password.py`.
- `token_login.{enabled, token_info_url, service_authorization, timeout_seconds, verify_ssl, ca_cert_path, mock_tokens}` - Token-exchange login (`POST /api/v1/auth/login/token`). The gateway POSTs to `token_info_url` with the caller-supplied token placed directly in the `Authorization: Bearer <token>` header (no cookie; the static `service_authorization` field is deprecated/unused, kept only for backward compat), reads `data.username` from a `{"state":"200",...}` body, then signs the matching local user in (auto-registering if absent). `token_info_url` is HTTPS ? `verify_ssl: false` skips cert verification for self-signed/internal UAT hosts, or set `ca_cert_path` to a PEM bundle (takes precedence, keeps verification on). **`mock_tokens` (offline test mapping `{token: username}`)**: when an incoming token matches a key, `_fetch_username_from_token` returns the mapped username and **skips the `token_info_url` call entirely** — for environments where the upstream `getTokenInfo` host is unreachable (e.g. external networks). Empty in prod (no effect); seeded `{"123ewq": "admin"}` in this deployment so `/login/任意?authToken=123ewq` signs in as admin. **`UsernameLoginResponse` now also returns `username`** (the raw resolved login name, pre email-normalization) on both `/login/token` and `/login/username`, so the frontend can append a faithful `?username=` on jump links (按钮管理 追加登录态) instead of the lossy email local-part. See [docs/AUTH_LOGIN.md](docs/AUTH_LOGIN.md) for the full request/response contract.
Both endpoints are whitelisted in `auth_middleware._PUBLIC_EXACT_PATHS`. The shared `_get_or_create_user_by_username()` helper (in `routers/auth.py`) backs both `/login/username` and `/login/token`.
- `log_system` - Admin-only external log-system entry (`deerflow/config/log_system_config.py`): `{enabled, url, label}`. Surfaced read-only by `GET /api/system-settings/log-system` (admin only); the frontend renders it as a "????" item in the bottom-right settings menu (`workspace-nav-menu.tsx`) that opens `url` in a new tab. Disabled by default.
**`extensions_config.json`**:
- `mcpServers` - Map of server name ? config (enabled, type, command, args, env, url, headers, oauth, description)
- `skills` - Map of skill name ? state (enabled). 常驻(免压缩)/always_on lives in the `skills` DB table, not here.
Both can be modified at runtime via Gateway API endpoints or `DeerFlowClient` methods.
### Conditional WeKnora LLMWiki
Ordinary knowledge-base cards expose an owner/admin-only name/description edit dialog in `WeKnoraKnowledgeWorkspace.tsx`, reusing `useLlmWikiActions().update` and `PATCH /api/llmwiki/knowledge-bases/{id}`. The route trims names, rejects blank or reserved conversation-deposit names, writes to WeKnora before updating local metadata, and retains existing write authorization. A successful remote list is authoritative for every list scope, including `personal`: absent remote IDs are hidden without deleting mappings/references. Upstream failures remain errors. The grid also filters cached `remote_status=missing` rows and no longer renders stale-record cleanup cards. Regression coverage: `tests/test_llmwiki_knowledge_management.py`.
`llmwiki.weknora.api_base_url` is the provider switch: non-empty enables WeKnora, empty preserves the legacy LLMWiki UI. WeKnora is independently deployed and owns PostgreSQL/Redis/object storage/vector storage/parsers/models; DeerFlow config contains only `api_base_url` and optional `web_base_url`. DeerFlow deliberately sends no WeKnora credential, so the WeKnora endpoint must be restricted to a trusted network and allow DeerFlow traffic itself or through its reverse proxy. There is intentionally no `default_embedding_model_id`: `WeKnoraClient` queries `/api/v1/models` and uses WeKnora's active Embedding model when creating a KB; Wiki creation additionally resolves an active KnowledgeQA model inside WeKnora for synthesis. The native BFF is `app/gateway/routers/llmwiki.py`; ownership/publication mappings live in `deerflow.persistence.llmwiki`; the browser only sees DeerFlow mapping IDs. Knowledge-base owners publish immediately (`private` -> `published`) without a review queue, and can unpublish ordinary bases. The conversation-deposit KB is system-owned, always published, shared by all DeerFlow users, excluded from My Knowledge, and immutable through user/admin detail operations; recognition requires both the reserved name and `system` owner. Existing same-name remote mappings are adopted in place as the system deposit, preserving mapping ids and document links. Its user-facing description, generated Markdown source and tags use `cmzs` branding; list/detail reads synchronize stale remote descriptions, and search results, source previews and embedded-detail JSON recursively replace legacy brand text in historical conversation-deposit payloads. Deposit requests must reference a thread accessible to the current account. Personal knowledge-base list/search views are scoped to the current DeerFlow account even for admins; frontend LLMWiki caches include the live account id, and admin write power does not make other users' private KBs appear as owned. Only `system_role === "admin"` sees the WeKnora connection-config tab. Knowledge cards navigate to a dedicated WeKnora-style detail route. Wiki-enabled bases render and default to `Wiki -> documents -> graph`; standard bases omit Wiki and render `documents -> graph`. The detail page supports sequential multi-file upload with an aggregate result for writable ordinary bases, hides it for public read-only/system-deposit bases, and uses a generic loading label while the embed session starts. It also supports URL and published manual-Markdown imports, browses real Wiki pages/content, proxies authenticated original-file previews, pages through `/chunks/{knowledge_id}`, and renders the Wiki page graph from `/knowledgebase/{kb_id}/wiki/graph`; document details use a right sheet rather than the removed KB modal. The credential-free proxy uses one WeKnora service identity, so per-user separation is enforced by DeerFlow mappings while all remote KBs physically share that identity's WeKnora workspace. WeKnora iframe detail pages first call `/api/llmwiki/knowledge-bases/{id}/weknora/session` through the normal DeerFlow API auth path, then load `/platform/knowledge-bases/{weknora_id}` with a signed `deerflow_weknora_embed` cookie for `/platform`, `/assets`, `/locales`, root assets, and API calls. Injected fetch/XHR rewriting sends all iframe `/api/v1/*`, including colliding `/api/v1/auth/*`, through `/api/llmwiki/weknora-embed/api/v1/*`; Gateway auth/CSRF bypass only applies to those signed iframe proxy requests, the proxy still revalidates the mapped KB, and cookie security honors `X-Forwarded-Proto` for HTTPS reverse proxies. `llmwiki_search` is conditionally registered in `tools.py`, is read-only, and rechecks the current DeerFlow user, agent `llmwiki_knowledge_base_ids`, and thread-context selection on every call. Source previews prefer the cleaned Wiki/LLMWiki page text when a retrieved chunk maps to a Wiki page, with raw split chunks as the fallback. Agent bindings use an inline WeKnora-style searchable/grouped multi-select in the standard agent create/detail UI; new chats use the same interaction pattern as a compact composer dropdown, and selections are written to thread metadata/context and restored on reopen. The DeerFlow frontend is native React and the WeKnora frontend is not modified or embedded. Full deployment and behavior notes: `docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md`. Tests: `tests/test_llmwiki_weknora.py`.
The embedded detail proxy hides WeKnora's outer `.main > .aside_box` application sidebar and expands the route outlet, but deliberately keeps the Wiki page's internal directory/index sidebar. Do not add an HTML `sandbox` attribute to this trusted, permission-filtered detail iframe: a sandboxed ancestor causes Chromium's built-in Blob PDF viewer to render the browser-blocked page. The Gateway follows upstream redirects only for `/knowledge/{id}/preview` and `/knowledge-bases/{id}/files`, keeping PDF bytes and protected Wiki images inside the signed proxy instead of sending the browser to WeKnora object storage.
Native/private and public Wiki readers must rewrite protected Markdown image
sources to their mapping-scoped DeerFlow `/files?file_path=` endpoints and keep
real HTTP URLs on rendered image elements. The embedded guard also rewrites
protected fetch/XHR values and hydrates late `img[data-protected-src]`
placeholders with the signed iframe HTTP proxy URL. Observe both
`data-protected-src` and `src`, because WeKnora may replace the URL after first
paint. Never create/revoke a temporary object URL here: its lightbox reuses the
thumbnail `src` after click. Never expose the service token or relax the
current-KB authorization to make images work.
**Embedded native knowledge-base switching:** The original WeKnora breadcrumb picker remains available inside the detail iframe. A signed 15-minute embed cookie carries only the current DeerFlow user's visible mappings (their own plus published bases); the proxied `/api/v1/knowledge-bases` response is filtered to that allowlist. Choosing another entry performs a full iframe navigation, rechecks the DeerFlow mapping and read/write permission, then reissues the cookie for the selected base. Never expose the upstream service-admin list directly or permit an arbitrary remote knowledge-base ID.
The independent enterprise-research workbench is **temporarily disabled**. Its source remains under `app/gateway/routers/enterprise_research.py` and `deerflow.persistence.enterprise_research`, but Gateway does not mount `/api/enterprise-research`, initialize its stores/executor, or register its ORM rows in `Base.metadata`. This ensures startup does not create or alter `enterprise_research_tasks` / `enterprise_research_report_jobs` on MySQL. It remains separate from `/api/deep-research`; do not re-enable it incidentally while changing the active Deep Research workflow. Regression coverage for its preserved implementation remains in `tests/test_enterprise_research_tasks.py`.
### Embedded Client (`packages/harness/deerflow/client.py`)
`DeerFlowClient` provides direct in-process access to all DeerFlow capabilities without HTTP services. All return types align with the Gateway API response schemas, so consumer code works identically in HTTP and embedded modes.
**Architecture**: Imports the same `deerflow` modules that Gateway API uses. Shares the same config files and data directories. No FastAPI dependency.
**Agent Conversation**:
- `chat(message, thread_id)` ? synchronous, accumulates streaming deltas per message-id and returns the final AI text
- `stream(message, thread_id)` ? subscribes to LangGraph `stream_mode=["values", "messages", "custom"]` and yields `StreamEvent`:
- `"values"` ? full state snapshot (title, messages, artifacts); AI text already delivered via `messages` mode is **not** re-synthesized here to avoid duplicate deliveries
- `"messages-tuple"` ? per-chunk update: for AI text this is a **delta** (concat per `id` to rebuild the full message); tool calls and tool results are emitted once each
- `"custom"` ? forwarded from `StreamWriter`
- `"end"` ? stream finished (carries cumulative `usage` counted once per message id)
- Agent created lazily via `create_agent()` + `_build_middlewares()`, same as `make_lead_agent`
- Supports `checkpointer` parameter for state persistence across turns
- `reset_agent()` forces agent recreation (e.g. after memory or skill changes)
- See [docs/STREAMING.md](docs/STREAMING.md) for the full design: why Gateway and DeerFlowClient are parallel paths, LangGraph's `stream_mode` semantics, the per-id dedup invariants, and regression testing strategy
**Gateway Equivalent Methods** (replaces Gateway API):
| Category | Methods | Return format |
|----------|---------|---------------|
| Models | `list_models()`, `get_model(name)` | `{"models": [...]}`, `{name, display_name, ...}` |
| MCP | `get_mcp_config()`, `update_mcp_config(servers)` | `{"mcp_servers": {...}}` |
| Skills | `list_skills()`, `get_skill(name)`, `update_skill(name, enabled)`, `install_skill(path)` | `{"skills": [...]}` |
| Memory | `get_memory()`, `reload_memory()`, `get_memory_config()`, `get_memory_status()` | dict |
| Uploads | `upload_files(thread_id, files)`, `list_uploads(thread_id)`, `delete_upload(thread_id, filename)` | `{"success": true, "files": [...]}`, `{"files": [...], "count": N}` |
| Artifacts | `get_artifact(thread_id, path)` ? `(bytes, mime_type)` | tuple |
**Key difference from Gateway**: Upload accepts local `Path` objects instead of HTTP `UploadFile`, rejects directory paths before copying, and reuses a single worker when document conversion must run inside an active event loop. Artifact returns `(bytes, mime_type)` instead of HTTP Response. The new Gateway-only thread cleanup route deletes `.deer-flow/threads/{thread_id}` after LangGraph thread deletion; there is no matching `DeerFlowClient` method yet. `update_mcp_config()` and `update_skill()` automatically invalidate the cached agent.
**Tests**: `tests/test_client.py` (77 unit tests including `TestGatewayConformance`), `tests/test_client_live.py` (live integration tests, requires config.yaml)
**Gateway Conformance Tests** (`TestGatewayConformance`): Validate that every dict-returning client method conforms to the corresponding Gateway Pydantic response model. Each test parses the client output through the Gateway model ? if Gateway adds a required field that the client doesn't provide, Pydantic raises `ValidationError` and CI catches the drift. Covers: `ModelsListResponse`, `ModelResponse`, `SkillsListResponse`, `SkillResponse`, `SkillInstallResponse`, `McpConfigResponse`, `UploadResponse`, `MemoryConfigResponse`, `MemoryStatusResponse`.
## Development Workflow
### Test-Driven Development (TDD) ? MANDATORY
**Every new feature or bug fix MUST be accompanied by unit tests. No exceptions.**
- Write tests in `backend/tests/` following the existing naming convention `test_<feature>.py`
- Run the full suite before and after your change: `make test`
- Tests must pass before a feature is considered complete
- For lightweight config/utility modules, prefer pure unit tests with no external dependencies
- If a module causes circular import issues in tests, add a `sys.modules` mock in `tests/conftest.py` (see existing example for `deerflow.subagents.executor`)
```bash
# Run all tests
make test
# Run a specific test file
PYTHONPATH=. uv run pytest tests/test_<feature>.py -v
```
### Running the Full Application
From the **project root** directory:
```bash
make dev
```
This starts all services and makes the application available at `http://localhost:2026`.
**All startup modes:**
| | **Local Foreground** | **Local Daemon** | **Docker Dev** | **Docker Prod** |
|---|---|---|---|---|
| **Dev** | `./scripts/serve.sh --dev`<br/>`make dev` | `./scripts/serve.sh --dev --daemon`<br/>`make dev-daemon` | `./scripts/docker.sh start`<br/>`make docker-start` | ? |
| **Prod** | `./scripts/serve.sh --prod`<br/>`make start` | `./scripts/serve.sh --prod --daemon`<br/>`make start-daemon` | ? | `./scripts/deploy.sh`<br/>`make up` |
| Action | Local | Docker Dev | Docker Prod |
|---|---|---|---|
| **Stop** | `./scripts/serve.sh --stop`<br/>`make stop` | `./scripts/docker.sh stop`<br/>`make docker-stop` | `./scripts/deploy.sh down`<br/>`make down` |
| **Restart** | `./scripts/serve.sh --restart [flags]` | `./scripts/docker.sh restart` | ? |
**Nginx routing**:
- `/api/langgraph/*` ? Gateway embedded runtime (8001), rewritten to `/api/*`
- `/api/*` (other) ? Gateway API (8001)
- `/` (non-API) ? Frontend (3000)
### Running Backend Services Separately
From the **backend** directory:
```bash
# Gateway API
make gateway
```
Direct access (without nginx):
- Gateway: `http://localhost:8001`
### Browser CORS
The Gateway always permits local Vite origins on ports 5173, 5174, 3000, and
8080 for both `localhost` and `127.0.0.1` with credentials. A deployed
frontend on another origin must be listed exactly in `GATEWAY_CORS_ORIGINS`
(for example `http://47.88.25.99:7010`); explicit origins are merged with the
local development origins. Regression: `tests/test_taskcop_cors.py`.
### Frontend Configuration
The frontend uses environment variables to connect to backend services:
- `NEXT_PUBLIC_LANGGRAPH_BASE_URL` - Defaults to `/api/langgraph` (through nginx)
- `NEXT_PUBLIC_BACKEND_BASE_URL` - Defaults to empty string (through nginx)
When using `make dev` from root, the frontend automatically connects through nginx.
## Key Features
### Administrator-managed custom themes
`system_settings.json` → `appearance` now stores `custom_themes` alongside
the built-in default theme. A row is `CustomThemePreset` (`id`, `name`,
six-digit `primary`, optional primary-button `text_color`, `mode`, `overrides`, `enabled`,
`sort_order`). `overrides` is intentionally a short interaction-only list:
`button_hover`, `navigation_gradient_start`, `navigation_gradient_end`,
`selected_background`, `border`, `muted`, `muted_text`, `destructive`,
`brand_title`, and `primary_foreground`. The navigation gradient endpoints
default from `primary` and are only stored when an administrator overrides them.
All values must be six-digit hex strings; the frontend derives every remaining
semantic CSS/TDesign token so admins do not maintain a full raw token sheet.
Ids must be `custom-theme-*`, names and ids are unique, and the router limits
the list to 12. `GET/PUT /api/system-settings/appearance` is the single
read/write surface: reads are available to all users, writes require admin.
The router re-orders rows deterministically and only permits an enabled custom
theme as `default_theme`; removing or disabling the default safely falls back
to `light`. Regression coverage: `tests/test_appearance_theme.py`.
Compact top-bar shortcuts live in the same `appearance` setting. Their route
parameters support `login`, `theme`, and keyless `path`. A `path` parameter has
no query-key input; the operator selects its value source from the right-hand
dropdown and the frontend appends the resolved, URL-encoded value to the
external URL path (or Hash route) in configured order. Password stays an
ordinary `login` query parameter with the key `password`, so an external login
link can become `#/login/<user>?password=…`. Regression coverage:
`tests/test_appearance_theme.py`.
### File Upload
Multi-file upload with automatic document conversion:
- Endpoint: `POST /api/threads/{thread_id}/uploads`
- Supports: PDF, PPT, Excel, Word documents (converted via `markitdown`)
- Rejects directory inputs before copying so uploads stay all-or-nothing
- Reuses one conversion worker per request when called from an active event loop
- Files stored in thread-isolated directories
- Agent receives uploaded file list via `UploadsMiddleware`
See [docs/FILE_UPLOAD.md](docs/FILE_UPLOAD.md) for details.
### Plan Mode
TodoList middleware for complex multi-step tasks:
- Controlled via runtime config: `config.configurable.is_plan_mode = True`
- Provides `write_todos` tool for task tracking
- One task in_progress at a time, real-time updates
See [docs/plan_mode_usage.md](docs/plan_mode_usage.md) for details.
### Context Summarization
Automatic conversation summarization when approaching token limits:
- **On/off switch is per-request, not the config file.** The chat input box has a
「上下文压缩」 toggle (below 参考文献, **default off**) → frontend sends
`context.summarization_enabled` → whitelisted in `_CONTEXT_CONFIGURABLE_KEYS`
(`app/gateway/services.py`) → `_build_middlewares` reads it from the run config
and calls `_create_summarization_middleware(force_enabled=summarization_enabled)`.
`force_enabled=True` mounts the middleware even when `config.summarization.enabled`
is false; when the flag is absent (channels / scheduled runs) it falls back to
the config flag. `config.summarization.enabled` now defaults to **false**.
- `config.yaml` `summarization` still supplies the **tuning**: trigger types
(tokens, messages, or fraction of max input) and the keep policy.
- A compressed request must retain the latest non-empty `HumanMessage`, even
when it has already been folded into the hidden summary. This keeps strict
OpenAI-compatible gateways from rejecting a tool-continuation call with
`No user query found in message`.
- Keeps recent messages while summarizing older ones.
Tests: `tests/test_summarization_toggle.py` (gating + whitelist),
`tests/test_summarization_middleware.py` (cache behavior).
See [docs/summarization.md](docs/summarization.md) for details.
### Vision Support / Image Recognition (??)
Two paths, selected by `config.yaml` ? `vision`:
**Dedicated vision model (delegated recognition)** ? set `vision.model_name` to a model from `models[]`. The `view_image` tool sends the image to that model and returns its **text description** as the tool result, so a text-only main chat model can still "see" images. `view_image` accepts an optional `question` argument for targeted queries; the default prompt is `vision.prompt` (overridable), and `vision.max_tokens` caps the recognition output. Resolution/gating lives in `deerflow/models/vision.py` (`resolve_vision_model_name`, `vision_enabled`, `is_delegated_vision`, `recognize_image`).
**Native vision (fallback)** ? when `vision.model_name` is unset, behavior falls back to the main chat model, which only works if it has `supports_vision: true`:
- `ViewImageMiddleware` injects base64 image data into the conversation before the LLM call
- `view_image` stores the image in `state.viewed_images` for that middleware
Both the `view_image` tool and `ViewImageMiddleware` are registered whenever `vision_enabled(model_name)` is true (dedicated model configured **or** main model supports vision). Config schema: `deerflow/config/vision_config.py` (`VisionConfig`); tests: `tests/test_vision.py`.
### Upload symlink-safety (cross-platform)
`deerflow/uploads/manager.py:open_upload_file_no_symlink` writes uploads without following a pre-existing destination symlink. `O_NOFOLLOW` is applied where available (POSIX) but is **optional** ? Windows lacks `os.O_NOFOLLOW`/`O_NONBLOCK`, so requiring it previously made every Windows upload fail with `UnsafeUploadPathError`. Pre-existing symlinks are still rejected via `lstat`, and the opened fd is re-checked with `fstat`. `O_BINARY` is added on Windows. Regression: `tests/test_uploads_windows_safe.py`.
## Code Style
- Uses `ruff` for linting and formatting
- Line length: 240 characters
- Python 3.12+ with type hints
- Double quotes, space indentation
## Documentation
### WeKnora iframe behind `/deerflow`
The LLMWiki WeKnora detail iframe is publicly served through the Gateway at
`/deerflow`, independently of the frontend deployment prefix (for example
`/magentweb`). Register `llmwiki.proxy_router` with the `/deerflow` prefix
before its unprefixed fallback router, otherwise the unprefixed catch-all
captures `/deerflow/*`. HTML root-relative assets, API strings, and Vite's
`__vite__mapDeps` table are rewritten for `/deerflow`; generic JavaScript
regular-expression literals must not be rewritten. Auth and CSRF bypass only
applies when the short-lived signed WeKnora embed cookie is valid. Keep static
path recognition in `app/gateway/weknora_embed.py` synchronized with the
proxy's static fallback paths in `app/gateway/routers/llmwiki.py`.
### Position-roundtable role directory
`position_roles` is the global shared directory for position-roundtable work
views. It is deliberately separate from organisation/RBAC `positions`.
`app.gateway.deps.langgraph_runtime()` seeds the four stable built-in ids from
`app/gateway/position_role_defaults.py` via `ensure_seed()` (missing rows only;
never overwrite operator edits). The API is `GET/PUT /api/position-roles`;
writes are admin-only. Exactly one enabled `intent` role is required and owns
the initial intent workspace plus the summary/action-plan closing agents. Do
not delete a built-in id because historical chain nodes persist those ids;
custom role ids may be added and removed.
### Position-roundtable task-scoped session history
`position_roundtable_sessions.external_task_id` is optional. When populated from the frontend route `taskId`, new rows set `task_scoped = true`: a task can have multiple sessions and all users opening that task see the same sessions, nodes and persisted conversation snapshots. When it is NULL/empty, keep the original `user_id`-scoped personal workspace behavior. `20260808_01` adds `task_scoped` with default `false` (so all historical records remain personal) plus the task/update-time index; it must not backfill or make `external_task_id` non-null.
### Roundtable artifact submission status
`roundtable_artifact_submissions` keeps task-level submission state for a
generated deliverable and is deliberately separate from task details and any
external task-system delivery. The API is
`GET /api/roundtable/tasks/{taskId}/artifacts?submitted=true` and
`PUT /api/roundtable/tasks/{taskId}/artifacts/submission`. All callers use the
normal Gateway Bearer authentication; records are shared by task while
`created_by`/`submitted_by`/`cancelled_by` are audit-only. The server derives a
SHA-256 idempotency key from `taskId + sourceKey + path + revision`: submitting
the same revision is an upsert, while cancellation only flips `submitted` and
sets `cancelled_at` (it never deletes history). Store only the sandbox file
reference (`threadId`, path, source/node/session fields, name/MIME and revision)
and timestamps; do not persist Markdown content or call a formal delivery
interface. Previews continue to read the DeerFlow artifact by `threadId + path`.
See `docs/` directory for detailed documentation:
- [../README.md](../README.md) - Offline Docker wheelhouse preparation and startup workflow. The pinned versions in `../offline-runtime-requirements.txt` and `../start_backend.sh` must stay synchronized. Startup checks `croniter`, `hindsight-client`, the aiomysql compatibility bundle, and Redis, installing missing or incompatible packages from wheelhouse only. The bundle installs `aiomysql==0.3.2`, `PyMySQL==1.2.0`, and `SQLAlchemy==2.0.51` together; the startup probe also verifies that SQLAlchemy's adapted `ping(reconnect=False)` signature is present, preventing the old-adapter missing-`reconnect` failure. The download target defaults to CPython 3.12 on Linux x86_64 and verifies dependency closure using wheelhouse only.
- [CONFIGURATION.md](docs/CONFIGURATION.md) - Configuration options
- [ARCHITECTURE.md](docs/ARCHITECTURE.md) - Architecture details
- [API.md](docs/API.md) - API reference
- [SETUP.md](docs/SETUP.md) - Setup guide
- [FILE_UPLOAD.md](docs/FILE_UPLOAD.md) - File upload feature
- [PATH_EXAMPLES.md](docs/PATH_EXAMPLES.md) - Path types and usage
- [summarization.md](docs/summarization.md) - Context summarization
- [plan_mode_usage.md](docs/plan_mode_usage.md) - Plan mode with TodoList
- [BACKEND_RUN_STOP_ZH.md](docs/BACKEND_RUN_STOP_ZH.md) - Windows-first backend start/stop and troubleshooting guide
### Deep Research retrieval
`deerflow.agents.deep_research.adapters.material_provider.DeerFlowMaterialProvider`
must remain on the Q&A retrieval path. Enabled `deep-search`, `web-research`,
and knowledge-base skills run through the existing `SubagentExecutor` skill
runtime only when their required Q&A tool is registered; then the provider may
use configured `web_search`. A disabled placeholder endpoint is not a valid
search channel: use the maintained keyless DuckDuckGo provider and a relevant
real-web fallback as read-only sources, keep source provenance, and never
enable or call a placeholder just to make the Deep Research UI look configured. Regression coverage is in
`tests/test_deep_research_material_provider.py`.
When a queued Deep Research job is cancelled, the Gateway must also clear the
owning session's `active_job_id` and mark it `cancelled`; otherwise the stale
session incorrectly consumes that user's concurrency limit.
Report prose is the one Deep Research LLM path that must use
`DeerFlowCompletionBackend.stream_complete()` rather than `complete()`: it
drives `model.astream()`, sends native `report_delta` frames through
`ReportChunkEmitter`, and persists bounded `report_chunk` checkpoints. The
app-layer `DeepResearchLiveHub` fans out those live frames inside one Gateway
worker only; the database event log still owns `?after=<seq>` reconnect replay
and multi-worker recovery. Inject the hub callback into `PersistingEventSink`
and never import `app.*` from `deerflow.*`. Keep structured plan/review calls
on `complete()`. Before the search fan-out, basic mode emits one
`queries_planned` event and detailed/multi-agent modes emit one per subtopic;
they are durable user-visible search terms, not chain-of-thought. On completion, retain
the terminal `lastJobId` in `usage_snapshot` so a refreshed conversation page
can replay its trace from the event log. Detailed mode uses stable targets (`introduction`,
`section:N`, `conclusion`) because section writers are concurrent;
multi-agent revisions must emit `report_reset` before their next report
version. Do not persist raw model reasoning or one persisted event per provider
token. Models that expose thinking inside ordinary `<think>...</think>` text
must be split statefully before a report delta, checkpoint, or Markdown
projection is created; an optional live-only `report_thinking` frame can drive
the message-list thought trace, while the sandbox receives answer Markdown
only. The post-report digest (`summarize_report`) must likewise forward
provider reasoning through live-only `summary_thinking` frames so the
chat timeline can render thought while waiting for the first `summary_delta`.
The Vite page starts Word downloads through
`frontend-web/src/core/scheduled-tasks/markdown-docx.ts`, but the Gateway owns
generation through `POST /api/writing/export/docx`. Keep the standard-library
OOXML builder and its embedded-font validation; do not add a Python document-
conversion dependency merely to export text.
All non-streaming Deep Research control-plane calls (`complete()` for query /
section / plan / review JSON) run through `asyncio.wait_for` using the frozen
`DeepResearchConfig.control_plane_timeout_seconds` (45 seconds by default,
clamped to 120). Do not apply that total-duration cap to `stream_complete()`:
reports may legitimately be long but must continue emitting visible deltas. A
basic query-planning timeout/failure emits `query_planning_fallback` and searches
the original topic; detailed and multi-agent plan creation similarly fall back
to their original topic/section, rather than leaving a durable job stuck in
`planning`. Tests: `tests/test_deep_research_llm_adapter.py` and
`tests/test_deep_research_runners.py`.
The per-user Deep Research concurrency guard counts only session statuses
`running` and `awaiting_input`. A `draft` has no durable job and must never
consume a slot; otherwise saved drafts make `POST /sessions/{id}/jobs` return
429 permanently. Regression: `test_saved_draft_sessions_do_not_consume_research_concurrency`.
`POST /api/deep-research/sessions/{id}/jobs` is retry-safe under MySQL
read/write splitting. Deep Research session/job/event/source/message writes
must serialise their changed row while the write transaction is still open
(`flush` for inserts, select-after-update **before** commit), never through a
post-commit `session.refresh()` or a fresh pooled read. Replica lag otherwise
turns a committed job into a misleading 500, can make dispatcher `claim_job()`
lose its snapshot, or can fail report/follow-up persistence after the model has
already completed. Always nudge the dispatcher for both a newly created job and
an idempotently reused queued job; otherwise a retry after such an interrupted
response returns the old job but leaves it waiting for the periodic scan. The
SSE replay generator must catch temporary event/job-store reads, emit a
recoverable warning, and keep the connection open so the next poll can repair
the cursor instead of surfacing an ASGI `ExceptionGroup`. Regressions:
`test_deep_research_writes_never_depend_on_post_commit_refresh`,
`test_job_create_never_reads_back_after_commit`,
`test_retry_of_existing_queued_job_nudges_dispatcher`, and
`test_stream_stays_open_when_durable_replay_has_a_transient_failure`.
At ordinary report completion, persist and confirm the session's
`report_markdown` projection **before** marking the Job `completed`. Retry a
short transient projection failure; if it cannot be confirmed, let the job
become failed rather than reporting a report that vanishes on refresh.
`test_report_is_persisted_before_job_is_marked_completed` and
`test_report_projection_failure_never_marks_job_completed` cover the ordering.
The chat-collection frontend must keep report-writing progress in the existing
workspace `MessageList` / `MessageGroup` tool-step renderer: inject only
virtual `deep_research_progress` tool calls for the durable phases, keep the
completed phase visible after the job ends, and append the digest as a normal
assistant message. Do not recreate a second deep-research-only conversation
timeline. On `summarizing`, insert that same normal assistant message with its
streaming marker and update it for every SSE `summary_delta`; when the report
completes, replace its content in place with the final completion/digest
message rather than waiting to append a summary. Completed chat-collection
sessions also inject a normal `present_files` card into the transcript, named
from the report Markdown H1 rather than its internal `report.md` path.
Replaying a completed history session must not auto-open its
sandbox; opening that card selects its virtual write-file report instead, and
its file-card download action delegates to the Deep Research Markdown/Word
export dialog. The default in-job concurrency and per-user active-job cap are
both three. Before the user confirms the report configuration, derive its
displayed material count from unique collector tool-result rows because the
durable Deep Research source store is populated when the writing job harvests
that thread. The Deep Research composer model picker must use the installed
TDesign `Select` backed by `/api/models`; its explicit choice is sent as
`model_name` on collector turns and applied to `fast_model`, `smart_model`, and
`strategic_model` when creating a session or starting a report job. Disable it
while collection or report writing is active so one durable run cannot switch
models mid-stream.
Report structures (`/api/report-structures`) may store an optional
`retrieval_directions` list (检索方向) and an optional `retrieval_skills`
list (检索来源 / 指定技能). Empty / unset keeps today's collector
query-planning, default tools, and runner sub-query generation. A non-empty
directions list is frozen into `DeepResearchConfig.retrieval_directions`: the
chat collector's first user message gets a `<retrieval-directions>` block that
asks it to **split concrete retrievable terms along those facets** (never search
the direction labels themselves), and legacy `basic`/`detailed`/`multi_agent`/
`deep` runners pass the same constraint into query-planning instead of using
`课题 + 方向` as the search string. A non-empty skills list is frozen
into `DeepResearchConfig.retrieval_skills`: the chat collector's first user
message gets a `<retrieval-skills>` block, the collector run context receives
`extra_skills` so `skill_view`/`skill_list` can see those names, and legacy
material collection runs those skills instead of the default Q&A research-skill
trio. Helpers live in `deerflow.agents.deep_research.retrieval`.
Completed-report Q&A has two endpoints: retain `POST /sessions/{id}/chat` as the
legacy JSON response for compatibility, and use `POST /sessions/{id}/chat/stream`
for report-grounded Q&A SSE. The workbench must send a strong scoped edit
instruction (for example “修改第一段” / “重新生成一下第一段” / “重写第二节”) to
the separate `POST /sessions/{id}/rewrite/stream` report-section writing pipeline.
That endpoint requires a deterministic current-report range (otherwise 422), so
its output is always a replacement candidate rather than a Q&A answer. The
candidate is streamed in the ordinary message row
while the sandbox preserves the current report. When generation ends, inject a
confirmation card. Only the owner’s explicit
`POST /sessions/{id}/messages/{messageId}/rewrite-proposal` action `apply` may
persist the assembled `report_markdown`; `dismiss` preserves the original. The
proposal is bound to a hash of the report it was generated from, so stale
candidates fail with 409 rather than overwrite a later edit. Ambiguous or
ordinary evidence questions must remain non-mutating, and cancellation or a
stream error must never save a partial report. The rewrite prompt receives only
the selected range plus the selected evidence set, never a new web search.
### DeerFlow-local LLMWiki Wiki index
`llmwiki.local_wiki_index` is an opt-in local-vector subsystem. WeKnora is the
Wiki generation/sync source and the Wiki-only fallback while an entire mapping
has not completed vectorization. The readiness gate requires `state=idle`, a
successful full scan, matching embedding fingerprint, no failed pages,
`ready_page_count == local_page_count`, and vectors for a non-empty library.
Ready mappings read DeerFlow's `llmwiki_wiki_pages` /
`llmwiki_wiki_vectors` / `llmwiki_wiki_sync_states` tables; incomplete mappings
call only WeKnora `/knowledgebase/{id}/wiki/pages?query=...` plus Wiki-page
detail. Never call `/knowledge-search` from internal search, RAG, or the
`llmwiki_search` tool, and never return raw chunks as the fallback. Core code lives under
`deerflow.integrations.weknora.local_index` and
`deerflow.persistence.llmwiki_index`; preserve the harness → app import
firewall. Gateway composition injects the store/search/sync services through
`local_index.runtime`. `enabled` turns on local vector search and `auto_sync`
defaults to the lease-protected background scheduler. The scheduler performs
an immediate incremental scan at startup, then rescans every configured
interval so new WeKnora libraries and asynchronously generated Wiki pages are
picked up without a manual vectorization click. Unchanged content must reuse
existing vectors.
Vectors are normalized little-endian float32 BLOBs. A page mirror and all of
its vectors are replaced in one transaction only after strict batched
embedding succeeds. Never delete the old generation before a remote embedding
call and never fall back to WeKnora raw chunks while the local feature flag is
enabled. NumPy matrices are per-mapping LRU snapshots invalidated by
`index_revision`; large decode/matrix work runs through `asyncio.to_thread`.
`embedding.dimensions` is a local response-size invariant and fingerprint
input, not an outbound API field: fixed-dimension BGE/OpenAI-compatible
endpoints can reject a `dimensions` request parameter with HTTP 400. When the
value is omitted, the embedding client probes and freezes the first response's
dimension before sync lease/rebuild decisions.
Embedding errors must retain the upstream status and bounded response body in
logs, without exposing API keys. SQL vector writes are split into statements
of at most `database_batch_size` rows (default 50) inside the same atomic page
replacement transaction.
During a fingerprint rebuild the mapping uses the Wiki-page API until every
page is successfully rebuilt. `WikiRetrievalService` partitions mixed requests
per mapping, so ready libraries stay local while incomplete libraries use that
Wiki-only fallback.
The frontend must keep Wiki citations visibly distinct from raw-document
citations while hiding provider-brand wording from user-facing labels, status
text, and errors. Any click on a Wiki source card opens the shared right-side
source drawer, which renders the complete Markdown page. When the local index
is disabled, the drawer reads the remote WeKnora Wiki-page endpoint directly
and must not surface the local-index-disabled response. A Wiki target must not
also load or render the legacy original-document preview, even when its citation
still carries a document id. The drawer supports single-page Markdown and Word
downloads, and its **Open in Wiki** link carries both the DeerFlow knowledge-base
mapping id and the exact `wikiSlug`. Wiki search-result cards use this same
drawer instead of navigating immediately. Exact-page navigation uses the native
`knowledge/{mapping_id}/wiki/view?wikiSlug=...` reader, not the WeKnora detail
iframe. Keep legacy `knowledge/{mapping_id}?tab=wiki&wikiSlug=...` links fast by
redirecting them in `LlmWikiKnowledgePage` before its runtime-loading gate.
Drawer-to-reader navigation stays in the current SPA and passes the fetched
Wiki page in router state so the article can render before the background query
refresh completes. The native reader must retain a searchable/collapsible left
Wiki directory, reconstruct hierarchy from `parent_slug`, `wiki_path`,
`category_path`, and slug fallback, expand/highlight the active page, and keep
directory navigation on the same lightweight route rather than mounting the
WeKnora iframe.
Generated HashRouter URLs must retain `window.location.pathname` (for example
the `/web/` reverse-proxy prefix); never emit root-absolute `/#/...` citation,
reader, or public-embed links. Protected Wiki image rendering must use real HTTP
URLs, never `URL.createObjectURL`. Public readers use the published file proxy;
authenticated readers obtain a short-lived signed capability from
`/files/access-url` and load `/api/public/llmwiki/files/{token}`. The signed
payload is restricted to `kind=wiki_file`, mapping id, remote id, validated
protected path, and expiry, and responses provide an inline UTF-8 filename for
new-tab viewing and browser Save As. The proxied WeKnora detail iframe must
likewise leave its scoped HTTP file-proxy URL on every protected image and
restore it if upstream rendering replaces `src`; its click preview reads that
same DOM value, so a fetched-and-revoked object URL is not acceptable.
Third-party pages embed published Wiki content through the chrome-free route
`/#/embed/knowledge/{mapping_id}` (optional exact-page `wikiSlug`). It defaults
to a recognized Wiki index page (title `索引`/`index`, page type `index`, or an
`index`/`home`/`readme` slug) and promotes that page to the first public-directory
entry. When upstream pages exist without an explicit index shell, the frontend
injects a virtual `__index__` page and builds its catalogue from the page list;
an all-empty dynamic index response must also fall back to that page list. It retains
directory navigation and supports the shared public-embed `theme` / `bg` /
`text` appearance inputs. WeKnora dynamically renders index page-type counts,
nested category headings, page summaries, and `data-slug` links even though the
stored index Markdown only contains its introduction; the native public reader
reads the same ordered groups through the published-only
`/api/public/llmwiki/knowledge-bases/{id}/wiki/index` proxy and renders them with
`buildWikiIndexCatalogMarkdown`. The Gateway paginates every WeKnora index type
(`summary`, `entity`, `concept`, `synthesis`, `comparison`) to completion while
preserving upstream group/item order. Index Markdown
may also contain WeKnora shell-only placeholder links;
`resolveWikiIndexPageLinks` matches their visible labels to
the loaded page title, slug, or alias and rewrites them to exact embedded
`wikiSlug` routes. Preserve WeKnora's `wiki-content-link` interaction: primary
brand color, medium weight, dashed underline becoming solid on hover, focus
outline, and in-reader SPA navigation rather than a page reload. `GET
/api/public/llmwiki/knowledge-bases` is the anonymous discovery endpoint. It
validates published ordinary mappings against WeKnora, drops stale mappings,
and returns safe comprehensive metadata from the WeKnora list/detail/page APIs:
document/chunk/processing/share counts, non-secret configuration and timestamps,
Wiki page type/status counts and explicit/generated index entry, plus aggregate
local-vector state/counts/timestamps. Never expose remote WeKnora ids,
tenant/creator identity, storage credentials, API credentials, raw vectors, or
per-page indexing errors. All anonymous reader calls use that
public namespace and the `/knowledge-bases/{id}` Wiki-page children. Those
endpoints must resolve with a synthetic anonymous user and `is_admin=False`,
explicitly require `publication_status=published`, and hard-deny the system
conversation-deposit mapping. Never point this page at the authenticated
`/api/llmwiki` reads or weaken the published-only gate merely because the
mapping id is hard to guess.
Owners manually force a full sync/vectorization through `POST
/api/llmwiki/knowledge-bases/{id}/vectorize`; status is available through `GET
/api/llmwiki/knowledge-bases/{id}/wiki-index-status`. That response includes
aggregate `progress` and per-page `pages` details; current-scan progress is
derived from `active_sync_id` / `last_seen_sync_id`, while each page exposes its
index state, vector count, timestamps, and bounded error. The Wiki management
page must show the progress bar and a status badge beside every article. It
also exports all remote processed Wiki pages via `GET
/api/llmwiki/knowledge-bases/{id}/wiki/export.xlsx`. Directory export must keep
the full path, current directory, parent path, and dynamic per-level columns;
long Markdown is split into Excel-safe continuation columns instead of being
truncated. Workbook creation runs in `asyncio.to_thread`.
Internal search authorization always starts from DeerFlow mapping ids and is
rechecked by the Gateway. External endpoints use the isolated
`/api/external/llmwiki/wiki` namespace and `X-API-Key`; require published +
external-enabled + index-enabled on every request and response, and hard-deny
the `system` / `对话沉淀` mapping. Those API-key endpoints never return remote
WeKnora ids, provenance refs, credentials, or vector bytes.
`POST /api/knowledge/vector-search` is a deliberately separate anonymous API.
It has no token/API-key requirement or rate limit, applies route-scoped
`Access-Control-Allow-Origin: *`, searches every published local Wiki index
regardless of `external_search_enabled`, and still hard-denies the `system` /
`对话沉淀` mapping before and after search. Its contract is section-level (one
result per matched vector, so Wiki-page fields and URLs may repeat), default
`top_k=10`, with optional `include_page_content`. Unlike the protected legacy
external namespace, this endpoint intentionally returns DeerFlow/WeKnora ids
and raw source/chunk provenance, but never credentials or vector bytes. Wiki
links use `llmwiki.local_wiki_index.frontend_base_url` (development default
`http://127.0.0.1:5174`) and point to the native
`/#/page/workspace/knowledge/{mapping_id}/wiki/view?wikiSlug=...` route.
Regression tests:
`tests/test_llmwiki_local_index_repository.py`,
`tests/test_llmwiki_wiki_sync.py`, `tests/test_llmwiki_vector_search.py`,
`tests/test_external_llmwiki_search.py`,
`tests/test_public_knowledge_vector_search.py`, and
`tests/test_llmwiki_wiki_retrieval.py`.
The sandbox whole-report rewrite may carry an optional `reportOutline` Markdown
template from the shared report-structure library. Persist the original template
in the durable input snapshot and append it as an explicit chapter-structure
constraint before requirements planning and prose generation; a template must
never be used as permission to invent facts absent from the current report.
`DocumentRewriteVersionRepository.create` backs the reversible before-image for
both artifact and whole-report rewrites. It must flush and serialize the new row
on the writer before commit, and must never call `session.refresh()` after
commit: a read/write-splitting MySQL endpoint can route that refresh to a
lagging replica, falsely report a version-save failure, and trigger the caller's
intentional restore of the rewritten document.
Version snapshots are an undo enhancement, not a condition for preserving an
already committed rewrite. If snapshot creation fails after the artifact/report
write, log the exception, complete the rewrite with
`versionSnapshotFailed: true`, and do not emit a user-facing failure or restore
the original. The sandbox uses that signal only to enable its ordinary manual
save fallback.
### Offline local-sandbox Office runtime
The DeerFlow-only offline deployment uses `LocalSandboxProvider`, so document
dependencies belong in the backend image rather than a separate sandbox image.
`deploy/linux-docker-offline/Dockerfile.backend` installs python-docx,
LibreOffice Writer/Calc/Impress, Pandoc, Poppler, and CJK fonts in addition to
the MarkItDown Office extras. `file_conversion.py` normalizes legacy
DOC/XLS/PPT and ODT/ODS/ODP through a per-conversion LibreOffice user profile
before MarkItDown extraction, avoiding concurrent profile locks. The image
build and `OFFLINE_INSTALL.sh --office-check` both run
`scripts/office_runtime_selfcheck.py`; that smoke test creates and re-parses
minimal DOCX/XLSX/PPTX files and does not need the application database.
## Skill/assistant knowledge invariants
Skill distillation is implemented in `deerflow.skill_knowledge` and
`deerflow.persistence.skill_knowledge`; native assistant bases are in
`deerflow.assistant_knowledge` and
`deerflow.persistence.assistant_knowledge`. Gateway composition, remote adapters,
the lease dispatcher, and APIs remain in `app.gateway` so the harness-to-app
import firewall is preserved.
Never execute a skill or read/upload/summarize its code and script files.
Knowledge fingerprints exclude those files and Markdown code fences. All
resyncs must stage and validate the knowledge digest before activation; the old
snapshot remains authoritative on failure. A binding may have only one active
item (`active_dedupe_key` plus atomic lease claims). Skill distillation targets
only writable WeKnora Wiki knowledge bases in this product version; reject direct
assistant-base targets, non-Wiki/FAQ/conversation/readonly targets, and direct
Neo4j writes. The create/list/resync/retry/detach flow is available to ordinary
users for their own visible skills, writable knowledge bases, jobs and bindings;
administrator-only auto-sync/review APIs remain privileged. Detach/delete operates on source contributions, not canonical rows,
and shared entities/relations are retired only when their final contribution is
gone.
Assistant knowledge has one global full-library base (`知识梳理总库`) and any
number of admin-created custom bases. Ordinary users have read/search access and
administrators own all mutations. Assistant bases accept local documents independently
of WeKnora via `/api/assistant-knowledge/imports/files`, as well as WeKnora export
packages or imports triggered from a WeKnora mapping. Local documents use the
Gateway `assistant_file_ingest` pipeline: persisted original files, chunked LLM Wiki
extraction, cached body embeddings, and import-job status/retry. Do not add direct skill
imports. Importing into a
custom assistant base must also mirror the package into the global base for
cross-source entity alignment. Admins may rename any assistant base and delete
custom bases. Deletion removes the custom base's Wiki pages, vectors, entities,
relations, import audit rows, and package files; the global base is never
deletable because it owns full-library aggregation.
### Wiki package and local encoder invariants (2026-09)
Ordinary knowledge detail landing is globally configured by `ordinary_knowledge_landing` (`wiki` default / `weknora`), with admin-only PUT `/api/system-settings/ordinary-knowledge-landing`. Explicit document/graph/article links bypass this default. Wiki reading and the embedded WeKnora view provide reciprocal navigation; `?view=weknora` explicitly selects the iframe and must not loop back to the default Wiki route.
- `knowledge_transfer.sync_to_global` is called after successful explicit local
Wiki vectorization; it skips conversation/system mappings. Do not enable a new
watcher or automatic remote-write scheduler as a side effect.
- `wiki_packages` exposes file-backed assistant exports and restores into empty
ordinary Wiki libraries. Validate vectors before any remote mutation; recreate
folder IDs from category paths using provider APIs (a raw category_path is
discarded by WeKnora without a matching folder_id). Preserve page metadata and
rebuild backlinks after restoring all pages.
- `PackageArchive` is a replayable, disk-backed JSONL staging index. Always close
it in finally. Runtime imports must not use the legacy in-memory read_package.
- Preserve normalized float32 base64 bytes exactly. Same dimension alone is not
model compatibility. Assistant vector retrieval is fingerprint-gated cosine
search, and any keyword fallback must be labeled as such.
Preserve a valid zero section_index. Export/search vectors belonging to each
current source snapshot (`source.last_import_job_id`), not merely the page's
current revision: one aligned page may intentionally contain current vectors
from several sources.
- The `local` embedding provider uses FastEmbed with official BAAI/bge-m3 ONNX,
1024 dimensions, bounded CPU batches and a per-client async lock. Run inference
off the event loop. Existing page vectors are reusable; zero-vector nonempty
pages need backfill, not a forced rebuild of every page.
FastEmbed is a standard backend dependency. Encoder configuration, directory
upload and warmup live only in the administrator global settings dialog.
A runtime `system_settings.json` path override wins over config.yaml; uploaded
files are streamed into `.deer-flow/models/bge-m3`, never the source tree.
Encoder warmup is explicitly started from the settings UI/API and must never block application startup
or turn a model-load failure into a DeerFlow service failure. Local provider
is offline-only: require local_model_path or DEERFLOW_BGE_M3_MODEL_PATH and
never let the application download model weights. Keep
`local_process_isolation=true` by default: inference runs in the bundled JSONL
worker subprocess, and shutdown must terminate that worker without waiting on
model teardown.
- Every successful mapping snapshot is mirrored to the global assistant base,
including empty snapshots so deleted source pages are retired. Re-import is
source-replacement semantics: absent Wiki/chunk/entity/relation pages must not
remain visible merely because they existed in an older source snapshot.
- Provider graph nodes bind to their original wiki_slug. They must not create
duplicate synthetic entity/relation articles. Store original Wiki metadata
separately from assistant workflow status. Count Wiki pages separately from
raw document chunks.
- The global base aligns entities by normalized display name and aliases across
source records, and aligns content Wiki pages by normalized title. Do not use
source slug as the global canonical identity. Merge unique knowledge blocks
while retaining source contributions/origins. Uploaded packages must rebind
the temporary upload source to the stable WeKnora knowledge-base id in the
package, so re-upload is source-replacement rather than append. Retrieval
candidates exclude summary/index/directory pages; keyword matching, vector
snippets, citations and prompts must use knowledge body/sections, never page
summary as a substitute.
- Migration `20260903_01` adds nullable page metadata and merges the pre-existing
knowledge/workflow Alembic heads. Regression coverage is in
`tests/test_wiki_package_roundtrip.py`, `test_llmwiki_wiki_sync.py` and
`test_assistant_knowledge_repository.py`.
## openGauss / GaussDB persistence invariant
`database.backend=gauss` is an application-ORM backend, parallel to `mysql`.
It normalizes `database.gauss_url` to `opengauss+asyncpg://` and requires the
`gauss` optional dependency (`asyncpg` plus `opengauss-sqlalchemy`). Never route
a Gauss URL through SQLAlchemy's generic PostgreSQL dialect: the wire protocol
may connect, but dialect initialization and schema DDL can fail against the
Gauss server/version behavior. Keep the LangGraph checkpointer on the separately
configured SQLite backend; the PostgreSQL checkpointer schema is not part of the
Gauss compatibility contract. Treat both `postgresql` and `opengauss` dialect
names as timezone-aware PostgreSQL-family engines in portable types and
dialect-specific session settings.