1847 lines
269 KiB
Markdown
1847 lines
269 KiB
Markdown
# CLAUDE.md
|
||
??
|
||
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||
|
||
## Project Overview
|
||
|
||
DeerFlow is a LangGraph-based AI super agent system with a full-stack architecture. The backend provides a "super agent" with sandbox execution, persistent memory, subagent delegation, and extensible tool integration - all operating in per-thread isolated environments.
|
||
|
||
**Architecture**:
|
||
- **Gateway API** (port 8001): REST API plus embedded LangGraph-compatible agent runtime
|
||
- **Frontend** (port 3000): Next.js web interface
|
||
- **Nginx** (port 2026): Unified reverse proxy entry point
|
||
- **Provisioner** (port 8002, optional in Docker dev): Started only when sandbox is configured for provisioner/Kubernetes mode
|
||
|
||
**Runtime**:
|
||
- `make dev`, Docker dev, and production all run the agent runtime in Gateway via `RunManager` + `run_agent()` + `StreamBridge` (`packages/harness/deerflow/runtime/`). Nginx exposes that runtime at `/api/langgraph/*` and rewrites it to Gateway's native `/api/*` routers.
|
||
|
||
**Project Structure**:
|
||
```
|
||
deer-flow/
|
||
??? Makefile # Root commands (check, install, dev, stop)
|
||
??? config.yaml # Main application configuration
|
||
??? extensions_config.json # MCP servers and skills configuration
|
||
??? backend/ # Backend application (this directory)
|
||
? ??? Makefile # Backend-only commands (dev, gateway, lint)
|
||
? ??? langgraph.json # LangGraph Studio graph configuration
|
||
? ??? packages/
|
||
? ? ??? harness/ # deerflow-harness package (import: deerflow.*)
|
||
? ? ??? pyproject.toml
|
||
? ? ??? deerflow/
|
||
? ? ??? agents/ # LangGraph agent system
|
||
? ? ? ??? lead_agent/ # Main agent (factory + system prompt)
|
||
? ? ? ??? middlewares/ # 10 middleware components
|
||
? ? ? ??? memory/ # Memory V2: provider/manager/builtin/hindsight
|
||
? ? ? ??? thread_state.py # ThreadState schema
|
||
? ? ??? sandbox/ # Sandbox execution system
|
||
? ? ? ??? local/ # Local filesystem provider
|
||
? ? ? ??? sandbox.py # Abstract Sandbox interface
|
||
? ? ? ??? tools.py # bash, ls, read/write/str_replace
|
||
? ? ? ??? middleware.py # Sandbox lifecycle management
|
||
? ? ??? subagents/ # Subagent delegation system
|
||
? ? ? ??? builtins/ # general-purpose, bash agents
|
||
? ? ? ??? executor.py # Background execution engine
|
||
? ? ? ??? registry.py # Agent registry
|
||
? ? ??? tools/builtins/ # Built-in tools (present_files, ask_clarification, view_image)
|
||
? ? ??? mcp/ # MCP integration (tools, cache, client)
|
||
? ? ??? models/ # Model factory with thinking/vision support
|
||
? ? ??? skills/ # Skills discovery, loading, parsing
|
||
? ? ??? config/ # Configuration system (app, model, sandbox, tool, etc.)
|
||
? ? ??? community/ # Community tools (tavily, jina_ai, firecrawl, image_search, aio_sandbox)
|
||
? ? ??? reflection/ # Dynamic module loading (resolve_variable, resolve_class)
|
||
? ? ??? utils/ # Utilities (network, readability)
|
||
? ? ??? client.py # Embedded Python client (DeerFlowClient)
|
||
? ??? app/ # Application layer (import: app.*)
|
||
? ? ??? gateway/ # FastAPI Gateway API
|
||
? ? ? ??? app.py # FastAPI application
|
||
? ? ? ??? routers/ # FastAPI route modules (models, mcp, memory, skills, uploads, threads, artifacts, agents, suggestions, channels)
|
||
? ? ??? channels/ # IM platform integrations
|
||
? ??? tests/ # Test suite
|
||
? ??? docs/ # Documentation
|
||
??? frontend/ # Next.js frontend application
|
||
??? skills/ # Agent skills directory
|
||
??? public/ # Public skills (committed)
|
||
??? custom/ # Custom skills (gitignored)
|
||
```
|
||
|
||
## Important Development Guidelines
|
||
|
||
### Documentation Update Policy
|
||
**CRITICAL: Always update README.md and CLAUDE.md after every code change**
|
||
|
||
When making code changes, you MUST update the relevant documentation:
|
||
- Update `README.md` for user-facing changes (features, setup, usage instructions)
|
||
- Update `CLAUDE.md` for development changes (architecture, commands, workflows, internal systems)
|
||
- Keep documentation synchronized with the codebase at all times
|
||
- Ensure accuracy and timeliness of all documentation
|
||
|
||
### Workflow Studio stream and canvas compatibility
|
||
|
||
- `app/gateway/workflow_agent_runner.py` is the boundary between LangGraph token chunks and durable workflow events. Provider reasoning in `AIMessage.additional_kwargs.reasoning_content`/`reasoning` is emitted to the UI as balanced inline `<think>` blocks; only visible answer content becomes the node's `text` output for downstream nodes. It must also emit every original `[AIMessage|ToolMessage, metadata]` tuple as `node.message.data.messages` with no whitelist, preview or length crop. Tool results and `additional_kwargs.clarification` are needed by the Studio's DeerFlow-compatible message UI and must survive durable SSE replay.
|
||
- An agent `ask_clarification` ToolMessage pauses the workflow at that agent node (`awaiting_input`), not as a terminal run. The pause records the source `nodeRunId`/`toolCallId`; resuming turns the submitted reply into the next user message on the same `wf-{run}-{node}` LangGraph thread. Keep the raw ToolMessage authoritative; do not invent a second card payload.
|
||
- `app/gateway/workflow_proposal_planner.py` converts the workflow controller's selected catalog roles into per-agent `mission`/`deliverable`/`scope`/`handoff` contracts. Only a selected visible `agentId` may contribute a contract; every omitted or malformed field falls back to a deterministic role-aware contract. Preserve this prompt-level worker boundary when adding strategies: researchers must not silently become final writers, and sequential roles must receive their upstream handoff binding plus an explicit next handoff.
|
||
- `app/gateway/routers/workflows_coze_compat.py` must return both `creator.self: true` and editable VCS metadata for an owned standalone workflow. Coze Playground derives its preview/read-only mode from those fields, so omitting `creator.self` disables node dragging even when the standalone page passes `readonly={false}`.
|
||
|
||
## Commands
|
||
|
||
**Root directory** (for full application):
|
||
```bash
|
||
make check # Check system requirements
|
||
make install # Install all dependencies (frontend + backend)
|
||
make dev # Start all services (Gateway + Frontend + Nginx), with config.yaml preflight
|
||
make start # Start production services locally
|
||
make stop # Stop all services
|
||
```
|
||
|
||
**Backend directory** (for backend development only):
|
||
```bash
|
||
make install # Install backend dependencies
|
||
make dev # Run Gateway API with reload (port 8001)
|
||
make gateway # Run Gateway API only (port 8001)
|
||
make test # Run all backend tests
|
||
make lint # Lint with ruff
|
||
make format # Format code with ruff
|
||
```
|
||
|
||
Regression tests related to Docker/provisioner behavior:
|
||
- `tests/test_docker_sandbox_mode_detection.py` (mode detection from `config.yaml`)
|
||
- `tests/test_provisioner_kubeconfig.py` (kubeconfig file/directory handling)
|
||
|
||
Boundary check (harness ? app import firewall):
|
||
- `tests/test_harness_boundary.py` ? ensures `packages/harness/deerflow/` never imports from `app.*`
|
||
|
||
CI runs these regression tests for every pull request via [.github/workflows/backend-unit-tests.yml](../.github/workflows/backend-unit-tests.yml).
|
||
|
||
## Architecture
|
||
|
||
### Position roundtable recovery persistence
|
||
|
||
`/api/position-roundtable` uses the shared `/api/intent/*` LangGraph flow as
|
||
the live Step-1 execution source and stores an additional SQL recovery copy in
|
||
`position_roundtable_sessions`/`position_roundtable_nodes`. The recovery copy
|
||
contains complete intent/node/summary/action-plan conversations, unsent
|
||
composer drafts, workflow snapshots, and bounded full text for generated
|
||
Markdown/JSON artifacts. This lets history and downstream closing agents keep
|
||
working when container-local checkpoint or sandbox volumes are replaced.
|
||
Revision `20260723_05` adds the conversation columns. Session deletion
|
||
explicitly removes child nodes so SQLite deployments do not depend on FK
|
||
cascade pragmas. The router is mounted canonically below
|
||
`/api/position-roundtable` and, for intranet gateway allowlists, below the
|
||
schema-hidden compatibility prefix `/api/multi-agent/position-roundtable`.
|
||
Clients may retry the compatibility route only for a framework/proxy-level
|
||
`404 Not Found`; domain-level missing-session responses must remain real 404s.
|
||
|
||
The position-roundtable state machine is multi-worker safe: node commands
|
||
(`complete`, `bind`, conversation save, reject) and session commands
|
||
(`activate`, archive, delete) lock the parent session and affected nodes in one
|
||
transaction, then record a unique command id for retry replay. The task and
|
||
intent snapshot writes use the same parent-row lock; they are permitted only
|
||
while `status=intent_pending`, so a late “confirm intent” write cannot change a
|
||
chain that another user has already activated. Activation itself rechecks the
|
||
stored confirmed intent after taking the lock. `tests/test_position_roundtable_
|
||
commands.py` and `tests/test_position_roundtable_concurrency.py` cover the
|
||
state transitions; `tests/test_roundtable_postgres_concurrency.py` is the real
|
||
PostgreSQL acceptance suite (set `DEERFLOW_POSTGRES_CONCURRENCY_TEST_URL`).
|
||
|
||
Roundtable backend jobs use a database unique active-dedupe key plus a leased
|
||
dispatcher, never process-local start deduplication. Cancellation is two
|
||
phase (`cancel_requested` then `cancelled`): the lease owner’s watcher normally
|
||
finalizes it, while the dispatcher reaps a cancellation whose lease has
|
||
expired after a worker crash. A worker that loses lease stops its local model
|
||
run immediately. Do not reintroduce direct `executor.cancel()` from the HTTP
|
||
cancel route, because it can cancel the watcher before the database state is
|
||
finalized. The boolean defaults on these new rows must use SQLAlchemy
|
||
`false()` rather than numeric `0`, so fresh PostgreSQL schemas can be created.
|
||
|
||
Position roundtable selects the separately seeded built-in agent
|
||
`position-roundtable-intent` by passing `intent_agent_id` to the shared
|
||
`/api/intent/init` and `/api/intent/stream` endpoints. Its request lifecycle,
|
||
files, and SSE protocol match `roundtable-intent`, but its cap is five
|
||
**user-assistance rounds**. A single round may emit several independent
|
||
`ask_clarification` tool calls; those calls are separate cards, not one card
|
||
containing several numbered questions. Count that as one round by the
|
||
tool-calling AI message, never by card count. Topic-only requests must request
|
||
all material missing decisions as separate cards rather than being treated as a
|
||
resolved intent; after the fifth set is answered it must emit completed final
|
||
JSON without unresolved placeholders. The required arrays are
|
||
`coreGoals`, `riskWarnings`, `keyPoints`, and `strategicSignificance`. Do not
|
||
treat a fresh dedicated-agent `ask_clarification` as stale: it must win over
|
||
a same-run premature `[INTENT_READY]` and leave the UI waiting for the user.
|
||
Conversely, the prior turn’s historical clarification must not block the user
|
||
reply from completing. Do not
|
||
change `roundtable-intent`'s `{objective,constraints,assumptions}` contract for
|
||
this feature. The position router accepts the new schema for activation and
|
||
injects all four sections into downstream position-node and closing-agent
|
||
context; legacy stored summaries are mapped read-only for compatibility.
|
||
`/api/intent/init` must create the empty Step-1 thread through
|
||
`multi_agent._create_thread_direct`, exactly like multi-agent initialization;
|
||
do not restore the old HTTP loopback to `/api/threads`, because deployed
|
||
Gateway ports/proxies may make the automatic page initialization fail while
|
||
later manual turns still work. If checkpoint creation fails or the parallel seat
|
||
task is cancelled, the helper compensates by deleting any partial checkpoint and
|
||
metadata row (the cancellation branch runs under `asyncio.shield` so a second
|
||
cancellation cannot abort it half-way). Streaming stays on the existing loopback
|
||
path.
|
||
|
||
`ThreadMetaStore.create` never issues a post-commit `session.refresh()`. Every
|
||
`ThreadMetaRow` column is assigned before commit and the session factory uses
|
||
`expire_on_commit=False`, so the returned dict is built from the in-memory row.
|
||
The removed read-back ran on a *different* pooled connection than the INSERT,
|
||
which a MySQL read/write-splitting endpoint can route to a lagging replica — the
|
||
row comes back empty and SQLAlchemy raises `Could not refresh instance`, surfacing
|
||
as `500 "Failed to create thread"`. Concurrency makes that far more likely, so a
|
||
single-user deployment never sees it. No caller uses the return value, and all
|
||
seven call sites (`multi_agent._create_thread_direct`, `services.start_run`,
|
||
`routers/threads.create_thread`, `roundtable_inprocess_gateway`,
|
||
`scheduler/service`, `thread_shares`, `authz`) benefit. Keep new `create`
|
||
implementations read-back-free; `tests/test_position_roundtable_sessions.py::
|
||
test_thread_meta_create_never_reads_back_after_commit` guards this.
|
||
|
||
### Harness / App Split
|
||
|
||
The backend is split into two layers with a strict dependency direction:
|
||
|
||
- **Harness** (`packages/harness/deerflow/`): Publishable agent framework package (`deerflow-harness`). Import prefix: `deerflow.*`. Contains agent orchestration, tools, sandbox, models, MCP, skills, config ? everything needed to build and run agents.
|
||
- **App** (`app/`): Unpublished application code. Import prefix: `app.*`. Contains the FastAPI Gateway API and IM channel integrations (Feishu, Slack, Telegram, DingTalk).
|
||
|
||
**Dependency rule**: App imports deerflow, but deerflow never imports app. This boundary is enforced by `tests/test_harness_boundary.py` which runs in CI.
|
||
|
||
### Report Collaboration (AgentScope, in progress)
|
||
|
||
Independent of Workflow Studio and the existing roundtable. Plan: repo-root `docs/agentscope-多智能体报告协作工作台-后端实施计划.md`(`RC-BE-000`~`018` 已落地,下一票 `RC-BE-019`)。Spike baseline: `docs/AGENTSCOPE_BASELINE.md`.
|
||
|
||
- **Config**: `config.yaml → report_collaboration` (`deerflow.config.report_collaboration_config.ReportCollaborationConfig`). Default **off** (`enabled` / `worker_enabled`).
|
||
- **Contracts**: `app/report_collaboration/contracts/` — snake_case Pydantic models matching `frontend-web/src/report-collaboration/api/types.ts`. SSE envelope uses camelCase (`eventId` / `sessionId` / `runId`). `node_run_id` is always `${node_id}-attempt{attempt}`.
|
||
- **Persistence**: `deerflow.persistence.report_collaboration` (`report_collaboration_*` tables, migrations `20260905_01` / `20260906_01` / `20260907_01`). SQL + memory stores; one active run per session via `active_dedupe_key`; command/message/run/version idempotency keys refuse cross-session reuse; event `seq` taken from the run row. Session delete removes children explicitly.
|
||
- **Gateway**: `/api/report-collaboration` (`app/gateway/routers/report_collaboration.py`). `enabled=false` → 503. Plan/run writes enqueue durable jobs and return immediately. Pre-run `POST /sessions/{id}/messages` runs a rule-based `IntentRouterAgent` + `RequirementResolver` (`app/report_collaboration/requirement/`): writes a requirement revision and clarification messages, never auto-starts planning. `POST /sessions/{id}/plan-requests` then runs `PlanProposalService` (`app/report_collaboration/planning/`): catalog projection (no SOUL/keys) → selector → 2–3 strategy RoleDemand graphs → assembler binds catalog agents/tools/quality gates → validator (acyclic DAG, artifact types, allowlisted tools). Select and start-run stay independent commands. `POST /runs` freezes role models, report-structure template and run budget. Rewrite / apply / restore are live (`reporting/`). `GET /runs/{id}/stream` is a read-only durable SSE tail (`execution/live_hub.py`): persist then publish, replay + `after_seq` / `Last-Event-ID`, heartbeat comments, token-delta coalesce before seq assignment. Disconnect does **not** cancel the run. Lifespan mounts `ReportCollaborationDispatcher` only when `enabled` **and** `worker_enabled` are both true; default both false.
|
||
- **Runtime**: AgentScope adapters live in `app/report_collaboration/agentscope_runtime/` (not the publishable harness). `ConversationReplica` maps AgentScope 2.0.5 events 1:1 onto collaboration SSE types; TeamSay is `HINT_BLOCK`. `DeerFlowChatModelAdapter` wraps `create_chat_model` (no second Provider config); role models are frozen onto node `agent_snapshot` at `POST /runs` and missing/incapable models fail with 422 before the run is created. Fallback is recorded on the snapshot plus a durable `heartbeat` audit event. Nodes complete only after a validated artifact (`app/report_collaboration/execution/completion.py`). `TaskLedger` (`execution/ledger.py`) owns ready waves, attempts and terminal node status; AgentScope / coordinator suggestions never pick the next node. A writer-bearing run is `completed` only after `ReportFinalizer` publishes a formal version. `WaveExecutor` + `QualityRoleKernel` (`execution/quality_roles.py`) remain for fixtures: Verifier marks stale/contested claims, Analyst/Writer only cite supported IDs, Reviewer `revise` replans only the target write node plus review. Production worker injects `AgentScopeMemberKernel` (`agentscope_runtime/member_kernel.py`) with DeerFlow tools and per-run budget. `QualityGate` (`execution/quality_gate.py`) is the claim-level gate: conflicts/missing units, unverified facts, coverage gaps, and unmapped review targets fail before `validated`; `ArtifactBoard.supersede_previous` hides old attempts from downstream. `ReportCollaborationLiveHub` (`execution/live_hub.py`) fans out persisted envelopes to SSE viewers; `PublishingReportCollaborationStore` sanitizes secrets/privacy/hidden reasoning without truncating visible tool results, and writes redaction + event audits. Dispatcher (`execution/dispatcher.py`) atomically claims queued/recovering/expired-lease runs, heartbeats via `renew_lease`, two-phase cancels (`cancel_requested` → `finalize_cancel`), and recovers orphans (`execution/recovery.py`): keep completed work, mark leftover in-flight `interrupted`/`superseded`, open a new `planned` attempt. Late artifacts after cancel/reclaim are `superseded` and must not change terminal run status. Tests drive `dispatch_once()`; do not rely on the background loop. Per-run budgets (`security/budget.py`) fail-closed; per-user active-run caps are enforced at `create_run`.
|
||
- **Quality eval**: `app/report_collaboration/eval/` (`RC-BE-018`) — 20 annotated tasks (policy/market/enterprise/event/comparison ×4) with required angles, key facts, forbidden fabrications, templates and review rules. Deterministic scorer (no LLM judge) gates citation/angle/template/directed-mod; gold reports must pass, canned failures (including a roundtable-style essay) must fail and stay reproducible. Recommended runtime knobs live in `eval/thresholds.py` and match current `ReportCollaborationConfig` defaults (`worker_enabled` remains false). Tests: `tests/test_report_collaboration_eval.py`.
|
||
- **Runtime intervention**: `app/report_collaboration/interventions/` owns run-time natural-language commands. `RuntimeIntentRouterAgent` uses explicit rules first and a tool-free coordinator-model fallback only for ambiguity; its Pydantic decision cannot dispatch or mutate anything. `ImpactAnalyzer` validates server-side node/agent ownership, resolves angle-specific lineage, and enforces confidence/cost/parallel-branch confirmation (`intent_confirmation_threshold`, default `0.8`). `RuntimeCommandService` mirrors the human message, emits frozen command events, CAS-claims accepted commands, and applies them only through `TaskLedger` safe points. Rework persists command overlays, supersedes affected attempts/artifacts, creates monotonic attempts, and prevents late output from entering downstream. Waiting-node answers resume only those nodes. `rerun_node` without a legal target does not expand to the whole graph; requirement patches apply only after rework succeeds; stuck `executing` commands are reclaimed on takeover or lease timeout; `command.cancelled` is a durable SSE event. In-run `POST /sessions/{id}/messages` returns 409 (`RUN_ACTIVE`). Formal rewrite candidates and version apply/restore live in `reporting/` (`RC-BE-015`).
|
||
- **Dependency**: optional extra `report-collaboration` → `agentscope==2.0.5`. Do not add AgentScope to the harness `pyproject.toml`.
|
||
|
||
Tests: `tests/test_report_collaboration_*.py`.
|
||
|
||
### Workflow Studio
|
||
|
||
Coze Studio is canvas-only; ChatDev is reference-only. Plan + status per phase: `docs/WORKFLOW_STUDIO_BACKEND_DEV_ZH.md`; decisions: `docs/adr/0001-workflow-studio-runtime.md`; contracts: `docs/WORKFLOW_SCHEMA_V1_ZH.md` + `docs/WORKFLOW_SSE_V1_ZH.md` + `docs/WORKFLOW_OPENAPI_DRAFT_V1.yaml`. Master switch + all limits: `config.yaml` → `workflows` (`deerflow.config.workflow_config.WorkflowConfig`). `workflows.agent_recursion_limit` defaults to 250: LangGraph counts each model→tool research exchange as about two super-steps, so workflow agent/skill nodes must not fall back to a chat-sized 60-step limit. In the local Studio proxy, keep Rsbuild `server.compress: false`; dev gzip buffers small SSE frames and makes live planning/run output arrive only at completion despite the backend's `no-transform` / `X-Accel-Buffering: no` headers. The composer’s optional `modelName` must be validated against `AppConfig.models`: it selects the controller model and is copied only into newly generated candidate agent configs (or the projected direct-agent task), never injected into an already-confirmed proposal/run.
|
||
|
||
**Execution kernel is a custom DAG scheduler, not a compiled LangGraph.** LangGraph stays the kernel *inside* an `agent`/`skill` node. The reason is single-source-of-truth recovery: a re-leased or resumed run replays completed node outputs from `workflow_node_runs` plus the `workflow_run_events` log, so a second (checkpoint) copy of workflow-level state would have no defensible tie-breaker. See the ADR's phase-2 amendment.
|
||
|
||
**Harness** (`packages/harness/deerflow/workflows/`, never imports `app.*`):
|
||
- `schemas` / `events` / `errors` / `ports` / `validator` / `samples` — frozen v1 contracts.
|
||
- `expressions.py` — `{{ inputs.x }}` / `{{ nodes.n.data.y }}` templates and conditions resolved by path walking + a small comparator set. **No `eval`.** Roots are restricted to `inputs` / `nodes` / `run` / `loop` / `env`.
|
||
- `runtime/engine.py` — the scheduler. Every edge is `pending` → `active` (source chose this port) or `pruned`; a node is ready when its non-back incoming edges are all resolved and at least one is `active`, and is skipped when all are `pruned` (skips propagate). That one rule covers fan-out, condition branching, and merge joins with no topological pre-pass. The `loop` node's `continue` back edge is the only permitted cycle; re-entry clears the body's results and invalidates its node runs.
|
||
- `runtime/{sink,sql_runner,code_runner,context}.py`, `nodes/*` (15 executors, including deterministic `evidence_normalizer` and durable `deep_research_write`), `security/{http_policy,sql_policy,secret_box}.py`.
|
||
|
||
**Persistence**: `workflow_definitions` / `workflow_versions` (migration `20260828_01`), `workflow_runs` / `workflow_node_runs` / `workflow_run_events` / `workflow_run_artifacts` (`20260828_02`), `workflow_data_sources` (`20260828_03`), and short-lived-but-durable `workflow_planning_sessions` / `workflow_planning_proposals` (`20260831_01`). Stores under `deerflow.persistence.workflow_{runs,events,data_sources,planning}` with SQL + in-memory implementations. Because planning proposal rows have a scalar FK rather than an ORM relationship, `SqlWorkflowPlanningStore.create_session` must flush the parent session before adding its children; otherwise SQLite can insert the child first and return a 500. A planning proposal is distinct from a run: its graph can be edited with optimistic `graph_revision` checks, but it becomes executable only once selected and atomically bound to one run.
|
||
|
||
**App layer**: `workflow_dispatcher.py` (lease scan → atomic claim; an expired lease is how a dead worker's run is reclaimed, so every worker can run the loop), `workflow_executor.py` (one lease's lifecycle: replay → engine under a run timeout with heartbeat + cancel watcher → exactly one terminal transition; also runs subworkflows inline, depth ≤ 3), `workflow_retention.py` (hourly sweep of `workflows.retention.event_retention_days` / `run_retention_days`), `workflow_live_hub.py` (in-process SSE fan-out), `workflow_agent_runner.py` (agent/skill nodes → lead agent; `disable_tools=True` means an empty bound-tool set for control-plane callers), `workflow_deep_research_adapter.py` (Evidence Pack → selected Deep Research sources/job/event projection; no internal HTTP call), and `workflow_proposal_planner.py` (calls the dedicated `workflow-planner` controller over visible-agent metadata, then constructs safe candidate graphs server-side). `workflow-planner` is a separate Workflow Studio control-plane agent; never reuse or alias `roundtable-coordinator` for workflow planning, and never let it dispatch a run, call research tools, or return arbitrary graph JSON. It always runs with tools and model thinking disabled.
|
||
|
||
**Gateway routes** (register static prefixes before `workflows.router`'s `/{workflow_id}` catch-alls): `/api/workflows` CRUD/draft/validate/publish/copy, `/api/workflows/node-types` + `/resources/*`, `/api/workflows/{id}/planning-sessions{/stream}` + `/api/workflows/planning-sessions/{session_id}{,/proposals/{proposal_id},/proposals/{proposal_id}/graph}`, `/api/workflows/{id}/runs` + `/api/workflows/runs/{run_id}{,/events,/stream,/cancel,/resume,/retry,/feedback,/artifacts}`, `/api/workflows/data-sources` (+ `/{id}/introspect`, `/{id}/validate-query`), `/api/workflows/embed/{tickets,verify}`, `/api/workflow_api/*` Coze compat. The planning `/stream` endpoint sends product-safe real controller milestones (`catalog`, `controller`, `selection`, `assembly`, `validation`, `complete`) then a final `result` session; never stream agent catalog descriptions or raw planner JSON. Formal runs accept the selected `planningSessionId` and `proposalId`; creating the same pair again returns the already-bound run rather than launching another execution.
|
||
|
||
**Guarantees worth not breaking**: events are persisted *then* published, and the DB is authoritative for `seq` ordering (live frames at or below the replay cursor are dropped and re-read on the next poll, so a reconnect sees no gap; `Last-Event-ID` wins over `?after=`); a run reaching a terminal status always gets a terminal event, synthesised if necessary, or SSE clients hang; `resumeToken` appears only in the `run.awaiting_input` event, never in the run view; `POST /retry` on a `failed` run copies completed node outputs onto a new run marked `retry_of_run_id` so the engine resumes from the failed node; `POST /feedback` never mutates its source run—on a terminal acyclic run it may target only a completed `agent`/`skill`, reuses completed nodes outside that node's downstream closure, records those `reusedNodeIds` alongside the feedback summary, and puts the bounded feedback only in `RunContext.env.nodeFeedback[target]` before creating a revision run; data-source secrets are write-only (encrypted via `WORKFLOW_SECRET_KEY`, decrypted only inside the executor, reads expose `masked_target`); the `http` node denies private/loopback/link-local targets by default (SSRF), `sql_read` allows one read-only statement with bound parameters, and the `code` node is **off by default** because its runner is a hardened subprocess rather than a container.
|
||
|
||
HTTP interface resources persist a non-secret base_url plus an allowed_methods allowlist (migration
|
||
20260829_01); request headers stay only in the encrypted secret blob. The resource catalog may return
|
||
the address and methods to the owning user, but must never return headers or any decrypted secret. The
|
||
executor receives HTTP metadata only through its data-source resolver.
|
||
|
||
The Workflow Studio agent/skill resource catalog is a read-only display projection over the existing
|
||
agent and skill ownership stores. It returns stable `agentId` / `skillId`, an agent's cached skill names,
|
||
and a `createdBy` display label. Resolve human labels from the immutable owner id at request time and
|
||
fall back to that id if lookup fails; built-in records must return `系统内置`. Do not create a parallel
|
||
creator table or accept creator data from the browser.
|
||
|
||
`evidence_normalizer` is a deterministic data node, not an LLM: it consumes only configured/upstream
|
||
node results and emits a bounded `evidencePack` of claims with node provenance, role labels and gaps.
|
||
The planner inserts it between parallel research roles and the synthesis Agent. It must not treat
|
||
source text as verified fact or fetch any external resource, and its explicit `sources[].nodeId` values
|
||
must stay upstream-only under `validate_workflow_graph`.
|
||
|
||
`deep_research_write` is a durable report node, not an Agent prompt shortcut. It accepts only an
|
||
Evidence Pack, derives a session identity from the workflow run/node plus the topic, evidence snapshot and writing config, writes
|
||
the pack as selected `other` source rows with upstream-node provenance, and queues/reuses the existing
|
||
Deep Research `regenerate` job. Its live `report_delta` and persisted `report_chunk` frames are
|
||
deduplicated before becoming workflow `node.output.delta` events; only report prose is forwarded (never
|
||
planning/reasoning deltas). Workflow cancellation requests cancellation on that same job. Do not make it
|
||
self-call `/api/deep-research/*`, use Deep Research `multi_agent` inside this node, or treat Evidence
|
||
Pack claims as verified facts. The parallel-research candidate grants the node an explicit 1800-second
|
||
timeout; deployments that override `workflows.node_timeout_seconds` must keep it at least that high.
|
||
|
||
Tests: `tests/test_workflow_schemas.py`, `tests/test_workflow_definitions.py`, `tests/test_workflow_engine.py`, `tests/test_workflow_runs.py`.
|
||
|
||
**Import conventions**:
|
||
```python
|
||
# Harness internal
|
||
from deerflow.agents import make_lead_agent
|
||
from deerflow.models import create_chat_model
|
||
|
||
# App internal
|
||
from app.gateway.app import app
|
||
from app.channels.service import start_channel_service
|
||
|
||
# App ? Harness (allowed)
|
||
from deerflow.config import get_app_config
|
||
|
||
# Harness ? App (FORBIDDEN ? enforced by test_harness_boundary.py)
|
||
# from app.gateway.routers.uploads import ... # ? will fail CI
|
||
```
|
||
|
||
### Agent System
|
||
|
||
**Lead Agent** (`packages/harness/deerflow/agents/lead_agent/agent.py`):
|
||
- Entry point: `make_lead_agent(config: RunnableConfig)` registered in `langgraph.json`
|
||
- Dynamic model selection via `create_chat_model()` with thinking/vision support
|
||
- Tools loaded via `get_available_tools()` - combines sandbox, built-in, MCP, community, and subagent tools
|
||
- System prompt generated by `apply_prompt_template()` with skills, memory, and subagent instructions
|
||
|
||
**agentfx built-in analyst + report convert skills.** `agentfx-analyst` (「任务研判报告助手」) is seeded from `app/gateway/routers/_agentfx_seed_assets/` before `_sync_legacy_agents`. After mandatory `web_search`, the agent writes each tool JSON to `/mnt/user-data/workspace/hits/qN.json` and runs **one** command — `task-report-build` (group `all`) — to fill the 14-category `task-reports.json` for the dashboard slots; the four per-page skills `task-report-{enemy,our,env,judge}` remain for re-running a single dashboard page. All five are thin CLIs over the shared stdlib engine `skills/public/task-report-lib/convert_lib.py` (no SKILL.md on the lib, so it never enters the skill list). The model must not `write_file` that JSON. Because the sandbox maps `/mnt/user-data/...` onto a host dir, a virtual absolute hits path can be `isdir`-true yet list empty on a Windows host — `resolve_hit_files` therefore tries the path as given, then the cwd-mapped sandbox path (`discover_user_data_root` + `resolve_sandbox_path` — bash `cd`s into the thread `workspace`, so `/mnt/user-data/outputs/task-reports.json` must land in that thread's `outputs/`, never `F:\mnt\...` on a Windows host), then the virtual-prefix-stripped relative form, then the bare dir name, and picks the first candidate that actually contains `.json`; a miss returns a `warnings` entry listing every tried form so the agent never has to write a debug script. Assemble of the 14 categories is the same function (`assemble_records`); the model must not copy or merge files. If the model is interrupted after convert, `task-reports.json` is already in outputs and the frontend import card also scans `GET /api/threads/{id}/artifacts` so入库 still works without `present_files`. Import remains `task-report-import`. Existing agent/skill directories are never overwritten. Tests: `tests/test_task_report_convert.py`, `tests/test_agentfx_seed.py`.
|
||
|
||
**Forced-research built-in agent.** `forced-research-responder` (「强制检索输出助手」) is seeded from `app/gateway/routers/_forced_research_seed_assets/` before `_sync_legacy_agents`. `ForcedResearchMiddleware` is mounted only for this stable id. On every visible user turn it requires a successful information-gathering tool result before the model may finish; `tool_choice=required` is backed by an `after_model -> model` retry for gateways that ignore tool choice, with an eight-attempt failure cap that permits only an honest failure response. Skill discovery, reading `SKILL.md`, and writing a helper script do not satisfy the gate; web/knowledge/MCP results, non-skill uploaded-file reads, or successful Python skill-script execution do. Tool filtering removes `present_files`, `skill_manage`, task/schedule/memory/clarification surfaces. The middleware permits `write_file`/`str_replace` only for `/mnt/user-data/workspace/*.py` and permits `bash` only for `.py` execution without redirection/document paths, so skills can run scripts but the agent cannot create Markdown/document artifacts or write `outputs`. Tests: `tests/test_forced_research_agent.py`.
|
||
|
||
**ThreadState** (`packages/harness/deerflow/agents/thread_state.py`):
|
||
- Extends `AgentState` with: `sandbox`, `thread_data`, `title`, `artifacts`, `todos`, `uploaded_files`, `viewed_images`
|
||
- Uses custom reducers: `merge_artifacts` (deduplicate), `merge_viewed_images` (merge/clear)
|
||
|
||
**Runtime Configuration** (via `config.configurable`):
|
||
- `thinking_enabled` - Enable model's extended thinking
|
||
- `model_name` - Select specific LLM model
|
||
- `is_plan_mode` - Enable TodoList middleware
|
||
- `subagent_enabled` - Enable task delegation tool
|
||
|
||
### Middleware Chain
|
||
|
||
Lead-agent middlewares are assembled in strict append order across `packages/harness/deerflow/agents/middlewares/tool_error_handling_middleware.py` (`build_lead_runtime_middlewares`) and `packages/harness/deerflow/agents/lead_agent/agent.py` (`_build_middlewares`):
|
||
|
||
1. **ThreadDataMiddleware** - Creates per-thread directories under the creator's stable isolation scope (`backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/{workspace,uploads,outputs}`); resolves persisted threads through `resolve_path_user_id(thread_id)` (falls back to `get_effective_user_id()` / `"default"` when metadata is unavailable); Web UI thread deletion now follows LangGraph thread removal with Gateway cleanup of the local thread directory
|
||
2. **UploadsMiddleware** - Tracks and injects newly uploaded files into conversation
|
||
3. **SandboxMiddleware** - Acquires sandbox, stores `sandbox_id` in state
|
||
4. **DanglingToolCallMiddleware** - Injects placeholder ToolMessages for AIMessage tool_calls that lack responses (e.g., due to user interruption), including raw provider tool-call payloads preserved only in `additional_kwargs["tool_calls"]`
|
||
5. **LLMErrorHandlingMiddleware** - Normalizes provider/model invocation failures into recoverable assistant-facing errors before later middleware/tool stages run
|
||
6. **GuardrailMiddleware** - Pre-tool-call authorization via pluggable `GuardrailProvider` protocol (optional, if `guardrails.enabled` in config). Evaluates each tool call and returns error ToolMessage on deny. Three provider options: built-in `AllowlistProvider` (zero deps), OAP policy providers (e.g. `aport-agent-guardrails`), or custom providers. See [docs/GUARDRAILS.md](docs/GUARDRAILS.md) for setup, usage, and how to implement a provider.
|
||
7. **SandboxAuditMiddleware** - Audits sandboxed shell/file operations for security logging before tool execution continues
|
||
8. **ToolErrorHandlingMiddleware** - Converts tool exceptions into error `ToolMessage`s so the run can continue instead of aborting
|
||
9. **SummarizationMiddleware** (`DeerFlowSummarizationMiddleware`) - Context reduction when approaching token limits (optional, if enabled). **Non-destructive**: overrides the stock `before_model` (which persists a `RemoveMessage(REMOVE_ALL_MESSAGES)`) to a no-op and instead compresses **transiently** in `wrap_model_call` ? the model request is replaced with `[summary, *recent]` while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). The `summary` message is `hide_from_ui` and never persisted; summaries are cached **per-thread in a process-wide, thread-safe LRU** (`_GLOBAL_SUMMARY_CACHE`, bounded to 512 threads, sticky boundary). The cache is module-level — **not** on the middleware instance — because `make_lead_agent` rebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed by `thread_id` (resolved like `ThreadDataMiddleware`), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. **Threshold-first gating**: `_plan_compression` returns passthrough (no compression at all) whenever the *full* current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. **Stickiness headroom** (`_apply_stickiness_headroom`, `_KEEP_HEADROOM_FRACTION=0.75`): after the `keep` policy picks a cutoff, the kept suffix is forced below ~75% of *every* trigger before summarizing. Without it a `keep` in *messages* (e.g. 20) can itself exceed a *token* trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a `{"type":"context_compacting","message":"正在压缩上下文…"}` event on LangGraph's `custom` stream channel (the cheap reuse path stays silent); the chat frontend's `onCustomEvent` renders it as a transient toast. Tests: `tests/test_summarization_middleware.py`
|
||
10. **TodoListMiddleware** - Task tracking with `write_todos` tool (optional, if plan_mode)
|
||
11. **TokenUsageMiddleware** - Records token usage metrics when token tracking is enabled (optional)
|
||
12. **TitleMiddleware** - Two-phase auto-title (never shows "untitled"): `before_model` sets an instant question-derived *provisional* title the moment the first user message arrives (no LLM, zero latency), then `after_model` asynchronously generates the final LLM title and replaces it (tracked via the `title_provisional` state flag). Normalizes structured message content before prompting the title model
|
||
13. **MemoryMiddleware** - Memory V2: `wrap_model_call` injects Hindsight recall transiently; `after_agent` auto-retains the turn to Hindsight
|
||
14. **ViewImageMiddleware** - Injects base64 image data before LLM call (conditional on vision support)
|
||
15. **DeferredToolFilterMiddleware** - Hides deferred tool schemas from the bound model until tool search is enabled (optional)
|
||
16. **SubagentLimitMiddleware** - Truncates excess `task` tool calls from model response to enforce `MAX_CONCURRENT_SUBAGENTS` limit (optional, if `subagent_enabled`)
|
||
17. **LoopDetectionMiddleware** - Detects repeated tool-call loops; hard-stop responses clear both structured `tool_calls` and raw provider tool-call metadata before forcing a final text answer
|
||
18. **ClarificationMiddleware** - Intercepts `ask_clarification` tool calls, interrupts via `Command(goto=END)` (must be last)
|
||
8a. **ToolMetricsMiddleware** - Appended immediately after `ToolErrorHandlingMiddleware` in `_build_runtime_middlewares`, so it sits *deepest* in the `wrap_tool_call` chain. Records one `tool_call_metrics` row per tool/skill/MCP invocation ? timing duration and bucketing failures into a coarse `error_category` (rate_limited / ip_blocked / auth_error / timeout / network_error / not_found / empty_result / tool_error). It catches both raised exceptions (before #8 converts them) and *self-handled* errors returned as error `ToolMessage`s (search tools that swallow HTTP 429/403). It also captures the targeted Agent Skill into `tool_call_metrics.skill_name`: the lead agent invokes a skill by `read_file`-ing its `/mnt/skills/<category>/<name>/SKILL.md` path (per the `<skill_system>` prompt block ? there is no dedicated "use skill" tool), so `skill_name` is recovered by scanning every tool call's string arguments for the `skills/{public,custom}/{name}/` path shape (`_SKILL_PATH_RE`), plus the explicit `name` argument of `skill_view`/`skill_manage`/`skill_archive`. `skill_view`-specific success detection handles its swallowed failures (`Skill 'x' not found.` returned as a plain string rather than raised). Regression coverage: `tests/test_tool_metrics_middleware.py`. Writes are best-effort fire-and-forget. **Per-run skill merging**: `SqlToolMetricsStore.record` upserts skill rows by `(run_id, skill_name)` instead of inserting one row per call ? a skill touched by several tool calls / retries inside one run collapses to a single row. Any success in the run wins (a skill that failed then finally succeeded is recorded as one `success`); a skill that only ever failed yields one `error` row whose `error_message` accumulates every attempt's message; durations are summed. Non-skill tool calls, and skill calls with no `run_id`, keep one row each. The read?merge?write is guarded by a per-event-loop `asyncio.Lock`. Surfaced by the admin `GET /api/admin/tool-metrics` router (which has a `kind=skill` filter and a per-skill `/summary` breakdown).
|
||
9. **PromptPrefixMiddleware** - Injects the admin-configured prompt prefix into the model request only. The prefix is tagged onto the first human message's `additional_kwargs['prompt_prefix']` by the Gateway (`app.gateway.services`) and prepended to content here at call time, so the persisted/displayed message stays prefix-free
|
||
10. **SummarizationMiddleware** (`DeerFlowSummarizationMiddleware`) - Context reduction when approaching token limits (optional, if enabled). **Non-destructive**: overrides the stock `before_model` (which persists a `RemoveMessage(REMOVE_ALL_MESSAGES)`) to a no-op and instead compresses **transiently** in `wrap_model_call` ? the model request is replaced with `[summary, *recent]` while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). The `summary` message is `hide_from_ui` and never persisted; summaries are cached **per-thread in a process-wide, thread-safe LRU** (`_GLOBAL_SUMMARY_CACHE`, bounded to 512 threads, sticky boundary). The cache is module-level — **not** on the middleware instance — because `make_lead_agent` rebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed by `thread_id` (resolved like `ThreadDataMiddleware`), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. **Threshold-first gating**: `_plan_compression` returns passthrough (no compression at all) whenever the *full* current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. **Stickiness headroom** (`_apply_stickiness_headroom`, `_KEEP_HEADROOM_FRACTION=0.75`): after the `keep` policy picks a cutoff, the kept suffix is forced below ~75% of *every* trigger before summarizing. Without it a `keep` in *messages* (e.g. 20) can itself exceed a *token* trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a `{"type":"context_compacting","message":"正在压缩上下文…"}` event on LangGraph's `custom` stream channel (the cheap reuse path stays silent); the chat frontend's `onCustomEvent` renders it as a transient toast. Tests: `tests/test_summarization_middleware.py`
|
||
11. **TodoListMiddleware** - Task tracking with `write_todos` tool (optional, if plan_mode)
|
||
12. **TokenUsageMiddleware** - Records token usage metrics when token tracking is enabled (optional)
|
||
13. **TitleMiddleware** - Two-phase auto-title (never shows "untitled"): `before_model` sets an instant question-derived *provisional* title the moment the first user message arrives (no LLM, zero latency), then `after_model` asynchronously generates the final LLM title and replaces it (tracked via the `title_provisional` state flag). Normalizes structured message content before prompting the title model
|
||
14. **MemoryMiddleware** - Queues conversations for async memory update (filters to user + final AI responses)
|
||
15. **ViewImageMiddleware** - Injects base64 image data before LLM call (conditional on vision support)
|
||
16. **DeferredToolFilterMiddleware** - Hides deferred tool schemas from the bound model until tool search is enabled (optional)
|
||
17. **SubagentLimitMiddleware** - Truncates excess `task` tool calls from model response to enforce `MAX_CONCURRENT_SUBAGENTS` limit (optional, if `subagent_enabled`)
|
||
18. **LoopDetectionMiddleware** - Detects repeated tool-call loops; hard-stop responses clear both structured `tool_calls` and raw provider tool-call metadata before forcing a final text answer
|
||
19. **ClarificationMiddleware** - Intercepts `ask_clarification` tool calls, interrupts via `Command(goto=END)` (must be last)
|
||
|
||
### Configuration System
|
||
|
||
**Main Configuration** (`config.yaml`):
|
||
|
||
Setup: Copy `config.example.yaml` to `config.yaml` in the **project root** directory.
|
||
|
||
**Config Versioning**: `config.example.yaml` has a `config_version` field. On startup, `AppConfig.from_file()` compares user version vs example version and emits a warning if outdated. Missing `config_version` = version 0. Run `make config-upgrade` to auto-merge missing fields. When changing the config schema, bump `config_version` in `config.example.yaml`.
|
||
|
||
**Config Caching**: `get_app_config()` caches the parsed config, but automatically reloads it when the resolved config path changes or the file's mtime increases. This keeps Gateway and LangGraph reads aligned with `config.yaml` edits without requiring a manual process restart.
|
||
|
||
Configuration priority:
|
||
1. Explicit `config_path` argument
|
||
2. `DEER_FLOW_CONFIG_PATH` environment variable
|
||
3. `config.yaml` in current directory (backend/)
|
||
4. `config.yaml` in parent directory (project root - **recommended location**)
|
||
|
||
Config values starting with `$` are resolved as environment variables (e.g., `$OPENAI_API_KEY`).
|
||
`ModelConfig` also declares `use_responses_api` and `output_version` so OpenAI `/v1/responses` can be enabled explicitly while still using `langchain_openai:ChatOpenAI`.
|
||
|
||
**Extensions Configuration** (`extensions_config.json`):
|
||
|
||
MCP servers and skills are configured together in `extensions_config.json` in project root:
|
||
|
||
Configuration priority:
|
||
1. Explicit `config_path` argument
|
||
2. `DEER_FLOW_EXTENSIONS_CONFIG_PATH` environment variable
|
||
3. `extensions_config.json` in current directory (backend/)
|
||
4. `extensions_config.json` in parent directory (project root - **recommended location**)
|
||
|
||
### Gateway API (`app/gateway/`)
|
||
|
||
FastAPI application on port 8001 with health checks at `GET|HEAD /health`, `/api/health`, and `/api/langgraph/health`. These exact paths are always public in `auth_middleware.PUBLIC_HEALTH_PATHS`, including when authentication is enabled; reverse-proxy prefixes and trailing slashes are normalized before the whitelist check. Set `GATEWAY_ENABLE_DOCS=false` to disable `/docs`, `/redoc`, and `/openapi.json` in production (default: enabled). Regression coverage: `tests/test_health_auth_bypass.py`.
|
||
|
||
**Routers**:
|
||
|
||
| Router | Endpoints |
|
||
|--------|-----------|
|
||
| **Models** (`/api/models`) | `GET /` - list models; `GET /{name}` - model details |
|
||
| **MCP** (`/api/mcp`) | `GET /config` - get config; `PUT /config` - update config (saves to extensions_config.json) |
|
||
| **Skills** (`/api/skills`) | `GET /` - list skills; `GET /{name}` - details; `PUT /{name}` - update enabled; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
|
||
| **Memory V2** (`/api/memory/v2`) | `GET /` - builtin data (USER.md/MEMORY.md); `POST/PUT/DELETE /entries` - add/replace/remove entry; `GET/PATCH/DELETE /config` - effective config + per-user overrides; `POST /search` - Hindsight recall/reflect proxy |
|
||
| **Agents** (`/api/agents`) | `GET /` - list built-in, owned, and published agents; `POST /` - create user-owned agent; `GET/PUT/DELETE /{id}` - read/update/delete by stable agent id; `GET /check` - validate non-empty display name (duplicates allowed) |
|
||
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags`); `GET /{name}` - details; `PUT /{name}` - update enabled; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
|
||
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
|
||
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + `tag_id` filters; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `GET /{name}/content` - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
|
||
| **Skills** (`/api/skills`) | `GET /` - list skills (optional `search` fuzzy-name + repeatable `tag_ids` multi-select filter, OR-matched; each item carries `tags` + owner-authored `detail`); `GET /{name}` - details; `PUT /{name}` - update enabled; `PUT /{name}/detail` - set the long-form `detail` text (custom skill: owner or admin; built-in/public: admin only ? materializes a `skills` row on demand); `GET /{name}/content` - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; `POST /install` - install from .skill archive (accepts standard optional frontmatter like `version`, `author`, `compatibility`) |
|
||
| **Memory** (`/api/memory`) | `GET /` - memory data; `POST /reload` - force reload; `GET /config` - config; `GET /status` - config + data |
|
||
| **Agents** (`/api/agents`) | `GET /` - list built-in, owned, and published agents (optional `search` + repeatable `tag_ids` multi-select filter, OR-matched; each item carries `tags` and `featured_order`); `POST /` - create user-owned agent; `GET/PUT/DELETE /{id}` - read/update/delete by stable agent id; `GET /check` - validate non-empty display name (duplicates allowed); `GET /selector` - homepage agent picker (single admin-curated list, visibility-filtered, sorted by `featured_order` asc); `PUT /{id}/featured` - **admin** sets/clears the homepage `featured_order` (must be `published=true` to feature; passing `order: null` clears unconditionally); `PUT /featured/order` - **admin** bulk-reorders the homepage list (body `{agent_ids: [...]}` ? indices become 0..N-1) |
|
||
| **Tags** (`/api/tags`) | `GET /` - list tags (optional `tag_type` filter; open to all); `POST /` - create tag (admin); `PUT /{id}` / `DELETE /{id}` - update/delete tag (admin); `GET /assignments/{target_type}/{target_id}` - tags on a resource; `PUT /assignments/{target_type}/{target_id}` - replace a resource's tags (admin). Tags are typed `agent`/`skill`/`scheduled_task` and may only tag resources of their own type. |
|
||
| **Uploads** (`/api/threads/{id}/uploads`) | `POST /` - upload files (auto-converts PDF/PPT/Excel/Word); `GET /list` - list; `DELETE /{filename}` - delete |
|
||
| **Threads** (`/api/threads/{id}`) | `DELETE /` - remove DeerFlow-managed local thread data after LangGraph thread deletion; unexpected failures are logged server-side and return a generic 500 detail |
|
||
| **Threads search** (`POST /threads/search`, LangGraph-compat) | Conversation listing. **System threads are excluded by default**: rows whose `metadata.system` is truthy (scheduler runs, roundtable workers) never crowd the user's list; a caller whose `metadata` filter explicitly references `thread_type`/`system` (e.g. studio notebooks) opts out of the exclusion. New `query` field = case-insensitive title (`display_name`) substring search, used by the chats page's server-side pagination. `POST /threads/count` takes the same request body (limit/offset ignored) and returns `{total}` under identical filtering/exclusion semantics ? powers the chats page's ??/??? display without fetching rows (`ThreadMetaStore.count`). `ThreadMetaStore.search` gained `exclude_system`/`query` params; JSON-blob filters batch-scan until the requested `limit`/`offset` window is full (so pagination over a heavily-filtered set never under-fills), with a `thread_id` tie-breaker on the `updated_at DESC` ordering for stable OFFSET paging. Tests: `tests/test_thread_meta_search.py` |
|
||
| **Artifacts** (`/api/threads/{id}/artifacts`, `/api/artifact-library`) | `GET /{path}` - serve artifacts; active content types (`text/html`, `application/xhtml+xml`, `image/svg+xml`) are always forced as download attachments to reduce XSS risk; `?download=true` still forces download for other file types. `GET /api/artifact-library` scans every current-user normal thread's authoritative `outputs/` directory (including files omitted from `present_files`), excludes uploads and image/audio/video media, joins the owning conversation title/agent route metadata, and exposes search/filter/pagination with a bounded 15-second per-user cache plus `refresh=true` bypass. Frontend route `/page/workspace/library` links each item back to the source chat with `?artifact=<path>` so the chat opens the existing artifact panel directly. |
|
||
| **Word export** (`POST /api/writing/export/docx`) | Shared Markdown → DOCX generator used by conversation, artifact, scheduled-task, Deep Research, and AI-writing download surfaces. `app/gateway/word_export.py` builds OOXML with no Python conversion dependency: A4 portrait, 3.7/3.5/2.8/2.6 cm margins, real multilevel heading numbering (`一、` / `(一)` / `1.`), two-character first-line indent on heading levels 1–3, formal body/table styles, page numbers beginning at 1 with odd pages left-aligned and even pages right-aligned, and embedded `app/gateway/assets/fonts/方正小标宋简体.ttf`. Before numbering it removes complete manual heading prefixes (including `2.1` / `2.1.1`) and drops Markdown horizontal rules from Word output. The ordinary-Q&A Lead Agent can emit the matching Markdown heading forms and avoid unordered bullet lists when the admin enables `prompt_prefix.ordinary_qa_markdown_format_enabled` in `system_settings.json`; the switch defaults off, while writing setup, writing mode, notebook, scheduled runs, and SOUL-driven custom agents keep their independent prompts. The font loader validates its internal family, SHA-256, and OpenType embedding permission (`OS/2.fsType`) before export. Tests: `tests/test_word_export.py`, `tests/test_es_query_routing_prompt.py`. |
|
||
| **Deep Research** (`/api/deep-research`) | Durable, user-owned research sessions/jobs with replayable SSE events and hidden-thread artifacts. `basic`/`quick` use the compact runner; `detailed` creates bounded subtopics, searches/writes sections concurrently, and assembles a report; `deep` recursively searches follow-up questions under a deterministic breadth/depth query budget. `multi_agent` is a LangGraph 1.x editor/researcher/writer/reviewer graph: its initial plan emits `awaiting_input`, the recorder stores the interrupt and releases the lease, and `POST /jobs/{id}/resume` rebuilds the graph from the persisted approved/revised plan. A bounded `custom_outline` is frozen into the session; headings/list entries deterministically supply the detailed and multi-agent plan, and the shared writer receives it only as a guarded content constraint. Completed reports support owner-scoped follow-up (`GET /sessions/{id}/messages`, `POST /sessions/{id}/chat`): context is only the current report, selected sources and recent messages; source-id citations are persisted, while `allow_new_research=true` is explicitly rejected. Owners may change that evidence set only after the job stops via `PATCH /sessions/{id}/sources/{source_id}`; this applies to future follow-ups rather than rewriting a completed report. Optional report illustrations are gated by `config.yaml -> deep_research.image_generation`: the Harness `OpenAICompatibleImageProvider` requires base64 output, the shared runner step writes only bounded `research-image-01..04.(png|jpg|webp)` artifacts and Markdown uses private `deep-research://` links. `GET /sessions/{id}/artifacts/{name}` checks session ownership before serving an image; an unavailable provider yields a warning but never fails the text report. Runners live in `packages/harness/deerflow/agents/deep_research/runners/`, may never import `app.*`, and use only the `AdapterBundle`. |
|
||
| **Gateway auth** | Owner-scoped Deep Research routes await the platform's asynchronous `get_current_user` resolver before they read or write persistence, ensuring each session and job operation receives a concrete user ID. Deep Research artifacts are always resolved from a whitelisted filename via `/mnt/user-data/outputs/<name>` before mapping to a host path, preserving the sandbox guard on native Windows and Linux hosts. |
|
||
| **Suggestions** (`/api/threads/{id}/suggestions`) | `POST /` - generate follow-up questions; rich list/block model content is normalized before JSON parsing |
|
||
| **Thread Runs** (`/api/threads/{id}/runs`) | `POST /` - create background run; `POST /stream` - create + SSE stream; `POST /wait` - create + block; `GET /` - list runs; `GET /{rid}` - run details; `POST /{rid}/cancel` - cancel; `GET /{rid}/join` - join SSE; `GET /{rid}/messages` - paginated messages `{data, has_more}`; `GET /{rid}/events` - full event stream; `GET /../messages` - thread messages with feedback; `GET /../token-usage` - aggregate tokens |
|
||
| **Feedback** (`/api/threads/{id}/runs/{rid}/feedback`) | `PUT /` - upsert feedback; `DELETE /` - delete user feedback; `POST /` - create feedback; `GET /` - list feedback; `GET /stats` - aggregate stats; `DELETE /{fid}` - delete specific |
|
||
| **Runs** (`/api/runs`) | `POST /stream` - stateless run + SSE; `POST /wait` - stateless run + block; `GET /{rid}/messages` - paginated messages by run_id `{data, has_more}` (cursor: `after_seq`/`before_seq`); `GET /{rid}/feedback` - list feedback by run_id |
|
||
| **Open Chat** (`/api/open/chat`) | **Unauthenticated** open Q&A entry for external systems (the `/api/open/` prefix is whitelisted in `auth_middleware._PUBLIC_PATH_PREFIXES` + CSRF-exempt in `csrf_middleware.should_check_csrf`). `POST /chat` `{message \| messages[], model_name?, agent_name?, agent_id?, thinking_enabled=false, strip_think=true, thread_id?, recursion_limit=100}` → resolves the agent by **Chinese display name** (`agent_name`, matched against each `.deer-flow/agents/*/config.yaml` `name` via `_resolve_agent_id`; explicit `agent_id` wins if given), builds the lead agent via `make_lead_agent` (so the chosen **agent's** SOUL/tools/skills all apply), `ainvoke`s it once (stateless — no checkpointer, no thread persistence; runs under the `"default"` user bucket), and returns **only the final** `{answer, model_name, agent_name, agent_id, thinking_enabled, thread_id}` (no intermediate tool-call/process). **`POST /chat/stream`** is the **SSE streaming variant** (same request body) for **long tasks** — the blocking `/chat` produces no downstream bytes for the whole run, so a front nginx `proxy_read_timeout` (default 60s) cuts it off with a **504**; the stream keeps the connection warm (a `: keepalive` SSE comment every `_HEARTBEAT_SECONDS=15`s via a background producer-task + queue, plus `event: delta` AI-text increments and `event: progress`), and ends with **`event: result`** carrying the same final `{answer, …}` JSON (or `event: error`). Clients only need to read `event: result`; intermediate events are ignorable. Both endpoints share `_build_run_context` (config/agent/thread). `GET /agents` lists `{name, id}` so callers know what `agent_name` to send. **Thinking off is reliable**: when `thinking_enabled=false` the route also sets `thinking_force_disabled=true` so the lead agent calls `create_chat_model(force_disable_thinking=True)`, which injects **three** OpenAI-compatible "thinking off" shapes into `extra_body` — top-level `enable_thinking=false` (**SiliconFlow / DashScope hybrid models** like `deepseek-ai/DeepSeek-V4-*` on `api.siliconflow.cn` — the knob the nested `chat_template_kwargs` shape does NOT reach), `chat_template_kwargs.enable_thinking=false` (vLLM/Qwen native), and `thinking.type="disabled"` (Zhipu GLM) — even for models that only declare `supports_thinking: true`. For the **web-UI path** (which sends `thinking_enabled=false` but NOT `force_disable`), SiliconFlow models need `when_thinking_disabled: {extra_body: {enable_thinking: false}}` in their `config.yaml` entry (added to `deepseek-ai/DeepSeek-V4-Flash`) for thinking-off to take effect. Sets `is_scheduled_run=true` internally to skip the auto-title LLM call. **Generated-file return (`include_files=true`, default on)**: an agent that produces its report as a file (writes to `/mnt/user-data/outputs/` + `present_files`) returns only a short summary in `answer`; to recover the real deliverable both endpoints now also return `files: [{name, virtual_path, encoding(text|base64), content}]` — `_collect_output_files` scans the run's `state['thread_data']['outputs_path']` (the open thread is stateless so artifacts can't be fetched by thread, but the endpoint runs server-side and the sandbox already wrote the files to disk), text decoded as utf-8 else base64, capped by `_MAX_OUTPUT_FILES=20` / `_MAX_OUTPUT_FILE_BYTES=5MB`. `/chat` puts `files` on the response; `/chat/stream` puts it in the `event: result` payload. A run with no answer **and** no files now 502s ("…文本答案或文件"). Router `app/gateway/routers/open_chat.py`; standalone caller `test_open_chat_api.py` (repo root, streaming by default); tests `tests/test_open_chat.py`. |
|
||
| **Open 3qfx Ask** (`/api/open/3qfx/ask`) | **Unauthenticated** open entry that, given just a **任务id**, directly kicks off the 3qfx(三情分析)「第一次问答」as a **real persisted conversation** (sibling of `/api/open/chat`; same `/api/open/` auth-whitelist + CSRF-exempt). `POST /ask` `{task_id, message?, model_name?, thinking_enabled=false}` → (1) resolves the `3Q` agent from `app.state.business_mapping_store.get_mapping("3Q")` (the **same** agent the frontend deep-link `goPath=3qfx` uses; 503 if unconfigured); (2) **fetches the task detail server-side** (consumer `cop-task-detail`) and builds the opening text **identical to the frontend** `buildOpeningText` 3qfx branch — `任务名称:…\n任务内容:…\n任务id(task_id)为{task_id}` (an explicit `message` overrides and **skips** the fetch); (3) creates a fresh thread tagged `metadata.{taskId, agent_id}` (the exact filter `useTaskThreads` uses, and `metadata.taskId` makes it task-deeplink-shared / creator-path-resolved) and **starts a background persisted run** via `start_run(..., enforce_access=False)` (writes the checkpointer, runs under the `"default"` bucket); (4) **returns immediately** `{thread_id, run_id, agent_id, status, opening_text}` — the caller just needs "started", then watches the run in the 3qfx 任务工作区 conversation list. `start_run` gained a keyword-only `enforce_access` (default `True`); the open entry passes `False` so the admin-configured agent isn't rejected by the per-(missing)-login visibility check. **Consumer fetch config** lives in `config.yaml → task_deeplink`: `cop_task_detail_url` (full URL, appends `?taskId=`), `cop_task_detail_authorization` (default `Bearer admin`), `cop_task_detail_test` + `mock.cop_task_detail` (offline fake when the upstream is unreachable — on in this deployment). Router `app/gateway/routers/open_3qfx.py`; tests `tests/test_open_3qfx.py`. |
|
||
| **Thread Shares** (`/api/thread-shares`) | Per-conversation sharing with two modes (`thread_shares.mode`). **`import`**: `POST /` `{thread_id, mode}` - create (or reuse) a revocable share code; `GET /{code}/preview` + `POST /{code}/import` - copy that conversation into the caller's account (fresh `thread_id` via `app/gateway/thread_copy.py` + new `threads_meta` row + `user-data` copy; dedup on `metadata.shared_origin`). **`view`**: a public, read-only static page ? see Public Shares below. Both modes: `GET /` - list my codes; `DELETE /{code}` - revoke. |
|
||
| **Public Shares** (`/api/public/shares`) | **Unauthenticated** (the `/api/public/` prefix is whitelisted in `auth_middleware._PUBLIC_PATH_PREFIXES`). `GET /{code}` - read-only conversation snapshot (title + serialized messages + artifacts) for a `view`-mode share; messages are read straight from the latest checkpoint. `GET /{code}/files/{path}` - serves a file referenced by that conversation (images, Markdown, text, PDF, ?) so the shared page's links are openable; path resolution is confined to the owner's thread directory. Active content (`text/html`, `application/xhtml+xml`, `image/svg+xml`) is forced to download (`_public_share_disposition`); everything else is served inline. Missing/revoked/non-`view` codes all return 404. |
|
||
| **Roundtable Drafts** (`/api/roundtable-drafts`) | **Strictly per-user** persistence for the roundtable-planning page (replaces the old browser `localStorage`). `GET /` - list drafts (lightweight meta — incl. `chain_title`, the business-chain name used this session, cheaply regex-extracted from the in-memory `step2` blob in `_draft_to_meta`/`_extract_chain_title`; null = no chain / 自由模式; surfaced on the 第一步「分析记录」history card + sidebar list instead of the old「第 X 步」); `POST /` - create; `GET/PUT/DELETE /{id}` - full read / partial update / delete (all owner-scoped). A draft = one full session (Step 1 intent + Step 2 roundtable), step snapshots stored as JSON in `PortableLongText`. Recommendation history bound to a draft: `GET /{id}/recommendations` - list newest-first; `POST /{id}/recommendations` - append one run (objective/status/rationale/picks/candidates). taskId 深链会商记录 are **not** here — they live in the separate `/api/roundtable-task-drafts` store (see below) so personal history never mixes with task records. Backed by `deerflow.persistence.roundtable_drafts` (tables `roundtable_drafts`, `roundtable_recommend_history`); store wired in `deps.py` as `app.state.roundtable_draft_store`. Tests: `tests/test_roundtable_task_drafts.py`. |
|
||
| **Roundtable Task Drafts** (`/api/roundtable-task-drafts`) | **Task-scoped, unpartitioned** store for taskId-deeplink roundtable 聊天记录 (无界嵌入抽屉). Keyed **only by `task_id`** (one task → many records), **never partitioned by user** — any authenticated caller may read/write/delete a task's records (`created_by` is audit-only). Physically separate from `roundtable_drafts` so the two never mix. `GET /by-task/{task_id}` - all drafts for a task, newest-first (history dropdown), enriched with the latest background-job status per draft via `RoundtableJobRepository.status_map_by_drafts` (each meta also carries `chain_title`, same `_extract_chain_title` regex as the personal store); `GET /by-task/{task_id}/latest` - newest draft (回显最新一条); `POST /` - create (`task_id` required); `GET/PUT/DELETE /{id}` - full read / partial update / delete (all user-agnostic). Backed by `deerflow.persistence.roundtable_task_drafts` (table `roundtable_task_drafts`, auto-created by `create_all`; migration `20260618_01`); store wired in `deps.py` as `app.state.roundtable_task_draft_store`. Existing `task_id`-bearing `roundtable_drafts` rows are migrated over by `scripts/migrate_roundtable_task_drafts.py`. Tests: `tests/test_roundtable_task_drafts.py`, `tests/test_roundtable_drafts_by_task.py`. |
|
||
| **Roundtable Diagnostics** (`/api/roundtable-diagnostics`) | 会商全流程诊断日志 (报错 + 关键事件). One row per event in `roundtable_diagnostics`: `{scope (background = 后台挂起 / foreground = 前台服务端观测 / client = 前端上报), stage (step1_intent/step2_init/step2_leader/step2_seat/step2_recommend/step2_orchestration/step3_summary/step3_report/step3_dashboard/step3_structure/upstream/job), level (warning/error — **`info` 级运行日志不再收集**:`record_diagnostic` 在写库前丢弃 `level=="info"`,只留报错;begin/consensus/gather/job_started 等纯进度事件即使仍调 `.diag(level="info")` 也成 no-op,故诊断页 = 报错日志页), event, message, detail (JSON), job_id/draft_id/task_id/user_id/agent_id/agent_name/cycle, created_at}`. **Reads are admin-only**: `GET /` - filter (scope/stage/level/event/job_id/draft_id/task_id/user_id/q/since/until; `scope` accepts a comma-separated set so 前端「前台调用」= `foreground,client`) + paginate (newest-first), each item enriched with `userName` (user_id → 邮箱 `@`-prefix via `get_local_provider().get_user`, best-effort batched cache); `GET /facets` - distinct values for filter dropdowns; `POST /admin/cleanup` - retention purge. **`POST /report` is NOT admin-gated** — any authenticated user's frontend posts the errors *the user actually hit* (HTTP 502 before any SSE frame / network drop / SSE error frame) so they're collected too; stored as `scope="client"`. Backed by `deerflow.persistence.roundtable_diagnostics` (table `roundtable_diagnostics`, auto-created by `create_all`; migration `20260619_01`); store wired in `deps.py` as `app.state.roundtable_diagnostic_store` + a process-level default store (`register_default_store`) so the harness-side engine/gateway can write across the app/harness boundary. **Writers** (all best-effort — diagnostics never break a run): background path uses `DiagnosticsRecorder` (bound to job ctx) threaded through `RoundtableJobExecutor` → `JobProgressWriter.diag(...)` (engine transitions/step3 failures/empty-delivery) + `InProcessRoundtableGateway` (per-turn timeout/report·summary·dashboard truncation+missing); foreground path uses `app/gateway/roundtable_diag.py::record_foreground_diag` in `multi_agent.py`/`intent.py`/`recommend.py` error frames. **Bug-fix shipped alongside**: `InProcessRoundtableGateway._run_agent_turn` guards each agent run with an **inactivity (idle) timeout** (`_await_run_with_idle_guard`, `ROUNDTABLE_SEAT_IDLE_TIMEOUT_SECONDS` default 1200s = 20min, tuned high for unstable intranet models) — **not** a total-duration cap, so a slow-but-alive run that keeps streaming tokens runs as long as it needs; only a run that emits **zero StreamBridge events for `idle_timeout` seconds straight** (a true freeze: 内网模型连接池挂死 / sqlite 锁在挂载卷上挂死) is cancelled + raises (顶层把 job 标 error,而非永久冻结) + records a `seat_turn_timeout` diagnostic. Activity is observed via `app.state.stream_bridge` (the worker publishes every chunk there). Optional absolute cap `ROUNDTABLE_SEAT_MAX_TOTAL_SECONDS` (default 0 = off). Frontend: admin page `pages/RoundtableDiagnosticsPage.tsx` (route `/page/strategy/admin/roundtable-diagnostics`), opened from the **会商页右上角设置按钮** (admin only — `handleSettingsClick` in `roundtable-planning/pages/RoundtablePlanningPage.tsx`); admin read client `strategy-components/api/roundtable-diagnostics.ts`. **Client error reporter** `roundtable-planning/api/diagnostics-report.ts` (`reportRoundtableError` fire-and-forget + `keepalive`, `setRoundtableErrorContext` for draft/task attribution set by the page, `stageForRoundtableAgent`) is called at every client-observable failure: `multi-agent.ts` (`streamMultiAgent` — covers Step2 leader/seat AND Step3 report/summary/dashboard/structure — reports HTTP error + SSE `{status:"error"}` frame + worker `event:error` frame [模型调用没成功] + fetch-reject [网络不可达] + mid-stream drop [连接重置/超时], deduped via a `reported` flag; `initMultiAgent` HTTP/error-frame/incomplete/mid-stream), `intent.ts` (Step1 HTTP+error-frame), `recommend.ts` (Step2 recommend HTTP+error-frame), `roundtable-jobs.ts` (job api/stream). The admin page's first filter is a **后台挂起 / 前台调用** toggle (frontend = `foreground,client`); the stage dropdown options adapt to it (background → step2+step3+job; frontend → step1+step2+step3) and each row shows the reporting user's name + 👤. Tests: `tests/test_roundtable_diagnostics.py`, `tests/test_roundtable_inprocess_timeout.py`. **Seat skill-recognition (弱模型补偿)**: a seat = one custom agent run whose configured skills (`agent config.yaml → skills`) are injected **only** via the system prompt's `<skill_system>` block (model must `read_file` the SKILL.md itself) — weak models (deepseek-chat 等) ignore it and never use their skills. `app/gateway/roundtable_seat_skills.py::append_seat_skill_directive(agent_id, task, *, chain_enabled=False)` appends an explicit Chinese directive (each configured-and-enabled skill's name + description + SKILL.md container path + "先 read_file 再按其流程取信息/执行") to the **seat's per-turn task (human turn)** — far stronger pull on weak models than the system prompt. Wired into **both** seat paths: foreground `multi_agent._special_run` (only `role in {seat, seat_parallel}`; Step3 singletons report/summary/dashboard/ingest excluded) + background `InProcessRoundtableGateway.run_seat`. No-op (returns task unchanged) when the seat has no configured skills. **2026-06-20: default-off + two opt-in switches (OR semantics)** — was global-on, now injects only when **either** the seat agent's own switch (`AgentConfig.seat_skill_directive`, edited in the agent edit dialog `agent-card.tsx` + `agent-detail-panel.tsx`) **or** the **per-seat** switch on that agent **within the selected business chain** is on. The per-seat chain switch lives **in the chain's `seats` JSON** (`ChainSeatSchema.seat_skill_directive`, no DB column — rides the existing JSON like the per-seat `model` override), edited via a small ✨ icon → dropdown (Switch + description) on each agent card in `BusinessChainEditorPage`. It is threaded **per-seat** to each seat run: foreground `multi_agent` per-call field `seat_skill_directive` (orchestration looks it up per agentId from a `{agentId: bool}` session map) → `_special_run` → `chain_enabled`; background as the job field `seat_skill_directives: dict[str,bool]` (StartJobRequest → StartParams → deps factory → `InProcessRoundtableGateway._seat_skill_directives`) → `run_seat` looks up `agent_id` → `chain_enabled`. `ROUNDTABLE_SEAT_SKILL_DIRECTIVE` is only a **master kill-switch** (default allow; set `0/false` to hard-disable regardless of the switches — never force-enables). Tests: `tests/test_roundtable_seat_skills.py`, `tests/test_agents_backend.py`. **2026-06-20 Step2 提速三件套**:(1) **每席位推理深度覆盖** —— 业务链条 `seats` JSON 里逐席位可配 `seat_mode`(flash/thinking/pro/ultra,覆盖链条级 `seatMode`;编辑器席位卡右侧 ⏱Gauge 图标下拉,null=跟随链条默认)。前台经 `useStep2Orchestration` 三处 dispatch 的 `seatModeToRequestFields(seatModes[id] ?? selectedMode)` 透传;后台经 job `seat_modes: dict[str,str]` → `StartParams` → `InProcessRoundtableGateway._seat_modes`,`run_seat` 用 `_SEAT_MODE_PARAMS`(与前端 `seatModeToRunParams` 逐字对齐)解析覆盖作业级 seat_*(未命中→跟随作业级→policy 默认)。无 DB 列(ride `seats` JSON,但 `ChainSeatSchema.seat_mode` 须显式声明否则 Pydantic 剥掉)。Tests: `tests/test_roundtable_seat_modes.py`。(2) **recommend 批内并行** —— 后端 `engine._run_recommend` 由 for-await 串行改 `asyncio.gather`(对齐前端 2026-06-13 已并行;总控一轮派多席位=判定相互独立→并发;批内用派活前 deliveries 快照、批后统一并入,语义不变)。(3) **研讨前并行『取数』两段式**(业务链条 `gather_first` 列,逐条 opt-in,migration `20260620_02`,editor「研讨前并行取数」开关)—— `gather_first=True` 时 `run_orchestration` 在分发前先 `engine._run_gather_phase`(全席位 `asyncio.gather` 跑 `_GATHER_TASK`「只取数不研讨」),产出 `gathered: dict` 经 `_seat_message(gathered=...)` 作为「已收集资料」块注入 recommend/chain/dag 各研讨轮席位任务(与「前序交付」块分开),把最慢的技能延迟前置并行、给(尤其 chain 串行的)研讨段提速;研讨阶段不限制再调技能。前台对称实现:`useStep2Orchestration.runGatherPhase` 在 `runActiveOrchestration` 分发前跑一次(`gatherPhaseDoneRef` 闸门,resetInitGate 复位),产出存 `gatheredDeliveriesRef` 经 `peerDeliveries`/初始 `priorDeliveries` 注入研讨席位。链条级 `gatherFirst`(camel)↔`gather_first`(snake) 经 `api/chains.ts` 映射;job 经 `StartJobRequest.gather_first`→`StartParams`→`run_orchestration(gather_first=…)`。Tests: `tests/test_roundtable_gather_phase.py`。(4) **削减重复广播**(前端 only,多智能体 foreground 路径)—— `suppress_pre_broadcast`(仅 leader:抑制「总控本轮输入→各席位 thread」预广播 + 「总控给X派活」→兄弟席位的 post 广播,二者一个 flag)此前只在 DAG 派活轮开;现扩到 recommend(`safetyCounter > 1`,即第 2 轮起的续轮/综合)+ chain 逐席位(`i > 0`,第 2 个席位起)+ chain 收口轮(恒 true)。**保留每种模式的首轮预广播**(它为席位打底任务背景),只 suppress 重复的续轮广播(N 次读写巨大席位 thread 的纯浪费 + 弱模型模仿派活);席位的前序交付另经「席位交付广播」`broadcastSeatDeliveryToOthers` 送达,不受影响。后台引擎路径无此广播(`_seat_message` 直接拼上下文),故 #4 仅前端。**2026-06-20 深链接业务链抽取规范**:深链接会商(taskId+goPath,businessCode 如 rwfx→6BF)下,结构化抽取(flow-json)与方案总结报告都按所选业务链**完整层级**强制校正。结构化抽取实为 `roundtable-structure` 智能体(`flow_extract.py` 的 `_system_prompt()=load_agent_soul(STRUCTURE_AGENT_ID)`,抽取→校验两段式 `model.astream`);其 SOUL 全局一份,旧实现把某业务层级写死进 SOUL→所有业务都按「目的→行为体」抽(bug)。修法:**单份领域知识、两角色用法**的 `deerflow.agents.roundtable_orchestrator.extraction_spec`(放 harness 层,app+harness 都可 import;`build_structure_directive`/`build_summary_directive` 按 businessCode 取;rwfx=6BF 写实「任务→目的→行为体→预期行为→驱动因素→叙事,预期行为↔驱动因素 1:1、余 1:N、label 不含步骤/阶段名、优先报告回退席位」,其余业务通用规范=强约束完整性+命名+取数且明令不套固定模板;business_code 空→""零副作用)。注入:(1) 结构化抽取 `flow_extract._hierarchy_hint`+`_validate_messages` 用 `build_structure_directive` 权威覆盖默认层级(前端 `streamFlowExtract` 传 `business_code`,源 `ingestBusinessCode`);(2) 方案总结报告 `build_summary_directive`——前台 `multi_agent._special_run`(role==summary 且 business_code) 追加、后台 `engine.build_summary_prompt(business_code=…)`(StartParams/StartJobRequest 透传)。**仅深链接生效**。种子 `roundtable-structure/SOUL.md` 层级顺序改 spec-driven、连接示例去 rwfx 写死(存量 live SOUL 不被 seed 覆盖,但注入规范权威可纠正)。Tests: `tests/test_roundtable_extraction_spec.py`。 |
|
||
| **Roundtable Draft Shares** (`/api/roundtable-draft-shares`) | Per-draft sharing for the roundtable-planning page, mirroring Thread Shares but the shared resource is a `roundtable_drafts` row. Two modes (`roundtable_draft_shares.mode`). **`import`**: `POST /` `{draft_id, mode}` - create (or reuse) a revocable code; `GET /{code}/preview` + `POST /{code}/import` - copy that draft (step1/2/3) into the caller's account under a deterministic per-(code,recipient) id (`imp_<sha1>` ? re-import is a no-op). **`view`**: a public, read-only result page ? see Public Roundtable Shares. Both: `GET /` - list my codes; `DELETE /{code}` - revoke. Backed by `deerflow.persistence.roundtable_draft_shares` (table `roundtable_draft_shares`, migration `20260606_01`); store wired in `deps.py` as `app.state.roundtable_draft_share_store`. Tests: `tests/test_roundtable_draft_shares.py`. |
|
||
| **Public Roundtable Shares** (`/api/public/roundtable-shares`) | **Unauthenticated** (`/api/public/` whitelisted). `GET /{code}` - read-only snapshot for a `view`-mode draft share: returns the Step 3??????report HTML (`step3.html`, a self-contained document) + title/summary/`has_report`. The frontend renders it in a sandboxed `iframe` (no `allow-scripts`) at `#/roundtable-share/<code>`. Missing/revoked/non-`view` codes all return 404. |
|
||
| **Scheduled Tasks** (`/api/scheduled-tasks`) | Per-user cron-like agent tasks. `GET/POST /` - list/create owned tasks; `PATCH/DELETE /{id}` - update/delete (owner); `POST /{id}/{pause,resume,trigger}`; `GET /{id}/runs`, `GET /runs/{rid}`, run-file preview/edit/download; `GET /runs/{rid}/messages` - agent messages for the run's underlying agent run (used to inspect a stuck/running run; authorized via `_load_run`); `GET /runs/{rid}/html` - generated HTML document for an `html_page` run; `DELETE /runs/{rid}` - owner-scoped, with an admin fallback so admins can delete any user's run; publish/subscribe under `/{id}/{publish,unpublish,subscribe}` and `GET /published`. **Admin (cross-user):** `GET /admin/all` - every task; `PATCH/DELETE /admin/{id}` - edit/delete any task; `POST /admin/{id}/trigger`; `GET /admin/{id}/runs`. Admin routes require `system_role == "admin"`. |
|
||
| **HTML Page Favorites** (`/api/html-page-favorites`) | Bookmarked generated HTML pages, owner-scoped. `GET /` - list; `POST /` `{page_id, title?, description?, tags?}` - bookmark a `scheduled_html_pages` row (extracts a reusable style/layout summary via GLM5 at favorite time); a page can only be favorited once per user (`DuplicateFavoriteError` ? HTTP 409). `GET/DELETE /{id}` - read/delete (the create-task reference picker deletes favorites here). A favorite can be selected as a style reference when creating a new `html_page` task. |
|
||
| **Public Scheduled Tasks** (`/api/public/scheduled-tasks`) | **Unauthenticated** mirror for embedding. Adds `GET /runs/{rid}/html` and `GET /{task_id}/latest-html` for the `html_page` preview; the run result keeps `result.html` (only file paths are redacted). |
|
||
| **Light Apps** (`/api/light-apps`) | Light-app registry (????? / ????), **global/shared** (no user scoping; `created_by`/`updated_by` audit-only). Two kinds: `window_open` (launched in a new tab from the ???? card grid) and `iframe` (embedded under a parent sidebar menu via `mount_location` and dynamically injected into the sidebar). Both carry `route_params` ? a list of `{key, paramType: login\|theme, source, customValue, themeValue}` descriptors the **frontend** resolves at launch time (login params read from `localStorage.userInfo`; theme params use a literal value) and appends to `url` as query string. `GET /` - list (filters `app_type`/`mount_location`/`enabled_only`/`search`); `GET /{app_id}` - read one (both open to any authenticated user); `POST /`, `PUT /{app_id}`, `DELETE /{app_id}` - **admin only**; iframe apps must carry a `mount_location` (400 otherwise); the unique `name` collides as HTTP 409. Backed by `deerflow.persistence.light_apps` (table `light_apps`, auto-created by `create_all`); store wired in `deps.py` as `app.state.light_app_store`. Tests: `tests/test_light_apps.py`. |
|
||
| **Menu Overrides** (`/api/menu-overrides`) | Sidebar-menu overrides (菜单管理), **global/shared** (no user scoping; `updated_by` audit-only). The static menu tree (icons/routes/`adminOnly`) lives in the **frontend** code (`core/page-layout/sidebar-menu.ts`); iframe light-apps are injected at render time. This router persists a per-node **overlay** the frontend merges onto that tree: `parent_id` (re-parent, moves the whole subtree; NULL = top level), `sort_order` (reorder among siblings), `label` (rename; NULL keeps code label), `disabled` (停用 — hides the node **and its subtree**). A node with no row renders exactly as code defines it, so new code menus / newly-registered iframe apps appear automatically and are immediately configurable. `GET /` - list all overrides (open to any authenticated user); `PUT /` - **admin only**, atomically **replaces the whole set** (`{overrides: [...]}`); rejects re-parenting that would exceed **3 levels** or form a cycle (`_validate_depth`, HTTP 400). Backed by `deerflow.persistence.menu_overrides` (table `menu_overrides`, auto-created by `create_all`); store wired in `deps.py` as `app.state.menu_override_store`. Frontend: merge logic `core/page-layout/menu-overrides.ts` (`applyMenuOverrides` in `use-sidebar-menu.ts`), client `strategy-components/api/menu-overrides.ts`, admin page `pages/MenuManagementPage.tsx` (@dnd-kit drag-reorder/re-parent + 「移动到」parent selector + ↑↓ arrows + 改名/停用), entry in the bottom-left settings dropdown (`components/page-sidebar.tsx`, admin-only). Tests: `tests/test_menu_overrides.py`, `frontend-web/src/core/page-layout/menu-overrides.test.ts`. |
|
||
| **Compact Menu Overrides** (`/api/compact-menu-overrides`) | Compact/sidebar-summary left-nav overlay (简介模式菜单), **global/shared**. Two independent sets keyed by `menu_type` (`A` default / `B`); saving one type does not wipe the other. Type B rows use a `B::` node_id prefix so they stay unique on databases that still use `node_id` as the sole PK. `GET ?menu_type=` lists one type (open); `PUT {menu_type, overrides}` admin-only replaces that type. Frontend: `MenuManagementPage` 简介模式 type switcher + `runtime-config.js` `VITE_COMPACT_MENU_TYPE`. Tests: `tests/test_compact_menu_overrides.py`. |
|
||
| **Sensitive Words** (`/api/sensitive-words`) | Sensitive-word / desensitize map (敏感词管理 / 脱敏词), **global/shared** (no user scoping; `updated_by` audit-only). Each row maps an original `word` → `replacement`, with `enabled` (停用 parks a word without deleting) + `sort_order`. The **frontend** ships a seed map in code (`src/lib/desensitize.ts` `SEED_DESENSITIZE_MAP`) used as the default on a fresh deployment; once an admin saves, this table is the source of truth and the frontend loads it at startup (`DesensitizeWordsLoader` → `replaceDesensitizeMap`, **enabled** rows only). `GET /` - list all words (open to any authenticated user — every page desensitizes); `PUT /` - **admin OR the `lqq` account** (`_require_manager`: `system_role == "admin"` or email `@`-prefix `== "lqq"`), atomically **replaces the whole set** (`{words: [...]}`, de-duped by `word`, last wins). Backed by `deerflow.persistence.sensitive_words` (table `sensitive_words`, PK `word` `String(191)`, auto-created by `create_all`); store wired in `deps.py` as `app.state.sensitive_word_store`. Frontend: client `strategy-components/api/sensitive-words.ts`, page `pages/SensitiveWordsPage.tsx` (route `/page/strategy/admin/sensitive-words`, entry in the bottom-left settings dropdown — **gated on username `lqq`, not admin role**, present in both sidebar modes; 「导入种子词」 merges missing seed entries). Tests: `tests/test_sensitive_words.py`. |
|
||
| **Sentiment Agent Proxy** (`/api/sentiment-agent`) | BFF for the frontend 舆情分析 virtual agent. `POST /stream` accepts the AG-UI body plus `apiUrl`/`authUsername`/`authPassword` from `runtime-config.js`, forwards to that URL with TLS verify off, and byte-pipes upstream SSE. Connection/HTTP failures return the full error in `detail`; mid-stream failures emit AG-UI `RUN_ERROR`. Distinct from `skill-packages/sentiment-analysis.skill`. Tests: `tests/test_sentiment_agent_proxy.py`. |
|
||
| **Dashboard Sessions** (`/api/dashboard-sessions`) | 「大屏绘制智能体」独立页面会话,**按外部 `task_id` 存取,不按 user 分权**(open read + shared writes;`user_id` audit-only)。一个 task 一行(upsert):`transcript`(聊天记录 JSON)+ `report_json`(最新结构化数据,可解析 = 「有效结果」)。`GET /by-task/{task_id}` - 取该 task 会话(无则 404);`PUT /by-task/{task_id}` - upsert(`{transcript?, report_json?}`,只覆盖传入项)。Backed by `deerflow.persistence.dashboard_sessions`(table `dashboard_sessions`, auto-created by `create_all`,无迁移);store wired in `deps.py` as `app.state.dashboard_session_store`。前端独立页 `roundtable-planning/pages/DashboardAgentPage.tsx`(路由 `/page/strategy/dashboard-agent?taskId=`,侧边栏「问答管理 → 大屏绘制」)让智能体分析任务报告 JSON → report-json → 渲染 `DashboardApp`。Tests: `tests/test_dashboard_sessions.py`。 |
|
||
| **Task Reports** (`/api/task-reports`) | 任务级研判情报报告详情,**按 `task_id` 共享、不按 user 分权**(写入者只留 `updated_by` 审计)。`GET /by-task/{task_id}` 始终返回原 Consumer 兼容外壳 `{state:"200", msg:"操作成功!", data:[...]}`(未保存为 `data: []`);`POST /by-task/{task_id}` 首次保存,已有报告返回 409;`PUT /by-task/{task_id}` 原子替换完整 `data` 数组。`POST /import/by-task/{task_id}` 面向 TaskCOP 导入页,接收 `sq-report-mock.json` 兼容的 `data`/`records` 数组并完整覆盖目标任务;路径中的选定任务 id 优先于文件内的 `taskId`,避免误写。每一项严格是 `{id, taskId, createTime, updateTime, sessionId, contentJson, categoryType}`;`contentJson` 以 JSON **字符串**原样落在 `records_json`,避免破坏大屏按类目解析的输入格式。保存或导入非空 `data` 后,router 会通过 `app.state.taskcop_task_store.mark_analysis_completed` 将同 id 的 TaskCOP 任务升级至状态 `25`,使旧任务列表显示并筛选为“已完成”;普通保存时不存在匹配任务的报告仍可独立保存,导入则要求目标任务存在。Backed by `deerflow.persistence.task_reports`(table `task_reports`, auto-created by `create_all`);store wired in `deps.py` as `app.state.task_report_store`。前端客户端 `strategy-components/api/task-reports.ts`,`core/auth/dashboard-input.ts` 与 rwfx 3q 完成检查均直接读此接口,不再依赖 Consumer 或本地 sq-report mock。Tests: `tests/test_task_reports.py`。 |
|
||
| **Task Buttons** (`/api/task-buttons`) | Task-workspace configurable jump buttons (按钮管理), **global/shared** (no user scoping; `updated_by` audit-only). Each row = one jump button in a task deep-link workspace's bottom 快捷跳转 bar: `{id, business (rwfx/3qfx/xdfx), label, link_type (url/business/purpose/host), target, append_task_id, id_param (editable query-param name, e.g. id/task_id), id_kind (task/action — only xdfx differs, its deep-link id is an 行动id), open_mode (blank/self), login_params, append_auth, auth_token_param, auth_name_param, enabled, sort_order}`. **Two independent, coexisting login-append mechanisms for `url` buttons** — both resolved frontend-side: (1) `login_params` (JSON column `login_params_json`, `PortableJSON`, migration `20260625_01`) = a list of `{key, source, custom_value}` resolved at click time (`source` = a `userInfo` field or `other`→literal `custom_value`) and appended to the target URL — mirrors the light-app login params; empty = nothing appended. (2) `append_auth` + `auth_token_param` + `auth_name_param` (columns, migration `20260625_02`, default off) — appends `?{auth_token_param}={token1}&{auth_name_param}={username}` so the target system can identify the current user; gated on **token login** — the frontend only stashes `token1` (the `?authToken=`) + the raw `getTokenInfo` username (`login.token1`/`login.username`) when login went through `/login/token`, so a non-token-login session appends neither. The **frontend** holds code-seed defaults (rwfx's 4 historical buttons built from the live `task_deeplink` config); only when this table is empty do the workspace + admin page fall back to them, else this table is the sole source of truth. `GET /` - list all (open to any authenticated user — the workspace renders them); `PUT /` - **admin only** (`_require_admin`), atomically **replaces the whole set** (`{buttons: [...]}`, de-duped by `id`, last wins). Backed by `deerflow.persistence.task_buttons` (table `task_buttons`, auto-created by `create_all`); store wired in `deps.py` as `app.state.task_button_store`. Frontend: helpers `core/auth/task-buttons.ts`, client `strategy-components/api/task-buttons.ts`, admin page `pages/TaskButtonsPage.tsx` (route `/page/strategy/admin/task-buttons`, entry「按钮管理」in the bottom-left settings dropdown, admin-only, both sidebar modes; per-business tabs + 「导入默认」). Tests: `tests/test_task_buttons.py`. |
|
||
| **Fixed Questions** (`/api/fixed-questions`) | Administrator-managed exact-match fixed Q&A (问题配置), **global/shared**. Each row is `{id, question, answer, enabled, tokens_per_second, sort_order}`; stream speed is configurable per row from `1–1000 token/s` and defaults to `100`. `GET /` and whole-set atomic `PUT /` are both admin-only. At the start of an external chat run, `services._find_fixed_answer` checks the latest eligible plain-text human message with exact, case/space/punctuation-sensitive equality. Enabled matches switch the run to `make_fixed_answer_agent`: a local `BaseChatModel` tokenizes with `cl100k_base`, replays the answer in rate-limited token chunks, and keeps the existing LangGraph SSE `messages`/`values` protocol, run journal, checkpoint history, deterministic local title, and refresh behavior intact while no provider/model request or model-concurrency slot is used. File/multimodal/context-prefixed turns, `canvas_agent`, `ai_writing`, writing mode, resume commands, and internal orchestration/background child runs bypass matching. Storage failures fail open to the normal model path. Backed by `deerflow.persistence.fixed_questions` (table `fixed_questions`, SHA-256 indexed lookup plus exact original-string collision check, migrations `20260821_01` + `20260821_02`), wired as `app.state.fixed_question_store`. Frontend: `strategy-components/api/fixed-questions.ts`, admin page `pages/FixedQuestionsPage.tsx`, route `/page/strategy/admin/fixed-questions`, entries in both full-layout and compact-workspace settings menus. Tests: `tests/test_fixed_questions.py`. |
|
||
| **Knowledge** (`/api/knowledge`) | Product knowledge base, **global/shared** (no user isolation; writes record `created_by`/`updated_by` for audit). `POST /ingest/thread/{thread_id}` - sediment one thread into a note (reads messages from the checkpointer, `threads:read` owner-checked); `POST /ingest/threads/batch` - batch-sediment the caller's recent threads (`skip_existing` via `.manifest.json` dedup). **Defaults to `background=true`**: the slow per-thread extraction loop is deferred to a FastAPI background task and the endpoint returns `{queued:true, processed:N}` immediately (each thread's transcript is truncated to `max_input_chars` before the LLM call). Running the loop inline for a large batch outlives the nginx gateway timeout (504), which the frontend's `parseResponse` surfaces as a generic "Request failed"; `background=false` keeps the blocking path (full per-thread `created/skipped/failed` tally) for small batches/scripts. Shared loop: `_run_batch_ingest`. **Optional model selection**: both ingest endpoints accept `model_name` (overrides `knowledge.extract_model_name` for that capture; threaded through `KnowledgeService.capture_thread` ? `extract_thread_llm_multi`). **Model-error surfacing**: when the chosen model produces nothing usable (timeout / API error / non-JSON) capture still **falls back to rule-based extraction and lands a draft** (`llm_fallback:true`); the single-thread (non-background) response and the batch response (`llm_fallback` count) carry this so the UI tells the user "???????". The "??????" chat button stays background/fire-and-forget (passes `model_name`, no per-save error toast); the batch-import dialog is synchronous and shows the fallback count. `POST /ingest/file` - **import an uploaded file** (multipart: `file` + optional `title`/`tags`/`status`/`model_name`) into wiki note(s): the file is converted to Markdown (md/txt read directly, pdf/word/ppt/excel via `convert_file_to_markdown`), distilled by the same wiki-capture LLM pipeline as thread capture (`source_type="file"`), and falls back to a verbatim **draft** note when the model returns nothing usable (`llm_fallback:true`); `POST /ingest/search` - sediment search/tool results as a reference note; `POST /notes` - manual note; `GET /notes` (keyword/tag/source_type/status + pagination); `GET/PUT/DELETE /notes/{id}` (DELETE is soft = archive, `?hard=true` removes the vault file); `POST /search` - `mode=keyword\|vector\|hybrid` (vector?keyword fallback when no embedding endpoint); `GET /notes/{id}/versions` - edit history; `GET /duplicates` + `POST /merge` - dedup; `POST /reindex` - rebuild vector index; `GET /graph` - product graph data (notes + entities, graph-blacklist filtered); `GET/POST /graph/blacklist` + `DELETE /graph/blacklist/{id}` - graph node blacklist (labels = entity names / note titles, case-insensitive match; blacklisted nodes and their edges are dropped from `/graph` and the wiki-export artifacts; table `knowledge_graph_blacklist`, auto-created); `POST /export/graph` + `GET /export/files/{graph.json\|graph.html\|graph.graphml\|cypher.txt}` - obsidian-wiki `wiki-export` artifacts (self-contained `graph.html` for iframe); `POST /resolve` `{targets:[?]}` + `GET /notes/{id}/wikilinks` - resolve `[[wikilink]]` targets to a note id / vault page / missing (powers clickable wikilinks); `GET /vault-page?path=` - read a raw vault meta/hub Markdown page (no DB note). `GET/POST /folders` + `PUT/DELETE /folders/{id}` - user-managed directories (notes carry a `folder` path; `GET /notes?folder=` filters a dir + sub-dirs); `GET/POST /extract-templates` + `GET/PUT/DELETE /extract-templates/{id}` - editable extraction-prompt templates (thread/document ingest accept `template_id`; capture/PUT-note accept `folder`). Backed by the `deerflow.knowledge` module (see below). |
|
||
|
||
Proxied through nginx: `/api/langgraph/*` ? LangGraph, all other `/api/*` ? Gateway.
|
||
|
||
### TaskCOP Situation-Overview Compatibility
|
||
|
||
The standalone legacy Vue page calls the Gateway URL in its runtime `b1ConsumerUrl` configuration directly (no Vite `/api` reverse proxy). The Gateway CORS middleware merges standard local Vite origins (`localhost`/`127.0.0.1` on ports 5173, 5174, 3000, and 8080) even when `GATEWAY_CORS_ORIGINS` was set for another frontend. This avoids a restart regression when the existing Vite server is on 5173 but the launcher sets 5174. Deployments must set `GATEWAY_CORS_ORIGINS` to any additional exact frontend origins that are allowed to call the Gateway.
|
||
|
||
`app/gateway/routers/taskcop_tasks.py` is a temporary compatibility layer for the separately deployed legacy Vue page at `/situation-overview/action/sentiment` and its task import page at `/situation-overview/action/task/import`. It preserves the old Consumer paths and response envelopes: `POST /cop/saveSuperiorTask` creates a row with a sequential four-digit string id (`0001`, `0002`, ...) and returns `{state: 201, msg, data}`; `POST /api/taskcop/import/tasks` accepts a JSON array or `tasks`/`records`/legacy list envelope, uses each supplied `id` or `taskId` as the primary key, and completely replaces duplicate IDs; `DELETE /cop/tasks/{task_id}` removes that task and its shared report document; `GET /taskAnalyseSearch/cop-task-three-list` (all), `GET /taskAnalyseSearch/cop-my-tasks` (creator-scoped), and `GET /taskAnalyseSearch/cop-task-list` return `{state: 200, data: {total, records}}`; `GET /taskAnalyseSearch/cop-task-detail?taskId=` returns one task. Existing UUID-style task rows are retained, and do not affect the new numeric sequence. An empty TaskCOP store remains empty—there are no built-in task or report seed records. It supports the page's existing `content`, `taskStatus`, `taskDirection`, `startTime`, `endTime`, `pageNum`, and `pageSize` query fields. Tasks are stored in `taskcop_tasks` through `deerflow.persistence.taskcop_tasks`, registered by `persistence.models`, and initialized as `app.state.taskcop_task_store` in `deps.py`. A non-empty save/update/import through `/api/task-reports/*` calls `mark_analysis_completed(task_id)`: it changes only pre-completion task states to `25`, so the list renders and filters the task as “已完成” without downgrading later workflow states. `trendSucc` is derived from the same status for the legacy sentiment list. Zero-padded task ids stay strings in the report API, so `0001` is never converted to `1`. The Vue router calls DeerFlow's `/api/v1/auth/login/username` directly with an incoming `username`/`userId`/`yUserId` parameter; it does not depend on the retired Consumer login service, and sends the returned DeerFlow JWT to every TaskCOP call. The TaskCOP routes are no longer public or CSRF-exempt. The Gateway fallback CORS allow-list includes TaskCOP Vite origins `http://localhost:8080` and `http://127.0.0.1:8080`; deployments must use explicit `GATEWAY_CORS_ORIGINS`. Tests: `tests/test_taskcop_tasks.py` and `tests/test_taskcop_cors.py`.
|
||
|
||
### AI Writing Intent Router (对话驱动写作台意图路由)
|
||
|
||
Backs the frontend「AI 写作·对话版」page's unified bottom composer. `POST /api/ai-writing/sessions/{session_id}/intent` (owner-checked; in `routers/ai_writing.py`) reads the session's graph checkpoint via the shared `app.state.checkpointer` — the **pending interrupt is the authoritative pause point** (the frontend's claimed `pause_point` is only a fallback when the checkpoint can't be read) — then routes the user's free text through `deerflow/agents/ai_writing/intent_router.py::parse_user_intent`: an LLM call (session's `model_name` + `fallback_chain` model fault-tolerance) that returns `{intent, pause_point, action, payload, need_research, research_query, answer, confidence, clarify}`. Safety boundaries in `normalize_intent_result`: **per-pause action whitelist strictly aligned to graph.py's real resume actions** (`ALLOWED_ACTIONS`), no pause → no action, high-cost actions (`finalize` / `force_finalize` / query-less覆盖式 `re_search`) require confidence ≥ 0.85 else downgrade to `clarify`, payload key-whitelist cleaning (material-id filtering, review-issue index bounds, section-decision back-fill with the majority decision). QA intents return `answer` directly from the context snapshot (`build_context_snapshot` — stage-tailored: materials with ids / outline / draft body capped / review issues with indices). Model failures never 5xx — they degrade to clarify. Prompts in `agents/ai_writing/prompts/intent_router_prompts.py`; tests `tests/test_ai_writing_intent.py`.
|
||
|
||
### AI Writing Built-in Functional Agents (????????)
|
||
|
||
The AI-writing graph (`packages/harness/deerflow/agents/ai_writing/`) has four stage "brains" ? ?????? / ????? / ?? / ?? ? that can each run in one of two modes, switched per-stage by `config.yaml ? ai_writing.use_builtin_agents.{researcher,outliner,writer,editor}` (default **false**, mtime hot-reloaded, new sessions pick up changes immediately):
|
||
|
||
- **Legacy (flag off)**: hard-coded prompts in `prompts/*.py`, direct LLM call (unchanged behavior).
|
||
- **Builtin-agent (flag on)**: the node loads the corresponding functional agent from `.deer-flow/agents/ai-writing-<stage>/` ? `SOUL.md` becomes the system prompt and `config.yaml`'s `model` overrides the stage model (priority: `agent.model > node explicit fallback (keyword_model/rank_model) > state.model_name`). Admins edit SOUL/model from the agent management page. Runner: `nodes/_agent_runner.py` (`run_writing_agent` / `stream_writing_agent` / `parse_marker_json`). Output contracts (marker + ```json``` blocks, intent.py-style): researcher `[KEYWORDS_READY]` / `[MATERIALS_RANKED]`, outliner `[OUTLINE_READY]`, editor (per-dimension) `[REVIEW_READY]`; writer emits raw section Markdown (keeps per-section streaming) with the legacy `[[SECTION_BLOCKED]]` sentinel in strict mode. **Orchestration stays in the nodes** (search fan-out, per-section loop, 3-dimension review loop, all pause points) ? the agent only does single structured generations, and a missing/empty SOUL silently falls back to the legacy path (`WritingAgentUnavailable`).
|
||
|
||
Seeding mirrors the roundtable pattern: verified `config.yaml` + `SOUL.md` ship in `app/gateway/routers/_ai_writing_seed_assets/<id>/`; `_ai_writing_seed.py::ensure_ai_writing_functional_agents()` copies them to `.deer-flow/agents/` when missing (never overwrites local edits), called from the startup lifespan (before `_sync_legacy_agents`, so the agents appear in `/api/agents` on first boot) and self-heals from `GET /api/ai-writing/sessions` + `GET /api/ai-writing/article-types`. Tests: `tests/test_ai_writing_seed.py`, `tests/test_ai_writing_agent_runner.py`.
|
||
|
||
**Sample imitation mode (样文仿写).** The initial writing form supports `writing_mode_type=imitate`: the user can paste sample text or upload Word/PDF/Markdown/TXT. `POST /api/ai-writing/sample/extract` converts supported document files to Markdown via `deerflow.utils.file_conversion` and returns capped plain text to the frontend. The LangGraph entry point now starts at `sample_analyzer` before `intent_parser`; non-imitation sessions return immediately, while imitation sessions analyze up to 12k chars into `sample_style_profile` (`SampleStyleProfile` in `state.py`) and emit `sample_profile_ready`. `writer_outline` and `writer_draft` call `render_sample_style_section(state)` so both outline planning and section drafting receive the same style profile, imitation strength (`light|medium|strong`), optional structure-preservation preference, and guardrails: learn structure/tone/paragraph rhythm/sentence style, but do not copy sample sentences, proprietary facts, people, organizations, cases, or data. Frontend state mirrors this with `analyzing_sample`, uploaded/pasted sample controls, history/timeline restoration of `sampleStyleProfile`, and the setup-chat direct-start path preserves the same fields. Tests live with the knowledge-base AI-writing tests (`tests/test_ai_writing_knowledge_base.py`).
|
||
|
||
**Skill-driven retrieval (检索类型).** `material_source` has three values: `general` (通用检索, default), `knowledge_base` (知识库检索), and `notebook` (我的空间); legacy `skill` is normalized to `general` by the researcher node. **`knowledge_base` 直连内网 ES 接口(= 旧版「内网直连检索」)** — it bypasses skill orchestration entirely and calls `IntranetSearchProvider` over `ai_writing.intranet_search_url` (concurrent over the generated keywords; `intranet_knowledge_base`/`intranet_search_timeout` apply). **An empty/unconfigured `intranet_search_url` does NOT disable it** — the empty string is passed straight to `IntranetSearchProvider`, which falls back to its built-in default endpoint (`intranet_stub._DEFAULT_ENDPOINT`), exactly like the legacy `get_search_provider(search_provider=intranet)` path. So 内网直连检索 works out of the box without configuring an address (configure `intranet_search_url` only to point at a different endpoint). This is the "弱模型识别不了技能" escape hatch — the old no-skill retrieval path restored as an explicit option (weak models like deepseek-chat often fail to invoke configured skills). Tests: `tests/test_ai_writing_knowledge_base.py`. General retrieval is orchestrated from the **skills configured on the `ai-writing-researcher` agent** (`config.yaml → skills`, editable from the agent management page): the built-in `knowledge-base-search` skill (知识库检索) runs the native keyword fan-out web/intranet search (the old default path), while any other configured skill searches with a `description`-targeted query; skills run in configured order with early-stop at `max_materials`, and an empty/unresolvable skills config falls back to the built-in knowledge search so retrieval always works. The skill itself is seeded from `_ai_writing_seed_assets/skills/knowledge-base-search/SKILL.md` into `skills/public/` (never overwrites), and a pre-existing researcher `config.yaml` lacking the `skills` key is patched in place with the default (explicit `skills: []` is respected). **PAUSE-1 input append flow**: when the user types natural language into the material-confirm input (`re_search`/legacy `enable_skill_search` + `user_query`), the node matches the most relevant configured skill via `SkillPicker(candidates=…)` (skipped when only one skill is configured), searches with the user's words, and **appends** the new materials to the existing package (URL-dedup, original order preserved, no LLM re-rank/truncation, 100-item hard cap). **Skill-failure fault tolerance**: a configured skill that times out / raises / has all its keyword searches fail (vs. legitimately returning 0 results) no longer just silently moves on — the researcher node emits a `skill_search_step` of type `skill_error` (carrying the error `reason`), and if the configured skills collected **no** new materials at all it runs a `fallback` pass over other available skills (`_fallback_candidate_skills`: the built-in `knowledge-base-search` native path first, then other enabled catalog skills not yet tried) instead of erroring out. The frontend `SkillStepsCard` renders `skill_error` (rose) and `fallback` (amber) so the timeline shows "技能报错 → 改用其他技能搜集". Tests: `tests/test_ai_writing_material_source.py`.
|
||
|
||
The intranet client uses `trust_env=False` and defaults `intranet_verify_ssl=false`, so internal HTTPS endpoints with self-signed/internal-CA certificates do not fail on certificate verification; set `ai_writing.intranet_verify_ssl: true` only for environments that require strict certificate checks.
|
||
|
||
**Skill-call timeout** is hardcoded **300s (5 min)** — `researcher.py` `_SKILL_SEARCH_TIMEOUT_SECONDS` (per-skill wrapper) and `_skill_agent.py` `_REAL_AGENT_TIMEOUT_SECONDS` (whole real-agent skill run). The intranet ES fallback (`IntranetSearchProvider`) gets a longer **600s (10 min)** default (`intranet_stub.py`; overridable via `ai_writing.intranet_search_timeout`) — that interface is the slow one, so our own fallback timeout is generous to avoid being blamed for "检不出".
|
||
|
||
**Parallel intranet fallback (技能 0 条 → 直接用并行预跑的内网结果).** When `ai_writing.intranet_search_url` is configured, the researcher node launches the intranet ES search (`IntranetSearchProvider` over the generated keywords) as a **concurrent `asyncio.Task` the moment the skill loop starts** (not sequentially after all skills fail). If the configured + fallback skills collect **nothing**, the node awaits that already-running task and uses its results (`_collect`) — no extra wait. If the skills **did** get materials, the task is cancelled (`_cancel_intranet_task`, suppresses CancelledError) so the intranet call isn't wasted. URL empty → no task, behavior unchanged. Tests: `tests/test_ai_writing_intranet_parallel.py`.
|
||
|
||
**Partial-result harvest + query simplification (检不出/超时的容错).** When `researcher_skill_agent: true` runs a non-native skill via `run_skill_via_real_agent` (the deployment's real path — the skill, e.g. `knowledge-search-v2`, is invoked by the real agent through `bash python3 scripts/search.py "<query>"`), one writing run issues many queries and **some time out** (the skill script prints `{"success":false,"error":"请求超时…","count":0}`) while others succeed (`{"success":true,"count":30,"results":[…]}`). Two fixes so successful data is never dropped: (1) `_drive()` captures every **ToolMessage** stdout and `_harvest_tool_results()` parses the skill's JSON **by reusing the lead-agent 参考文献 mode's proven extractor** (`deerflow.runtime.references`: `_parse_json_like` + `_first_list` + `_normalize_reference_item`) — the same logic that already extracts skill JSON into citations 100% reliably on ChatPage/AgentChatPage, with its huge field-name union (title/m_title/name, content/page_content/content_preview/snippet/…, recUuid/uuid/docId/…) and per-skill display-config support — plus skips `success:false` outputs — so even if the model's final `{materials:[]}` summary gives up (confused by the timeouts), the successful queries' real results are still collected; the function merges model materials + harvested results (dedup by url/recUuid/title, capped at `max_results`). **The harvest+merge (`_finalize`) runs on EVERY exit path — normal end, overall timeout, AND `await task` exceptions (e.g. `GraphRecursionError`)** — so when the run is cut off before the model writes its summary (observed: 10 queries × ~2 graph steps > the old `recursion_limit`, so the agent errored out with the real 30+ results already in `tool_outputs`), those results are still salvaged instead of returning `[]`. `recursion_limit` was also raised from `_REAL_AGENT_MAX_TURNS` to `_REAL_AGENT_MAX_TURNS*2+4` (it counts graph super-steps ≈ 2 per model→tool round, not model turns) so the model normally has room to finish its summary. **Early-stop on `max_results`:** `_drive` re-harvests after every tool output and, once the deduped harvested count reaches `max_results` (= `per_skill` = `max_materials`), **breaks the nested `astream`** and emits a 「已获取 N 条素材,达到上限,停止本技能检索」 step — without it a weak model keeps issuing new queries forever (observed: 100+ step-bar entries, materials far over the limit). Because the skill now returns exactly `max_materials`, the researcher loop's own `collect_limit` early-stop fires right after, so later configured skills don't run either. (2) The model is instructed (in `_MATERIAL_OUTPUT_CONTRACT`) that a 0-result/timeout query must be **retried with a simpler, core query** (drop topic suffixes like 竞选/诉讼/政策, keep the core entity: 「特朗普竞选」→「特朗普」) and that any data already retrieved must always be kept. The same deterministic simplifier `search/base.py::simplify_query` is wired into `IntranetSearchProvider.search`: an empty first attempt retries once with the simplified query. **(3) Root cause of the instability — bash output middle-truncation.** The skill prints its results as JSON to stdout, and the `bash` tool **middle-truncates** any output over `sandbox.bash_output_max_chars` (default **20000**, inserting a `... [middle truncated …] ...` marker). A `count:30` skill response (~30 records × ~600 chars) exceeds 20000 → the JSON is cut in the middle → **unparseable as a whole** by both the model and `_try_load_json` → all materials lost (while a small `count:9` response parses fine — hence the flapping). Fix: when whole-JSON parse yields no items, `_harvest_tool_results` falls back to `_scan_balanced_json_objects` — a brace-depth/string-aware scanner that recovers **every individually-complete `{…}` record object** from the truncated text (the records on each side of the truncation marker are still valid JSON), so ~16+16 records survive a 20000-char cut, far above `max_materials`. This is why the lead-agent chat pages (`/page/workspace/chats/new`, agent chat) call the same skill with 100% success — they only need the model to summarize prose from whatever it sees, never to parse a structured list, so truncation doesn't hurt them; ai-writing needs the structured records, so it must salvage them itself. Tests: `tests/test_ai_writing_skill_real_agent.py`, `tests/test_ai_writing_query_simplify.py`.
|
||
|
||
**Agentic skill execution (`ai_writing.researcher_skill_agent`).** When this flag is on (the deployment `config.yaml` enables it), the directed-skill branch runs `nodes/_skill_agent.py::run_skill_via_real_agent`: it reuses `SubagentExecutor` with the full tool stack, sandbox middleware, and a one-skill whitelist so the skill can `read_file` its `SKILL.md`, call `bash`, MCP tools, or whatever the skill instructions require. The nested run now passes `agent_id=ai-writing-researcher` in both `config.configurable` and runtime `context`, matching normal AgentChat skill visibility/whitelist behavior. It also appends a human-turn hard directive listing the exact SKILL.md path and skill directory, requiring weak models to first `read_file` the SKILL.md and then `cd <skill-dir> && python/python3 scripts/*.py ...` when the skill is script-backed; the model is explicitly forbidden from returning `materials` before executing the skill's required tools. If the real-agent path fails or yields no materials, `_execute_skill` still falls back to the legacy description→query search so materials are never empty. The built-in `knowledge-base-search` native concurrent path is unchanged. Tests: `tests/test_ai_writing_skill_real_agent.py`.
|
||
|
||
**Targeted user revision (局部修改 — 只改命中章节).** When the user submits a free-text 修改意见 on the draft-confirm card (`action == "user_revise"`), `writer_draft_node` no longer always rewrites the whole article section-by-section. Before the full-rewrite loop it tries `_targeted_user_revision` (`nodes/writer_draft.py`): `_split_draft_sections` parses the previous `current_draft.full_markdown` into `(article_title, [(section_title, body)], references_block)` (level-1 `#` title, level-2 `##` sections, trailing `## 参考文献` pulled out and kept verbatim), then `_detect_revision_scope` runs a small `llm_json` classifier (prompts `DRAFT_REVISE_SCOPE_{SYSTEM,USER}`) returning **per-section targets** `[{index, new_title, rewrite_body}]`: it maps the note to the section indices it targets ("第一段/开头/引言"→0, "结尾/结论"→last, by title/topic) **and** distinguishes a **章节标题改名** (`new_title` set, `rewrite_body=false` → rename only, body untouched) from a **正文修改** (`rewrite_body=true`). Title-only renames apply directly to the assembled `## {new_title}` heading **and are written back into `current_outline.sections[i].section_title`** (so the renamed title persists through confirm/finalize/history — this fixes "改章节标题不生效"). Body rewrites run `_stream_revise_section` (prompt `DRAFT_REVISE_SECTION_USER`, streamed as `draft_chunk`). **Only hit sections are touched; unhit sections are preserved byte-for-byte and are NOT re-emitted to the stream** — the frontend clears the right canvas to its initial placeholder on entering `revising` and shows only the live rewrite of the changed section(s), so the user sees「清空 → 流式重写 → 终稿」rather than the whole article rebuilding. The classifier returns `whole_article=true` (整体性意见 like 「整体语气更正式」/「全文都改」, **structure changes** 加章/删章/调整顺序, or a **文章总标题** change) → `_detect_revision_scope` yields `None` → `_targeted_user_revision` returns `None` → the node **falls back to the whole-article rewrite** (which still runs `_revise_outline_with_notes` for title/structure changes). Also falls back when there is no prior draft, an empty note, or scope-detection errors. Editor-review revisions (`accept_review`) keep the holistic rewrite. Reassembly preserves the article title + structure and re-appends the original references block; word count is computed on the body (excluding references). Frontend: `isRevising` in `AIWritingState` (`types/ai-writing.ts`) is set in `useAIWritingStream.ts`'s `intervention_submitted` (`ns === 'revising'`) and cleared on `draft_ready`; `AIWritingDraftPanel.tsx` shows the initial `DraftWaitingPlaceholder` while `isRevising && streamingDraftSections.length === 0`, then the streaming preview once chunks arrive. Tests: `tests/test_ai_writing_targeted_revision.py`.
|
||
|
||
**History draft editing (历史回看可编辑成稿).** The right-side draft in 历史回看 is no longer read-only — the user can edit it like the live 写作完成态 (Tiptap + 目录 + 格式工具栏 + AI 润色) and **persist** the change back to the session. `PUT /api/ai-writing/sessions/{id}/draft` `{draft_markdown, draft_title?}` (owner-checked → 404 otherwise; `draft_title` omitted ⇒ keep原标题) updates the row via the existing `AIWritingSessionRepository.update`. Frontend: `AIWritingDraftPanel` now passes `readOnly={false}` + an `onSaveDraft` handler (only in history mode) → `TextArtifactRenderer` → `WritingToolbar` 「保存」button (loading + toast); the handler is `AIWritingHistoryContext.saveHistoryDraft(markdown, title?)` → `updateAIWritingDraft` (`api/ai-writing-sessions.ts`), which write-backs the returned summary to `historySession`. Edits to the **live** completed draft stay ephemeral (export-only) as before — only history mode shows 保存. Tests: `tests/test_ai_writing_draft_update.py`.
|
||
|
||
**Background suspend (后台挂起).** A writing session can be handed off to the server to auto-drive the remaining flow to finalize, surviving page leave / refresh / account switch. `POST /api/ai-writing/sessions/{id}/background` (owner-checked, idempotent) starts `AIWritingJobExecutor` (`app/gateway/ai_writing_job_executor.py`), wired in `deps.py` as `app.state.ai_writing_job_executor`. The executor **drives the same `ai_writing` graph in-process** — it builds `make_ai_writing_graph()`, attaches the shared `app.state.checkpointer` (the very saver the runtime worker uses, so it continues the real thread state the frontend already wrote), sets the owning user via `set_current_user`, and loops `aget_state` → `ainvoke(Command(resume=…))` until `status==done`. **Auto-action per interrupt**: material/outline confirm → `confirm`; section_help → every blocked section `loose` (continue with general knowledge, never re-loop to research); draft_confirm → `to_editor` until reviewed, then `finalize` (on pass, or when `revision_count >= max_revisions`); review_confirm → `accept_review` (revise & re-review until pass or budget exhausted, then forced finalize). To avoid racing a still-running runtime worker (suspend mid-node), the driver **waits for an actionable state** (`_wait_for_actionable`): it only acts when the graph is at an interrupt / done, or when the checkpoint stops advancing (no worker → abandoned/restart). No new table — the graph checkpoint + `ai_writing_sessions` already hold all state; `status="background"` marks 挂起中 (surfaced in the history dropdown), nodes write `draft`/`review_result`/`status=done` as usual, so the user reopens the finished article from history. Per-user concurrency cap `AI_WRITING_BACKGROUND_SLOTS_PER_USER` (default 2, same SQLite-lock rationale as the roundtable executor). Startup `reconcile` (in `deps.py`) relaunches drivers for sessions left `background` after a restart (`AIWritingSessionRepository.list_by_status`). **Progress flowchart popup**: `GET /api/ai-writing/sessions/{id}/progress` (owner-checked) reads the graph checkpoint via `executor.get_progress` and derives a canonical pipeline `stage` key (`intent/research/material_confirm/outline/outline_confirm/draft/draft_confirm/review/done`, from the pending interrupt's `pause_point` or `snapshot.next` node) + `note` (素材不足求助 / 编辑打回) + review progress + `isRunning`. Frontend: 「后台挂起」 button + auto-suspend on route unmount / `beforeunload` (`AIWritingPanel.tsx`, `suspendAIWritingToBackground[Keepalive]` in `api/ai-writing-sessions.ts`); 挂起中 badge in `AIWritingHistoryDropdown.tsx`, and **clicking a 挂起中 history record opens a flowchart dialog** (`AIWritingFlowDialog.tsx` polls `getAIWritingProgress` every 2.5s, renders the fixed vertical pipeline `AIWritingFlowChart.tsx` lit by the live `stage`) so the user sees how far the background run has gotten without reopening/taking over the session. **Transcript consistency**: the rich history timeline (`AIWritingHistoryView` → `AIWritingTimeline`) replays `transcript = {progressEvents, completedInterventions}`, normally assembled+saved client-side. A background run has no client, so the executor rebuilds it at the end (`_save_transcript`/`_build_transcript_from_checkpoint`) **entirely from the final checkpoint** — both `progressEvents` and `completedInterventions` are derived from the same set of `resumed` events (every pause-resume emits exactly one), so they are always 1:1 aligned. For each `resumed`: classify its message → `(pausePoint, action)` (`_classify_resumed`, matched against `graph.py`'s resume messages), insert an `await_user` (with the pause prompt) before it, and build a **rich** intervention card enriched from final state (material_confirm→`materials`, outline_confirm→`outline`, draft_confirm→`draftTitle`/`draftWordCount`/`reviewResult`, review_confirm→`reviewResult`); `materials_ready` is inserted before the first material-confirm. (Earlier code merged the frontend's pre-suspend `completedInterventions` with auto ones and aligned by index — but the frontend's last save can predate its most recent decision, so the counts drifted and cards rendered off-by-one with empty data; deriving everything from the checkpoint root-fixes both.) Raw (snake_case) events **and** intervention data are normalized on load by the **exported** `normalizeProgressDict` / `normalize{Material,Outline,ReviewResult}` (idempotent on already-camelCase live transcripts — one normalize path for both sources), so reopening a background-completed session shows the same cards as a frontend-driven one. Tests: `tests/test_ai_writing_background.py`.
|
||
|
||
### Admin Leaderboard Snapshots (`/api/admin/users/leaderboard*`)
|
||
|
||
The admin ????? is served from **precomputed per-day snapshots** in `admin_leaderboard_daily_stats` (keyed by `(stat_date, settings_hash)`), not by aggregating raw rows on every request. A background scheduler inside the Gateway (`_leaderboard_snapshot_scheduler_loop` in `app/gateway/app.py`, singleton file-locked) ticks every `ADMIN_LEADERBOARD_BACKGROUND_INTERVAL_SECONDS` (default 300s): it prewarms recent days, then builds any `queued`/`error`/stale-`running` day. `GET /leaderboard` merges the per-day snapshots for the requested range. **Today** is rebuilt whenever its snapshot is older than `ADMIN_LEADERBOARD_TODAY_TTL_SECONDS` (default 43200s = half a day, so the leaderboard refreshes ~twice daily); a past day, once `ready`, is treated as fresh forever.
|
||
|
||
Two on-demand/automatic refresh paths layer on top:
|
||
- **Manual "????"** ? `POST /leaderboard/refresh-today` force-rebuilds *today's* snapshot **synchronously** (bounded by the same semaphore/timeout the scheduler uses), then returns the merged range. Lets an admin pull the latest real-time records into the stats without waiting for the TTL. Frontend: a "????" button on `AdminUserLeaderboardPage` ? `useRefreshTodayLeaderboard` (`core/admin/hooks.ts`), which primes + invalidates the `["admin-activity","leaderboard"]` query.
|
||
- **Daily finalize of the previous day** ? once per day, on/after `ADMIN_LEADERBOARD_FINALIZE_HOUR` (local Beijing, default `1` = ?? 1 ?), the scheduler tick force-recomputes *yesterday* a single time so runs/tool-calls that only landed near midnight are fully captured. The decision is the pure `deerflow.persistence.admin_stats.finalize.should_finalize_previous_day` (status/generated_at/now/hour) ? the rebuild stamps `generated_at` to today, which self-debounces it for the rest of the day (no extra state file). Tests: `tests/test_leaderboard_finalize.py`.
|
||
|
||
### Admin Concurrency Monitor (实时并发监控 / 系统压力)
|
||
|
||
A live admin monitor embedded **at the top of `AdminUserLeaderboardPage`** (`components/workspace/admin/concurrency-monitor.tsx`, `ConcurrencyMonitorPanel`) that answers "当前有几个并发的大模型调用、是谁、能不能直接停掉" and charts system pressure over time. Built on the existing in-memory `RunManager` (no new run plumbing).
|
||
|
||
- **Live concurrency + 真正停止** — `RunManager.list_active()` returns every run whose status is `pending`/`running` (the registry keeps finished records around until `cleanup`, so it filters by status). `len(list_active())` is the current concurrent-call count. Router `app/gateway/routers/admin_active_runs.py` (`/api/admin/active-runs`, **admin only** via the shared `admin_users._require_admin`): `GET /` returns `{count, server_time, runs:[{run_id, thread_id, status, created_at, elapsed_seconds, user_id, user_name, agent_id, agent_name, thread_title, kind, multitask_strategy}]}` — each run best-effort enriched (batched DB) with the owning user's email (`threads_meta.user_id`→`users.email`) and the agent's display name (`agents.name`); `kind` is derived from thread metadata (圆桌会商 / 系统/后台 / 任务对话 / 对话). `POST /{run_id}/cancel` calls `RunManager.cancel(run_id, action="interrupt")`, which sets the abort event **and cancels the run's asyncio task** — i.e. it truly aborts the in-flight model stream, freeing capacity. `POST /cancel-all` stops every in-flight run.
|
||
- **并发量折线图 (按分钟/小时)** — a lightweight background **sampler** (`_concurrency_sampler_loop` in `app/gateway/app.py`, singleton file-locked, started in the lifespan) records `len(list_active())` every `CONCURRENCY_SAMPLE_INTERVAL_SECONDS` (default 15s) into the new `concurrency_samples` table (`deerflow.persistence.concurrency`, `ConcurrencySampleRow{sampled_at, active_count}`, auto-created by `create_all`, **no migration**; store wired as `app.state.concurrency_sample_store`). ~Hourly it purges rows older than `CONCURRENCY_SAMPLE_RETENTION_HOURS` (default 168h). `GET /concurrency?granularity=minute|hour&hours=N` aggregates samples into per-bucket **peak + avg** (`ConcurrencySampleStore.aggregate`, bucketed in Python so it's identical across sqlite/mysql/postgres; fetch capped at 60k rows). Default lookback: minute→3h, hour→48h.
|
||
- **大模型出字速度折线图 (tokens/sec, 按分钟/小时)** — `GET /token-speed?granularity=&hours=` aggregates the persisted `llm_call_metrics` (`status="success"`) into per-bucket avg/min/max tokens/sec (deriving tokens/sec from `output_tokens÷duration_ms` when the stored value is null). A dip pinpoints the model/provider as the bottleneck for that window — "知道具体原因出现在哪里".
|
||
- **On/off + 不影响线上** — the whole feature is designed to never impact production: sampling is one tiny INSERT, retention-purged, all paths best-effort/guarded, queries bounded. A **runtime switch** lives in `system_settings.json → concurrency_monitor.enabled` (`ConcurrencyMonitorSettings`, **default false / opt-in**, mtime-hot-reloaded): the background loop always runs (just sleeps) but only samples when an admin turns it on, and it re-reads the flag **every tick** so toggling it on/off **takes effect immediately without a restart** (the on-demand live snapshot + 停止 stay available regardless; only the historical chart needs sampling on). `GET|PUT /api/admin/active-runs/settings` reads/writes it (admin only); the panel's 「后台采样」 Switch drives it. Env `CONCURRENCY_SAMPLE_ENABLED=0` is a hard master kill (never starts the loop). Frontend client/hooks: `core/admin/active-runs.ts` + `core/admin/hooks.ts` (`useActiveRuns` polls 4s while the panel is open, `useConcurrencySeries`/`useTokenSpeedSeries` poll 15s, `useCancelActiveRun`/`useCancelAllActiveRuns`, `useMonitorSettings`/`useUpdateMonitorSettings`). Tests: `tests/test_concurrency_monitor.py`.
|
||
|
||
### Scheduled Task Runtime
|
||
|
||
`ScheduledTaskService` (`deerflow/runtime/scheduler/service.py`) polls due tasks every `poll_interval_seconds` and executes each in its own throwaway thread, serialized by a shared `_agent_run_lock`.
|
||
|
||
**Stuck-run recovery**: each execution attempt is bounded by `_RUN_ATTEMPT_TIMEOUT_SECONDS` (30 min) via `asyncio.wait_for`. On timeout the run is retried up to `_MAX_RUN_RETRIES` (3) times, then marked `failed`. A watchdog (`_reconcile_stuck_runs`, run every poll) recovers runs left in `running` with no in-process owner ? e.g. orphaned by a crash or restart ? re-driving them through the same retry budget so a task can never stay perpetually "executing". The `scheduled_task_runs.retry_count` column tracks consumed retries.
|
||
|
||
**Unique task names**: task names are unique per user (`uq_scheduled_tasks_user_name`). `create_task`/`update_task`/`admin_update_task` proactively check via `_assert_name_available()` and raise `DuplicateTaskNameError` on a collision; the router translates that into HTTP `409` instead of letting the raw DB `IntegrityError` surface as a `500`.
|
||
|
||
**HTML page tasks (`task_kind == "html_page"`)**: a second task kind that renders a full HTML document each run instead of a Markdown report. The kind + its config (`html_config`: `page_prompt`, `negative_prompt`, `reference_favorite_id`, `model_name`) live inside `execution_context` ? no schema change to `scheduled_tasks`. Data gathering is delegated to the agent's own search skills (driven by `page_prompt`), so there is no explicit search-query/data-source field; the render model is user-selectable (`model_name`, default `zai-org/GLM-5-FP8`). `_run_task_attempt()` dispatches by kind; `_run_html_page_attempt()` runs the pipeline: gather real data via the lead agent (its search tools) ? optionally load a reference favorite's reusable style/layout ? render with the chosen model (thinking off) ? fault-tolerant JSON parse ? sanitize (BeautifulSoup, strips `<script>`/`<iframe>`/`on*`/`javascript:`) ? structural validate ? one GLM5 repair pass if invalid ? **completeness gate**: if the HTML is still not a full valid document (missing doctype/html/head/body, empty body, or residual scripts) after the repair pass, raise `HtmlGenerationError` so the run is marked **failed** (a page only counts as successful when it is a complete, previewable HTML document) ? best-effort Playwright screenshot (no-op when Playwright absent) ? archive `index.html` ? persist a `scheduled_html_pages` row ? return `result` with `format: "html"` and the full HTML inline. Pure helpers live in `deerflow/runtime/scheduler/html_page.py`. The frontend previews the HTML in a sandboxed `iframe` (no `allow-scripts`) on the run-detail and public pages; a user can bookmark a page (style summary extracted via GLM5) and later pick it as a style reference. New tables `scheduled_html_pages` + `html_page_favorites` (`deerflow/persistence/html_pages/`, wired as `app.state.html_page_store`) are created by `Base.metadata.create_all` with no migration. Tests: `tests/test_html_page_tasks.py`.
|
||
|
||
**Unified output switches (`require_html` + `require_markdown`).** The task form no longer splits "???? / HTML ??" into separate kinds. `execution_context` now carries two **independent** booleans that may both be on: `require_html` (???? HTML) and `require_markdown` (???? Markdown). `_task_requires_html()` returns true for `require_html` **or** legacy `task_kind == "html_page"` (so old tasks need no data migration), and `_run_task_attempt()` dispatches to `_run_html_page_attempt` when it is true, else the regular agent attempt. When **both** switches are on, `_gather_html_data()` returns `(text, files)` ? the gather agent run is additionally told to save a `.md` file ? and `_run_html_page_attempt` validates it (`MarkdownRequirementError` on failure) and surfaces it alongside the generated page, so the run delivers both artifacts. The single main `prompt` doubles as the HTML page requirement (`html_config.page_prompt`); there is no separate page-prompt field. Tests: `tests/test_scheduled_task_require_html.py` + the markdown-coexistence cases in `tests/test_html_page_tasks.py`.
|
||
|
||
**Template-fill mode (JSON → 固定 HTML 模板填充).** A third execution mode, distinct from the free-form `require_html` render pipeline: convert a user-supplied JSON into a **fixed** HTML template's data structure and fill it. Triggered when `execution_context.template_config.template_id` is set (chosen in the task form after selecting the built-in **「模板数据转换器」** agent `template-json-builder`). `_task_template_id()` reads it and `_run_task_attempt()` dispatches to `_run_template_fill_attempt` **before** the `require_html` check. Pipeline: take the source (the task `prompt`, pasted or uploaded .json) — which may be **pure JSON, plain text, or text-with-embedded-JSON** (`template_fill.extract_embedded_json` pre-extracts any embedded `{…}`/`[…]` with a string-aware balanced scan and surfaces it as a labeled block while keeping the full original text; plain text → no block, the model reads it semantically) → load the selected agent's `SOUL.md`+model and run a **direct** `_invoke_conversion_model` call (system=SOUL, human=conversion prompt carrying the template's compact **schema skeleton** + the recognized-JSON block + source) → `extract_json_object` → `template_fill.validate_data` (top-level keys + container types) → **on failure feed the structure errors + `field_coverage_gaps` back and retry up to `_MAX_TEMPLATE_FILL_RETRIES`=2** → `template_fill.fill_template` replaces the template's `const DATA={…}` line (and refreshes `MANIFEST.系统.更新时间`) → archive `index.html` + persist a `scheduled_html_pages` row → return `format:"html"` (reusing the existing HTML preview at `/runs/{id}/html`). Any terminal failure raises `TemplateFillError` → run marked failed (no retry, message shown). **Pure engine** `deerflow/runtime/scheduler/template_fill.py` (harness, no `app.*`): the template registry (`TEMPLATES`: id=`no1` name=`no.1` 台湾滨海防卫大屏, and id=`no2` name=`no.2` 台湾飞行部队平台 with **`convertible=False`** — a 「只展示」 template that renders its existing HTML **as-is**, no model/validation/fill, until conversion is wired in later; `_run_template_fill_attempt` short-circuits to `read_template_text` for it, and `list_templates`/`is_convertible` expose the flag so the form hides the JSON input), `list_templates`/`reference_data`/`schema_skeleton_json`/`validate_data`/`field_coverage_gaps`/`fill_template` — **the template's own embedded `DATA` is the schema source of truth** (parsed at load, 1-sample-per-array skeleton fed to the LLM, top-level structure as the validation basis); template HTML files live under `runtime/scheduler/templates/<id>/`, add a `TemplateSpec` to register a no.2. The agent is seeded like the AI-writing ones: `app/gateway/routers/_template_builder_seed.py` + `_template_builder_seed_assets/template-json-builder/{config.yaml,SOUL.md}`, called from the lifespan before `_sync_legacy_agents` (which upserts it into the `agents` table as built-in/published so it shows in `/api/agents` → the task form's 执行智能体 dropdown). Template list API: `GET /api/scheduled-tasks/html-templates` → `{templates:[{id,name,description,convertible}]}` (open). **Debug-render API** `POST /api/scheduled-tasks/template-fill/render` `{template_id, text?|data?}` → `{ok, display_only?, errors, gaps, html}` (open): no model call — `extract_embedded_json(text)` (or `data`) → `validate_data` → `fill_template`, returning the filled HTML **plus** the validation errors/field-gaps (fills even when invalid so you can see the effect). Powers the chat-page debug button below; non-convertible templates return their raw HTML. **Templates as `require_html` 参考收藏页面 (use the REAL template, all sub-pages)**: the 参考收藏页面 list in the task form also offers the templates (value `tpl:<id>`, e.g. `tpl:no1`). These templates are **JS-driven multi-page SPAs** (sub-pages/tabs are rendered by script from `DATA`), so a free-form model render would only ever produce the main page and **drop every sub-page**. Therefore, when `_run_html_page_attempt` sees a `tpl:<id>` reference (`template_fill.reference_template_id` parses it; distinguishes a template marker from a real favorite id), it short-circuits **before** the free-form pipeline to `_run_html_template_reference`, which renders the **actual template file** (all sub-pages intact): convertible templates gather real data (`_gather_html_data`) → convert to the template `DATA` (`_convert_to_template_data`, shared with template-fill mode) → **`merge_with_reference`** (per top-level key = per sub-page: any sub-page/field with no extracted content falls back to the template's **original** data, so sub-pages never render blank) → `fill_template`; non-convertible templates (no.2) output the raw template as-is; if the task prompt **and** description are both empty, or gather/conversion yields nothing, it uses the template's original data outright. **数据抽取开关** `html_config.template_extract` (default **off**, shown in the form only when a `tpl:` reference is picked): it **only** controls whether the HTML uses **extracted** data vs the template's **default** data — it does **not** affect anything else. Off → HTML fills the template's default data (no convert); on → gather+convert+merge. **Markdown is independent**: when `require_markdown` is also on, the gather agent run still runs and the `.md` is produced/validated/surfaced **regardless** of the extract switch (so turning extraction off never blocks the Markdown). Display-only templates (no.2) never extract; they still honor `require_markdown`. `_run_html_template_reference` runs gather once when `do_extract` (extract on + has prompt/desc + convertible) **or** `require_markdown`, and threads any `.md` through `_finalize_template_html(extra_files=…)`. `merge_with_reference` is also applied in template-fill mode (`_run_template_fill_attempt`) so empty sub-pages there fall back too. Archive/persist/result are shared via `_finalize_template_html`. (The earlier `style_reference` CSS-extraction approach was removed — it lost the sub-pages.) Frontend: when the agent dropdown = `template-json-builder` the form hides the HTML/Markdown switches and shows a 「HTML模板」 select (`useHtmlTemplates`) + a .json upload (reads file text into the prompt box). Tests: `tests/test_template_fill.py`, `tests/test_template_fill_attempt.py`, `tests/test_template_builder_seed.py`.
|
||
|
||
### Per-user Prompts & UI Preferences
|
||
|
||
**Custom prompts can be scoped to one agent.** `user_custom_prompts` gains a nullable `agent_id` column (migration `20260603_02`): `NULL` = global (the chip shows in every chat), otherwise the prompt only surfaces when that agent is selected. The visibility filter is applied in the frontend `CustomPromptBar` (it receives the current agent id from `ChatPage`'s selector or `AgentChatPage`'s `assistantId`); the backend (`/api/user-prompts` CRUD) just stores/returns/clears `agent_id` (update accepts an explicit `null` to clear). Tests: `tests/test_user_prompts_agent_id.py`.
|
||
|
||
**Per-user UI preferences** (`deerflow/persistence/user_preferences/`, wired as `app.state.user_preferences_store`; migration `20260603_03`, also auto-created by `create_all`): one JSON-blob row per user so new toggles need no schema change. Router `/api/user-preferences` (`GET`/`PUT`, account-scoped). Current key: `hide_admin_prompt_prefix` ? when true the chat input hides the admin-configured prompt-prefix switch and never applies the prefix (the user uses only their own prompts; enforced in `ChatPage` for both `shouldApplyPrefix` and `promptPrefixProp`). Toggled from the "???????" dialog. Tests: `tests/test_user_preferences.py`.
|
||
|
||
**HTML skills (tools the render step uses).** Two offline skills under `skills/public/` shape the generated HTML and are toggled in the skills settings page (enabled by default): `html-page-builder` (a self-contained inline-CSS design system ? colors/cards/tables/KPI/responsive) and `html-charts` (inline-SVG charts that render in the no-`allow-scripts` preview iframe; optional local ECharts documented in its `assets/README.md`). A skill opts into HTML injection via `metadata.kind: html_page` in its SKILL.md frontmatter; `_collect_html_skill_instructions()` reads the bodies of **enabled** such skills and `build_generation_prompt(skill_instructions=?)` injects them, so the (raw-model) render step honors the user's skill selection. The data-gathering step and the render/repair/style passes all use the task's selected `model_name` (no silent fallback to the default model ? and there is **no hardcoded GLM-5**; an unset selection falls back to config `models[0]`). To stop auxiliary LLM calls from leaking onto the default model, **`TitleMiddleware` is skipped when `is_scheduled_run` is set** (`_build_middlewares` in `lead_agent/agent.py`) ? scheduled/HTML runs use throwaway threads whose titles are never shown, and `config.title.model_name` is usually unset (? would call `models[0]`, not the chosen model). The completeness gate (`validate_html`) also requires a non-empty `<title>` and forbids external stylesheets, on top of the doctype/html/head/body + non-empty-body + no-script checks.
|
||
|
||
### Knowledge Base (`packages/harness/deerflow/knowledge/`)
|
||
|
||
Self-contained product knowledge base. **All knowledge capability lives in this one module** (harness layer, `deerflow.*` only ? no `app.*` imports). Wired as `app.state.knowledge_service` in `deps.py` via `make_knowledge_service(sf, config.knowledge)` (returns `None` on the memory backend or when `knowledge.enabled=false` ? router 503). Surfaced by `app/gateway/routers/knowledge.py` (`/api/knowledge`). Config: `config.yaml` ? `knowledge` (`deerflow/config/knowledge_config.py`).
|
||
|
||
**Two stores, kept in sync on every write:**
|
||
- **DB mirror** ? tables `knowledge_notes` / `knowledge_sources` / `knowledge_tags` / `knowledge_note_tags` (`knowledge/models.py`, registered in `persistence/models/__init__.py`, auto-created by `Base.metadata.create_all`; portable types `PortableLongText`/`BeijingDateTime`). Powers listing/filter/search/audit. Global/shared ? no `user_id` column; `created_by`/`updated_by` are audit-only (resolved via `get_effective_user_id()`).
|
||
- **Markdown vault** ? obsidian-wiki compatible, default `{$DEER_FLOW_HOME}/data/knowledge/obsidian-vault`. `ObsidianWikiAdapter` (`obsidian_adapter.py`) writes wiki-capture frontmatter+body, and after **every** write syncs `index.md` / `log.md` (CAPTURE/INGEST/UPDATE/ARCHIVE/EXPORT) / `hot.md` / `.manifest.json` (source-hash dedup for `skip_existing`). `markdown_vault.py` handles frontmatter build/parse, slugging, wikilink extraction and **path-traversal-proof** resolution.
|
||
|
||
**Module layout:** `models.py` (ORM: notes/sources/tags/note_tags + embeddings/versions/entities/relations) ? `repository.py` (DB only, no LLM) ? `schemas.py` (NoteDraft/SourceDraft DTOs) ? `skill_templates.py` ? `markdown_vault.py` ? `manifest.py` ? `extractor.py` (rule-based fallback) ? `llm_extractor.py` (**default**: distills declarative knowledge per wiki-capture, structured JSON incl. entities/relations; falls back to `extractor.py`) ? `sources.py` ? `exporter.py` (`graph.json`/`graph.graphml`/`cypher.txt` + self-contained vanilla-JS `graph.html`; entity nodes + relation edges. **`graph.html` is theme-aware** ? it reads `?theme=light|dark` from its own URL since the embedding iframe can't inherit the app's CSS; `KnowledgeGraphPage.tsx` appends the active day/night mode and re-skins on toggle. The layout cools down and **freezes** so dense graphs stay clickable, node radius scales with degree, and hovering/clicking a node focuses it + its direct neighbours while dimming the rest + a detail panel ? taming heavy node/edge clutter) ? `chunking.py` + `embeddings.py` (OpenAI-compatible embedding client, keyword fallback) + `search.py` (keyword/vector/hybrid) ? `sensitive.py` (secret/PII redaction) ? `dedup.py` (Jaccard + hash dedup) ? `auto_ingest.py` (background queue/worker) ? `obsidian_adapter.py` ? `service.py` (`KnowledgeService` orchestrator ? the only thing the router touches; capture entry points `capture_thread` / `capture_search` / `capture_file` / `create_manual_note`, all converging on `_persist_draft`). Tests: `tests/test_knowledge.py` (20 cases). Frontend: `frontend-web/src/core/knowledge/*` + `pages/KnowledgeBasePage.tsx` / `KnowledgeGraphPage.tsx` + chat `knowledge-save-trigger.tsx` / `knowledge-rag-toggle.tsx`; spec at `frontend-web/docs/product-knowledge-base-dev.md`.
|
||
|
||
**Wikilinks & multi-note capture** (`knowledge.{related_notes_enabled,related_notes_limit,multi_note_enabled}`, all default-on; `project_name` defaults to `zncm`):
|
||
- **Navigable `[[wikilinks]]`** ? `service.resolve_wikilink(s)` maps a `[[target]]` to a knowledge note (by id ? `vault_path` ? title) or a vault meta page (`read_vault_page`, path-traversal-proof), else `missing`. `markdown_vault.normalize_wikilink_target` canonicalizes targets (drops `|alias`, slashes, `.md`). The frontend `MarkdownRenderer` renders non-numeric `[[...]]` as clickable links from the resolver map (numeric `[[123]]` stay citation badges); note links navigate, vault links open a `VaultPageDialog`, missing links are greyed.
|
||
- **Real related notes only (no structural meta-links)** ? `extractor.related_links()` now returns `[]`: notes no longer auto-link `index` / `_meta/taxonomy` / a project hub (they were noise in every `## Related`). On capture, `service._find_related_notes` cross-links **only by shared *distinctive* tags** (project/`knowledge-base` defaults excluded; free-text title/summary matching was removed because CJK tokenization linked unrelated notes via generic words like ??/????) and `_inject_related_notes` writes real `[[Title]]` links into `## Related` (plus note?note `related` graph edges); no distinctive tags ? no related notes. The renderers omit the `## Related` section entirely when there are no related notes and no sources. The vault's own `index.md`/`_meta/taxonomy.md` still exist (obsidian scaffold) and remain resolvable for manually-written links. The router enriches `GET /notes/{id}` with `created_by_name`/`updated_by_name` (auth-provider email lookup) so the UI shows a person, not a UUID. The frontend `MarkdownRenderer` renders wiki-capture provenance markers `^[inferred]`/`^[ambiguous]` as small "??"/"??" badges.
|
||
- **LLM-extraction outcomes** ? capture distinguishes three LLM outcomes: (1) **success** ? declarative note(s); (2) **deliberate rejection** ? the model returns `should_ingest:false` with no notes (greeting / meta chatter) ? `extract_thread_llm_multi` raises `LLMRejectedIngest` and `capture_thread` raises `EmptyThreadError` (skip ? nothing persisted, so worthless chats never enter the KB); (3) **genuine failure** ? API error / non-JSON ? returns `[]`, capture falls back to the rule-based verbatim extractor and lands the note as a **`draft`** (not `approved`), logs a WARNING, and returns `llm_fallback:true` (surfaced by the router). `llm_extractor` logs an actionable WARNING/INFO for each non-success path so extraction issues are diagnosable.
|
||
- **Multi-note split** ? `llm_extractor.extract_thread_llm_multi` lets one conversation distill into 1?4 cross-linked notes (`notes:[?]` JSON array; legacy single-object output still accepted). `capture_thread` persists each, sibling-cross-links them by title, and returns `{note, created, notes:[?]}`. Gated by `multi_note_enabled` (false ? first note only). Tests: `tests/test_knowledge_wikilinks.py`.
|
||
|
||
**Phases 2?4 (all config-gated, defaults safe):**
|
||
- **Phase 2 ? vector RAG**: `knowledge.embedding.{enabled,base_url,api_key,model,...}` enables an OpenAI-compatible embedding endpoint; notes are chunked + embedded on write (`knowledge_embeddings`), `POST /search` supports `mode=keyword|vector|hybrid` (vector degrades to keyword when unconfigured). `KnowledgeRagMiddleware` (lead-agent chain) injects a transient `<knowledge-context>` recall before the model call, gated **per-conversation** by the `knowledge_rag_enabled` runtime flag (frontend chat toggle) falling back to global `knowledge.rag_enabled`. `POST /reindex` rebuilds the index.
|
||
- **Phase 3 ? governance**: `sensitive.py` redacts secrets/PII before persistence (`sensitive_redaction_enabled`); every update snapshots `knowledge_note_versions` (`GET /notes/{id}/versions`); `GET /duplicates` + `POST /merge` for dedup; background `AutoIngestQueue` (wired in `deps.py`, started/stopped in the gateway lifespan) auto-sediments finished conversations **off the chat path** when `auto_ingest_enabled` ? enqueued from `KnowledgeRagMiddleware.after_agent`, value-judged by min-chars + `skip_existing`, landing as `draft`.
|
||
- **Phase 4 ? graph**: LLM extraction also yields entities/relations ? `knowledge_entities`/`knowledge_relations`; `exporter.build_graph` adds entity nodes + note?entity/entity?entity edges (reflected in `graph.html`/`graph.json`); `GET /graph` returns product graph data.
|
||
|
||
**User directories + editable extraction templates + display polish (knowledge UX overhaul):**
|
||
- **Custom directories (`knowledge_notes.folder` + `knowledge_folders`)** ? notes carry a free-form, slash-joined `folder` path (e.g. `投研/行业`) independent of the wiki-capture `category` and of `vault_path`. `knowledge_folders` is the manageable directory list (supports empty folders + rename/move that re-points descendant folders **and** filed notes). Router CRUD under `/api/knowledge/folders` (`GET`/`POST`/`PUT/{id}`/`DELETE/{id}?reassign_to=`); `GET /notes?folder=` filters a directory + its sub-folders (`folder=""` ? unfiled). Capture (`capture_thread`/`capture_file`/`create_manual_note`) and `PUT /notes/{id}` accept `folder` (the note PUT uses `model_fields_set` so an absent key leaves the folder untouched, a present `null` clears it). Frontend `buildKnowledgeTree(notes, folderPaths)` builds the left tree from `folder` (no longer from `vault_path`); folder management dialog + `FolderSelect` in the create/import dialogs and the note detail header.
|
||
- **Editable extraction prompts (`knowledge_extract_templates`)** ? the hard-coded distillation prompts are seeded as editable templates on startup (`KnowledgeService.seed_default_templates`, idempotent — `默认 · 对话提炼` / `默认 · 文档提炼`, each `is_default` for its scope). `scope ∈ {thread, document, both}`; capture resolves an explicit `template_id` → the scope default → the built-in constant, and passes it to `llm_extractor.extract_thread_llm_multi` / `extract_document_llm_multi` as `system_prompt` (used verbatim). Router CRUD under `/api/knowledge/extract-templates` (409 on duplicate name). Frontend: extraction-template management dialog + `TemplateSelect` in the batch/file import dialogs. Editing a template is the supported way to fix "提炼太精简" (loosen the prompt to keep more detail) and to add alternate extraction logics.
|
||
- **Single source list** ? `extractor._related_section_md` no longer emits a `### Sources` sub-section into the body; the note detail shows sources once via the structured `来源` block (`note.sources`). `## Related` stays pure wikilinks so the frontend extracts it cleanly into the related-docs module.
|
||
- **Chinese section headings (display-only)** ? the frontend `localizeHeadings` (wikiUtils) maps `Context/Finding / Decision/Reasoning/Implications/Related/Sources` to 中文 at render time (applied to the body fed to both `extractToc` and the viewer, so anchors stay consistent). No data migration — historical English-heading notes render in Chinese too.
|
||
- **Historical-data cleanup** ? `scripts/cleanup_knowledge_related.py` (one-off, idempotent, `--dry-run`/`--note-id`) strips legacy `## Related` structural links (vault scaffold: `index`/`_meta/taxonomy`/`projects/…`, any slash-bearing target) and the old `### Sources` sub-section from existing notes (DB `content_md` + vault files, via `KnowledgeService.update_note`).
|
||
- Migration `20260616_02` + idempotent `scripts/sql/column_additions_mysql.sql` cover `knowledge_notes.folder` (+ index); the two new tables are auto-created by `create_all` / `_ensure_orm_columns_sync` on startup. Tests: `tests/test_knowledge.py`.
|
||
|
||
### Custom Agent Persistence
|
||
|
||
Custom agent metadata lives in the `agents` ORM table (`deerflow.persistence.agents`). `agents.id` is the stable runtime id and the filesystem directory name under `.deer-flow/agents/{id}/`; `agents.name` is display-only, can contain Chinese characters, and is not unique. `user_id IS NULL` marks built-in agents, user-owned agents are visible only to their owner unless `published = true`, and update/delete operations require ownership. The runtime accepts `agent_id` in context/configurable, and non-default `assistant_id` values are treated as agent ids after visibility checks.
|
||
|
||
**Listing cache + pagination** ? to keep `GET /api/agents` fast, the gallery's per-card extras (skills + tool_groups) are denormalized into two cache tables instead of reading every agent's `config.yaml` on each request: `agent_skills` (relational, one row per (agent, skill) with `position`) and `agent_extras` (one row per agent: `tool_groups` JSON + `model` + `skills_synced_at`). `config.yaml` remains authoritative for the agent runtime; these tables are a read cache, write-through on create/update and backfilled from disk by `_sync_legacy_agents` (mtime-vs-`skills_synced_at` drift check). The list endpoint enriches via `AgentStore.get_extras_for([ids])` (two batched `IN` queries). Both tables are auto-created by `Base.metadata.create_all` (no migration needed). `GET /api/agents` supports **opt-in server-side pagination**: `scope` (`mine` = owned + favorited / `square` = built-in/published public), `page`, `page_size`, returning `{agents, total, page, page_size}`; omitting `scope`/`page` preserves the legacy "return everything" behavior. Name `search` matches the raw OR desensitized name (the frontend's static term?abbrev map is mirrored server-side in `deerflow/config/agent_name_desensitize.py`); `tag_ids` is an SQL OR-filter; `owner` keeps only agents published by that user id (the card's ??? ? built-in agents never match), powering the gallery's click-???-to-filter. `POST /api/agents/private-square/list` takes the same pagination params (incl. `owner`). Repository entry point: `AgentStore.list_paginated(...)`. Tests: `tests/test_agent_listing_cache.py::test_list_paginated_owner_filter_keeps_one_publisher`.
|
||
|
||
**Homepage selector (agent picker on the new-chat page)** ? single admin-curated list backed by `agents.featured_order` (nullable BIGINT). NULL = not featured. Only `published = true` agents can be featured (router enforces on writes); any published agent is eligible regardless of ownership. Cleared with `order: null`. The list rendered to end users is `WHERE featured_order IS NOT NULL` + visibility-filter, sorted by `featured_order ASC`. Admins reorder by sending the desired ids to `PUT /api/agents/featured/order` (assigns indices 0..N-1). BIGINT is used rather than INT so the management UI can seed initial values from `Date.now()` (overflows INT on MySQL).
|
||
|
||
### Publish Squares (????)
|
||
|
||
Admin-configurable, **access-control-free** categories that published agents/skills are filed under (distinct from the password-gated ????). The square list lives in `system_settings.json` ? `publish_squares` (`PublishSquaresSettings`: ordered `squares: [{id,name,description}]` + `default_square_id`); helpers `list_publish_squares` / `resolve_default_square_id` / `normalize_square_id` always yield a valid square (built-in "????" fallback when none configured). Admin CRUD: `GET/PUT /api/system-settings/publish-squares` (GET readable by any user ? the gallery sub-tabs + publish picker need it; PUT admin-only, auto-slugs missing ids and re-anchors a dangling default).
|
||
|
||
Each agent/skill carries a single `square_id` column (`AgentRow` / `SkillRow`, `String(64)` NOT NULL default `''`, migration `20260611_01`, idempotent SQL in `column_additions_mysql.sql`). One column = at most one square, so an item can never be "published twice". `''` = unassigned ? resolved to the default square at read/filter time. Writes normalize via `normalize_square_id`: agents on create/update (`square_id` field), skills on publish (`PUT /api/skills/custom/{name}/published` accepts optional `square_id`). Listing filters: `GET /api/agents?square=` and `GET /api/skills?square=` keep only that square's items; the **default** square additionally sweeps in legacy `''` rows (`square_is_default`). Tests: `tests/test_publish_squares.py`. Frontend: `core/system-settings` (`usePublishSquares`/`useUpdatePublishSquares`), admin `PublishSquaresSettingsPage` (embedded in the appearance/settings hub), square picker in the agent edit dialog + per-skill card, and a ??-tab square sub-filter in both galleries.
|
||
|
||
### Reference-mode Entry Switch
|
||
|
||
The 参考文献/RAG citation-mode composer entry is system-controlled through
|
||
`system_settings.json` → `citation_display.reference_mode_enabled` (default
|
||
`false`). `GET /api/system-settings/citation-display` remains readable so the
|
||
frontend can hide or show the entry consistently; `PUT` is admin-only and
|
||
partially updates `reference_mode_enabled` and/or `cleanup_words`. The normal
|
||
workspace input box, iframe chat toolbar, send context, right reference panel,
|
||
and appearance settings all use this flag, while already locked citation
|
||
threads and selected LLMWiki knowledge bases still render their source
|
||
documents in the reference panel. Tests:
|
||
`tests/test_prompt_prefix.py` and
|
||
`frontend-web/src/pages/iframe-chat-mode-buttons.test.ts`.
|
||
|
||
### Agent Module Tag Config (zzJC / zzXD 按类型聚合)
|
||
|
||
Backs the zzJC / zzXD agent-chat left sidebar (「按类型聚合智能体 + 聊天记录」). An admin picks which agent **tags** belong to a module; the module's left sidebar then shows only agents under those tags (empty/unconfigured = show all visible agents). Stored in `system_settings.json` → `agent_modules` (`AgentModuleSettings`: `modules: {module_key → [tag_id, ...]}`); pure helpers `get_module_tag_ids` / `set_module_tag_ids` (de-dupe + strip). Endpoints (in `app/gateway/routers/system_settings.py`): `GET /api/system-settings/agent-modules/{key}` (readable by any user — the sidebar needs it to filter), `PUT /api/system-settings/agent-modules/{key}` (admin only, full replacement of the tag-id list). Module keys are frontend-defined (`AGENT_MODULES` in `components/workspace/agents/agent-module-sidebar.tsx`, keyed by the module's landing agent id: `6bf`→`zzjc`, `10d69…`→`zzxd`). Tests: `tests/test_agent_modules.py`. Frontend: `core/system-settings` (`useAgentModuleConfig`/`useUpdateAgentModuleConfig`), admin `AgentModuleSettingsDialog` (gear button in the agent-chat header).
|
||
|
||
### Business-mapping authorization
|
||
|
||
`/api/business-mapping` remains globally readable. Updates to 3Q, 6BF, and 7BF require `system_role == "admin"`; the 8BF public business-chain mapping is intentionally different and can be updated only by the trusted authenticated account whose email prefix is `lqq`. Do not rely on a username supplied in an external login URL. The frontend follows the same email-prefix rule when deciding whether to render the 8BF configuration button, while the router (`app/gateway/routers/business_mapping.py`) remains the enforcement point. Tests: `tests/test_business_mapping.py`.
|
||
|
||
### Position-collaboration foundation
|
||
|
||
### Roundtable business-chain enable/disable
|
||
|
||
`roundtable_chains.enabled` (BOOLEAN NOT NULL DEFAULT 1, migration `20260819_02`)
|
||
is the per-chain 启停 switch. New and existing chains default to **enabled**.
|
||
Owner or admin can flip it via `PUT /api/roundtable-chains/{id}` `{enabled}`.
|
||
Disabled chains stay visible in the management list (with an 启用/停用 filter)
|
||
and still show `updated_at`; the frontend hides them from the picker / public
|
||
entry rail so they cannot be selected into a roundtable. Tests:
|
||
`tests/test_roundtable_chain_stages.py` (`test_create_defaults_enabled_then_toggle`).
|
||
|
||
### Roundtable business-chain human validation
|
||
|
||
The existing `/api/roundtable-chains` `seats` JSON also accepts the optional
|
||
boolean `human_validation` field on each seat (`ChainSeatSchema`). It is a
|
||
frontend Step 2 checkpoint, not a server-side approval task: after the marked
|
||
seat completes (or, in a DAG, after its stage completes), the client pauses and
|
||
uses the existing resume button or intervention composer to continue. No new
|
||
database column, migration, or validation card is required.
|
||
|
||
The position-collaboration workspace is a separate frontend route and must not
|
||
use the original roundtable coordinator or jobs. Its shared configuration
|
||
integration point is the existing `/api/roundtable-chains` store: each `seats` JSON item
|
||
may carry an optional `position_id` (validated by `ChainSeatSchema`, no schema
|
||
migration). The original roundtable only reads its established seat fields and
|
||
therefore remains unchanged. The separate `/api/position-roundtable` router
|
||
(also mounted at `/api/multi-agent/position-roundtable` for intranet gateway
|
||
compatibility)
|
||
persists task/intent/frozen-chain snapshots in `position_roundtable_sessions`
|
||
and direct-agent node indexes in `position_roundtable_nodes` (migration
|
||
`20260721_01`). Message history remains in the established LangGraph thread
|
||
store; a node completion unlocks its next chain stage and a later upstream
|
||
revision marks downstream records `stale` for refresh.
|
||
`POST /api/position-roundtable/sessions/{id}/nodes/{node_key}/reject` uses the
|
||
same cascade: its reason is intentionally optional, the selected node becomes
|
||
`rejected`, and downstream material becomes `stale` (or `locked` when no
|
||
material exists) with `invalidated_by` set. Clients must replace their complete
|
||
session snapshot from this response so the flow, delivery list, and active
|
||
node state stay consistent. A rejected/stale node returns to `done` only after
|
||
a fresh Markdown delivery is recorded.
|
||
For direct-agent handoff, the router resolves artifact paths only inside the
|
||
completed upstream node's own thread, then includes bounded UTF-8 excerpts for
|
||
text outputs; binary artifacts stay as metadata-only references. Because each
|
||
seat runs on an isolated LangGraph thread, a downstream `read_file` on the
|
||
upstream artifact path would fail with "File not found"; the handoff therefore
|
||
also **mirrors** every upstream `outputs/` deliverable into the current seat's
|
||
own `workspace/upstream/` directory (`_mirror_upstream_artifacts`, bytes from
|
||
the upstream sandbox file with the SQL recovery copy as fallback, idempotent
|
||
per turn via size comparison, capped at 8 files / 2 MB each / 8 MB total) and
|
||
rewrites each handoff artifact ref's `path` to the mirror (`upstream_path`
|
||
keeps the original; `mirrored_to_sandbox: true`), with a matching note in the
|
||
hidden instructions telling the agent it may `read_file` the full text. The
|
||
mirror deliberately stays outside `outputs/` so it never appears in the seat's
|
||
own delivery manifest.
|
||
`POST /sessions/{id}/archive` turns the session into a read-only historical
|
||
record: task/intent updates, node binding, hidden-context construction, and
|
||
turn completion are rejected, while read endpoints remain available.
|
||
Built-in summary and action-plan completion also preserves a visible report
|
||
answer when the model misses its final `write_file` call: the router writes
|
||
that exact answer as the required Markdown artifact in the owning thread's
|
||
`outputs/` directory, so users receive both chat text and a downloadable file.
|
||
**Per-turn outputs write cap (弱模型防重复产物)**: ordinary position-node chats
|
||
and the summary builtin send run-context `position_artifact_cap: 1` (whitelisted
|
||
in `services.merge_run_context_overrides`). `PositionArtifactCapMiddleware`
|
||
then keeps at most one successful `/mnt/user-data/outputs/` `write_file` **and**
|
||
one `present_files` per user turn — truncates excess tool calls in the same
|
||
model response (and syncs raw provider `additional_kwargs.tool_calls`), rejects
|
||
further writes/presents after an `OK` / `Successfully presented files` or an
|
||
earlier sibling in the same batch, ends the turn when only excess calls remain,
|
||
and reminds the model to present once (or wrap up if already presented).
|
||
**Action-plan does not send the flag** (it must write report `.md` +
|
||
`action-plan-subtasks.json`). Tests: `tests/test_position_artifact_cap.py`.
|
||
|
||
### 岗位 (Position) Subscription System
|
||
|
||
An admin-managed aggregation layer that decides, per user, which agents / skills
|
||
/ scheduled tasks / recommended questions they are subscribed to or can see.
|
||
**A user has at most one position** (`users.position_id`, 1:1); unassigned users
|
||
inherit the single `is_default` position. Note: the admin-configured prompts
|
||
shown to users ARE the **recommended questions** (`recommended_questions` table)
|
||
— there is no separate prompt-template store; the per-user `user_custom_prompts`
|
||
(personal 常用提示词 bar) is untouched by this system.
|
||
|
||
**Persistence** (`deerflow/persistence/positions/`, wired as
|
||
`app.state.position_store`; tables auto-created by `create_all`, migration
|
||
`20260616_01`):
|
||
- `positions` (`PositionRow`): `id`, `name` (unique), `description`,
|
||
`is_default` (only one at a time — the store clears the flag on others).
|
||
- `position_tag_bindings` (`PositionTagBindingRow`): one row per
|
||
`(position, resource_type, tag)`. **No binding for a resource_type = "use the
|
||
default set".** `POSITION_RESOURCE_TYPES = (agent, skill, scheduled_task,
|
||
recommended_question)` — the same set drives the tag types (`TAG_TYPES` was
|
||
extended with `recommended_question`).
|
||
|
||
**Sync engine** (`deerflow/services/position_sync.py`, wired as
|
||
`app.state.position_sync`). `sync_user(user)` **materializes** three resource
|
||
types into the existing "我的" tables, idempotently:
|
||
- agents → `agent_favorites` rows with `origin='position'`
|
||
- **skills → `skill_favorites` rows with `origin='position'`** (table
|
||
`SkillFavoriteRow`, migration `20260616_04`, auto-created by `create_all`)
|
||
- scheduled tasks → `scheduled_task_subscriptions` rows with `origin='position'`
|
||
|
||
Skills behave **exactly like agents**: a position only *adds* its tagged skills
|
||
to the member's 我的技能 — the browse-all 广场 is never restricted. (This replaced
|
||
the earlier query-time visibility filter, which wrongly shrank the square to the
|
||
bound skill.) `PositionSyncService.resolve_visible(user, type)` returns
|
||
`set | None` (None = unrestricted) and is now used **only** for the one remaining
|
||
query-time type, `recommended_question`.
|
||
|
||
Sync only ever adds/removes `origin='position'` rows, so a user's own
|
||
favorites/subscriptions are never clobbered. Default sets (no binding for a
|
||
type): agents = built-ins (`user_id IS NULL`); skills = empty (square stays
|
||
full, nothing added to 我的); scheduled tasks = empty; questions = **untagged
|
||
only** (默认岗位 / 无标签绑定 只看到未打标的题目;打标的题目只在绑定了对应标签的
|
||
岗位展示 — the router interprets `resolve_visible → None` as "show only untagged",
|
||
not "show everything").
|
||
`origin`/`position_id` columns were added to `agent_favorites` and
|
||
`scheduled_task_subscriptions` — see migration `20260616_01` +
|
||
`scripts/sql/column_additions_mysql.sql`. The running app also auto-adds these
|
||
via `_ensure_orm_columns_sync` on startup.
|
||
|
||
**Triggers**: assigning/removing members, changing tag bindings (re-syncs all
|
||
members), becoming the default position (`sync_default_members`), and login
|
||
(`auth.py::_sync_user_position`, best-effort so a brand-new user inheriting the
|
||
default gets their grants materialized).
|
||
|
||
**API** (`app/gateway/routers/positions.py`, `/api/positions`, **admin only**):
|
||
position CRUD; `GET/PUT /{id}/tag-bindings`; `GET/POST/DELETE /{id}/members`
|
||
(batch assign/unassign → triggers sync); `POST /{id}/resync`;
|
||
`GET /users` (all users + their position); `PUT /users/{user_id}` (set one
|
||
user's position). **`GET /me` is the one non-admin route** — it returns the
|
||
caller's own effective position (explicit assignment, else the default position
|
||
unassigned users inherit, else `null`); powers the top-bar 岗位 label
|
||
(`strategy-components/Header.tsx` user cluster, via `useMyPosition`). Related
|
||
changes: `GET /api/recommended-questions` is now
|
||
filtered by the caller's position (`?all=1` bypasses for admin management) and
|
||
admin can tag questions via `/api/tags/assignments/recommended_question/{id}`.
|
||
**Default-position rule**: tags are bulk-attached to **all** questions first,
|
||
then the visibility filter applies — a position binding recommended_question 标签
|
||
sees only those tagged questions; a 默认岗位 / 无绑定 (`resolve_visible → None`)
|
||
sees only the **untagged** questions (打标的题目不再泄漏到默认岗位). `?all=1`
|
||
(admin management) skips this filter and returns every question with its tags.
|
||
The list also accepts a repeatable `tag_ids`
|
||
OR-filter — so the 推荐问题配置 page can display per-question 标签 and search by
|
||
them (frontend `RecommendedQuestionsPage` + `TagSearchBar`). Tests:
|
||
`tests/test_recommended_questions_tags.py`.
|
||
`GET /api/skills` no longer restricts by position; each skill carries a
|
||
`favorited` flag (true = in the caller's 我的技能, e.g. position-granted) so the
|
||
skills page's 我的 tab includes owned ∪ position-granted skills while the 广场
|
||
stays complete. Frontend: `core/positions`
|
||
+ admin `PositionManagementPage` (侧边栏「综合管理 → 岗位管理」 + 右下角设置菜单).
|
||
Recommended questions stay rendered **above** the chat input (their original
|
||
position), just filtered by the caller's position. Tests:
|
||
`tests/test_positions.py`.
|
||
|
||
### Portable Column Types (`deerflow/persistence/types.py`)
|
||
|
||
Two `TypeDecorator`s keep ORM columns correct across the `sqlite` / `postgres` / `mysql` backends. **All ORM timestamp/JSON columns must use these, not the raw SQLAlchemy types.**
|
||
|
||
- **`PortableJSON`** ? native `JSON` on SQLite/Postgres; `TEXT` (with manual `json.dumps`/`loads`) on MySQL, where older servers lack JSON support.
|
||
- **`PortableLongText`** ? a large text column: maps to MySQL `LONGTEXT` (4 GB) and plain `TEXT` on SQLite/Postgres (already unbounded). Use for big JSON/markdown blobs that would overflow MySQL's 64 KB `TEXT` cap ? e.g. `ai_writing_sessions.transcript`.
|
||
- **`BeijingDateTime`** ? stores timestamps as Beijing (UTC+8) wall-clock time. The app records time with `datetime.now(UTC)`, but MySQL `DATETIME` / SQLite have no timezone and persist whatever wall clock the driver serializes ? so a raw `DateTime` column would store UTC and read 8 hours behind Beijing. `BeijingDateTime` converts UTC?Beijing on bind and re-attaches `+08:00` on result, so the database holds Beijing time while Python still gets timezone-aware datetimes (instant comparisons stay correct, including in WHERE clauses). PostgreSQL `timestamptz` stores the instant unambiguously, so values pass through unchanged there. Uses a fixed UTC+8 offset (China has no DST), so no `tzdata` dependency.
|
||
|
||
### Schema Migrations / Manual Column-Add SQL (`scripts/sql/`)
|
||
|
||
Gateway startup runs Alembic through `app.gateway.app._run_db_migrations` after ORM table creation. The migration graph must always have unique revision ids and exactly one head; `tests/test_migration_graph.py` enforces both. Historical duplicate revision `20260821_01` is repaired by the `20260821_03` / `20260821_04` fixed-question chain and merge revision `20260825_01`. Startup DDL has a bounded timeout and failures are logged without preventing the app from serving.
|
||
|
||
**Convention (required):** whenever code adds a field to an existing table (a new `mapped_column` on an ORM model), you must, in the same change:
|
||
1. Keep the matching alembic version file under `persistence/migrations/versions/` (guarded with `_has_table`/`_has_column`), and
|
||
2. Append the idempotent `ADD COLUMN` statement to **`scripts/sql/column_additions_mysql.sql`**, with a definition matching the ORM column exactly (type, length, nullability).
|
||
|
||
Operators upgrading an existing DB may run the whole file (`USE <db>;` then execute) as a recovery path; it is idempotent (existing columns are skipped via `information_schema.columns`). Index-only changes go in the sibling `scripts/sql/leaderboard_indexes_mysql.sql` (same idempotent pattern) and stay out of startup migrations because online index creation on large metric tables can delay service boot.
|
||
|
||
Leaderboard scheduled-run exclusion uses correlated `NOT EXISTS` against the
|
||
indexed `scheduled_task_runs.agent_run_id`; do not restore `NOT IN`. Legacy
|
||
scheduler detection for tool metrics is also a correlated existence check and
|
||
must not outer-join the complete `runs` table into metric aggregations. Keep the
|
||
created-at + dimension compound indexes in ORM metadata and the idempotent
|
||
MySQL index script synchronized.
|
||
|
||
### Sandbox System (`packages/harness/deerflow/sandbox/`)
|
||
|
||
**Interface**: Abstract `Sandbox` with `execute_command`, `read_file`, `write_file`, `list_dir`
|
||
**Provider Pattern**: `SandboxProvider` with `acquire`, `get`, `release` lifecycle
|
||
**Implementations**:
|
||
- `LocalSandboxProvider` - Singleton local filesystem execution with path mappings
|
||
- `AioSandboxProvider` (`packages/harness/deerflow/community/`) - Docker-based isolation
|
||
|
||
**Virtual Path System**:
|
||
- Agent sees: `/mnt/user-data/{workspace,uploads,outputs}`, `/mnt/skills`
|
||
- Physical: `backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/...`, `deer-flow/skills/`
|
||
- Translation: `replace_virtual_path()` / `replace_virtual_paths_in_command()`
|
||
- Detection: `is_local_sandbox()` checks `sandbox_id == "local"`
|
||
- **`{user_id}` for path resolution = `resolve_path_user_id(thread_id)`, using the persisted creator for every tracked thread.** The route's regular owner check still authorizes access first; creator resolution only selects the stable physical directory (`threads_meta.user_id`) and avoids a lost request context falling back to `users/default`. This also preserves cross-account task-deeplink conversations (rwfx/3qfx/xdfx) and roundtable system threads. Positive creator mappings are cached process-wide and DB/memory/missing-row failures fall back to the current login. It is wired at the thread-path consumers including `ThreadDataMiddleware.before_agent`, `present_file_tool`, uploads, artifact serving, and embedded `client.get_artifact`. **Rewrite-history recovery is metadata-only and job-owner scoped** (`list_by_session(..., user_id=current_user)`), so its GET/list endpoints intentionally skip redundant thread-owner checks and return an empty history instead of a history-only 404; the start/commit endpoints retain strict thread write authorization. Tests: `tests/test_thread_paths_owner.py`, `tests/test_thread_paths_roundtable.py`.
|
||
|
||
- **Durable document-rewrite commit:** `document_rewrite_pipeline.atomic_write_text` writes through a deliberately short sibling temporary filename so long Windows artifact paths stay below the path-length limit, and guarantees the target `outputs` parent exists. If the sandbox recreates that directory in the narrow commit window, it repairs the directory and retries; the original file is otherwise only changed through `os.replace`.
|
||
|
||
- **Deep Research structural regeneration:** `POST /api/deep-research/sessions/{id}/report-variants` creates a distinct completed session from the source report and its selected evidence. It rebases both saved `[来源:source_id]` and legacy streamed `[[source:source_id]]` markers to cloned ids, and removes inherited markers with no selected clone. The frontend uses this endpoint only for “选择结构再生成”, then queues the durable writer with `operation=generate`; generation receives the confirmed outline and complete selected evidence set but never receives the inherited report as text to edit. The new report's completed card can subsequently start a separate in-place `AI 全文改写`. Requirement plans are deterministic and operation-specific, model reasoning is forwarded as live SSE before the first Markdown token, and generation prompts emit display-ready `[来源:id]` citations. Citation validation repairs only a unique near-complete `drs_src_*` prefix (provider-truncated suffix) and still rejects genuinely unknown ids. The report-only stream requests 8192 output tokens (except Codex Responses models, which reject token-limit kwargs) and rejects a malformed trailing citation instead of committing a provider-truncated report. Never restore the former second planning-model call unless its latency and streaming contract are redesigned together.
|
||
|
||
**Sandbox Tools** (in `packages/harness/deerflow/sandbox/tools.py`):
|
||
- `bash` - Execute commands with path translation and error handling
|
||
- `ls` - Directory listing (tree format, max 2 levels)
|
||
- `read_file` - Read file contents with optional line range
|
||
- `write_file` - Write/append to files, creates directories
|
||
- `str_replace` - Substring replacement (single or all occurrences); same-path serialization is scoped to `(sandbox.id, path)` so isolated sandboxes do not contend on identical virtual paths inside one process
|
||
|
||
### Subagent System (`packages/harness/deerflow/subagents/`)
|
||
|
||
**Built-in Agents**: `general-purpose` (all tools except `task`) and `bash` (command specialist)
|
||
**Execution**: Dual thread pool - `_scheduler_pool` (3 workers) + `_execution_pool` (3 workers)
|
||
**Concurrency**: `MAX_CONCURRENT_SUBAGENTS = 3` enforced by `SubagentLimitMiddleware` (truncates excess tool calls in `after_model`), 15-minute timeout
|
||
**Flow**: `task()` tool ? `SubagentExecutor` ? background thread ? poll 5s ? SSE events ? result
|
||
**Events**: `task_started`, `task_running`, `task_completed`/`task_failed`/`task_timed_out`
|
||
|
||
### Tool System (`packages/harness/deerflow/tools/`)
|
||
|
||
`get_available_tools(groups, include_mcp, model_name, subagent_enabled)` assembles:
|
||
1. **Config-defined tools** - Enabled entries are resolved from `config.yaml` via `resolve_variable()`; entries with `enabled: false` are omitted before import and are not registered with the agent
|
||
2. **MCP tools** - From enabled MCP servers (lazy initialized, cached with mtime invalidation)
|
||
3. **Built-in tools**:
|
||
- `present_files` - Make output files visible to user (only `/mnt/user-data/outputs`)
|
||
- `ask_clarification` - Request clarification (intercepted by ClarificationMiddleware ? interrupts)
|
||
- `view_image` - Read image as base64 (added only if model supports vision)
|
||
4. **Subagent tool** (if enabled):
|
||
- `task` - Delegate to subagent (description, prompt, subagent_type, max_turns)
|
||
|
||
**Community tools** (`packages/harness/deerflow/community/`):
|
||
- `tavily/` - Web search (5 results default) and web fetch (4KB limit)
|
||
- `jina_ai/` - Web fetch via Jina reader API with readability extraction
|
||
- `firecrawl/` - Web scraping via Firecrawl API
|
||
- `configurable_search/` - Config-driven `POST` JSON `web_search` replacement. It deep-copies the configured payload and overwrites only `query`, supports HTTP/HTTPS plus configurable `verify_ssl`, timeout and result limit, normalizes nullable `recUuid`/`content`/`title`/`url` fields, and rewrites a result URL through `result_url_template` only when `recUuid` is non-empty. The required response envelope is `{"results": [...]}`. Tests use `httpx.MockTransport` and never require the deployment endpoint.
|
||
|
||
**ACP agent tools**:
|
||
- `invoke_acp_agent` - Invokes external ACP-compatible agents from `config.yaml`
|
||
- ACP launchers must be real ACP adapters. The standard `codex` CLI is not ACP-compatible by itself; configure a wrapper such as `npx -y @zed-industries/codex-acp` or an installed `codex-acp` binary
|
||
- Missing ACP executables now return an actionable error message instead of a raw `[Errno 2]`
|
||
- Each ACP agent uses a per-thread workspace at `{base_dir}/users/{user_id}/threads/{thread_id}/acp-workspace/`. The workspace is accessible to the lead agent via the virtual path `/mnt/acp-workspace/` (read-only). In docker sandbox mode, the directory is volume-mounted into the container at `/mnt/acp-workspace` (read-only); in local sandbox mode, path translation is handled by `tools.py`
|
||
- `image_search/` - Image search via DuckDuckGo
|
||
|
||
### MCP System (`packages/harness/deerflow/mcp/`)
|
||
|
||
- Uses `langchain-mcp-adapters` `MultiServerMCPClient` for multi-server management
|
||
- **Lazy initialization**: Tools loaded on first use via `get_cached_mcp_tools()`
|
||
- **Cache invalidation**: Detects config file changes via mtime comparison
|
||
- **Transports**: stdio (command-based), SSE, HTTP
|
||
- **OAuth (HTTP/SSE)**: Supports token endpoint flows (`client_credentials`, `refresh_token`) with automatic token refresh + Authorization header injection
|
||
- **Runtime updates**: Gateway API saves to extensions_config.json; LangGraph detects via mtime
|
||
|
||
### Skills System (`packages/harness/deerflow/skills/`)
|
||
|
||
- **Location**: `deer-flow/skills/{public,custom}/`
|
||
- **Format**: Directory with `SKILL.md` (YAML frontmatter: name, description, license, allowed-tools)
|
||
- **Loading**: `load_skills()` recursively scans `skills/{public,custom}` for `SKILL.md`, parses metadata, and reads enabled state from extensions_config.json. **A directory containing `SKILL.md` is a self-contained leaf** — discovery does not descend into its subdirectories, so a skill *package* that bundles its own `examples/`/`references/` sub-skills (each with a SKILL.md) surfaces as **one** skill, not a swarm of phantom entries.
|
||
- **Discovery↔delete consistency**: a skill is keyed for discovery by its frontmatter `name`, but its on-disk directory name may differ (hand-dropped / mis-extracted package). `LocalSkillStorage._resolve_skill_dir(name, category)` resolves the *discovered* directory (canonical `<category>/<name>` fast-path, else scan-by-frontmatter-name); `custom_skill_exists`/`read_custom_skill`/`delete_custom_skill` all go through it, so **anything that gets listed can also be read and deleted** regardless of directory name (the old code assumed `custom/<name>/` and 404'd "Custom skill 'x' not found." on any non-canonical layout). Delete is path-guarded to never escape the `custom/` root. Tests: `tests/test_skill_storage_delete.py`.
|
||
- **Injection**: Enabled skills listed in agent system prompt with container paths
|
||
- **Configurable `es_query` routing prompt**: `config.yaml → skills.es_query_routing`
|
||
contains an `enabled` switch and editable `prompt`. When enabled,
|
||
`lead_agent.prompt._build_es_query_routing_section` injects the rule into the
|
||
default lead-agent system prompt. Custom-agent-only prompts receive it only
|
||
when their explicit skill allowlist contains `es_query`; disabling the switch
|
||
or leaving the prompt blank emits no block. Immediately before an ordinary Q&A
|
||
graph is created, `lead_agent.agent._log_es_query_routing_registration` checks
|
||
the exact configured block against the final system-prompt string and logs
|
||
`enabled`, allowlist eligibility, `registered`, character count, and a short
|
||
SHA-256 fingerprint (never the prompt body). Tests:
|
||
`tests/test_es_query_routing_prompt.py`.
|
||
- **Installation**: `POST /api/skills/install` extracts .skill ZIP archive to custom/ directory. Direct uploads use the ownership-aware `POST /api/skills/validate?target_category=...` preflight plus `POST /api/skills/install-upload`: a same-owner custom conflict may be uploaded again with `overwrite=true`, while a different owner's conflict is non-overwritable and returns that owner's username (falling back to user id). The install endpoint repeats the ownership check, including for admins targeting `custom`, so callers cannot bypass the preflight. Public-skill overwrite remains admin-only.
|
||
- **Installable sentiment analysis package**: `skill-packages/sentiment-analysis.skill` is deliberately separate from the frontend virtual-agent chat page. The chat page goes through Gateway `POST /api/sentiment-agent/stream` (runtime-config URL/Basic Auth in the body, TLS verify off, raw AG-UI SSE passthrough). The skill package's self-contained `config/sentiment-agent.json` owns the external AG-UI address for agent-skill callers; Basic Auth is resolved from `SENTIMENT_AGENT_AUTH_USERNAME` / `SENTIMENT_AGENT_AUTH_PASSWORD` at runtime (not committed into the package). `scripts/sentiment_stream.py` uses only the Python standard library and converts upstream SSE to stable NDJSON events (`thinking_delta`, `answer_delta`, `done`), with a Markdown convenience mode for normal agent calls. Installation/calling guide: `skill-packages/sentiment-analysis/docs/调用与安装说明.md`.
|
||
- **Chinese name (`name_zh`)**: a Chinese display name stored in the **`skills` DB table** (`SkillRow.name_zh`, `String(128)`, migration `20260609_01`) ? never written to the on-disk SKILL.md, so it works for built-in skills too. Threaded through `SkillStore` create/ensure_legacy/update/update_admin exactly like `detail`; surfaced on `SkillResponse.name_zh` from the ownership record in `_skill_to_response`. Authorization mirrors `detail`: custom skill ? owner or admin; built-in/public ? admin only (row materialized on demand via `ensure_legacy`). Endpoints (in `app/gateway/routers/skills.py`): `PUT /api/skills/{name}/name-zh` (set/clear), `POST /api/skills/{name}/name-zh/generate` (LLM generates one), `POST /api/skills/name-zh/generate-missing` (????: **admin-only** bulk-generate for every skill that lacks one, bounded-concurrency `asyncio.gather`). **Optional model selection**: both generate endpoints accept an optional JSON body `{model_name}` ? an explicit name wins, else the curator model setting, else the main model (`_resolve_skill_name_zh_model_name`). **Model-error surfacing**: `_generate_skill_name_zh` raises `SkillNameZhModelError` when the model itself fails (build/timeout/API error) ? distinct from the model returning an unusable name (? `None`). The single endpoint maps that to `502 ?????,?????????:?`; the bulk endpoint fails fast on a model-build error (502) and otherwise carries a representative `model_error` string in `SkillNameZhBulkResponse` so the UI can say "N ?????????". The frontend (`skill-settings-page.tsx` list + `skill-detail-dialog.tsx` editor, plus the agent skill pickers/badges in `NewAgentPage.tsx` + `agent-card.tsx`) shows `name_zh || name`, offers per-skill "AI ??" plus an **admin-only** header "???????" button, and a compact model dropdown (`components/workspace/model-name-select.tsx`, default = configured model) beside each action.
|
||
|
||
#### Skill Address Management (技能地址管理 — URL/IP 批量替换)
|
||
|
||
Admin-only maintenance tool to find every `http(s)://` URL and IPv4 (`ip[:port]`)
|
||
address in the **running** skills (`skills/{public,custom}`, all `.md` + `.py`,
|
||
recursing into `scripts/`/`references/`/`templates/`) and rewrite them in bulk —
|
||
e.g. swapping a UAT endpoint for production across all skills at once. The same
|
||
address commonly lives in both a skill's `SKILL.md` and its `scripts/*.py`, so
|
||
results are **aggregated by unique address** with every occurrence (`file:line` +
|
||
snippet) and replaced together.
|
||
|
||
- **Pure engine** — `deerflow/skills/address_scan.py` (harness, no `app.*`):
|
||
`find_spans` / `scan_occurrences` (URL regex with trailing-punctuation strip;
|
||
IPv4 with 0–255 octet check + optional port; an IP inside a URL is *not*
|
||
separately surfaced), `apply_replacements` (**substring-safe**: re-detects exact
|
||
token spans and only rewrites a span whose text equals a requested `from`, so
|
||
`1.2.3.4` is never rewritten inside `11.2.3.40`), `infer_kind`,
|
||
`is_valid_replacement`. Tested in `tests/test_skill_addresses.py`.
|
||
- **Router** — `app/gateway/routers/skill_addresses.py` (`/api/skill-addresses`,
|
||
**admin only**). **Availability at scale (150+ skills):** every filesystem
|
||
walk/read/write runs via `asyncio.to_thread` so the event loop is never blocked.
|
||
Routes: `GET /scan` (aggregated, one-shot) + `GET /scan/stream` (per-skill NDJSON
|
||
progress `{type:start|progress|result}`); `POST /preview` (dry-run diff +
|
||
validation warnings); `POST /apply` + `POST /apply/stream` (per-file NDJSON
|
||
progress); `GET /backups`; `POST /rollback/{backup_id}`. **Apply safety:**
|
||
optimistic lock via `expected_total` (409 if files changed since the scan),
|
||
path-guard (`is_relative_to` the skills root), auto disk backup to
|
||
`{skills_root}/.address-edits/{timestamp}/` (the `.`-prefixed dir is excluded
|
||
from scans), then `refresh_skills_system_prompt_cache_async()` so edits take
|
||
effect immediately. Registered in `app.py` after `sensitive_words.router`.
|
||
**Per-edit history:** each apply also writes a sibling manifest
|
||
`{skills_root}/.address-edits/{timestamp}.manifest.json` (best-effort) capturing
|
||
`{created_at, replaced_total, replacements:[{from,to,kind}], changed_files}`; the
|
||
sibling-not-child placement keeps it out of both the restore set and the file
|
||
count. `GET /backups` returns each record enriched with `replaced_total` +
|
||
`replacements` (from the manifest; falls back to dir mtime when absent), and
|
||
`POST /rollback/{backup_id}` restores a single record — so the UI shows a full
|
||
modification history and rolls back any one entry independently.
|
||
- **Frontend** — the panel is the reusable `SkillAddressesPanel` (exported from
|
||
`pages/SkillAddressesPage.tsx`; its default export keeps the standalone route
|
||
`/page/strategy/admin/skill-addresses` as a direct-link fallback). It is
|
||
**embedded as a「技能地址管理」tab inside the admin「技能管理」page**
|
||
`pages/AdminSkillsPage.tsx` (Tabs: 技能列表 + 技能地址管理; reached from the
|
||
「技能管理」item in the「设置和更多」dropdown `workspace-nav-menu.tsx` →
|
||
`/page/workspace/admin/skills`) — it is no longer a separate entry in the
|
||
`components/page-sidebar.tsx` settings dropdown. The panel:
|
||
auto-scan on open with a **progress bar** (扫描逐技能 / 应用逐文件 X/N), URL/IP
|
||
type filter, search, a「片段批量套用」helper (e.g. all `uat.4.cn` → `prod.4.cn`),
|
||
per-address「替换为」input with expandable occurrences, inline 预览 panel, 应用,
|
||
**multi-keyword search** (space-separated, all-must-hit), and a **修改记录**
|
||
(modification history) panel — each apply is one record (time / which addresses
|
||
changed from→to / file & replacement counts) with a per-record 回退此次
|
||
(single-record rollback, with inline confirm). Client + NDJSON stream parsing in
|
||
`strategy-components/api/skill-addresses.ts`. Scope is the running skills only;
|
||
the `deploy/minimal/skills/` copy is not touched (sync separately).
|
||
|
||
#### Skill Compression (技能压缩)
|
||
|
||
When the skill catalogue grows large, injecting every enabled skill full-text
|
||
into the lead agent's system prompt bloats it and wastes tokens. **Skill
|
||
compression** is an admin-gated mode that keeps only *常驻(免压缩)* skills in the
|
||
prompt body and surfaces the rest on demand via keyword retrieval.
|
||
|
||
- **Master switch** — `system_settings.json → skill_compression`
|
||
(`SkillCompressionSettings`: `enabled` default **false**, `search_top_n`
|
||
default 5, `keep_index_in_prompt` default true). Read/written by
|
||
`GET`/`PUT /api/system-settings/skill-compression` (GET open to any user — the
|
||
skill detail page reads it; PUT admin-only, refreshes the prompt cache).
|
||
- **Per-skill 常驻 flag** — stored in the **DB**: `SkillRow.always_on`
|
||
(`Boolean`, migration `20260616_03`, idempotent SQL in
|
||
`column_additions_mysql.sql`), **not** in `extensions_config.json` (the
|
||
source-tree config file is read-only / unsafe to rewrite in some deployments,
|
||
and built-in skills carry no file entry). Threaded through `SkillStore`
|
||
create/ensure_legacy/update/update_admin + `list_always_on_names()` exactly
|
||
like `name_zh`/`detail`; surfaced on `SkillResponse.always_on` from the
|
||
ownership record in `_skill_to_response`. Toggled via
|
||
`PUT /api/skills/{name}/always-on` (custom skill → owner or admin; built-in /
|
||
public → admin only, materializes a row on demand via `ensure_legacy`). The
|
||
prompt builder stamps `Skill.always_on` from `list_always_on_names()` via a
|
||
sync DB bridge (`_apply_always_on` in `lead_agent/prompt.py`).
|
||
- **Prompt split** — `get_skills_prompt_section` (`lead_agent/prompt.py`) reads
|
||
the compression settings (skipped for restricted custom agents — they already
|
||
have a curated allowlist). When **off**, behaviour is unchanged (all enabled
|
||
skills full-text). When **on**, `always_on` skills render as full `<skill>`
|
||
blocks; the rest are dropped from the body, replaced by a `search_skills`
|
||
usage note plus (when `keep_index_in_prompt`) a name-only `<compressed_skills>`
|
||
index. The `_get_cached_skills_prompt_section` cache key includes the
|
||
`always_on` flag + compression flags so changes invalidate correctly.
|
||
- **Retrieval tool** — `search_skills` (`tools/builtins/skill_tools.py`,
|
||
`search_skills_impl`) keyword-scores `name + description` over enabled,
|
||
visible, non-archived skills (same visibility filter as `skill_list`) and
|
||
returns top-N `{name, description, category, location}`. It is registered in
|
||
`get_available_tools` **only when compression is enabled** (offline,
|
||
deterministic, zero external deps). The agent then `read_file`s the returned
|
||
`location`.
|
||
|
||
Tests: `tests/test_skill_compression.py`.
|
||
|
||
#### Skill Curator (`packages/harness/deerflow/skills/curator.py`)
|
||
|
||
Background skill lifecycle manager. The Gateway `lifespan` starts a resident
|
||
scheduler task (`_curator_scheduler_loop` in `app/gateway/app.py`) that ? after a
|
||
5-minute startup delay ? calls `maybe_run_curator()` every hour. A cross-process
|
||
file lock (`.curator_scheduler.lock`) elects a single leader worker so the
|
||
scheduler runs once per deployment. Per run:
|
||
|
||
- **Master switch** ? `skill_evolution.enabled` (default **`false`**) gates the
|
||
whole feature; `curator.enabled` (default **`false`**) gates the curator
|
||
specifically. Both must be `true` for the curator to act. `should_run_now()`
|
||
and `_curator_thread_target()` both enforce this ? so manual `POST
|
||
/api/curator/run` also respects the master switch. Config is
|
||
`config.yaml.skill_evolution` deep-merged with per-user
|
||
`users/{user_id}/skill_evolution.json`.
|
||
- **Per-user run lock** ? `.curator_state` carries `running` / `started_at`;
|
||
`_try_acquire_run_lock()` prevents the same user's curator from running
|
||
concurrently (scheduler vs. manual trigger, or double manual clicks). A lock
|
||
whose `started_at` is older than `_CURATOR_STALE_LOCK_HOURS` (6h) is treated as
|
||
process-crash residue and can be taken over.
|
||
- **Concurrency cap** ? `maybe_run_curator()` collects all due users and hands
|
||
them to `_dispatch_curator_runs()`, which runs them through a
|
||
`ThreadPoolExecutor` capped at `_CURATOR_MAX_CONCURRENCY` (5) ? avoids a
|
||
thundering herd of LLM calls when many users come due at once (e.g. first
|
||
deploy).
|
||
- **Failure back-off** ? `_release_run_lock()` always stamps `last_attempt_at`.
|
||
When a first run keeps crashing (`last_run_at` never set), `should_run_now()`
|
||
retries on `_CURATOR_RETRY_HOURS` (6h) instead of every hourly tick.
|
||
- **`should_run_now()`** is the real gate: the hourly tick only checks "is the
|
||
per-user `interval_hours` (default 168h) elapsed". Timestamp parsing tolerates
|
||
naive (tz-less) strings ? a corrupt `last_run_at` triggers a re-run rather than
|
||
a permanent deadlock.
|
||
- **Published-skill filter** ? `SkillStore.list_published_names()` is a batch
|
||
query; the curator fetches the published-name set once per run instead of one
|
||
DB round-trip per candidate skill.
|
||
- **LLM review timeout** ? `_run_llm_review()` runs `agent.invoke()` in a worker
|
||
thread bounded by `curator.review_timeout_seconds` (default 600s). The model's
|
||
per-request `request_timeout` can't bound a multi-turn React agent loop, so on
|
||
timeout the run is abandoned (a warning is written to `curator_log`, visible in
|
||
the Web UI) and the curator finishes cleanly + releases the run lock. The
|
||
worker context is copied via `contextvars.copy_context()` so tool calls still
|
||
resolve the correct `user_id`.
|
||
- **Review model** ? `_run_llm_review()` uses `curator.model_name` (a
|
||
`config.yaml` `models[]` entry name) if set, else the primary model; it is
|
||
editable from the Web UI evolution panel. The skill security scanner
|
||
(`scan_skill_content` in `security_scanner.py`) reuses the same
|
||
`skill_evolution.curator.model_name` setting ? review and moderation share one
|
||
model. There is no separate moderation-model config field.
|
||
|
||
### Model Factory (`packages/harness/deerflow/models/factory.py`)
|
||
|
||
- `create_chat_model(name, thinking_enabled)` instantiates LLM from config via reflection
|
||
- Supports `thinking_enabled` flag with per-model `when_thinking_enabled` overrides
|
||
- `force_disable_thinking=True` (kwarg) hard-disables thinking for models that explicitly declare `supports_thinking: true` even when the model config declares no `when_thinking_*`/`thinking` shape — many internal vLLM/Qwen models only set that flag, so `thinking_enabled=False` alone injects nothing and they keep reasoning on. The flag injects the OpenAI-compatible disable shapes (`extra_body.enable_thinking=false`, `extra_body.chat_template_kwargs.enable_thinking=false`, and `extra_body.thinking.type="disabled"`) only for classes that accept `extra_body`. Models with `supports_thinking: false`, especially generic OpenAI-compatible proxies such as LiteLLM, receive no vendor-specific thinking fields because those gateways may reject unknown parameters. Mirrors `TitleMiddleware._force_disable_thinking`. The lead agent reads the runtime flag `thinking_force_disabled` and forwards it (used by `/api/open/chat`)
|
||
- Supports vLLM-style thinking toggles via `when_thinking_enabled.extra_body.chat_template_kwargs.enable_thinking` for Qwen reasoning models, while normalizing legacy `thinking` configs for backward compatibility
|
||
- Supports `supports_vision` flag for image understanding models
|
||
- Config values starting with `$` resolved as environment variables
|
||
- Missing provider modules surface actionable install hints from reflection resolvers (for example `uv add langchain-google-genai`)
|
||
|
||
### Roundtable Model Fault Tolerance (会商模型自动容错 / 换下一个模型)
|
||
|
||
会商第二步(leader/seat)+ 第三步(summary/report/dashboard + 结构化抽取)单轮执行时,**某模型出问题就自动切到下一个模型继续**,并告知用户「『X』模型异常 → 已切换『Y』」。前台实时(`multi_agent.py`)与后台挂起(`engine.py` + `InProcessRoundtableGateway`)两条路径都覆盖。
|
||
|
||
- **可检测的模型失败**:`LLMErrorHandlingMiddleware` 在同一模型重试 3 次后**不抛异常**,而是返回兜底 `AIMessage` —— 现在给它打 `additional_kwargs[LLM_ERROR_MARKER="deerflow_llm_error"]=reason`(model_dump 保留,经 checkpointer 序列化仍在)。`is_llm_error_message(msg)` 据标记(回退按已知兜底文案正文)识别,返回 reason。
|
||
- **候选模型链**:`app/gateway/roundtable_model_fallback.py::fallback_model_chain(current)` 按 `config.yaml models[]` 顺序给出 `[当前模型, …其余]`(去重保序)。`reason_text(reason)` 翻译成简短中文。总开关 env `ROUNDTABLE_MODEL_FALLBACK`(默认开,`0/false/off` 关 → 单模型旧行为);`ROUNDTABLE_MODEL_FALLBACK_MAX`(0=试完全部)。
|
||
- **触发换模型的情形**:① 单轮 run 抛异常(start_run 失败 / idle 卡死 → `timeout`,其余 `exception`);② 兜底错误文案(`is_llm_error_message`);③ 交付无效(仅 seat/leader;leader 还含「既无正文又无派活」)。席位**只认剥掉 `<think>` 后的可见正文**——思考-only / 空串 / 断流都不算已交付(`deerflow.agents.roundtable_orchestrator.delivery`:`visible_seat_text` / `classify_seat_delivery`)。前台 `_special_run` 对此 yield `status:"invalid_delivery"` 且不广播;`GET /last-ai-message` 回 `{delivered, content, length, invalid_reason}`(思考-only 标 `thinking_only`,总控席位清单标 ⚠️)。后台引擎只把可见正文写入 `deliveries`,续轮用 `build_continue_prompt`:有无效席位时强制重派,禁止声称「已完成交付」。Step3 单例(report/summary/dashboard/ingest/结构化)的产物是落盘文件,**不**把「正文为空」当失败,只对真模型错误换模型。
|
||
- **后台**:`InProcessRoundtableGateway._run_turn_with_fallback(...)` 包住 `_run_agent_turn`,逐候选模型试;换模型落 `model_switch` 诊断、全失败落 `model_fallback_exhausted`。`run_leader/run_seat/run_report/run_summary/run_dashboard_data` 全走它(report/summary/dashboard 的产物自愈重试保留,并记住已切到的可用模型用于后续续写轮)。各 Turn/Result 带回 `switches: list[ModelSwitch]`,引擎 `_emit_switches` → `JobProgressWriter.model_switch(...)` 往研讨时间线插一条 `role="system"` 系统提示气泡。
|
||
- **前台**:`_special_run`/`_leader_run` stream 收尾后检测失败 → yield `{status:"model_switch", role, agent_name, failed_model, next_model, reason}` 帧并用下一个模型重新流式;终态帧带 `model_used`。结构化抽取 `flow_extract.py` 的 round1 `model.astream` 也包了同款换模型循环(emit `event:"model_switch"` 帧)。
|
||
- **Step2 单席位隔离**:`engine._run_recommend/_run_chain/_run_dag` 的每个席位执行包 try/except —— 一个席位彻底失败(落 `seat_failed` 诊断 + 置空交付)不拖垮整批/整链/整个作业。
|
||
- **前端**:`multi-agent.ts` 解析 `model_switch` 帧 → `onModelSwitch` 回调(清空失败尝试的半截气泡 + 累计的 tool_call delta,phase 退回 connected);`useStep2Orchestration` 弹 toast + 时间线插「系统提示」气泡;`useStep3Summary/Report/Dashboard` + `flow-extract.ts` 重置对应气泡并在步骤卡留可见提示;后台路径经 `job-live.ts::toStep2Dialogue` 把 `role="system"` 渲染成 amber「系统提示」气泡。
|
||
- Tests: `tests/test_roundtable_model_fallback.py`。
|
||
|
||
### vLLM Provider (`packages/harness/deerflow/models/vllm_provider.py`)
|
||
|
||
- `VllmChatModel` subclasses `langchain_openai:ChatOpenAI` for vLLM 0.19.0 OpenAI-compatible endpoints
|
||
- Preserves vLLM's non-standard assistant `reasoning` field on full responses, streaming deltas, and follow-up tool-call turns
|
||
- Designed for configs that enable thinking through `extra_body.chat_template_kwargs.enable_thinking` on vLLM 0.19.0 Qwen reasoning models, while accepting the older `thinking` alias
|
||
|
||
### IM Channels System (`app/channels/`)
|
||
|
||
Bridges external messaging platforms (Feishu, Slack, Telegram, DingTalk) to the DeerFlow agent via the LangGraph Server.
|
||
|
||
|
||
**Architecture**: Channels communicate with Gateway through the `langgraph-sdk` HTTP client (same as the frontend), ensuring threads are created and managed server-side. The internal SDK client injects process-local internal auth plus a matching CSRF cookie/header pair so Gateway accepts state-changing thread/run requests from channel workers without relying on browser session cookies.
|
||
|
||
**Components**:
|
||
- `message_bus.py` - Async pub/sub hub (`InboundMessage` ? queue ? dispatcher; `OutboundMessage` ? callbacks ? channels)
|
||
- `store.py` - JSON-file persistence mapping `channel_name:chat_id[:topic_id]` ? `thread_id` (keys are `channel:chat` for root conversations and `channel:chat:topic` for threaded conversations)
|
||
- `manager.py` - Core dispatcher: creates threads via `client.threads.create()`, routes commands, keeps Slack/Telegram on `client.runs.wait()`, and uses `client.runs.stream(["messages-tuple", "values"])` for Feishu incremental outbound updates
|
||
- `base.py` - Abstract `Channel` base class (start/stop/send lifecycle)
|
||
- `service.py` - Manages lifecycle of all configured channels from `config.yaml`
|
||
- `slack.py` / `feishu.py` / `telegram.py` / `dingtalk.py` - Platform-specific implementations (`feishu.py` tracks the running card `message_id` in memory and patches the same card in place; `dingtalk.py` optionally uses AI Card streaming for in-place updates when `card_template_id` is configured)
|
||
|
||
**Message Flow**:
|
||
1. External platform -> Channel impl -> `MessageBus.publish_inbound()`
|
||
2. `ChannelManager._dispatch_loop()` consumes from queue
|
||
3. For chat: look up/create thread through Gateway's LangGraph-compatible API
|
||
4. Feishu chat: `runs.stream()` ? accumulate AI text ? publish multiple outbound updates (`is_final=False`) ? publish final outbound (`is_final=True`)
|
||
5. Slack/Telegram chat: `runs.wait()` ? extract final response ? publish outbound
|
||
6. Feishu channel sends one running reply card up front, then patches the same card for each outbound update (card JSON sets `config.update_multi=true` for Feishu's patch API requirement)
|
||
7. DingTalk AI Card mode (when `card_template_id` configured): `runs.stream()` ? create card with initial text ? stream updates via `PUT /v1.0/card/streaming` ? finalize on `is_final=True`. Falls back to `sampleMarkdown` if card creation or streaming fails
|
||
8. For commands (`/new`, `/status`, `/models`, `/memory`, `/help`): handle locally or query Gateway API
|
||
9. Outbound ? channel callbacks ? platform reply
|
||
|
||
**Configuration** (`config.yaml` -> `channels`):
|
||
- `langgraph_url` - LangGraph-compatible Gateway API base URL (default: `http://localhost:8001/api`)
|
||
- `gateway_url` - Gateway API URL for auxiliary commands (default: `http://localhost:8001`)
|
||
- In Docker Compose, IM channels run inside the `gateway` container, so `localhost` points back to that container. Use `http://gateway:8001/api` for `langgraph_url` and `http://gateway:8001` for `gateway_url`, or set `DEER_FLOW_CHANNELS_LANGGRAPH_URL` / `DEER_FLOW_CHANNELS_GATEWAY_URL`.
|
||
- Per-channel configs: `feishu` (app_id, app_secret), `slack` (bot_token, app_token), `telegram` (bot_token), `dingtalk` (client_id, client_secret, optional `card_template_id` for AI Card streaming)
|
||
|
||
|
||
### Memory System ? Memory V2 (`packages/harness/deerflow/agents/memory/`)
|
||
|
||
Pluggable memory architecture (Hermes-style). See `docs/MEMORY_V2_DESIGN_ZH.md`
|
||
and `docs/HINDSIGHT_SETUP_ZH.md` for the full design.
|
||
|
||
**Components**:
|
||
- `provider.py` - `MemoryProvider` abstract base (12 lifecycle/hook methods)
|
||
- `manager.py` - `MemoryManager` orchestrates builtin + ?1 external provider; `build_memory_manager()` is the construction entry point
|
||
- `providers/builtin.py` - `BuiltinFileProvider`: local USER.md / MEMORY.md, ? delimiter, char quotas, file locking, atomic writes, frozen snapshot
|
||
- `providers/hindsight.py` - `HindsightProvider`: HTTP client to a self-hosted Hindsight Docker instance, dual-bank routing
|
||
- `security.py` - injection/exfiltration scan + invisible-unicode block for memory content
|
||
- `message_processing.py` - `filter_messages_for_memory` (last-turn extraction for Hindsight retain)
|
||
|
||
**Two stores (binary classification)**:
|
||
- `USER.md` at `{base_dir}/users/{user_id}/USER.md` ? user profile, shared across agents
|
||
- `MEMORY.md` at `{base_dir}/users/{user_id}/agents/{agent_id}/MEMORY.md` ? agent-private working notes
|
||
- `user_id` resolved via `get_effective_user_id()`; `"default"` in no-auth mode
|
||
|
||
**Write paths**:
|
||
- LLM actively calls the `memory` tool (`tools/builtins/memory_tool.py`) ? writes USER.md/MEMORY.md, mirrors to Hindsight
|
||
- LLM calls `hindsight_recall`/`reflect`/`retain` tools (`tools/builtins/hindsight_tools.py`) ? only in `tools`/`hybrid` memory_mode
|
||
- Hindsight `auto_retain` ? every turn auto-stored to the work bank (via `MemoryMiddleware.after_agent`)
|
||
|
||
**Injection**:
|
||
- builtin frozen snapshot ? injected into the system prompt via `lead_agent/prompt.py` `_get_memory_context`
|
||
- Hindsight dynamic recall ? injected per model call via `MemoryMiddleware.wrap_model_call` (transient ? not persisted to state, so it never leaks to the frontend)
|
||
|
||
**Hindsight dual-bank isolation** (B-scheme):
|
||
- global user bank `deerflow-user-{user_id}` ? cross-agent
|
||
- agent work bank `deerflow-work-{user_id}-{agent_id}` ? per-agent
|
||
- `hindsight-client` is an optional dependency; when missing, `HindsightProvider.is_available()` returns False and the system falls back to builtin-only
|
||
|
||
**Migration**: `PYTHONPATH=. python scripts/migrate_memory_to_v2.py [--all|--user-id X] [--dry-run]` converts legacy `memory.json` into USER.md/MEMORY.md.
|
||
|
||
**Configuration** (`config.yaml` ? `memory`, nested): `enabled`, `provider` (builtin/hindsight),
|
||
`builtin.{memory_char_limit,user_char_limit,...}`, `hindsight.{mode,api_url,api_key,bank_id_template,work_bank_id_template,memory_mode,...}`,
|
||
`injection.{enabled,max_tokens,include_builtin,include_hindsight,context_tag}`, `security.{scan_content,block_invisible_unicode}`.
|
||
Per-user runtime overrides (whitelisted fields) live at `{base_dir}/users/{user_id}/memory_config.json` ? see `get_effective_memory_config()`.
|
||
|
||
**Gateway API** ? see the Memory router below (`/api/memory/v2*`).
|
||
|
||
### Reflection System (`packages/harness/deerflow/reflection/`)
|
||
|
||
- `resolve_variable(path)` - Import module and return variable (e.g., `module.path:variable_name`)
|
||
- `resolve_class(path, base_class)` - Import and validate class against base class
|
||
|
||
### Config Schema
|
||
|
||
**`config.yaml`** key sections:
|
||
- `models[]` - LLM configs with `use` class path, `supports_thinking`, `supports_vision`, provider-specific fields
|
||
- vLLM reasoning models should use `deerflow.models.vllm_provider:VllmChatModel`; for Qwen-style parsers prefer `when_thinking_enabled.extra_body.chat_template_kwargs.enable_thinking`, and DeerFlow will also normalize the older `thinking` alias
|
||
- `tools[]` - Tool configs with `use` variable path and `group`
|
||
- `tool_groups[]` - Logical groupings for tools
|
||
- `sandbox.use` - Sandbox provider class path
|
||
- `skills.path` / `skills.container_path` - Host and container paths to skills directory
|
||
- `title` - Auto-title generation (enabled, max_words, max_chars, prompt_template)
|
||
- `summarization` - Context summarization (enabled, trigger conditions, keep policy)
|
||
- `subagents.enabled` - Master switch for subagent delegation
|
||
- `memory` - Memory system (enabled, storage_path, debounce_seconds, model_name, max_facts, fact_confidence_threshold, injection_enabled, max_injection_tokens)
|
||
- `auth_login` - Custom login modes layered on the base auth (`deerflow/config/login_config.py`):
|
||
- `username_login.{require_password, password}` - **Shared** password gate for `POST /api/v1/auth/login/username`. When `require_password=true`, the (otherwise passwordless) username login additionally requires `?password=<shared-secret>`; empty `password` falls back to `123ewq`. Gate only ? it does not change which account logs in. **Per-account override (账号级登录口令)**: independent of this shared gate, an admin can assign **one specific account** an individual **plaintext** password via the 白名单管理 page (`users.login_password` column, migration `20260625_04`, nullable). When the target account carries a `login_password`, `/login/username` **requires** `?password=` to match it **exactly** and that personal password **takes precedence over** the shared gate (the shared secret will not unlock that account); accounts with no personal password keep the shared-gate / passwordless behavior. The check happens before auto-register via the lookup-only helper `_lookup_user_by_username` (so a brand-new username is unaffected). Surfaced/managed admin-only: `GET /api/v1/auth/users` returns `UserListItem.login_password` (plaintext — the list is already admin-gated), and `PUT /api/v1/auth/users/{user_id}/login-password` `{password: str|null}` sets (non-empty) / clears (null·empty) it without touching `token_version`. Plaintext storage is intentional (the password travels in the login URL query anyway, so it stays re-copyable). The user logs in via the existing `#/login/<username>?password=xxxx`. Tests: `tests/test_login_personal_password.py`.
|
||
- `token_login.{enabled, token_info_url, service_authorization, timeout_seconds, verify_ssl, ca_cert_path, mock_tokens}` - Token-exchange login (`POST /api/v1/auth/login/token`). The gateway POSTs to `token_info_url` with the caller-supplied token placed directly in the `Authorization: Bearer <token>` header (no cookie; the static `service_authorization` field is deprecated/unused, kept only for backward compat), reads `data.username` from a `{"state":"200",...}` body, then signs the matching local user in (auto-registering if absent). `token_info_url` is HTTPS ? `verify_ssl: false` skips cert verification for self-signed/internal UAT hosts, or set `ca_cert_path` to a PEM bundle (takes precedence, keeps verification on). **`mock_tokens` (offline test mapping `{token: username}`)**: when an incoming token matches a key, `_fetch_username_from_token` returns the mapped username and **skips the `token_info_url` call entirely** — for environments where the upstream `getTokenInfo` host is unreachable (e.g. external networks). Empty in prod (no effect); seeded `{"123ewq": "admin"}` in this deployment so `/login/任意?authToken=123ewq` signs in as admin. **`UsernameLoginResponse` now also returns `username`** (the raw resolved login name, pre email-normalization) on both `/login/token` and `/login/username`, so the frontend can append a faithful `?username=` on jump links (按钮管理 追加登录态) instead of the lossy email local-part. See [docs/AUTH_LOGIN.md](docs/AUTH_LOGIN.md) for the full request/response contract.
|
||
|
||
Both endpoints are whitelisted in `auth_middleware._PUBLIC_EXACT_PATHS`. The shared `_get_or_create_user_by_username()` helper (in `routers/auth.py`) backs both `/login/username` and `/login/token`.
|
||
|
||
- `log_system` - Admin-only external log-system entry (`deerflow/config/log_system_config.py`): `{enabled, url, label}`. Surfaced read-only by `GET /api/system-settings/log-system` (admin only); the frontend renders it as a "????" item in the bottom-right settings menu (`workspace-nav-menu.tsx`) that opens `url` in a new tab. Disabled by default.
|
||
|
||
**`extensions_config.json`**:
|
||
- `mcpServers` - Map of server name ? config (enabled, type, command, args, env, url, headers, oauth, description)
|
||
- `skills` - Map of skill name ? state (enabled). 常驻(免压缩)/always_on lives in the `skills` DB table, not here.
|
||
|
||
Both can be modified at runtime via Gateway API endpoints or `DeerFlowClient` methods.
|
||
|
||
### Conditional WeKnora LLMWiki
|
||
|
||
Ordinary knowledge-base cards expose an owner/admin-only name/description edit dialog in `WeKnoraKnowledgeWorkspace.tsx`, reusing `useLlmWikiActions().update` and `PATCH /api/llmwiki/knowledge-bases/{id}`. The route trims names, rejects blank or reserved conversation-deposit names, writes to WeKnora before updating local metadata, and retains existing write authorization. A successful remote list is authoritative for every list scope, including `personal`: absent remote IDs are hidden without deleting mappings/references. Upstream failures remain errors. The grid also filters cached `remote_status=missing` rows and no longer renders stale-record cleanup cards. Regression coverage: `tests/test_llmwiki_knowledge_management.py`.
|
||
|
||
`llmwiki.weknora.api_base_url` is the provider switch: non-empty enables WeKnora, empty preserves the legacy LLMWiki UI. WeKnora is independently deployed and owns PostgreSQL/Redis/object storage/vector storage/parsers/models; DeerFlow config contains only `api_base_url` and optional `web_base_url`. DeerFlow deliberately sends no WeKnora credential, so the WeKnora endpoint must be restricted to a trusted network and allow DeerFlow traffic itself or through its reverse proxy. There is intentionally no `default_embedding_model_id`: `WeKnoraClient` queries `/api/v1/models` and uses WeKnora's active Embedding model when creating a KB; Wiki creation additionally resolves an active KnowledgeQA model inside WeKnora for synthesis. The native BFF is `app/gateway/routers/llmwiki.py`; ownership/publication mappings live in `deerflow.persistence.llmwiki`; the browser only sees DeerFlow mapping IDs. Knowledge-base owners publish immediately (`private` -> `published`) without a review queue, and can unpublish ordinary bases. The conversation-deposit KB is system-owned, always published, shared by all DeerFlow users, excluded from My Knowledge, and immutable through user/admin detail operations; recognition requires both the reserved name and `system` owner. Existing same-name remote mappings are adopted in place as the system deposit, preserving mapping ids and document links. Its user-facing description, generated Markdown source and tags use `cmzs` branding; list/detail reads synchronize stale remote descriptions, and search results, source previews and embedded-detail JSON recursively replace legacy brand text in historical conversation-deposit payloads. Deposit requests must reference a thread accessible to the current account. Personal knowledge-base list/search views are scoped to the current DeerFlow account even for admins; frontend LLMWiki caches include the live account id, and admin write power does not make other users' private KBs appear as owned. Only `system_role === "admin"` sees the WeKnora connection-config tab. Knowledge cards navigate to a dedicated WeKnora-style detail route. Wiki-enabled bases render and default to `Wiki -> documents -> graph`; standard bases omit Wiki and render `documents -> graph`. The detail page supports sequential multi-file upload with an aggregate result for writable ordinary bases, hides it for public read-only/system-deposit bases, and uses a generic loading label while the embed session starts. It also supports URL and published manual-Markdown imports, browses real Wiki pages/content, proxies authenticated original-file previews, pages through `/chunks/{knowledge_id}`, and renders the Wiki page graph from `/knowledgebase/{kb_id}/wiki/graph`; document details use a right sheet rather than the removed KB modal. The credential-free proxy uses one WeKnora service identity, so per-user separation is enforced by DeerFlow mappings while all remote KBs physically share that identity's WeKnora workspace. WeKnora iframe detail pages first call `/api/llmwiki/knowledge-bases/{id}/weknora/session` through the normal DeerFlow API auth path, then load `/platform/knowledge-bases/{weknora_id}` with a signed `deerflow_weknora_embed` cookie for `/platform`, `/assets`, `/locales`, root assets, and API calls. Injected fetch/XHR rewriting sends all iframe `/api/v1/*`, including colliding `/api/v1/auth/*`, through `/api/llmwiki/weknora-embed/api/v1/*`; Gateway auth/CSRF bypass only applies to those signed iframe proxy requests, the proxy still revalidates the mapped KB, and cookie security honors `X-Forwarded-Proto` for HTTPS reverse proxies. `llmwiki_search` is conditionally registered in `tools.py`, is read-only, and rechecks the current DeerFlow user, agent `llmwiki_knowledge_base_ids`, and thread-context selection on every call. Source previews prefer the cleaned Wiki/LLMWiki page text when a retrieved chunk maps to a Wiki page, with raw split chunks as the fallback. Agent bindings use an inline WeKnora-style searchable/grouped multi-select in the standard agent create/detail UI; new chats use the same interaction pattern as a compact composer dropdown, and selections are written to thread metadata/context and restored on reopen. The DeerFlow frontend is native React and the WeKnora frontend is not modified or embedded. Full deployment and behavior notes: `docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md`. Tests: `tests/test_llmwiki_weknora.py`.
|
||
|
||
The embedded detail proxy hides WeKnora's outer `.main > .aside_box` application sidebar and expands the route outlet, but deliberately keeps the Wiki page's internal directory/index sidebar. Do not add an HTML `sandbox` attribute to this trusted, permission-filtered detail iframe: a sandboxed ancestor causes Chromium's built-in Blob PDF viewer to render the browser-blocked page. The Gateway follows upstream redirects only for `/knowledge/{id}/preview` and `/knowledge-bases/{id}/files`, keeping PDF bytes and protected Wiki images inside the signed proxy instead of sending the browser to WeKnora object storage.
|
||
Native/private and public Wiki readers must rewrite protected Markdown image
|
||
sources to their mapping-scoped DeerFlow `/files?file_path=` endpoints and keep
|
||
real HTTP URLs on rendered image elements. The embedded guard also rewrites
|
||
protected fetch/XHR values and hydrates late `img[data-protected-src]`
|
||
placeholders with the signed iframe HTTP proxy URL. Observe both
|
||
`data-protected-src` and `src`, because WeKnora may replace the URL after first
|
||
paint. Never create/revoke a temporary object URL here: its lightbox reuses the
|
||
thumbnail `src` after click. Never expose the service token or relax the
|
||
current-KB authorization to make images work.
|
||
|
||
**Embedded native knowledge-base switching:** The original WeKnora breadcrumb picker remains available inside the detail iframe. A signed 15-minute embed cookie carries only the current DeerFlow user's visible mappings (their own plus published bases); the proxied `/api/v1/knowledge-bases` response is filtered to that allowlist. Choosing another entry performs a full iframe navigation, rechecks the DeerFlow mapping and read/write permission, then reissues the cookie for the selected base. Never expose the upstream service-admin list directly or permit an arbitrary remote knowledge-base ID.
|
||
|
||
The independent enterprise-research workbench is **temporarily disabled**. Its source remains under `app/gateway/routers/enterprise_research.py` and `deerflow.persistence.enterprise_research`, but Gateway does not mount `/api/enterprise-research`, initialize its stores/executor, or register its ORM rows in `Base.metadata`. This ensures startup does not create or alter `enterprise_research_tasks` / `enterprise_research_report_jobs` on MySQL. It remains separate from `/api/deep-research`; do not re-enable it incidentally while changing the active Deep Research workflow. Regression coverage for its preserved implementation remains in `tests/test_enterprise_research_tasks.py`.
|
||
|
||
### Embedded Client (`packages/harness/deerflow/client.py`)
|
||
|
||
`DeerFlowClient` provides direct in-process access to all DeerFlow capabilities without HTTP services. All return types align with the Gateway API response schemas, so consumer code works identically in HTTP and embedded modes.
|
||
|
||
**Architecture**: Imports the same `deerflow` modules that Gateway API uses. Shares the same config files and data directories. No FastAPI dependency.
|
||
|
||
**Agent Conversation**:
|
||
- `chat(message, thread_id)` ? synchronous, accumulates streaming deltas per message-id and returns the final AI text
|
||
- `stream(message, thread_id)` ? subscribes to LangGraph `stream_mode=["values", "messages", "custom"]` and yields `StreamEvent`:
|
||
- `"values"` ? full state snapshot (title, messages, artifacts); AI text already delivered via `messages` mode is **not** re-synthesized here to avoid duplicate deliveries
|
||
- `"messages-tuple"` ? per-chunk update: for AI text this is a **delta** (concat per `id` to rebuild the full message); tool calls and tool results are emitted once each
|
||
- `"custom"` ? forwarded from `StreamWriter`
|
||
- `"end"` ? stream finished (carries cumulative `usage` counted once per message id)
|
||
- Agent created lazily via `create_agent()` + `_build_middlewares()`, same as `make_lead_agent`
|
||
- Supports `checkpointer` parameter for state persistence across turns
|
||
- `reset_agent()` forces agent recreation (e.g. after memory or skill changes)
|
||
- See [docs/STREAMING.md](docs/STREAMING.md) for the full design: why Gateway and DeerFlowClient are parallel paths, LangGraph's `stream_mode` semantics, the per-id dedup invariants, and regression testing strategy
|
||
|
||
**Gateway Equivalent Methods** (replaces Gateway API):
|
||
|
||
| Category | Methods | Return format |
|
||
|----------|---------|---------------|
|
||
| Models | `list_models()`, `get_model(name)` | `{"models": [...]}`, `{name, display_name, ...}` |
|
||
| MCP | `get_mcp_config()`, `update_mcp_config(servers)` | `{"mcp_servers": {...}}` |
|
||
| Skills | `list_skills()`, `get_skill(name)`, `update_skill(name, enabled)`, `install_skill(path)` | `{"skills": [...]}` |
|
||
| Memory | `get_memory()`, `reload_memory()`, `get_memory_config()`, `get_memory_status()` | dict |
|
||
| Uploads | `upload_files(thread_id, files)`, `list_uploads(thread_id)`, `delete_upload(thread_id, filename)` | `{"success": true, "files": [...]}`, `{"files": [...], "count": N}` |
|
||
| Artifacts | `get_artifact(thread_id, path)` ? `(bytes, mime_type)` | tuple |
|
||
|
||
**Key difference from Gateway**: Upload accepts local `Path` objects instead of HTTP `UploadFile`, rejects directory paths before copying, and reuses a single worker when document conversion must run inside an active event loop. Artifact returns `(bytes, mime_type)` instead of HTTP Response. The new Gateway-only thread cleanup route deletes `.deer-flow/threads/{thread_id}` after LangGraph thread deletion; there is no matching `DeerFlowClient` method yet. `update_mcp_config()` and `update_skill()` automatically invalidate the cached agent.
|
||
|
||
**Tests**: `tests/test_client.py` (77 unit tests including `TestGatewayConformance`), `tests/test_client_live.py` (live integration tests, requires config.yaml)
|
||
|
||
**Gateway Conformance Tests** (`TestGatewayConformance`): Validate that every dict-returning client method conforms to the corresponding Gateway Pydantic response model. Each test parses the client output through the Gateway model ? if Gateway adds a required field that the client doesn't provide, Pydantic raises `ValidationError` and CI catches the drift. Covers: `ModelsListResponse`, `ModelResponse`, `SkillsListResponse`, `SkillResponse`, `SkillInstallResponse`, `McpConfigResponse`, `UploadResponse`, `MemoryConfigResponse`, `MemoryStatusResponse`.
|
||
|
||
## Development Workflow
|
||
|
||
### Test-Driven Development (TDD) ? MANDATORY
|
||
|
||
**Every new feature or bug fix MUST be accompanied by unit tests. No exceptions.**
|
||
|
||
- Write tests in `backend/tests/` following the existing naming convention `test_<feature>.py`
|
||
- Run the full suite before and after your change: `make test`
|
||
- Tests must pass before a feature is considered complete
|
||
- For lightweight config/utility modules, prefer pure unit tests with no external dependencies
|
||
- If a module causes circular import issues in tests, add a `sys.modules` mock in `tests/conftest.py` (see existing example for `deerflow.subagents.executor`)
|
||
|
||
```bash
|
||
# Run all tests
|
||
make test
|
||
|
||
# Run a specific test file
|
||
PYTHONPATH=. uv run pytest tests/test_<feature>.py -v
|
||
```
|
||
|
||
### Running the Full Application
|
||
|
||
From the **project root** directory:
|
||
```bash
|
||
make dev
|
||
```
|
||
|
||
This starts all services and makes the application available at `http://localhost:2026`.
|
||
|
||
**All startup modes:**
|
||
|
||
| | **Local Foreground** | **Local Daemon** | **Docker Dev** | **Docker Prod** |
|
||
|---|---|---|---|---|
|
||
| **Dev** | `./scripts/serve.sh --dev`<br/>`make dev` | `./scripts/serve.sh --dev --daemon`<br/>`make dev-daemon` | `./scripts/docker.sh start`<br/>`make docker-start` | ? |
|
||
| **Prod** | `./scripts/serve.sh --prod`<br/>`make start` | `./scripts/serve.sh --prod --daemon`<br/>`make start-daemon` | ? | `./scripts/deploy.sh`<br/>`make up` |
|
||
|
||
| Action | Local | Docker Dev | Docker Prod |
|
||
|---|---|---|---|
|
||
| **Stop** | `./scripts/serve.sh --stop`<br/>`make stop` | `./scripts/docker.sh stop`<br/>`make docker-stop` | `./scripts/deploy.sh down`<br/>`make down` |
|
||
| **Restart** | `./scripts/serve.sh --restart [flags]` | `./scripts/docker.sh restart` | ? |
|
||
|
||
**Nginx routing**:
|
||
- `/api/langgraph/*` ? Gateway embedded runtime (8001), rewritten to `/api/*`
|
||
- `/api/*` (other) ? Gateway API (8001)
|
||
- `/` (non-API) ? Frontend (3000)
|
||
|
||
### Running Backend Services Separately
|
||
|
||
From the **backend** directory:
|
||
|
||
```bash
|
||
# Gateway API
|
||
make gateway
|
||
```
|
||
|
||
Direct access (without nginx):
|
||
- Gateway: `http://localhost:8001`
|
||
|
||
### Browser CORS
|
||
|
||
The Gateway always permits local Vite origins on ports 5173, 5174, 3000, and
|
||
8080 for both `localhost` and `127.0.0.1` with credentials. A deployed
|
||
frontend on another origin must be listed exactly in `GATEWAY_CORS_ORIGINS`
|
||
(for example `http://47.88.25.99:7010`); explicit origins are merged with the
|
||
local development origins. Regression: `tests/test_taskcop_cors.py`.
|
||
|
||
### Frontend Configuration
|
||
|
||
The frontend uses environment variables to connect to backend services:
|
||
- `NEXT_PUBLIC_LANGGRAPH_BASE_URL` - Defaults to `/api/langgraph` (through nginx)
|
||
- `NEXT_PUBLIC_BACKEND_BASE_URL` - Defaults to empty string (through nginx)
|
||
|
||
When using `make dev` from root, the frontend automatically connects through nginx.
|
||
|
||
## Key Features
|
||
|
||
### Administrator-managed custom themes
|
||
|
||
`system_settings.json` → `appearance` now stores `custom_themes` alongside
|
||
the built-in default theme. A row is `CustomThemePreset` (`id`, `name`,
|
||
six-digit `primary`, optional primary-button `text_color`, `mode`, `overrides`, `enabled`,
|
||
`sort_order`). `overrides` is intentionally a short interaction-only list:
|
||
`button_hover`, `navigation_gradient_start`, `navigation_gradient_end`,
|
||
`selected_background`, `border`, `muted`, `muted_text`, `destructive`,
|
||
`brand_title`, and `primary_foreground`. The navigation gradient endpoints
|
||
default from `primary` and are only stored when an administrator overrides them.
|
||
All values must be six-digit hex strings; the frontend derives every remaining
|
||
semantic CSS/TDesign token so admins do not maintain a full raw token sheet.
|
||
Ids must be `custom-theme-*`, names and ids are unique, and the router limits
|
||
the list to 12. `GET/PUT /api/system-settings/appearance` is the single
|
||
read/write surface: reads are available to all users, writes require admin.
|
||
The router re-orders rows deterministically and only permits an enabled custom
|
||
theme as `default_theme`; removing or disabling the default safely falls back
|
||
to `light`. Regression coverage: `tests/test_appearance_theme.py`.
|
||
|
||
Compact top-bar shortcuts live in the same `appearance` setting. Their route
|
||
parameters support `login`, `theme`, and keyless `path`. A `path` parameter has
|
||
no query-key input; the operator selects its value source from the right-hand
|
||
dropdown and the frontend appends the resolved, URL-encoded value to the
|
||
external URL path (or Hash route) in configured order. Password stays an
|
||
ordinary `login` query parameter with the key `password`, so an external login
|
||
link can become `#/login/<user>?password=…`. Regression coverage:
|
||
`tests/test_appearance_theme.py`.
|
||
|
||
### File Upload
|
||
|
||
Multi-file upload with automatic document conversion:
|
||
- Endpoint: `POST /api/threads/{thread_id}/uploads`
|
||
- Supports: PDF, PPT, Excel, Word documents (converted via `markitdown`)
|
||
- Rejects directory inputs before copying so uploads stay all-or-nothing
|
||
- Reuses one conversion worker per request when called from an active event loop
|
||
- Files stored in thread-isolated directories
|
||
- Agent receives uploaded file list via `UploadsMiddleware`
|
||
|
||
See [docs/FILE_UPLOAD.md](docs/FILE_UPLOAD.md) for details.
|
||
|
||
### Plan Mode
|
||
|
||
TodoList middleware for complex multi-step tasks:
|
||
- Controlled via runtime config: `config.configurable.is_plan_mode = True`
|
||
- Provides `write_todos` tool for task tracking
|
||
- One task in_progress at a time, real-time updates
|
||
|
||
See [docs/plan_mode_usage.md](docs/plan_mode_usage.md) for details.
|
||
|
||
### Context Summarization
|
||
|
||
Automatic conversation summarization when approaching token limits:
|
||
- **On/off switch is per-request, not the config file.** The chat input box has a
|
||
「上下文压缩」 toggle (below 参考文献, **default off**) → frontend sends
|
||
`context.summarization_enabled` → whitelisted in `_CONTEXT_CONFIGURABLE_KEYS`
|
||
(`app/gateway/services.py`) → `_build_middlewares` reads it from the run config
|
||
and calls `_create_summarization_middleware(force_enabled=summarization_enabled)`.
|
||
`force_enabled=True` mounts the middleware even when `config.summarization.enabled`
|
||
is false; when the flag is absent (channels / scheduled runs) it falls back to
|
||
the config flag. `config.summarization.enabled` now defaults to **false**.
|
||
- `config.yaml` `summarization` still supplies the **tuning**: trigger types
|
||
(tokens, messages, or fraction of max input) and the keep policy.
|
||
- A compressed request must retain the latest non-empty `HumanMessage`, even
|
||
when it has already been folded into the hidden summary. This keeps strict
|
||
OpenAI-compatible gateways from rejecting a tool-continuation call with
|
||
`No user query found in message`.
|
||
- Keeps recent messages while summarizing older ones.
|
||
|
||
Tests: `tests/test_summarization_toggle.py` (gating + whitelist),
|
||
`tests/test_summarization_middleware.py` (cache behavior).
|
||
See [docs/summarization.md](docs/summarization.md) for details.
|
||
|
||
### Vision Support / Image Recognition (??)
|
||
|
||
Two paths, selected by `config.yaml` ? `vision`:
|
||
|
||
**Dedicated vision model (delegated recognition)** ? set `vision.model_name` to a model from `models[]`. The `view_image` tool sends the image to that model and returns its **text description** as the tool result, so a text-only main chat model can still "see" images. `view_image` accepts an optional `question` argument for targeted queries; the default prompt is `vision.prompt` (overridable), and `vision.max_tokens` caps the recognition output. Resolution/gating lives in `deerflow/models/vision.py` (`resolve_vision_model_name`, `vision_enabled`, `is_delegated_vision`, `recognize_image`).
|
||
|
||
**Native vision (fallback)** ? when `vision.model_name` is unset, behavior falls back to the main chat model, which only works if it has `supports_vision: true`:
|
||
- `ViewImageMiddleware` injects base64 image data into the conversation before the LLM call
|
||
- `view_image` stores the image in `state.viewed_images` for that middleware
|
||
|
||
Both the `view_image` tool and `ViewImageMiddleware` are registered whenever `vision_enabled(model_name)` is true (dedicated model configured **or** main model supports vision). Config schema: `deerflow/config/vision_config.py` (`VisionConfig`); tests: `tests/test_vision.py`.
|
||
|
||
### Upload symlink-safety (cross-platform)
|
||
|
||
`deerflow/uploads/manager.py:open_upload_file_no_symlink` writes uploads without following a pre-existing destination symlink. `O_NOFOLLOW` is applied where available (POSIX) but is **optional** ? Windows lacks `os.O_NOFOLLOW`/`O_NONBLOCK`, so requiring it previously made every Windows upload fail with `UnsafeUploadPathError`. Pre-existing symlinks are still rejected via `lstat`, and the opened fd is re-checked with `fstat`. `O_BINARY` is added on Windows. Regression: `tests/test_uploads_windows_safe.py`.
|
||
|
||
## Code Style
|
||
|
||
- Uses `ruff` for linting and formatting
|
||
- Line length: 240 characters
|
||
- Python 3.12+ with type hints
|
||
- Double quotes, space indentation
|
||
|
||
## Documentation
|
||
|
||
### WeKnora iframe behind `/deerflow`
|
||
|
||
The LLMWiki WeKnora detail iframe is publicly served through the Gateway at
|
||
`/deerflow`, independently of the frontend deployment prefix (for example
|
||
`/magentweb`). Register `llmwiki.proxy_router` with the `/deerflow` prefix
|
||
before its unprefixed fallback router, otherwise the unprefixed catch-all
|
||
captures `/deerflow/*`. HTML root-relative assets, API strings, and Vite's
|
||
`__vite__mapDeps` table are rewritten for `/deerflow`; generic JavaScript
|
||
regular-expression literals must not be rewritten. Auth and CSRF bypass only
|
||
applies when the short-lived signed WeKnora embed cookie is valid. Keep static
|
||
path recognition in `app/gateway/weknora_embed.py` synchronized with the
|
||
proxy's static fallback paths in `app/gateway/routers/llmwiki.py`.
|
||
|
||
### Position-roundtable role directory
|
||
|
||
`position_roles` is the global shared directory for position-roundtable work
|
||
views. It is deliberately separate from organisation/RBAC `positions`.
|
||
`app.gateway.deps.langgraph_runtime()` seeds the four stable built-in ids from
|
||
`app/gateway/position_role_defaults.py` via `ensure_seed()` (missing rows only;
|
||
never overwrite operator edits). The API is `GET/PUT /api/position-roles`;
|
||
writes are admin-only. Exactly one enabled `intent` role is required and owns
|
||
the initial intent workspace plus the summary/action-plan closing agents. Do
|
||
not delete a built-in id because historical chain nodes persist those ids;
|
||
custom role ids may be added and removed.
|
||
|
||
### Position-roundtable task-scoped session history
|
||
|
||
`position_roundtable_sessions.external_task_id` is optional. When populated from the frontend route `taskId`, new rows set `task_scoped = true`: a task can have multiple sessions and all users opening that task see the same sessions, nodes and persisted conversation snapshots. When it is NULL/empty, keep the original `user_id`-scoped personal workspace behavior. `20260808_01` adds `task_scoped` with default `false` (so all historical records remain personal) plus the task/update-time index; it must not backfill or make `external_task_id` non-null.
|
||
|
||
### Roundtable artifact submission status
|
||
|
||
`roundtable_artifact_submissions` keeps task-level submission state for a
|
||
generated deliverable and is deliberately separate from task details and any
|
||
external task-system delivery. The API is
|
||
`GET /api/roundtable/tasks/{taskId}/artifacts?submitted=true` and
|
||
`PUT /api/roundtable/tasks/{taskId}/artifacts/submission`. All callers use the
|
||
normal Gateway Bearer authentication; records are shared by task while
|
||
`created_by`/`submitted_by`/`cancelled_by` are audit-only. The server derives a
|
||
SHA-256 idempotency key from `taskId + sourceKey + path + revision`: submitting
|
||
the same revision is an upsert, while cancellation only flips `submitted` and
|
||
sets `cancelled_at` (it never deletes history). Store only the sandbox file
|
||
reference (`threadId`, path, source/node/session fields, name/MIME and revision)
|
||
and timestamps; do not persist Markdown content or call a formal delivery
|
||
interface. Previews continue to read the DeerFlow artifact by `threadId + path`.
|
||
|
||
See `docs/` directory for detailed documentation:
|
||
- [../README.md](../README.md) - Offline Docker wheelhouse preparation and startup workflow. The pinned versions in `../offline-runtime-requirements.txt` and `../start_backend.sh` must stay synchronized. Startup checks `croniter`, `hindsight-client`, the aiomysql compatibility bundle, and Redis, installing missing or incompatible packages from wheelhouse only. The bundle installs `aiomysql==0.3.2`, `PyMySQL==1.2.0`, and `SQLAlchemy==2.0.51` together; the startup probe also verifies that SQLAlchemy's adapted `ping(reconnect=False)` signature is present, preventing the old-adapter missing-`reconnect` failure. The download target defaults to CPython 3.12 on Linux x86_64 and verifies dependency closure using wheelhouse only.
|
||
- [CONFIGURATION.md](docs/CONFIGURATION.md) - Configuration options
|
||
- [ARCHITECTURE.md](docs/ARCHITECTURE.md) - Architecture details
|
||
- [API.md](docs/API.md) - API reference
|
||
- [SETUP.md](docs/SETUP.md) - Setup guide
|
||
- [FILE_UPLOAD.md](docs/FILE_UPLOAD.md) - File upload feature
|
||
- [PATH_EXAMPLES.md](docs/PATH_EXAMPLES.md) - Path types and usage
|
||
- [summarization.md](docs/summarization.md) - Context summarization
|
||
- [plan_mode_usage.md](docs/plan_mode_usage.md) - Plan mode with TodoList
|
||
- [BACKEND_RUN_STOP_ZH.md](docs/BACKEND_RUN_STOP_ZH.md) - Windows-first backend start/stop and troubleshooting guide
|
||
|
||
### Deep Research retrieval
|
||
|
||
`deerflow.agents.deep_research.adapters.material_provider.DeerFlowMaterialProvider`
|
||
must remain on the Q&A retrieval path. Enabled `deep-search`, `web-research`,
|
||
and knowledge-base skills run through the existing `SubagentExecutor` skill
|
||
runtime only when their required Q&A tool is registered; then the provider may
|
||
use configured `web_search`. A disabled placeholder endpoint is not a valid
|
||
search channel: use the maintained keyless DuckDuckGo provider and a relevant
|
||
real-web fallback as read-only sources, keep source provenance, and never
|
||
enable or call a placeholder just to make the Deep Research UI look configured. Regression coverage is in
|
||
`tests/test_deep_research_material_provider.py`.
|
||
When a queued Deep Research job is cancelled, the Gateway must also clear the
|
||
owning session's `active_job_id` and mark it `cancelled`; otherwise the stale
|
||
session incorrectly consumes that user's concurrency limit.
|
||
|
||
Report prose is the one Deep Research LLM path that must use
|
||
`DeerFlowCompletionBackend.stream_complete()` rather than `complete()`: it
|
||
drives `model.astream()`, sends native `report_delta` frames through
|
||
`ReportChunkEmitter`, and persists bounded `report_chunk` checkpoints. The
|
||
app-layer `DeepResearchLiveHub` fans out those live frames inside one Gateway
|
||
worker only; the database event log still owns `?after=<seq>` reconnect replay
|
||
and multi-worker recovery. Inject the hub callback into `PersistingEventSink`
|
||
and never import `app.*` from `deerflow.*`. Keep structured plan/review calls
|
||
on `complete()`. Before the search fan-out, basic mode emits one
|
||
`queries_planned` event and detailed/multi-agent modes emit one per subtopic;
|
||
they are durable user-visible search terms, not chain-of-thought. On completion, retain
|
||
the terminal `lastJobId` in `usage_snapshot` so a refreshed conversation page
|
||
can replay its trace from the event log. Detailed mode uses stable targets (`introduction`,
|
||
`section:N`, `conclusion`) because section writers are concurrent;
|
||
multi-agent revisions must emit `report_reset` before their next report
|
||
version. Do not persist raw model reasoning or one persisted event per provider
|
||
token. Models that expose thinking inside ordinary `<think>...</think>` text
|
||
must be split statefully before a report delta, checkpoint, or Markdown
|
||
projection is created; an optional live-only `report_thinking` frame can drive
|
||
the message-list thought trace, while the sandbox receives answer Markdown
|
||
only. The post-report digest (`summarize_report`) must likewise forward
|
||
provider reasoning through live-only `summary_thinking` frames so the
|
||
chat timeline can render thought while waiting for the first `summary_delta`.
|
||
The Vite page starts Word downloads through
|
||
`frontend-web/src/core/scheduled-tasks/markdown-docx.ts`, but the Gateway owns
|
||
generation through `POST /api/writing/export/docx`. Keep the standard-library
|
||
OOXML builder and its embedded-font validation; do not add a Python document-
|
||
conversion dependency merely to export text.
|
||
All non-streaming Deep Research control-plane calls (`complete()` for query /
|
||
section / plan / review JSON) run through `asyncio.wait_for` using the frozen
|
||
`DeepResearchConfig.control_plane_timeout_seconds` (45 seconds by default,
|
||
clamped to 120). Do not apply that total-duration cap to `stream_complete()`:
|
||
reports may legitimately be long but must continue emitting visible deltas. A
|
||
basic query-planning timeout/failure emits `query_planning_fallback` and searches
|
||
the original topic; detailed and multi-agent plan creation similarly fall back
|
||
to their original topic/section, rather than leaving a durable job stuck in
|
||
`planning`. Tests: `tests/test_deep_research_llm_adapter.py` and
|
||
`tests/test_deep_research_runners.py`.
|
||
The per-user Deep Research concurrency guard counts only session statuses
|
||
`running` and `awaiting_input`. A `draft` has no durable job and must never
|
||
consume a slot; otherwise saved drafts make `POST /sessions/{id}/jobs` return
|
||
429 permanently. Regression: `test_saved_draft_sessions_do_not_consume_research_concurrency`.
|
||
|
||
`POST /api/deep-research/sessions/{id}/jobs` is retry-safe under MySQL
|
||
read/write splitting. Deep Research session/job/event/source/message writes
|
||
must serialise their changed row while the write transaction is still open
|
||
(`flush` for inserts, select-after-update **before** commit), never through a
|
||
post-commit `session.refresh()` or a fresh pooled read. Replica lag otherwise
|
||
turns a committed job into a misleading 500, can make dispatcher `claim_job()`
|
||
lose its snapshot, or can fail report/follow-up persistence after the model has
|
||
already completed. Always nudge the dispatcher for both a newly created job and
|
||
an idempotently reused queued job; otherwise a retry after such an interrupted
|
||
response returns the old job but leaves it waiting for the periodic scan. The
|
||
SSE replay generator must catch temporary event/job-store reads, emit a
|
||
recoverable warning, and keep the connection open so the next poll can repair
|
||
the cursor instead of surfacing an ASGI `ExceptionGroup`. Regressions:
|
||
`test_deep_research_writes_never_depend_on_post_commit_refresh`,
|
||
`test_job_create_never_reads_back_after_commit`,
|
||
`test_retry_of_existing_queued_job_nudges_dispatcher`, and
|
||
`test_stream_stays_open_when_durable_replay_has_a_transient_failure`.
|
||
|
||
At ordinary report completion, persist and confirm the session's
|
||
`report_markdown` projection **before** marking the Job `completed`. Retry a
|
||
short transient projection failure; if it cannot be confirmed, let the job
|
||
become failed rather than reporting a report that vanishes on refresh.
|
||
`test_report_is_persisted_before_job_is_marked_completed` and
|
||
`test_report_projection_failure_never_marks_job_completed` cover the ordering.
|
||
|
||
The chat-collection frontend must keep report-writing progress in the existing
|
||
workspace `MessageList` / `MessageGroup` tool-step renderer: inject only
|
||
virtual `deep_research_progress` tool calls for the durable phases, keep the
|
||
completed phase visible after the job ends, and append the digest as a normal
|
||
assistant message. Do not recreate a second deep-research-only conversation
|
||
timeline. On `summarizing`, insert that same normal assistant message with its
|
||
streaming marker and update it for every SSE `summary_delta`; when the report
|
||
completes, replace its content in place with the final completion/digest
|
||
message rather than waiting to append a summary. Completed chat-collection
|
||
sessions also inject a normal `present_files` card into the transcript, named
|
||
from the report Markdown H1 rather than its internal `report.md` path.
|
||
Replaying a completed history session must not auto-open its
|
||
sandbox; opening that card selects its virtual write-file report instead, and
|
||
its file-card download action delegates to the Deep Research Markdown/Word
|
||
export dialog. The default in-job concurrency and per-user active-job cap are
|
||
both three. Before the user confirms the report configuration, derive its
|
||
displayed material count from unique collector tool-result rows because the
|
||
durable Deep Research source store is populated when the writing job harvests
|
||
that thread. The Deep Research composer model picker must use the installed
|
||
TDesign `Select` backed by `/api/models`; its explicit choice is sent as
|
||
`model_name` on collector turns and applied to `fast_model`, `smart_model`, and
|
||
`strategic_model` when creating a session or starting a report job. Disable it
|
||
while collection or report writing is active so one durable run cannot switch
|
||
models mid-stream.
|
||
|
||
Report structures (`/api/report-structures`) may store an optional
|
||
`retrieval_directions` list (检索方向) and an optional `retrieval_skills`
|
||
list (检索来源 / 指定技能). Empty / unset keeps today's collector
|
||
query-planning, default tools, and runner sub-query generation. A non-empty
|
||
directions list is frozen into `DeepResearchConfig.retrieval_directions`: the
|
||
chat collector's first user message gets a `<retrieval-directions>` block that
|
||
asks it to **split concrete retrievable terms along those facets** (never search
|
||
the direction labels themselves), and legacy `basic`/`detailed`/`multi_agent`/
|
||
`deep` runners pass the same constraint into query-planning instead of using
|
||
`课题 + 方向` as the search string. A non-empty skills list is frozen
|
||
into `DeepResearchConfig.retrieval_skills`: the chat collector's first user
|
||
message gets a `<retrieval-skills>` block, the collector run context receives
|
||
`extra_skills` so `skill_view`/`skill_list` can see those names, and legacy
|
||
material collection runs those skills instead of the default Q&A research-skill
|
||
trio. Helpers live in `deerflow.agents.deep_research.retrieval`.
|
||
|
||
Completed-report Q&A has two endpoints: retain `POST /sessions/{id}/chat` as the
|
||
legacy JSON response for compatibility, and use `POST /sessions/{id}/chat/stream`
|
||
for report-grounded Q&A SSE. The workbench must send a strong scoped edit
|
||
instruction (for example “修改第一段” / “重新生成一下第一段” / “重写第二节”) to
|
||
the separate `POST /sessions/{id}/rewrite/stream` report-section writing pipeline.
|
||
That endpoint requires a deterministic current-report range (otherwise 422), so
|
||
its output is always a replacement candidate rather than a Q&A answer. The
|
||
candidate is streamed in the ordinary message row
|
||
while the sandbox preserves the current report. When generation ends, inject a
|
||
confirmation card. Only the owner’s explicit
|
||
`POST /sessions/{id}/messages/{messageId}/rewrite-proposal` action `apply` may
|
||
persist the assembled `report_markdown`; `dismiss` preserves the original. The
|
||
proposal is bound to a hash of the report it was generated from, so stale
|
||
candidates fail with 409 rather than overwrite a later edit. Ambiguous or
|
||
ordinary evidence questions must remain non-mutating, and cancellation or a
|
||
stream error must never save a partial report. The rewrite prompt receives only
|
||
the selected range plus the selected evidence set, never a new web search.
|
||
|
||
### DeerFlow-local LLMWiki Wiki index
|
||
|
||
`llmwiki.local_wiki_index` is an opt-in local-vector subsystem. WeKnora is the
|
||
Wiki generation/sync source and the Wiki-only fallback while an entire mapping
|
||
has not completed vectorization. The readiness gate requires `state=idle`, a
|
||
successful full scan, matching embedding fingerprint, no failed pages,
|
||
`ready_page_count == local_page_count`, and vectors for a non-empty library.
|
||
Ready mappings read DeerFlow's `llmwiki_wiki_pages` /
|
||
`llmwiki_wiki_vectors` / `llmwiki_wiki_sync_states` tables; incomplete mappings
|
||
call only WeKnora `/knowledgebase/{id}/wiki/pages?query=...` plus Wiki-page
|
||
detail. Never call `/knowledge-search` from internal search, RAG, or the
|
||
`llmwiki_search` tool, and never return raw chunks as the fallback. Core code lives under
|
||
`deerflow.integrations.weknora.local_index` and
|
||
`deerflow.persistence.llmwiki_index`; preserve the harness → app import
|
||
firewall. Gateway composition injects the store/search/sync services through
|
||
`local_index.runtime`. `enabled` turns on local vector search and `auto_sync`
|
||
defaults to the lease-protected background scheduler. The scheduler performs
|
||
an immediate incremental scan at startup, then rescans every configured
|
||
interval so new WeKnora libraries and asynchronously generated Wiki pages are
|
||
picked up without a manual vectorization click. Unchanged content must reuse
|
||
existing vectors.
|
||
|
||
Vectors are normalized little-endian float32 BLOBs. A page mirror and all of
|
||
its vectors are replaced in one transaction only after strict batched
|
||
embedding succeeds. Never delete the old generation before a remote embedding
|
||
call and never fall back to WeKnora raw chunks while the local feature flag is
|
||
enabled. NumPy matrices are per-mapping LRU snapshots invalidated by
|
||
`index_revision`; large decode/matrix work runs through `asyncio.to_thread`.
|
||
`embedding.dimensions` is a local response-size invariant and fingerprint
|
||
input, not an outbound API field: fixed-dimension BGE/OpenAI-compatible
|
||
endpoints can reject a `dimensions` request parameter with HTTP 400. When the
|
||
value is omitted, the embedding client probes and freezes the first response's
|
||
dimension before sync lease/rebuild decisions.
|
||
Embedding errors must retain the upstream status and bounded response body in
|
||
logs, without exposing API keys. SQL vector writes are split into statements
|
||
of at most `database_batch_size` rows (default 50) inside the same atomic page
|
||
replacement transaction.
|
||
During a fingerprint rebuild the mapping uses the Wiki-page API until every
|
||
page is successfully rebuilt. `WikiRetrievalService` partitions mixed requests
|
||
per mapping, so ready libraries stay local while incomplete libraries use that
|
||
Wiki-only fallback.
|
||
|
||
The frontend must keep Wiki citations visibly distinct from raw-document
|
||
citations while hiding provider-brand wording from user-facing labels, status
|
||
text, and errors. Any click on a Wiki source card opens the shared right-side
|
||
source drawer, which renders the complete Markdown page. When the local index
|
||
is disabled, the drawer reads the remote WeKnora Wiki-page endpoint directly
|
||
and must not surface the local-index-disabled response. A Wiki target must not
|
||
also load or render the legacy original-document preview, even when its citation
|
||
still carries a document id. The drawer supports single-page Markdown and Word
|
||
downloads, and its **Open in Wiki** link carries both the DeerFlow knowledge-base
|
||
mapping id and the exact `wikiSlug`. Wiki search-result cards use this same
|
||
drawer instead of navigating immediately. Exact-page navigation uses the native
|
||
`knowledge/{mapping_id}/wiki/view?wikiSlug=...` reader, not the WeKnora detail
|
||
iframe. Keep legacy `knowledge/{mapping_id}?tab=wiki&wikiSlug=...` links fast by
|
||
redirecting them in `LlmWikiKnowledgePage` before its runtime-loading gate.
|
||
Drawer-to-reader navigation stays in the current SPA and passes the fetched
|
||
Wiki page in router state so the article can render before the background query
|
||
refresh completes. The native reader must retain a searchable/collapsible left
|
||
Wiki directory, reconstruct hierarchy from `parent_slug`, `wiki_path`,
|
||
`category_path`, and slug fallback, expand/highlight the active page, and keep
|
||
directory navigation on the same lightweight route rather than mounting the
|
||
WeKnora iframe.
|
||
|
||
Generated HashRouter URLs must retain `window.location.pathname` (for example
|
||
the `/web/` reverse-proxy prefix); never emit root-absolute `/#/...` citation,
|
||
reader, or public-embed links. Protected Wiki image rendering must use real HTTP
|
||
URLs, never `URL.createObjectURL`. Public readers use the published file proxy;
|
||
authenticated readers obtain a short-lived signed capability from
|
||
`/files/access-url` and load `/api/public/llmwiki/files/{token}`. The signed
|
||
payload is restricted to `kind=wiki_file`, mapping id, remote id, validated
|
||
protected path, and expiry, and responses provide an inline UTF-8 filename for
|
||
new-tab viewing and browser Save As. The proxied WeKnora detail iframe must
|
||
likewise leave its scoped HTTP file-proxy URL on every protected image and
|
||
restore it if upstream rendering replaces `src`; its click preview reads that
|
||
same DOM value, so a fetched-and-revoked object URL is not acceptable.
|
||
|
||
Third-party pages embed published Wiki content through the chrome-free route
|
||
`/#/embed/knowledge/{mapping_id}` (optional exact-page `wikiSlug`). It defaults
|
||
to a recognized Wiki index page (title `索引`/`index`, page type `index`, or an
|
||
`index`/`home`/`readme` slug) and promotes that page to the first public-directory
|
||
entry. When upstream pages exist without an explicit index shell, the frontend
|
||
injects a virtual `__index__` page and builds its catalogue from the page list;
|
||
an all-empty dynamic index response must also fall back to that page list. It retains
|
||
directory navigation and supports the shared public-embed `theme` / `bg` /
|
||
`text` appearance inputs. WeKnora dynamically renders index page-type counts,
|
||
nested category headings, page summaries, and `data-slug` links even though the
|
||
stored index Markdown only contains its introduction; the native public reader
|
||
reads the same ordered groups through the published-only
|
||
`/api/public/llmwiki/knowledge-bases/{id}/wiki/index` proxy and renders them with
|
||
`buildWikiIndexCatalogMarkdown`. The Gateway paginates every WeKnora index type
|
||
(`summary`, `entity`, `concept`, `synthesis`, `comparison`) to completion while
|
||
preserving upstream group/item order. Index Markdown
|
||
may also contain WeKnora shell-only placeholder links;
|
||
`resolveWikiIndexPageLinks` matches their visible labels to
|
||
the loaded page title, slug, or alias and rewrites them to exact embedded
|
||
`wikiSlug` routes. Preserve WeKnora's `wiki-content-link` interaction: primary
|
||
brand color, medium weight, dashed underline becoming solid on hover, focus
|
||
outline, and in-reader SPA navigation rather than a page reload. `GET
|
||
/api/public/llmwiki/knowledge-bases` is the anonymous discovery endpoint. It
|
||
validates published ordinary mappings against WeKnora, drops stale mappings,
|
||
and returns safe comprehensive metadata from the WeKnora list/detail/page APIs:
|
||
document/chunk/processing/share counts, non-secret configuration and timestamps,
|
||
Wiki page type/status counts and explicit/generated index entry, plus aggregate
|
||
local-vector state/counts/timestamps. Never expose remote WeKnora ids,
|
||
tenant/creator identity, storage credentials, API credentials, raw vectors, or
|
||
per-page indexing errors. All anonymous reader calls use that
|
||
public namespace and the `/knowledge-bases/{id}` Wiki-page children. Those
|
||
endpoints must resolve with a synthetic anonymous user and `is_admin=False`,
|
||
explicitly require `publication_status=published`, and hard-deny the system
|
||
conversation-deposit mapping. Never point this page at the authenticated
|
||
`/api/llmwiki` reads or weaken the published-only gate merely because the
|
||
mapping id is hard to guess.
|
||
|
||
Owners manually force a full sync/vectorization through `POST
|
||
/api/llmwiki/knowledge-bases/{id}/vectorize`; status is available through `GET
|
||
/api/llmwiki/knowledge-bases/{id}/wiki-index-status`. That response includes
|
||
aggregate `progress` and per-page `pages` details; current-scan progress is
|
||
derived from `active_sync_id` / `last_seen_sync_id`, while each page exposes its
|
||
index state, vector count, timestamps, and bounded error. The Wiki management
|
||
page must show the progress bar and a status badge beside every article. It
|
||
also exports all remote processed Wiki pages via `GET
|
||
/api/llmwiki/knowledge-bases/{id}/wiki/export.xlsx`. Directory export must keep
|
||
the full path, current directory, parent path, and dynamic per-level columns;
|
||
long Markdown is split into Excel-safe continuation columns instead of being
|
||
truncated. Workbook creation runs in `asyncio.to_thread`.
|
||
|
||
Internal search authorization always starts from DeerFlow mapping ids and is
|
||
rechecked by the Gateway. External endpoints use the isolated
|
||
`/api/external/llmwiki/wiki` namespace and `X-API-Key`; require published +
|
||
external-enabled + index-enabled on every request and response, and hard-deny
|
||
the `system` / `对话沉淀` mapping. Those API-key endpoints never return remote
|
||
WeKnora ids, provenance refs, credentials, or vector bytes.
|
||
|
||
`POST /api/knowledge/vector-search` is a deliberately separate anonymous API.
|
||
It has no token/API-key requirement or rate limit, applies route-scoped
|
||
`Access-Control-Allow-Origin: *`, searches every published local Wiki index
|
||
regardless of `external_search_enabled`, and still hard-denies the `system` /
|
||
`对话沉淀` mapping before and after search. Its contract is section-level (one
|
||
result per matched vector, so Wiki-page fields and URLs may repeat), default
|
||
`top_k=10`, with optional `include_page_content`. Unlike the protected legacy
|
||
external namespace, this endpoint intentionally returns DeerFlow/WeKnora ids
|
||
and raw source/chunk provenance, but never credentials or vector bytes. Wiki
|
||
links use `llmwiki.local_wiki_index.frontend_base_url` (development default
|
||
`http://127.0.0.1:5174`) and point to the native
|
||
`/#/page/workspace/knowledge/{mapping_id}/wiki/view?wikiSlug=...` route.
|
||
Regression tests:
|
||
`tests/test_llmwiki_local_index_repository.py`,
|
||
`tests/test_llmwiki_wiki_sync.py`, `tests/test_llmwiki_vector_search.py`,
|
||
`tests/test_external_llmwiki_search.py`,
|
||
`tests/test_public_knowledge_vector_search.py`, and
|
||
`tests/test_llmwiki_wiki_retrieval.py`.
|
||
|
||
The sandbox whole-report rewrite may carry an optional `reportOutline` Markdown
|
||
template from the shared report-structure library. Persist the original template
|
||
in the durable input snapshot and append it as an explicit chapter-structure
|
||
constraint before requirements planning and prose generation; a template must
|
||
never be used as permission to invent facts absent from the current report.
|
||
|
||
`DocumentRewriteVersionRepository.create` backs the reversible before-image for
|
||
both artifact and whole-report rewrites. It must flush and serialize the new row
|
||
on the writer before commit, and must never call `session.refresh()` after
|
||
commit: a read/write-splitting MySQL endpoint can route that refresh to a
|
||
lagging replica, falsely report a version-save failure, and trigger the caller's
|
||
intentional restore of the rewritten document.
|
||
|
||
Version snapshots are an undo enhancement, not a condition for preserving an
|
||
already committed rewrite. If snapshot creation fails after the artifact/report
|
||
write, log the exception, complete the rewrite with
|
||
`versionSnapshotFailed: true`, and do not emit a user-facing failure or restore
|
||
the original. The sandbox uses that signal only to enable its ordinary manual
|
||
save fallback.
|
||
|
||
### Offline local-sandbox Office runtime
|
||
|
||
The DeerFlow-only offline deployment uses `LocalSandboxProvider`, so document
|
||
dependencies belong in the backend image rather than a separate sandbox image.
|
||
`deploy/linux-docker-offline/Dockerfile.backend` installs python-docx,
|
||
LibreOffice Writer/Calc/Impress, Pandoc, Poppler, and CJK fonts in addition to
|
||
the MarkItDown Office extras. `file_conversion.py` normalizes legacy
|
||
DOC/XLS/PPT and ODT/ODS/ODP through a per-conversion LibreOffice user profile
|
||
before MarkItDown extraction, avoiding concurrent profile locks. The image
|
||
build and `OFFLINE_INSTALL.sh --office-check` both run
|
||
`scripts/office_runtime_selfcheck.py`; that smoke test creates and re-parses
|
||
minimal DOCX/XLSX/PPTX files and does not need the application database.
|
||
|
||
## Skill/assistant knowledge invariants
|
||
|
||
Skill distillation is implemented in `deerflow.skill_knowledge` and
|
||
`deerflow.persistence.skill_knowledge`; native assistant bases are in
|
||
`deerflow.assistant_knowledge` and
|
||
`deerflow.persistence.assistant_knowledge`. Gateway composition, remote adapters,
|
||
the lease dispatcher, and APIs remain in `app.gateway` so the harness-to-app
|
||
import firewall is preserved.
|
||
|
||
Never execute a skill or read/upload/summarize its code and script files.
|
||
Knowledge fingerprints exclude those files and Markdown code fences. All
|
||
resyncs must stage and validate the knowledge digest before activation; the old
|
||
snapshot remains authoritative on failure. A binding may have only one active
|
||
item (`active_dedupe_key` plus atomic lease claims). Skill distillation targets
|
||
only writable WeKnora Wiki knowledge bases in this product version; reject direct
|
||
assistant-base targets, non-Wiki/FAQ/conversation/readonly targets, and direct
|
||
Neo4j writes. The create/list/resync/retry/detach flow is available to ordinary
|
||
users for their own visible skills, writable knowledge bases, jobs and bindings;
|
||
administrator-only auto-sync/review APIs remain privileged. Detach/delete operates on source contributions, not canonical rows,
|
||
and shared entities/relations are retired only when their final contribution is
|
||
gone.
|
||
|
||
Assistant knowledge has one global full-library base (`知识梳理总库`) and any
|
||
number of admin-created custom bases. Ordinary users have read/search access and
|
||
administrators own all mutations. Assistant bases accept local documents independently
|
||
of WeKnora via `/api/assistant-knowledge/imports/files`, as well as WeKnora export
|
||
packages or imports triggered from a WeKnora mapping. Local documents use the
|
||
Gateway `assistant_file_ingest` pipeline: persisted original files, chunked LLM Wiki
|
||
extraction, cached body embeddings, and import-job status/retry. Do not add direct skill
|
||
imports. Importing into a
|
||
custom assistant base must also mirror the package into the global base for
|
||
cross-source entity alignment. Admins may rename any assistant base and delete
|
||
custom bases. Deletion removes the custom base's Wiki pages, vectors, entities,
|
||
relations, import audit rows, and package files; the global base is never
|
||
deletable because it owns full-library aggregation.
|
||
|
||
### Wiki package and local encoder invariants (2026-09)
|
||
|
||
Ordinary knowledge detail landing is globally configured by `ordinary_knowledge_landing` (`wiki` default / `weknora`), with admin-only PUT `/api/system-settings/ordinary-knowledge-landing`. Explicit document/graph/article links bypass this default. Wiki reading and the embedded WeKnora view provide reciprocal navigation; `?view=weknora` explicitly selects the iframe and must not loop back to the default Wiki route.
|
||
|
||
- `knowledge_transfer.sync_to_global` is called after successful explicit local
|
||
Wiki vectorization; it skips conversation/system mappings. Do not enable a new
|
||
watcher or automatic remote-write scheduler as a side effect.
|
||
- `wiki_packages` exposes file-backed assistant exports and restores into empty
|
||
ordinary Wiki libraries. Validate vectors before any remote mutation; recreate
|
||
folder IDs from category paths using provider APIs (a raw category_path is
|
||
discarded by WeKnora without a matching folder_id). Preserve page metadata and
|
||
rebuild backlinks after restoring all pages.
|
||
- `PackageArchive` is a replayable, disk-backed JSONL staging index. Always close
|
||
it in finally. Runtime imports must not use the legacy in-memory read_package.
|
||
- Preserve normalized float32 base64 bytes exactly. Same dimension alone is not
|
||
model compatibility. Assistant vector retrieval is fingerprint-gated cosine
|
||
search, and any keyword fallback must be labeled as such.
|
||
Preserve a valid zero section_index. Export/search vectors belonging to each
|
||
current source snapshot (`source.last_import_job_id`), not merely the page's
|
||
current revision: one aligned page may intentionally contain current vectors
|
||
from several sources.
|
||
- The `local` embedding provider uses FastEmbed with official BAAI/bge-m3 ONNX,
|
||
1024 dimensions, bounded CPU batches and a per-client async lock. Run inference
|
||
off the event loop. Existing page vectors are reusable; zero-vector nonempty
|
||
pages need backfill, not a forced rebuild of every page.
|
||
FastEmbed is a standard backend dependency. Encoder configuration, directory
|
||
upload and warmup live only in the administrator global settings dialog.
|
||
A runtime `system_settings.json` path override wins over config.yaml; uploaded
|
||
files are streamed into `.deer-flow/models/bge-m3`, never the source tree.
|
||
Encoder warmup is explicitly started from the settings UI/API and must never block application startup
|
||
or turn a model-load failure into a DeerFlow service failure. Local provider
|
||
is offline-only: require local_model_path or DEERFLOW_BGE_M3_MODEL_PATH and
|
||
never let the application download model weights. Keep
|
||
`local_process_isolation=true` by default: inference runs in the bundled JSONL
|
||
worker subprocess, and shutdown must terminate that worker without waiting on
|
||
model teardown.
|
||
- Every successful mapping snapshot is mirrored to the global assistant base,
|
||
including empty snapshots so deleted source pages are retired. Re-import is
|
||
source-replacement semantics: absent Wiki/chunk/entity/relation pages must not
|
||
remain visible merely because they existed in an older source snapshot.
|
||
- Provider graph nodes bind to their original wiki_slug. They must not create
|
||
duplicate synthetic entity/relation articles. Store original Wiki metadata
|
||
separately from assistant workflow status. Count Wiki pages separately from
|
||
raw document chunks.
|
||
- The global base aligns entities by normalized display name and aliases across
|
||
source records, and aligns content Wiki pages by normalized title. Do not use
|
||
source slug as the global canonical identity. Merge unique knowledge blocks
|
||
while retaining source contributions/origins. Uploaded packages must rebind
|
||
the temporary upload source to the stable WeKnora knowledge-base id in the
|
||
package, so re-upload is source-replacement rather than append. Retrieval
|
||
candidates exclude summary/index/directory pages; keyword matching, vector
|
||
snippets, citations and prompts must use knowledge body/sections, never page
|
||
summary as a substitute.
|
||
- Migration `20260903_01` adds nullable page metadata and merges the pre-existing
|
||
knowledge/workflow Alembic heads. Regression coverage is in
|
||
`tests/test_wiki_package_roundtrip.py`, `test_llmwiki_wiki_sync.py` and
|
||
`test_assistant_knowledge_repository.py`.
|
||
|
||
## openGauss / GaussDB persistence invariant
|
||
|
||
`database.backend=gauss` is an application-ORM backend, parallel to `mysql`.
|
||
It normalizes `database.gauss_url` to `opengauss+asyncpg://` and requires the
|
||
`gauss` optional dependency (`asyncpg` plus `opengauss-sqlalchemy`). Never route
|
||
a Gauss URL through SQLAlchemy's generic PostgreSQL dialect: the wire protocol
|
||
may connect, but dialect initialization and schema DDL can fail against the
|
||
Gauss server/version behavior. Keep the LangGraph checkpointer on the separately
|
||
configured SQLite backend; the PostgreSQL checkpointer schema is not part of the
|
||
Gauss compatibility contract. Treat both `postgresql` and `opengauss` dialect
|
||
names as timezone-aware PostgreSQL-family engines in portable types and
|
||
dialect-specific session settings.
|