269 KiB
CLAUDE.md
?? This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
DeerFlow is a LangGraph-based AI super agent system with a full-stack architecture. The backend provides a "super agent" with sandbox execution, persistent memory, subagent delegation, and extensible tool integration - all operating in per-thread isolated environments.
Architecture:
- Gateway API (port 8001): REST API plus embedded LangGraph-compatible agent runtime
- Frontend (port 3000): Next.js web interface
- Nginx (port 2026): Unified reverse proxy entry point
- Provisioner (port 8002, optional in Docker dev): Started only when sandbox is configured for provisioner/Kubernetes mode
Runtime:
make dev, Docker dev, and production all run the agent runtime in Gateway viaRunManager+run_agent()+StreamBridge(packages/harness/deerflow/runtime/). Nginx exposes that runtime at/api/langgraph/*and rewrites it to Gateway's native/api/*routers.
Project Structure:
deer-flow/
??? Makefile # Root commands (check, install, dev, stop)
??? config.yaml # Main application configuration
??? extensions_config.json # MCP servers and skills configuration
??? backend/ # Backend application (this directory)
? ??? Makefile # Backend-only commands (dev, gateway, lint)
? ??? langgraph.json # LangGraph Studio graph configuration
? ??? packages/
? ? ??? harness/ # deerflow-harness package (import: deerflow.*)
? ? ??? pyproject.toml
? ? ??? deerflow/
? ? ??? agents/ # LangGraph agent system
? ? ? ??? lead_agent/ # Main agent (factory + system prompt)
? ? ? ??? middlewares/ # 10 middleware components
? ? ? ??? memory/ # Memory V2: provider/manager/builtin/hindsight
? ? ? ??? thread_state.py # ThreadState schema
? ? ??? sandbox/ # Sandbox execution system
? ? ? ??? local/ # Local filesystem provider
? ? ? ??? sandbox.py # Abstract Sandbox interface
? ? ? ??? tools.py # bash, ls, read/write/str_replace
? ? ? ??? middleware.py # Sandbox lifecycle management
? ? ??? subagents/ # Subagent delegation system
? ? ? ??? builtins/ # general-purpose, bash agents
? ? ? ??? executor.py # Background execution engine
? ? ? ??? registry.py # Agent registry
? ? ??? tools/builtins/ # Built-in tools (present_files, ask_clarification, view_image)
? ? ??? mcp/ # MCP integration (tools, cache, client)
? ? ??? models/ # Model factory with thinking/vision support
? ? ??? skills/ # Skills discovery, loading, parsing
? ? ??? config/ # Configuration system (app, model, sandbox, tool, etc.)
? ? ??? community/ # Community tools (tavily, jina_ai, firecrawl, image_search, aio_sandbox)
? ? ??? reflection/ # Dynamic module loading (resolve_variable, resolve_class)
? ? ??? utils/ # Utilities (network, readability)
? ? ??? client.py # Embedded Python client (DeerFlowClient)
? ??? app/ # Application layer (import: app.*)
? ? ??? gateway/ # FastAPI Gateway API
? ? ? ??? app.py # FastAPI application
? ? ? ??? routers/ # FastAPI route modules (models, mcp, memory, skills, uploads, threads, artifacts, agents, suggestions, channels)
? ? ??? channels/ # IM platform integrations
? ??? tests/ # Test suite
? ??? docs/ # Documentation
??? frontend/ # Next.js frontend application
??? skills/ # Agent skills directory
??? public/ # Public skills (committed)
??? custom/ # Custom skills (gitignored)
Important Development Guidelines
Documentation Update Policy
CRITICAL: Always update README.md and CLAUDE.md after every code change
When making code changes, you MUST update the relevant documentation:
- Update
README.mdfor user-facing changes (features, setup, usage instructions) - Update
CLAUDE.mdfor development changes (architecture, commands, workflows, internal systems) - Keep documentation synchronized with the codebase at all times
- Ensure accuracy and timeliness of all documentation
Workflow Studio stream and canvas compatibility
app/gateway/workflow_agent_runner.pyis the boundary between LangGraph token chunks and durable workflow events. Provider reasoning inAIMessage.additional_kwargs.reasoning_content/reasoningis emitted to the UI as balanced inline<think>blocks; only visible answer content becomes the node'stextoutput for downstream nodes. It must also emit every original[AIMessage|ToolMessage, metadata]tuple asnode.message.data.messageswith no whitelist, preview or length crop. Tool results andadditional_kwargs.clarificationare needed by the Studio's DeerFlow-compatible message UI and must survive durable SSE replay.- An agent
ask_clarificationToolMessage pauses the workflow at that agent node (awaiting_input), not as a terminal run. The pause records the sourcenodeRunId/toolCallId; resuming turns the submitted reply into the next user message on the samewf-{run}-{node}LangGraph thread. Keep the raw ToolMessage authoritative; do not invent a second card payload. app/gateway/workflow_proposal_planner.pyconverts the workflow controller's selected catalog roles into per-agentmission/deliverable/scope/handoffcontracts. Only a selected visibleagentIdmay contribute a contract; every omitted or malformed field falls back to a deterministic role-aware contract. Preserve this prompt-level worker boundary when adding strategies: researchers must not silently become final writers, and sequential roles must receive their upstream handoff binding plus an explicit next handoff.app/gateway/routers/workflows_coze_compat.pymust return bothcreator.self: trueand editable VCS metadata for an owned standalone workflow. Coze Playground derives its preview/read-only mode from those fields, so omittingcreator.selfdisables node dragging even when the standalone page passesreadonly={false}.
Commands
Root directory (for full application):
make check # Check system requirements
make install # Install all dependencies (frontend + backend)
make dev # Start all services (Gateway + Frontend + Nginx), with config.yaml preflight
make start # Start production services locally
make stop # Stop all services
Backend directory (for backend development only):
make install # Install backend dependencies
make dev # Run Gateway API with reload (port 8001)
make gateway # Run Gateway API only (port 8001)
make test # Run all backend tests
make lint # Lint with ruff
make format # Format code with ruff
Regression tests related to Docker/provisioner behavior:
tests/test_docker_sandbox_mode_detection.py(mode detection fromconfig.yaml)tests/test_provisioner_kubeconfig.py(kubeconfig file/directory handling)
Boundary check (harness ? app import firewall):
tests/test_harness_boundary.py? ensurespackages/harness/deerflow/never imports fromapp.*
CI runs these regression tests for every pull request via .github/workflows/backend-unit-tests.yml.
Architecture
Position roundtable recovery persistence
/api/position-roundtable uses the shared /api/intent/* LangGraph flow as
the live Step-1 execution source and stores an additional SQL recovery copy in
position_roundtable_sessions/position_roundtable_nodes. The recovery copy
contains complete intent/node/summary/action-plan conversations, unsent
composer drafts, workflow snapshots, and bounded full text for generated
Markdown/JSON artifacts. This lets history and downstream closing agents keep
working when container-local checkpoint or sandbox volumes are replaced.
Revision 20260723_05 adds the conversation columns. Session deletion
explicitly removes child nodes so SQLite deployments do not depend on FK
cascade pragmas. The router is mounted canonically below
/api/position-roundtable and, for intranet gateway allowlists, below the
schema-hidden compatibility prefix /api/multi-agent/position-roundtable.
Clients may retry the compatibility route only for a framework/proxy-level
404 Not Found; domain-level missing-session responses must remain real 404s.
The position-roundtable state machine is multi-worker safe: node commands
(complete, bind, conversation save, reject) and session commands
(activate, archive, delete) lock the parent session and affected nodes in one
transaction, then record a unique command id for retry replay. The task and
intent snapshot writes use the same parent-row lock; they are permitted only
while status=intent_pending, so a late “confirm intent” write cannot change a
chain that another user has already activated. Activation itself rechecks the
stored confirmed intent after taking the lock. tests/test_position_roundtable_ commands.py and tests/test_position_roundtable_concurrency.py cover the
state transitions; tests/test_roundtable_postgres_concurrency.py is the real
PostgreSQL acceptance suite (set DEERFLOW_POSTGRES_CONCURRENCY_TEST_URL).
Roundtable backend jobs use a database unique active-dedupe key plus a leased
dispatcher, never process-local start deduplication. Cancellation is two
phase (cancel_requested then cancelled): the lease owner’s watcher normally
finalizes it, while the dispatcher reaps a cancellation whose lease has
expired after a worker crash. A worker that loses lease stops its local model
run immediately. Do not reintroduce direct executor.cancel() from the HTTP
cancel route, because it can cancel the watcher before the database state is
finalized. The boolean defaults on these new rows must use SQLAlchemy
false() rather than numeric 0, so fresh PostgreSQL schemas can be created.
Position roundtable selects the separately seeded built-in agent
position-roundtable-intent by passing intent_agent_id to the shared
/api/intent/init and /api/intent/stream endpoints. Its request lifecycle,
files, and SSE protocol match roundtable-intent, but its cap is five
user-assistance rounds. A single round may emit several independent
ask_clarification tool calls; those calls are separate cards, not one card
containing several numbered questions. Count that as one round by the
tool-calling AI message, never by card count. Topic-only requests must request
all material missing decisions as separate cards rather than being treated as a
resolved intent; after the fifth set is answered it must emit completed final
JSON without unresolved placeholders. The required arrays are
coreGoals, riskWarnings, keyPoints, and strategicSignificance. Do not
treat a fresh dedicated-agent ask_clarification as stale: it must win over
a same-run premature [INTENT_READY] and leave the UI waiting for the user.
Conversely, the prior turn’s historical clarification must not block the user
reply from completing. Do not
change roundtable-intent's {objective,constraints,assumptions} contract for
this feature. The position router accepts the new schema for activation and
injects all four sections into downstream position-node and closing-agent
context; legacy stored summaries are mapped read-only for compatibility.
/api/intent/init must create the empty Step-1 thread through
multi_agent._create_thread_direct, exactly like multi-agent initialization;
do not restore the old HTTP loopback to /api/threads, because deployed
Gateway ports/proxies may make the automatic page initialization fail while
later manual turns still work. If checkpoint creation fails or the parallel seat
task is cancelled, the helper compensates by deleting any partial checkpoint and
metadata row (the cancellation branch runs under asyncio.shield so a second
cancellation cannot abort it half-way). Streaming stays on the existing loopback
path.
ThreadMetaStore.create never issues a post-commit session.refresh(). Every
ThreadMetaRow column is assigned before commit and the session factory uses
expire_on_commit=False, so the returned dict is built from the in-memory row.
The removed read-back ran on a different pooled connection than the INSERT,
which a MySQL read/write-splitting endpoint can route to a lagging replica — the
row comes back empty and SQLAlchemy raises Could not refresh instance, surfacing
as 500 "Failed to create thread". Concurrency makes that far more likely, so a
single-user deployment never sees it. No caller uses the return value, and all
seven call sites (multi_agent._create_thread_direct, services.start_run,
routers/threads.create_thread, roundtable_inprocess_gateway,
scheduler/service, thread_shares, authz) benefit. Keep new create
implementations read-back-free; tests/test_position_roundtable_sessions.py:: test_thread_meta_create_never_reads_back_after_commit guards this.
Harness / App Split
The backend is split into two layers with a strict dependency direction:
- Harness (
packages/harness/deerflow/): Publishable agent framework package (deerflow-harness). Import prefix:deerflow.*. Contains agent orchestration, tools, sandbox, models, MCP, skills, config ? everything needed to build and run agents. - App (
app/): Unpublished application code. Import prefix:app.*. Contains the FastAPI Gateway API and IM channel integrations (Feishu, Slack, Telegram, DingTalk).
Dependency rule: App imports deerflow, but deerflow never imports app. This boundary is enforced by tests/test_harness_boundary.py which runs in CI.
Report Collaboration (AgentScope, in progress)
Independent of Workflow Studio and the existing roundtable. Plan: repo-root docs/agentscope-多智能体报告协作工作台-后端实施计划.md(RC-BE-000~018 已落地,下一票 RC-BE-019)。Spike baseline: docs/AGENTSCOPE_BASELINE.md.
- Config:
config.yaml → report_collaboration(deerflow.config.report_collaboration_config.ReportCollaborationConfig). Default off (enabled/worker_enabled). - Contracts:
app/report_collaboration/contracts/— snake_case Pydantic models matchingfrontend-web/src/report-collaboration/api/types.ts. SSE envelope uses camelCase (eventId/sessionId/runId).node_run_idis always${node_id}-attempt{attempt}. - Persistence:
deerflow.persistence.report_collaboration(report_collaboration_*tables, migrations20260905_01/20260906_01/20260907_01). SQL + memory stores; one active run per session viaactive_dedupe_key; command/message/run/version idempotency keys refuse cross-session reuse; eventseqtaken from the run row. Session delete removes children explicitly. - Gateway:
/api/report-collaboration(app/gateway/routers/report_collaboration.py).enabled=false→ 503. Plan/run writes enqueue durable jobs and return immediately. Pre-runPOST /sessions/{id}/messagesruns a rule-basedIntentRouterAgent+RequirementResolver(app/report_collaboration/requirement/): writes a requirement revision and clarification messages, never auto-starts planning.POST /sessions/{id}/plan-requeststhen runsPlanProposalService(app/report_collaboration/planning/): catalog projection (no SOUL/keys) → selector → 2–3 strategy RoleDemand graphs → assembler binds catalog agents/tools/quality gates → validator (acyclic DAG, artifact types, allowlisted tools). Select and start-run stay independent commands.POST /runsfreezes role models, report-structure template and run budget. Rewrite / apply / restore are live (reporting/).GET /runs/{id}/streamis a read-only durable SSE tail (execution/live_hub.py): persist then publish, replay +after_seq/Last-Event-ID, heartbeat comments, token-delta coalesce before seq assignment. Disconnect does not cancel the run. Lifespan mountsReportCollaborationDispatcheronly whenenabledandworker_enabledare both true; default both false. - Runtime: AgentScope adapters live in
app/report_collaboration/agentscope_runtime/(not the publishable harness).ConversationReplicamaps AgentScope 2.0.5 events 1:1 onto collaboration SSE types; TeamSay isHINT_BLOCK.DeerFlowChatModelAdapterwrapscreate_chat_model(no second Provider config); role models are frozen onto nodeagent_snapshotatPOST /runsand missing/incapable models fail with 422 before the run is created. Fallback is recorded on the snapshot plus a durableheartbeataudit event. Nodes complete only after a validated artifact (app/report_collaboration/execution/completion.py).TaskLedger(execution/ledger.py) owns ready waves, attempts and terminal node status; AgentScope / coordinator suggestions never pick the next node. A writer-bearing run iscompletedonly afterReportFinalizerpublishes a formal version.WaveExecutor+QualityRoleKernel(execution/quality_roles.py) remain for fixtures: Verifier marks stale/contested claims, Analyst/Writer only cite supported IDs, Reviewerrevisereplans only the target write node plus review. Production worker injectsAgentScopeMemberKernel(agentscope_runtime/member_kernel.py) with DeerFlow tools and per-run budget.QualityGate(execution/quality_gate.py) is the claim-level gate: conflicts/missing units, unverified facts, coverage gaps, and unmapped review targets fail beforevalidated;ArtifactBoard.supersede_previoushides old attempts from downstream.ReportCollaborationLiveHub(execution/live_hub.py) fans out persisted envelopes to SSE viewers;PublishingReportCollaborationStoresanitizes secrets/privacy/hidden reasoning without truncating visible tool results, and writes redaction + event audits. Dispatcher (execution/dispatcher.py) atomically claims queued/recovering/expired-lease runs, heartbeats viarenew_lease, two-phase cancels (cancel_requested→finalize_cancel), and recovers orphans (execution/recovery.py): keep completed work, mark leftover in-flightinterrupted/superseded, open a newplannedattempt. Late artifacts after cancel/reclaim aresupersededand must not change terminal run status. Tests drivedispatch_once(); do not rely on the background loop. Per-run budgets (security/budget.py) fail-closed; per-user active-run caps are enforced atcreate_run. - Quality eval:
app/report_collaboration/eval/(RC-BE-018) — 20 annotated tasks (policy/market/enterprise/event/comparison ×4) with required angles, key facts, forbidden fabrications, templates and review rules. Deterministic scorer (no LLM judge) gates citation/angle/template/directed-mod; gold reports must pass, canned failures (including a roundtable-style essay) must fail and stay reproducible. Recommended runtime knobs live ineval/thresholds.pyand match currentReportCollaborationConfigdefaults (worker_enabledremains false). Tests:tests/test_report_collaboration_eval.py. - Runtime intervention:
app/report_collaboration/interventions/owns run-time natural-language commands.RuntimeIntentRouterAgentuses explicit rules first and a tool-free coordinator-model fallback only for ambiguity; its Pydantic decision cannot dispatch or mutate anything.ImpactAnalyzervalidates server-side node/agent ownership, resolves angle-specific lineage, and enforces confidence/cost/parallel-branch confirmation (intent_confirmation_threshold, default0.8).RuntimeCommandServicemirrors the human message, emits frozen command events, CAS-claims accepted commands, and applies them only throughTaskLedgersafe points. Rework persists command overlays, supersedes affected attempts/artifacts, creates monotonic attempts, and prevents late output from entering downstream. Waiting-node answers resume only those nodes.rerun_nodewithout a legal target does not expand to the whole graph; requirement patches apply only after rework succeeds; stuckexecutingcommands are reclaimed on takeover or lease timeout;command.cancelledis a durable SSE event. In-runPOST /sessions/{id}/messagesreturns 409 (RUN_ACTIVE). Formal rewrite candidates and version apply/restore live inreporting/(RC-BE-015). - Dependency: optional extra
report-collaboration→agentscope==2.0.5. Do not add AgentScope to the harnesspyproject.toml.
Tests: tests/test_report_collaboration_*.py.
Workflow Studio
Coze Studio is canvas-only; ChatDev is reference-only. Plan + status per phase: docs/WORKFLOW_STUDIO_BACKEND_DEV_ZH.md; decisions: docs/adr/0001-workflow-studio-runtime.md; contracts: docs/WORKFLOW_SCHEMA_V1_ZH.md + docs/WORKFLOW_SSE_V1_ZH.md + docs/WORKFLOW_OPENAPI_DRAFT_V1.yaml. Master switch + all limits: config.yaml → workflows (deerflow.config.workflow_config.WorkflowConfig). workflows.agent_recursion_limit defaults to 250: LangGraph counts each model→tool research exchange as about two super-steps, so workflow agent/skill nodes must not fall back to a chat-sized 60-step limit. In the local Studio proxy, keep Rsbuild server.compress: false; dev gzip buffers small SSE frames and makes live planning/run output arrive only at completion despite the backend's no-transform / X-Accel-Buffering: no headers. The composer’s optional modelName must be validated against AppConfig.models: it selects the controller model and is copied only into newly generated candidate agent configs (or the projected direct-agent task), never injected into an already-confirmed proposal/run.
Execution kernel is a custom DAG scheduler, not a compiled LangGraph. LangGraph stays the kernel inside an agent/skill node. The reason is single-source-of-truth recovery: a re-leased or resumed run replays completed node outputs from workflow_node_runs plus the workflow_run_events log, so a second (checkpoint) copy of workflow-level state would have no defensible tie-breaker. See the ADR's phase-2 amendment.
Harness (packages/harness/deerflow/workflows/, never imports app.*):
schemas/events/errors/ports/validator/samples— frozen v1 contracts.expressions.py—{{ inputs.x }}/{{ nodes.n.data.y }}templates and conditions resolved by path walking + a small comparator set. Noeval. Roots are restricted toinputs/nodes/run/loop/env.runtime/engine.py— the scheduler. Every edge ispending→active(source chose this port) orpruned; a node is ready when its non-back incoming edges are all resolved and at least one isactive, and is skipped when all arepruned(skips propagate). That one rule covers fan-out, condition branching, and merge joins with no topological pre-pass. Theloopnode'scontinueback edge is the only permitted cycle; re-entry clears the body's results and invalidates its node runs.runtime/{sink,sql_runner,code_runner,context}.py,nodes/*(15 executors, including deterministicevidence_normalizerand durabledeep_research_write),security/{http_policy,sql_policy,secret_box}.py.
Persistence: workflow_definitions / workflow_versions (migration 20260828_01), workflow_runs / workflow_node_runs / workflow_run_events / workflow_run_artifacts (20260828_02), workflow_data_sources (20260828_03), and short-lived-but-durable workflow_planning_sessions / workflow_planning_proposals (20260831_01). Stores under deerflow.persistence.workflow_{runs,events,data_sources,planning} with SQL + in-memory implementations. Because planning proposal rows have a scalar FK rather than an ORM relationship, SqlWorkflowPlanningStore.create_session must flush the parent session before adding its children; otherwise SQLite can insert the child first and return a 500. A planning proposal is distinct from a run: its graph can be edited with optimistic graph_revision checks, but it becomes executable only once selected and atomically bound to one run.
App layer: workflow_dispatcher.py (lease scan → atomic claim; an expired lease is how a dead worker's run is reclaimed, so every worker can run the loop), workflow_executor.py (one lease's lifecycle: replay → engine under a run timeout with heartbeat + cancel watcher → exactly one terminal transition; also runs subworkflows inline, depth ≤ 3), workflow_retention.py (hourly sweep of workflows.retention.event_retention_days / run_retention_days), workflow_live_hub.py (in-process SSE fan-out), workflow_agent_runner.py (agent/skill nodes → lead agent; disable_tools=True means an empty bound-tool set for control-plane callers), workflow_deep_research_adapter.py (Evidence Pack → selected Deep Research sources/job/event projection; no internal HTTP call), and workflow_proposal_planner.py (calls the dedicated workflow-planner controller over visible-agent metadata, then constructs safe candidate graphs server-side). workflow-planner is a separate Workflow Studio control-plane agent; never reuse or alias roundtable-coordinator for workflow planning, and never let it dispatch a run, call research tools, or return arbitrary graph JSON. It always runs with tools and model thinking disabled.
Gateway routes (register static prefixes before workflows.router's /{workflow_id} catch-alls): /api/workflows CRUD/draft/validate/publish/copy, /api/workflows/node-types + /resources/*, /api/workflows/{id}/planning-sessions{/stream} + /api/workflows/planning-sessions/{session_id}{,/proposals/{proposal_id},/proposals/{proposal_id}/graph}, /api/workflows/{id}/runs + /api/workflows/runs/{run_id}{,/events,/stream,/cancel,/resume,/retry,/feedback,/artifacts}, /api/workflows/data-sources (+ /{id}/introspect, /{id}/validate-query), /api/workflows/embed/{tickets,verify}, /api/workflow_api/* Coze compat. The planning /stream endpoint sends product-safe real controller milestones (catalog, controller, selection, assembly, validation, complete) then a final result session; never stream agent catalog descriptions or raw planner JSON. Formal runs accept the selected planningSessionId and proposalId; creating the same pair again returns the already-bound run rather than launching another execution.
Guarantees worth not breaking: events are persisted then published, and the DB is authoritative for seq ordering (live frames at or below the replay cursor are dropped and re-read on the next poll, so a reconnect sees no gap; Last-Event-ID wins over ?after=); a run reaching a terminal status always gets a terminal event, synthesised if necessary, or SSE clients hang; resumeToken appears only in the run.awaiting_input event, never in the run view; POST /retry on a failed run copies completed node outputs onto a new run marked retry_of_run_id so the engine resumes from the failed node; POST /feedback never mutates its source run—on a terminal acyclic run it may target only a completed agent/skill, reuses completed nodes outside that node's downstream closure, records those reusedNodeIds alongside the feedback summary, and puts the bounded feedback only in RunContext.env.nodeFeedback[target] before creating a revision run; data-source secrets are write-only (encrypted via WORKFLOW_SECRET_KEY, decrypted only inside the executor, reads expose masked_target); the http node denies private/loopback/link-local targets by default (SSRF), sql_read allows one read-only statement with bound parameters, and the code node is off by default because its runner is a hardened subprocess rather than a container.
HTTP interface resources persist a non-secret base_url plus an allowed_methods allowlist (migration 20260829_01); request headers stay only in the encrypted secret blob. The resource catalog may return the address and methods to the owning user, but must never return headers or any decrypted secret. The executor receives HTTP metadata only through its data-source resolver.
The Workflow Studio agent/skill resource catalog is a read-only display projection over the existing
agent and skill ownership stores. It returns stable agentId / skillId, an agent's cached skill names,
and a createdBy display label. Resolve human labels from the immutable owner id at request time and
fall back to that id if lookup fails; built-in records must return 系统内置. Do not create a parallel
creator table or accept creator data from the browser.
evidence_normalizer is a deterministic data node, not an LLM: it consumes only configured/upstream
node results and emits a bounded evidencePack of claims with node provenance, role labels and gaps.
The planner inserts it between parallel research roles and the synthesis Agent. It must not treat
source text as verified fact or fetch any external resource, and its explicit sources[].nodeId values
must stay upstream-only under validate_workflow_graph.
deep_research_write is a durable report node, not an Agent prompt shortcut. It accepts only an
Evidence Pack, derives a session identity from the workflow run/node plus the topic, evidence snapshot and writing config, writes
the pack as selected other source rows with upstream-node provenance, and queues/reuses the existing
Deep Research regenerate job. Its live report_delta and persisted report_chunk frames are
deduplicated before becoming workflow node.output.delta events; only report prose is forwarded (never
planning/reasoning deltas). Workflow cancellation requests cancellation on that same job. Do not make it
self-call /api/deep-research/*, use Deep Research multi_agent inside this node, or treat Evidence
Pack claims as verified facts. The parallel-research candidate grants the node an explicit 1800-second
timeout; deployments that override workflows.node_timeout_seconds must keep it at least that high.
Tests: tests/test_workflow_schemas.py, tests/test_workflow_definitions.py, tests/test_workflow_engine.py, tests/test_workflow_runs.py.
Import conventions:
# Harness internal
from deerflow.agents import make_lead_agent
from deerflow.models import create_chat_model
# App internal
from app.gateway.app import app
from app.channels.service import start_channel_service
# App ? Harness (allowed)
from deerflow.config import get_app_config
# Harness ? App (FORBIDDEN ? enforced by test_harness_boundary.py)
# from app.gateway.routers.uploads import ... # ? will fail CI
Agent System
Lead Agent (packages/harness/deerflow/agents/lead_agent/agent.py):
- Entry point:
make_lead_agent(config: RunnableConfig)registered inlanggraph.json - Dynamic model selection via
create_chat_model()with thinking/vision support - Tools loaded via
get_available_tools()- combines sandbox, built-in, MCP, community, and subagent tools - System prompt generated by
apply_prompt_template()with skills, memory, and subagent instructions
agentfx built-in analyst + report convert skills. agentfx-analyst (「任务研判报告助手」) is seeded from app/gateway/routers/_agentfx_seed_assets/ before _sync_legacy_agents. After mandatory web_search, the agent writes each tool JSON to /mnt/user-data/workspace/hits/qN.json and runs one command — task-report-build (group all) — to fill the 14-category task-reports.json for the dashboard slots; the four per-page skills task-report-{enemy,our,env,judge} remain for re-running a single dashboard page. All five are thin CLIs over the shared stdlib engine skills/public/task-report-lib/convert_lib.py (no SKILL.md on the lib, so it never enters the skill list). The model must not write_file that JSON. Because the sandbox maps /mnt/user-data/... onto a host dir, a virtual absolute hits path can be isdir-true yet list empty on a Windows host — resolve_hit_files therefore tries the path as given, then the cwd-mapped sandbox path (discover_user_data_root + resolve_sandbox_path — bash cds into the thread workspace, so /mnt/user-data/outputs/task-reports.json must land in that thread's outputs/, never F:\mnt\... on a Windows host), then the virtual-prefix-stripped relative form, then the bare dir name, and picks the first candidate that actually contains .json; a miss returns a warnings entry listing every tried form so the agent never has to write a debug script. Assemble of the 14 categories is the same function (assemble_records); the model must not copy or merge files. If the model is interrupted after convert, task-reports.json is already in outputs and the frontend import card also scans GET /api/threads/{id}/artifacts so入库 still works without present_files. Import remains task-report-import. Existing agent/skill directories are never overwritten. Tests: tests/test_task_report_convert.py, tests/test_agentfx_seed.py.
Forced-research built-in agent. forced-research-responder (「强制检索输出助手」) is seeded from app/gateway/routers/_forced_research_seed_assets/ before _sync_legacy_agents. ForcedResearchMiddleware is mounted only for this stable id. On every visible user turn it requires a successful information-gathering tool result before the model may finish; tool_choice=required is backed by an after_model -> model retry for gateways that ignore tool choice, with an eight-attempt failure cap that permits only an honest failure response. Skill discovery, reading SKILL.md, and writing a helper script do not satisfy the gate; web/knowledge/MCP results, non-skill uploaded-file reads, or successful Python skill-script execution do. Tool filtering removes present_files, skill_manage, task/schedule/memory/clarification surfaces. The middleware permits write_file/str_replace only for /mnt/user-data/workspace/*.py and permits bash only for .py execution without redirection/document paths, so skills can run scripts but the agent cannot create Markdown/document artifacts or write outputs. Tests: tests/test_forced_research_agent.py.
ThreadState (packages/harness/deerflow/agents/thread_state.py):
- Extends
AgentStatewith:sandbox,thread_data,title,artifacts,todos,uploaded_files,viewed_images - Uses custom reducers:
merge_artifacts(deduplicate),merge_viewed_images(merge/clear)
Runtime Configuration (via config.configurable):
thinking_enabled- Enable model's extended thinkingmodel_name- Select specific LLM modelis_plan_mode- Enable TodoList middlewaresubagent_enabled- Enable task delegation tool
Middleware Chain
Lead-agent middlewares are assembled in strict append order across packages/harness/deerflow/agents/middlewares/tool_error_handling_middleware.py (build_lead_runtime_middlewares) and packages/harness/deerflow/agents/lead_agent/agent.py (_build_middlewares):
- ThreadDataMiddleware - Creates per-thread directories under the creator's stable isolation scope (
backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/{workspace,uploads,outputs}); resolves persisted threads throughresolve_path_user_id(thread_id)(falls back toget_effective_user_id()/"default"when metadata is unavailable); Web UI thread deletion now follows LangGraph thread removal with Gateway cleanup of the local thread directory - UploadsMiddleware - Tracks and injects newly uploaded files into conversation
- SandboxMiddleware - Acquires sandbox, stores
sandbox_idin state - DanglingToolCallMiddleware - Injects placeholder ToolMessages for AIMessage tool_calls that lack responses (e.g., due to user interruption), including raw provider tool-call payloads preserved only in
additional_kwargs["tool_calls"] - LLMErrorHandlingMiddleware - Normalizes provider/model invocation failures into recoverable assistant-facing errors before later middleware/tool stages run
- GuardrailMiddleware - Pre-tool-call authorization via pluggable
GuardrailProviderprotocol (optional, ifguardrails.enabledin config). Evaluates each tool call and returns error ToolMessage on deny. Three provider options: built-inAllowlistProvider(zero deps), OAP policy providers (e.g.aport-agent-guardrails), or custom providers. See docs/GUARDRAILS.md for setup, usage, and how to implement a provider. - SandboxAuditMiddleware - Audits sandboxed shell/file operations for security logging before tool execution continues
- ToolErrorHandlingMiddleware - Converts tool exceptions into error
ToolMessages so the run can continue instead of aborting - SummarizationMiddleware (
DeerFlowSummarizationMiddleware) - Context reduction when approaching token limits (optional, if enabled). Non-destructive: overrides the stockbefore_model(which persists aRemoveMessage(REMOVE_ALL_MESSAGES)) to a no-op and instead compresses transiently inwrap_model_call? the model request is replaced with[summary, *recent]while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). Thesummarymessage ishide_from_uiand never persisted; summaries are cached per-thread in a process-wide, thread-safe LRU (_GLOBAL_SUMMARY_CACHE, bounded to 512 threads, sticky boundary). The cache is module-level — not on the middleware instance — becausemake_lead_agentrebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed bythread_id(resolved likeThreadDataMiddleware), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. Threshold-first gating:_plan_compressionreturns passthrough (no compression at all) whenever the full current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. Stickiness headroom (_apply_stickiness_headroom,_KEEP_HEADROOM_FRACTION=0.75): after thekeeppolicy picks a cutoff, the kept suffix is forced below ~75% of every trigger before summarizing. Without it akeepin messages (e.g. 20) can itself exceed a token trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a{"type":"context_compacting","message":"正在压缩上下文…"}event on LangGraph'scustomstream channel (the cheap reuse path stays silent); the chat frontend'sonCustomEventrenders it as a transient toast. Tests:tests/test_summarization_middleware.py - TodoListMiddleware - Task tracking with
write_todostool (optional, if plan_mode) - TokenUsageMiddleware - Records token usage metrics when token tracking is enabled (optional)
- TitleMiddleware - Two-phase auto-title (never shows "untitled"):
before_modelsets an instant question-derived provisional title the moment the first user message arrives (no LLM, zero latency), thenafter_modelasynchronously generates the final LLM title and replaces it (tracked via thetitle_provisionalstate flag). Normalizes structured message content before prompting the title model - MemoryMiddleware - Memory V2:
wrap_model_callinjects Hindsight recall transiently;after_agentauto-retains the turn to Hindsight - ViewImageMiddleware - Injects base64 image data before LLM call (conditional on vision support)
- DeferredToolFilterMiddleware - Hides deferred tool schemas from the bound model until tool search is enabled (optional)
- SubagentLimitMiddleware - Truncates excess
tasktool calls from model response to enforceMAX_CONCURRENT_SUBAGENTSlimit (optional, ifsubagent_enabled) - LoopDetectionMiddleware - Detects repeated tool-call loops; hard-stop responses clear both structured
tool_callsand raw provider tool-call metadata before forcing a final text answer - ClarificationMiddleware - Intercepts
ask_clarificationtool calls, interrupts viaCommand(goto=END)(must be last) 8a. ToolMetricsMiddleware - Appended immediately afterToolErrorHandlingMiddlewarein_build_runtime_middlewares, so it sits deepest in thewrap_tool_callchain. Records onetool_call_metricsrow per tool/skill/MCP invocation ? timing duration and bucketing failures into a coarseerror_category(rate_limited / ip_blocked / auth_error / timeout / network_error / not_found / empty_result / tool_error). It catches both raised exceptions (before #8 converts them) and self-handled errors returned as errorToolMessages (search tools that swallow HTTP 429/403). It also captures the targeted Agent Skill intotool_call_metrics.skill_name: the lead agent invokes a skill byread_file-ing its/mnt/skills/<category>/<name>/SKILL.mdpath (per the<skill_system>prompt block ? there is no dedicated "use skill" tool), soskill_nameis recovered by scanning every tool call's string arguments for theskills/{public,custom}/{name}/path shape (_SKILL_PATH_RE), plus the explicitnameargument ofskill_view/skill_manage/skill_archive.skill_view-specific success detection handles its swallowed failures (Skill 'x' not found.returned as a plain string rather than raised). Regression coverage:tests/test_tool_metrics_middleware.py. Writes are best-effort fire-and-forget. Per-run skill merging:SqlToolMetricsStore.recordupserts skill rows by(run_id, skill_name)instead of inserting one row per call ? a skill touched by several tool calls / retries inside one run collapses to a single row. Any success in the run wins (a skill that failed then finally succeeded is recorded as onesuccess); a skill that only ever failed yields oneerrorrow whoseerror_messageaccumulates every attempt's message; durations are summed. Non-skill tool calls, and skill calls with norun_id, keep one row each. The read?merge?write is guarded by a per-event-loopasyncio.Lock. Surfaced by the adminGET /api/admin/tool-metricsrouter (which has akind=skillfilter and a per-skill/summarybreakdown). - PromptPrefixMiddleware - Injects the admin-configured prompt prefix into the model request only. The prefix is tagged onto the first human message's
additional_kwargs['prompt_prefix']by the Gateway (app.gateway.services) and prepended to content here at call time, so the persisted/displayed message stays prefix-free - SummarizationMiddleware (
DeerFlowSummarizationMiddleware) - Context reduction when approaching token limits (optional, if enabled). Non-destructive: overrides the stockbefore_model(which persists aRemoveMessage(REMOVE_ALL_MESSAGES)) to a no-op and instead compresses transiently inwrap_model_call? the model request is replaced with[summary, *recent]while the full original conversation stays in state (so the frontend keeps rendering the user's real Q&A). Thesummarymessage ishide_from_uiand never persisted; summaries are cached per-thread in a process-wide, thread-safe LRU (_GLOBAL_SUMMARY_CACHE, bounded to 512 threads, sticky boundary). The cache is module-level — not on the middleware instance — becausemake_lead_agentrebuilds a fresh middleware on every run; an instance cache would be cold each turn and force a blocking summary LLM call on every turn of a long conversation. Keyed bythread_id(resolved likeThreadDataMiddleware), so a summary computed on one turn is reused on the next, paying for an LLM call only on the first threshold crossing or a genuine re-compaction. Threshold-first gating:_plan_compressionreturns passthrough (no compression at all) whenever the full current conversation is still under the trigger — the cache/summary substitution only engages once the real context crosses the configured trigger. Stickiness headroom (_apply_stickiness_headroom,_KEEP_HEADROOM_FRACTION=0.75): after thekeeppolicy picks a cutoff, the kept suffix is forced below ~75% of every trigger before summarizing. Without it akeepin messages (e.g. 20) can itself exceed a token trigger (e.g. 60000) when recent messages are large (tool outputs / pasted docs), leaving the "compressed" context still over the trigger so it re-summarizes on every follow-up question; the headroom makes the sticky boundary actually stick (re-compaction only after genuinely new content). When the genuine (blocking) summarize path runs, it emits a{"type":"context_compacting","message":"正在压缩上下文…"}event on LangGraph'scustomstream channel (the cheap reuse path stays silent); the chat frontend'sonCustomEventrenders it as a transient toast. Tests:tests/test_summarization_middleware.py - TodoListMiddleware - Task tracking with
write_todostool (optional, if plan_mode) - TokenUsageMiddleware - Records token usage metrics when token tracking is enabled (optional)
- TitleMiddleware - Two-phase auto-title (never shows "untitled"):
before_modelsets an instant question-derived provisional title the moment the first user message arrives (no LLM, zero latency), thenafter_modelasynchronously generates the final LLM title and replaces it (tracked via thetitle_provisionalstate flag). Normalizes structured message content before prompting the title model - MemoryMiddleware - Queues conversations for async memory update (filters to user + final AI responses)
- ViewImageMiddleware - Injects base64 image data before LLM call (conditional on vision support)
- DeferredToolFilterMiddleware - Hides deferred tool schemas from the bound model until tool search is enabled (optional)
- SubagentLimitMiddleware - Truncates excess
tasktool calls from model response to enforceMAX_CONCURRENT_SUBAGENTSlimit (optional, ifsubagent_enabled) - LoopDetectionMiddleware - Detects repeated tool-call loops; hard-stop responses clear both structured
tool_callsand raw provider tool-call metadata before forcing a final text answer - ClarificationMiddleware - Intercepts
ask_clarificationtool calls, interrupts viaCommand(goto=END)(must be last)
Configuration System
Main Configuration (config.yaml):
Setup: Copy config.example.yaml to config.yaml in the project root directory.
Config Versioning: config.example.yaml has a config_version field. On startup, AppConfig.from_file() compares user version vs example version and emits a warning if outdated. Missing config_version = version 0. Run make config-upgrade to auto-merge missing fields. When changing the config schema, bump config_version in config.example.yaml.
Config Caching: get_app_config() caches the parsed config, but automatically reloads it when the resolved config path changes or the file's mtime increases. This keeps Gateway and LangGraph reads aligned with config.yaml edits without requiring a manual process restart.
Configuration priority:
- Explicit
config_pathargument DEER_FLOW_CONFIG_PATHenvironment variableconfig.yamlin current directory (backend/)config.yamlin parent directory (project root - recommended location)
Config values starting with $ are resolved as environment variables (e.g., $OPENAI_API_KEY).
ModelConfig also declares use_responses_api and output_version so OpenAI /v1/responses can be enabled explicitly while still using langchain_openai:ChatOpenAI.
Extensions Configuration (extensions_config.json):
MCP servers and skills are configured together in extensions_config.json in project root:
Configuration priority:
- Explicit
config_pathargument DEER_FLOW_EXTENSIONS_CONFIG_PATHenvironment variableextensions_config.jsonin current directory (backend/)extensions_config.jsonin parent directory (project root - recommended location)
Gateway API (app/gateway/)
FastAPI application on port 8001 with health checks at GET|HEAD /health, /api/health, and /api/langgraph/health. These exact paths are always public in auth_middleware.PUBLIC_HEALTH_PATHS, including when authentication is enabled; reverse-proxy prefixes and trailing slashes are normalized before the whitelist check. Set GATEWAY_ENABLE_DOCS=false to disable /docs, /redoc, and /openapi.json in production (default: enabled). Regression coverage: tests/test_health_auth_bypass.py.
Routers:
| Router | Endpoints |
|---|---|
Models (/api/models) |
GET / - list models; GET /{name} - model details |
MCP (/api/mcp) |
GET /config - get config; PUT /config - update config (saves to extensions_config.json) |
Skills (/api/skills) |
GET / - list skills; GET /{name} - details; PUT /{name} - update enabled; POST /install - install from .skill archive (accepts standard optional frontmatter like version, author, compatibility) |
Memory V2 (/api/memory/v2) |
GET / - builtin data (USER.md/MEMORY.md); POST/PUT/DELETE /entries - add/replace/remove entry; GET/PATCH/DELETE /config - effective config + per-user overrides; POST /search - Hindsight recall/reflect proxy |
Agents (/api/agents) |
GET / - list built-in, owned, and published agents; POST / - create user-owned agent; GET/PUT/DELETE /{id} - read/update/delete by stable agent id; GET /check - validate non-empty display name (duplicates allowed) |
Skills (/api/skills) |
GET / - list skills (optional search fuzzy-name + tag_id filters; each item carries tags); GET /{name} - details; PUT /{name} - update enabled; POST /install - install from .skill archive (accepts standard optional frontmatter like version, author, compatibility) |
Skills (/api/skills) |
GET / - list skills (optional search fuzzy-name + tag_id filters; each item carries tags + owner-authored detail); GET /{name} - details; PUT /{name} - update enabled; PUT /{name}/detail - set the long-form detail text (custom skill: owner or admin; built-in/public: admin only ? materializes a skills row on demand); POST /install - install from .skill archive (accepts standard optional frontmatter like version, author, compatibility) |
Skills (/api/skills) |
GET / - list skills (optional search fuzzy-name + tag_id filters; each item carries tags + owner-authored detail); GET /{name} - details; PUT /{name} - update enabled; PUT /{name}/detail - set the long-form detail text (custom skill: owner or admin; built-in/public: admin only ? materializes a skills row on demand); GET /{name}/content - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; POST /install - install from .skill archive (accepts standard optional frontmatter like version, author, compatibility) |
Skills (/api/skills) |
GET / - list skills (optional search fuzzy-name + repeatable tag_ids multi-select filter, OR-matched; each item carries tags + owner-authored detail); GET /{name} - details; PUT /{name} - update enabled; PUT /{name}/detail - set the long-form detail text (custom skill: owner or admin; built-in/public: admin only ? materializes a skills row on demand); GET /{name}/content - raw SKILL.md source, restricted to an admin or the publisher (custom-skill owner); other users get 403 and see only description + detail; POST /install - install from .skill archive (accepts standard optional frontmatter like version, author, compatibility) |
Memory (/api/memory) |
GET / - memory data; POST /reload - force reload; GET /config - config; GET /status - config + data |
Agents (/api/agents) |
GET / - list built-in, owned, and published agents (optional search + repeatable tag_ids multi-select filter, OR-matched; each item carries tags and featured_order); POST / - create user-owned agent; GET/PUT/DELETE /{id} - read/update/delete by stable agent id; GET /check - validate non-empty display name (duplicates allowed); GET /selector - homepage agent picker (single admin-curated list, visibility-filtered, sorted by featured_order asc); PUT /{id}/featured - admin sets/clears the homepage featured_order (must be published=true to feature; passing order: null clears unconditionally); PUT /featured/order - admin bulk-reorders the homepage list (body {agent_ids: [...]} ? indices become 0..N-1) |
Tags (/api/tags) |
GET / - list tags (optional tag_type filter; open to all); POST / - create tag (admin); PUT /{id} / DELETE /{id} - update/delete tag (admin); GET /assignments/{target_type}/{target_id} - tags on a resource; PUT /assignments/{target_type}/{target_id} - replace a resource's tags (admin). Tags are typed agent/skill/scheduled_task and may only tag resources of their own type. |
Uploads (/api/threads/{id}/uploads) |
POST / - upload files (auto-converts PDF/PPT/Excel/Word); GET /list - list; DELETE /{filename} - delete |
Threads (/api/threads/{id}) |
DELETE / - remove DeerFlow-managed local thread data after LangGraph thread deletion; unexpected failures are logged server-side and return a generic 500 detail |
Threads search (POST /threads/search, LangGraph-compat) |
Conversation listing. System threads are excluded by default: rows whose metadata.system is truthy (scheduler runs, roundtable workers) never crowd the user's list; a caller whose metadata filter explicitly references thread_type/system (e.g. studio notebooks) opts out of the exclusion. New query field = case-insensitive title (display_name) substring search, used by the chats page's server-side pagination. POST /threads/count takes the same request body (limit/offset ignored) and returns {total} under identical filtering/exclusion semantics ? powers the chats page's ??/??? display without fetching rows (ThreadMetaStore.count). ThreadMetaStore.search gained exclude_system/query params; JSON-blob filters batch-scan until the requested limit/offset window is full (so pagination over a heavily-filtered set never under-fills), with a thread_id tie-breaker on the updated_at DESC ordering for stable OFFSET paging. Tests: tests/test_thread_meta_search.py |
Artifacts (/api/threads/{id}/artifacts, /api/artifact-library) |
GET /{path} - serve artifacts; active content types (text/html, application/xhtml+xml, image/svg+xml) are always forced as download attachments to reduce XSS risk; ?download=true still forces download for other file types. GET /api/artifact-library scans every current-user normal thread's authoritative outputs/ directory (including files omitted from present_files), excludes uploads and image/audio/video media, joins the owning conversation title/agent route metadata, and exposes search/filter/pagination with a bounded 15-second per-user cache plus refresh=true bypass. Frontend route /page/workspace/library links each item back to the source chat with ?artifact=<path> so the chat opens the existing artifact panel directly. |
Word export (POST /api/writing/export/docx) |
Shared Markdown → DOCX generator used by conversation, artifact, scheduled-task, Deep Research, and AI-writing download surfaces. app/gateway/word_export.py builds OOXML with no Python conversion dependency: A4 portrait, 3.7/3.5/2.8/2.6 cm margins, real multilevel heading numbering (一、 / (一) / 1.), two-character first-line indent on heading levels 1–3, formal body/table styles, page numbers beginning at 1 with odd pages left-aligned and even pages right-aligned, and embedded app/gateway/assets/fonts/方正小标宋简体.ttf. Before numbering it removes complete manual heading prefixes (including 2.1 / 2.1.1) and drops Markdown horizontal rules from Word output. The ordinary-Q&A Lead Agent can emit the matching Markdown heading forms and avoid unordered bullet lists when the admin enables prompt_prefix.ordinary_qa_markdown_format_enabled in system_settings.json; the switch defaults off, while writing setup, writing mode, notebook, scheduled runs, and SOUL-driven custom agents keep their independent prompts. The font loader validates its internal family, SHA-256, and OpenType embedding permission (OS/2.fsType) before export. Tests: tests/test_word_export.py, tests/test_es_query_routing_prompt.py. |
Deep Research (/api/deep-research) |
Durable, user-owned research sessions/jobs with replayable SSE events and hidden-thread artifacts. basic/quick use the compact runner; detailed creates bounded subtopics, searches/writes sections concurrently, and assembles a report; deep recursively searches follow-up questions under a deterministic breadth/depth query budget. multi_agent is a LangGraph 1.x editor/researcher/writer/reviewer graph: its initial plan emits awaiting_input, the recorder stores the interrupt and releases the lease, and POST /jobs/{id}/resume rebuilds the graph from the persisted approved/revised plan. A bounded custom_outline is frozen into the session; headings/list entries deterministically supply the detailed and multi-agent plan, and the shared writer receives it only as a guarded content constraint. Completed reports support owner-scoped follow-up (GET /sessions/{id}/messages, POST /sessions/{id}/chat): context is only the current report, selected sources and recent messages; source-id citations are persisted, while allow_new_research=true is explicitly rejected. Owners may change that evidence set only after the job stops via PATCH /sessions/{id}/sources/{source_id}; this applies to future follow-ups rather than rewriting a completed report. Optional report illustrations are gated by config.yaml -> deep_research.image_generation: the Harness OpenAICompatibleImageProvider requires base64 output, the shared runner step writes only bounded `research-image-01..04.(png |
| Gateway auth | Owner-scoped Deep Research routes await the platform's asynchronous get_current_user resolver before they read or write persistence, ensuring each session and job operation receives a concrete user ID. Deep Research artifacts are always resolved from a whitelisted filename via /mnt/user-data/outputs/<name> before mapping to a host path, preserving the sandbox guard on native Windows and Linux hosts. |
Suggestions (/api/threads/{id}/suggestions) |
POST / - generate follow-up questions; rich list/block model content is normalized before JSON parsing |
Thread Runs (/api/threads/{id}/runs) |
POST / - create background run; POST /stream - create + SSE stream; POST /wait - create + block; GET / - list runs; GET /{rid} - run details; POST /{rid}/cancel - cancel; GET /{rid}/join - join SSE; GET /{rid}/messages - paginated messages {data, has_more}; GET /{rid}/events - full event stream; GET /../messages - thread messages with feedback; GET /../token-usage - aggregate tokens |
Feedback (/api/threads/{id}/runs/{rid}/feedback) |
PUT / - upsert feedback; DELETE / - delete user feedback; POST / - create feedback; GET / - list feedback; GET /stats - aggregate stats; DELETE /{fid} - delete specific |
Runs (/api/runs) |
POST /stream - stateless run + SSE; POST /wait - stateless run + block; GET /{rid}/messages - paginated messages by run_id {data, has_more} (cursor: after_seq/before_seq); GET /{rid}/feedback - list feedback by run_id |
Open Chat (/api/open/chat) |
Unauthenticated open Q&A entry for external systems (the /api/open/ prefix is whitelisted in auth_middleware._PUBLIC_PATH_PREFIXES + CSRF-exempt in csrf_middleware.should_check_csrf). POST /chat {message | messages[], model_name?, agent_name?, agent_id?, thinking_enabled=false, strip_think=true, thread_id?, recursion_limit=100} → resolves the agent by Chinese display name (agent_name, matched against each .deer-flow/agents/*/config.yaml name via _resolve_agent_id; explicit agent_id wins if given), builds the lead agent via make_lead_agent (so the chosen agent's SOUL/tools/skills all apply), ainvokes it once (stateless — no checkpointer, no thread persistence; runs under the "default" user bucket), and returns only the final {answer, model_name, agent_name, agent_id, thinking_enabled, thread_id} (no intermediate tool-call/process). POST /chat/stream is the SSE streaming variant (same request body) for long tasks — the blocking /chat produces no downstream bytes for the whole run, so a front nginx proxy_read_timeout (default 60s) cuts it off with a 504; the stream keeps the connection warm (a : keepalive SSE comment every _HEARTBEAT_SECONDS=15s via a background producer-task + queue, plus event: delta AI-text increments and event: progress), and ends with event: result carrying the same final {answer, …} JSON (or event: error). Clients only need to read event: result; intermediate events are ignorable. Both endpoints share _build_run_context (config/agent/thread). GET /agents lists {name, id} so callers know what agent_name to send. Thinking off is reliable: when thinking_enabled=false the route also sets thinking_force_disabled=true so the lead agent calls create_chat_model(force_disable_thinking=True), which injects three OpenAI-compatible "thinking off" shapes into extra_body — top-level enable_thinking=false (SiliconFlow / DashScope hybrid models like deepseek-ai/DeepSeek-V4-* on api.siliconflow.cn — the knob the nested chat_template_kwargs shape does NOT reach), chat_template_kwargs.enable_thinking=false (vLLM/Qwen native), and thinking.type="disabled" (Zhipu GLM) — even for models that only declare supports_thinking: true. For the web-UI path (which sends thinking_enabled=false but NOT force_disable), SiliconFlow models need when_thinking_disabled: {extra_body: {enable_thinking: false}} in their config.yaml entry (added to deepseek-ai/DeepSeek-V4-Flash) for thinking-off to take effect. Sets is_scheduled_run=true internally to skip the auto-title LLM call. Generated-file return (include_files=true, default on): an agent that produces its report as a file (writes to /mnt/user-data/outputs/ + present_files) returns only a short summary in answer; to recover the real deliverable both endpoints now also return `files: [{name, virtual_path, encoding(text |
Open 3qfx Ask (/api/open/3qfx/ask) |
Unauthenticated open entry that, given just a 任务id, directly kicks off the 3qfx(三情分析)「第一次问答」as a real persisted conversation (sibling of /api/open/chat; same /api/open/ auth-whitelist + CSRF-exempt). POST /ask {task_id, message?, model_name?, thinking_enabled=false} → (1) resolves the 3Q agent from app.state.business_mapping_store.get_mapping("3Q") (the same agent the frontend deep-link goPath=3qfx uses; 503 if unconfigured); (2) fetches the task detail server-side (consumer cop-task-detail) and builds the opening text identical to the frontend buildOpeningText 3qfx branch — 任务名称:…\n任务内容:…\n任务id(task_id)为{task_id} (an explicit message overrides and skips the fetch); (3) creates a fresh thread tagged metadata.{taskId, agent_id} (the exact filter useTaskThreads uses, and metadata.taskId makes it task-deeplink-shared / creator-path-resolved) and starts a background persisted run via start_run(..., enforce_access=False) (writes the checkpointer, runs under the "default" bucket); (4) returns immediately {thread_id, run_id, agent_id, status, opening_text} — the caller just needs "started", then watches the run in the 3qfx 任务工作区 conversation list. start_run gained a keyword-only enforce_access (default True); the open entry passes False so the admin-configured agent isn't rejected by the per-(missing)-login visibility check. Consumer fetch config lives in config.yaml → task_deeplink: cop_task_detail_url (full URL, appends ?taskId=), cop_task_detail_authorization (default Bearer admin), cop_task_detail_test + mock.cop_task_detail (offline fake when the upstream is unreachable — on in this deployment). Router app/gateway/routers/open_3qfx.py; tests tests/test_open_3qfx.py. |
Thread Shares (/api/thread-shares) |
Per-conversation sharing with two modes (thread_shares.mode). import: POST / {thread_id, mode} - create (or reuse) a revocable share code; GET /{code}/preview + POST /{code}/import - copy that conversation into the caller's account (fresh thread_id via app/gateway/thread_copy.py + new threads_meta row + user-data copy; dedup on metadata.shared_origin). view: a public, read-only static page ? see Public Shares below. Both modes: GET / - list my codes; DELETE /{code} - revoke. |
Public Shares (/api/public/shares) |
Unauthenticated (the /api/public/ prefix is whitelisted in auth_middleware._PUBLIC_PATH_PREFIXES). GET /{code} - read-only conversation snapshot (title + serialized messages + artifacts) for a view-mode share; messages are read straight from the latest checkpoint. GET /{code}/files/{path} - serves a file referenced by that conversation (images, Markdown, text, PDF, ?) so the shared page's links are openable; path resolution is confined to the owner's thread directory. Active content (text/html, application/xhtml+xml, image/svg+xml) is forced to download (_public_share_disposition); everything else is served inline. Missing/revoked/non-view codes all return 404. |
Roundtable Drafts (/api/roundtable-drafts) |
Strictly per-user persistence for the roundtable-planning page (replaces the old browser localStorage). GET / - list drafts (lightweight meta — incl. chain_title, the business-chain name used this session, cheaply regex-extracted from the in-memory step2 blob in _draft_to_meta/_extract_chain_title; null = no chain / 自由模式; surfaced on the 第一步「分析记录」history card + sidebar list instead of the old「第 X 步」); POST / - create; GET/PUT/DELETE /{id} - full read / partial update / delete (all owner-scoped). A draft = one full session (Step 1 intent + Step 2 roundtable), step snapshots stored as JSON in PortableLongText. Recommendation history bound to a draft: GET /{id}/recommendations - list newest-first; POST /{id}/recommendations - append one run (objective/status/rationale/picks/candidates). taskId 深链会商记录 are not here — they live in the separate /api/roundtable-task-drafts store (see below) so personal history never mixes with task records. Backed by deerflow.persistence.roundtable_drafts (tables roundtable_drafts, roundtable_recommend_history); store wired in deps.py as app.state.roundtable_draft_store. Tests: tests/test_roundtable_task_drafts.py. |
Roundtable Task Drafts (/api/roundtable-task-drafts) |
Task-scoped, unpartitioned store for taskId-deeplink roundtable 聊天记录 (无界嵌入抽屉). Keyed only by task_id (one task → many records), never partitioned by user — any authenticated caller may read/write/delete a task's records (created_by is audit-only). Physically separate from roundtable_drafts so the two never mix. GET /by-task/{task_id} - all drafts for a task, newest-first (history dropdown), enriched with the latest background-job status per draft via RoundtableJobRepository.status_map_by_drafts (each meta also carries chain_title, same _extract_chain_title regex as the personal store); GET /by-task/{task_id}/latest - newest draft (回显最新一条); POST / - create (task_id required); GET/PUT/DELETE /{id} - full read / partial update / delete (all user-agnostic). Backed by deerflow.persistence.roundtable_task_drafts (table roundtable_task_drafts, auto-created by create_all; migration 20260618_01); store wired in deps.py as app.state.roundtable_task_draft_store. Existing task_id-bearing roundtable_drafts rows are migrated over by scripts/migrate_roundtable_task_drafts.py. Tests: tests/test_roundtable_task_drafts.py, tests/test_roundtable_drafts_by_task.py. |
Roundtable Diagnostics (/api/roundtable-diagnostics) |
会商全流程诊断日志 (报错 + 关键事件). One row per event in roundtable_diagnostics: {scope (background = 后台挂起 / foreground = 前台服务端观测 / client = 前端上报), stage (step1_intent/step2_init/step2_leader/step2_seat/step2_recommend/step2_orchestration/step3_summary/step3_report/step3_dashboard/step3_structure/upstream/job), level (warning/error — **info 级运行日志不再收集**:record_diagnostic在写库前丢弃level=="info",只留报错;begin/consensus/gather/job_started 等纯进度事件即使仍调 .diag(level="info") 也成 no-op,故诊断页 = 报错日志页), event, message, detail (JSON), job_id/draft_id/task_id/user_id/agent_id/agent_name/cycle, created_at}. Reads are admin-only: GET / - filter (scope/stage/level/event/job_id/draft_id/task_id/user_id/q/since/until; scope accepts a comma-separated set so 前端「前台调用」= foreground,client) + paginate (newest-first), each item enriched with userName (user_id → 邮箱 @-prefix via get_local_provider().get_user, best-effort batched cache); GET /facets - distinct values for filter dropdowns; POST /admin/cleanup - retention purge. POST /report is NOT admin-gated — any authenticated user's frontend posts the errors the user actually hit (HTTP 502 before any SSE frame / network drop / SSE error frame) so they're collected too; stored as scope="client". Backed by deerflow.persistence.roundtable_diagnostics (table roundtable_diagnostics, auto-created by create_all; migration 20260619_01); store wired in deps.py as app.state.roundtable_diagnostic_store + a process-level default store (register_default_store) so the harness-side engine/gateway can write across the app/harness boundary. Writers (all best-effort — diagnostics never break a run): background path uses DiagnosticsRecorder (bound to job ctx) threaded through RoundtableJobExecutor → JobProgressWriter.diag(...) (engine transitions/step3 failures/empty-delivery) + InProcessRoundtableGateway (per-turn timeout/report·summary·dashboard truncation+missing); foreground path uses app/gateway/roundtable_diag.py::record_foreground_diag in multi_agent.py/intent.py/recommend.py error frames. Bug-fix shipped alongside: InProcessRoundtableGateway._run_agent_turn guards each agent run with an inactivity (idle) timeout (_await_run_with_idle_guard, ROUNDTABLE_SEAT_IDLE_TIMEOUT_SECONDS default 1200s = 20min, tuned high for unstable intranet models) — not a total-duration cap, so a slow-but-alive run that keeps streaming tokens runs as long as it needs; only a run that emits zero StreamBridge events for idle_timeout seconds straight (a true freeze: 内网模型连接池挂死 / sqlite 锁在挂载卷上挂死) is cancelled + raises (顶层把 job 标 error,而非永久冻结) + records a seat_turn_timeout diagnostic. Activity is observed via app.state.stream_bridge (the worker publishes every chunk there). Optional absolute cap ROUNDTABLE_SEAT_MAX_TOTAL_SECONDS (default 0 = off). Frontend: admin page pages/RoundtableDiagnosticsPage.tsx (route /page/strategy/admin/roundtable-diagnostics), opened from the 会商页右上角设置按钮 (admin only — handleSettingsClick in roundtable-planning/pages/RoundtablePlanningPage.tsx); admin read client strategy-components/api/roundtable-diagnostics.ts. Client error reporter roundtable-planning/api/diagnostics-report.ts (reportRoundtableError fire-and-forget + keepalive, setRoundtableErrorContext for draft/task attribution set by the page, stageForRoundtableAgent) is called at every client-observable failure: multi-agent.ts (streamMultiAgent — covers Step2 leader/seat AND Step3 report/summary/dashboard/structure — reports HTTP error + SSE {status:"error"} frame + worker event:error frame [模型调用没成功] + fetch-reject [网络不可达] + mid-stream drop [连接重置/超时], deduped via a reported flag; initMultiAgent HTTP/error-frame/incomplete/mid-stream), intent.ts (Step1 HTTP+error-frame), recommend.ts (Step2 recommend HTTP+error-frame), roundtable-jobs.ts (job api/stream). The admin page's first filter is a 后台挂起 / 前台调用 toggle (frontend = foreground,client); the stage dropdown options adapt to it (background → step2+step3+job; frontend → step1+step2+step3) and each row shows the reporting user's name + 👤. Tests: tests/test_roundtable_diagnostics.py, tests/test_roundtable_inprocess_timeout.py. Seat skill-recognition (弱模型补偿): a seat = one custom agent run whose configured skills (agent config.yaml → skills) are injected only via the system prompt's <skill_system> block (model must read_file the SKILL.md itself) — weak models (deepseek-chat 等) ignore it and never use their skills. app/gateway/roundtable_seat_skills.py::append_seat_skill_directive(agent_id, task, *, chain_enabled=False) appends an explicit Chinese directive (each configured-and-enabled skill's name + description + SKILL.md container path + "先 read_file 再按其流程取信息/执行") to the seat's per-turn task (human turn) — far stronger pull on weak models than the system prompt. Wired into both seat paths: foreground multi_agent._special_run (only role in {seat, seat_parallel}; Step3 singletons report/summary/dashboard/ingest excluded) + background InProcessRoundtableGateway.run_seat. No-op (returns task unchanged) when the seat has no configured skills. 2026-06-20: default-off + two opt-in switches (OR semantics) — was global-on, now injects only when either the seat agent's own switch (AgentConfig.seat_skill_directive, edited in the agent edit dialog agent-card.tsx + agent-detail-panel.tsx) or the per-seat switch on that agent within the selected business chain is on. The per-seat chain switch lives in the chain's seats JSON (ChainSeatSchema.seat_skill_directive, no DB column — rides the existing JSON like the per-seat model override), edited via a small ✨ icon → dropdown (Switch + description) on each agent card in BusinessChainEditorPage. It is threaded per-seat to each seat run: foreground multi_agent per-call field seat_skill_directive (orchestration looks it up per agentId from a {agentId: bool} session map) → _special_run → chain_enabled; background as the job field seat_skill_directives: dict[str,bool] (StartJobRequest → StartParams → deps factory → InProcessRoundtableGateway._seat_skill_directives) → run_seat looks up agent_id → chain_enabled. ROUNDTABLE_SEAT_SKILL_DIRECTIVE is only a master kill-switch (default allow; set 0/false to hard-disable regardless of the switches — never force-enables). Tests: tests/test_roundtable_seat_skills.py, tests/test_agents_backend.py. 2026-06-20 Step2 提速三件套:(1) 每席位推理深度覆盖 —— 业务链条 seats JSON 里逐席位可配 seat_mode(flash/thinking/pro/ultra,覆盖链条级 seatMode;编辑器席位卡右侧 ⏱Gauge 图标下拉,null=跟随链条默认)。前台经 useStep2Orchestration 三处 dispatch 的 seatModeToRequestFields(seatModes[id] ?? selectedMode) 透传;后台经 job seat_modes: dict[str,str] → StartParams → InProcessRoundtableGateway._seat_modes,run_seat 用 _SEAT_MODE_PARAMS(与前端 seatModeToRunParams 逐字对齐)解析覆盖作业级 seat_*(未命中→跟随作业级→policy 默认)。无 DB 列(ride seats JSON,但 ChainSeatSchema.seat_mode 须显式声明否则 Pydantic 剥掉)。Tests: tests/test_roundtable_seat_modes.py。(2) recommend 批内并行 —— 后端 engine._run_recommend 由 for-await 串行改 asyncio.gather(对齐前端 2026-06-13 已并行;总控一轮派多席位=判定相互独立→并发;批内用派活前 deliveries 快照、批后统一并入,语义不变)。(3) 研讨前并行『取数』两段式(业务链条 gather_first 列,逐条 opt-in,migration 20260620_02,editor「研讨前并行取数」开关)—— gather_first=True 时 run_orchestration 在分发前先 engine._run_gather_phase(全席位 asyncio.gather 跑 _GATHER_TASK「只取数不研讨」),产出 gathered: dict 经 _seat_message(gathered=...) 作为「已收集资料」块注入 recommend/chain/dag 各研讨轮席位任务(与「前序交付」块分开),把最慢的技能延迟前置并行、给(尤其 chain 串行的)研讨段提速;研讨阶段不限制再调技能。前台对称实现:useStep2Orchestration.runGatherPhase 在 runActiveOrchestration 分发前跑一次(gatherPhaseDoneRef 闸门,resetInitGate 复位),产出存 gatheredDeliveriesRef 经 peerDeliveries/初始 priorDeliveries 注入研讨席位。链条级 gatherFirst(camel)↔gather_first(snake) 经 api/chains.ts 映射;job 经 StartJobRequest.gather_first→StartParams→run_orchestration(gather_first=…)。Tests: tests/test_roundtable_gather_phase.py。(4) 削减重复广播(前端 only,多智能体 foreground 路径)—— suppress_pre_broadcast(仅 leader:抑制「总控本轮输入→各席位 thread」预广播 + 「总控给X派活」→兄弟席位的 post 广播,二者一个 flag)此前只在 DAG 派活轮开;现扩到 recommend(safetyCounter > 1,即第 2 轮起的续轮/综合)+ chain 逐席位(i > 0,第 2 个席位起)+ chain 收口轮(恒 true)。保留每种模式的首轮预广播(它为席位打底任务背景),只 suppress 重复的续轮广播(N 次读写巨大席位 thread 的纯浪费 + 弱模型模仿派活);席位的前序交付另经「席位交付广播」broadcastSeatDeliveryToOthers 送达,不受影响。后台引擎路径无此广播(_seat_message 直接拼上下文),故 #4 仅前端。2026-06-20 深链接业务链抽取规范:深链接会商(taskId+goPath,businessCode 如 rwfx→6BF)下,结构化抽取(flow-json)与方案总结报告都按所选业务链完整层级强制校正。结构化抽取实为 roundtable-structure 智能体(flow_extract.py 的 _system_prompt()=load_agent_soul(STRUCTURE_AGENT_ID),抽取→校验两段式 model.astream);其 SOUL 全局一份,旧实现把某业务层级写死进 SOUL→所有业务都按「目的→行为体」抽(bug)。修法:单份领域知识、两角色用法的 deerflow.agents.roundtable_orchestrator.extraction_spec(放 harness 层,app+harness 都可 import;build_structure_directive/build_summary_directive 按 businessCode 取;rwfx=6BF 写实「任务→目的→行为体→预期行为→驱动因素→叙事,预期行为↔驱动因素 1:1、余 1:N、label 不含步骤/阶段名、优先报告回退席位」,其余业务通用规范=强约束完整性+命名+取数且明令不套固定模板;business_code 空→""零副作用)。注入:(1) 结构化抽取 flow_extract._hierarchy_hint+_validate_messages 用 build_structure_directive 权威覆盖默认层级(前端 streamFlowExtract 传 business_code,源 ingestBusinessCode);(2) 方案总结报告 build_summary_directive——前台 multi_agent._special_run(role==summary 且 business_code) 追加、后台 engine.build_summary_prompt(business_code=…)(StartParams/StartJobRequest 透传)。仅深链接生效。种子 roundtable-structure/SOUL.md 层级顺序改 spec-driven、连接示例去 rwfx 写死(存量 live SOUL 不被 seed 覆盖,但注入规范权威可纠正)。Tests: tests/test_roundtable_extraction_spec.py。 |
Roundtable Draft Shares (/api/roundtable-draft-shares) |
Per-draft sharing for the roundtable-planning page, mirroring Thread Shares but the shared resource is a roundtable_drafts row. Two modes (roundtable_draft_shares.mode). import: POST / {draft_id, mode} - create (or reuse) a revocable code; GET /{code}/preview + POST /{code}/import - copy that draft (step1/2/3) into the caller's account under a deterministic per-(code,recipient) id (imp_<sha1> ? re-import is a no-op). view: a public, read-only result page ? see Public Roundtable Shares. Both: GET / - list my codes; DELETE /{code} - revoke. Backed by deerflow.persistence.roundtable_draft_shares (table roundtable_draft_shares, migration 20260606_01); store wired in deps.py as app.state.roundtable_draft_share_store. Tests: tests/test_roundtable_draft_shares.py. |
Public Roundtable Shares (/api/public/roundtable-shares) |
Unauthenticated (/api/public/ whitelisted). GET /{code} - read-only snapshot for a view-mode draft share: returns the Step 3??????report HTML (step3.html, a self-contained document) + title/summary/has_report. The frontend renders it in a sandboxed iframe (no allow-scripts) at #/roundtable-share/<code>. Missing/revoked/non-view codes all return 404. |
Scheduled Tasks (/api/scheduled-tasks) |
Per-user cron-like agent tasks. GET/POST / - list/create owned tasks; PATCH/DELETE /{id} - update/delete (owner); POST /{id}/{pause,resume,trigger}; GET /{id}/runs, GET /runs/{rid}, run-file preview/edit/download; GET /runs/{rid}/messages - agent messages for the run's underlying agent run (used to inspect a stuck/running run; authorized via _load_run); GET /runs/{rid}/html - generated HTML document for an html_page run; DELETE /runs/{rid} - owner-scoped, with an admin fallback so admins can delete any user's run; publish/subscribe under /{id}/{publish,unpublish,subscribe} and GET /published. Admin (cross-user): GET /admin/all - every task; PATCH/DELETE /admin/{id} - edit/delete any task; POST /admin/{id}/trigger; GET /admin/{id}/runs. Admin routes require system_role == "admin". |
HTML Page Favorites (/api/html-page-favorites) |
Bookmarked generated HTML pages, owner-scoped. GET / - list; POST / {page_id, title?, description?, tags?} - bookmark a scheduled_html_pages row (extracts a reusable style/layout summary via GLM5 at favorite time); a page can only be favorited once per user (DuplicateFavoriteError ? HTTP 409). GET/DELETE /{id} - read/delete (the create-task reference picker deletes favorites here). A favorite can be selected as a style reference when creating a new html_page task. |
Public Scheduled Tasks (/api/public/scheduled-tasks) |
Unauthenticated mirror for embedding. Adds GET /runs/{rid}/html and GET /{task_id}/latest-html for the html_page preview; the run result keeps result.html (only file paths are redacted). |
Light Apps (/api/light-apps) |
Light-app registry (????? / ????), global/shared (no user scoping; created_by/updated_by audit-only). Two kinds: window_open (launched in a new tab from the ???? card grid) and iframe (embedded under a parent sidebar menu via mount_location and dynamically injected into the sidebar). Both carry route_params ? a list of {key, paramType: login|theme, source, customValue, themeValue} descriptors the frontend resolves at launch time (login params read from localStorage.userInfo; theme params use a literal value) and appends to url as query string. GET / - list (filters app_type/mount_location/enabled_only/search); GET /{app_id} - read one (both open to any authenticated user); POST /, PUT /{app_id}, DELETE /{app_id} - admin only; iframe apps must carry a mount_location (400 otherwise); the unique name collides as HTTP 409. Backed by deerflow.persistence.light_apps (table light_apps, auto-created by create_all); store wired in deps.py as app.state.light_app_store. Tests: tests/test_light_apps.py. |
Menu Overrides (/api/menu-overrides) |
Sidebar-menu overrides (菜单管理), global/shared (no user scoping; updated_by audit-only). The static menu tree (icons/routes/adminOnly) lives in the frontend code (core/page-layout/sidebar-menu.ts); iframe light-apps are injected at render time. This router persists a per-node overlay the frontend merges onto that tree: parent_id (re-parent, moves the whole subtree; NULL = top level), sort_order (reorder among siblings), label (rename; NULL keeps code label), disabled (停用 — hides the node and its subtree). A node with no row renders exactly as code defines it, so new code menus / newly-registered iframe apps appear automatically and are immediately configurable. GET / - list all overrides (open to any authenticated user); PUT / - admin only, atomically replaces the whole set ({overrides: [...]}); rejects re-parenting that would exceed 3 levels or form a cycle (_validate_depth, HTTP 400). Backed by deerflow.persistence.menu_overrides (table menu_overrides, auto-created by create_all); store wired in deps.py as app.state.menu_override_store. Frontend: merge logic core/page-layout/menu-overrides.ts (applyMenuOverrides in use-sidebar-menu.ts), client strategy-components/api/menu-overrides.ts, admin page pages/MenuManagementPage.tsx (@dnd-kit drag-reorder/re-parent + 「移动到」parent selector + ↑↓ arrows + 改名/停用), entry in the bottom-left settings dropdown (components/page-sidebar.tsx, admin-only). Tests: tests/test_menu_overrides.py, frontend-web/src/core/page-layout/menu-overrides.test.ts. |
Compact Menu Overrides (/api/compact-menu-overrides) |
Compact/sidebar-summary left-nav overlay (简介模式菜单), global/shared. Two independent sets keyed by menu_type (A default / B); saving one type does not wipe the other. Type B rows use a B:: node_id prefix so they stay unique on databases that still use node_id as the sole PK. GET ?menu_type= lists one type (open); PUT {menu_type, overrides} admin-only replaces that type. Frontend: MenuManagementPage 简介模式 type switcher + runtime-config.js VITE_COMPACT_MENU_TYPE. Tests: tests/test_compact_menu_overrides.py. |
Sensitive Words (/api/sensitive-words) |
Sensitive-word / desensitize map (敏感词管理 / 脱敏词), global/shared (no user scoping; updated_by audit-only). Each row maps an original word → replacement, with enabled (停用 parks a word without deleting) + sort_order. The frontend ships a seed map in code (src/lib/desensitize.ts SEED_DESENSITIZE_MAP) used as the default on a fresh deployment; once an admin saves, this table is the source of truth and the frontend loads it at startup (DesensitizeWordsLoader → replaceDesensitizeMap, enabled rows only). GET / - list all words (open to any authenticated user — every page desensitizes); PUT / - admin OR the lqq account (_require_manager: system_role == "admin" or email @-prefix == "lqq"), atomically replaces the whole set ({words: [...]}, de-duped by word, last wins). Backed by deerflow.persistence.sensitive_words (table sensitive_words, PK word String(191), auto-created by create_all); store wired in deps.py as app.state.sensitive_word_store. Frontend: client strategy-components/api/sensitive-words.ts, page pages/SensitiveWordsPage.tsx (route /page/strategy/admin/sensitive-words, entry in the bottom-left settings dropdown — gated on username lqq, not admin role, present in both sidebar modes; 「导入种子词」 merges missing seed entries). Tests: tests/test_sensitive_words.py. |
Sentiment Agent Proxy (/api/sentiment-agent) |
BFF for the frontend 舆情分析 virtual agent. POST /stream accepts the AG-UI body plus apiUrl/authUsername/authPassword from runtime-config.js, forwards to that URL with TLS verify off, and byte-pipes upstream SSE. Connection/HTTP failures return the full error in detail; mid-stream failures emit AG-UI RUN_ERROR. Distinct from skill-packages/sentiment-analysis.skill. Tests: tests/test_sentiment_agent_proxy.py. |
Dashboard Sessions (/api/dashboard-sessions) |
「大屏绘制智能体」独立页面会话,按外部 task_id 存取,不按 user 分权(open read + shared writes;user_id audit-only)。一个 task 一行(upsert):transcript(聊天记录 JSON)+ report_json(最新结构化数据,可解析 = 「有效结果」)。GET /by-task/{task_id} - 取该 task 会话(无则 404);PUT /by-task/{task_id} - upsert({transcript?, report_json?},只覆盖传入项)。Backed by deerflow.persistence.dashboard_sessions(table dashboard_sessions, auto-created by create_all,无迁移);store wired in deps.py as app.state.dashboard_session_store。前端独立页 roundtable-planning/pages/DashboardAgentPage.tsx(路由 /page/strategy/dashboard-agent?taskId=,侧边栏「问答管理 → 大屏绘制」)让智能体分析任务报告 JSON → report-json → 渲染 DashboardApp。Tests: tests/test_dashboard_sessions.py。 |
Task Reports (/api/task-reports) |
任务级研判情报报告详情,按 task_id 共享、不按 user 分权(写入者只留 updated_by 审计)。GET /by-task/{task_id} 始终返回原 Consumer 兼容外壳 {state:"200", msg:"操作成功!", data:[...]}(未保存为 data: []);POST /by-task/{task_id} 首次保存,已有报告返回 409;PUT /by-task/{task_id} 原子替换完整 data 数组。POST /import/by-task/{task_id} 面向 TaskCOP 导入页,接收 sq-report-mock.json 兼容的 data/records 数组并完整覆盖目标任务;路径中的选定任务 id 优先于文件内的 taskId,避免误写。每一项严格是 {id, taskId, createTime, updateTime, sessionId, contentJson, categoryType};contentJson 以 JSON 字符串原样落在 records_json,避免破坏大屏按类目解析的输入格式。保存或导入非空 data 后,router 会通过 app.state.taskcop_task_store.mark_analysis_completed 将同 id 的 TaskCOP 任务升级至状态 25,使旧任务列表显示并筛选为“已完成”;普通保存时不存在匹配任务的报告仍可独立保存,导入则要求目标任务存在。Backed by deerflow.persistence.task_reports(table task_reports, auto-created by create_all);store wired in deps.py as app.state.task_report_store。前端客户端 strategy-components/api/task-reports.ts,core/auth/dashboard-input.ts 与 rwfx 3q 完成检查均直接读此接口,不再依赖 Consumer 或本地 sq-report mock。Tests: tests/test_task_reports.py。 |
Task Buttons (/api/task-buttons) |
Task-workspace configurable jump buttons (按钮管理), global/shared (no user scoping; updated_by audit-only). Each row = one jump button in a task deep-link workspace's bottom 快捷跳转 bar: {id, business (rwfx/3qfx/xdfx), label, link_type (url/business/purpose/host), target, append_task_id, id_param (editable query-param name, e.g. id/task_id), id_kind (task/action — only xdfx differs, its deep-link id is an 行动id), open_mode (blank/self), login_params, append_auth, auth_token_param, auth_name_param, enabled, sort_order}. Two independent, coexisting login-append mechanisms for url buttons — both resolved frontend-side: (1) login_params (JSON column login_params_json, PortableJSON, migration 20260625_01) = a list of {key, source, custom_value} resolved at click time (source = a userInfo field or other→literal custom_value) and appended to the target URL — mirrors the light-app login params; empty = nothing appended. (2) append_auth + auth_token_param + auth_name_param (columns, migration 20260625_02, default off) — appends ?{auth_token_param}={token1}&{auth_name_param}={username} so the target system can identify the current user; gated on token login — the frontend only stashes token1 (the ?authToken=) + the raw getTokenInfo username (login.token1/login.username) when login went through /login/token, so a non-token-login session appends neither. The frontend holds code-seed defaults (rwfx's 4 historical buttons built from the live task_deeplink config); only when this table is empty do the workspace + admin page fall back to them, else this table is the sole source of truth. GET / - list all (open to any authenticated user — the workspace renders them); PUT / - admin only (_require_admin), atomically replaces the whole set ({buttons: [...]}, de-duped by id, last wins). Backed by deerflow.persistence.task_buttons (table task_buttons, auto-created by create_all); store wired in deps.py as app.state.task_button_store. Frontend: helpers core/auth/task-buttons.ts, client strategy-components/api/task-buttons.ts, admin page pages/TaskButtonsPage.tsx (route /page/strategy/admin/task-buttons, entry「按钮管理」in the bottom-left settings dropdown, admin-only, both sidebar modes; per-business tabs + 「导入默认」). Tests: tests/test_task_buttons.py. |
Fixed Questions (/api/fixed-questions) |
Administrator-managed exact-match fixed Q&A (问题配置), global/shared. Each row is {id, question, answer, enabled, tokens_per_second, sort_order}; stream speed is configurable per row from 1–1000 token/s and defaults to 100. GET / and whole-set atomic PUT / are both admin-only. At the start of an external chat run, services._find_fixed_answer checks the latest eligible plain-text human message with exact, case/space/punctuation-sensitive equality. Enabled matches switch the run to make_fixed_answer_agent: a local BaseChatModel tokenizes with cl100k_base, replays the answer in rate-limited token chunks, and keeps the existing LangGraph SSE messages/values protocol, run journal, checkpoint history, deterministic local title, and refresh behavior intact while no provider/model request or model-concurrency slot is used. File/multimodal/context-prefixed turns, canvas_agent, ai_writing, writing mode, resume commands, and internal orchestration/background child runs bypass matching. Storage failures fail open to the normal model path. Backed by deerflow.persistence.fixed_questions (table fixed_questions, SHA-256 indexed lookup plus exact original-string collision check, migrations 20260821_01 + 20260821_02), wired as app.state.fixed_question_store. Frontend: strategy-components/api/fixed-questions.ts, admin page pages/FixedQuestionsPage.tsx, route /page/strategy/admin/fixed-questions, entries in both full-layout and compact-workspace settings menus. Tests: tests/test_fixed_questions.py. |
Knowledge (/api/knowledge) |
Product knowledge base, global/shared (no user isolation; writes record created_by/updated_by for audit). POST /ingest/thread/{thread_id} - sediment one thread into a note (reads messages from the checkpointer, threads:read owner-checked); POST /ingest/threads/batch - batch-sediment the caller's recent threads (skip_existing via .manifest.json dedup). Defaults to background=true: the slow per-thread extraction loop is deferred to a FastAPI background task and the endpoint returns {queued:true, processed:N} immediately (each thread's transcript is truncated to max_input_chars before the LLM call). Running the loop inline for a large batch outlives the nginx gateway timeout (504), which the frontend's parseResponse surfaces as a generic "Request failed"; background=false keeps the blocking path (full per-thread created/skipped/failed tally) for small batches/scripts. Shared loop: _run_batch_ingest. Optional model selection: both ingest endpoints accept model_name (overrides knowledge.extract_model_name for that capture; threaded through KnowledgeService.capture_thread ? extract_thread_llm_multi). Model-error surfacing: when the chosen model produces nothing usable (timeout / API error / non-JSON) capture still falls back to rule-based extraction and lands a draft (llm_fallback:true); the single-thread (non-background) response and the batch response (llm_fallback count) carry this so the UI tells the user "???????". The "??????" chat button stays background/fire-and-forget (passes model_name, no per-save error toast); the batch-import dialog is synchronous and shows the fallback count. POST /ingest/file - import an uploaded file (multipart: file + optional title/tags/status/model_name) into wiki note(s): the file is converted to Markdown (md/txt read directly, pdf/word/ppt/excel via convert_file_to_markdown), distilled by the same wiki-capture LLM pipeline as thread capture (source_type="file"), and falls back to a verbatim draft note when the model returns nothing usable (llm_fallback:true); POST /ingest/search - sediment search/tool results as a reference note; POST /notes - manual note; GET /notes (keyword/tag/source_type/status + pagination); GET/PUT/DELETE /notes/{id} (DELETE is soft = archive, ?hard=true removes the vault file); POST /search - mode=keyword|vector|hybrid (vector?keyword fallback when no embedding endpoint); GET /notes/{id}/versions - edit history; GET /duplicates + POST /merge - dedup; POST /reindex - rebuild vector index; GET /graph - product graph data (notes + entities, graph-blacklist filtered); GET/POST /graph/blacklist + DELETE /graph/blacklist/{id} - graph node blacklist (labels = entity names / note titles, case-insensitive match; blacklisted nodes and their edges are dropped from /graph and the wiki-export artifacts; table knowledge_graph_blacklist, auto-created); POST /export/graph + GET /export/files/{graph.json|graph.html|graph.graphml|cypher.txt} - obsidian-wiki wiki-export artifacts (self-contained graph.html for iframe); POST /resolve {targets:[?]} + GET /notes/{id}/wikilinks - resolve [[wikilink]] targets to a note id / vault page / missing (powers clickable wikilinks); GET /vault-page?path= - read a raw vault meta/hub Markdown page (no DB note). GET/POST /folders + PUT/DELETE /folders/{id} - user-managed directories (notes carry a folder path; GET /notes?folder= filters a dir + sub-dirs); GET/POST /extract-templates + GET/PUT/DELETE /extract-templates/{id} - editable extraction-prompt templates (thread/document ingest accept template_id; capture/PUT-note accept folder). Backed by the deerflow.knowledge module (see below). |
Proxied through nginx: /api/langgraph/* ? LangGraph, all other /api/* ? Gateway.
TaskCOP Situation-Overview Compatibility
The standalone legacy Vue page calls the Gateway URL in its runtime b1ConsumerUrl configuration directly (no Vite /api reverse proxy). The Gateway CORS middleware merges standard local Vite origins (localhost/127.0.0.1 on ports 5173, 5174, 3000, and 8080) even when GATEWAY_CORS_ORIGINS was set for another frontend. This avoids a restart regression when the existing Vite server is on 5173 but the launcher sets 5174. Deployments must set GATEWAY_CORS_ORIGINS to any additional exact frontend origins that are allowed to call the Gateway.
app/gateway/routers/taskcop_tasks.py is a temporary compatibility layer for the separately deployed legacy Vue page at /situation-overview/action/sentiment and its task import page at /situation-overview/action/task/import. It preserves the old Consumer paths and response envelopes: POST /cop/saveSuperiorTask creates a row with a sequential four-digit string id (0001, 0002, ...) and returns {state: 201, msg, data}; POST /api/taskcop/import/tasks accepts a JSON array or tasks/records/legacy list envelope, uses each supplied id or taskId as the primary key, and completely replaces duplicate IDs; DELETE /cop/tasks/{task_id} removes that task and its shared report document; GET /taskAnalyseSearch/cop-task-three-list (all), GET /taskAnalyseSearch/cop-my-tasks (creator-scoped), and GET /taskAnalyseSearch/cop-task-list return {state: 200, data: {total, records}}; GET /taskAnalyseSearch/cop-task-detail?taskId= returns one task. Existing UUID-style task rows are retained, and do not affect the new numeric sequence. An empty TaskCOP store remains empty—there are no built-in task or report seed records. It supports the page's existing content, taskStatus, taskDirection, startTime, endTime, pageNum, and pageSize query fields. Tasks are stored in taskcop_tasks through deerflow.persistence.taskcop_tasks, registered by persistence.models, and initialized as app.state.taskcop_task_store in deps.py. A non-empty save/update/import through /api/task-reports/* calls mark_analysis_completed(task_id): it changes only pre-completion task states to 25, so the list renders and filters the task as “已完成” without downgrading later workflow states. trendSucc is derived from the same status for the legacy sentiment list. Zero-padded task ids stay strings in the report API, so 0001 is never converted to 1. The Vue router calls DeerFlow's /api/v1/auth/login/username directly with an incoming username/userId/yUserId parameter; it does not depend on the retired Consumer login service, and sends the returned DeerFlow JWT to every TaskCOP call. The TaskCOP routes are no longer public or CSRF-exempt. The Gateway fallback CORS allow-list includes TaskCOP Vite origins http://localhost:8080 and http://127.0.0.1:8080; deployments must use explicit GATEWAY_CORS_ORIGINS. Tests: tests/test_taskcop_tasks.py and tests/test_taskcop_cors.py.
AI Writing Intent Router (对话驱动写作台意图路由)
Backs the frontend「AI 写作·对话版」page's unified bottom composer. POST /api/ai-writing/sessions/{session_id}/intent (owner-checked; in routers/ai_writing.py) reads the session's graph checkpoint via the shared app.state.checkpointer — the pending interrupt is the authoritative pause point (the frontend's claimed pause_point is only a fallback when the checkpoint can't be read) — then routes the user's free text through deerflow/agents/ai_writing/intent_router.py::parse_user_intent: an LLM call (session's model_name + fallback_chain model fault-tolerance) that returns {intent, pause_point, action, payload, need_research, research_query, answer, confidence, clarify}. Safety boundaries in normalize_intent_result: per-pause action whitelist strictly aligned to graph.py's real resume actions (ALLOWED_ACTIONS), no pause → no action, high-cost actions (finalize / force_finalize / query-less覆盖式 re_search) require confidence ≥ 0.85 else downgrade to clarify, payload key-whitelist cleaning (material-id filtering, review-issue index bounds, section-decision back-fill with the majority decision). QA intents return answer directly from the context snapshot (build_context_snapshot — stage-tailored: materials with ids / outline / draft body capped / review issues with indices). Model failures never 5xx — they degrade to clarify. Prompts in agents/ai_writing/prompts/intent_router_prompts.py; tests tests/test_ai_writing_intent.py.
AI Writing Built-in Functional Agents (????????)
The AI-writing graph (packages/harness/deerflow/agents/ai_writing/) has four stage "brains" ? ?????? / ????? / ?? / ?? ? that can each run in one of two modes, switched per-stage by config.yaml ? ai_writing.use_builtin_agents.{researcher,outliner,writer,editor} (default false, mtime hot-reloaded, new sessions pick up changes immediately):
- Legacy (flag off): hard-coded prompts in
prompts/*.py, direct LLM call (unchanged behavior). - Builtin-agent (flag on): the node loads the corresponding functional agent from
.deer-flow/agents/ai-writing-<stage>/?SOUL.mdbecomes the system prompt andconfig.yaml'smodeloverrides the stage model (priority:agent.model > node explicit fallback (keyword_model/rank_model) > state.model_name). Admins edit SOUL/model from the agent management page. Runner:nodes/_agent_runner.py(run_writing_agent/stream_writing_agent/parse_marker_json). Output contracts (marker +jsonblocks, intent.py-style): researcher[KEYWORDS_READY]/[MATERIALS_RANKED], outliner[OUTLINE_READY], editor (per-dimension)[REVIEW_READY]; writer emits raw section Markdown (keeps per-section streaming) with the legacy[[SECTION_BLOCKED]]sentinel in strict mode. Orchestration stays in the nodes (search fan-out, per-section loop, 3-dimension review loop, all pause points) ? the agent only does single structured generations, and a missing/empty SOUL silently falls back to the legacy path (WritingAgentUnavailable).
Seeding mirrors the roundtable pattern: verified config.yaml + SOUL.md ship in app/gateway/routers/_ai_writing_seed_assets/<id>/; _ai_writing_seed.py::ensure_ai_writing_functional_agents() copies them to .deer-flow/agents/ when missing (never overwrites local edits), called from the startup lifespan (before _sync_legacy_agents, so the agents appear in /api/agents on first boot) and self-heals from GET /api/ai-writing/sessions + GET /api/ai-writing/article-types. Tests: tests/test_ai_writing_seed.py, tests/test_ai_writing_agent_runner.py.
Sample imitation mode (样文仿写). The initial writing form supports writing_mode_type=imitate: the user can paste sample text or upload Word/PDF/Markdown/TXT. POST /api/ai-writing/sample/extract converts supported document files to Markdown via deerflow.utils.file_conversion and returns capped plain text to the frontend. The LangGraph entry point now starts at sample_analyzer before intent_parser; non-imitation sessions return immediately, while imitation sessions analyze up to 12k chars into sample_style_profile (SampleStyleProfile in state.py) and emit sample_profile_ready. writer_outline and writer_draft call render_sample_style_section(state) so both outline planning and section drafting receive the same style profile, imitation strength (light|medium|strong), optional structure-preservation preference, and guardrails: learn structure/tone/paragraph rhythm/sentence style, but do not copy sample sentences, proprietary facts, people, organizations, cases, or data. Frontend state mirrors this with analyzing_sample, uploaded/pasted sample controls, history/timeline restoration of sampleStyleProfile, and the setup-chat direct-start path preserves the same fields. Tests live with the knowledge-base AI-writing tests (tests/test_ai_writing_knowledge_base.py).
Skill-driven retrieval (检索类型). material_source has three values: general (通用检索, default), knowledge_base (知识库检索), and notebook (我的空间); legacy skill is normalized to general by the researcher node. knowledge_base 直连内网 ES 接口(= 旧版「内网直连检索」) — it bypasses skill orchestration entirely and calls IntranetSearchProvider over ai_writing.intranet_search_url (concurrent over the generated keywords; intranet_knowledge_base/intranet_search_timeout apply). An empty/unconfigured intranet_search_url does NOT disable it — the empty string is passed straight to IntranetSearchProvider, which falls back to its built-in default endpoint (intranet_stub._DEFAULT_ENDPOINT), exactly like the legacy get_search_provider(search_provider=intranet) path. So 内网直连检索 works out of the box without configuring an address (configure intranet_search_url only to point at a different endpoint). This is the "弱模型识别不了技能" escape hatch — the old no-skill retrieval path restored as an explicit option (weak models like deepseek-chat often fail to invoke configured skills). Tests: tests/test_ai_writing_knowledge_base.py. General retrieval is orchestrated from the skills configured on the ai-writing-researcher agent (config.yaml → skills, editable from the agent management page): the built-in knowledge-base-search skill (知识库检索) runs the native keyword fan-out web/intranet search (the old default path), while any other configured skill searches with a description-targeted query; skills run in configured order with early-stop at max_materials, and an empty/unresolvable skills config falls back to the built-in knowledge search so retrieval always works. The skill itself is seeded from _ai_writing_seed_assets/skills/knowledge-base-search/SKILL.md into skills/public/ (never overwrites), and a pre-existing researcher config.yaml lacking the skills key is patched in place with the default (explicit skills: [] is respected). PAUSE-1 input append flow: when the user types natural language into the material-confirm input (re_search/legacy enable_skill_search + user_query), the node matches the most relevant configured skill via SkillPicker(candidates=…) (skipped when only one skill is configured), searches with the user's words, and appends the new materials to the existing package (URL-dedup, original order preserved, no LLM re-rank/truncation, 100-item hard cap). Skill-failure fault tolerance: a configured skill that times out / raises / has all its keyword searches fail (vs. legitimately returning 0 results) no longer just silently moves on — the researcher node emits a skill_search_step of type skill_error (carrying the error reason), and if the configured skills collected no new materials at all it runs a fallback pass over other available skills (_fallback_candidate_skills: the built-in knowledge-base-search native path first, then other enabled catalog skills not yet tried) instead of erroring out. The frontend SkillStepsCard renders skill_error (rose) and fallback (amber) so the timeline shows "技能报错 → 改用其他技能搜集". Tests: tests/test_ai_writing_material_source.py.
The intranet client uses trust_env=False and defaults intranet_verify_ssl=false, so internal HTTPS endpoints with self-signed/internal-CA certificates do not fail on certificate verification; set ai_writing.intranet_verify_ssl: true only for environments that require strict certificate checks.
Skill-call timeout is hardcoded 300s (5 min) — researcher.py _SKILL_SEARCH_TIMEOUT_SECONDS (per-skill wrapper) and _skill_agent.py _REAL_AGENT_TIMEOUT_SECONDS (whole real-agent skill run). The intranet ES fallback (IntranetSearchProvider) gets a longer 600s (10 min) default (intranet_stub.py; overridable via ai_writing.intranet_search_timeout) — that interface is the slow one, so our own fallback timeout is generous to avoid being blamed for "检不出".
Parallel intranet fallback (技能 0 条 → 直接用并行预跑的内网结果). When ai_writing.intranet_search_url is configured, the researcher node launches the intranet ES search (IntranetSearchProvider over the generated keywords) as a concurrent asyncio.Task the moment the skill loop starts (not sequentially after all skills fail). If the configured + fallback skills collect nothing, the node awaits that already-running task and uses its results (_collect) — no extra wait. If the skills did get materials, the task is cancelled (_cancel_intranet_task, suppresses CancelledError) so the intranet call isn't wasted. URL empty → no task, behavior unchanged. Tests: tests/test_ai_writing_intranet_parallel.py.
Partial-result harvest + query simplification (检不出/超时的容错). When researcher_skill_agent: true runs a non-native skill via run_skill_via_real_agent (the deployment's real path — the skill, e.g. knowledge-search-v2, is invoked by the real agent through bash python3 scripts/search.py "<query>"), one writing run issues many queries and some time out (the skill script prints {"success":false,"error":"请求超时…","count":0}) while others succeed ({"success":true,"count":30,"results":[…]}). Two fixes so successful data is never dropped: (1) _drive() captures every ToolMessage stdout and _harvest_tool_results() parses the skill's JSON by reusing the lead-agent 参考文献 mode's proven extractor (deerflow.runtime.references: _parse_json_like + _first_list + _normalize_reference_item) — the same logic that already extracts skill JSON into citations 100% reliably on ChatPage/AgentChatPage, with its huge field-name union (title/m_title/name, content/page_content/content_preview/snippet/…, recUuid/uuid/docId/…) and per-skill display-config support — plus skips success:false outputs — so even if the model's final {materials:[]} summary gives up (confused by the timeouts), the successful queries' real results are still collected; the function merges model materials + harvested results (dedup by url/recUuid/title, capped at max_results). The harvest+merge (_finalize) runs on EVERY exit path — normal end, overall timeout, AND await task exceptions (e.g. GraphRecursionError) — so when the run is cut off before the model writes its summary (observed: 10 queries × ~2 graph steps > the old recursion_limit, so the agent errored out with the real 30+ results already in tool_outputs), those results are still salvaged instead of returning []. recursion_limit was also raised from _REAL_AGENT_MAX_TURNS to _REAL_AGENT_MAX_TURNS*2+4 (it counts graph super-steps ≈ 2 per model→tool round, not model turns) so the model normally has room to finish its summary. Early-stop on max_results: _drive re-harvests after every tool output and, once the deduped harvested count reaches max_results (= per_skill = max_materials), breaks the nested astream and emits a 「已获取 N 条素材,达到上限,停止本技能检索」 step — without it a weak model keeps issuing new queries forever (observed: 100+ step-bar entries, materials far over the limit). Because the skill now returns exactly max_materials, the researcher loop's own collect_limit early-stop fires right after, so later configured skills don't run either. (2) The model is instructed (in _MATERIAL_OUTPUT_CONTRACT) that a 0-result/timeout query must be retried with a simpler, core query (drop topic suffixes like 竞选/诉讼/政策, keep the core entity: 「特朗普竞选」→「特朗普」) and that any data already retrieved must always be kept. The same deterministic simplifier search/base.py::simplify_query is wired into IntranetSearchProvider.search: an empty first attempt retries once with the simplified query. (3) Root cause of the instability — bash output middle-truncation. The skill prints its results as JSON to stdout, and the bash tool middle-truncates any output over sandbox.bash_output_max_chars (default 20000, inserting a ... [middle truncated …] ... marker). A count:30 skill response (~30 records × ~600 chars) exceeds 20000 → the JSON is cut in the middle → unparseable as a whole by both the model and _try_load_json → all materials lost (while a small count:9 response parses fine — hence the flapping). Fix: when whole-JSON parse yields no items, _harvest_tool_results falls back to _scan_balanced_json_objects — a brace-depth/string-aware scanner that recovers every individually-complete {…} record object from the truncated text (the records on each side of the truncation marker are still valid JSON), so ~16+16 records survive a 20000-char cut, far above max_materials. This is why the lead-agent chat pages (/page/workspace/chats/new, agent chat) call the same skill with 100% success — they only need the model to summarize prose from whatever it sees, never to parse a structured list, so truncation doesn't hurt them; ai-writing needs the structured records, so it must salvage them itself. Tests: tests/test_ai_writing_skill_real_agent.py, tests/test_ai_writing_query_simplify.py.
Agentic skill execution (ai_writing.researcher_skill_agent). When this flag is on (the deployment config.yaml enables it), the directed-skill branch runs nodes/_skill_agent.py::run_skill_via_real_agent: it reuses SubagentExecutor with the full tool stack, sandbox middleware, and a one-skill whitelist so the skill can read_file its SKILL.md, call bash, MCP tools, or whatever the skill instructions require. The nested run now passes agent_id=ai-writing-researcher in both config.configurable and runtime context, matching normal AgentChat skill visibility/whitelist behavior. It also appends a human-turn hard directive listing the exact SKILL.md path and skill directory, requiring weak models to first read_file the SKILL.md and then cd <skill-dir> && python/python3 scripts/*.py ... when the skill is script-backed; the model is explicitly forbidden from returning materials before executing the skill's required tools. If the real-agent path fails or yields no materials, _execute_skill still falls back to the legacy description→query search so materials are never empty. The built-in knowledge-base-search native concurrent path is unchanged. Tests: tests/test_ai_writing_skill_real_agent.py.
Targeted user revision (局部修改 — 只改命中章节). When the user submits a free-text 修改意见 on the draft-confirm card (action == "user_revise"), writer_draft_node no longer always rewrites the whole article section-by-section. Before the full-rewrite loop it tries _targeted_user_revision (nodes/writer_draft.py): _split_draft_sections parses the previous current_draft.full_markdown into (article_title, [(section_title, body)], references_block) (level-1 # title, level-2 ## sections, trailing ## 参考文献 pulled out and kept verbatim), then _detect_revision_scope runs a small llm_json classifier (prompts DRAFT_REVISE_SCOPE_{SYSTEM,USER}) returning per-section targets [{index, new_title, rewrite_body}]: it maps the note to the section indices it targets ("第一段/开头/引言"→0, "结尾/结论"→last, by title/topic) and distinguishes a 章节标题改名 (new_title set, rewrite_body=false → rename only, body untouched) from a 正文修改 (rewrite_body=true). Title-only renames apply directly to the assembled ## {new_title} heading and are written back into current_outline.sections[i].section_title (so the renamed title persists through confirm/finalize/history — this fixes "改章节标题不生效"). Body rewrites run _stream_revise_section (prompt DRAFT_REVISE_SECTION_USER, streamed as draft_chunk). Only hit sections are touched; unhit sections are preserved byte-for-byte and are NOT re-emitted to the stream — the frontend clears the right canvas to its initial placeholder on entering revising and shows only the live rewrite of the changed section(s), so the user sees「清空 → 流式重写 → 终稿」rather than the whole article rebuilding. The classifier returns whole_article=true (整体性意见 like 「整体语气更正式」/「全文都改」, structure changes 加章/删章/调整顺序, or a 文章总标题 change) → _detect_revision_scope yields None → _targeted_user_revision returns None → the node falls back to the whole-article rewrite (which still runs _revise_outline_with_notes for title/structure changes). Also falls back when there is no prior draft, an empty note, or scope-detection errors. Editor-review revisions (accept_review) keep the holistic rewrite. Reassembly preserves the article title + structure and re-appends the original references block; word count is computed on the body (excluding references). Frontend: isRevising in AIWritingState (types/ai-writing.ts) is set in useAIWritingStream.ts's intervention_submitted (ns === 'revising') and cleared on draft_ready; AIWritingDraftPanel.tsx shows the initial DraftWaitingPlaceholder while isRevising && streamingDraftSections.length === 0, then the streaming preview once chunks arrive. Tests: tests/test_ai_writing_targeted_revision.py.
History draft editing (历史回看可编辑成稿). The right-side draft in 历史回看 is no longer read-only — the user can edit it like the live 写作完成态 (Tiptap + 目录 + 格式工具栏 + AI 润色) and persist the change back to the session. PUT /api/ai-writing/sessions/{id}/draft {draft_markdown, draft_title?} (owner-checked → 404 otherwise; draft_title omitted ⇒ keep原标题) updates the row via the existing AIWritingSessionRepository.update. Frontend: AIWritingDraftPanel now passes readOnly={false} + an onSaveDraft handler (only in history mode) → TextArtifactRenderer → WritingToolbar 「保存」button (loading + toast); the handler is AIWritingHistoryContext.saveHistoryDraft(markdown, title?) → updateAIWritingDraft (api/ai-writing-sessions.ts), which write-backs the returned summary to historySession. Edits to the live completed draft stay ephemeral (export-only) as before — only history mode shows 保存. Tests: tests/test_ai_writing_draft_update.py.
Background suspend (后台挂起). A writing session can be handed off to the server to auto-drive the remaining flow to finalize, surviving page leave / refresh / account switch. POST /api/ai-writing/sessions/{id}/background (owner-checked, idempotent) starts AIWritingJobExecutor (app/gateway/ai_writing_job_executor.py), wired in deps.py as app.state.ai_writing_job_executor. The executor drives the same ai_writing graph in-process — it builds make_ai_writing_graph(), attaches the shared app.state.checkpointer (the very saver the runtime worker uses, so it continues the real thread state the frontend already wrote), sets the owning user via set_current_user, and loops aget_state → ainvoke(Command(resume=…)) until status==done. Auto-action per interrupt: material/outline confirm → confirm; section_help → every blocked section loose (continue with general knowledge, never re-loop to research); draft_confirm → to_editor until reviewed, then finalize (on pass, or when revision_count >= max_revisions); review_confirm → accept_review (revise & re-review until pass or budget exhausted, then forced finalize). To avoid racing a still-running runtime worker (suspend mid-node), the driver waits for an actionable state (_wait_for_actionable): it only acts when the graph is at an interrupt / done, or when the checkpoint stops advancing (no worker → abandoned/restart). No new table — the graph checkpoint + ai_writing_sessions already hold all state; status="background" marks 挂起中 (surfaced in the history dropdown), nodes write draft/review_result/status=done as usual, so the user reopens the finished article from history. Per-user concurrency cap AI_WRITING_BACKGROUND_SLOTS_PER_USER (default 2, same SQLite-lock rationale as the roundtable executor). Startup reconcile (in deps.py) relaunches drivers for sessions left background after a restart (AIWritingSessionRepository.list_by_status). Progress flowchart popup: GET /api/ai-writing/sessions/{id}/progress (owner-checked) reads the graph checkpoint via executor.get_progress and derives a canonical pipeline stage key (intent/research/material_confirm/outline/outline_confirm/draft/draft_confirm/review/done, from the pending interrupt's pause_point or snapshot.next node) + note (素材不足求助 / 编辑打回) + review progress + isRunning. Frontend: 「后台挂起」 button + auto-suspend on route unmount / beforeunload (AIWritingPanel.tsx, suspendAIWritingToBackground[Keepalive] in api/ai-writing-sessions.ts); 挂起中 badge in AIWritingHistoryDropdown.tsx, and clicking a 挂起中 history record opens a flowchart dialog (AIWritingFlowDialog.tsx polls getAIWritingProgress every 2.5s, renders the fixed vertical pipeline AIWritingFlowChart.tsx lit by the live stage) so the user sees how far the background run has gotten without reopening/taking over the session. Transcript consistency: the rich history timeline (AIWritingHistoryView → AIWritingTimeline) replays transcript = {progressEvents, completedInterventions}, normally assembled+saved client-side. A background run has no client, so the executor rebuilds it at the end (_save_transcript/_build_transcript_from_checkpoint) entirely from the final checkpoint — both progressEvents and completedInterventions are derived from the same set of resumed events (every pause-resume emits exactly one), so they are always 1:1 aligned. For each resumed: classify its message → (pausePoint, action) (_classify_resumed, matched against graph.py's resume messages), insert an await_user (with the pause prompt) before it, and build a rich intervention card enriched from final state (material_confirm→materials, outline_confirm→outline, draft_confirm→draftTitle/draftWordCount/reviewResult, review_confirm→reviewResult); materials_ready is inserted before the first material-confirm. (Earlier code merged the frontend's pre-suspend completedInterventions with auto ones and aligned by index — but the frontend's last save can predate its most recent decision, so the counts drifted and cards rendered off-by-one with empty data; deriving everything from the checkpoint root-fixes both.) Raw (snake_case) events and intervention data are normalized on load by the exported normalizeProgressDict / normalize{Material,Outline,ReviewResult} (idempotent on already-camelCase live transcripts — one normalize path for both sources), so reopening a background-completed session shows the same cards as a frontend-driven one. Tests: tests/test_ai_writing_background.py.
Admin Leaderboard Snapshots (/api/admin/users/leaderboard*)
The admin ????? is served from precomputed per-day snapshots in admin_leaderboard_daily_stats (keyed by (stat_date, settings_hash)), not by aggregating raw rows on every request. A background scheduler inside the Gateway (_leaderboard_snapshot_scheduler_loop in app/gateway/app.py, singleton file-locked) ticks every ADMIN_LEADERBOARD_BACKGROUND_INTERVAL_SECONDS (default 300s): it prewarms recent days, then builds any queued/error/stale-running day. GET /leaderboard merges the per-day snapshots for the requested range. Today is rebuilt whenever its snapshot is older than ADMIN_LEADERBOARD_TODAY_TTL_SECONDS (default 43200s = half a day, so the leaderboard refreshes ~twice daily); a past day, once ready, is treated as fresh forever.
Two on-demand/automatic refresh paths layer on top:
- Manual "????" ?
POST /leaderboard/refresh-todayforce-rebuilds today's snapshot synchronously (bounded by the same semaphore/timeout the scheduler uses), then returns the merged range. Lets an admin pull the latest real-time records into the stats without waiting for the TTL. Frontend: a "????" button onAdminUserLeaderboardPage?useRefreshTodayLeaderboard(core/admin/hooks.ts), which primes + invalidates the["admin-activity","leaderboard"]query. - Daily finalize of the previous day ? once per day, on/after
ADMIN_LEADERBOARD_FINALIZE_HOUR(local Beijing, default1= ?? 1 ?), the scheduler tick force-recomputes yesterday a single time so runs/tool-calls that only landed near midnight are fully captured. The decision is the puredeerflow.persistence.admin_stats.finalize.should_finalize_previous_day(status/generated_at/now/hour) ? the rebuild stampsgenerated_atto today, which self-debounces it for the rest of the day (no extra state file). Tests:tests/test_leaderboard_finalize.py.
Admin Concurrency Monitor (实时并发监控 / 系统压力)
A live admin monitor embedded at the top of AdminUserLeaderboardPage (components/workspace/admin/concurrency-monitor.tsx, ConcurrencyMonitorPanel) that answers "当前有几个并发的大模型调用、是谁、能不能直接停掉" and charts system pressure over time. Built on the existing in-memory RunManager (no new run plumbing).
- Live concurrency + 真正停止 —
RunManager.list_active()returns every run whose status ispending/running(the registry keeps finished records around untilcleanup, so it filters by status).len(list_active())is the current concurrent-call count. Routerapp/gateway/routers/admin_active_runs.py(/api/admin/active-runs, admin only via the sharedadmin_users._require_admin):GET /returns{count, server_time, runs:[{run_id, thread_id, status, created_at, elapsed_seconds, user_id, user_name, agent_id, agent_name, thread_title, kind, multitask_strategy}]}— each run best-effort enriched (batched DB) with the owning user's email (threads_meta.user_id→users.email) and the agent's display name (agents.name);kindis derived from thread metadata (圆桌会商 / 系统/后台 / 任务对话 / 对话).POST /{run_id}/cancelcallsRunManager.cancel(run_id, action="interrupt"), which sets the abort event and cancels the run's asyncio task — i.e. it truly aborts the in-flight model stream, freeing capacity.POST /cancel-allstops every in-flight run. - 并发量折线图 (按分钟/小时) — a lightweight background sampler (
_concurrency_sampler_loopinapp/gateway/app.py, singleton file-locked, started in the lifespan) recordslen(list_active())everyCONCURRENCY_SAMPLE_INTERVAL_SECONDS(default 15s) into the newconcurrency_samplestable (deerflow.persistence.concurrency,ConcurrencySampleRow{sampled_at, active_count}, auto-created bycreate_all, no migration; store wired asapp.state.concurrency_sample_store). ~Hourly it purges rows older thanCONCURRENCY_SAMPLE_RETENTION_HOURS(default 168h).GET /concurrency?granularity=minute|hour&hours=Naggregates samples into per-bucket peak + avg (ConcurrencySampleStore.aggregate, bucketed in Python so it's identical across sqlite/mysql/postgres; fetch capped at 60k rows). Default lookback: minute→3h, hour→48h. - 大模型出字速度折线图 (tokens/sec, 按分钟/小时) —
GET /token-speed?granularity=&hours=aggregates the persistedllm_call_metrics(status="success") into per-bucket avg/min/max tokens/sec (deriving tokens/sec fromoutput_tokens÷duration_mswhen the stored value is null). A dip pinpoints the model/provider as the bottleneck for that window — "知道具体原因出现在哪里". - On/off + 不影响线上 — the whole feature is designed to never impact production: sampling is one tiny INSERT, retention-purged, all paths best-effort/guarded, queries bounded. A runtime switch lives in
system_settings.json → concurrency_monitor.enabled(ConcurrencyMonitorSettings, default false / opt-in, mtime-hot-reloaded): the background loop always runs (just sleeps) but only samples when an admin turns it on, and it re-reads the flag every tick so toggling it on/off takes effect immediately without a restart (the on-demand live snapshot + 停止 stay available regardless; only the historical chart needs sampling on).GET|PUT /api/admin/active-runs/settingsreads/writes it (admin only); the panel's 「后台采样」 Switch drives it. EnvCONCURRENCY_SAMPLE_ENABLED=0is a hard master kill (never starts the loop). Frontend client/hooks:core/admin/active-runs.ts+core/admin/hooks.ts(useActiveRunspolls 4s while the panel is open,useConcurrencySeries/useTokenSpeedSeriespoll 15s,useCancelActiveRun/useCancelAllActiveRuns,useMonitorSettings/useUpdateMonitorSettings). Tests:tests/test_concurrency_monitor.py.
Scheduled Task Runtime
ScheduledTaskService (deerflow/runtime/scheduler/service.py) polls due tasks every poll_interval_seconds and executes each in its own throwaway thread, serialized by a shared _agent_run_lock.
Stuck-run recovery: each execution attempt is bounded by _RUN_ATTEMPT_TIMEOUT_SECONDS (30 min) via asyncio.wait_for. On timeout the run is retried up to _MAX_RUN_RETRIES (3) times, then marked failed. A watchdog (_reconcile_stuck_runs, run every poll) recovers runs left in running with no in-process owner ? e.g. orphaned by a crash or restart ? re-driving them through the same retry budget so a task can never stay perpetually "executing". The scheduled_task_runs.retry_count column tracks consumed retries.
Unique task names: task names are unique per user (uq_scheduled_tasks_user_name). create_task/update_task/admin_update_task proactively check via _assert_name_available() and raise DuplicateTaskNameError on a collision; the router translates that into HTTP 409 instead of letting the raw DB IntegrityError surface as a 500.
HTML page tasks (task_kind == "html_page"): a second task kind that renders a full HTML document each run instead of a Markdown report. The kind + its config (html_config: page_prompt, negative_prompt, reference_favorite_id, model_name) live inside execution_context ? no schema change to scheduled_tasks. Data gathering is delegated to the agent's own search skills (driven by page_prompt), so there is no explicit search-query/data-source field; the render model is user-selectable (model_name, default zai-org/GLM-5-FP8). _run_task_attempt() dispatches by kind; _run_html_page_attempt() runs the pipeline: gather real data via the lead agent (its search tools) ? optionally load a reference favorite's reusable style/layout ? render with the chosen model (thinking off) ? fault-tolerant JSON parse ? sanitize (BeautifulSoup, strips <script>/<iframe>/on*/javascript:) ? structural validate ? one GLM5 repair pass if invalid ? completeness gate: if the HTML is still not a full valid document (missing doctype/html/head/body, empty body, or residual scripts) after the repair pass, raise HtmlGenerationError so the run is marked failed (a page only counts as successful when it is a complete, previewable HTML document) ? best-effort Playwright screenshot (no-op when Playwright absent) ? archive index.html ? persist a scheduled_html_pages row ? return result with format: "html" and the full HTML inline. Pure helpers live in deerflow/runtime/scheduler/html_page.py. The frontend previews the HTML in a sandboxed iframe (no allow-scripts) on the run-detail and public pages; a user can bookmark a page (style summary extracted via GLM5) and later pick it as a style reference. New tables scheduled_html_pages + html_page_favorites (deerflow/persistence/html_pages/, wired as app.state.html_page_store) are created by Base.metadata.create_all with no migration. Tests: tests/test_html_page_tasks.py.
Unified output switches (require_html + require_markdown). The task form no longer splits "???? / HTML ??" into separate kinds. execution_context now carries two independent booleans that may both be on: require_html (???? HTML) and require_markdown (???? Markdown). _task_requires_html() returns true for require_html or legacy task_kind == "html_page" (so old tasks need no data migration), and _run_task_attempt() dispatches to _run_html_page_attempt when it is true, else the regular agent attempt. When both switches are on, _gather_html_data() returns (text, files) ? the gather agent run is additionally told to save a .md file ? and _run_html_page_attempt validates it (MarkdownRequirementError on failure) and surfaces it alongside the generated page, so the run delivers both artifacts. The single main prompt doubles as the HTML page requirement (html_config.page_prompt); there is no separate page-prompt field. Tests: tests/test_scheduled_task_require_html.py + the markdown-coexistence cases in tests/test_html_page_tasks.py.
Template-fill mode (JSON → 固定 HTML 模板填充). A third execution mode, distinct from the free-form require_html render pipeline: convert a user-supplied JSON into a fixed HTML template's data structure and fill it. Triggered when execution_context.template_config.template_id is set (chosen in the task form after selecting the built-in 「模板数据转换器」 agent template-json-builder). _task_template_id() reads it and _run_task_attempt() dispatches to _run_template_fill_attempt before the require_html check. Pipeline: take the source (the task prompt, pasted or uploaded .json) — which may be pure JSON, plain text, or text-with-embedded-JSON (template_fill.extract_embedded_json pre-extracts any embedded {…}/[…] with a string-aware balanced scan and surfaces it as a labeled block while keeping the full original text; plain text → no block, the model reads it semantically) → load the selected agent's SOUL.md+model and run a direct _invoke_conversion_model call (system=SOUL, human=conversion prompt carrying the template's compact schema skeleton + the recognized-JSON block + source) → extract_json_object → template_fill.validate_data (top-level keys + container types) → on failure feed the structure errors + field_coverage_gaps back and retry up to _MAX_TEMPLATE_FILL_RETRIES=2 → template_fill.fill_template replaces the template's const DATA={…} line (and refreshes MANIFEST.系统.更新时间) → archive index.html + persist a scheduled_html_pages row → return format:"html" (reusing the existing HTML preview at /runs/{id}/html). Any terminal failure raises TemplateFillError → run marked failed (no retry, message shown). Pure engine deerflow/runtime/scheduler/template_fill.py (harness, no app.*): the template registry (TEMPLATES: id=no1 name=no.1 台湾滨海防卫大屏, and id=no2 name=no.2 台湾飞行部队平台 with convertible=False — a 「只展示」 template that renders its existing HTML as-is, no model/validation/fill, until conversion is wired in later; _run_template_fill_attempt short-circuits to read_template_text for it, and list_templates/is_convertible expose the flag so the form hides the JSON input), list_templates/reference_data/schema_skeleton_json/validate_data/field_coverage_gaps/fill_template — the template's own embedded DATA is the schema source of truth (parsed at load, 1-sample-per-array skeleton fed to the LLM, top-level structure as the validation basis); template HTML files live under runtime/scheduler/templates/<id>/, add a TemplateSpec to register a no.2. The agent is seeded like the AI-writing ones: app/gateway/routers/_template_builder_seed.py + _template_builder_seed_assets/template-json-builder/{config.yaml,SOUL.md}, called from the lifespan before _sync_legacy_agents (which upserts it into the agents table as built-in/published so it shows in /api/agents → the task form's 执行智能体 dropdown). Template list API: GET /api/scheduled-tasks/html-templates → {templates:[{id,name,description,convertible}]} (open). Debug-render API POST /api/scheduled-tasks/template-fill/render {template_id, text?|data?} → {ok, display_only?, errors, gaps, html} (open): no model call — extract_embedded_json(text) (or data) → validate_data → fill_template, returning the filled HTML plus the validation errors/field-gaps (fills even when invalid so you can see the effect). Powers the chat-page debug button below; non-convertible templates return their raw HTML. Templates as require_html 参考收藏页面 (use the REAL template, all sub-pages): the 参考收藏页面 list in the task form also offers the templates (value tpl:<id>, e.g. tpl:no1). These templates are JS-driven multi-page SPAs (sub-pages/tabs are rendered by script from DATA), so a free-form model render would only ever produce the main page and drop every sub-page. Therefore, when _run_html_page_attempt sees a tpl:<id> reference (template_fill.reference_template_id parses it; distinguishes a template marker from a real favorite id), it short-circuits before the free-form pipeline to _run_html_template_reference, which renders the actual template file (all sub-pages intact): convertible templates gather real data (_gather_html_data) → convert to the template DATA (_convert_to_template_data, shared with template-fill mode) → merge_with_reference (per top-level key = per sub-page: any sub-page/field with no extracted content falls back to the template's original data, so sub-pages never render blank) → fill_template; non-convertible templates (no.2) output the raw template as-is; if the task prompt and description are both empty, or gather/conversion yields nothing, it uses the template's original data outright. 数据抽取开关 html_config.template_extract (default off, shown in the form only when a tpl: reference is picked): it only controls whether the HTML uses extracted data vs the template's default data — it does not affect anything else. Off → HTML fills the template's default data (no convert); on → gather+convert+merge. Markdown is independent: when require_markdown is also on, the gather agent run still runs and the .md is produced/validated/surfaced regardless of the extract switch (so turning extraction off never blocks the Markdown). Display-only templates (no.2) never extract; they still honor require_markdown. _run_html_template_reference runs gather once when do_extract (extract on + has prompt/desc + convertible) or require_markdown, and threads any .md through _finalize_template_html(extra_files=…). merge_with_reference is also applied in template-fill mode (_run_template_fill_attempt) so empty sub-pages there fall back too. Archive/persist/result are shared via _finalize_template_html. (The earlier style_reference CSS-extraction approach was removed — it lost the sub-pages.) Frontend: when the agent dropdown = template-json-builder the form hides the HTML/Markdown switches and shows a 「HTML模板」 select (useHtmlTemplates) + a .json upload (reads file text into the prompt box). Tests: tests/test_template_fill.py, tests/test_template_fill_attempt.py, tests/test_template_builder_seed.py.
Per-user Prompts & UI Preferences
Custom prompts can be scoped to one agent. user_custom_prompts gains a nullable agent_id column (migration 20260603_02): NULL = global (the chip shows in every chat), otherwise the prompt only surfaces when that agent is selected. The visibility filter is applied in the frontend CustomPromptBar (it receives the current agent id from ChatPage's selector or AgentChatPage's assistantId); the backend (/api/user-prompts CRUD) just stores/returns/clears agent_id (update accepts an explicit null to clear). Tests: tests/test_user_prompts_agent_id.py.
Per-user UI preferences (deerflow/persistence/user_preferences/, wired as app.state.user_preferences_store; migration 20260603_03, also auto-created by create_all): one JSON-blob row per user so new toggles need no schema change. Router /api/user-preferences (GET/PUT, account-scoped). Current key: hide_admin_prompt_prefix ? when true the chat input hides the admin-configured prompt-prefix switch and never applies the prefix (the user uses only their own prompts; enforced in ChatPage for both shouldApplyPrefix and promptPrefixProp). Toggled from the "???????" dialog. Tests: tests/test_user_preferences.py.
HTML skills (tools the render step uses). Two offline skills under skills/public/ shape the generated HTML and are toggled in the skills settings page (enabled by default): html-page-builder (a self-contained inline-CSS design system ? colors/cards/tables/KPI/responsive) and html-charts (inline-SVG charts that render in the no-allow-scripts preview iframe; optional local ECharts documented in its assets/README.md). A skill opts into HTML injection via metadata.kind: html_page in its SKILL.md frontmatter; _collect_html_skill_instructions() reads the bodies of enabled such skills and build_generation_prompt(skill_instructions=?) injects them, so the (raw-model) render step honors the user's skill selection. The data-gathering step and the render/repair/style passes all use the task's selected model_name (no silent fallback to the default model ? and there is no hardcoded GLM-5; an unset selection falls back to config models[0]). To stop auxiliary LLM calls from leaking onto the default model, TitleMiddleware is skipped when is_scheduled_run is set (_build_middlewares in lead_agent/agent.py) ? scheduled/HTML runs use throwaway threads whose titles are never shown, and config.title.model_name is usually unset (? would call models[0], not the chosen model). The completeness gate (validate_html) also requires a non-empty <title> and forbids external stylesheets, on top of the doctype/html/head/body + non-empty-body + no-script checks.
Knowledge Base (packages/harness/deerflow/knowledge/)
Self-contained product knowledge base. All knowledge capability lives in this one module (harness layer, deerflow.* only ? no app.* imports). Wired as app.state.knowledge_service in deps.py via make_knowledge_service(sf, config.knowledge) (returns None on the memory backend or when knowledge.enabled=false ? router 503). Surfaced by app/gateway/routers/knowledge.py (/api/knowledge). Config: config.yaml ? knowledge (deerflow/config/knowledge_config.py).
Two stores, kept in sync on every write:
- DB mirror ? tables
knowledge_notes/knowledge_sources/knowledge_tags/knowledge_note_tags(knowledge/models.py, registered inpersistence/models/__init__.py, auto-created byBase.metadata.create_all; portable typesPortableLongText/BeijingDateTime). Powers listing/filter/search/audit. Global/shared ? nouser_idcolumn;created_by/updated_byare audit-only (resolved viaget_effective_user_id()). - Markdown vault ? obsidian-wiki compatible, default
{$DEER_FLOW_HOME}/data/knowledge/obsidian-vault.ObsidianWikiAdapter(obsidian_adapter.py) writes wiki-capture frontmatter+body, and after every write syncsindex.md/log.md(CAPTURE/INGEST/UPDATE/ARCHIVE/EXPORT) /hot.md/.manifest.json(source-hash dedup forskip_existing).markdown_vault.pyhandles frontmatter build/parse, slugging, wikilink extraction and path-traversal-proof resolution.
Module layout: models.py (ORM: notes/sources/tags/note_tags + embeddings/versions/entities/relations) ? repository.py (DB only, no LLM) ? schemas.py (NoteDraft/SourceDraft DTOs) ? skill_templates.py ? markdown_vault.py ? manifest.py ? extractor.py (rule-based fallback) ? llm_extractor.py (default: distills declarative knowledge per wiki-capture, structured JSON incl. entities/relations; falls back to extractor.py) ? sources.py ? exporter.py (graph.json/graph.graphml/cypher.txt + self-contained vanilla-JS graph.html; entity nodes + relation edges. graph.html is theme-aware ? it reads ?theme=light|dark from its own URL since the embedding iframe can't inherit the app's CSS; KnowledgeGraphPage.tsx appends the active day/night mode and re-skins on toggle. The layout cools down and freezes so dense graphs stay clickable, node radius scales with degree, and hovering/clicking a node focuses it + its direct neighbours while dimming the rest + a detail panel ? taming heavy node/edge clutter) ? chunking.py + embeddings.py (OpenAI-compatible embedding client, keyword fallback) + search.py (keyword/vector/hybrid) ? sensitive.py (secret/PII redaction) ? dedup.py (Jaccard + hash dedup) ? auto_ingest.py (background queue/worker) ? obsidian_adapter.py ? service.py (KnowledgeService orchestrator ? the only thing the router touches; capture entry points capture_thread / capture_search / capture_file / create_manual_note, all converging on _persist_draft). Tests: tests/test_knowledge.py (20 cases). Frontend: frontend-web/src/core/knowledge/* + pages/KnowledgeBasePage.tsx / KnowledgeGraphPage.tsx + chat knowledge-save-trigger.tsx / knowledge-rag-toggle.tsx; spec at frontend-web/docs/product-knowledge-base-dev.md.
Wikilinks & multi-note capture (knowledge.{related_notes_enabled,related_notes_limit,multi_note_enabled}, all default-on; project_name defaults to zncm):
- Navigable
[[wikilinks]]?service.resolve_wikilink(s)maps a[[target]]to a knowledge note (by id ?vault_path? title) or a vault meta page (read_vault_page, path-traversal-proof), elsemissing.markdown_vault.normalize_wikilink_targetcanonicalizes targets (drops|alias, slashes,.md). The frontendMarkdownRendererrenders non-numeric[[...]]as clickable links from the resolver map (numeric[[123]]stay citation badges); note links navigate, vault links open aVaultPageDialog, missing links are greyed. - Real related notes only (no structural meta-links) ?
extractor.related_links()now returns[]: notes no longer auto-linkindex/_meta/taxonomy/ a project hub (they were noise in every## Related). On capture,service._find_related_notescross-links only by shared distinctive tags (project/knowledge-basedefaults excluded; free-text title/summary matching was removed because CJK tokenization linked unrelated notes via generic words like ??/????) and_inject_related_noteswrites real[[Title]]links into## Related(plus note?noterelatedgraph edges); no distinctive tags ? no related notes. The renderers omit the## Relatedsection entirely when there are no related notes and no sources. The vault's ownindex.md/_meta/taxonomy.mdstill exist (obsidian scaffold) and remain resolvable for manually-written links. The router enrichesGET /notes/{id}withcreated_by_name/updated_by_name(auth-provider email lookup) so the UI shows a person, not a UUID. The frontendMarkdownRendererrenders wiki-capture provenance markers^[inferred]/^[ambiguous]as small "??"/"??" badges. - LLM-extraction outcomes ? capture distinguishes three LLM outcomes: (1) success ? declarative note(s); (2) deliberate rejection ? the model returns
should_ingest:falsewith no notes (greeting / meta chatter) ?extract_thread_llm_multiraisesLLMRejectedIngestandcapture_threadraisesEmptyThreadError(skip ? nothing persisted, so worthless chats never enter the KB); (3) genuine failure ? API error / non-JSON ? returns[], capture falls back to the rule-based verbatim extractor and lands the note as adraft(notapproved), logs a WARNING, and returnsllm_fallback:true(surfaced by the router).llm_extractorlogs an actionable WARNING/INFO for each non-success path so extraction issues are diagnosable. - Multi-note split ?
llm_extractor.extract_thread_llm_multilets one conversation distill into 1?4 cross-linked notes (notes:[?]JSON array; legacy single-object output still accepted).capture_threadpersists each, sibling-cross-links them by title, and returns{note, created, notes:[?]}. Gated bymulti_note_enabled(false ? first note only). Tests:tests/test_knowledge_wikilinks.py.
Phases 2?4 (all config-gated, defaults safe):
- Phase 2 ? vector RAG:
knowledge.embedding.{enabled,base_url,api_key,model,...}enables an OpenAI-compatible embedding endpoint; notes are chunked + embedded on write (knowledge_embeddings),POST /searchsupportsmode=keyword|vector|hybrid(vector degrades to keyword when unconfigured).KnowledgeRagMiddleware(lead-agent chain) injects a transient<knowledge-context>recall before the model call, gated per-conversation by theknowledge_rag_enabledruntime flag (frontend chat toggle) falling back to globalknowledge.rag_enabled.POST /reindexrebuilds the index. - Phase 3 ? governance:
sensitive.pyredacts secrets/PII before persistence (sensitive_redaction_enabled); every update snapshotsknowledge_note_versions(GET /notes/{id}/versions);GET /duplicates+POST /mergefor dedup; backgroundAutoIngestQueue(wired indeps.py, started/stopped in the gateway lifespan) auto-sediments finished conversations off the chat path whenauto_ingest_enabled? enqueued fromKnowledgeRagMiddleware.after_agent, value-judged by min-chars +skip_existing, landing asdraft. - Phase 4 ? graph: LLM extraction also yields entities/relations ?
knowledge_entities/knowledge_relations;exporter.build_graphadds entity nodes + note?entity/entity?entity edges (reflected ingraph.html/graph.json);GET /graphreturns product graph data.
User directories + editable extraction templates + display polish (knowledge UX overhaul):
- Custom directories (
knowledge_notes.folder+knowledge_folders) ? notes carry a free-form, slash-joinedfolderpath (e.g.投研/行业) independent of the wiki-capturecategoryand ofvault_path.knowledge_foldersis the manageable directory list (supports empty folders + rename/move that re-points descendant folders and filed notes). Router CRUD under/api/knowledge/folders(GET/POST/PUT/{id}/DELETE/{id}?reassign_to=);GET /notes?folder=filters a directory + its sub-folders (folder=""? unfiled). Capture (capture_thread/capture_file/create_manual_note) andPUT /notes/{id}acceptfolder(the note PUT usesmodel_fields_setso an absent key leaves the folder untouched, a presentnullclears it). FrontendbuildKnowledgeTree(notes, folderPaths)builds the left tree fromfolder(no longer fromvault_path); folder management dialog +FolderSelectin the create/import dialogs and the note detail header. - Editable extraction prompts (
knowledge_extract_templates) ? the hard-coded distillation prompts are seeded as editable templates on startup (KnowledgeService.seed_default_templates, idempotent —默认 · 对话提炼/默认 · 文档提炼, eachis_defaultfor its scope).scope ∈ {thread, document, both}; capture resolves an explicittemplate_id→ the scope default → the built-in constant, and passes it tollm_extractor.extract_thread_llm_multi/extract_document_llm_multiassystem_prompt(used verbatim). Router CRUD under/api/knowledge/extract-templates(409 on duplicate name). Frontend: extraction-template management dialog +TemplateSelectin the batch/file import dialogs. Editing a template is the supported way to fix "提炼太精简" (loosen the prompt to keep more detail) and to add alternate extraction logics. - Single source list ?
extractor._related_section_mdno longer emits a### Sourcessub-section into the body; the note detail shows sources once via the structured来源block (note.sources).## Relatedstays pure wikilinks so the frontend extracts it cleanly into the related-docs module. - Chinese section headings (display-only) ? the frontend
localizeHeadings(wikiUtils) mapsContext/Finding / Decision/Reasoning/Implications/Related/Sourcesto 中文 at render time (applied to the body fed to bothextractTocand the viewer, so anchors stay consistent). No data migration — historical English-heading notes render in Chinese too. - Historical-data cleanup ?
scripts/cleanup_knowledge_related.py(one-off, idempotent,--dry-run/--note-id) strips legacy## Relatedstructural links (vault scaffold:index/_meta/taxonomy/projects/…, any slash-bearing target) and the old### Sourcessub-section from existing notes (DBcontent_md+ vault files, viaKnowledgeService.update_note). - Migration
20260616_02+ idempotentscripts/sql/column_additions_mysql.sqlcoverknowledge_notes.folder(+ index); the two new tables are auto-created bycreate_all/_ensure_orm_columns_syncon startup. Tests:tests/test_knowledge.py.
Custom Agent Persistence
Custom agent metadata lives in the agents ORM table (deerflow.persistence.agents). agents.id is the stable runtime id and the filesystem directory name under .deer-flow/agents/{id}/; agents.name is display-only, can contain Chinese characters, and is not unique. user_id IS NULL marks built-in agents, user-owned agents are visible only to their owner unless published = true, and update/delete operations require ownership. The runtime accepts agent_id in context/configurable, and non-default assistant_id values are treated as agent ids after visibility checks.
Listing cache + pagination ? to keep GET /api/agents fast, the gallery's per-card extras (skills + tool_groups) are denormalized into two cache tables instead of reading every agent's config.yaml on each request: agent_skills (relational, one row per (agent, skill) with position) and agent_extras (one row per agent: tool_groups JSON + model + skills_synced_at). config.yaml remains authoritative for the agent runtime; these tables are a read cache, write-through on create/update and backfilled from disk by _sync_legacy_agents (mtime-vs-skills_synced_at drift check). The list endpoint enriches via AgentStore.get_extras_for([ids]) (two batched IN queries). Both tables are auto-created by Base.metadata.create_all (no migration needed). GET /api/agents supports opt-in server-side pagination: scope (mine = owned + favorited / square = built-in/published public), page, page_size, returning {agents, total, page, page_size}; omitting scope/page preserves the legacy "return everything" behavior. Name search matches the raw OR desensitized name (the frontend's static term?abbrev map is mirrored server-side in deerflow/config/agent_name_desensitize.py); tag_ids is an SQL OR-filter; owner keeps only agents published by that user id (the card's ??? ? built-in agents never match), powering the gallery's click-???-to-filter. POST /api/agents/private-square/list takes the same pagination params (incl. owner). Repository entry point: AgentStore.list_paginated(...). Tests: tests/test_agent_listing_cache.py::test_list_paginated_owner_filter_keeps_one_publisher.
Homepage selector (agent picker on the new-chat page) ? single admin-curated list backed by agents.featured_order (nullable BIGINT). NULL = not featured. Only published = true agents can be featured (router enforces on writes); any published agent is eligible regardless of ownership. Cleared with order: null. The list rendered to end users is WHERE featured_order IS NOT NULL + visibility-filter, sorted by featured_order ASC. Admins reorder by sending the desired ids to PUT /api/agents/featured/order (assigns indices 0..N-1). BIGINT is used rather than INT so the management UI can seed initial values from Date.now() (overflows INT on MySQL).
Publish Squares (????)
Admin-configurable, access-control-free categories that published agents/skills are filed under (distinct from the password-gated ????). The square list lives in system_settings.json ? publish_squares (PublishSquaresSettings: ordered squares: [{id,name,description}] + default_square_id); helpers list_publish_squares / resolve_default_square_id / normalize_square_id always yield a valid square (built-in "????" fallback when none configured). Admin CRUD: GET/PUT /api/system-settings/publish-squares (GET readable by any user ? the gallery sub-tabs + publish picker need it; PUT admin-only, auto-slugs missing ids and re-anchors a dangling default).
Each agent/skill carries a single square_id column (AgentRow / SkillRow, String(64) NOT NULL default '', migration 20260611_01, idempotent SQL in column_additions_mysql.sql). One column = at most one square, so an item can never be "published twice". '' = unassigned ? resolved to the default square at read/filter time. Writes normalize via normalize_square_id: agents on create/update (square_id field), skills on publish (PUT /api/skills/custom/{name}/published accepts optional square_id). Listing filters: GET /api/agents?square= and GET /api/skills?square= keep only that square's items; the default square additionally sweeps in legacy '' rows (square_is_default). Tests: tests/test_publish_squares.py. Frontend: core/system-settings (usePublishSquares/useUpdatePublishSquares), admin PublishSquaresSettingsPage (embedded in the appearance/settings hub), square picker in the agent edit dialog + per-skill card, and a ??-tab square sub-filter in both galleries.
Reference-mode Entry Switch
The 参考文献/RAG citation-mode composer entry is system-controlled through
system_settings.json → citation_display.reference_mode_enabled (default
false). GET /api/system-settings/citation-display remains readable so the
frontend can hide or show the entry consistently; PUT is admin-only and
partially updates reference_mode_enabled and/or cleanup_words. The normal
workspace input box, iframe chat toolbar, send context, right reference panel,
and appearance settings all use this flag, while already locked citation
threads and selected LLMWiki knowledge bases still render their source
documents in the reference panel. Tests:
tests/test_prompt_prefix.py and
frontend-web/src/pages/iframe-chat-mode-buttons.test.ts.
Agent Module Tag Config (zzJC / zzXD 按类型聚合)
Backs the zzJC / zzXD agent-chat left sidebar (「按类型聚合智能体 + 聊天记录」). An admin picks which agent tags belong to a module; the module's left sidebar then shows only agents under those tags (empty/unconfigured = show all visible agents). Stored in system_settings.json → agent_modules (AgentModuleSettings: modules: {module_key → [tag_id, ...]}); pure helpers get_module_tag_ids / set_module_tag_ids (de-dupe + strip). Endpoints (in app/gateway/routers/system_settings.py): GET /api/system-settings/agent-modules/{key} (readable by any user — the sidebar needs it to filter), PUT /api/system-settings/agent-modules/{key} (admin only, full replacement of the tag-id list). Module keys are frontend-defined (AGENT_MODULES in components/workspace/agents/agent-module-sidebar.tsx, keyed by the module's landing agent id: 6bf→zzjc, 10d69…→zzxd). Tests: tests/test_agent_modules.py. Frontend: core/system-settings (useAgentModuleConfig/useUpdateAgentModuleConfig), admin AgentModuleSettingsDialog (gear button in the agent-chat header).
Business-mapping authorization
/api/business-mapping remains globally readable. Updates to 3Q, 6BF, and 7BF require system_role == "admin"; the 8BF public business-chain mapping is intentionally different and can be updated only by the trusted authenticated account whose email prefix is lqq. Do not rely on a username supplied in an external login URL. The frontend follows the same email-prefix rule when deciding whether to render the 8BF configuration button, while the router (app/gateway/routers/business_mapping.py) remains the enforcement point. Tests: tests/test_business_mapping.py.
Position-collaboration foundation
Roundtable business-chain enable/disable
roundtable_chains.enabled (BOOLEAN NOT NULL DEFAULT 1, migration 20260819_02)
is the per-chain 启停 switch. New and existing chains default to enabled.
Owner or admin can flip it via PUT /api/roundtable-chains/{id} {enabled}.
Disabled chains stay visible in the management list (with an 启用/停用 filter)
and still show updated_at; the frontend hides them from the picker / public
entry rail so they cannot be selected into a roundtable. Tests:
tests/test_roundtable_chain_stages.py (test_create_defaults_enabled_then_toggle).
Roundtable business-chain human validation
The existing /api/roundtable-chains seats JSON also accepts the optional
boolean human_validation field on each seat (ChainSeatSchema). It is a
frontend Step 2 checkpoint, not a server-side approval task: after the marked
seat completes (or, in a DAG, after its stage completes), the client pauses and
uses the existing resume button or intervention composer to continue. No new
database column, migration, or validation card is required.
The position-collaboration workspace is a separate frontend route and must not
use the original roundtable coordinator or jobs. Its shared configuration
integration point is the existing /api/roundtable-chains store: each seats JSON item
may carry an optional position_id (validated by ChainSeatSchema, no schema
migration). The original roundtable only reads its established seat fields and
therefore remains unchanged. The separate /api/position-roundtable router
(also mounted at /api/multi-agent/position-roundtable for intranet gateway
compatibility)
persists task/intent/frozen-chain snapshots in position_roundtable_sessions
and direct-agent node indexes in position_roundtable_nodes (migration
20260721_01). Message history remains in the established LangGraph thread
store; a node completion unlocks its next chain stage and a later upstream
revision marks downstream records stale for refresh.
POST /api/position-roundtable/sessions/{id}/nodes/{node_key}/reject uses the
same cascade: its reason is intentionally optional, the selected node becomes
rejected, and downstream material becomes stale (or locked when no
material exists) with invalidated_by set. Clients must replace their complete
session snapshot from this response so the flow, delivery list, and active
node state stay consistent. A rejected/stale node returns to done only after
a fresh Markdown delivery is recorded.
For direct-agent handoff, the router resolves artifact paths only inside the
completed upstream node's own thread, then includes bounded UTF-8 excerpts for
text outputs; binary artifacts stay as metadata-only references. Because each
seat runs on an isolated LangGraph thread, a downstream read_file on the
upstream artifact path would fail with "File not found"; the handoff therefore
also mirrors every upstream outputs/ deliverable into the current seat's
own workspace/upstream/ directory (_mirror_upstream_artifacts, bytes from
the upstream sandbox file with the SQL recovery copy as fallback, idempotent
per turn via size comparison, capped at 8 files / 2 MB each / 8 MB total) and
rewrites each handoff artifact ref's path to the mirror (upstream_path
keeps the original; mirrored_to_sandbox: true), with a matching note in the
hidden instructions telling the agent it may read_file the full text. The
mirror deliberately stays outside outputs/ so it never appears in the seat's
own delivery manifest.
POST /sessions/{id}/archive turns the session into a read-only historical
record: task/intent updates, node binding, hidden-context construction, and
turn completion are rejected, while read endpoints remain available.
Built-in summary and action-plan completion also preserves a visible report
answer when the model misses its final write_file call: the router writes
that exact answer as the required Markdown artifact in the owning thread's
outputs/ directory, so users receive both chat text and a downloadable file.
Per-turn outputs write cap (弱模型防重复产物): ordinary position-node chats
and the summary builtin send run-context position_artifact_cap: 1 (whitelisted
in services.merge_run_context_overrides). PositionArtifactCapMiddleware
then keeps at most one successful /mnt/user-data/outputs/ write_file and
one present_files per user turn — truncates excess tool calls in the same
model response (and syncs raw provider additional_kwargs.tool_calls), rejects
further writes/presents after an OK / Successfully presented files or an
earlier sibling in the same batch, ends the turn when only excess calls remain,
and reminds the model to present once (or wrap up if already presented).
Action-plan does not send the flag (it must write report .md +
action-plan-subtasks.json). Tests: tests/test_position_artifact_cap.py.
岗位 (Position) Subscription System
An admin-managed aggregation layer that decides, per user, which agents / skills
/ scheduled tasks / recommended questions they are subscribed to or can see.
A user has at most one position (users.position_id, 1:1); unassigned users
inherit the single is_default position. Note: the admin-configured prompts
shown to users ARE the recommended questions (recommended_questions table)
— there is no separate prompt-template store; the per-user user_custom_prompts
(personal 常用提示词 bar) is untouched by this system.
Persistence (deerflow/persistence/positions/, wired as
app.state.position_store; tables auto-created by create_all, migration
20260616_01):
positions(PositionRow):id,name(unique),description,is_default(only one at a time — the store clears the flag on others).position_tag_bindings(PositionTagBindingRow): one row per(position, resource_type, tag). No binding for a resource_type = "use the default set".POSITION_RESOURCE_TYPES = (agent, skill, scheduled_task, recommended_question)— the same set drives the tag types (TAG_TYPESwas extended withrecommended_question).
Sync engine (deerflow/services/position_sync.py, wired as
app.state.position_sync). sync_user(user) materializes three resource
types into the existing "我的" tables, idempotently:
- agents →
agent_favoritesrows withorigin='position' - skills →
skill_favoritesrows withorigin='position'(tableSkillFavoriteRow, migration20260616_04, auto-created bycreate_all) - scheduled tasks →
scheduled_task_subscriptionsrows withorigin='position'
Skills behave exactly like agents: a position only adds its tagged skills
to the member's 我的技能 — the browse-all 广场 is never restricted. (This replaced
the earlier query-time visibility filter, which wrongly shrank the square to the
bound skill.) PositionSyncService.resolve_visible(user, type) returns
set | None (None = unrestricted) and is now used only for the one remaining
query-time type, recommended_question.
Sync only ever adds/removes origin='position' rows, so a user's own
favorites/subscriptions are never clobbered. Default sets (no binding for a
type): agents = built-ins (user_id IS NULL); skills = empty (square stays
full, nothing added to 我的); scheduled tasks = empty; questions = untagged
only (默认岗位 / 无标签绑定 只看到未打标的题目;打标的题目只在绑定了对应标签的
岗位展示 — the router interprets resolve_visible → None as "show only untagged",
not "show everything").
origin/position_id columns were added to agent_favorites and
scheduled_task_subscriptions — see migration 20260616_01 +
scripts/sql/column_additions_mysql.sql. The running app also auto-adds these
via _ensure_orm_columns_sync on startup.
Triggers: assigning/removing members, changing tag bindings (re-syncs all
members), becoming the default position (sync_default_members), and login
(auth.py::_sync_user_position, best-effort so a brand-new user inheriting the
default gets their grants materialized).
API (app/gateway/routers/positions.py, /api/positions, admin only):
position CRUD; GET/PUT /{id}/tag-bindings; GET/POST/DELETE /{id}/members
(batch assign/unassign → triggers sync); POST /{id}/resync;
GET /users (all users + their position); PUT /users/{user_id} (set one
user's position). GET /me is the one non-admin route — it returns the
caller's own effective position (explicit assignment, else the default position
unassigned users inherit, else null); powers the top-bar 岗位 label
(strategy-components/Header.tsx user cluster, via useMyPosition). Related
changes: GET /api/recommended-questions is now
filtered by the caller's position (?all=1 bypasses for admin management) and
admin can tag questions via /api/tags/assignments/recommended_question/{id}.
Default-position rule: tags are bulk-attached to all questions first,
then the visibility filter applies — a position binding recommended_question 标签
sees only those tagged questions; a 默认岗位 / 无绑定 (resolve_visible → None)
sees only the untagged questions (打标的题目不再泄漏到默认岗位). ?all=1
(admin management) skips this filter and returns every question with its tags.
The list also accepts a repeatable tag_ids
OR-filter — so the 推荐问题配置 page can display per-question 标签 and search by
them (frontend RecommendedQuestionsPage + TagSearchBar). Tests:
tests/test_recommended_questions_tags.py.
GET /api/skills no longer restricts by position; each skill carries a
favorited flag (true = in the caller's 我的技能, e.g. position-granted) so the
skills page's 我的 tab includes owned ∪ position-granted skills while the 广场
stays complete. Frontend: core/positions
- admin
PositionManagementPage(侧边栏「综合管理 → 岗位管理」 + 右下角设置菜单). Recommended questions stay rendered above the chat input (their original position), just filtered by the caller's position. Tests:tests/test_positions.py.
Portable Column Types (deerflow/persistence/types.py)
Two TypeDecorators keep ORM columns correct across the sqlite / postgres / mysql backends. All ORM timestamp/JSON columns must use these, not the raw SQLAlchemy types.
PortableJSON? nativeJSONon SQLite/Postgres;TEXT(with manualjson.dumps/loads) on MySQL, where older servers lack JSON support.PortableLongText? a large text column: maps to MySQLLONGTEXT(4 GB) and plainTEXTon SQLite/Postgres (already unbounded). Use for big JSON/markdown blobs that would overflow MySQL's 64 KBTEXTcap ? e.g.ai_writing_sessions.transcript.BeijingDateTime? stores timestamps as Beijing (UTC+8) wall-clock time. The app records time withdatetime.now(UTC), but MySQLDATETIME/ SQLite have no timezone and persist whatever wall clock the driver serializes ? so a rawDateTimecolumn would store UTC and read 8 hours behind Beijing.BeijingDateTimeconverts UTC?Beijing on bind and re-attaches+08:00on result, so the database holds Beijing time while Python still gets timezone-aware datetimes (instant comparisons stay correct, including in WHERE clauses). PostgreSQLtimestamptzstores the instant unambiguously, so values pass through unchanged there. Uses a fixed UTC+8 offset (China has no DST), so notzdatadependency.
Schema Migrations / Manual Column-Add SQL (scripts/sql/)
Gateway startup runs Alembic through app.gateway.app._run_db_migrations after ORM table creation. The migration graph must always have unique revision ids and exactly one head; tests/test_migration_graph.py enforces both. Historical duplicate revision 20260821_01 is repaired by the 20260821_03 / 20260821_04 fixed-question chain and merge revision 20260825_01. Startup DDL has a bounded timeout and failures are logged without preventing the app from serving.
Convention (required): whenever code adds a field to an existing table (a new mapped_column on an ORM model), you must, in the same change:
- Keep the matching alembic version file under
persistence/migrations/versions/(guarded with_has_table/_has_column), and - Append the idempotent
ADD COLUMNstatement toscripts/sql/column_additions_mysql.sql, with a definition matching the ORM column exactly (type, length, nullability).
Operators upgrading an existing DB may run the whole file (USE <db>; then execute) as a recovery path; it is idempotent (existing columns are skipped via information_schema.columns). Index-only changes go in the sibling scripts/sql/leaderboard_indexes_mysql.sql (same idempotent pattern) and stay out of startup migrations because online index creation on large metric tables can delay service boot.
Leaderboard scheduled-run exclusion uses correlated NOT EXISTS against the
indexed scheduled_task_runs.agent_run_id; do not restore NOT IN. Legacy
scheduler detection for tool metrics is also a correlated existence check and
must not outer-join the complete runs table into metric aggregations. Keep the
created-at + dimension compound indexes in ORM metadata and the idempotent
MySQL index script synchronized.
Sandbox System (packages/harness/deerflow/sandbox/)
Interface: Abstract Sandbox with execute_command, read_file, write_file, list_dir
Provider Pattern: SandboxProvider with acquire, get, release lifecycle
Implementations:
LocalSandboxProvider- Singleton local filesystem execution with path mappingsAioSandboxProvider(packages/harness/deerflow/community/) - Docker-based isolation
Virtual Path System:
-
Agent sees:
/mnt/user-data/{workspace,uploads,outputs},/mnt/skills -
Physical:
backend/.deer-flow/users/{user_id}/threads/{thread_id}/user-data/...,deer-flow/skills/ -
Translation:
replace_virtual_path()/replace_virtual_paths_in_command() -
Detection:
is_local_sandbox()checkssandbox_id == "local" -
{user_id}for path resolution =resolve_path_user_id(thread_id), using the persisted creator for every tracked thread. The route's regular owner check still authorizes access first; creator resolution only selects the stable physical directory (threads_meta.user_id) and avoids a lost request context falling back tousers/default. This also preserves cross-account task-deeplink conversations (rwfx/3qfx/xdfx) and roundtable system threads. Positive creator mappings are cached process-wide and DB/memory/missing-row failures fall back to the current login. It is wired at the thread-path consumers includingThreadDataMiddleware.before_agent,present_file_tool, uploads, artifact serving, and embeddedclient.get_artifact. Rewrite-history recovery is metadata-only and job-owner scoped (list_by_session(..., user_id=current_user)), so its GET/list endpoints intentionally skip redundant thread-owner checks and return an empty history instead of a history-only 404; the start/commit endpoints retain strict thread write authorization. Tests:tests/test_thread_paths_owner.py,tests/test_thread_paths_roundtable.py. -
Durable document-rewrite commit:
document_rewrite_pipeline.atomic_write_textwrites through a deliberately short sibling temporary filename so long Windows artifact paths stay below the path-length limit, and guarantees the targetoutputsparent exists. If the sandbox recreates that directory in the narrow commit window, it repairs the directory and retries; the original file is otherwise only changed throughos.replace. -
Deep Research structural regeneration:
POST /api/deep-research/sessions/{id}/report-variantscreates a distinct completed session from the source report and its selected evidence. It rebases both saved[来源:source_id]and legacy streamed[[source:source_id]]markers to cloned ids, and removes inherited markers with no selected clone. The frontend uses this endpoint only for “选择结构再生成”, then queues the durable writer withoperation=generate; generation receives the confirmed outline and complete selected evidence set but never receives the inherited report as text to edit. The new report's completed card can subsequently start a separate in-placeAI 全文改写. Requirement plans are deterministic and operation-specific, model reasoning is forwarded as live SSE before the first Markdown token, and generation prompts emit display-ready[来源:id]citations. Citation validation repairs only a unique near-completedrs_src_*prefix (provider-truncated suffix) and still rejects genuinely unknown ids. The report-only stream requests 8192 output tokens (except Codex Responses models, which reject token-limit kwargs) and rejects a malformed trailing citation instead of committing a provider-truncated report. Never restore the former second planning-model call unless its latency and streaming contract are redesigned together.
Sandbox Tools (in packages/harness/deerflow/sandbox/tools.py):
bash- Execute commands with path translation and error handlingls- Directory listing (tree format, max 2 levels)read_file- Read file contents with optional line rangewrite_file- Write/append to files, creates directoriesstr_replace- Substring replacement (single or all occurrences); same-path serialization is scoped to(sandbox.id, path)so isolated sandboxes do not contend on identical virtual paths inside one process
Subagent System (packages/harness/deerflow/subagents/)
Built-in Agents: general-purpose (all tools except task) and bash (command specialist)
Execution: Dual thread pool - _scheduler_pool (3 workers) + _execution_pool (3 workers)
Concurrency: MAX_CONCURRENT_SUBAGENTS = 3 enforced by SubagentLimitMiddleware (truncates excess tool calls in after_model), 15-minute timeout
Flow: task() tool ? SubagentExecutor ? background thread ? poll 5s ? SSE events ? result
Events: task_started, task_running, task_completed/task_failed/task_timed_out
Tool System (packages/harness/deerflow/tools/)
get_available_tools(groups, include_mcp, model_name, subagent_enabled) assembles:
- Config-defined tools - Enabled entries are resolved from
config.yamlviaresolve_variable(); entries withenabled: falseare omitted before import and are not registered with the agent - MCP tools - From enabled MCP servers (lazy initialized, cached with mtime invalidation)
- Built-in tools:
present_files- Make output files visible to user (only/mnt/user-data/outputs)ask_clarification- Request clarification (intercepted by ClarificationMiddleware ? interrupts)view_image- Read image as base64 (added only if model supports vision)
- Subagent tool (if enabled):
task- Delegate to subagent (description, prompt, subagent_type, max_turns)
Community tools (packages/harness/deerflow/community/):
tavily/- Web search (5 results default) and web fetch (4KB limit)jina_ai/- Web fetch via Jina reader API with readability extractionfirecrawl/- Web scraping via Firecrawl APIconfigurable_search/- Config-drivenPOSTJSONweb_searchreplacement. It deep-copies the configured payload and overwrites onlyquery, supports HTTP/HTTPS plus configurableverify_ssl, timeout and result limit, normalizes nullablerecUuid/content/title/urlfields, and rewrites a result URL throughresult_url_templateonly whenrecUuidis non-empty. The required response envelope is{"results": [...]}. Tests usehttpx.MockTransportand never require the deployment endpoint.
ACP agent tools:
invoke_acp_agent- Invokes external ACP-compatible agents fromconfig.yaml- ACP launchers must be real ACP adapters. The standard
codexCLI is not ACP-compatible by itself; configure a wrapper such asnpx -y @zed-industries/codex-acpor an installedcodex-acpbinary - Missing ACP executables now return an actionable error message instead of a raw
[Errno 2] - Each ACP agent uses a per-thread workspace at
{base_dir}/users/{user_id}/threads/{thread_id}/acp-workspace/. The workspace is accessible to the lead agent via the virtual path/mnt/acp-workspace/(read-only). In docker sandbox mode, the directory is volume-mounted into the container at/mnt/acp-workspace(read-only); in local sandbox mode, path translation is handled bytools.py image_search/- Image search via DuckDuckGo
MCP System (packages/harness/deerflow/mcp/)
- Uses
langchain-mcp-adaptersMultiServerMCPClientfor multi-server management - Lazy initialization: Tools loaded on first use via
get_cached_mcp_tools() - Cache invalidation: Detects config file changes via mtime comparison
- Transports: stdio (command-based), SSE, HTTP
- OAuth (HTTP/SSE): Supports token endpoint flows (
client_credentials,refresh_token) with automatic token refresh + Authorization header injection - Runtime updates: Gateway API saves to extensions_config.json; LangGraph detects via mtime
Skills System (packages/harness/deerflow/skills/)
- Location:
deer-flow/skills/{public,custom}/ - Format: Directory with
SKILL.md(YAML frontmatter: name, description, license, allowed-tools) - Loading:
load_skills()recursively scansskills/{public,custom}forSKILL.md, parses metadata, and reads enabled state from extensions_config.json. A directory containingSKILL.mdis a self-contained leaf — discovery does not descend into its subdirectories, so a skill package that bundles its ownexamples//references/sub-skills (each with a SKILL.md) surfaces as one skill, not a swarm of phantom entries. - Discovery↔delete consistency: a skill is keyed for discovery by its frontmatter
name, but its on-disk directory name may differ (hand-dropped / mis-extracted package).LocalSkillStorage._resolve_skill_dir(name, category)resolves the discovered directory (canonical<category>/<name>fast-path, else scan-by-frontmatter-name);custom_skill_exists/read_custom_skill/delete_custom_skillall go through it, so anything that gets listed can also be read and deleted regardless of directory name (the old code assumedcustom/<name>/and 404'd "Custom skill 'x' not found." on any non-canonical layout). Delete is path-guarded to never escape thecustom/root. Tests:tests/test_skill_storage_delete.py. - Injection: Enabled skills listed in agent system prompt with container paths
- Configurable
es_queryrouting prompt:config.yaml → skills.es_query_routingcontains anenabledswitch and editableprompt. When enabled,lead_agent.prompt._build_es_query_routing_sectioninjects the rule into the default lead-agent system prompt. Custom-agent-only prompts receive it only when their explicit skill allowlist containses_query; disabling the switch or leaving the prompt blank emits no block. Immediately before an ordinary Q&A graph is created,lead_agent.agent._log_es_query_routing_registrationchecks the exact configured block against the final system-prompt string and logsenabled, allowlist eligibility,registered, character count, and a short SHA-256 fingerprint (never the prompt body). Tests:tests/test_es_query_routing_prompt.py. - Installation:
POST /api/skills/installextracts .skill ZIP archive to custom/ directory. Direct uploads use the ownership-awarePOST /api/skills/validate?target_category=...preflight plusPOST /api/skills/install-upload: a same-owner custom conflict may be uploaded again withoverwrite=true, while a different owner's conflict is non-overwritable and returns that owner's username (falling back to user id). The install endpoint repeats the ownership check, including for admins targetingcustom, so callers cannot bypass the preflight. Public-skill overwrite remains admin-only. - Installable sentiment analysis package:
skill-packages/sentiment-analysis.skillis deliberately separate from the frontend virtual-agent chat page. The chat page goes through GatewayPOST /api/sentiment-agent/stream(runtime-config URL/Basic Auth in the body, TLS verify off, raw AG-UI SSE passthrough). The skill package's self-containedconfig/sentiment-agent.jsonowns the external AG-UI address for agent-skill callers; Basic Auth is resolved fromSENTIMENT_AGENT_AUTH_USERNAME/SENTIMENT_AGENT_AUTH_PASSWORDat runtime (not committed into the package).scripts/sentiment_stream.pyuses only the Python standard library and converts upstream SSE to stable NDJSON events (thinking_delta,answer_delta,done), with a Markdown convenience mode for normal agent calls. Installation/calling guide:skill-packages/sentiment-analysis/docs/调用与安装说明.md. - Chinese name (
name_zh): a Chinese display name stored in theskillsDB table (SkillRow.name_zh,String(128), migration20260609_01) ? never written to the on-disk SKILL.md, so it works for built-in skills too. Threaded throughSkillStorecreate/ensure_legacy/update/update_admin exactly likedetail; surfaced onSkillResponse.name_zhfrom the ownership record in_skill_to_response. Authorization mirrorsdetail: custom skill ? owner or admin; built-in/public ? admin only (row materialized on demand viaensure_legacy). Endpoints (inapp/gateway/routers/skills.py):PUT /api/skills/{name}/name-zh(set/clear),POST /api/skills/{name}/name-zh/generate(LLM generates one),POST /api/skills/name-zh/generate-missing(????: admin-only bulk-generate for every skill that lacks one, bounded-concurrencyasyncio.gather). Optional model selection: both generate endpoints accept an optional JSON body{model_name}? an explicit name wins, else the curator model setting, else the main model (_resolve_skill_name_zh_model_name). Model-error surfacing:_generate_skill_name_zhraisesSkillNameZhModelErrorwhen the model itself fails (build/timeout/API error) ? distinct from the model returning an unusable name (?None). The single endpoint maps that to502 ?????,?????????:?; the bulk endpoint fails fast on a model-build error (502) and otherwise carries a representativemodel_errorstring inSkillNameZhBulkResponseso the UI can say "N ?????????". The frontend (skill-settings-page.tsxlist +skill-detail-dialog.tsxeditor, plus the agent skill pickers/badges inNewAgentPage.tsx+agent-card.tsx) showsname_zh || name, offers per-skill "AI ??" plus an admin-only header "???????" button, and a compact model dropdown (components/workspace/model-name-select.tsx, default = configured model) beside each action.
Skill Address Management (技能地址管理 — URL/IP 批量替换)
Admin-only maintenance tool to find every http(s):// URL and IPv4 (ip[:port])
address in the running skills (skills/{public,custom}, all .md + .py,
recursing into scripts//references//templates/) and rewrite them in bulk —
e.g. swapping a UAT endpoint for production across all skills at once. The same
address commonly lives in both a skill's SKILL.md and its scripts/*.py, so
results are aggregated by unique address with every occurrence (file:line +
snippet) and replaced together.
- Pure engine —
deerflow/skills/address_scan.py(harness, noapp.*):find_spans/scan_occurrences(URL regex with trailing-punctuation strip; IPv4 with 0–255 octet check + optional port; an IP inside a URL is not separately surfaced),apply_replacements(substring-safe: re-detects exact token spans and only rewrites a span whose text equals a requestedfrom, so1.2.3.4is never rewritten inside11.2.3.40),infer_kind,is_valid_replacement. Tested intests/test_skill_addresses.py. - Router —
app/gateway/routers/skill_addresses.py(/api/skill-addresses, admin only). Availability at scale (150+ skills): every filesystem walk/read/write runs viaasyncio.to_threadso the event loop is never blocked. Routes:GET /scan(aggregated, one-shot) +GET /scan/stream(per-skill NDJSON progress{type:start|progress|result});POST /preview(dry-run diff + validation warnings);POST /apply+POST /apply/stream(per-file NDJSON progress);GET /backups;POST /rollback/{backup_id}. Apply safety: optimistic lock viaexpected_total(409 if files changed since the scan), path-guard (is_relative_tothe skills root), auto disk backup to{skills_root}/.address-edits/{timestamp}/(the.-prefixed dir is excluded from scans), thenrefresh_skills_system_prompt_cache_async()so edits take effect immediately. Registered inapp.pyaftersensitive_words.router. Per-edit history: each apply also writes a sibling manifest{skills_root}/.address-edits/{timestamp}.manifest.json(best-effort) capturing{created_at, replaced_total, replacements:[{from,to,kind}], changed_files}; the sibling-not-child placement keeps it out of both the restore set and the file count.GET /backupsreturns each record enriched withreplaced_total+replacements(from the manifest; falls back to dir mtime when absent), andPOST /rollback/{backup_id}restores a single record — so the UI shows a full modification history and rolls back any one entry independently. - Frontend — the panel is the reusable
SkillAddressesPanel(exported frompages/SkillAddressesPage.tsx; its default export keeps the standalone route/page/strategy/admin/skill-addressesas a direct-link fallback). It is embedded as a「技能地址管理」tab inside the admin「技能管理」pagepages/AdminSkillsPage.tsx(Tabs: 技能列表 + 技能地址管理; reached from the 「技能管理」item in the「设置和更多」dropdownworkspace-nav-menu.tsx→/page/workspace/admin/skills) — it is no longer a separate entry in thecomponents/page-sidebar.tsxsettings dropdown. The panel: auto-scan on open with a progress bar (扫描逐技能 / 应用逐文件 X/N), URL/IP type filter, search, a「片段批量套用」helper (e.g. alluat.4.cn→prod.4.cn), per-address「替换为」input with expandable occurrences, inline 预览 panel, 应用, multi-keyword search (space-separated, all-must-hit), and a 修改记录 (modification history) panel — each apply is one record (time / which addresses changed from→to / file & replacement counts) with a per-record 回退此次 (single-record rollback, with inline confirm). Client + NDJSON stream parsing instrategy-components/api/skill-addresses.ts. Scope is the running skills only; thedeploy/minimal/skills/copy is not touched (sync separately).
Skill Compression (技能压缩)
When the skill catalogue grows large, injecting every enabled skill full-text into the lead agent's system prompt bloats it and wastes tokens. Skill compression is an admin-gated mode that keeps only 常驻(免压缩) skills in the prompt body and surfaces the rest on demand via keyword retrieval.
- Master switch —
system_settings.json → skill_compression(SkillCompressionSettings:enableddefault false,search_top_ndefault 5,keep_index_in_promptdefault true). Read/written byGET/PUT /api/system-settings/skill-compression(GET open to any user — the skill detail page reads it; PUT admin-only, refreshes the prompt cache). - Per-skill 常驻 flag — stored in the DB:
SkillRow.always_on(Boolean, migration20260616_03, idempotent SQL incolumn_additions_mysql.sql), not inextensions_config.json(the source-tree config file is read-only / unsafe to rewrite in some deployments, and built-in skills carry no file entry). Threaded throughSkillStorecreate/ensure_legacy/update/update_admin +list_always_on_names()exactly likename_zh/detail; surfaced onSkillResponse.always_onfrom the ownership record in_skill_to_response. Toggled viaPUT /api/skills/{name}/always-on(custom skill → owner or admin; built-in / public → admin only, materializes a row on demand viaensure_legacy). The prompt builder stampsSkill.always_onfromlist_always_on_names()via a sync DB bridge (_apply_always_oninlead_agent/prompt.py). - Prompt split —
get_skills_prompt_section(lead_agent/prompt.py) reads the compression settings (skipped for restricted custom agents — they already have a curated allowlist). When off, behaviour is unchanged (all enabled skills full-text). When on,always_onskills render as full<skill>blocks; the rest are dropped from the body, replaced by asearch_skillsusage note plus (whenkeep_index_in_prompt) a name-only<compressed_skills>index. The_get_cached_skills_prompt_sectioncache key includes thealways_onflag + compression flags so changes invalidate correctly. - Retrieval tool —
search_skills(tools/builtins/skill_tools.py,search_skills_impl) keyword-scoresname + descriptionover enabled, visible, non-archived skills (same visibility filter asskill_list) and returns top-N{name, description, category, location}. It is registered inget_available_toolsonly when compression is enabled (offline, deterministic, zero external deps). The agent thenread_files the returnedlocation.
Tests: tests/test_skill_compression.py.
Skill Curator (packages/harness/deerflow/skills/curator.py)
Background skill lifecycle manager. The Gateway lifespan starts a resident
scheduler task (_curator_scheduler_loop in app/gateway/app.py) that ? after a
5-minute startup delay ? calls maybe_run_curator() every hour. A cross-process
file lock (.curator_scheduler.lock) elects a single leader worker so the
scheduler runs once per deployment. Per run:
- Master switch ?
skill_evolution.enabled(defaultfalse) gates the whole feature;curator.enabled(defaultfalse) gates the curator specifically. Both must betruefor the curator to act.should_run_now()and_curator_thread_target()both enforce this ? so manualPOST /api/curator/runalso respects the master switch. Config isconfig.yaml.skill_evolutiondeep-merged with per-userusers/{user_id}/skill_evolution.json. - Per-user run lock ?
.curator_statecarriesrunning/started_at;_try_acquire_run_lock()prevents the same user's curator from running concurrently (scheduler vs. manual trigger, or double manual clicks). A lock whosestarted_atis older than_CURATOR_STALE_LOCK_HOURS(6h) is treated as process-crash residue and can be taken over. - Concurrency cap ?
maybe_run_curator()collects all due users and hands them to_dispatch_curator_runs(), which runs them through aThreadPoolExecutorcapped at_CURATOR_MAX_CONCURRENCY(5) ? avoids a thundering herd of LLM calls when many users come due at once (e.g. first deploy). - Failure back-off ?
_release_run_lock()always stampslast_attempt_at. When a first run keeps crashing (last_run_atnever set),should_run_now()retries on_CURATOR_RETRY_HOURS(6h) instead of every hourly tick. should_run_now()is the real gate: the hourly tick only checks "is the per-userinterval_hours(default 168h) elapsed". Timestamp parsing tolerates naive (tz-less) strings ? a corruptlast_run_attriggers a re-run rather than a permanent deadlock.- Published-skill filter ?
SkillStore.list_published_names()is a batch query; the curator fetches the published-name set once per run instead of one DB round-trip per candidate skill. - LLM review timeout ?
_run_llm_review()runsagent.invoke()in a worker thread bounded bycurator.review_timeout_seconds(default 600s). The model's per-requestrequest_timeoutcan't bound a multi-turn React agent loop, so on timeout the run is abandoned (a warning is written tocurator_log, visible in the Web UI) and the curator finishes cleanly + releases the run lock. The worker context is copied viacontextvars.copy_context()so tool calls still resolve the correctuser_id. - Review model ?
_run_llm_review()usescurator.model_name(aconfig.yamlmodels[]entry name) if set, else the primary model; it is editable from the Web UI evolution panel. The skill security scanner (scan_skill_contentinsecurity_scanner.py) reuses the sameskill_evolution.curator.model_namesetting ? review and moderation share one model. There is no separate moderation-model config field.
Model Factory (packages/harness/deerflow/models/factory.py)
create_chat_model(name, thinking_enabled)instantiates LLM from config via reflection- Supports
thinking_enabledflag with per-modelwhen_thinking_enabledoverrides force_disable_thinking=True(kwarg) hard-disables thinking for models that explicitly declaresupports_thinking: trueeven when the model config declares nowhen_thinking_*/thinkingshape — many internal vLLM/Qwen models only set that flag, sothinking_enabled=Falsealone injects nothing and they keep reasoning on. The flag injects the OpenAI-compatible disable shapes (extra_body.enable_thinking=false,extra_body.chat_template_kwargs.enable_thinking=false, andextra_body.thinking.type="disabled") only for classes that acceptextra_body. Models withsupports_thinking: false, especially generic OpenAI-compatible proxies such as LiteLLM, receive no vendor-specific thinking fields because those gateways may reject unknown parameters. MirrorsTitleMiddleware._force_disable_thinking. The lead agent reads the runtime flagthinking_force_disabledand forwards it (used by/api/open/chat)- Supports vLLM-style thinking toggles via
when_thinking_enabled.extra_body.chat_template_kwargs.enable_thinkingfor Qwen reasoning models, while normalizing legacythinkingconfigs for backward compatibility - Supports
supports_visionflag for image understanding models - Config values starting with
$resolved as environment variables - Missing provider modules surface actionable install hints from reflection resolvers (for example
uv add langchain-google-genai)
Roundtable Model Fault Tolerance (会商模型自动容错 / 换下一个模型)
会商第二步(leader/seat)+ 第三步(summary/report/dashboard + 结构化抽取)单轮执行时,某模型出问题就自动切到下一个模型继续,并告知用户「『X』模型异常 → 已切换『Y』」。前台实时(multi_agent.py)与后台挂起(engine.py + InProcessRoundtableGateway)两条路径都覆盖。
- 可检测的模型失败:
LLMErrorHandlingMiddleware在同一模型重试 3 次后不抛异常,而是返回兜底AIMessage—— 现在给它打additional_kwargs[LLM_ERROR_MARKER="deerflow_llm_error"]=reason(model_dump 保留,经 checkpointer 序列化仍在)。is_llm_error_message(msg)据标记(回退按已知兜底文案正文)识别,返回 reason。 - 候选模型链:
app/gateway/roundtable_model_fallback.py::fallback_model_chain(current)按config.yaml models[]顺序给出[当前模型, …其余](去重保序)。reason_text(reason)翻译成简短中文。总开关 envROUNDTABLE_MODEL_FALLBACK(默认开,0/false/off关 → 单模型旧行为);ROUNDTABLE_MODEL_FALLBACK_MAX(0=试完全部)。 - 触发换模型的情形:① 单轮 run 抛异常(start_run 失败 / idle 卡死 →
timeout,其余exception);② 兜底错误文案(is_llm_error_message);③ 交付无效(仅 seat/leader;leader 还含「既无正文又无派活」)。席位只认剥掉<think>后的可见正文——思考-only / 空串 / 断流都不算已交付(deerflow.agents.roundtable_orchestrator.delivery:visible_seat_text/classify_seat_delivery)。前台_special_run对此 yieldstatus:"invalid_delivery"且不广播;GET /last-ai-message回{delivered, content, length, invalid_reason}(思考-only 标thinking_only,总控席位清单标 ⚠️)。后台引擎只把可见正文写入deliveries,续轮用build_continue_prompt:有无效席位时强制重派,禁止声称「已完成交付」。Step3 单例(report/summary/dashboard/ingest/结构化)的产物是落盘文件,不把「正文为空」当失败,只对真模型错误换模型。 - 后台:
InProcessRoundtableGateway._run_turn_with_fallback(...)包住_run_agent_turn,逐候选模型试;换模型落model_switch诊断、全失败落model_fallback_exhausted。run_leader/run_seat/run_report/run_summary/run_dashboard_data全走它(report/summary/dashboard 的产物自愈重试保留,并记住已切到的可用模型用于后续续写轮)。各 Turn/Result 带回switches: list[ModelSwitch],引擎_emit_switches→JobProgressWriter.model_switch(...)往研讨时间线插一条role="system"系统提示气泡。 - 前台:
_special_run/_leader_runstream 收尾后检测失败 → yield{status:"model_switch", role, agent_name, failed_model, next_model, reason}帧并用下一个模型重新流式;终态帧带model_used。结构化抽取flow_extract.py的 round1model.astream也包了同款换模型循环(emitevent:"model_switch"帧)。 - Step2 单席位隔离:
engine._run_recommend/_run_chain/_run_dag的每个席位执行包 try/except —— 一个席位彻底失败(落seat_failed诊断 + 置空交付)不拖垮整批/整链/整个作业。 - 前端:
multi-agent.ts解析model_switch帧 →onModelSwitch回调(清空失败尝试的半截气泡 + 累计的 tool_call delta,phase 退回 connected);useStep2Orchestration弹 toast + 时间线插「系统提示」气泡;useStep3Summary/Report/Dashboard+flow-extract.ts重置对应气泡并在步骤卡留可见提示;后台路径经job-live.ts::toStep2Dialogue把role="system"渲染成 amber「系统提示」气泡。 - Tests:
tests/test_roundtable_model_fallback.py。
vLLM Provider (packages/harness/deerflow/models/vllm_provider.py)
VllmChatModelsubclasseslangchain_openai:ChatOpenAIfor vLLM 0.19.0 OpenAI-compatible endpoints- Preserves vLLM's non-standard assistant
reasoningfield on full responses, streaming deltas, and follow-up tool-call turns - Designed for configs that enable thinking through
extra_body.chat_template_kwargs.enable_thinkingon vLLM 0.19.0 Qwen reasoning models, while accepting the olderthinkingalias
IM Channels System (app/channels/)
Bridges external messaging platforms (Feishu, Slack, Telegram, DingTalk) to the DeerFlow agent via the LangGraph Server.
Architecture: Channels communicate with Gateway through the langgraph-sdk HTTP client (same as the frontend), ensuring threads are created and managed server-side. The internal SDK client injects process-local internal auth plus a matching CSRF cookie/header pair so Gateway accepts state-changing thread/run requests from channel workers without relying on browser session cookies.
Components:
message_bus.py- Async pub/sub hub (InboundMessage? queue ? dispatcher;OutboundMessage? callbacks ? channels)store.py- JSON-file persistence mappingchannel_name:chat_id[:topic_id]?thread_id(keys arechannel:chatfor root conversations andchannel:chat:topicfor threaded conversations)manager.py- Core dispatcher: creates threads viaclient.threads.create(), routes commands, keeps Slack/Telegram onclient.runs.wait(), and usesclient.runs.stream(["messages-tuple", "values"])for Feishu incremental outbound updatesbase.py- AbstractChannelbase class (start/stop/send lifecycle)service.py- Manages lifecycle of all configured channels fromconfig.yamlslack.py/feishu.py/telegram.py/dingtalk.py- Platform-specific implementations (feishu.pytracks the running cardmessage_idin memory and patches the same card in place;dingtalk.pyoptionally uses AI Card streaming for in-place updates whencard_template_idis configured)
Message Flow:
- External platform -> Channel impl ->
MessageBus.publish_inbound() ChannelManager._dispatch_loop()consumes from queue- For chat: look up/create thread through Gateway's LangGraph-compatible API
- Feishu chat:
runs.stream()? accumulate AI text ? publish multiple outbound updates (is_final=False) ? publish final outbound (is_final=True) - Slack/Telegram chat:
runs.wait()? extract final response ? publish outbound - Feishu channel sends one running reply card up front, then patches the same card for each outbound update (card JSON sets
config.update_multi=truefor Feishu's patch API requirement) - DingTalk AI Card mode (when
card_template_idconfigured):runs.stream()? create card with initial text ? stream updates viaPUT /v1.0/card/streaming? finalize onis_final=True. Falls back tosampleMarkdownif card creation or streaming fails - For commands (
/new,/status,/models,/memory,/help): handle locally or query Gateway API - Outbound ? channel callbacks ? platform reply
Configuration (config.yaml -> channels):
langgraph_url- LangGraph-compatible Gateway API base URL (default:http://localhost:8001/api)gateway_url- Gateway API URL for auxiliary commands (default:http://localhost:8001)- In Docker Compose, IM channels run inside the
gatewaycontainer, solocalhostpoints back to that container. Usehttp://gateway:8001/apiforlanggraph_urlandhttp://gateway:8001forgateway_url, or setDEER_FLOW_CHANNELS_LANGGRAPH_URL/DEER_FLOW_CHANNELS_GATEWAY_URL. - Per-channel configs:
feishu(app_id, app_secret),slack(bot_token, app_token),telegram(bot_token),dingtalk(client_id, client_secret, optionalcard_template_idfor AI Card streaming)
Memory System ? Memory V2 (packages/harness/deerflow/agents/memory/)
Pluggable memory architecture (Hermes-style). See docs/MEMORY_V2_DESIGN_ZH.md
and docs/HINDSIGHT_SETUP_ZH.md for the full design.
Components:
provider.py-MemoryProviderabstract base (12 lifecycle/hook methods)manager.py-MemoryManagerorchestrates builtin + ?1 external provider;build_memory_manager()is the construction entry pointproviders/builtin.py-BuiltinFileProvider: local USER.md / MEMORY.md, ? delimiter, char quotas, file locking, atomic writes, frozen snapshotproviders/hindsight.py-HindsightProvider: HTTP client to a self-hosted Hindsight Docker instance, dual-bank routingsecurity.py- injection/exfiltration scan + invisible-unicode block for memory contentmessage_processing.py-filter_messages_for_memory(last-turn extraction for Hindsight retain)
Two stores (binary classification):
USER.mdat{base_dir}/users/{user_id}/USER.md? user profile, shared across agentsMEMORY.mdat{base_dir}/users/{user_id}/agents/{agent_id}/MEMORY.md? agent-private working notesuser_idresolved viaget_effective_user_id();"default"in no-auth mode
Write paths:
- LLM actively calls the
memorytool (tools/builtins/memory_tool.py) ? writes USER.md/MEMORY.md, mirrors to Hindsight - LLM calls
hindsight_recall/reflect/retaintools (tools/builtins/hindsight_tools.py) ? only intools/hybridmemory_mode - Hindsight
auto_retain? every turn auto-stored to the work bank (viaMemoryMiddleware.after_agent)
Injection:
- builtin frozen snapshot ? injected into the system prompt via
lead_agent/prompt.py_get_memory_context - Hindsight dynamic recall ? injected per model call via
MemoryMiddleware.wrap_model_call(transient ? not persisted to state, so it never leaks to the frontend)
Hindsight dual-bank isolation (B-scheme):
- global user bank
deerflow-user-{user_id}? cross-agent - agent work bank
deerflow-work-{user_id}-{agent_id}? per-agent hindsight-clientis an optional dependency; when missing,HindsightProvider.is_available()returns False and the system falls back to builtin-only
Migration: PYTHONPATH=. python scripts/migrate_memory_to_v2.py [--all|--user-id X] [--dry-run] converts legacy memory.json into USER.md/MEMORY.md.
Configuration (config.yaml ? memory, nested): enabled, provider (builtin/hindsight),
builtin.{memory_char_limit,user_char_limit,...}, hindsight.{mode,api_url,api_key,bank_id_template,work_bank_id_template,memory_mode,...},
injection.{enabled,max_tokens,include_builtin,include_hindsight,context_tag}, security.{scan_content,block_invisible_unicode}.
Per-user runtime overrides (whitelisted fields) live at {base_dir}/users/{user_id}/memory_config.json ? see get_effective_memory_config().
Gateway API ? see the Memory router below (/api/memory/v2*).
Reflection System (packages/harness/deerflow/reflection/)
resolve_variable(path)- Import module and return variable (e.g.,module.path:variable_name)resolve_class(path, base_class)- Import and validate class against base class
Config Schema
config.yaml key sections:
models[]- LLM configs withuseclass path,supports_thinking,supports_vision, provider-specific fields- vLLM reasoning models should use
deerflow.models.vllm_provider:VllmChatModel; for Qwen-style parsers preferwhen_thinking_enabled.extra_body.chat_template_kwargs.enable_thinking, and DeerFlow will also normalize the olderthinkingalias tools[]- Tool configs withusevariable path andgrouptool_groups[]- Logical groupings for toolssandbox.use- Sandbox provider class pathskills.path/skills.container_path- Host and container paths to skills directorytitle- Auto-title generation (enabled, max_words, max_chars, prompt_template)summarization- Context summarization (enabled, trigger conditions, keep policy)subagents.enabled- Master switch for subagent delegationmemory- Memory system (enabled, storage_path, debounce_seconds, model_name, max_facts, fact_confidence_threshold, injection_enabled, max_injection_tokens)auth_login- Custom login modes layered on the base auth (deerflow/config/login_config.py):username_login.{require_password, password}- Shared password gate forPOST /api/v1/auth/login/username. Whenrequire_password=true, the (otherwise passwordless) username login additionally requires?password=<shared-secret>; emptypasswordfalls back to123ewq. Gate only ? it does not change which account logs in. Per-account override (账号级登录口令): independent of this shared gate, an admin can assign one specific account an individual plaintext password via the 白名单管理 page (users.login_passwordcolumn, migration20260625_04, nullable). When the target account carries alogin_password,/login/usernamerequires?password=to match it exactly and that personal password takes precedence over the shared gate (the shared secret will not unlock that account); accounts with no personal password keep the shared-gate / passwordless behavior. The check happens before auto-register via the lookup-only helper_lookup_user_by_username(so a brand-new username is unaffected). Surfaced/managed admin-only:GET /api/v1/auth/usersreturnsUserListItem.login_password(plaintext — the list is already admin-gated), andPUT /api/v1/auth/users/{user_id}/login-password{password: str|null}sets (non-empty) / clears (null·empty) it without touchingtoken_version. Plaintext storage is intentional (the password travels in the login URL query anyway, so it stays re-copyable). The user logs in via the existing#/login/<username>?password=xxxx. Tests:tests/test_login_personal_password.py.token_login.{enabled, token_info_url, service_authorization, timeout_seconds, verify_ssl, ca_cert_path, mock_tokens}- Token-exchange login (POST /api/v1/auth/login/token). The gateway POSTs totoken_info_urlwith the caller-supplied token placed directly in theAuthorization: Bearer <token>header (no cookie; the staticservice_authorizationfield is deprecated/unused, kept only for backward compat), readsdata.usernamefrom a{"state":"200",...}body, then signs the matching local user in (auto-registering if absent).token_info_urlis HTTPS ?verify_ssl: falseskips cert verification for self-signed/internal UAT hosts, or setca_cert_pathto a PEM bundle (takes precedence, keeps verification on).mock_tokens(offline test mapping{token: username}): when an incoming token matches a key,_fetch_username_from_tokenreturns the mapped username and skips thetoken_info_urlcall entirely — for environments where the upstreamgetTokenInfohost is unreachable (e.g. external networks). Empty in prod (no effect); seeded{"123ewq": "admin"}in this deployment so/login/任意?authToken=123ewqsigns in as admin.UsernameLoginResponsenow also returnsusername(the raw resolved login name, pre email-normalization) on both/login/tokenand/login/username, so the frontend can append a faithful?username=on jump links (按钮管理 追加登录态) instead of the lossy email local-part. See docs/AUTH_LOGIN.md for the full request/response contract.
Both endpoints are whitelisted in auth_middleware._PUBLIC_EXACT_PATHS. The shared _get_or_create_user_by_username() helper (in routers/auth.py) backs both /login/username and /login/token.
log_system- Admin-only external log-system entry (deerflow/config/log_system_config.py):{enabled, url, label}. Surfaced read-only byGET /api/system-settings/log-system(admin only); the frontend renders it as a "????" item in the bottom-right settings menu (workspace-nav-menu.tsx) that opensurlin a new tab. Disabled by default.
extensions_config.json:
mcpServers- Map of server name ? config (enabled, type, command, args, env, url, headers, oauth, description)skills- Map of skill name ? state (enabled). 常驻(免压缩)/always_on lives in theskillsDB table, not here.
Both can be modified at runtime via Gateway API endpoints or DeerFlowClient methods.
Conditional WeKnora LLMWiki
Ordinary knowledge-base cards expose an owner/admin-only name/description edit dialog in WeKnoraKnowledgeWorkspace.tsx, reusing useLlmWikiActions().update and PATCH /api/llmwiki/knowledge-bases/{id}. The route trims names, rejects blank or reserved conversation-deposit names, writes to WeKnora before updating local metadata, and retains existing write authorization. A successful remote list is authoritative for every list scope, including personal: absent remote IDs are hidden without deleting mappings/references. Upstream failures remain errors. The grid also filters cached remote_status=missing rows and no longer renders stale-record cleanup cards. Regression coverage: tests/test_llmwiki_knowledge_management.py.
llmwiki.weknora.api_base_url is the provider switch: non-empty enables WeKnora, empty preserves the legacy LLMWiki UI. WeKnora is independently deployed and owns PostgreSQL/Redis/object storage/vector storage/parsers/models; DeerFlow config contains only api_base_url and optional web_base_url. DeerFlow deliberately sends no WeKnora credential, so the WeKnora endpoint must be restricted to a trusted network and allow DeerFlow traffic itself or through its reverse proxy. There is intentionally no default_embedding_model_id: WeKnoraClient queries /api/v1/models and uses WeKnora's active Embedding model when creating a KB; Wiki creation additionally resolves an active KnowledgeQA model inside WeKnora for synthesis. The native BFF is app/gateway/routers/llmwiki.py; ownership/publication mappings live in deerflow.persistence.llmwiki; the browser only sees DeerFlow mapping IDs. Knowledge-base owners publish immediately (private -> published) without a review queue, and can unpublish ordinary bases. The conversation-deposit KB is system-owned, always published, shared by all DeerFlow users, excluded from My Knowledge, and immutable through user/admin detail operations; recognition requires both the reserved name and system owner. Existing same-name remote mappings are adopted in place as the system deposit, preserving mapping ids and document links. Its user-facing description, generated Markdown source and tags use cmzs branding; list/detail reads synchronize stale remote descriptions, and search results, source previews and embedded-detail JSON recursively replace legacy brand text in historical conversation-deposit payloads. Deposit requests must reference a thread accessible to the current account. Personal knowledge-base list/search views are scoped to the current DeerFlow account even for admins; frontend LLMWiki caches include the live account id, and admin write power does not make other users' private KBs appear as owned. Only system_role === "admin" sees the WeKnora connection-config tab. Knowledge cards navigate to a dedicated WeKnora-style detail route. Wiki-enabled bases render and default to Wiki -> documents -> graph; standard bases omit Wiki and render documents -> graph. The detail page supports sequential multi-file upload with an aggregate result for writable ordinary bases, hides it for public read-only/system-deposit bases, and uses a generic loading label while the embed session starts. It also supports URL and published manual-Markdown imports, browses real Wiki pages/content, proxies authenticated original-file previews, pages through /chunks/{knowledge_id}, and renders the Wiki page graph from /knowledgebase/{kb_id}/wiki/graph; document details use a right sheet rather than the removed KB modal. The credential-free proxy uses one WeKnora service identity, so per-user separation is enforced by DeerFlow mappings while all remote KBs physically share that identity's WeKnora workspace. WeKnora iframe detail pages first call /api/llmwiki/knowledge-bases/{id}/weknora/session through the normal DeerFlow API auth path, then load /platform/knowledge-bases/{weknora_id} with a signed deerflow_weknora_embed cookie for /platform, /assets, /locales, root assets, and API calls. Injected fetch/XHR rewriting sends all iframe /api/v1/*, including colliding /api/v1/auth/*, through /api/llmwiki/weknora-embed/api/v1/*; Gateway auth/CSRF bypass only applies to those signed iframe proxy requests, the proxy still revalidates the mapped KB, and cookie security honors X-Forwarded-Proto for HTTPS reverse proxies. llmwiki_search is conditionally registered in tools.py, is read-only, and rechecks the current DeerFlow user, agent llmwiki_knowledge_base_ids, and thread-context selection on every call. Source previews prefer the cleaned Wiki/LLMWiki page text when a retrieved chunk maps to a Wiki page, with raw split chunks as the fallback. Agent bindings use an inline WeKnora-style searchable/grouped multi-select in the standard agent create/detail UI; new chats use the same interaction pattern as a compact composer dropdown, and selections are written to thread metadata/context and restored on reopen. The DeerFlow frontend is native React and the WeKnora frontend is not modified or embedded. Full deployment and behavior notes: docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md. Tests: tests/test_llmwiki_weknora.py.
The embedded detail proxy hides WeKnora's outer .main > .aside_box application sidebar and expands the route outlet, but deliberately keeps the Wiki page's internal directory/index sidebar. Do not add an HTML sandbox attribute to this trusted, permission-filtered detail iframe: a sandboxed ancestor causes Chromium's built-in Blob PDF viewer to render the browser-blocked page. The Gateway follows upstream redirects only for /knowledge/{id}/preview and /knowledge-bases/{id}/files, keeping PDF bytes and protected Wiki images inside the signed proxy instead of sending the browser to WeKnora object storage.
Native/private and public Wiki readers must rewrite protected Markdown image
sources to their mapping-scoped DeerFlow /files?file_path= endpoints and keep
real HTTP URLs on rendered image elements. The embedded guard also rewrites
protected fetch/XHR values and hydrates late img[data-protected-src]
placeholders with the signed iframe HTTP proxy URL. Observe both
data-protected-src and src, because WeKnora may replace the URL after first
paint. Never create/revoke a temporary object URL here: its lightbox reuses the
thumbnail src after click. Never expose the service token or relax the
current-KB authorization to make images work.
Embedded native knowledge-base switching: The original WeKnora breadcrumb picker remains available inside the detail iframe. A signed 15-minute embed cookie carries only the current DeerFlow user's visible mappings (their own plus published bases); the proxied /api/v1/knowledge-bases response is filtered to that allowlist. Choosing another entry performs a full iframe navigation, rechecks the DeerFlow mapping and read/write permission, then reissues the cookie for the selected base. Never expose the upstream service-admin list directly or permit an arbitrary remote knowledge-base ID.
The independent enterprise-research workbench is temporarily disabled. Its source remains under app/gateway/routers/enterprise_research.py and deerflow.persistence.enterprise_research, but Gateway does not mount /api/enterprise-research, initialize its stores/executor, or register its ORM rows in Base.metadata. This ensures startup does not create or alter enterprise_research_tasks / enterprise_research_report_jobs on MySQL. It remains separate from /api/deep-research; do not re-enable it incidentally while changing the active Deep Research workflow. Regression coverage for its preserved implementation remains in tests/test_enterprise_research_tasks.py.
Embedded Client (packages/harness/deerflow/client.py)
DeerFlowClient provides direct in-process access to all DeerFlow capabilities without HTTP services. All return types align with the Gateway API response schemas, so consumer code works identically in HTTP and embedded modes.
Architecture: Imports the same deerflow modules that Gateway API uses. Shares the same config files and data directories. No FastAPI dependency.
Agent Conversation:
chat(message, thread_id)? synchronous, accumulates streaming deltas per message-id and returns the final AI textstream(message, thread_id)? subscribes to LangGraphstream_mode=["values", "messages", "custom"]and yieldsStreamEvent:"values"? full state snapshot (title, messages, artifacts); AI text already delivered viamessagesmode is not re-synthesized here to avoid duplicate deliveries"messages-tuple"? per-chunk update: for AI text this is a delta (concat peridto rebuild the full message); tool calls and tool results are emitted once each"custom"? forwarded fromStreamWriter"end"? stream finished (carries cumulativeusagecounted once per message id)
- Agent created lazily via
create_agent()+_build_middlewares(), same asmake_lead_agent - Supports
checkpointerparameter for state persistence across turns reset_agent()forces agent recreation (e.g. after memory or skill changes)- See docs/STREAMING.md for the full design: why Gateway and DeerFlowClient are parallel paths, LangGraph's
stream_modesemantics, the per-id dedup invariants, and regression testing strategy
Gateway Equivalent Methods (replaces Gateway API):
| Category | Methods | Return format |
|---|---|---|
| Models | list_models(), get_model(name) |
{"models": [...]}, {name, display_name, ...} |
| MCP | get_mcp_config(), update_mcp_config(servers) |
{"mcp_servers": {...}} |
| Skills | list_skills(), get_skill(name), update_skill(name, enabled), install_skill(path) |
{"skills": [...]} |
| Memory | get_memory(), reload_memory(), get_memory_config(), get_memory_status() |
dict |
| Uploads | upload_files(thread_id, files), list_uploads(thread_id), delete_upload(thread_id, filename) |
{"success": true, "files": [...]}, {"files": [...], "count": N} |
| Artifacts | get_artifact(thread_id, path) ? (bytes, mime_type) |
tuple |
Key difference from Gateway: Upload accepts local Path objects instead of HTTP UploadFile, rejects directory paths before copying, and reuses a single worker when document conversion must run inside an active event loop. Artifact returns (bytes, mime_type) instead of HTTP Response. The new Gateway-only thread cleanup route deletes .deer-flow/threads/{thread_id} after LangGraph thread deletion; there is no matching DeerFlowClient method yet. update_mcp_config() and update_skill() automatically invalidate the cached agent.
Tests: tests/test_client.py (77 unit tests including TestGatewayConformance), tests/test_client_live.py (live integration tests, requires config.yaml)
Gateway Conformance Tests (TestGatewayConformance): Validate that every dict-returning client method conforms to the corresponding Gateway Pydantic response model. Each test parses the client output through the Gateway model ? if Gateway adds a required field that the client doesn't provide, Pydantic raises ValidationError and CI catches the drift. Covers: ModelsListResponse, ModelResponse, SkillsListResponse, SkillResponse, SkillInstallResponse, McpConfigResponse, UploadResponse, MemoryConfigResponse, MemoryStatusResponse.
Development Workflow
Test-Driven Development (TDD) ? MANDATORY
Every new feature or bug fix MUST be accompanied by unit tests. No exceptions.
- Write tests in
backend/tests/following the existing naming conventiontest_<feature>.py - Run the full suite before and after your change:
make test - Tests must pass before a feature is considered complete
- For lightweight config/utility modules, prefer pure unit tests with no external dependencies
- If a module causes circular import issues in tests, add a
sys.modulesmock intests/conftest.py(see existing example fordeerflow.subagents.executor)
# Run all tests
make test
# Run a specific test file
PYTHONPATH=. uv run pytest tests/test_<feature>.py -v
Running the Full Application
From the project root directory:
make dev
This starts all services and makes the application available at http://localhost:2026.
All startup modes:
| Local Foreground | Local Daemon | Docker Dev | Docker Prod | |
|---|---|---|---|---|
| Dev | ./scripts/serve.sh --devmake dev |
./scripts/serve.sh --dev --daemonmake dev-daemon |
./scripts/docker.sh startmake docker-start |
? |
| Prod | ./scripts/serve.sh --prodmake start |
./scripts/serve.sh --prod --daemonmake start-daemon |
? | ./scripts/deploy.shmake up |
| Action | Local | Docker Dev | Docker Prod |
|---|---|---|---|
| Stop | ./scripts/serve.sh --stopmake stop |
./scripts/docker.sh stopmake docker-stop |
./scripts/deploy.sh downmake down |
| Restart | ./scripts/serve.sh --restart [flags] |
./scripts/docker.sh restart |
? |
Nginx routing:
/api/langgraph/*? Gateway embedded runtime (8001), rewritten to/api/*/api/*(other) ? Gateway API (8001)/(non-API) ? Frontend (3000)
Running Backend Services Separately
From the backend directory:
# Gateway API
make gateway
Direct access (without nginx):
- Gateway:
http://localhost:8001
Browser CORS
The Gateway always permits local Vite origins on ports 5173, 5174, 3000, and
8080 for both localhost and 127.0.0.1 with credentials. A deployed
frontend on another origin must be listed exactly in GATEWAY_CORS_ORIGINS
(for example http://47.88.25.99:7010); explicit origins are merged with the
local development origins. Regression: tests/test_taskcop_cors.py.
Frontend Configuration
The frontend uses environment variables to connect to backend services:
NEXT_PUBLIC_LANGGRAPH_BASE_URL- Defaults to/api/langgraph(through nginx)NEXT_PUBLIC_BACKEND_BASE_URL- Defaults to empty string (through nginx)
When using make dev from root, the frontend automatically connects through nginx.
Key Features
Administrator-managed custom themes
system_settings.json → appearance now stores custom_themes alongside
the built-in default theme. A row is CustomThemePreset (id, name,
six-digit primary, optional primary-button text_color, mode, overrides, enabled,
sort_order). overrides is intentionally a short interaction-only list:
button_hover, navigation_gradient_start, navigation_gradient_end,
selected_background, border, muted, muted_text, destructive,
brand_title, and primary_foreground. The navigation gradient endpoints
default from primary and are only stored when an administrator overrides them.
All values must be six-digit hex strings; the frontend derives every remaining
semantic CSS/TDesign token so admins do not maintain a full raw token sheet.
Ids must be custom-theme-*, names and ids are unique, and the router limits
the list to 12. GET/PUT /api/system-settings/appearance is the single
read/write surface: reads are available to all users, writes require admin.
The router re-orders rows deterministically and only permits an enabled custom
theme as default_theme; removing or disabling the default safely falls back
to light. Regression coverage: tests/test_appearance_theme.py.
Compact top-bar shortcuts live in the same appearance setting. Their route
parameters support login, theme, and keyless path. A path parameter has
no query-key input; the operator selects its value source from the right-hand
dropdown and the frontend appends the resolved, URL-encoded value to the
external URL path (or Hash route) in configured order. Password stays an
ordinary login query parameter with the key password, so an external login
link can become #/login/<user>?password=…. Regression coverage:
tests/test_appearance_theme.py.
File Upload
Multi-file upload with automatic document conversion:
- Endpoint:
POST /api/threads/{thread_id}/uploads - Supports: PDF, PPT, Excel, Word documents (converted via
markitdown) - Rejects directory inputs before copying so uploads stay all-or-nothing
- Reuses one conversion worker per request when called from an active event loop
- Files stored in thread-isolated directories
- Agent receives uploaded file list via
UploadsMiddleware
See docs/FILE_UPLOAD.md for details.
Plan Mode
TodoList middleware for complex multi-step tasks:
- Controlled via runtime config:
config.configurable.is_plan_mode = True - Provides
write_todostool for task tracking - One task in_progress at a time, real-time updates
See docs/plan_mode_usage.md for details.
Context Summarization
Automatic conversation summarization when approaching token limits:
- On/off switch is per-request, not the config file. The chat input box has a
「上下文压缩」 toggle (below 参考文献, default off) → frontend sends
context.summarization_enabled→ whitelisted in_CONTEXT_CONFIGURABLE_KEYS(app/gateway/services.py) →_build_middlewaresreads it from the run config and calls_create_summarization_middleware(force_enabled=summarization_enabled).force_enabled=Truemounts the middleware even whenconfig.summarization.enabledis false; when the flag is absent (channels / scheduled runs) it falls back to the config flag.config.summarization.enablednow defaults to false. config.yamlsummarizationstill supplies the tuning: trigger types (tokens, messages, or fraction of max input) and the keep policy.- A compressed request must retain the latest non-empty
HumanMessage, even when it has already been folded into the hidden summary. This keeps strict OpenAI-compatible gateways from rejecting a tool-continuation call withNo user query found in message. - Keeps recent messages while summarizing older ones.
Tests: tests/test_summarization_toggle.py (gating + whitelist),
tests/test_summarization_middleware.py (cache behavior).
See docs/summarization.md for details.
Vision Support / Image Recognition (??)
Two paths, selected by config.yaml ? vision:
Dedicated vision model (delegated recognition) ? set vision.model_name to a model from models[]. The view_image tool sends the image to that model and returns its text description as the tool result, so a text-only main chat model can still "see" images. view_image accepts an optional question argument for targeted queries; the default prompt is vision.prompt (overridable), and vision.max_tokens caps the recognition output. Resolution/gating lives in deerflow/models/vision.py (resolve_vision_model_name, vision_enabled, is_delegated_vision, recognize_image).
Native vision (fallback) ? when vision.model_name is unset, behavior falls back to the main chat model, which only works if it has supports_vision: true:
ViewImageMiddlewareinjects base64 image data into the conversation before the LLM callview_imagestores the image instate.viewed_imagesfor that middleware
Both the view_image tool and ViewImageMiddleware are registered whenever vision_enabled(model_name) is true (dedicated model configured or main model supports vision). Config schema: deerflow/config/vision_config.py (VisionConfig); tests: tests/test_vision.py.
Upload symlink-safety (cross-platform)
deerflow/uploads/manager.py:open_upload_file_no_symlink writes uploads without following a pre-existing destination symlink. O_NOFOLLOW is applied where available (POSIX) but is optional ? Windows lacks os.O_NOFOLLOW/O_NONBLOCK, so requiring it previously made every Windows upload fail with UnsafeUploadPathError. Pre-existing symlinks are still rejected via lstat, and the opened fd is re-checked with fstat. O_BINARY is added on Windows. Regression: tests/test_uploads_windows_safe.py.
Code Style
- Uses
rufffor linting and formatting - Line length: 240 characters
- Python 3.12+ with type hints
- Double quotes, space indentation
Documentation
WeKnora iframe behind /deerflow
The LLMWiki WeKnora detail iframe is publicly served through the Gateway at
/deerflow, independently of the frontend deployment prefix (for example
/magentweb). Register llmwiki.proxy_router with the /deerflow prefix
before its unprefixed fallback router, otherwise the unprefixed catch-all
captures /deerflow/*. HTML root-relative assets, API strings, and Vite's
__vite__mapDeps table are rewritten for /deerflow; generic JavaScript
regular-expression literals must not be rewritten. Auth and CSRF bypass only
applies when the short-lived signed WeKnora embed cookie is valid. Keep static
path recognition in app/gateway/weknora_embed.py synchronized with the
proxy's static fallback paths in app/gateway/routers/llmwiki.py.
Position-roundtable role directory
position_roles is the global shared directory for position-roundtable work
views. It is deliberately separate from organisation/RBAC positions.
app.gateway.deps.langgraph_runtime() seeds the four stable built-in ids from
app/gateway/position_role_defaults.py via ensure_seed() (missing rows only;
never overwrite operator edits). The API is GET/PUT /api/position-roles;
writes are admin-only. Exactly one enabled intent role is required and owns
the initial intent workspace plus the summary/action-plan closing agents. Do
not delete a built-in id because historical chain nodes persist those ids;
custom role ids may be added and removed.
Position-roundtable task-scoped session history
position_roundtable_sessions.external_task_id is optional. When populated from the frontend route taskId, new rows set task_scoped = true: a task can have multiple sessions and all users opening that task see the same sessions, nodes and persisted conversation snapshots. When it is NULL/empty, keep the original user_id-scoped personal workspace behavior. 20260808_01 adds task_scoped with default false (so all historical records remain personal) plus the task/update-time index; it must not backfill or make external_task_id non-null.
Roundtable artifact submission status
roundtable_artifact_submissions keeps task-level submission state for a
generated deliverable and is deliberately separate from task details and any
external task-system delivery. The API is
GET /api/roundtable/tasks/{taskId}/artifacts?submitted=true and
PUT /api/roundtable/tasks/{taskId}/artifacts/submission. All callers use the
normal Gateway Bearer authentication; records are shared by task while
created_by/submitted_by/cancelled_by are audit-only. The server derives a
SHA-256 idempotency key from taskId + sourceKey + path + revision: submitting
the same revision is an upsert, while cancellation only flips submitted and
sets cancelled_at (it never deletes history). Store only the sandbox file
reference (threadId, path, source/node/session fields, name/MIME and revision)
and timestamps; do not persist Markdown content or call a formal delivery
interface. Previews continue to read the DeerFlow artifact by threadId + path.
See docs/ directory for detailed documentation:
- ../README.md - Offline Docker wheelhouse preparation and startup workflow. The pinned versions in
../offline-runtime-requirements.txtand../start_backend.shmust stay synchronized. Startup checkscroniter,hindsight-client, the aiomysql compatibility bundle, and Redis, installing missing or incompatible packages from wheelhouse only. The bundle installsaiomysql==0.3.2,PyMySQL==1.2.0, andSQLAlchemy==2.0.51together; the startup probe also verifies that SQLAlchemy's adaptedping(reconnect=False)signature is present, preventing the old-adapter missing-reconnectfailure. The download target defaults to CPython 3.12 on Linux x86_64 and verifies dependency closure using wheelhouse only. - CONFIGURATION.md - Configuration options
- ARCHITECTURE.md - Architecture details
- API.md - API reference
- SETUP.md - Setup guide
- FILE_UPLOAD.md - File upload feature
- PATH_EXAMPLES.md - Path types and usage
- summarization.md - Context summarization
- plan_mode_usage.md - Plan mode with TodoList
- BACKEND_RUN_STOP_ZH.md - Windows-first backend start/stop and troubleshooting guide
Deep Research retrieval
deerflow.agents.deep_research.adapters.material_provider.DeerFlowMaterialProvider
must remain on the Q&A retrieval path. Enabled deep-search, web-research,
and knowledge-base skills run through the existing SubagentExecutor skill
runtime only when their required Q&A tool is registered; then the provider may
use configured web_search. A disabled placeholder endpoint is not a valid
search channel: use the maintained keyless DuckDuckGo provider and a relevant
real-web fallback as read-only sources, keep source provenance, and never
enable or call a placeholder just to make the Deep Research UI look configured. Regression coverage is in
tests/test_deep_research_material_provider.py.
When a queued Deep Research job is cancelled, the Gateway must also clear the
owning session's active_job_id and mark it cancelled; otherwise the stale
session incorrectly consumes that user's concurrency limit.
Report prose is the one Deep Research LLM path that must use
DeerFlowCompletionBackend.stream_complete() rather than complete(): it
drives model.astream(), sends native report_delta frames through
ReportChunkEmitter, and persists bounded report_chunk checkpoints. The
app-layer DeepResearchLiveHub fans out those live frames inside one Gateway
worker only; the database event log still owns ?after=<seq> reconnect replay
and multi-worker recovery. Inject the hub callback into PersistingEventSink
and never import app.* from deerflow.*. Keep structured plan/review calls
on complete(). Before the search fan-out, basic mode emits one
queries_planned event and detailed/multi-agent modes emit one per subtopic;
they are durable user-visible search terms, not chain-of-thought. On completion, retain
the terminal lastJobId in usage_snapshot so a refreshed conversation page
can replay its trace from the event log. Detailed mode uses stable targets (introduction,
section:N, conclusion) because section writers are concurrent;
multi-agent revisions must emit report_reset before their next report
version. Do not persist raw model reasoning or one persisted event per provider
token. Models that expose thinking inside ordinary <think>...</think> text
must be split statefully before a report delta, checkpoint, or Markdown
projection is created; an optional live-only report_thinking frame can drive
the message-list thought trace, while the sandbox receives answer Markdown
only. The post-report digest (summarize_report) must likewise forward
provider reasoning through live-only summary_thinking frames so the
chat timeline can render thought while waiting for the first summary_delta.
The Vite page starts Word downloads through
frontend-web/src/core/scheduled-tasks/markdown-docx.ts, but the Gateway owns
generation through POST /api/writing/export/docx. Keep the standard-library
OOXML builder and its embedded-font validation; do not add a Python document-
conversion dependency merely to export text.
All non-streaming Deep Research control-plane calls (complete() for query /
section / plan / review JSON) run through asyncio.wait_for using the frozen
DeepResearchConfig.control_plane_timeout_seconds (45 seconds by default,
clamped to 120). Do not apply that total-duration cap to stream_complete():
reports may legitimately be long but must continue emitting visible deltas. A
basic query-planning timeout/failure emits query_planning_fallback and searches
the original topic; detailed and multi-agent plan creation similarly fall back
to their original topic/section, rather than leaving a durable job stuck in
planning. Tests: tests/test_deep_research_llm_adapter.py and
tests/test_deep_research_runners.py.
The per-user Deep Research concurrency guard counts only session statuses
running and awaiting_input. A draft has no durable job and must never
consume a slot; otherwise saved drafts make POST /sessions/{id}/jobs return
429 permanently. Regression: test_saved_draft_sessions_do_not_consume_research_concurrency.
POST /api/deep-research/sessions/{id}/jobs is retry-safe under MySQL
read/write splitting. Deep Research session/job/event/source/message writes
must serialise their changed row while the write transaction is still open
(flush for inserts, select-after-update before commit), never through a
post-commit session.refresh() or a fresh pooled read. Replica lag otherwise
turns a committed job into a misleading 500, can make dispatcher claim_job()
lose its snapshot, or can fail report/follow-up persistence after the model has
already completed. Always nudge the dispatcher for both a newly created job and
an idempotently reused queued job; otherwise a retry after such an interrupted
response returns the old job but leaves it waiting for the periodic scan. The
SSE replay generator must catch temporary event/job-store reads, emit a
recoverable warning, and keep the connection open so the next poll can repair
the cursor instead of surfacing an ASGI ExceptionGroup. Regressions:
test_deep_research_writes_never_depend_on_post_commit_refresh,
test_job_create_never_reads_back_after_commit,
test_retry_of_existing_queued_job_nudges_dispatcher, and
test_stream_stays_open_when_durable_replay_has_a_transient_failure.
At ordinary report completion, persist and confirm the session's
report_markdown projection before marking the Job completed. Retry a
short transient projection failure; if it cannot be confirmed, let the job
become failed rather than reporting a report that vanishes on refresh.
test_report_is_persisted_before_job_is_marked_completed and
test_report_projection_failure_never_marks_job_completed cover the ordering.
The chat-collection frontend must keep report-writing progress in the existing
workspace MessageList / MessageGroup tool-step renderer: inject only
virtual deep_research_progress tool calls for the durable phases, keep the
completed phase visible after the job ends, and append the digest as a normal
assistant message. Do not recreate a second deep-research-only conversation
timeline. On summarizing, insert that same normal assistant message with its
streaming marker and update it for every SSE summary_delta; when the report
completes, replace its content in place with the final completion/digest
message rather than waiting to append a summary. Completed chat-collection
sessions also inject a normal present_files card into the transcript, named
from the report Markdown H1 rather than its internal report.md path.
Replaying a completed history session must not auto-open its
sandbox; opening that card selects its virtual write-file report instead, and
its file-card download action delegates to the Deep Research Markdown/Word
export dialog. The default in-job concurrency and per-user active-job cap are
both three. Before the user confirms the report configuration, derive its
displayed material count from unique collector tool-result rows because the
durable Deep Research source store is populated when the writing job harvests
that thread. The Deep Research composer model picker must use the installed
TDesign Select backed by /api/models; its explicit choice is sent as
model_name on collector turns and applied to fast_model, smart_model, and
strategic_model when creating a session or starting a report job. Disable it
while collection or report writing is active so one durable run cannot switch
models mid-stream.
Report structures (/api/report-structures) may store an optional
retrieval_directions list (检索方向) and an optional retrieval_skills
list (检索来源 / 指定技能). Empty / unset keeps today's collector
query-planning, default tools, and runner sub-query generation. A non-empty
directions list is frozen into DeepResearchConfig.retrieval_directions: the
chat collector's first user message gets a <retrieval-directions> block that
asks it to split concrete retrievable terms along those facets (never search
the direction labels themselves), and legacy basic/detailed/multi_agent/
deep runners pass the same constraint into query-planning instead of using
课题 + 方向 as the search string. A non-empty skills list is frozen
into DeepResearchConfig.retrieval_skills: the chat collector's first user
message gets a <retrieval-skills> block, the collector run context receives
extra_skills so skill_view/skill_list can see those names, and legacy
material collection runs those skills instead of the default Q&A research-skill
trio. Helpers live in deerflow.agents.deep_research.retrieval.
Completed-report Q&A has two endpoints: retain POST /sessions/{id}/chat as the
legacy JSON response for compatibility, and use POST /sessions/{id}/chat/stream
for report-grounded Q&A SSE. The workbench must send a strong scoped edit
instruction (for example “修改第一段” / “重新生成一下第一段” / “重写第二节”) to
the separate POST /sessions/{id}/rewrite/stream report-section writing pipeline.
That endpoint requires a deterministic current-report range (otherwise 422), so
its output is always a replacement candidate rather than a Q&A answer. The
candidate is streamed in the ordinary message row
while the sandbox preserves the current report. When generation ends, inject a
confirmation card. Only the owner’s explicit
POST /sessions/{id}/messages/{messageId}/rewrite-proposal action apply may
persist the assembled report_markdown; dismiss preserves the original. The
proposal is bound to a hash of the report it was generated from, so stale
candidates fail with 409 rather than overwrite a later edit. Ambiguous or
ordinary evidence questions must remain non-mutating, and cancellation or a
stream error must never save a partial report. The rewrite prompt receives only
the selected range plus the selected evidence set, never a new web search.
DeerFlow-local LLMWiki Wiki index
llmwiki.local_wiki_index is an opt-in local-vector subsystem. WeKnora is the
Wiki generation/sync source and the Wiki-only fallback while an entire mapping
has not completed vectorization. The readiness gate requires state=idle, a
successful full scan, matching embedding fingerprint, no failed pages,
ready_page_count == local_page_count, and vectors for a non-empty library.
Ready mappings read DeerFlow's llmwiki_wiki_pages /
llmwiki_wiki_vectors / llmwiki_wiki_sync_states tables; incomplete mappings
call only WeKnora /knowledgebase/{id}/wiki/pages?query=... plus Wiki-page
detail. Never call /knowledge-search from internal search, RAG, or the
llmwiki_search tool, and never return raw chunks as the fallback. Core code lives under
deerflow.integrations.weknora.local_index and
deerflow.persistence.llmwiki_index; preserve the harness → app import
firewall. Gateway composition injects the store/search/sync services through
local_index.runtime. enabled turns on local vector search and auto_sync
defaults to the lease-protected background scheduler. The scheduler performs
an immediate incremental scan at startup, then rescans every configured
interval so new WeKnora libraries and asynchronously generated Wiki pages are
picked up without a manual vectorization click. Unchanged content must reuse
existing vectors.
Vectors are normalized little-endian float32 BLOBs. A page mirror and all of
its vectors are replaced in one transaction only after strict batched
embedding succeeds. Never delete the old generation before a remote embedding
call and never fall back to WeKnora raw chunks while the local feature flag is
enabled. NumPy matrices are per-mapping LRU snapshots invalidated by
index_revision; large decode/matrix work runs through asyncio.to_thread.
embedding.dimensions is a local response-size invariant and fingerprint
input, not an outbound API field: fixed-dimension BGE/OpenAI-compatible
endpoints can reject a dimensions request parameter with HTTP 400. When the
value is omitted, the embedding client probes and freezes the first response's
dimension before sync lease/rebuild decisions.
Embedding errors must retain the upstream status and bounded response body in
logs, without exposing API keys. SQL vector writes are split into statements
of at most database_batch_size rows (default 50) inside the same atomic page
replacement transaction.
During a fingerprint rebuild the mapping uses the Wiki-page API until every
page is successfully rebuilt. WikiRetrievalService partitions mixed requests
per mapping, so ready libraries stay local while incomplete libraries use that
Wiki-only fallback.
The frontend must keep Wiki citations visibly distinct from raw-document
citations while hiding provider-brand wording from user-facing labels, status
text, and errors. Any click on a Wiki source card opens the shared right-side
source drawer, which renders the complete Markdown page. When the local index
is disabled, the drawer reads the remote WeKnora Wiki-page endpoint directly
and must not surface the local-index-disabled response. A Wiki target must not
also load or render the legacy original-document preview, even when its citation
still carries a document id. The drawer supports single-page Markdown and Word
downloads, and its Open in Wiki link carries both the DeerFlow knowledge-base
mapping id and the exact wikiSlug. Wiki search-result cards use this same
drawer instead of navigating immediately. Exact-page navigation uses the native
knowledge/{mapping_id}/wiki/view?wikiSlug=... reader, not the WeKnora detail
iframe. Keep legacy knowledge/{mapping_id}?tab=wiki&wikiSlug=... links fast by
redirecting them in LlmWikiKnowledgePage before its runtime-loading gate.
Drawer-to-reader navigation stays in the current SPA and passes the fetched
Wiki page in router state so the article can render before the background query
refresh completes. The native reader must retain a searchable/collapsible left
Wiki directory, reconstruct hierarchy from parent_slug, wiki_path,
category_path, and slug fallback, expand/highlight the active page, and keep
directory navigation on the same lightweight route rather than mounting the
WeKnora iframe.
Generated HashRouter URLs must retain window.location.pathname (for example
the /web/ reverse-proxy prefix); never emit root-absolute /#/... citation,
reader, or public-embed links. Protected Wiki image rendering must use real HTTP
URLs, never URL.createObjectURL. Public readers use the published file proxy;
authenticated readers obtain a short-lived signed capability from
/files/access-url and load /api/public/llmwiki/files/{token}. The signed
payload is restricted to kind=wiki_file, mapping id, remote id, validated
protected path, and expiry, and responses provide an inline UTF-8 filename for
new-tab viewing and browser Save As. The proxied WeKnora detail iframe must
likewise leave its scoped HTTP file-proxy URL on every protected image and
restore it if upstream rendering replaces src; its click preview reads that
same DOM value, so a fetched-and-revoked object URL is not acceptable.
Third-party pages embed published Wiki content through the chrome-free route
/#/embed/knowledge/{mapping_id} (optional exact-page wikiSlug). It defaults
to a recognized Wiki index page (title 索引/index, page type index, or an
index/home/readme slug) and promotes that page to the first public-directory
entry. When upstream pages exist without an explicit index shell, the frontend
injects a virtual __index__ page and builds its catalogue from the page list;
an all-empty dynamic index response must also fall back to that page list. It retains
directory navigation and supports the shared public-embed theme / bg /
text appearance inputs. WeKnora dynamically renders index page-type counts,
nested category headings, page summaries, and data-slug links even though the
stored index Markdown only contains its introduction; the native public reader
reads the same ordered groups through the published-only
/api/public/llmwiki/knowledge-bases/{id}/wiki/index proxy and renders them with
buildWikiIndexCatalogMarkdown. The Gateway paginates every WeKnora index type
(summary, entity, concept, synthesis, comparison) to completion while
preserving upstream group/item order. Index Markdown
may also contain WeKnora shell-only placeholder links;
resolveWikiIndexPageLinks matches their visible labels to
the loaded page title, slug, or alias and rewrites them to exact embedded
wikiSlug routes. Preserve WeKnora's wiki-content-link interaction: primary
brand color, medium weight, dashed underline becoming solid on hover, focus
outline, and in-reader SPA navigation rather than a page reload. GET /api/public/llmwiki/knowledge-bases is the anonymous discovery endpoint. It
validates published ordinary mappings against WeKnora, drops stale mappings,
and returns safe comprehensive metadata from the WeKnora list/detail/page APIs:
document/chunk/processing/share counts, non-secret configuration and timestamps,
Wiki page type/status counts and explicit/generated index entry, plus aggregate
local-vector state/counts/timestamps. Never expose remote WeKnora ids,
tenant/creator identity, storage credentials, API credentials, raw vectors, or
per-page indexing errors. All anonymous reader calls use that
public namespace and the /knowledge-bases/{id} Wiki-page children. Those
endpoints must resolve with a synthetic anonymous user and is_admin=False,
explicitly require publication_status=published, and hard-deny the system
conversation-deposit mapping. Never point this page at the authenticated
/api/llmwiki reads or weaken the published-only gate merely because the
mapping id is hard to guess.
Owners manually force a full sync/vectorization through POST /api/llmwiki/knowledge-bases/{id}/vectorize; status is available through GET /api/llmwiki/knowledge-bases/{id}/wiki-index-status. That response includes
aggregate progress and per-page pages details; current-scan progress is
derived from active_sync_id / last_seen_sync_id, while each page exposes its
index state, vector count, timestamps, and bounded error. The Wiki management
page must show the progress bar and a status badge beside every article. It
also exports all remote processed Wiki pages via GET /api/llmwiki/knowledge-bases/{id}/wiki/export.xlsx. Directory export must keep
the full path, current directory, parent path, and dynamic per-level columns;
long Markdown is split into Excel-safe continuation columns instead of being
truncated. Workbook creation runs in asyncio.to_thread.
Internal search authorization always starts from DeerFlow mapping ids and is
rechecked by the Gateway. External endpoints use the isolated
/api/external/llmwiki/wiki namespace and X-API-Key; require published +
external-enabled + index-enabled on every request and response, and hard-deny
the system / 对话沉淀 mapping. Those API-key endpoints never return remote
WeKnora ids, provenance refs, credentials, or vector bytes.
POST /api/knowledge/vector-search is a deliberately separate anonymous API.
It has no token/API-key requirement or rate limit, applies route-scoped
Access-Control-Allow-Origin: *, searches every published local Wiki index
regardless of external_search_enabled, and still hard-denies the system /
对话沉淀 mapping before and after search. Its contract is section-level (one
result per matched vector, so Wiki-page fields and URLs may repeat), default
top_k=10, with optional include_page_content. Unlike the protected legacy
external namespace, this endpoint intentionally returns DeerFlow/WeKnora ids
and raw source/chunk provenance, but never credentials or vector bytes. Wiki
links use llmwiki.local_wiki_index.frontend_base_url (development default
http://127.0.0.1:5174) and point to the native
/#/page/workspace/knowledge/{mapping_id}/wiki/view?wikiSlug=... route.
Regression tests:
tests/test_llmwiki_local_index_repository.py,
tests/test_llmwiki_wiki_sync.py, tests/test_llmwiki_vector_search.py,
tests/test_external_llmwiki_search.py,
tests/test_public_knowledge_vector_search.py, and
tests/test_llmwiki_wiki_retrieval.py.
The sandbox whole-report rewrite may carry an optional reportOutline Markdown
template from the shared report-structure library. Persist the original template
in the durable input snapshot and append it as an explicit chapter-structure
constraint before requirements planning and prose generation; a template must
never be used as permission to invent facts absent from the current report.
DocumentRewriteVersionRepository.create backs the reversible before-image for
both artifact and whole-report rewrites. It must flush and serialize the new row
on the writer before commit, and must never call session.refresh() after
commit: a read/write-splitting MySQL endpoint can route that refresh to a
lagging replica, falsely report a version-save failure, and trigger the caller's
intentional restore of the rewritten document.
Version snapshots are an undo enhancement, not a condition for preserving an
already committed rewrite. If snapshot creation fails after the artifact/report
write, log the exception, complete the rewrite with
versionSnapshotFailed: true, and do not emit a user-facing failure or restore
the original. The sandbox uses that signal only to enable its ordinary manual
save fallback.
Offline local-sandbox Office runtime
The DeerFlow-only offline deployment uses LocalSandboxProvider, so document
dependencies belong in the backend image rather than a separate sandbox image.
deploy/linux-docker-offline/Dockerfile.backend installs python-docx,
LibreOffice Writer/Calc/Impress, Pandoc, Poppler, and CJK fonts in addition to
the MarkItDown Office extras. file_conversion.py normalizes legacy
DOC/XLS/PPT and ODT/ODS/ODP through a per-conversion LibreOffice user profile
before MarkItDown extraction, avoiding concurrent profile locks. The image
build and OFFLINE_INSTALL.sh --office-check both run
scripts/office_runtime_selfcheck.py; that smoke test creates and re-parses
minimal DOCX/XLSX/PPTX files and does not need the application database.
Skill/assistant knowledge invariants
Skill distillation is implemented in deerflow.skill_knowledge and
deerflow.persistence.skill_knowledge; native assistant bases are in
deerflow.assistant_knowledge and
deerflow.persistence.assistant_knowledge. Gateway composition, remote adapters,
the lease dispatcher, and APIs remain in app.gateway so the harness-to-app
import firewall is preserved.
Never execute a skill or read/upload/summarize its code and script files.
Knowledge fingerprints exclude those files and Markdown code fences. All
resyncs must stage and validate the knowledge digest before activation; the old
snapshot remains authoritative on failure. A binding may have only one active
item (active_dedupe_key plus atomic lease claims). Skill distillation targets
only writable WeKnora Wiki knowledge bases in this product version; reject direct
assistant-base targets, non-Wiki/FAQ/conversation/readonly targets, and direct
Neo4j writes. The create/list/resync/retry/detach flow is available to ordinary
users for their own visible skills, writable knowledge bases, jobs and bindings;
administrator-only auto-sync/review APIs remain privileged. Detach/delete operates on source contributions, not canonical rows,
and shared entities/relations are retired only when their final contribution is
gone.
Assistant knowledge has one global full-library base (知识梳理总库) and any
number of admin-created custom bases. Ordinary users have read/search access and
administrators own all mutations. Assistant bases accept local documents independently
of WeKnora via /api/assistant-knowledge/imports/files, as well as WeKnora export
packages or imports triggered from a WeKnora mapping. Local documents use the
Gateway assistant_file_ingest pipeline: persisted original files, chunked LLM Wiki
extraction, cached body embeddings, and import-job status/retry. Do not add direct skill
imports. Importing into a
custom assistant base must also mirror the package into the global base for
cross-source entity alignment. Admins may rename any assistant base and delete
custom bases. Deletion removes the custom base's Wiki pages, vectors, entities,
relations, import audit rows, and package files; the global base is never
deletable because it owns full-library aggregation.
Wiki package and local encoder invariants (2026-09)
Ordinary knowledge detail landing is globally configured by ordinary_knowledge_landing (wiki default / weknora), with admin-only PUT /api/system-settings/ordinary-knowledge-landing. Explicit document/graph/article links bypass this default. Wiki reading and the embedded WeKnora view provide reciprocal navigation; ?view=weknora explicitly selects the iframe and must not loop back to the default Wiki route.
knowledge_transfer.sync_to_globalis called after successful explicit local Wiki vectorization; it skips conversation/system mappings. Do not enable a new watcher or automatic remote-write scheduler as a side effect.wiki_packagesexposes file-backed assistant exports and restores into empty ordinary Wiki libraries. Validate vectors before any remote mutation; recreate folder IDs from category paths using provider APIs (a raw category_path is discarded by WeKnora without a matching folder_id). Preserve page metadata and rebuild backlinks after restoring all pages.PackageArchiveis a replayable, disk-backed JSONL staging index. Always close it in finally. Runtime imports must not use the legacy in-memory read_package.- Preserve normalized float32 base64 bytes exactly. Same dimension alone is not
model compatibility. Assistant vector retrieval is fingerprint-gated cosine
search, and any keyword fallback must be labeled as such.
Preserve a valid zero section_index. Export/search vectors belonging to each
current source snapshot (
source.last_import_job_id), not merely the page's current revision: one aligned page may intentionally contain current vectors from several sources. - The
localembedding provider uses FastEmbed with official BAAI/bge-m3 ONNX, 1024 dimensions, bounded CPU batches and a per-client async lock. Run inference off the event loop. Existing page vectors are reusable; zero-vector nonempty pages need backfill, not a forced rebuild of every page. FastEmbed is a standard backend dependency. Encoder configuration, directory upload and warmup live only in the administrator global settings dialog. A runtimesystem_settings.jsonpath override wins over config.yaml; uploaded files are streamed into.deer-flow/models/bge-m3, never the source tree. Encoder warmup is explicitly started from the settings UI/API and must never block application startup or turn a model-load failure into a DeerFlow service failure. Local provider is offline-only: require local_model_path or DEERFLOW_BGE_M3_MODEL_PATH and never let the application download model weights. Keeplocal_process_isolation=trueby default: inference runs in the bundled JSONL worker subprocess, and shutdown must terminate that worker without waiting on model teardown. - Every successful mapping snapshot is mirrored to the global assistant base, including empty snapshots so deleted source pages are retired. Re-import is source-replacement semantics: absent Wiki/chunk/entity/relation pages must not remain visible merely because they existed in an older source snapshot.
- Provider graph nodes bind to their original wiki_slug. They must not create duplicate synthetic entity/relation articles. Store original Wiki metadata separately from assistant workflow status. Count Wiki pages separately from raw document chunks.
- The global base aligns entities by normalized display name and aliases across source records, and aligns content Wiki pages by normalized title. Do not use source slug as the global canonical identity. Merge unique knowledge blocks while retaining source contributions/origins. Uploaded packages must rebind the temporary upload source to the stable WeKnora knowledge-base id in the package, so re-upload is source-replacement rather than append. Retrieval candidates exclude summary/index/directory pages; keyword matching, vector snippets, citations and prompts must use knowledge body/sections, never page summary as a substitute.
- Migration
20260903_01adds nullable page metadata and merges the pre-existing knowledge/workflow Alembic heads. Regression coverage is intests/test_wiki_package_roundtrip.py,test_llmwiki_wiki_sync.pyandtest_assistant_knowledge_repository.py.
openGauss / GaussDB persistence invariant
database.backend=gauss is an application-ORM backend, parallel to mysql.
It normalizes database.gauss_url to opengauss+asyncpg:// and requires the
gauss optional dependency (asyncpg plus opengauss-sqlalchemy). Never route
a Gauss URL through SQLAlchemy's generic PostgreSQL dialect: the wire protocol
may connect, but dialect initialization and schema DDL can fail against the
Gauss server/version behavior. Keep the LangGraph checkpointer on the separately
configured SQLite backend; the PostgreSQL checkpointer schema is not part of the
Gauss compatibility contract. Treat both postgresql and opengauss dialect
names as timezone-aware PostgreSQL-family engines in portable types and
dialect-specific session settings.