deerflow-code/offline-backend-20260512/backend
2026-09-07 18:24:55 +08:00
..
.model-config-backups 初始化 2026-09-07 18:24:55 +08:00
app 初始化 2026-09-07 18:24:55 +08:00
docs 初始化 2026-09-07 18:24:55 +08:00
knowledge 初始化 2026-09-07 18:24:55 +08:00
offline-wheels 初始化 2026-09-07 18:24:55 +08:00
packages/harness 初始化 2026-09-07 18:24:55 +08:00
scripts 初始化 2026-09-07 18:24:55 +08:00
skill-packages 初始化 2026-09-07 18:24:55 +08:00
tests 初始化 2026-09-07 18:24:55 +08:00
wheelhouse 初始化 2026-09-07 18:24:55 +08:00
_reqs.txt 初始化 2026-09-07 18:24:55 +08:00
.dockerignore 初始化 2026-09-07 18:24:55 +08:00
.python-version 初始化 2026-09-07 18:24:55 +08:00
AGENTS.md 初始化 2026-09-07 18:24:55 +08:00
build-backend-image.sh 初始化 2026-09-07 18:24:55 +08:00
CLAUDE.md 初始化 2026-09-07 18:24:55 +08:00
config.yaml 初始化 2026-09-07 18:24:55 +08:00
CONTRIBUTING.md 初始化 2026-09-07 18:24:55 +08:00
debug.py 初始化 2026-09-07 18:24:55 +08:00
DOCKER_OFFLINE.md 初始化 2026-09-07 18:24:55 +08:00
Dockerfile 初始化 2026-09-07 18:24:55 +08:00
Dockerfile.backend-offline 初始化 2026-09-07 18:24:55 +08:00
extensions_config.json 初始化 2026-09-07 18:24:55 +08:00
langgraph.json 初始化 2026-09-07 18:24:55 +08:00
Makefile 初始化 2026-09-07 18:24:55 +08:00
memory_user_overrides.json 初始化 2026-09-07 18:24:55 +08:00
pyproject.toml 初始化 2026-09-07 18:24:55 +08:00
README.md 初始化 2026-09-07 18:24:55 +08:00
requirements-offline.txt 初始化 2026-09-07 18:24:55 +08:00
ruff.toml 初始化 2026-09-07 18:24:55 +08:00
save-backend-image.sh 初始化 2026-09-07 18:24:55 +08:00
start-backend-docker.sh 初始化 2026-09-07 18:24:55 +08:00
tests.rar 初始化 2026-09-07 18:24:55 +08:00
后台离线部署说明.md 初始化 2026-09-07 18:24:55 +08:00
离线镜像部署简明指南.md 初始化 2026-09-07 18:24:55 +08:00

DeerFlow Backend

普通知识库默认详情页面:管理员可在知识库顶部选择“本系统 Wiki 索引”或“嵌入 WeKnora”,对应 ordinary_knowledge_landing(默认 wiki)。Wiki 库首次进入默认展示本系统索引,右上角“打开 WeKnora”切换到内嵌管理页;内嵌页可返回索引。非 Wiki 库仍进入 WeKnora 文档页。设置通过 /api/system-settings/ordinary-knowledge-landing 保存,仅管理员可修改;指定文章、文档、图谱的直达链接优先于默认设置。

助手知识库可独立于 WeKnora 使用:管理员进入助手知识库,选择或创建库后点击“上传文件生成 Wiki”。支持 Markdown/TXT/CSV、PDF、Word、Excel、PowerPoint(单文件 100 MB,解析正文上限 400 万字符)。文件先写入运行目录 assistant-knowledge/files/{job_id},后台分段调用已配置的语言模型生成分类 Wiki、实体关系,再通过 DeerFlow 编码器对正文分段向量化,保存到助手库并汇入知识梳理总库。无需配置 WeKnora URL;配置后仍可直接上传,也保留 WeKnora 数据包导入。文件处理沿用管理员写入权限,所有用户可以浏览和搜索。

导入记录显示解析、Wiki 生成和向量化阶段;失败或服务中断后可“重试 / 恢复处理”,复用已落盘的 Wiki 和相同模型指纹的向量。离线编码器须预先放置模型并在管理员设置中配置;模型或编码器不可用会明确失败,不把摘要冒充 Wiki 或虚构向量。扫描件是否可解析取决于部署的文档解析/OCR 能力。

DeerFlow is a LangGraph-based AI super agent with sandbox execution, persistent memory, and extensible tool integration. The backend enables AI agents to execute code, browse the web, manage files, delegate tasks to subagents, and retain context across conversations - all in isolated, per-thread environments.


Architecture

                        ┌──────────────────────────────────────┐
                        │          Nginx (Port 2026)           │
                        │      Unified reverse proxy           │
                        └───────┬──────────────────┬───────────┘
                                │                  │
              /api/langgraph/*  │                  │  /api/* (other)
                                ▼                  ▼
               ┌────────────────────┐  ┌────────────────────────┐
               │ LangGraph Server   │  │   Gateway API (8001)   │
               │    (Port 2024)     │  │   FastAPI REST         │
               │                    │  │                        │
               │ ┌────────────────┐ │  │ Models, MCP, Skills,   │
               │ │  Lead Agent    │ │  │ Memory, Uploads,       │
               │ │  ┌──────────┐  │ │  │ Artifacts              │
               │ │  │Middleware│  │ │  └────────────────────────┘
               │ │  │  Chain   │  │ │
               │ │  └──────────┘  │ │
               │ │  ┌──────────┐  │ │
               │ │  │  Tools   │  │ │
               │ │  └──────────┘  │ │
               │ │  ┌──────────┐  │ │
               │ │  │Subagents │  │ │
               │ │  └──────────┘  │ │
               │ └────────────────┘ │
               └────────────────────┘

Request Routing (via Nginx):

  • /api/langgraph/* → LangGraph Server - agent interactions, threads, streaming
  • /api/* (other) → Gateway API - models, MCP, skills, memory, artifacts, uploads, thread-local cleanup
  • / (non-API) → Frontend - Next.js web interface

Health probes are always unauthenticated, even when API authentication is enabled: GET|HEAD /health, /api/health, and /api/langgraph/health all return the Gateway health status without requiring a cookie or bearer token.


Core Components

Lead Agent

The single LangGraph agent (lead_agent) is the runtime entry point, created via make_lead_agent(config). It combines:

  • Dynamic model selection with thinking and vision support
  • Middleware chain for cross-cutting concerns (9 middlewares)
  • Tool system with sandbox, MCP, community, and built-in tools
  • Subagent delegation for parallel task execution
  • System prompt with skills injection, memory context, and working directory guidance

Workflow Studio runtime

  • Candidate workflow graphs remain editable after they are loaded into the Coze-compatible canvas; the compatibility response explicitly marks the current owner as editable so node dragging is not mistaken for a read-only preview.
  • Workflow-agent SSE forwards provider reasoning_content as an inline <think>…</think> trace for the UI while keeping the executable node output limited to the visible answer text. Every native LangGraph [AIMessage|ToolMessage, metadata] frame is additionally persisted and sent as node.message.data.messages without projection or length cropping, so the Studio can replay DeerFlow's tool-result, search-result and ask_clarification cards exactly after reconnecting.

AgentScope report collaboration (in progress)

The independent /api/report-collaboration workbench supports durable run-time intervention commands and immutable formal report versions. Its production worker now runs a real AgentScope 2.0.5 Agent per TaskLedger assignment; AgentScope owns the member ReAct/tool/structured-output loop, while TaskLedger alone owns DAG readiness, retries and completion. Ready nodes in one wave run concurrently up to max_parallel_tasks; cancellation and member exceptions close agent runs as terminal states, and lease recovery remains independent of in-memory AgentScope objects. Retrieval uses the deployment's configured DeerFlow web_search/web_fetch tools through the frozen role permission snapshot. Every visible AgentScope reply, TeamSay handoff, tool call and tool result is persisted immediately into the existing SSE/message protocol; hidden reasoning and system prompts are not exposed. A constrained intent agent classifies explicit answers, angle additions, re-search, re-analysis, node reruns, section rewrites, polishing, report questions, replanning, and cancellation. Server-side impact analysis validates node/agent ownership and computes the minimum downstream lineage; low-confidence, high-cost, or multi-branch changes require confirmation. Commands are CAS-claimed and applied at TaskLedger safe points, where old attempts and artifacts become superseded so late results cannot contaminate the revised report. Run creation freezes the selected /api/report-structures row and all server-enforced budgets. Per-user active-run limits and per-run model-call, retrieval, source-byte, token, cost, node-time, and total-time limits fail closed. User-visible event payloads are centrally redacted before durable persistence without size truncation; external source content is explicitly untrusted, and hidden reasoning/system prompts are never exposed. Append-only audits cover agent selection, tools, sources, user commands, confirmations, report versions, redactions, and budget failures; owner-scoped reads are available at GET /runs/{id}/audits. Temporary Markdown is emitted as report.delta; only a template/citation/coverage-clean report with an independent passing review is atomically persisted and emitted as report.version.created before the run becomes completed. Section rewrites are replacement candidates; apply and restore append a new head version and never destroy history. Configure runtime_mode, concurrency and limits under report_collaboration; both the feature and worker remain disabled by default. Install the report-collaboration project extra when enabling the worker.

Middleware Chain

Middlewares execute in strict order, each handling a specific concern:

# Middleware Purpose
1 ThreadDataMiddleware Creates per-thread isolated directories (workspace, uploads, outputs)
2 UploadsMiddleware Injects newly uploaded files into conversation context
3 SandboxMiddleware Acquires sandbox environment for code execution
4 SummarizationMiddleware Reduces context when approaching token limits (optional)
5 TodoListMiddleware Tracks multi-step tasks in plan mode (optional)
6 TitleMiddleware Auto-generates conversation titles after first exchange
7 MemoryMiddleware Queues conversations for async memory extraction
8 ViewImageMiddleware Injects image data for vision-capable models (conditional)
9 ClarificationMiddleware Intercepts clarification requests and interrupts execution (must be last)

Sandbox System

Per-thread isolated execution with virtual path translation:

  • Abstract interface: execute_command, read_file, write_file, list_dir
  • Providers: LocalSandboxProvider (filesystem) and AioSandboxProvider (Docker, in community/)
  • Virtual paths: /mnt/user-data/{workspace,uploads,outputs} → thread-specific physical directories
  • Skills path: /mnt/skills → deer-flow/skills/ directory
  • Skills loading: Recursively discovers nested SKILL.md files under skills/{public,custom} and preserves nested container paths
  • Conditional WeKnora LLMWiki: a non-empty llmwiki.weknora.api_base_url enables the DeerFlow-native WeKnora BFF and UI; an empty address preserves the existing LLMWiki page. WeKnora remains independently deployed and owns its database, Redis, storage, parsers and models. DeerFlow keeps only the API/admin addresses, user/publication mappings, agent knowledge-base bindings and per-conversation selections; it stores and sends no WeKnora credential, so the endpoint must be restricted to a trusted network and allow DeerFlow traffic. Owners publish directly to the public LLMWiki without an approval queue for ordinary bases; the system-owned conversation-deposit base is always published, shared by all users, excluded from My Knowledge, immutable through user/admin detail operations, and recognized by both its reserved name and system owner. An existing same-name remote mapping is adopted in place as the system deposit so its mapping id and document links stay valid. Its user-facing knowledge-base description and generated document metadata use the cmzs brand; list/detail reads synchronize stale remote descriptions, while conversation-deposit search results, source previews and embedded-detail JSON rewrite legacy brand text to cmzs. Deposit requests must reference a thread accessible to the current account. Personal knowledge-base lists are scoped to the current DeerFlow account even for admins, while frontend LLMWiki query caches also include the active account id so switching accounts cannot reuse another account's list/detail data. Admins remain the only role that sees the WeKnora connection-config tab. Knowledge cards open a dedicated WeKnora-style detail route: Wiki-enabled bases default to Wiki, then documents and graph; standard bases hide Wiki and default to documents. Writable ordinary knowledge-base details expose a native multi-file upload action that submits files sequentially and reports aggregate success/failure; read-only public bases and the system deposit base do not. It supports file/URL/manual-Markdown imports, Wiki page browsing, document preview, real parsed chunks and Wiki graph nodes/edges. New chats use a compact WeKnora-style searchable/grouped multi-select inside the composer, while agent create/edit pages bind spaces through an inline panel. A single proxy service identity means DeerFlow users are isolated logically while their WeKnora entities share that identity's workspace. The read-only llmwiki_search tool reauthorizes the current DeerFlow user on every call and renders credential-free source previews; source previews prefer cleaned Wiki/LLMWiki text when a retrieved chunk maps to a Wiki page, with raw chunks as the fallback. The detail page first creates a signed iframe session through the normal DeerFlow API auth path, then points the iframe directly at the proxied WeKnora /platform/... route using the signed deerflow_weknora_embed cookie; global auth/CSRF bypass is limited to those signed iframe proxy requests, and HTTPS cookie attributes honor X-Forwarded-Proto behind reverse proxies. The iframe proxy keeps WeKnora's root-relative /platform, /assets, /locales, and /api/v1/* paths, but is registered last through llmwiki.proxy_router so it cannot intercept first-party Gateway APIs such as POST /api/v1/auth/login/username. See docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md.
  • Enterprise Research workbench (temporarily disabled): the independent /api/enterprise-research router, its report executor, and its ORM table registration are intentionally inactive. The implementation remains in the repository, but startup does not create or alter enterprise_research_tasks or enterprise_research_report_jobs; this keeps the active Deep Research workflow independent of the paused workbench and avoids MySQL schema errors.
  • Embedded WeKnora detail chrome: the proxy hides WeKnora's application sidebar (.main > .aside_box) and expands the detail outlet to the full iframe width, while preserving the Wiki page's own index/navigation panel. The host does not sandbox this trusted, permission-filtered detail frame because browser PDF viewers are disabled inside sandboxed ancestor frames. Preview and knowledge-base file redirects are followed by the Gateway so PDF bytes and protected Wiki images remain inside the authorized proxy instead of leaking to an unreachable object-storage URL.
  • Embedded WeKnora knowledge-base switching: the original breadcrumb picker works inside the iframe, but lists only the signed-in user's own and published knowledge bases. Every switch is reauthorized through DeerFlow and receives a fresh short-lived iframe session, so it never exposes the upstream service-admin knowledge-base list.
  • Prefixed WeKnora deployment: the frontend may be published under a separate prefix such as /magentweb, while embedded WeKnora detail traffic is served by the Gateway under /deerflow. The Gateway rewrites HTML root assets, WeKnora API URLs, redirects, and Vite dynamic dependency maps so page assets stay under /deerflow; the prefixed proxy routes must be registered before the unprefixed fallback. The signed embed cookie covers /deerflow, and auth/CSRF recognition includes all supported static directories.
  • File-write safety: str_replace serializes read-modify-write per (sandbox.id, path) so isolated sandboxes keep concurrency even when virtual paths match
  • Tools: bash, ls, read_file, write_file, str_replace (bash is disabled by default when using LocalSandboxProvider; use AioSandboxProvider for isolated shell access)

Subagent System

Async task delegation with concurrent execution:

  • Built-in agents: general-purpose (full toolset) and bash (command specialist, exposed only when shell access is available)
  • Concurrency: Max 3 subagents per turn, 15-minute timeout
  • Execution: Background thread pools with status tracking and SSE events
  • Flow: Agent calls task() tool → executor runs subagent in background → polls for completion → returns result

Memory System

LLM-powered persistent context retention across conversations:

  • Automatic extraction: Analyzes conversations for user context, facts, and preferences
  • Structured storage: User context (work, personal, top-of-mind), history, and confidence-scored facts
  • Debounced updates: Batches updates to minimize LLM calls (configurable wait time)
  • System prompt injection: Top facts + context injected into agent prompts
  • Storage: JSON file with mtime-based cache invalidation

Tool Ecosystem

Category Tools
Sandbox bash, ls, read_file, write_file, str_replace
Built-in present_files, ask_clarification, view_image, task (subagent)
Community Tavily (web search), Jina AI (web fetch), Firecrawl (scraping), DuckDuckGo (image search)
MCP Any Model Context Protocol server (stdio, SSE, HTTP transports)
Skills Domain-specific workflows injected via system prompt

The config-defined web_search tool can call a deployment-specific JSON API instead of a public search provider. Set its endpoint, static payload, SSL verification, timeout, result limit, and result_url_template in config.yaml. At runtime only payload.query is overwritten with the model's search text. The API response must use {"results": [{"recUuid": "", "content": "", "title": "", "url": ""}]}; missing or null result fields are normalized to empty strings. When recUuid is present, {recUuid} in result_url_template replaces the original result URL; otherwise the original URL is retained. Set enabled: false to keep the tool completely out of the agent's registered tool list. verify_ssl: false supports trusted internal HTTPS services with self-signed certificates.

Gateway API

FastAPI application providing REST endpoints for frontend integration:

Route Purpose
GET /api/models List available LLM models
GET/PUT /api/mcp/config Manage MCP server configurations
GET/PUT /api/skills List and manage skills
POST /api/skills/install Install skill from .skill archive
POST /api/skills/validate, POST /api/skills/install-upload Validate and upload .zip/.skill packages with owner-aware conflicts: owners may explicitly overwrite their own copy; other users' copies report the uploader and cannot be overwritten
GET/POST /api/agents List visible agents or create a user-owned agent
GET/PUT/DELETE /api/agents/{id} Read, update, or delete an id-addressed agent
GET/POST/PUT/DELETE /api/roundtable-chains Manage user-owned roundtable business chains; each seat may optionally carry position_id for the separate position-collaboration workspace, plus human_validation to mark a Step 2 client-side pause checkpoint after that seat/stage completes (the original roundtable has no extra validation card)
GET/PUT /api/business-mapping Read global business mappings. Updates to 3Q/6BF/7BF require an administrator; the 8BF public business-chain mapping is maintained exclusively by the trusted lqq account (derived from the authenticated email prefix).
GET/POST/PUT /api/position-roundtable/sessions Standalone position-collaboration sessions: freeze task/intent/chain snapshots, index one direct-agent node per seat, build upstream handoff context (including bounded UTF-8 excerpts of text artifacts), and make archived sessions read-only without invoking the original roundtable coordinator or jobs. Built-in summary/action-plan completion mirrors visible report text into a Markdown artifact when a model omits its final write_file call. POST /sessions/{id}/nodes/{nodeKey}/reject accepts an optional reason, marks the rejected delivery for rework, and invalidates affected downstream deliveries so they are regenerated in order. The same router is also mounted at /api/multi-agent/position-roundtable/* for intranet gateways that only forward the established multi-agent namespace; the compatibility mount is intentionally omitted from OpenAPI.
GET /api/memory Retrieve memory data
POST /api/memory/reload Force memory reload
GET /api/memory/config Memory configuration
GET /api/memory/status Combined config + data
POST /api/threads/{id}/uploads Upload files (auto-converts PDF/PPT/Excel/Word to Markdown, rejects directory paths)
GET /api/threads/{id}/uploads/list List uploaded files
DELETE /api/threads/{id} Delete DeerFlow-managed local thread data after LangGraph thread deletion; unexpected failures are logged server-side and return a generic 500 detail
GET /api/artifact-library List the current user's generated conversation files across normal threads, with filename/conversation search, file-kind filtering, pagination, and an optional cache-bypassing refresh. The filesystem outputs/ directories remain authoritative; uploads and image/audio/video files are excluded.
GET /api/threads/{id}/artifacts/{path} Serve generated artifacts
GET/POST /api/deep-research/* Durable Deep Research sessions and jobs. quick/basic run the compact report flow; detailed plans subtopics and writes independent sections; deep recursively investigates bounded follow-up branches. All modes reuse DeerFlow's configured material provider, LLM factory, hidden-thread artifacts, and SSE event log. Deep Research session/job/event/source/message writes serialise their rows in the writer transaction, rather than doing a post-commit ORM refresh or readback; this prevents MySQL read/write-splitting replica lag from failing “start writing”, dispatcher claim, report persistence, or follow-up history. Whole-document artifact/report rewrites likewise create their reversible version snapshots from writer-local ORM data and never refresh after commit, so replica lag cannot restore an otherwise successful rewrite. A report is written to the session and confirmed before its Job becomes completed; an unconfirmed write is retried and never exposed as a false completion. Retrying an already-created queued job re-nudges the dispatcher, recovering an interrupted HTTP response without creating a second report. If progress-event persistence, durable replay, or live fan-out is briefly unavailable, the worker/stream logs the degradation and continues or reconnects rather than failing the whole job. Whole-report rewrite accepts an optional reportOutline Markdown template; the durable job persists it and applies it as the chapter structure constraint. Optional report illustrations use a separately configured OpenAI-compatible image endpoint and are served only through the owning session's authenticated artifact route.
/api/enterprise-research/* Temporarily disabled. Its router and persistence initialization are not registered, so these endpoints return 404 and no enterprise-research tables are created or migrated during startup.
POST /api/ai-writing/sample/extract Extract text from Word/PDF/Markdown/TXT samples for AI-writing imitation mode
POST /api/writing/export/docx Generate the shared formal Word download from Markdown. Uses the standard-library OOXML builder in app/gateway/word_export.py, removes manual/compound heading prefixes and Markdown horizontal rules, applies the prescribed A4/margin/heading/body/table/footer profile, and embeds the server-bundled 方正小标宋简体.ttf; the client computer does not need that font installed.
POST /cop/saveSuperiorTask TaskCOP compatibility API (DeerFlow Bearer/session authentication required): create a situation-overview task with a sequential four-digit id (0001, 0002, ...) and return the legacy {state, msg, data} envelope
POST /api/taskcop/import/tasks TaskCOP compatibility API (authenticated): import tasks from a JSON array or tasks/records envelope; every row's supplied id/taskId is authoritative and a duplicate id completely replaces the stored task
DELETE /cop/tasks/{task_id} TaskCOP compatibility API (authenticated): delete one task and its task-scoped situation-report document
GET /taskAnalyseSearch/cop-task-three-list TaskCOP compatibility API (authenticated): paginated task list for the situation-overview sentiment page (content, taskStatus, taskDirection, startTime, endTime)
GET/POST/PUT /api/task-reports/by-task/{task_id} Shared task-scoped situation-report detail. Query, first save, and full update all use the legacy-compatible {state, msg, data:[{id, taskId, createTime, updateTime, sessionId, contentJson, categoryType}]} shape; contentJson is preserved as a JSON string. Saving a non-empty report automatically moves the matching TaskCOP task to status 25 (the legacy list's “已完成” state).
POST /api/task-reports/import/by-task/{task_id} TaskCOP detail import (authenticated): accept a sq-report-mock.json-compatible data/records array and atomically replace the selected task's full report. The path task id overrides task ids inside the file. The agentfx built-in agent fills that file with the task-report-build convert skill in one call (search JSON → 14 categories; task-report-{enemy,our,env,judge} re-run a single dashboard page) then imports via task-report-import.
POST /api/sentiment-agent/stream Authenticated BFF for the frontend sentiment-analysis virtual agent. Forwards the AG-UI JSON body to the apiUrl supplied from runtime-config.js (Basic Auth from the same payload). TLS certificate verification is off. Upstream SSE is piped through unchanged; connection/HTTP failures return the complete error in detail.

The separately deployed TaskCOP Vue app uses POST /api/v1/auth/login/username directly (its username, userId, or yUserId URL parameter identifies the user), so this compatibility flow does not depend on the retired Consumer login service. It calls the Gateway address configured as b1ConsumerUrl directly; it does not use a Vite /api reverse proxy. An empty TaskCOP store remains empty; task data is introduced through normal creation or the JSON import API. The Gateway CORS middleware always merges local Vite origins http://localhost|127.0.0.1:{5173,5174,3000,8080}, including when GATEWAY_CORS_ORIGINS is set for another frontend. This keeps a Vite server that falls back from 5173 to 5174 (or the reverse) working after a backend restart. For a deployment, set GATEWAY_CORS_ORIGINS to the exact additional frontend origins that are allowed to call the Gateway.

Workflow Studio interface resources

The Workflow Studio data-source API also registers reusable HTTP interface resources. Create them through POST /api/workflows/data-sources with kind set to http; provide a non-secret baseUrl, an allowedMethods allowlist, and optional request headers. The service encrypts the headers immediately and never returns them through list, detail, resource-catalog, run-event, or canvas APIs. The canvas resource catalog exposes only the name, description, HTTP address, and allowed methods. HTTP execution continues to enforce the workflow target security policy, including the default SSRF restrictions.

The same catalog's agent and skill entries are display projections, not a second ownership store: /api/workflows/resources/agents returns agentId, cached available skills, and createdBy; the skills endpoint returns skillId, direct-call capability, and createdBy. Creator names are resolved from the immutable owner id at read time (falling back to the id if an account was removed), while built-in resources explicitly return 系统内置.

Conversation-driven workflow planning

The Workflow Studio chat composer does not silently start a template. A normal user message first creates a durable planning session with POST /api/workflows/{workflow_id}/planning-sessions/stream (SSE; the legacy non-streaming POST /api/workflows/{workflow_id}/planning-sessions remains for integrations). The stream relays safe DeerFlow controller milestones—catalog read, controller parsing, role selection, graph assembly and validation—before returning the final session, so the UI never has to fabricate a progress bar. The constrained planner builds two or three validated candidate graphs from the current draft and the visible business agents the user can use. It is a dedicated built-in 工作流总控 (workflow-planner), not the multi-agent roundtable coordinator: it only reads the user requirement and catalog metadata, selects permitted agent ids, and ranks the supported strategies. It does not dispatch a roundtable, use research tools, or execute a formal run. The server constructs the executable graph from those selected, visible resources; a controller failure returns an explicit planning error rather than a keyword/template fallback. The controller call has no callable tools and disables model thinking, so it adds one bounded planning inference without starting research or worker-agent costs. A browser selects one candidate, automatically applies the editor's DAG layout when it loads, then may explicitly save an edited graph through PATCH /api/workflows/planning-sessions/{session_id}/proposals/{proposal_id}/graph, then creates the one formal run through the existing POST /api/workflows/{workflow_id}/runs endpoint with planningSessionId and proposalId. The gateway rejects unselected or invalid candidates and makes a repeated confirmation return the already-created run, so a double click cannot execute a second workflow. Planning sessions are persisted separately from runs (migration 20260831_01); their parent row is flushed before candidate rows so the foreign-key transaction works consistently across SQLite, MySQL, and PostgreSQL. A candidate is not consuming execution resources until it is confirmed.

The development Studio proxy must not gzip /api responses: compressed SSE can hold small planning and run-event frames until the response ends. The backend already marks both streams no-transform and X-Accel-Buffering: no; apps/coze-studio also sets its Rsbuild development server compress: false. Agent/skill nodes use the bounded workflows.agent_recursion_limit (default 250), rather than a chat-sized 60-step limit, because a model-to-tool research turn consumes about two LangGraph super-steps. The composer may also send a configured modelName: the server validates it against the safe model catalog, uses it for the workflow controller, and writes it into the generated candidate agent nodes. For a direct single-agent task it overrides only that projected node; confirmation still executes the immutable selected proposal snapshot. The controller additionally creates a per-selected-agent task contract (mission, deliverable, scope, handoff). The graph builder accepts contracts only for catalog-visible, selected agent ids and injects each one into that worker's prompt, so parallel branches receive complementary research assignments and sequential branches have an explicit review/handoff boundary rather than all workers attempting the whole user task. Missing controller fields receive a deterministic role-aware contract. When an agent emits DeerFlow's native ask_clarification ToolMessage, the workflow enters awaiting_input rather than completing or failing. The pause event references the originating nodeRunId and toolCallId; the browser renders that original card inside the agent message and POST /runs/{run_id}/resume supplies the user's reply as the next turn of the same agent thread.

The current planner is deliberately constrained: it can rank supported strategies and focus resources, but the server builds the graph from the already configured agent nodes. It cannot invent a node type, external target, or credential from chat text. For a parallel-research candidate, the configured research agents first flow into the deterministic evidence_normalizer node. It emits a bounded Evidence Pack with the originating node, role label, explicit gaps, and no invented verification claim before the synthesis agent sees it. The candidate then runs deep_research_write: its app-layer adapter converts exactly that Evidence Pack into selected, provenance-tagged Deep Research source rows, starts or reattaches to the existing durable Deep Research job, forwards report deltas as workflow node.output.delta events, and exposes the final Markdown as a run artifact. Cancellation is forwarded to the same research job; the node never calls the Deep Research HTTP API or launches a second writer. The session key includes the topic, evidence snapshot and writing configuration, so a retry reuses the same job while a changed upstream pack cannot return stale report prose. multi_agent Deep Research is intentionally not nested yet because its separate plan-review pause would conflict with workflow human_input. A completed agent or skill node in a terminal, acyclic run can now receive targeted feedback through POST /api/workflows/runs/{run_id}/feedback: the original remains immutable, a revision run reuses completed nodes outside the target's downstream closure, and only the target prompt receives the bounded feedback block before it and its downstream nodes stream again. The public revision run retains the affected and reused node ids in its feedback summary for the conversation/history UI. Running and looped workflows deliberately continue to use human_input or a new full run; they are not modified in place. See docs/WORKFLOW_CONVERSATIONAL_MULTI_AGENT_ORCHESTRATION_ZH.md for the roadmap.

Deep Research modes

Deep Research is a separate Gateway feature rather than part of AI Writing. Its detailed runner adapts GPT Researcher's detailed-report control flow (plan subtopics, collect materials concurrently, write sections concurrently, then assemble the report). Its deep runner adapts the recursive DeepResearchSkill flow (branch, extract learnings, investigate follow-ups) with a deterministic query budget derived from deep_breadth and deep_depth. The runners live in packages/harness/deerflow/agents/deep_research/runners/ and only use injected DeerFlow adapters. multi_agent is a LangGraph 1.x editor/researcher/writer/ reviewer graph: it persists a plan-review pause, releases the worker lease, and resumes from the approved or revised plan through the existing durable-job API. Report prose uses the configured LangChain model's astream() path. Each native model delta is fanned out as an in-process report_delta SSE frame so the right-side report.md sandbox can write smoothly, while bounded, persisted report_chunk events remain reconnect checkpoints. The durable log is still authoritative: a browser reconnect, slow consumer, or another worker falls back to ?after=<seq> replay rather than starting a second model call. Reasoning exposed by providers either through structured reasoning fields or inline <think>...</think> text is separated from report deltas, checkpoints, and the persisted Markdown projection. It may be forwarded as an ephemeral live report_thinking frame for the message-list thought trace, but is never persisted or written to the sandbox; only answer Markdown reaches the sandbox. The short non-streaming control-plane calls used to plan queries, sections, and reviews are independently capped at 45 seconds (maximum 120 seconds). If query or section planning times out, the runner emits a visible recoverable warning and searches the original research topic instead of leaving the job at the planning step indefinitely. Report-prose streaming is intentionally not capped by this control-plane limit. Planning also persists queries_planned before each basic, detailed, or multi-agent parallel search fan-out; the completed session stores its terminal lastJobId in the usage snapshot so the conversation workbench can replay the full planning/search/source/curation trace after a page refresh. These are observable execution facts only, never hidden model reasoning. Detailed reports use stable introduction/section/conclusion targets; a multi-agent review revision emits report_reset before the replacement draft. Every session may also provide a bounded custom_outline. Markdown headings or numbered entries become the deterministic section plan for detailed and multi_agent; basic and deep receive the same outline as a guarded report-writing constraint. The outline is frozen in the session snapshot and never treated as a system instruction.

All owner-scoped Deep Research routes resolve the platform's asynchronous get_current_user identity before querying or writing persistence, so session and job records are always bound to a concrete user ID.

Report artifacts are server-named and written through the canonical sandbox virtual path (/mnt/user-data/outputs/<name>), which the path layer maps to the host filesystem. This keeps native Windows deployments compatible without relaxing the sandbox traversal guard.

Deep Research reuses the DeerFlow Q&A retrieval stack instead of a separate scraper. It first runs enabled research skills (deep-search, web-research, and enabled knowledge-base skills) through the same SubagentExecutor runtime used by Q&A, then uses a configured web_search provider when present. If this offline deployment still has the shipped placeholder/disabled web_search configuration, it falls back to DeerFlow's existing keyless DuckDuckGo provider and, when that provider returns only irrelevant engine noise, a second real-web search source. It persists only real returned sources. A missing intranet endpoint therefore cannot by itself produce an empty Deep Research report. The capabilities API reports skill only when its required Q&A tool is available, and reports web_search when configured search or the online fallback can run. Cancelling a queued research job finalizes its job row and immediately marks the owning session cancelled with no active job, so it cannot retain a per-user concurrency slot. Saved draft sessions do not consume a concurrency slot; the two-job limit only applies to sessions with a running or awaiting_input durable job.

Whole-report rewrite builds its requirement plan locally from the selected style and instruction, publishes that plan immediately, and then opens only one model stream for the actual Markdown writer. This avoids making users wait for a separate model-thinking/planning pass before sandbox output begins. POST /api/deep-research/sessions/{id}/report-variants is the structural regeneration boundary: it copies the completed report and selected evidence into a new session (including both saved [来源:id] and legacy streamed [[source:id]] citation markers remapped to the cloned ids). The following durable writer is explicitly queued with operation=generate: it does not send the inherited report to the model as text to edit, and instead writes a fresh report from the complete selected evidence set plus the newly confirmed outline. Model reasoning remains a live SSE event before the first Markdown token. New output uses display-ready [来源:id] citations; validation repairs only unique near-complete drs_src_* ids (for example a provider-truncated suffix) and still rejects genuinely unknown evidence ids. The report-only model stream requests an 8192-token completion budget so a normal multi-section report and its reference list do not stop at the common 4096-token provider default; malformed trailing citation syntax is rejected rather than committed as a partial document. The original report therefore remains independently available in history even if the new report is cancelled or fails.

Optional report illustrations

DeerFlow's stock configuration includes image recognition but no image-generation provider. Deep Research therefore keeps illustrations disabled until an operator fills config.yaml -> deep_research.image_generation with an image endpoint, API key, and model. provider: openai_compatible uses the standard /images/generations protocol and requires b64_json; provider: dashscope_native supports Qwen Image 3.0 and Wan 2.7 on DashScope / Token Plan with the native multimodal-generation protocol. Native provider image URLs are downloaded immediately, so both providers produce protected local artifacts. When configured, the capabilities endpoint enables the page's “生成报告配图” switch. A selected job writes at most four research-image-*.png|jpg|webp artifacts, injects deep-research:// links into the Markdown report, and emits image SSE events. A failed/unconfigured image provider only emits a recoverable warning; the text report still completes.

Completed reports also expose GET /api/deep-research/sessions/{id}/messages and POST /api/deep-research/sessions/{id}/chat. Follow-up answers are grounded only in that session's report, selected sources, and recent follow-up history; the assistant returns persisted source ids for every accepted citation. Passing allow_new_research=true is rejected rather than silently triggering another search job. Owners can change the follow-up evidence set with PATCH /api/deep-research/sessions/{id}/sources/{source_id} after a research job stops. That selection is applied immediately to later follow-up prompts; it never rewrites the completed report, and edits are rejected while the job is running so they cannot race the automatic curation projection.

The Deep Research page is a conversation workspace: the original research question, observable retrieval steps, completion summary, and report-grounded follow-ups appear in chronological order. The report itself is written in the adjacent sandbox as /mnt/user-data/outputs/report.md, with Word, Markdown, and HTML export actions. Word export calls the shared Gateway POST /api/writing/export/docx generator, embeds the licensed title font, and includes the collected reference list without adding a backend Python package. In chat-collection sessions the report configuration card shows the unique material count inferred from the collector's visible tool results before writing begins. The write lifecycle itself is rendered as ordinary workspace message-list tool steps (not a second timeline component), and remains in the transcript after completion; the generated digest is a normal assistant message below those steps. As soon as the runner enters its summary phase, that same assistant message is inserted in a streaming state and receives each SSE summary_delta; completion only changes it in place to the final report-complete wording rather than appending a delayed second summary card.

Each completed chat-collection report also contributes a normal file card to the message history, named from the report Markdown title rather than the internal report.md artifact path. Historical sessions keep the sandbox closed on replay, and opening that card selects the session's virtual report artifact in the sandbox; its download action uses the existing Markdown/Word export dialog. The per-user active-job limit and the default in-job retrieval/write concurrency are both three. The Deep Research composer uses the installed TDesign Select, populated from /api/models, for an optional per-research model override. It is disabled during active collection/writing, is sent as model_name to collector turns, and is applied to the fast_model, smart_model, and strategic_model roles when a session or report job starts.

Completed-report Q&A uses POST /api/deep-research/sessions/{id}/chat/stream for native SSE deltas; the earlier /chat JSON endpoint remains for API compatibility. An explicit scoped edit such as “修改第一段”、 “重新生成一下第一段” or “重写第二节” uses the separate POST /api/deep-research/sessions/{id}/rewrite/stream report-section writing pipeline. That route requires a deterministic Markdown range and returns 422 instead of falling back to Q&A when no target is supplied. Ordinary follow-up questions remain report/source-grounded and never mutate the report. The streamed replacement candidate is shown in the message list while the sandbox keeps the current report unchanged. The completed candidate creates a confirmation card; only POST /sessions/{id}/messages/{messageId}/rewrite-proposal with {"action":"apply"} substitutes the matching range and saves the assembled Markdown. The proposal carries the source report hash, so a stale candidate is rejected instead of overwriting a later edit; dismissing/cancelling leaves the persisted report unchanged. This does not initiate new web research. Legacy unmarked candidates from before this route split are inferred only when their immediately preceding user message contains the same explicit scoped edit.

Position-roundtable history supports two modes. A non-empty external_task_id from the frontend route's taskId creates a row with task_scoped = true, so all sessions, nodes and persisted conversation snapshots for that task are shared across users; user_id is then creator audit metadata only. When the field is absent or empty, the original per-user personal-session behavior remains unchanged. Alembic revision 20260808_01 adds task_scoped with the safe default false for historical rows plus the task/update-time index.

Roundtable concurrency guarantees

Position-roundtable node completion, rejection, thread binding and activation are database transactions with row locking and retry-stable command ids. This keeps simultaneous seat completion from leaving a downstream stage locked. Task and intent snapshots are frozen when the chain becomes active, so a late request from another user cannot alter the confirmed input. Background roundtable jobs use a database unique active-job key and renewable worker leases; cancelled jobs are finalized by the lease holder or recovered after an expired lease, allowing users to start a new job after a worker failure.

The ordinary unit tests use SQLite. To verify deployed database semantics with real PostgreSQL row locks and independent connections, set DEERFLOW_POSTGRES_CONCURRENCY_TEST_URL and run:

PYTHONPATH=. uv run --no-sync pytest tests/test_roundtable_postgres_concurrency.py -q

AI writing supports a sample-imitation mode: the frontend can upload or paste a sample article, the Gateway extracts text for document files, and the ai_writing graph runs a sample_analyzer node before normal intent/research/outline/draft stages. The resulting style profile is injected into outline and draft prompts with explicit guardrails against copying source sentences, facts, names, or data.

Custom Agents

Custom agents are indexed in the agents database table. The id column is the stable runtime identifier and filesystem directory name (.deer-flow/agents/{id}/SOUL.md); name is display-only, may be Chinese, and may be duplicated. Rows with user_id = NULL are built-in agents, user rows are private unless published = true, and users can only update or delete agents they created.

The built-in forced-research-responder (「强制检索输出助手」) is seeded at Gateway startup. Unlike an ordinary skill-enabled custom agent, it has a runtime-enforced per-turn collection gate: the model cannot finish a response until a retrieval/knowledge/MCP tool or a skill Python script returns usable information. It then answers directly in the user's requested chat format rather than generating an artifact. The agent may create or edit only temporary .py files below /mnt/user-data/workspace/ and may execute .py skill scripts through bash; writes to outputs, Markdown/document creation, present_files, general shell commands, and skill mutation are denied. Administrators can edit its model and skill allowlist from agent management; the runtime safety policy remains fixed. Regression coverage: tests/test_forced_research_agent.py.

AI Writing Intranet Retrieval

AI 写作的「知识库检索」会直连 ai_writing.intranet_search_url 内网 ES/知识库接口,并默认跳过代理环境变量和 HTTPS 证书校验,适配自签名或内部 CA 证书的内网服务。需要强校验时可在 config.yaml 里设置 ai_writing.intranet_verify_ssl: true。通用检索里的非内置技能在 ai_writing.researcher_skill_agent: true 时会用真实 agent 运行时执行技能:以 ai-writing-researcher 身份读取该技能的 SKILL.md,并按技能要求通过 bash 运行脚本,弱模型也会在本轮任务里收到明确的读文件和脚本执行指令。

Installable Sentiment Analysis Skill

skill-packages/sentiment-analysis.skill is a separately installable skill package. It keeps the external AG-UI gateway address in config/sentiment-agent.json, resolves Basic Auth from deployment environment variables, and normalizes the upstream SSE stream into NDJSON (thinking_delta / answer_delta / done) for callers that need incremental rendering. It does not alter the existing frontend sentiment-analysis page. See skill-packages/sentiment-analysis/docs/调用与安装说明.md for installation and invocation.

IM Channels

The IM bridge supports Feishu, Slack, and Telegram. Slack and Telegram still use the final runs.wait() response path, while Feishu now streams through runs.stream(["messages-tuple", "values"]) and updates a single in-thread card in place.

For Feishu card updates, DeerFlow stores the running card's message_id per inbound message and patches that same card until the run finishes, preserving the existing OK / DONE reaction flow.


Quick Start

Prerequisites

  • Python 3.12+
  • uv package manager
  • API keys for your chosen LLM provider

Installation

cd deer-flow

# Copy configuration files
cp config.example.yaml config.yaml

# Install backend dependencies
cd backend
make install

Configuration

Edit config.yaml in the project root:

models:
  - name: gpt-4o
    display_name: GPT-4o
    use: langchain_openai:ChatOpenAI
    model: gpt-4o
    api_key: $OPENAI_API_KEY
    supports_thinking: false
    supports_vision: true

  - name: gpt-5-responses
    display_name: GPT-5 (Responses API)
    use: langchain_openai:ChatOpenAI
    model: gpt-5
    api_key: $OPENAI_API_KEY
    use_responses_api: true
    output_version: responses/v1
    supports_vision: true

Set your API keys:

export OPENAI_API_KEY="your-api-key-here"

For a LiteLLM OpenAI-compatible proxy, keep supports_thinking: false unless the configured model has an explicit, proxy-supported thinking control. This prevents DeerFlow's connection check from sending vendor-specific thinking fields that a generic OpenAI-compatible gateway may reject.

Running

Full Application (from project root):

make dev  # Starts LangGraph + Gateway + Frontend + Nginx

Access at: http://localhost:2026

Backend Only (from backend directory):

# Terminal 1: LangGraph server
make dev

# Terminal 2: Gateway API
make gateway

Direct access: LangGraph at http://localhost:2024, Gateway at http://localhost:8001


Project Structure

backend/
├── src/
│   ├── agents/                  # Agent system
│   │   ├── lead_agent/         # Main agent (factory, prompts)
│   │   ├── middlewares/        # 9 middleware components
│   │   ├── memory/             # Memory extraction & storage
│   │   └── thread_state.py    # ThreadState schema
│   ├── gateway/                # FastAPI Gateway API
│   │   ├── app.py             # Application setup
│   │   └── routers/           # 6 route modules
│   ├── sandbox/                # Sandbox execution
│   │   ├── local/             # Local filesystem provider
│   │   ├── sandbox.py         # Abstract interface
│   │   ├── tools.py           # bash, ls, read/write/str_replace
│   │   └── middleware.py      # Sandbox lifecycle
│   ├── subagents/              # Subagent delegation
│   │   ├── builtins/          # general-purpose, bash agents
│   │   ├── executor.py        # Background execution engine
│   │   └── registry.py        # Agent registry
│   ├── tools/builtins/         # Built-in tools
│   ├── mcp/                    # MCP protocol integration
│   ├── models/                 # Model factory
│   ├── skills/                 # Skill discovery & loading
│   ├── config/                 # Configuration system
│   ├── community/              # Community tools & providers
│   ├── reflection/             # Dynamic module loading
│   └── utils/                  # Utilities
├── docs/                       # Documentation
├── tests/                      # Test suite
├── langgraph.json              # LangGraph server configuration
├── pyproject.toml              # Python dependencies
├── Makefile                    # Development commands
└── Dockerfile                  # Container build

Configuration

Main Configuration (config.yaml)

Place in project root. Config values starting with $ resolve as environment variables.

Key sections:

  • models - LLM configurations with class paths, API keys, thinking/vision flags
  • tools - Tool definitions with module paths and groups
  • tool_groups - Logical tool groupings
  • sandbox - Execution environment provider
  • skills - Skills directory paths plus optional prompt routing; skills.es_query_routing can inject an editable system-prompt rule that sends flexible Elasticsearch Query DSL requests to the es_query skill, and can be disabled without changing code. Each ordinary Q&A agent build logs ES query routing prompt registration with registered, eligibility, character count, and a short SHA-256 fingerprint so operators can verify the exact configured block reached the final model system prompt without logging its contents
  • title - Auto-title generation settings
  • summarization - Context summarization settings
  • subagents - Subagent system (enabled/disabled)
  • memory - Memory system settings (enabled, storage, debounce, facts limits)

Provider note:

  • models[*].use references provider classes by module path (for example langchain_openai:ChatOpenAI).
  • If a provider module is missing, DeerFlow now returns an actionable error with install guidance (for example uv add langchain-google-genai).

Browser CORS

The Gateway permits the local Vite development origins on ports 5173, 5174, 3000, and 8080 for both localhost and 127.0.0.1 with credentials, even when GATEWAY_CORS_ORIGINS is set. For a separate deployed frontend, set GATEWAY_CORS_ORIGINS to its exact origin (including scheme and port), for example http://47.88.25.99:7010.

Extensions Configuration (extensions_config.json)

MCP servers and skill states in a single file:

{
  "mcpServers": {
    "github": {
      "enabled": true,
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-github"],
      "env": {"GITHUB_TOKEN": "$GITHUB_TOKEN"}
    },
    "secure-http": {
      "enabled": true,
      "type": "http",
      "url": "https://api.example.com/mcp",
      "oauth": {
        "enabled": true,
        "token_url": "https://auth.example.com/oauth/token",
        "grant_type": "client_credentials",
        "client_id": "$MCP_OAUTH_CLIENT_ID",
        "client_secret": "$MCP_OAUTH_CLIENT_SECRET"
      }
    }
  },
  "skills": {
    "pdf-processing": {"enabled": true}
  }
}

Environment Variables

  • DEER_FLOW_CONFIG_PATH - Override config.yaml location
  • DEER_FLOW_EXTENSIONS_CONFIG_PATH - Override extensions_config.json location
  • Model API keys: OPENAI_API_KEY, ANTHROPIC_API_KEY, DEEPSEEK_API_KEY, etc.
  • Tool API keys: TAVILY_API_KEY, GITHUB_TOKEN, etc.

LangSmith Tracing

DeerFlow has built-in LangSmith integration for observability. When enabled, all LLM calls, agent runs, tool executions, and middleware processing are traced and visible in the LangSmith dashboard.

Setup:

  1. Sign up at smith.langchain.com and create a project.
  2. Add the following to your .env file in the project root:
LANGSMITH_TRACING=true
LANGSMITH_ENDPOINT=https://api.smith.langchain.com
LANGSMITH_API_KEY=lsv2_pt_xxxxxxxxxxxxxxxx
LANGSMITH_PROJECT=xxx

Legacy variables: The LANGCHAIN_TRACING_V2, LANGCHAIN_API_KEY, LANGCHAIN_PROJECT, and LANGCHAIN_ENDPOINT variables are also supported for backward compatibility. LANGSMITH_* variables take precedence when both are set.

Langfuse Tracing

DeerFlow also supports Langfuse observability for LangChain-compatible runs.

Add the following to your .env file:

LANGFUSE_TRACING=true
LANGFUSE_PUBLIC_KEY=pk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_SECRET_KEY=sk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_BASE_URL=https://cloud.langfuse.com

If you are using a self-hosted Langfuse deployment, set LANGFUSE_BASE_URL to your Langfuse host.

Dual Provider Behavior

If both LangSmith and Langfuse are enabled, DeerFlow initializes and attaches both callbacks so the same run data is reported to both systems.

If a provider is explicitly enabled but required credentials are missing, or the provider callback cannot be initialized, DeerFlow raises an error when tracing is initialized during model creation instead of silently disabling tracing.

Docker: In docker-compose.yaml, tracing is disabled by default (LANGSMITH_TRACING=false). Set LANGSMITH_TRACING=true and/or LANGFUSE_TRACING=true in your .env, together with the required credentials, to enable tracing in containerized deployments.


Development

Commands

make install    # Install dependencies
make dev        # Run LangGraph server (port 2024)
make gateway    # Run Gateway API (port 8001)
make lint       # Run linter (ruff)
make format     # Format code (ruff)

Code Style

  • Linter/Formatter: ruff
  • Line length: 240 characters
  • Python: 3.12+ with type hints
  • Quotes: Double quotes
  • Indentation: 4 spaces

Testing

uv run pytest

Runtime appearance themes

The admin appearance settings are persisted in .deer-flow/system_settings.json. Besides the built-in themes, an administrator can create up to 12 runtime themes through PUT /api/system-settings/appearance. Each theme stores a safe six-digit primary colour, an optional primary-button text_color, light/dark mode, name, enabled flag and sort order. The frontend derives the rest of the visual tokens; the optional overrides object only exposes ten interaction colours (button_hover, navigation_gradient_start, navigation_gradient_end, selected_background, border, muted, muted_text, destructive, brand_title, primary_foreground) for exceptional brand requirements. The two navigation values override the automatically derived top-bar gradient. A custom theme id must use the custom-theme- prefix. Disabled or removed custom themes cannot remain the system default and automatically fall back to light.

Compact top-bar shortcuts use the same appearance endpoint. In addition to login and theme query parameters, each shortcut may include ordered keyless path parameters. The frontend resolves each selected login field and appends it as an encoded path segment (including inside a #/… hash route); password remains an ordinary login query parameter with the key password.

Citation display settings

Reference-mode entry points are controlled by citation_display in .deer-flow/system_settings.json. reference_mode_enabled defaults to false; when disabled, the frontend hides the "参考文献" composer button in both the normal workspace and iframe chat, and hides the per-user RAG citation toggle from appearance settings. Conversations with selected LLMWiki knowledge bases still show their source documents in the reference panel even when the manual entry is hidden. GET /api/system-settings/citation-display is readable by logged-in users so the UI can decide whether to expose the entry; PUT /api/system-settings/citation-display is admin-only and supports partial updates for both reference_mode_enabled and cleanup_words.

Ordinary Q&A Markdown format switch

Administrators can opt into the formal long-answer Markdown prompt from the admin "问答提示词配置" panel. The value is persisted as prompt_prefix.ordinary_qa_markdown_format_enabled in .deer-flow/system_settings.json and defaults to false. When enabled, only ordinary Q&A receives the # 主标题 / ## 一、标题 / ### (一) 标题 / #### 1. 标题 rules; writing mode, notebook mode and scheduled runs remain unaffected. The change takes effect when the next prompt is built.


Technology Stack

  • LangGraph (1.0.6+) - Agent framework and multi-agent orchestration
  • LangChain (1.2.3+) - LLM abstractions and tool system
  • FastAPI (0.115.0+) - Gateway REST API
  • langchain-mcp-adapters - Model Context Protocol support
  • agent-sandbox - Sandboxed code execution
  • markitdown - Multi-format document conversion
  • tavily-python / firecrawl-py - Web search and scraping

Documentation


License

See the LICENSE file in the project root.

Contributing

See CONTRIBUTING.md for contribution guidelines.

Position roundtable durable history

The position-roundtable workflow stores database-owned recovery snapshots for Step-1 intent chat, every position-node chat, the summary/action-plan chats, unsent composer drafts, and generated text deliverables. LangGraph checkpoints remain the execution source, but SQL history continues to render after a checkpoint or sandbox-output volume is replaced. Apply Alembic revision 20260723_05 (or simply upgrade to head) before deploying this version.

Its live intent conversation still uses the shared /api/intent/init and /api/intent/stream endpoints used by the original multi-agent roundtable. Position roundtable identifies itself with intent_agent_id: "position-roundtable-intent"; this is a separately seeded built-in agent, not an alias of roundtable-intent. It preserves the same thread, SSE streaming, and file-reading lifecycle as the generic Step-1 agent, but has its own at-most-five user-assistance-round brake. A round is one AI turn plus the user reply, not one tool call: the dedicated agent may emit several independent ask_clarification calls in the same turn, which the frontend renders as several separate selection cards and submits together. A topic-only request is not considered resolved: the agent asks for every independent missing decision (such as scope, decision use, or task constraints) as separate cards. For the dedicated agent, a fresh ask_clarification emitted in a stream is a strict user-input pause and wins over any same-run premature [INTENT_READY]; a historical clarification from the prior turn does not block the user answer from completing. Its final [INTENT_READY] JSON contains four string arrays: coreGoals, riskWarnings, keyPoints, and strategicSignificance; it must contain completed conclusions, not “待确认” or other unresolved placeholders. The multi-agent endpoint remains backwards compatible when this optional field is omitted and continues to emit objective, constraints, and assumptions. /api/intent/init creates its empty thread directly through the in-process thread store/checkpointer, matching multi-agent initialization and avoiding a deployment-sensitive 127.0.0.1:<gateway-port> loopback. If initial checkpoint creation fails or is cancelled, bootstrap removes both the partial checkpoint and its thread metadata row. Thread metadata creation issues no post-commit ORM read-back on any code path: all returned fields are assigned before commit, and skipping the extra read keeps a MySQL read/write-splitting endpoint from sending it to a lagging replica (which fails the request with Could not refresh instance). Only the recovery snapshot layer is position-specific. Deployments must expose either /api/position-roundtable/* or the compatibility mirror /api/multi-agent/position-roundtable/*; updated clients probe both and only retry when the first response is the framework/proxy-level 404 Not Found.

Position-roundtable role directory

The position-roundtable work views are persisted separately from organisation positions/RBAC in position_roles. Gateway startup idempotently inserts the four built-in ids (testing analysis, organisation test, solution design and comprehensive test), but only when an id is missing: administrator changes to the name, order, enabled state or role type are never overwritten.

GET /api/position-roles is consumed by the workbench and PUT /api/position-roles is administrator-only. Exactly one enabled intent role is required. It is the initial intent-identification workspace and owns the summary/action-plan closing agents. Built-in rows retain stable ids so historical chain nodes remain readable; they can be edited or disabled but not deleted. Custom roles may be freely added or removed.

Roundtable artifact submission status

The task-scoped roundtable workbench exposes a lightweight submission-status API for generated deliverables:

  • GET /api/roundtable/tasks/{taskId}/artifacts?submitted=true returns the submitted artifact references for a task (or items: []).
  • PUT /api/roundtable/tasks/{taskId}/artifacts/submission records submitted: true or cancels it with submitted: false.

Both routes use the normal Gateway Authorization: Bearer <token> middleware. The record is shared by taskId; the authenticated user is retained only in audit columns. Its server-owned idempotency key is taskId + sourceKey + path + revision. Repeating the same submitted payload returns the existing record; cancellation updates its state and timestamp without deleting the audit row. DeerFlow stores only the artifact reference (threadId + sandbox path, plus the optional session/node/name/MIME fields), never the Markdown body and never a task-system delivery. Continue to preview the file through the existing thread-artifact endpoint.

普通知识库管理

在「普通知识库」的「我的知识库」或「公共 LLMWiki」中,知识库所有者和管理员可以点击卡片上的「编辑」,修改名称和描述(描述可清空)。保存后同步更新知识服务和本地记录,并刷新列表;系统「对话沉淀」库不提供编辑入口。

远端已经删除的知识库会在下次列表刷新时自动隐藏,包括个人、公共列表和知识库选择器;页面每 15 秒自动刷新,也可以手动点击「刷新」。远端服务请求失败时展示错误,不把连接故障当成删除,不自动删除本地映射或已有引用。

DeerFlow-local LLMWiki Wiki vector index

LLMWiki can mirror processed WeKnora Wiki pages into DeerFlow and search only processed Wiki content. A knowledge base uses exact NumPy cosine similarity only after its whole Wiki has completed one coherent vectorization pass; until then internal search calls WeKnora's Wiki-page search/list API and reads the matched Wiki pages. It never falls back to raw documents or chunks. Configure llmwiki.local_wiki_index.embedding, then enable the local index in config.yaml. auto_sync defaults to true: the scheduler performs one immediate incremental scan at startup and then scans every configured interval (30 seconds by default). Wiki changes are content-hash checked, so unchanged pages reuse their vectors instead of being encoded again. database_batch_size caps each vector insert statement (default 50) while the page replacement remains atomic, preventing a large Wiki from flooding the SQL connection in one statement. The optional embedding.dimensions value is used only to validate the returned vector size and build the index fingerprint; it is deliberately not sent in the embedding request because fixed-dimension BGE/OpenAI-compatible endpoints may reject that extra field with HTTP 400. Omit it to discover the dimension from the first successful response.

Embedding failures log the upstream HTTP status and a bounded response body. The client never sends a dimensions request parameter.

The migration creates llmwiki_wiki_pages, llmwiki_wiki_vectors, and llmwiki_wiki_sync_states. Wiki Markdown and normalized float32 vectors are stored in DeerFlow's configured SQL database. Multi-worker sync is protected by database leases; searches use revision-checked per-knowledge-base LRU snapshots after the full-library readiness gate succeeds. Knowledge-base owners can click 手动向量化 in Wiki management (POST /api/llmwiki/knowledge-bases/{id}/vectorize) and follow its status at GET /api/llmwiki/knowledge-bases/{id}/wiki-index-status. The response includes aggregate progress plus a pages list with every Wiki page's index state, vector-section count, timestamps, and last error. The management page renders the percentage and labels each article as 已向量化 / 待向量化 / 失败, so it is clear exactly what is searchable. Existing chunk citations remain readable while all new retrieval citations open Wiki pages.

The frontend labels these results simply as Wiki references; provider-brand wording is intentionally hidden from user-facing labels and status text. Clicking any part of a Wiki reference card opens one right-side drawer (rather than a second detail dialog), renders the complete page as Markdown, and reads the remote WeKnora Wiki page directly when the local index is disabled. Even if a Wiki citation carries a legacy document id, the drawer does not mix in the uploaded original-file preview. A single page can be downloaded as Markdown or Word from the drawer. Open in Wiki navigates to the exact DeerFlow knowledge-base mapping and Wiki slug. Wiki search-result cards use the same drawer interaction. Exact Wiki links now open a lightweight native reader at knowledge/{mapping_id}/wiki/view?wikiSlug=...; it fetches only the selected page and never creates a WeKnora iframe session. The former knowledge/{mapping_id}?tab=wiki&wikiSlug=... links redirect to the reader before waiting for the LLMWiki runtime check. Drawer navigation stays in the same SPA and passes the already-loaded page as route state for immediate paint. The native reader also keeps a searchable, collapsible Wiki directory on the left. It reconstructs hierarchy from parent_slug, wiki_path, category_path, and the slug fallback, highlights and expands the current page, and switches directory entries through the same lightweight reader route without mounting the WeKnora iframe.

All generated SPA deep links preserve the deployment document path (for example /web/) before the # route, including Q&A references and Wiki index links. Protected Wiki images no longer render from temporary browser blob: URLs. Public images use the published file proxy directly; private images first obtain a 24-hour signed capability and then load through GET /api/public/llmwiki/files/{token} with a real HTTP URL and inline filename, allowing zoom, new-tab viewing, and browser Save As without exposing the user's bearer token. The embedded WeKnora detail page follows the same rule: late img[data-protected-src] placeholders keep their scoped HTTP proxy URL on the DOM node, so WeKnora's click-to-enlarge viewer never receives an already-revoked blob: URL.

Leaderboard query and schema startup safety

The leaderboard excludes scheduled runs with indexed correlated NOT EXISTS checks. Tool metrics no longer outer-join the full runs table merely to recognize legacy scheduler messages. New databases receive dimension-aware analytics indexes from ORM metadata; existing large MySQL deployments should run scripts/sql/leaderboard_indexes_mysql.sql during a low-traffic window so index creation never delays application startup.

Startup Alembic migrations now have one unambiguous head. The historical duplicate 20260821_01 identifier is repaired by unique fixed-question revisions and the 20260825_01 merge revision, which also fills the report-structure column for databases stamped by either old variant.

Published knowledge bases also have a login-free, chrome-free reader at /#/embed/knowledge/{mapping_id} for direct third-party iframe use. With no wikiSlug, it selects the Wiki 索引 page first (recognized by title, page type, or index/home/readme slug) and pins that page to the top of the public directory. If WeKnora has pages but no explicit index page, the reader creates a virtual __index__ landing page from the public directory; empty dynamic-index groups also fall back to those page summaries; ?wikiSlug=... deep-links one page. The reader supports theme=light|dark, bg, and text embed appearance parameters. Index-article catalogue sections (page-type counts, nested categories, page summaries) are read from the published-only /wiki/index proxy because WeKnora renders that part dynamically rather than storing it in the index page body; group and item ordering therefore stays identical to WeKnora. Links generated as knowledge-base-shell placeholders are matched against the public page title, slug, or alias and rewritten to exact embedded wikiSlug links. They retain WeKnora's brand-color medium text and dashed underline (solid on hover), and switch articles inside the current reader without a full-page reload. Data comes only from the read-only /api/public/llmwiki/knowledge-bases/{id}/... namespace, which rechecks publication_status=published on every request and denies the system conversation-deposit knowledge base. Private or unpublished bases are never exposed by this route. Anonymous consumers can discover the available public Wiki libraries through GET /api/public/llmwiki/knowledge-bases. The endpoint validates each published mapping against WeKnora, omits stale remote mappings, and returns safe comprehensive metadata: description/type, document/chunk/processing/share counts, non-secret configuration and timestamps, Wiki page/type/status counts plus explicit/generated index mode and default embed route, and aggregate local-vector state/counts/timestamps. It never returns the WeKnora id, tenant/creator identity, storage credentials, API credentials, raw vectors, or per-page indexing errors. Protected resource:// and supported object-storage images are fetched through the current mapping instead of being exposed as upstream URLs. Authenticated readers use GET /api/llmwiki/knowledge-bases/{id}/files?file_path=...; public readers use the corresponding /api/public/llmwiki/... route, which requires a published mapping. The embedded WeKnora guard applies the same mapping-scoped proxy to document previews and Markdown images, including images inserted after the drawer/page first renders.

Wiki management also provides 导出 Excel. GET /api/llmwiki/knowledge-bases/{id}/wiki/export.xlsx exports every processed Wiki page, including its full directory path, current directory, parent-directory path, one column per directory level, metadata, and complete Markdown content (split across continuation columns when Excel's per-cell limit is reached).

Authenticated operations are available below /api/llmwiki/wiki-index and POST /api/llmwiki/wiki-vector-search. The service-to-service endpoints below /api/external/llmwiki/wiki require X-API-Key, are rate-limited, and only expose mappings that are both internally published and explicitly marked external_search_enabled; the system 对话沉淀 mapping is always denied.

A separate browser-friendly endpoint, POST /api/knowledge/vector-search, is fully anonymous: it requires neither a DeerFlow token nor X-API-Key, has no rate limit, and responds with Access-Control-Allow-Origin: *. It searches the local vectors of every published knowledge base regardless of external_search_enabled, while still excluding the system 对话沉淀 base. The minimal request is {"query":"如何部署"}; optional fields are top_k (default 10, maximum 100) and include_page_content (default false). Each matched vector section is returned as its own ranked result, including the matched text, score, DeerFlow and WeKnora knowledge/Wiki ids, page metadata, raw source/chunk references, best-effort document/chunk ids, and a directly usable wiki_url. Set llmwiki.local_wiki_index.frontend_base_url to the deployed frontend origin/base path; its development default is http://127.0.0.1:5174.

See docs/LLMWIKI_DEERFLOW_LOCAL_VECTOR_INDEX_ZH.md for configuration, security boundaries, rollout, and API contracts.

Document rewrite fallback

Whole-document rewrite commits the artifact/report before it attempts to store the optional reversible before-image. If that version-snapshot write fails, the rewritten content remains available and Gateway logs the exception; it never restores the original solely because undo history is unavailable. The durable result marks versionSnapshotFailed, allowing the sandbox to enable its normal manual-save action without showing the user an error notification.

The optional rewrite-history hydration endpoints scope results to the authenticated job owner. They deliberately return an empty list rather than performing a second thread-owner check, so a legacy guest/SSO display alias cannot make an otherwise usable sandbox report fail with a history-only 404.

Artifact and rewrite path resolution uses the persisted thread creator's directory when metadata exists, after the route's normal ownership check. This prevents a missing request context from incorrectly looking under users/default and returning 404 for an existing sandbox file.

The final rewrite commit uses a deliberately short sibling temporary filename, which keeps long Windows artifact paths below the path-length limit. It also recreates a missing outputs parent before retrying, so transient sandbox-directory recreation cannot discard a completed rewrite.

Offline local-sandbox Office runtime

The pure DeerFlow offline image treats the backend container as the default LocalSandboxProvider runtime. It bundles MarkItDown, python-docx, openpyxl, python-pptx, LibreOffice Headless, Pandoc, Poppler, and CJK fonts. Modern OOXML uploads are parsed directly; legacy DOC/XLS/PPT and ODT/ODS/ODP files are normalized through an isolated LibreOffice profile before Markdown extraction. Run ./OFFLINE_INSTALL.sh --office-check from the offline bundle to generate and re-parse Word, Excel, and PowerPoint samples without requiring a database.

Skill knowledge distillation and assistant knowledge

The Gateway now exposes an administrator-only skill distillation pipeline at /api/skill-knowledge. Each request can send up to 200 skills to one writable WeKnora Wiki knowledge base; multi-base distillation is done by repeated operations so the local binding remains skill_name + target_id. The scanner never reads or uploads code/script files, strips Markdown code fences, blocks credential-like files and content, safely inspects ZIP members, and converts supported Office/PDF/spreadsheet files off the event loop. Jobs use database claims and renewable leases, stage snapshots before activation, retain the prior active version on failure, support per-item retry, and keep per-binding contribution records for safe detach/delete. Direct assistant-base targets, non-Wiki/FAQ/conversation/readonly targets, and direct Neo4j writes are rejected for this product version.

/api/assistant-knowledge manages DeerFlow-native assistant knowledge bases. There is one global full-library base named 知识梳理总库, plus admin-created custom assistant bases. Reads and search are available to authenticated users; global initialization, custom-base creation, WeKnora export/import, source reimport, and source deletion are admin-only. Assistant bases only accept WeKnora export packages or imports triggered from a WeKnora mapping; direct skill import and arbitrary local-file import return 410. Importing into a custom assistant base also mirrors the same package into the global base, where canonical entities and relations are aligned across all imported sources. The database schema is introduced by Alembic revision 20260901_01.

Wiki 流转与本地小批量编码

所有登录用户都可在「技能 → 归纳到知识库」中多选自己可见的技能,并归纳到一个 自己可写的 WeKnora Wiki 知识库;用户只能看到和维护自己发起的归纳记录,管理员 接口仍可查看全量记录。技能归纳上传清理后的知识正文,等待 WeKnora 完成 Wiki 分析,再建立本地 Wiki 向量并通过落盘导出包自动沉淀到「知识梳理」。内置技能没有数据库所有权记录 也可以归纳。若解析结束却没有 Wiki,会提示检查生成模型和资料;文档解析完成 不等于 Wiki 生成成功,例如生成服务余额不足仍可能表现为文档已完成。

管理员可以在普通知识库直接生成向量数据包,每次操作有独立后台记录和下载入口。 助手知识库支持上传包、异步导入和导出;新建空的普通 Wiki 库可上传助手导出包。管理员可编辑助手知识库名称与描述,也可删除自建助手知识库及其 Wiki、向量、实体、关系、导入记录和落盘数据包;承担全库沉淀的「知识梳理总库」允许编辑但禁止删除。 包保留正文、别名/标签、目录分类、双向链接、实体关系及 float32 向量。 回流时通过 WeKnora API 重建目录/页面,向量进入 DeerFlow 本地 Wiki 索引, 不直接写 WeKnora 的底层向量或图数据库。普通与助手 Wiki 使用同一个阅读组件。

小批量 CPU 编码可配置如下,无需额外部署 HTTP 编码服务:

llmwiki:
  local_wiki_index:
    enabled: true
    auto_sync: true
    sync_interval_seconds: 30
    embedding:
      provider: local
      model: bge-embedding-m3
      dimensions: 1024
      batch_size: 2
      local_threads: 2
      local_process_isolation: true
      # 离线部署时可指定已下载的 BAAI/bge-m3 模型目录:
      # local_model_path: /opt/models/bge-m3

本地编码依赖随标准后端安装。将完整 BGE-M3 模型目录提前放到离线服务器,并通过 local_model_path 或 DEERFLOW_BGE_M3_MODEL_PATH 指向它。管理员也可在 「设置和更多 → 设置 → 本地编码器」中选择服务器目录,或从浏览器上传完整模型目录; 页面配置持久化到运行目录的 system_settings.json,上传文件落到 .deer-flow/models/bge-m3,均不进入源码仓库。系统强制 local_files_only, 不会访问互联网或下载模型。管理员可在同一设置页点击「一键启动编码器」异步加载;该操作不挂在应用启动流程,加载失败不会 中断 DeerFlow 主服务。默认由独立低优先级子进程承载模型内存,正常停止 DeerFlow 时会一并回收该子进程;编码进程异常退出只会将当前向量任务标记失败。模型目录应包含 onnx/model.onnx(官方权重约 2.3 GB)及对应 tokenizer/config 文件。权重不进入业务源码仓库。 浏览器目录上传采用磁盘暂存和 4 MB 分块复制,页面显示上传进度;经 Nginx 等反向 代理部署时仍需把请求体上限和读写超时配置到可承载完整模型目录。 相同模型指纹、相同内容的页面复用现有向量;短正文/摘要至少保留一个非空分段。 数据包导入不重新编码,模型指纹或维度不一致时明确拒绝,不能只按维度混用模型。 问答编码查询本身仍需可用的同一模型。WeKnora 的原始文档向量不能当作生成后 Wiki 正文的向量使用,两种文本需要分别建立各自的首次索引。

Wiki 页面新增、编辑、删除及文档上传/重解析都会立即触发一次增量检查;WeKnora 后台稍后才完成的 Wiki 抽取由 30 秒周期补扫接管。每次成功向量化后自动把该普通 知识库的最新完整快照同步到「知识梳理总库」,包括服务重启前已有库和新建库;来源 页面被删除时也会在总库中退役,不保留失效 Wiki。

「知识梳理总库」不是按来源简单堆叠:导入时会规范化全/半角、空白、标点及常见实体 类型,并利用实体别名对齐同一人物、组织、产品或概念;不同来源中同标题的实体 Wiki 只保留一个规范页面,正文按去重后的知识块合并,来源与实体/关系贡献仍分别留痕。 重复上传同一个 WeKnora 导出包时,以包内稳定知识库 ID 识别同一来源并执行快照替换, 不会再产生一批新的片段副本。已有历史重复数据会在总库首次读取、检索、导出或下次 同步时自动归并。助手知识检索排除目录、索引和摘要页;关键词只查标题与知识正文, 向量检索返回实际命中的正文片段,问答上下文和检索卡片均不再用摘要代替知识。

大包采用流式 JSONL 落盘、同盘临时 SQLite 分阶段读取及异步导入,避免一次性将 整个包加载进内存。部署需预留包、暂存索引和业务数据库空间,并配置代理上传大小 与超时。当前不宣称断点续传或 GB 级压力验收已完成。

openGauss / GaussDB application database

The application ORM can use a PG-compatible openGauss/GaussDB database through the official opengauss-sqlalchemy async dialect. Install the optional driver set before enabling it:

uv sync --extra gauss

Configure the business database independently from the LangGraph checkpointer:

database:
  backend: gauss
  gauss_url: $GAUSS_DATABASE_URL

checkpointer:
  type: sqlite
  connection_string: .deer-flow/data/deerflow.db

GAUSS_DATABASE_URL should normally be opengauss+asyncpg://user:password@host:26000/deerflow. Plain opengauss://, gaussdb://, and PostgreSQL-style URLs are accepted and normalized to the openGauss dialect. Do not configure the LangGraph PostgreSQL checkpointer against GaussDB: its migrations are PostgreSQL-specific. For a production image, build with --build-arg UV_EXTRAS=gauss.