# DeerFlow Backend
普通知识库默认详情页面:管理员可在知识库顶部选择“本系统 Wiki 索引”或“嵌入 WeKnora”,对应 `ordinary_knowledge_landing`(默认 `wiki`)。Wiki 库首次进入默认展示本系统索引,右上角“打开 WeKnora”切换到内嵌管理页;内嵌页可返回索引。非 Wiki 库仍进入 WeKnora 文档页。设置通过 `/api/system-settings/ordinary-knowledge-landing` 保存,仅管理员可修改;指定文章、文档、图谱的直达链接优先于默认设置。
助手知识库可独立于 WeKnora 使用:管理员进入助手知识库,选择或创建库后点击“上传文件生成 Wiki”。支持 Markdown/TXT/CSV、PDF、Word、Excel、PowerPoint(单文件 100 MB,解析正文上限 400 万字符)。文件先写入运行目录 `assistant-knowledge/files/{job_id}`,后台分段调用已配置的语言模型生成分类 Wiki、实体关系,再通过 DeerFlow 编码器对正文分段向量化,保存到助手库并汇入知识梳理总库。无需配置 WeKnora URL;配置后仍可直接上传,也保留 WeKnora 数据包导入。文件处理沿用管理员写入权限,所有用户可以浏览和搜索。
导入记录显示解析、Wiki 生成和向量化阶段;失败或服务中断后可“重试 / 恢复处理”,复用已落盘的 Wiki 和相同模型指纹的向量。离线编码器须预先放置模型并在管理员设置中配置;模型或编码器不可用会明确失败,不把摘要冒充 Wiki 或虚构向量。扫描件是否可解析取决于部署的文档解析/OCR 能力。
DeerFlow is a LangGraph-based AI super agent with sandbox execution, persistent memory, and extensible tool integration. The backend enables AI agents to execute code, browse the web, manage files, delegate tasks to subagents, and retain context across conversations - all in isolated, per-thread environments.
---
## Architecture
```
┌──────────────────────────────────────┐
│ Nginx (Port 2026) │
│ Unified reverse proxy │
└───────┬──────────────────┬───────────┘
│ │
/api/langgraph/* │ │ /api/* (other)
▼ ▼
┌────────────────────┐ ┌────────────────────────┐
│ LangGraph Server │ │ Gateway API (8001) │
│ (Port 2024) │ │ FastAPI REST │
│ │ │ │
│ ┌────────────────┐ │ │ Models, MCP, Skills, │
│ │ Lead Agent │ │ │ Memory, Uploads, │
│ │ ┌──────────┐ │ │ │ Artifacts │
│ │ │Middleware│ │ │ └────────────────────────┘
│ │ │ Chain │ │ │
│ │ └──────────┘ │ │
│ │ ┌──────────┐ │ │
│ │ │ Tools │ │ │
│ │ └──────────┘ │ │
│ │ ┌──────────┐ │ │
│ │ │Subagents │ │ │
│ │ └──────────┘ │ │
│ └────────────────┘ │
└────────────────────┘
```
**Request Routing** (via Nginx):
- `/api/langgraph/*` → LangGraph Server - agent interactions, threads, streaming
- `/api/*` (other) → Gateway API - models, MCP, skills, memory, artifacts, uploads, thread-local cleanup
- `/` (non-API) → Frontend - Next.js web interface
Health probes are always unauthenticated, even when API authentication is enabled: `GET|HEAD /health`, `/api/health`, and `/api/langgraph/health` all return the Gateway health status without requiring a cookie or bearer token.
---
## Core Components
### Lead Agent
The single LangGraph agent (`lead_agent`) is the runtime entry point, created via `make_lead_agent(config)`. It combines:
- **Dynamic model selection** with thinking and vision support
- **Middleware chain** for cross-cutting concerns (9 middlewares)
- **Tool system** with sandbox, MCP, community, and built-in tools
- **Subagent delegation** for parallel task execution
- **System prompt** with skills injection, memory context, and working directory guidance
### Workflow Studio runtime
- Candidate workflow graphs remain editable after they are loaded into the Coze-compatible canvas; the compatibility response explicitly marks the current owner as editable so node dragging is not mistaken for a read-only preview.
- Workflow-agent SSE forwards provider `reasoning_content` as an inline `…` trace for the UI while keeping the executable node output limited to the visible answer text. Every native LangGraph `[AIMessage|ToolMessage, metadata]` frame is additionally persisted and sent as `node.message.data.messages` without projection or length cropping, so the Studio can replay DeerFlow's tool-result, search-result and `ask_clarification` cards exactly after reconnecting.
### AgentScope report collaboration (in progress)
The independent `/api/report-collaboration` workbench supports durable run-time intervention commands and immutable formal report versions. Its production worker now runs a real AgentScope 2.0.5 `Agent` per TaskLedger assignment; AgentScope owns the member ReAct/tool/structured-output loop, while TaskLedger alone owns DAG readiness, retries and completion. Ready nodes in one wave run concurrently up to `max_parallel_tasks`; cancellation and member exceptions close agent runs as terminal states, and lease recovery remains independent of in-memory AgentScope objects. Retrieval uses the deployment's configured DeerFlow `web_search`/`web_fetch` tools through the frozen role permission snapshot. Every visible AgentScope reply, TeamSay handoff, tool call and tool result is persisted immediately into the existing SSE/message protocol; hidden reasoning and system prompts are not exposed. A constrained intent agent classifies explicit answers, angle additions, re-search, re-analysis, node reruns, section rewrites, polishing, report questions, replanning, and cancellation. Server-side impact analysis validates node/agent ownership and computes the minimum downstream lineage; low-confidence, high-cost, or multi-branch changes require confirmation. Commands are CAS-claimed and applied at TaskLedger safe points, where old attempts and artifacts become `superseded` so late results cannot contaminate the revised report. Run creation freezes the selected `/api/report-structures` row and all server-enforced budgets. Per-user active-run limits and per-run model-call, retrieval, source-byte, token, cost, node-time, and total-time limits fail closed. User-visible event payloads are centrally redacted before durable persistence without size truncation; external source content is explicitly untrusted, and hidden reasoning/system prompts are never exposed. Append-only audits cover agent selection, tools, sources, user commands, confirmations, report versions, redactions, and budget failures; owner-scoped reads are available at `GET /runs/{id}/audits`. Temporary Markdown is emitted as `report.delta`; only a template/citation/coverage-clean report with an independent passing review is atomically persisted and emitted as `report.version.created` before the run becomes completed. Section rewrites are replacement candidates; apply and restore append a new head version and never destroy history. Configure `runtime_mode`, concurrency and limits under `report_collaboration`; both the feature and worker remain disabled by default. Install the `report-collaboration` project extra when enabling the worker.
### Middleware Chain
Middlewares execute in strict order, each handling a specific concern:
| # | Middleware | Purpose |
|---|-----------|---------|
| 1 | **ThreadDataMiddleware** | Creates per-thread isolated directories (workspace, uploads, outputs) |
| 2 | **UploadsMiddleware** | Injects newly uploaded files into conversation context |
| 3 | **SandboxMiddleware** | Acquires sandbox environment for code execution |
| 4 | **SummarizationMiddleware** | Reduces context when approaching token limits (optional) |
| 5 | **TodoListMiddleware** | Tracks multi-step tasks in plan mode (optional) |
| 6 | **TitleMiddleware** | Auto-generates conversation titles after first exchange |
| 7 | **MemoryMiddleware** | Queues conversations for async memory extraction |
| 8 | **ViewImageMiddleware** | Injects image data for vision-capable models (conditional) |
| 9 | **ClarificationMiddleware** | Intercepts clarification requests and interrupts execution (must be last) |
### Sandbox System
Per-thread isolated execution with virtual path translation:
- **Abstract interface**: `execute_command`, `read_file`, `write_file`, `list_dir`
- **Providers**: `LocalSandboxProvider` (filesystem) and `AioSandboxProvider` (Docker, in community/)
- **Virtual paths**: `/mnt/user-data/{workspace,uploads,outputs}` → thread-specific physical directories
- **Skills path**: `/mnt/skills` → `deer-flow/skills/` directory
- **Skills loading**: Recursively discovers nested `SKILL.md` files under `skills/{public,custom}` and preserves nested container paths
- **Conditional WeKnora LLMWiki**: a non-empty `llmwiki.weknora.api_base_url` enables the DeerFlow-native WeKnora BFF and UI; an empty address preserves the existing LLMWiki page. WeKnora remains independently deployed and owns its database, Redis, storage, parsers and models. DeerFlow keeps only the API/admin addresses, user/publication mappings, agent knowledge-base bindings and per-conversation selections; it stores and sends no WeKnora credential, so the endpoint must be restricted to a trusted network and allow DeerFlow traffic. Owners publish directly to the public LLMWiki without an approval queue for ordinary bases; the system-owned conversation-deposit base is always published, shared by all users, excluded from My Knowledge, immutable through user/admin detail operations, and recognized by both its reserved name and `system` owner. An existing same-name remote mapping is adopted in place as the system deposit so its mapping id and document links stay valid. Its user-facing knowledge-base description and generated document metadata use the `cmzs` brand; list/detail reads synchronize stale remote descriptions, while conversation-deposit search results, source previews and embedded-detail JSON rewrite legacy brand text to `cmzs`. Deposit requests must reference a thread accessible to the current account. Personal knowledge-base lists are scoped to the current DeerFlow account even for admins, while frontend LLMWiki query caches also include the active account id so switching accounts cannot reuse another account's list/detail data. Admins remain the only role that sees the WeKnora connection-config tab. Knowledge cards open a dedicated WeKnora-style detail route: Wiki-enabled bases default to Wiki, then documents and graph; standard bases hide Wiki and default to documents. Writable ordinary knowledge-base details expose a native multi-file upload action that submits files sequentially and reports aggregate success/failure; read-only public bases and the system deposit base do not. It supports file/URL/manual-Markdown imports, Wiki page browsing, document preview, real parsed chunks and Wiki graph nodes/edges. New chats use a compact WeKnora-style searchable/grouped multi-select inside the composer, while agent create/edit pages bind spaces through an inline panel. A single proxy service identity means DeerFlow users are isolated logically while their WeKnora entities share that identity's workspace. The read-only `llmwiki_search` tool reauthorizes the current DeerFlow user on every call and renders credential-free source previews; source previews prefer cleaned Wiki/LLMWiki text when a retrieved chunk maps to a Wiki page, with raw chunks as the fallback. The detail page first creates a signed iframe session through the normal DeerFlow API auth path, then points the iframe directly at the proxied WeKnora `/platform/...` route using the signed `deerflow_weknora_embed` cookie; global auth/CSRF bypass is limited to those signed iframe proxy requests, and HTTPS cookie attributes honor `X-Forwarded-Proto` behind reverse proxies. The iframe proxy keeps WeKnora's root-relative `/platform`, `/assets`, `/locales`, and `/api/v1/*` paths, but is registered last through `llmwiki.proxy_router` so it cannot intercept first-party Gateway APIs such as `POST /api/v1/auth/login/username`. See [docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md](docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md).
- **Enterprise Research workbench (temporarily disabled)**: the independent `/api/enterprise-research` router, its report executor, and its ORM table registration are intentionally inactive. The implementation remains in the repository, but startup does not create or alter `enterprise_research_tasks` or `enterprise_research_report_jobs`; this keeps the active Deep Research workflow independent of the paused workbench and avoids MySQL schema errors.
- **Embedded WeKnora detail chrome**: the proxy hides WeKnora's application sidebar (`.main > .aside_box`) and expands the detail outlet to the full iframe width, while preserving the Wiki page's own index/navigation panel. The host does not sandbox this trusted, permission-filtered detail frame because browser PDF viewers are disabled inside sandboxed ancestor frames. Preview and knowledge-base file redirects are followed by the Gateway so PDF bytes and protected Wiki images remain inside the authorized proxy instead of leaking to an unreachable object-storage URL.
- **Embedded WeKnora knowledge-base switching**: the original breadcrumb picker works inside the iframe, but lists only the signed-in user's own and published knowledge bases. Every switch is reauthorized through DeerFlow and receives a fresh short-lived iframe session, so it never exposes the upstream service-admin knowledge-base list.
- **Prefixed WeKnora deployment**: the frontend may be published under a separate prefix such as `/magentweb`, while embedded WeKnora detail traffic is served by the Gateway under `/deerflow`. The Gateway rewrites HTML root assets, WeKnora API URLs, redirects, and Vite dynamic dependency maps so page assets stay under `/deerflow`; the prefixed proxy routes must be registered before the unprefixed fallback. The signed embed cookie covers `/deerflow`, and auth/CSRF recognition includes all supported static directories.
- **File-write safety**: `str_replace` serializes read-modify-write per `(sandbox.id, path)` so isolated sandboxes keep concurrency even when virtual paths match
- **Tools**: `bash`, `ls`, `read_file`, `write_file`, `str_replace` (`bash` is disabled by default when using `LocalSandboxProvider`; use `AioSandboxProvider` for isolated shell access)
### Subagent System
Async task delegation with concurrent execution:
- **Built-in agents**: `general-purpose` (full toolset) and `bash` (command specialist, exposed only when shell access is available)
- **Concurrency**: Max 3 subagents per turn, 15-minute timeout
- **Execution**: Background thread pools with status tracking and SSE events
- **Flow**: Agent calls `task()` tool → executor runs subagent in background → polls for completion → returns result
### Memory System
LLM-powered persistent context retention across conversations:
- **Automatic extraction**: Analyzes conversations for user context, facts, and preferences
- **Structured storage**: User context (work, personal, top-of-mind), history, and confidence-scored facts
- **Debounced updates**: Batches updates to minimize LLM calls (configurable wait time)
- **System prompt injection**: Top facts + context injected into agent prompts
- **Storage**: JSON file with mtime-based cache invalidation
### Tool Ecosystem
| Category | Tools |
|----------|-------|
| **Sandbox** | `bash`, `ls`, `read_file`, `write_file`, `str_replace` |
| **Built-in** | `present_files`, `ask_clarification`, `view_image`, `task` (subagent) |
| **Community** | Tavily (web search), Jina AI (web fetch), Firecrawl (scraping), DuckDuckGo (image search) |
| **MCP** | Any Model Context Protocol server (stdio, SSE, HTTP transports) |
| **Skills** | Domain-specific workflows injected via system prompt |
#### Configurable HTTP web search
The config-defined `web_search` tool can call a deployment-specific JSON API
instead of a public search provider. Set its `endpoint`, static `payload`, SSL
verification, timeout, result limit, and `result_url_template` in `config.yaml`.
At runtime only `payload.query` is overwritten with the model's search text.
The API response must use `{"results": [{"recUuid": "", "content": "",
"title": "", "url": ""}]}`; missing or null result fields are normalized to empty strings. When
`recUuid` is present, `{recUuid}` in `result_url_template` replaces the original
result URL; otherwise the original URL is retained. Set `enabled: false` to keep
the tool completely out of the agent's registered tool list. `verify_ssl: false`
supports trusted internal HTTPS services with self-signed certificates.
### Gateway API
FastAPI application providing REST endpoints for frontend integration:
| Route | Purpose |
|-------|---------|
| `GET /api/models` | List available LLM models |
| `GET/PUT /api/mcp/config` | Manage MCP server configurations |
| `GET/PUT /api/skills` | List and manage skills |
| `POST /api/skills/install` | Install skill from `.skill` archive |
| `POST /api/skills/validate`, `POST /api/skills/install-upload` | Validate and upload `.zip`/`.skill` packages with owner-aware conflicts: owners may explicitly overwrite their own copy; other users' copies report the uploader and cannot be overwritten |
| `GET/POST /api/agents` | List visible agents or create a user-owned agent |
| `GET/PUT/DELETE /api/agents/{id}` | Read, update, or delete an id-addressed agent |
| `GET/POST/PUT/DELETE /api/roundtable-chains` | Manage user-owned roundtable business chains; each seat may optionally carry `position_id` for the separate position-collaboration workspace, plus `human_validation` to mark a Step 2 client-side pause checkpoint after that seat/stage completes (the original roundtable has no extra validation card) |
| `GET/PUT /api/business-mapping` | Read global business mappings. Updates to 3Q/6BF/7BF require an administrator; the 8BF public business-chain mapping is maintained exclusively by the trusted `lqq` account (derived from the authenticated email prefix). |
| `GET/POST/PUT /api/position-roundtable/sessions` | Standalone position-collaboration sessions: freeze task/intent/chain snapshots, index one direct-agent node per seat, build upstream handoff context (including bounded UTF-8 excerpts of text artifacts), and make archived sessions read-only without invoking the original roundtable coordinator or jobs. Built-in summary/action-plan completion mirrors visible report text into a Markdown artifact when a model omits its final `write_file` call. `POST /sessions/{id}/nodes/{nodeKey}/reject` accepts an optional reason, marks the rejected delivery for rework, and invalidates affected downstream deliveries so they are regenerated in order. The same router is also mounted at `/api/multi-agent/position-roundtable/*` for intranet gateways that only forward the established multi-agent namespace; the compatibility mount is intentionally omitted from OpenAPI. |
| `GET /api/memory` | Retrieve memory data |
| `POST /api/memory/reload` | Force memory reload |
| `GET /api/memory/config` | Memory configuration |
| `GET /api/memory/status` | Combined config + data |
| `POST /api/threads/{id}/uploads` | Upload files (auto-converts PDF/PPT/Excel/Word to Markdown, rejects directory paths) |
| `GET /api/threads/{id}/uploads/list` | List uploaded files |
| `DELETE /api/threads/{id}` | Delete DeerFlow-managed local thread data after LangGraph thread deletion; unexpected failures are logged server-side and return a generic 500 detail |
| `GET /api/artifact-library` | List the current user's generated conversation files across normal threads, with filename/conversation search, file-kind filtering, pagination, and an optional cache-bypassing refresh. The filesystem `outputs/` directories remain authoritative; uploads and image/audio/video files are excluded. |
| `GET /api/threads/{id}/artifacts/{path}` | Serve generated artifacts |
| `GET/POST /api/deep-research/*` | Durable Deep Research sessions and jobs. `quick`/`basic` run the compact report flow; `detailed` plans subtopics and writes independent sections; `deep` recursively investigates bounded follow-up branches. All modes reuse DeerFlow's configured material provider, LLM factory, hidden-thread artifacts, and SSE event log. Deep Research session/job/event/source/message writes serialise their rows in the writer transaction, rather than doing a post-commit ORM refresh or readback; this prevents MySQL read/write-splitting replica lag from failing “start writing”, dispatcher claim, report persistence, or follow-up history. Whole-document artifact/report rewrites likewise create their reversible version snapshots from writer-local ORM data and never refresh after commit, so replica lag cannot restore an otherwise successful rewrite. A report is written to the session and confirmed before its Job becomes `completed`; an unconfirmed write is retried and never exposed as a false completion. Retrying an already-created queued job re-nudges the dispatcher, recovering an interrupted HTTP response without creating a second report. If progress-event persistence, durable replay, or live fan-out is briefly unavailable, the worker/stream logs the degradation and continues or reconnects rather than failing the whole job. Whole-report rewrite accepts an optional `reportOutline` Markdown template; the durable job persists it and applies it as the chapter structure constraint. Optional report illustrations use a separately configured OpenAI-compatible image endpoint and are served only through the owning session's authenticated artifact route. |
| `/api/enterprise-research/*` | Temporarily disabled. Its router and persistence initialization are not registered, so these endpoints return 404 and no enterprise-research tables are created or migrated during startup. |
| `POST /api/ai-writing/sample/extract` | Extract text from Word/PDF/Markdown/TXT samples for AI-writing imitation mode |
| `POST /api/writing/export/docx` | Generate the shared formal Word download from Markdown. Uses the standard-library OOXML builder in `app/gateway/word_export.py`, removes manual/compound heading prefixes and Markdown horizontal rules, applies the prescribed A4/margin/heading/body/table/footer profile, and embeds the server-bundled `方正小标宋简体.ttf`; the client computer does not need that font installed. |
| `POST /cop/saveSuperiorTask` | TaskCOP compatibility API (DeerFlow Bearer/session authentication required): create a situation-overview task with a sequential four-digit id (`0001`, `0002`, ...) and return the legacy `{state, msg, data}` envelope |
| `POST /api/taskcop/import/tasks` | TaskCOP compatibility API (authenticated): import tasks from a JSON array or `tasks`/`records` envelope; every row's supplied `id`/`taskId` is authoritative and a duplicate id completely replaces the stored task |
| `DELETE /cop/tasks/{task_id}` | TaskCOP compatibility API (authenticated): delete one task and its task-scoped situation-report document |
| `GET /taskAnalyseSearch/cop-task-three-list` | TaskCOP compatibility API (authenticated): paginated task list for the situation-overview sentiment page (`content`, `taskStatus`, `taskDirection`, `startTime`, `endTime`) |
| `GET/POST/PUT /api/task-reports/by-task/{task_id}` | Shared task-scoped situation-report detail. Query, first save, and full update all use the legacy-compatible `{state, msg, data:[{id, taskId, createTime, updateTime, sessionId, contentJson, categoryType}]}` shape; `contentJson` is preserved as a JSON string. Saving a non-empty report automatically moves the matching TaskCOP task to status `25` (the legacy list's “已完成” state). |
| `POST /api/task-reports/import/by-task/{task_id}` | TaskCOP detail import (authenticated): accept a `sq-report-mock.json`-compatible `data`/`records` array and atomically replace the selected task's full report. The path task id overrides task ids inside the file. The agentfx built-in agent fills that file with the `task-report-build` convert skill in one call (search JSON → 14 categories; `task-report-{enemy,our,env,judge}` re-run a single dashboard page) then imports via `task-report-import`. |
| `POST /api/sentiment-agent/stream` | Authenticated BFF for the frontend sentiment-analysis virtual agent. Forwards the AG-UI JSON body to the `apiUrl` supplied from `runtime-config.js` (Basic Auth from the same payload). TLS certificate verification is off. Upstream SSE is piped through unchanged; connection/HTTP failures return the complete error in `detail`. |
The separately deployed TaskCOP Vue app uses `POST /api/v1/auth/login/username` directly (its `username`, `userId`, or `yUserId` URL parameter identifies the user), so this compatibility flow does not depend on the retired Consumer login service. It calls the Gateway address configured as `b1ConsumerUrl` directly; it does not use a Vite `/api` reverse proxy. An empty TaskCOP store remains empty; task data is introduced through normal creation or the JSON import API. The Gateway CORS middleware always merges local Vite origins `http://localhost|127.0.0.1:{5173,5174,3000,8080}`, including when `GATEWAY_CORS_ORIGINS` is set for another frontend. This keeps a Vite server that falls back from 5173 to 5174 (or the reverse) working after a backend restart. For a deployment, set `GATEWAY_CORS_ORIGINS` to the exact additional frontend origins that are allowed to call the Gateway.
### Workflow Studio interface resources
The Workflow Studio data-source API also registers reusable HTTP interface resources. Create them through
POST /api/workflows/data-sources with kind set to http; provide a non-secret baseUrl, an allowedMethods
allowlist, and optional request headers. The service encrypts the headers immediately and never returns
them through list, detail, resource-catalog, run-event, or canvas APIs. The canvas resource catalog
exposes only the name, description, HTTP address, and allowed methods. HTTP execution continues to
enforce the workflow target security policy, including the default SSRF restrictions.
The same catalog's agent and skill entries are display projections, not a second ownership store:
`/api/workflows/resources/agents` returns `agentId`, cached available `skills`, and `createdBy`; the
skills endpoint returns `skillId`, direct-call capability, and `createdBy`. Creator names are resolved
from the immutable owner id at read time (falling back to the id if an account was removed), while
built-in resources explicitly return `系统内置`.
### Conversation-driven workflow planning
The Workflow Studio chat composer does not silently start a template. A normal user
message first creates a durable planning session with
`POST /api/workflows/{workflow_id}/planning-sessions/stream` (SSE; the legacy
non-streaming `POST /api/workflows/{workflow_id}/planning-sessions` remains for
integrations). The stream relays safe DeerFlow controller milestones—catalog read,
controller parsing, role selection, graph assembly and validation—before returning
the final session, so the UI never has to fabricate a progress bar. The constrained planner builds
two or three validated candidate graphs from the current draft and the visible business
agents the user can use. It is a dedicated built-in **工作流总控** (`workflow-planner`),
not the multi-agent roundtable coordinator: it only reads the user requirement and
catalog metadata, selects permitted agent ids, and ranks the supported strategies. It
does not dispatch a roundtable, use research tools, or execute a formal run. The server
constructs the executable graph from those selected, visible resources; a controller
failure returns an explicit planning error rather than a keyword/template fallback. The
controller call has no callable tools and disables model thinking, so it adds one bounded
planning inference without starting research or worker-agent costs. A
browser selects one candidate, automatically applies the editor's DAG layout when
it loads, then may explicitly save an edited graph through
`PATCH /api/workflows/planning-sessions/{session_id}/proposals/{proposal_id}/graph`,
then creates the one formal run through the existing
`POST /api/workflows/{workflow_id}/runs` endpoint with `planningSessionId` and
`proposalId`. The gateway rejects unselected or invalid candidates and makes a repeated
confirmation return the already-created run, so a double click cannot execute a second
workflow. Planning sessions are persisted separately from runs (migration
`20260831_01`); their parent row is flushed before candidate rows so the foreign-key
transaction works consistently across SQLite, MySQL, and PostgreSQL. A candidate is
not consuming execution resources until it is confirmed.
The development Studio proxy must not gzip `/api` responses: compressed SSE can hold
small planning and run-event frames until the response ends. The backend already marks
both streams `no-transform` and `X-Accel-Buffering: no`; `apps/coze-studio` also sets
its Rsbuild development server `compress: false`. Agent/skill nodes use the bounded
`workflows.agent_recursion_limit` (default `250`), rather than a chat-sized 60-step
limit, because a model-to-tool research turn consumes about two LangGraph super-steps.
The composer may also send a configured `modelName`: the server validates it against
the safe model catalog, uses it for the workflow controller, and writes it into the
generated candidate agent nodes. For a direct single-agent task it overrides only that
projected node; confirmation still executes the immutable selected proposal snapshot.
The controller additionally creates a per-selected-agent task contract (`mission`,
`deliverable`, `scope`, `handoff`). The graph builder accepts contracts only for
catalog-visible, selected agent ids and injects each one into that worker's prompt, so
parallel branches receive complementary research assignments and sequential branches
have an explicit review/handoff boundary rather than all workers attempting the whole
user task. Missing controller fields receive a deterministic role-aware contract.
When an agent emits DeerFlow's native `ask_clarification` ToolMessage, the workflow
enters `awaiting_input` rather than completing or failing. The pause event references
the originating `nodeRunId` and `toolCallId`; the browser renders that original card
inside the agent message and `POST /runs/{run_id}/resume` supplies the user's reply as
the next turn of the same agent thread.
The current planner is deliberately constrained: it can rank supported strategies and
focus resources, but the server builds the graph from the already configured agent
nodes. It cannot invent a node type, external target, or credential from chat text.
For a parallel-research candidate, the configured research agents first flow into the
deterministic `evidence_normalizer` node. It emits a bounded Evidence Pack with the
originating node, role label, explicit gaps, and no invented verification claim before
the synthesis agent sees it. The candidate then runs `deep_research_write`: its
app-layer adapter converts exactly that Evidence Pack into selected, provenance-tagged
Deep Research source rows, starts or reattaches to the existing durable Deep Research
job, forwards report deltas as workflow `node.output.delta` events, and exposes the
final Markdown as a run artifact. Cancellation is forwarded to the same research job;
the node never calls the Deep Research HTTP API or launches a second writer. The
session key includes the topic, evidence snapshot and writing configuration, so a retry
reuses the same job while a changed upstream pack cannot return stale report prose. `multi_agent` Deep Research is
intentionally not nested yet because its separate plan-review pause would conflict
with workflow `human_input`. A completed `agent` or `skill` node in a terminal,
acyclic run can now receive targeted feedback through
`POST /api/workflows/runs/{run_id}/feedback`: the original remains immutable, a
revision run reuses completed nodes outside the target's downstream closure, and
only the target prompt receives the bounded feedback block before it and its
downstream nodes stream again. The public revision run retains the affected and
reused node ids in its feedback summary for the conversation/history UI. Running and looped workflows deliberately continue
to use `human_input` or a new full run; they are not modified in place. See
`docs/WORKFLOW_CONVERSATIONAL_MULTI_AGENT_ORCHESTRATION_ZH.md` for the roadmap.
### Deep Research modes
Deep Research is a separate Gateway feature rather than part of AI Writing. Its
`detailed` runner adapts GPT Researcher's detailed-report control flow (plan
subtopics, collect materials concurrently, write sections concurrently, then
assemble the report). Its `deep` runner adapts the recursive DeepResearchSkill
flow (branch, extract learnings, investigate follow-ups) with a deterministic
query budget derived from `deep_breadth` and `deep_depth`. The runners live in
`packages/harness/deerflow/agents/deep_research/runners/` and only use injected
DeerFlow adapters. `multi_agent` is a LangGraph 1.x editor/researcher/writer/
reviewer graph: it persists a plan-review pause, releases the worker lease, and
resumes from the approved or revised plan through the existing durable-job API.
Report prose uses the configured LangChain model's `astream()` path. Each
native model delta is fanned out as an in-process `report_delta` SSE frame so
the right-side `report.md` sandbox can write smoothly, while bounded,
persisted `report_chunk` events remain reconnect checkpoints. The durable log
is still authoritative: a browser reconnect, slow consumer, or another worker
falls back to `?after=` replay rather than starting a second model call.
Reasoning exposed by providers either through structured reasoning fields or
inline `...` text is separated from report deltas, checkpoints,
and the persisted Markdown projection. It may be forwarded as an ephemeral
live `report_thinking` frame for the message-list thought trace, but is never
persisted or written to the sandbox; only answer Markdown reaches the sandbox.
The short non-streaming control-plane calls used to plan queries, sections, and
reviews are independently capped at 45 seconds (maximum 120 seconds). If query
or section planning times out, the runner emits a visible recoverable warning
and searches the original research topic instead of leaving the job at the
planning step indefinitely. Report-prose streaming is intentionally not capped
by this control-plane limit.
Planning also persists `queries_planned` before each basic, detailed, or
multi-agent parallel search fan-out; the completed session stores its terminal
`lastJobId` in the usage snapshot so
the conversation workbench can replay the full planning/search/source/curation
trace after a page refresh. These are observable execution facts only, never
hidden model reasoning.
Detailed reports use stable introduction/section/conclusion targets; a
multi-agent review revision emits `report_reset` before the replacement draft.
Every session may also provide a bounded `custom_outline`. Markdown headings or
numbered entries become the deterministic section plan for `detailed` and
`multi_agent`; `basic` and `deep` receive the same outline as a guarded
report-writing constraint. The outline is frozen in the session snapshot and
never treated as a system instruction.
All owner-scoped Deep Research routes resolve the platform's asynchronous
`get_current_user` identity before querying or writing persistence, so session
and job records are always bound to a concrete user ID.
Report artifacts are server-named and written through the canonical sandbox
virtual path (`/mnt/user-data/outputs/`), which the path layer maps to
the host filesystem. This keeps native Windows deployments compatible without
relaxing the sandbox traversal guard.
Deep Research reuses the DeerFlow Q&A retrieval stack instead of a separate
scraper. It first runs enabled research skills (`deep-search`, `web-research`,
and enabled knowledge-base skills) through the same `SubagentExecutor` runtime
used by Q&A, then uses a configured `web_search` provider when present. If this
offline deployment still has the shipped placeholder/disabled `web_search`
configuration, it falls back to DeerFlow's existing keyless DuckDuckGo provider
and, when that provider returns only irrelevant engine noise, a second real-web
search source. It persists only real returned sources. A missing intranet
endpoint therefore cannot by itself produce an empty Deep Research report. The capabilities API
reports `skill` only when its required Q&A tool is available, and reports
`web_search` when configured search or the online fallback can run.
Cancelling a queued research job finalizes its job row and immediately marks
the owning session `cancelled` with no active job, so it cannot retain a
per-user concurrency slot.
Saved `draft` sessions do not consume a concurrency slot; the two-job limit
only applies to sessions with a `running` or `awaiting_input` durable job.
Whole-report rewrite builds its requirement plan locally from the selected
style and instruction, publishes that plan immediately, and then opens only
one model stream for the actual Markdown writer. This avoids making users wait
for a separate model-thinking/planning pass before sandbox output begins.
`POST /api/deep-research/sessions/{id}/report-variants` is the structural
regeneration boundary: it copies the completed report and selected evidence
into a new session (including both saved `[来源:id]` and legacy streamed
`[[source:id]]` citation markers remapped to the cloned ids). The following
durable writer is explicitly queued with `operation=generate`: it does not send
the inherited report to the model as text to edit, and instead writes a fresh
report from the complete selected evidence set plus the newly confirmed
outline. Model reasoning remains a live SSE event before the first Markdown
token. New output uses display-ready `[来源:id]` citations; validation repairs
only unique near-complete `drs_src_*` ids (for example a provider-truncated
suffix) and still rejects genuinely unknown evidence ids.
The report-only model stream requests an 8192-token completion budget so a
normal multi-section report and its reference list do not stop at the common
4096-token provider default; malformed trailing citation syntax is rejected
rather than committed as a partial document.
The original report therefore remains independently available in history even
if the new report is cancelled or fails.
#### Optional report illustrations
DeerFlow's stock configuration includes image recognition but no image-generation
provider. Deep Research therefore keeps illustrations disabled until an operator
fills `config.yaml -> deep_research.image_generation` with an
image endpoint, API key, and model. `provider: openai_compatible` uses the
standard `/images/generations` protocol and requires `b64_json`; `provider:
dashscope_native` supports Qwen Image 3.0 and Wan 2.7 on DashScope / Token Plan
with the native multimodal-generation protocol. Native provider image URLs are
downloaded immediately, so both providers produce protected local artifacts.
When configured, the capabilities endpoint enables the page's “生成报告配图” switch.
A selected job writes at most four `research-image-*.png|jpg|webp` artifacts,
injects `deep-research://` links into the Markdown report, and emits image SSE
events. A failed/unconfigured image provider only emits a recoverable warning;
the text report still completes.
Completed reports also expose `GET /api/deep-research/sessions/{id}/messages`
and `POST /api/deep-research/sessions/{id}/chat`. Follow-up answers are grounded
only in that session's report, selected sources, and recent follow-up history;
the assistant returns persisted source ids for every accepted citation. Passing
`allow_new_research=true` is rejected rather than silently triggering another
search job. Owners can change the follow-up evidence set with
`PATCH /api/deep-research/sessions/{id}/sources/{source_id}` after a research
job stops. That selection is applied immediately to later follow-up prompts;
it never rewrites the completed report, and edits are rejected while the job is
running so they cannot race the automatic curation projection.
The Deep Research page is a conversation workspace: the original research
question, observable retrieval steps, completion summary, and report-grounded
follow-ups appear in chronological order. The report itself is written in the
adjacent sandbox as `/mnt/user-data/outputs/report.md`, with Word, Markdown,
and HTML export actions. Word export calls the shared Gateway
`POST /api/writing/export/docx` generator, embeds the licensed title font, and
includes the collected reference list without adding a backend Python package.
In chat-collection sessions the report configuration card
shows the unique material count inferred from the collector's visible tool
results before writing begins. The write lifecycle itself is rendered as
ordinary workspace message-list tool steps (not a second timeline component),
and remains in the transcript after completion; the generated digest is a
normal assistant message below those steps. As soon as the runner enters its
summary phase, that same assistant message is inserted in a streaming state
and receives each SSE `summary_delta`; completion only changes it in place to
the final report-complete wording rather than appending a delayed second
summary card.
Each completed chat-collection report also contributes a normal file card to
the message history, named from the report Markdown title rather than the
internal `report.md` artifact path. Historical sessions keep the sandbox closed
on replay, and opening that card selects the session's virtual report artifact
in the sandbox; its download action uses the existing Markdown/Word export
dialog. The per-user active-job limit and the default in-job retrieval/write
concurrency are both three. The Deep Research composer uses the installed
TDesign `Select`, populated from `/api/models`, for an optional per-research
model override. It is disabled during active collection/writing, is sent as
`model_name` to collector turns, and is applied to the `fast_model`,
`smart_model`, and `strategic_model` roles when a session or report job starts.
Completed-report Q&A uses `POST /api/deep-research/sessions/{id}/chat/stream`
for native SSE deltas; the earlier `/chat` JSON endpoint remains for API
compatibility. An explicit scoped edit such as “修改第一段”、
“重新生成一下第一段” or “重写第二节” uses the separate
`POST /api/deep-research/sessions/{id}/rewrite/stream` report-section writing
pipeline. That route requires a deterministic Markdown range and returns 422
instead of falling back to Q&A when no target is supplied. Ordinary follow-up
questions remain report/source-grounded and never mutate the report. The streamed
replacement candidate is shown in the message list while the sandbox keeps the
current report unchanged. The completed candidate creates a confirmation card;
only `POST /sessions/{id}/messages/{messageId}/rewrite-proposal` with
`{"action":"apply"}` substitutes the matching range and saves the assembled
Markdown. The proposal carries the source report hash, so a stale candidate is
rejected instead of overwriting a later edit; dismissing/cancelling leaves the
persisted report unchanged. This does not initiate new web research. Legacy
unmarked candidates from before this route split are inferred only when their
immediately preceding user message contains the same explicit scoped edit.
Position-roundtable history supports two modes. A non-empty `external_task_id` from the frontend route's `taskId` creates a row with `task_scoped = true`, so all sessions, nodes and persisted conversation snapshots for that task are shared across users; `user_id` is then creator audit metadata only. When the field is absent or empty, the original per-user personal-session behavior remains unchanged. Alembic revision `20260808_01` adds `task_scoped` with the safe default `false` for historical rows plus the task/update-time index.
### Roundtable concurrency guarantees
Position-roundtable node completion, rejection, thread binding and activation
are database transactions with row locking and retry-stable command ids. This
keeps simultaneous seat completion from leaving a downstream stage locked.
Task and intent snapshots are frozen when the chain becomes active, so a late
request from another user cannot alter the confirmed input. Background
roundtable jobs use a database unique active-job key and renewable worker
leases; cancelled jobs are finalized by the lease holder or recovered after an
expired lease, allowing users to start a new job after a worker failure.
The ordinary unit tests use SQLite. To verify deployed database semantics with
real PostgreSQL row locks and independent connections, set
`DEERFLOW_POSTGRES_CONCURRENCY_TEST_URL` and run:
```bash
PYTHONPATH=. uv run --no-sync pytest tests/test_roundtable_postgres_concurrency.py -q
```
AI writing supports a sample-imitation mode: the frontend can upload or paste a sample article, the Gateway extracts text for document files, and the `ai_writing` graph runs a `sample_analyzer` node before normal intent/research/outline/draft stages. The resulting style profile is injected into outline and draft prompts with explicit guardrails against copying source sentences, facts, names, or data.
### Custom Agents
Custom agents are indexed in the `agents` database table. The `id` column is the stable runtime identifier and filesystem directory name (`.deer-flow/agents/{id}/SOUL.md`); `name` is display-only, may be Chinese, and may be duplicated. Rows with `user_id = NULL` are built-in agents, user rows are private unless `published = true`, and users can only update or delete agents they created.
The built-in `forced-research-responder` (「强制检索输出助手」) is seeded at Gateway startup. Unlike an ordinary skill-enabled custom agent, it has a runtime-enforced per-turn collection gate: the model cannot finish a response until a retrieval/knowledge/MCP tool or a skill Python script returns usable information. It then answers directly in the user's requested chat format rather than generating an artifact. The agent may create or edit only temporary `.py` files below `/mnt/user-data/workspace/` and may execute `.py` skill scripts through `bash`; writes to `outputs`, Markdown/document creation, `present_files`, general shell commands, and skill mutation are denied. Administrators can edit its model and skill allowlist from agent management; the runtime safety policy remains fixed. Regression coverage: `tests/test_forced_research_agent.py`.
### AI Writing Intranet Retrieval
AI 写作的「知识库检索」会直连 `ai_writing.intranet_search_url` 内网 ES/知识库接口,并默认跳过代理环境变量和 HTTPS 证书校验,适配自签名或内部 CA 证书的内网服务。需要强校验时可在 `config.yaml` 里设置 `ai_writing.intranet_verify_ssl: true`。通用检索里的非内置技能在 `ai_writing.researcher_skill_agent: true` 时会用真实 agent 运行时执行技能:以 `ai-writing-researcher` 身份读取该技能的 `SKILL.md`,并按技能要求通过 `bash` 运行脚本,弱模型也会在本轮任务里收到明确的读文件和脚本执行指令。
### Installable Sentiment Analysis Skill
skill-packages/sentiment-analysis.skill is a separately installable skill package. It keeps the external AG-UI gateway address in config/sentiment-agent.json, resolves Basic Auth from deployment environment variables, and normalizes the upstream SSE stream into NDJSON (thinking_delta / answer_delta / done) for callers that need incremental rendering. It does not alter the existing frontend sentiment-analysis page. See skill-packages/sentiment-analysis/docs/调用与安装说明.md for installation and invocation.
### IM Channels
The IM bridge supports Feishu, Slack, and Telegram. Slack and Telegram still use the final `runs.wait()` response path, while Feishu now streams through `runs.stream(["messages-tuple", "values"])` and updates a single in-thread card in place.
For Feishu card updates, DeerFlow stores the running card's `message_id` per inbound message and patches that same card until the run finishes, preserving the existing `OK` / `DONE` reaction flow.
---
## Quick Start
### Prerequisites
- Python 3.12+
- [uv](https://docs.astral.sh/uv/) package manager
- API keys for your chosen LLM provider
### Installation
```bash
cd deer-flow
# Copy configuration files
cp config.example.yaml config.yaml
# Install backend dependencies
cd backend
make install
```
### Configuration
Edit `config.yaml` in the project root:
```yaml
models:
- name: gpt-4o
display_name: GPT-4o
use: langchain_openai:ChatOpenAI
model: gpt-4o
api_key: $OPENAI_API_KEY
supports_thinking: false
supports_vision: true
- name: gpt-5-responses
display_name: GPT-5 (Responses API)
use: langchain_openai:ChatOpenAI
model: gpt-5
api_key: $OPENAI_API_KEY
use_responses_api: true
output_version: responses/v1
supports_vision: true
```
Set your API keys:
```bash
export OPENAI_API_KEY="your-api-key-here"
```
For a LiteLLM OpenAI-compatible proxy, keep `supports_thinking: false` unless
the configured model has an explicit, proxy-supported thinking control. This
prevents DeerFlow's connection check from sending vendor-specific `thinking`
fields that a generic OpenAI-compatible gateway may reject.
### Running
**Full Application** (from project root):
```bash
make dev # Starts LangGraph + Gateway + Frontend + Nginx
```
Access at: http://localhost:2026
**Backend Only** (from backend directory):
```bash
# Terminal 1: LangGraph server
make dev
# Terminal 2: Gateway API
make gateway
```
Direct access: LangGraph at http://localhost:2024, Gateway at http://localhost:8001
---
## Project Structure
```
backend/
├── src/
│ ├── agents/ # Agent system
│ │ ├── lead_agent/ # Main agent (factory, prompts)
│ │ ├── middlewares/ # 9 middleware components
│ │ ├── memory/ # Memory extraction & storage
│ │ └── thread_state.py # ThreadState schema
│ ├── gateway/ # FastAPI Gateway API
│ │ ├── app.py # Application setup
│ │ └── routers/ # 6 route modules
│ ├── sandbox/ # Sandbox execution
│ │ ├── local/ # Local filesystem provider
│ │ ├── sandbox.py # Abstract interface
│ │ ├── tools.py # bash, ls, read/write/str_replace
│ │ └── middleware.py # Sandbox lifecycle
│ ├── subagents/ # Subagent delegation
│ │ ├── builtins/ # general-purpose, bash agents
│ │ ├── executor.py # Background execution engine
│ │ └── registry.py # Agent registry
│ ├── tools/builtins/ # Built-in tools
│ ├── mcp/ # MCP protocol integration
│ ├── models/ # Model factory
│ ├── skills/ # Skill discovery & loading
│ ├── config/ # Configuration system
│ ├── community/ # Community tools & providers
│ ├── reflection/ # Dynamic module loading
│ └── utils/ # Utilities
├── docs/ # Documentation
├── tests/ # Test suite
├── langgraph.json # LangGraph server configuration
├── pyproject.toml # Python dependencies
├── Makefile # Development commands
└── Dockerfile # Container build
```
---
## Configuration
### Main Configuration (`config.yaml`)
Place in project root. Config values starting with `$` resolve as environment variables.
Key sections:
- `models` - LLM configurations with class paths, API keys, thinking/vision flags
- `tools` - Tool definitions with module paths and groups
- `tool_groups` - Logical tool groupings
- `sandbox` - Execution environment provider
- `skills` - Skills directory paths plus optional prompt routing; `skills.es_query_routing`
can inject an editable system-prompt rule that sends flexible Elasticsearch Query DSL
requests to the `es_query` skill, and can be disabled without changing code. Each ordinary
Q&A agent build logs `ES query routing prompt registration` with `registered`, eligibility,
character count, and a short SHA-256 fingerprint so operators can verify the exact configured
block reached the final model system prompt without logging its contents
- `title` - Auto-title generation settings
- `summarization` - Context summarization settings
- `subagents` - Subagent system (enabled/disabled)
- `memory` - Memory system settings (enabled, storage, debounce, facts limits)
Provider note:
- `models[*].use` references provider classes by module path (for example `langchain_openai:ChatOpenAI`).
- If a provider module is missing, DeerFlow now returns an actionable error with install guidance (for example `uv add langchain-google-genai`).
### Browser CORS
The Gateway permits the local Vite development origins on ports 5173, 5174,
3000, and 8080 for both `localhost` and `127.0.0.1` with credentials, even
when `GATEWAY_CORS_ORIGINS` is set. For a separate deployed frontend, set
`GATEWAY_CORS_ORIGINS` to its exact origin (including scheme and port), for
example `http://47.88.25.99:7010`.
### Extensions Configuration (`extensions_config.json`)
MCP servers and skill states in a single file:
```json
{
"mcpServers": {
"github": {
"enabled": true,
"type": "stdio",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {"GITHUB_TOKEN": "$GITHUB_TOKEN"}
},
"secure-http": {
"enabled": true,
"type": "http",
"url": "https://api.example.com/mcp",
"oauth": {
"enabled": true,
"token_url": "https://auth.example.com/oauth/token",
"grant_type": "client_credentials",
"client_id": "$MCP_OAUTH_CLIENT_ID",
"client_secret": "$MCP_OAUTH_CLIENT_SECRET"
}
}
},
"skills": {
"pdf-processing": {"enabled": true}
}
}
```
### Environment Variables
- `DEER_FLOW_CONFIG_PATH` - Override config.yaml location
- `DEER_FLOW_EXTENSIONS_CONFIG_PATH` - Override extensions_config.json location
- Model API keys: `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `DEEPSEEK_API_KEY`, etc.
- Tool API keys: `TAVILY_API_KEY`, `GITHUB_TOKEN`, etc.
### LangSmith Tracing
DeerFlow has built-in [LangSmith](https://smith.langchain.com) integration for observability. When enabled, all LLM calls, agent runs, tool executions, and middleware processing are traced and visible in the LangSmith dashboard.
**Setup:**
1. Sign up at [smith.langchain.com](https://smith.langchain.com) and create a project.
2. Add the following to your `.env` file in the project root:
```bash
LANGSMITH_TRACING=true
LANGSMITH_ENDPOINT=https://api.smith.langchain.com
LANGSMITH_API_KEY=lsv2_pt_xxxxxxxxxxxxxxxx
LANGSMITH_PROJECT=xxx
```
**Legacy variables:** The `LANGCHAIN_TRACING_V2`, `LANGCHAIN_API_KEY`, `LANGCHAIN_PROJECT`, and `LANGCHAIN_ENDPOINT` variables are also supported for backward compatibility. `LANGSMITH_*` variables take precedence when both are set.
### Langfuse Tracing
DeerFlow also supports [Langfuse](https://langfuse.com) observability for LangChain-compatible runs.
Add the following to your `.env` file:
```bash
LANGFUSE_TRACING=true
LANGFUSE_PUBLIC_KEY=pk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_SECRET_KEY=sk-lf-xxxxxxxxxxxxxxxx
LANGFUSE_BASE_URL=https://cloud.langfuse.com
```
If you are using a self-hosted Langfuse deployment, set `LANGFUSE_BASE_URL` to your Langfuse host.
### Dual Provider Behavior
If both LangSmith and Langfuse are enabled, DeerFlow initializes and attaches both callbacks so the same run data is reported to both systems.
If a provider is explicitly enabled but required credentials are missing, or the provider callback cannot be initialized, DeerFlow raises an error when tracing is initialized during model creation instead of silently disabling tracing.
**Docker:** In `docker-compose.yaml`, tracing is disabled by default (`LANGSMITH_TRACING=false`). Set `LANGSMITH_TRACING=true` and/or `LANGFUSE_TRACING=true` in your `.env`, together with the required credentials, to enable tracing in containerized deployments.
---
## Development
### Commands
```bash
make install # Install dependencies
make dev # Run LangGraph server (port 2024)
make gateway # Run Gateway API (port 8001)
make lint # Run linter (ruff)
make format # Format code (ruff)
```
### Code Style
- **Linter/Formatter**: `ruff`
- **Line length**: 240 characters
- **Python**: 3.12+ with type hints
- **Quotes**: Double quotes
- **Indentation**: 4 spaces
### Testing
```bash
uv run pytest
```
### Runtime appearance themes
The admin appearance settings are persisted in `.deer-flow/system_settings.json`.
Besides the built-in themes, an administrator can create up to 12 runtime
themes through `PUT /api/system-settings/appearance`. Each theme stores a
safe six-digit `primary` colour, an optional primary-button `text_color`,
`light`/`dark` mode, name, enabled flag and sort order. The frontend derives
the rest of the visual tokens; the optional `overrides` object only exposes
ten interaction colours (`button_hover`, `navigation_gradient_start`,
`navigation_gradient_end`, `selected_background`, `border`, `muted`,
`muted_text`, `destructive`, `brand_title`, `primary_foreground`) for
exceptional brand requirements. The two navigation values override the
automatically derived top-bar gradient. A custom theme id
must use the `custom-theme-` prefix. Disabled or removed custom themes cannot
remain the system default and automatically fall back to `light`.
Compact top-bar shortcuts use the same appearance endpoint. In addition to
login and theme query parameters, each shortcut may include ordered keyless
`path` parameters. The frontend resolves each selected login field and appends
it as an encoded path segment (including inside a `#/…` hash route); password
remains an ordinary login query parameter with the key `password`.
### Citation display settings
Reference-mode entry points are controlled by `citation_display` in
`.deer-flow/system_settings.json`. `reference_mode_enabled` defaults to `false`;
when disabled, the frontend hides the "参考文献" composer button in both the
normal workspace and iframe chat, and hides the per-user RAG citation toggle
from appearance settings. Conversations with selected LLMWiki knowledge bases
still show their source documents in the reference panel even when the manual
entry is hidden. `GET /api/system-settings/citation-display` is
readable by logged-in users so the UI can decide whether to expose the entry;
`PUT /api/system-settings/citation-display` is admin-only and supports partial
updates for both `reference_mode_enabled` and `cleanup_words`.
### Ordinary Q&A Markdown format switch
Administrators can opt into the formal long-answer Markdown prompt from the
admin "问答提示词配置" panel. The value is persisted as
`prompt_prefix.ordinary_qa_markdown_format_enabled` in
`.deer-flow/system_settings.json` and defaults to `false`. When enabled, only
ordinary Q&A receives the `# 主标题` / `## 一、标题` / `### (一) 标题` /
`#### 1. 标题` rules; writing mode, notebook mode and scheduled runs remain
unaffected. The change takes effect when the next prompt is built.
---
## Technology Stack
- **LangGraph** (1.0.6+) - Agent framework and multi-agent orchestration
- **LangChain** (1.2.3+) - LLM abstractions and tool system
- **FastAPI** (0.115.0+) - Gateway REST API
- **langchain-mcp-adapters** - Model Context Protocol support
- **agent-sandbox** - Sandboxed code execution
- **markitdown** - Multi-format document conversion
- **tavily-python** / **firecrawl-py** - Web search and scraping
---
## Documentation
- [Configuration Guide](docs/CONFIGURATION.md)
- [Architecture Details](docs/ARCHITECTURE.md)
- [API Reference](docs/API.md)
- [Offline Docker wheelhouse deployment (ZH)](../README.md)
- [File Upload](docs/FILE_UPLOAD.md)
- [Path Examples](docs/PATH_EXAMPLES.md)
- [Context Summarization](docs/summarization.md)
- [Plan Mode](docs/plan_mode_usage.md)
- [Backend Run/Stop Guide (ZH)](docs/BACKEND_RUN_STOP_ZH.md)
- [Setup Guide](docs/SETUP.md)
---
## License
See the [LICENSE](../LICENSE) file in the project root.
## Contributing
See [CONTRIBUTING.md](CONTRIBUTING.md) for contribution guidelines.
## Position roundtable durable history
The position-roundtable workflow stores database-owned recovery snapshots for
Step-1 intent chat, every position-node chat, the summary/action-plan chats,
unsent composer drafts, and generated text deliverables. LangGraph checkpoints
remain the execution source, but SQL history continues to render after a
checkpoint or sandbox-output volume is replaced. Apply Alembic revision
`20260723_05` (or simply upgrade to `head`) before deploying this version.
Its live intent conversation still uses the shared `/api/intent/init` and
`/api/intent/stream` endpoints used by the original multi-agent roundtable.
Position roundtable identifies itself with `intent_agent_id: "position-roundtable-intent"`;
this is a separately seeded built-in agent, not an alias of `roundtable-intent`.
It preserves the same thread, SSE streaming, and file-reading lifecycle as the
generic Step-1 agent, but has its own **at-most-five user-assistance-round**
brake. A round is one AI turn plus the user reply, not one tool call: the
dedicated agent may emit several independent `ask_clarification` calls in the
same turn, which the frontend renders as several separate selection cards and
submits together. A topic-only request is not considered resolved: the agent
asks for every independent missing decision (such as scope, decision use, or
task constraints) as separate cards. For the dedicated agent, a fresh
`ask_clarification` emitted in a stream is a strict user-input pause and wins over any same-run premature
`[INTENT_READY]`; a historical clarification from the prior turn does not
block the user answer from completing. Its final
`[INTENT_READY]` JSON contains four string arrays: `coreGoals`, `riskWarnings`,
`keyPoints`, and `strategicSignificance`; it must contain completed conclusions,
not “待确认” or other unresolved placeholders. The multi-agent endpoint remains
backwards compatible when this optional field is omitted and continues to emit
`objective`, `constraints`, and `assumptions`.
`/api/intent/init` creates its empty thread directly through the in-process
thread store/checkpointer, matching multi-agent initialization and avoiding a
deployment-sensitive `127.0.0.1:` loopback. If initial checkpoint
creation fails or is cancelled, bootstrap removes both the partial checkpoint and
its thread metadata row. Thread metadata creation issues no post-commit ORM
read-back on any code path: all returned fields are assigned before commit, and
skipping the extra read keeps a MySQL read/write-splitting endpoint from sending
it to a lagging replica (which fails the request with
`Could not refresh instance`). Only the
recovery snapshot layer is position-specific. Deployments must expose
either `/api/position-roundtable/*` or the compatibility mirror
`/api/multi-agent/position-roundtable/*`; updated clients probe both and only
retry when the first response is the framework/proxy-level `404 Not Found`.
## Position-roundtable role directory
The position-roundtable work views are persisted separately from organisation
positions/RBAC in `position_roles`. Gateway startup idempotently inserts the
four built-in ids (testing analysis, organisation test, solution design and
comprehensive test), but only when an id is missing: administrator changes to
the name, order, enabled state or role type are never overwritten.
`GET /api/position-roles` is consumed by the workbench and `PUT
/api/position-roles` is administrator-only. Exactly one enabled `intent` role
is required. It is the initial intent-identification workspace and owns the
summary/action-plan closing agents. Built-in rows retain stable ids so
historical chain nodes remain readable; they can be edited or disabled but not
deleted. Custom roles may be freely added or removed.
## Roundtable artifact submission status
The task-scoped roundtable workbench exposes a lightweight submission-status
API for generated deliverables:
- `GET /api/roundtable/tasks/{taskId}/artifacts?submitted=true` returns the
submitted artifact references for a task (or `items: []`).
- `PUT /api/roundtable/tasks/{taskId}/artifacts/submission` records
`submitted: true` or cancels it with `submitted: false`.
Both routes use the normal Gateway `Authorization: Bearer ` middleware.
The record is shared by `taskId`; the authenticated user is retained only in
audit columns. Its server-owned idempotency key is
`taskId + sourceKey + path + revision`. Repeating the same submitted payload
returns the existing record; cancellation updates its state and timestamp
without deleting the audit row. DeerFlow stores only the artifact reference
(`threadId` + sandbox `path`, plus the optional session/node/name/MIME fields),
never the Markdown body and never a task-system delivery. Continue to preview
the file through the existing thread-artifact endpoint.
## 普通知识库管理
在「普通知识库」的「我的知识库」或「公共 LLMWiki」中,知识库所有者和管理员可以点击卡片上的「编辑」,修改名称和描述(描述可清空)。保存后同步更新知识服务和本地记录,并刷新列表;系统「对话沉淀」库不提供编辑入口。
远端已经删除的知识库会在下次列表刷新时自动隐藏,包括个人、公共列表和知识库选择器;页面每 15 秒自动刷新,也可以手动点击「刷新」。远端服务请求失败时展示错误,不把连接故障当成删除,不自动删除本地映射或已有引用。
## DeerFlow-local LLMWiki Wiki vector index
LLMWiki can mirror processed WeKnora Wiki pages into DeerFlow and search only
processed Wiki content. A knowledge base uses exact NumPy cosine similarity
only after its whole Wiki has completed one coherent vectorization pass; until
then internal search calls WeKnora's Wiki-page search/list API and reads the
matched Wiki pages. It never falls back to raw documents or chunks. Configure
`llmwiki.local_wiki_index.embedding`, then enable the local index in
`config.yaml`. `auto_sync` defaults to true: the scheduler performs one
immediate incremental scan at startup and then scans every configured interval
(30 seconds by default). Wiki changes are content-hash checked, so unchanged
pages reuse their vectors instead of being encoded again.
`database_batch_size` caps each vector insert statement (default `50`) while
the page replacement remains atomic, preventing a large Wiki from flooding
the SQL connection in one statement.
The optional `embedding.dimensions` value is used only to validate the returned
vector size and build the index fingerprint; it is deliberately not sent in
the embedding request because fixed-dimension BGE/OpenAI-compatible endpoints
may reject that extra field with HTTP 400. Omit it to discover the dimension
from the first successful response.
Embedding failures log the upstream HTTP status and a bounded response body.
The client never sends a `dimensions` request parameter.
The migration creates `llmwiki_wiki_pages`, `llmwiki_wiki_vectors`, and
`llmwiki_wiki_sync_states`. Wiki Markdown and normalized float32 vectors are
stored in DeerFlow's configured SQL database. Multi-worker sync is protected by
database leases; searches use revision-checked per-knowledge-base LRU snapshots
after the full-library readiness gate succeeds. Knowledge-base owners can click
**手动向量化** in Wiki management (`POST
/api/llmwiki/knowledge-bases/{id}/vectorize`) and follow its status at
`GET /api/llmwiki/knowledge-bases/{id}/wiki-index-status`. The response includes
aggregate `progress` plus a `pages` list with every Wiki page's index state,
vector-section count, timestamps, and last error. The management page renders
the percentage and labels each article as 已向量化 / 待向量化 / 失败, so it is
clear exactly what is searchable. Existing chunk citations remain readable
while all new retrieval citations open Wiki pages.
The frontend labels these results simply as **Wiki** references; provider-brand
wording is intentionally hidden from user-facing labels and status text. Clicking any
part of a Wiki reference card opens one right-side drawer (rather than a second
detail dialog), renders the complete page as Markdown, and reads the remote
WeKnora Wiki page directly when the local index is disabled. Even if a Wiki
citation carries a legacy document id, the drawer does not mix in the uploaded
original-file preview. A single page can be downloaded as Markdown or Word from
the drawer. **Open in Wiki** navigates to the exact DeerFlow knowledge-base
mapping and Wiki slug. Wiki search-result cards use the same drawer interaction.
Exact Wiki links now open a lightweight native reader at
`knowledge/{mapping_id}/wiki/view?wikiSlug=...`; it fetches only the selected
page and never creates a WeKnora iframe session. The former
`knowledge/{mapping_id}?tab=wiki&wikiSlug=...` links redirect to the reader
before waiting for the LLMWiki runtime check. Drawer navigation stays in the
same SPA and passes the already-loaded page as route state for immediate paint.
The native reader also keeps a searchable, collapsible Wiki directory on the
left. It reconstructs hierarchy from `parent_slug`, `wiki_path`,
`category_path`, and the slug fallback, highlights and expands the current
page, and switches directory entries through the same lightweight reader route
without mounting the WeKnora iframe.
All generated SPA deep links preserve the deployment document path (for
example `/web/`) before the `#` route, including Q&A references and Wiki index
links. Protected Wiki images no longer render from temporary browser `blob:`
URLs. Public images use the published file proxy directly; private images first
obtain a 24-hour signed capability and then load through
`GET /api/public/llmwiki/files/{token}` with a real HTTP URL and inline
filename, allowing zoom, new-tab viewing, and browser Save As without exposing
the user's bearer token. The embedded WeKnora detail page follows the same
rule: late `img[data-protected-src]` placeholders keep their scoped HTTP proxy
URL on the DOM node, so WeKnora's click-to-enlarge viewer never receives an
already-revoked `blob:` URL.
## Leaderboard query and schema startup safety
The leaderboard excludes scheduled runs with indexed correlated `NOT EXISTS`
checks. Tool metrics no longer outer-join the full `runs` table merely to
recognize legacy scheduler messages. New databases receive dimension-aware
analytics indexes from ORM metadata; existing large MySQL deployments should
run `scripts/sql/leaderboard_indexes_mysql.sql` during a low-traffic window so
index creation never delays application startup.
Startup Alembic migrations now have one unambiguous head. The historical
duplicate `20260821_01` identifier is repaired by unique fixed-question
revisions and the `20260825_01` merge revision, which also fills the
report-structure column for databases stamped by either old variant.
Published knowledge bases also have a login-free, chrome-free reader at
`/#/embed/knowledge/{mapping_id}` for direct third-party iframe use. With no
`wikiSlug`, it selects the Wiki **索引** page first (recognized by title, page
type, or `index`/`home`/`readme` slug) and pins that page to the top of the
public directory. If WeKnora has pages but no explicit index page, the reader
creates a virtual `__index__` landing page from the public directory; empty
dynamic-index groups also fall back to those page summaries;
`?wikiSlug=...` deep-links one page. The reader supports
`theme=light|dark`, `bg`, and `text` embed appearance parameters. Index-article
catalogue sections (page-type counts, nested categories, page summaries) are
read from the published-only `/wiki/index` proxy because WeKnora renders that
part dynamically rather than storing it in the index page body; group and item
ordering therefore stays identical to WeKnora. Links generated as
knowledge-base-shell placeholders are matched
against the public page title, slug, or alias and rewritten to exact embedded
`wikiSlug` links. They retain WeKnora's brand-color medium text and dashed
underline (solid on hover), and switch articles inside the current reader
without a full-page reload. Data comes
only from the read-only `/api/public/llmwiki/knowledge-bases/{id}/...` namespace,
which rechecks `publication_status=published` on every request and denies the
system conversation-deposit knowledge base. Private or unpublished bases are
never exposed by this route. Anonymous consumers can discover the available
public Wiki libraries through `GET /api/public/llmwiki/knowledge-bases`. The
endpoint validates each published mapping against WeKnora, omits stale remote
mappings, and returns safe comprehensive metadata: description/type,
document/chunk/processing/share counts, non-secret configuration and
timestamps, Wiki page/type/status counts plus explicit/generated index mode and
default embed route, and aggregate local-vector state/counts/timestamps. It
never returns the WeKnora id, tenant/creator identity, storage credentials, API
credentials, raw vectors, or per-page indexing errors.
Protected `resource://` and supported object-storage images are fetched through
the current mapping instead of being exposed as upstream URLs. Authenticated
readers use `GET /api/llmwiki/knowledge-bases/{id}/files?file_path=...`; public
readers use the corresponding `/api/public/llmwiki/...` route, which requires a
published mapping. The embedded WeKnora guard applies the same mapping-scoped
proxy to document previews and Markdown images, including images inserted after
the drawer/page first renders.
Wiki management also provides **导出 Excel**. `GET
/api/llmwiki/knowledge-bases/{id}/wiki/export.xlsx` exports every processed Wiki
page, including its full directory path, current directory, parent-directory
path, one column per directory level, metadata, and complete Markdown content
(split across continuation columns when Excel's per-cell limit is reached).
Authenticated operations are available below `/api/llmwiki/wiki-index` and
`POST /api/llmwiki/wiki-vector-search`. The service-to-service endpoints below
`/api/external/llmwiki/wiki` require `X-API-Key`, are rate-limited, and only
expose mappings that are both internally published and explicitly marked
`external_search_enabled`; the system `对话沉淀` mapping is always denied.
A separate browser-friendly endpoint, `POST /api/knowledge/vector-search`, is
fully anonymous: it requires neither a DeerFlow token nor `X-API-Key`, has no
rate limit, and responds with `Access-Control-Allow-Origin: *`. It searches the
local vectors of every published knowledge base regardless of
`external_search_enabled`, while still excluding the system `对话沉淀` base.
The minimal request is `{"query":"如何部署"}`; optional fields are `top_k`
(default `10`, maximum `100`) and `include_page_content` (default `false`). Each
matched vector section is returned as its own ranked result, including the
matched text, score, DeerFlow and WeKnora knowledge/Wiki ids, page metadata,
raw source/chunk references, best-effort document/chunk ids, and a directly
usable `wiki_url`. Set `llmwiki.local_wiki_index.frontend_base_url` to the
deployed frontend origin/base path; its development default is
`http://127.0.0.1:5174`.
See `docs/LLMWIKI_DEERFLOW_LOCAL_VECTOR_INDEX_ZH.md` for configuration,
security boundaries, rollout, and API contracts.
## Document rewrite fallback
Whole-document rewrite commits the artifact/report before it attempts to store
the optional reversible before-image. If that version-snapshot write fails, the
rewritten content remains available and Gateway logs the exception; it never
restores the original solely because undo history is unavailable. The durable
result marks `versionSnapshotFailed`, allowing the sandbox to enable its normal
manual-save action without showing the user an error notification.
The optional rewrite-history hydration endpoints scope results to the
authenticated job owner. They deliberately return an empty list rather than
performing a second thread-owner check, so a legacy guest/SSO display alias
cannot make an otherwise usable sandbox report fail with a history-only 404.
Artifact and rewrite path resolution uses the persisted thread creator's
directory when metadata exists, after the route's normal ownership check. This
prevents a missing request context from incorrectly looking under
`users/default` and returning 404 for an existing sandbox file.
The final rewrite commit uses a deliberately short sibling temporary filename,
which keeps long Windows artifact paths below the path-length limit. It also
recreates a missing `outputs` parent before retrying, so transient
sandbox-directory recreation cannot discard a completed rewrite.
## Offline local-sandbox Office runtime
The pure DeerFlow offline image treats the backend container as the default
`LocalSandboxProvider` runtime. It bundles MarkItDown, python-docx, openpyxl,
python-pptx, LibreOffice Headless, Pandoc, Poppler, and CJK fonts. Modern OOXML
uploads are parsed directly; legacy DOC/XLS/PPT and ODT/ODS/ODP files are
normalized through an isolated LibreOffice profile before Markdown extraction.
Run `./OFFLINE_INSTALL.sh --office-check` from the offline bundle to generate
and re-parse Word, Excel, and PowerPoint samples without requiring a database.
## Skill knowledge distillation and assistant knowledge
The Gateway now exposes an administrator-only skill distillation pipeline at
`/api/skill-knowledge`. Each request can send up to 200 skills to one writable
WeKnora Wiki knowledge base; multi-base distillation is done by repeated
operations so the local binding remains `skill_name + target_id`. The scanner
never reads or uploads code/script files, strips Markdown code fences, blocks
credential-like files and content, safely inspects ZIP members, and converts
supported Office/PDF/spreadsheet files off the event loop. Jobs use database
claims and renewable leases, stage snapshots before activation, retain the
prior active version on failure, support per-item retry, and keep per-binding
contribution records for safe detach/delete. Direct assistant-base targets,
non-Wiki/FAQ/conversation/readonly targets, and direct Neo4j writes are rejected
for this product version.
`/api/assistant-knowledge` manages DeerFlow-native assistant knowledge bases.
There is one global full-library base named `知识梳理总库`, plus admin-created
custom assistant bases. Reads and search are available to authenticated users;
global initialization, custom-base creation, WeKnora export/import, source
reimport, and source deletion are admin-only. Assistant bases only accept
WeKnora export packages or imports triggered from a WeKnora mapping; direct
skill import and arbitrary local-file import return 410. Importing into a custom
assistant base also mirrors the same package into the global base, where
canonical entities and relations are aligned across all imported sources.
The database schema is introduced by Alembic revision `20260901_01`.
### Wiki 流转与本地小批量编码
所有登录用户都可在「技能 → 归纳到知识库」中多选自己可见的技能,并归纳到一个
自己可写的 WeKnora Wiki 知识库;用户只能看到和维护自己发起的归纳记录,管理员
接口仍可查看全量记录。技能归纳上传清理后的知识正文,等待 WeKnora 完成 Wiki 分析,再建立本地
Wiki 向量并通过落盘导出包自动沉淀到「知识梳理」。内置技能没有数据库所有权记录
也可以归纳。若解析结束却没有 Wiki,会提示检查生成模型和资料;文档解析完成
不等于 Wiki 生成成功,例如生成服务余额不足仍可能表现为文档已完成。
管理员可以在普通知识库直接生成向量数据包,每次操作有独立后台记录和下载入口。
助手知识库支持上传包、异步导入和导出;新建空的普通 Wiki 库可上传助手导出包。管理员可编辑助手知识库名称与描述,也可删除自建助手知识库及其 Wiki、向量、实体、关系、导入记录和落盘数据包;承担全库沉淀的「知识梳理总库」允许编辑但禁止删除。
包保留正文、别名/标签、目录分类、双向链接、实体关系及 float32 向量。
回流时通过 WeKnora API 重建目录/页面,向量进入 DeerFlow 本地 Wiki 索引,
**不直接写 WeKnora 的底层向量或图数据库**。普通与助手 Wiki 使用同一个阅读组件。
小批量 CPU 编码可配置如下,无需额外部署 HTTP 编码服务:
```yaml
llmwiki:
local_wiki_index:
enabled: true
auto_sync: true
sync_interval_seconds: 30
embedding:
provider: local
model: bge-embedding-m3
dimensions: 1024
batch_size: 2
local_threads: 2
local_process_isolation: true
# 离线部署时可指定已下载的 BAAI/bge-m3 模型目录:
# local_model_path: /opt/models/bge-m3
```
本地编码依赖随标准后端安装。将完整 BGE-M3 模型目录提前放到离线服务器,并通过
`local_model_path` 或 `DEERFLOW_BGE_M3_MODEL_PATH` 指向它。管理员也可在
「设置和更多 → 设置 → 本地编码器」中选择服务器目录,或从浏览器上传完整模型目录;
页面配置持久化到运行目录的 `system_settings.json`,上传文件落到
`.deer-flow/models/bge-m3`,均不进入源码仓库。系统强制 `local_files_only`,
**不会访问互联网或下载模型**。管理员可在同一设置页点击「一键启动编码器」异步加载;该操作不挂在应用启动流程,加载失败不会
中断 DeerFlow 主服务。默认由独立低优先级子进程承载模型内存,正常停止 DeerFlow
时会一并回收该子进程;编码进程异常退出只会将当前向量任务标记失败。模型目录应包含
`onnx/model.onnx`(官方权重约 2.3 GB)及对应 tokenizer/config 文件。权重不进入业务源码仓库。
浏览器目录上传采用磁盘暂存和 4 MB 分块复制,页面显示上传进度;经 Nginx 等反向
代理部署时仍需把请求体上限和读写超时配置到可承载完整模型目录。
相同模型指纹、相同内容的页面复用现有向量;短正文/摘要至少保留一个非空分段。
数据包导入不重新编码,模型指纹或维度不一致时明确拒绝,不能只按维度混用模型。
问答编码查询本身仍需可用的同一模型。WeKnora 的原始文档向量不能当作生成后
Wiki 正文的向量使用,两种文本需要分别建立各自的首次索引。
Wiki 页面新增、编辑、删除及文档上传/重解析都会立即触发一次增量检查;WeKnora
后台稍后才完成的 Wiki 抽取由 30 秒周期补扫接管。每次成功向量化后自动把该普通
知识库的最新完整快照同步到「知识梳理总库」,包括服务重启前已有库和新建库;来源
页面被删除时也会在总库中退役,不保留失效 Wiki。
「知识梳理总库」不是按来源简单堆叠:导入时会规范化全/半角、空白、标点及常见实体
类型,并利用实体别名对齐同一人物、组织、产品或概念;不同来源中同标题的实体 Wiki
只保留一个规范页面,正文按去重后的知识块合并,来源与实体/关系贡献仍分别留痕。
重复上传同一个 WeKnora 导出包时,以包内稳定知识库 ID 识别同一来源并执行快照替换,
不会再产生一批新的片段副本。已有历史重复数据会在总库首次读取、检索、导出或下次
同步时自动归并。助手知识检索排除目录、索引和摘要页;关键词只查标题与知识正文,
向量检索返回实际命中的正文片段,问答上下文和检索卡片均不再用摘要代替知识。
大包采用流式 JSONL 落盘、同盘临时 SQLite 分阶段读取及异步导入,避免一次性将
整个包加载进内存。部署需预留包、暂存索引和业务数据库空间,并配置代理上传大小
与超时。当前不宣称断点续传或 GB 级压力验收已完成。
## openGauss / GaussDB application database
The application ORM can use a PG-compatible openGauss/GaussDB database through
the official `opengauss-sqlalchemy` async dialect. Install the optional driver
set before enabling it:
```bash
uv sync --extra gauss
```
Configure the business database independently from the LangGraph checkpointer:
```yaml
database:
backend: gauss
gauss_url: $GAUSS_DATABASE_URL
checkpointer:
type: sqlite
connection_string: .deer-flow/data/deerflow.db
```
`GAUSS_DATABASE_URL` should normally be
`opengauss+asyncpg://user:password@host:26000/deerflow`. Plain
`opengauss://`, `gaussdb://`, and PostgreSQL-style URLs are accepted and
normalized to the openGauss dialect. Do not configure the LangGraph
PostgreSQL checkpointer against GaussDB: its migrations are PostgreSQL-specific.
For a production image, build with `--build-arg UV_EXTRAS=gauss`.