1226 lines
86 KiB
Markdown
1226 lines
86 KiB
Markdown
# DeerFlow Backend
|
||
|
||
普通知识库默认详情页面:管理员可在知识库顶部选择“本系统 Wiki 索引”或“嵌入 WeKnora”,对应 `ordinary_knowledge_landing`(默认 `wiki`)。Wiki 库首次进入默认展示本系统索引,右上角“打开 WeKnora”切换到内嵌管理页;内嵌页可返回索引。非 Wiki 库仍进入 WeKnora 文档页。设置通过 `/api/system-settings/ordinary-knowledge-landing` 保存,仅管理员可修改;指定文章、文档、图谱的直达链接优先于默认设置。
|
||
|
||
助手知识库可独立于 WeKnora 使用:管理员进入助手知识库,选择或创建库后点击“上传文件生成 Wiki”。支持 Markdown/TXT/CSV、PDF、Word、Excel、PowerPoint(单文件 100 MB,解析正文上限 400 万字符)。文件先写入运行目录 `assistant-knowledge/files/{job_id}`,后台分段调用已配置的语言模型生成分类 Wiki、实体关系,再通过 DeerFlow 编码器对正文分段向量化,保存到助手库并汇入知识梳理总库。无需配置 WeKnora URL;配置后仍可直接上传,也保留 WeKnora 数据包导入。文件处理沿用管理员写入权限,所有用户可以浏览和搜索。
|
||
|
||
导入记录显示解析、Wiki 生成和向量化阶段;失败或服务中断后可“重试 / 恢复处理”,复用已落盘的 Wiki 和相同模型指纹的向量。离线编码器须预先放置模型并在管理员设置中配置;模型或编码器不可用会明确失败,不把摘要冒充 Wiki 或虚构向量。扫描件是否可解析取决于部署的文档解析/OCR 能力。
|
||
|
||
DeerFlow is a LangGraph-based AI super agent with sandbox execution, persistent memory, and extensible tool integration. The backend enables AI agents to execute code, browse the web, manage files, delegate tasks to subagents, and retain context across conversations - all in isolated, per-thread environments.
|
||
|
||
---
|
||
|
||
## Architecture
|
||
|
||
```
|
||
┌──────────────────────────────────────┐
|
||
│ Nginx (Port 2026) │
|
||
│ Unified reverse proxy │
|
||
└───────┬──────────────────┬───────────┘
|
||
│ │
|
||
/api/langgraph/* │ │ /api/* (other)
|
||
▼ ▼
|
||
┌────────────────────┐ ┌────────────────────────┐
|
||
│ LangGraph Server │ │ Gateway API (8001) │
|
||
│ (Port 2024) │ │ FastAPI REST │
|
||
│ │ │ │
|
||
│ ┌────────────────┐ │ │ Models, MCP, Skills, │
|
||
│ │ Lead Agent │ │ │ Memory, Uploads, │
|
||
│ │ ┌──────────┐ │ │ │ Artifacts │
|
||
│ │ │Middleware│ │ │ └────────────────────────┘
|
||
│ │ │ Chain │ │ │
|
||
│ │ └──────────┘ │ │
|
||
│ │ ┌──────────┐ │ │
|
||
│ │ │ Tools │ │ │
|
||
│ │ └──────────┘ │ │
|
||
│ │ ┌──────────┐ │ │
|
||
│ │ │Subagents │ │ │
|
||
│ │ └──────────┘ │ │
|
||
│ └────────────────┘ │
|
||
└────────────────────┘
|
||
```
|
||
|
||
**Request Routing** (via Nginx):
|
||
- `/api/langgraph/*` → LangGraph Server - agent interactions, threads, streaming
|
||
- `/api/*` (other) → Gateway API - models, MCP, skills, memory, artifacts, uploads, thread-local cleanup
|
||
- `/` (non-API) → Frontend - Next.js web interface
|
||
|
||
Health probes are always unauthenticated, even when API authentication is enabled: `GET|HEAD /health`, `/api/health`, and `/api/langgraph/health` all return the Gateway health status without requiring a cookie or bearer token.
|
||
|
||
---
|
||
|
||
## Core Components
|
||
|
||
### Lead Agent
|
||
|
||
The single LangGraph agent (`lead_agent`) is the runtime entry point, created via `make_lead_agent(config)`. It combines:
|
||
|
||
- **Dynamic model selection** with thinking and vision support
|
||
- **Middleware chain** for cross-cutting concerns (9 middlewares)
|
||
- **Tool system** with sandbox, MCP, community, and built-in tools
|
||
- **Subagent delegation** for parallel task execution
|
||
- **System prompt** with skills injection, memory context, and working directory guidance
|
||
|
||
### Workflow Studio runtime
|
||
|
||
- Candidate workflow graphs remain editable after they are loaded into the Coze-compatible canvas; the compatibility response explicitly marks the current owner as editable so node dragging is not mistaken for a read-only preview.
|
||
- Workflow-agent SSE forwards provider `reasoning_content` as an inline `<think>…</think>` trace for the UI while keeping the executable node output limited to the visible answer text. Every native LangGraph `[AIMessage|ToolMessage, metadata]` frame is additionally persisted and sent as `node.message.data.messages` without projection or length cropping, so the Studio can replay DeerFlow's tool-result, search-result and `ask_clarification` cards exactly after reconnecting.
|
||
|
||
### AgentScope report collaboration (in progress)
|
||
|
||
The independent `/api/report-collaboration` workbench supports durable run-time intervention commands and immutable formal report versions. Its production worker now runs a real AgentScope 2.0.5 `Agent` per TaskLedger assignment; AgentScope owns the member ReAct/tool/structured-output loop, while TaskLedger alone owns DAG readiness, retries and completion. Ready nodes in one wave run concurrently up to `max_parallel_tasks`; cancellation and member exceptions close agent runs as terminal states, and lease recovery remains independent of in-memory AgentScope objects. Retrieval uses the deployment's configured DeerFlow `web_search`/`web_fetch` tools through the frozen role permission snapshot. Every visible AgentScope reply, TeamSay handoff, tool call and tool result is persisted immediately into the existing SSE/message protocol; hidden reasoning and system prompts are not exposed. A constrained intent agent classifies explicit answers, angle additions, re-search, re-analysis, node reruns, section rewrites, polishing, report questions, replanning, and cancellation. Server-side impact analysis validates node/agent ownership and computes the minimum downstream lineage; low-confidence, high-cost, or multi-branch changes require confirmation. Commands are CAS-claimed and applied at TaskLedger safe points, where old attempts and artifacts become `superseded` so late results cannot contaminate the revised report. Run creation freezes the selected `/api/report-structures` row and all server-enforced budgets. Per-user active-run limits and per-run model-call, retrieval, source-byte, token, cost, node-time, and total-time limits fail closed. User-visible event payloads are centrally redacted before durable persistence without size truncation; external source content is explicitly untrusted, and hidden reasoning/system prompts are never exposed. Append-only audits cover agent selection, tools, sources, user commands, confirmations, report versions, redactions, and budget failures; owner-scoped reads are available at `GET /runs/{id}/audits`. Temporary Markdown is emitted as `report.delta`; only a template/citation/coverage-clean report with an independent passing review is atomically persisted and emitted as `report.version.created` before the run becomes completed. Section rewrites are replacement candidates; apply and restore append a new head version and never destroy history. Configure `runtime_mode`, concurrency and limits under `report_collaboration`; both the feature and worker remain disabled by default. Install the `report-collaboration` project extra when enabling the worker.
|
||
|
||
### Middleware Chain
|
||
|
||
Middlewares execute in strict order, each handling a specific concern:
|
||
|
||
| # | Middleware | Purpose |
|
||
|---|-----------|---------|
|
||
| 1 | **ThreadDataMiddleware** | Creates per-thread isolated directories (workspace, uploads, outputs) |
|
||
| 2 | **UploadsMiddleware** | Injects newly uploaded files into conversation context |
|
||
| 3 | **SandboxMiddleware** | Acquires sandbox environment for code execution |
|
||
| 4 | **SummarizationMiddleware** | Reduces context when approaching token limits (optional) |
|
||
| 5 | **TodoListMiddleware** | Tracks multi-step tasks in plan mode (optional) |
|
||
| 6 | **TitleMiddleware** | Auto-generates conversation titles after first exchange |
|
||
| 7 | **MemoryMiddleware** | Queues conversations for async memory extraction |
|
||
| 8 | **ViewImageMiddleware** | Injects image data for vision-capable models (conditional) |
|
||
| 9 | **ClarificationMiddleware** | Intercepts clarification requests and interrupts execution (must be last) |
|
||
|
||
### Sandbox System
|
||
|
||
Per-thread isolated execution with virtual path translation:
|
||
|
||
- **Abstract interface**: `execute_command`, `read_file`, `write_file`, `list_dir`
|
||
- **Providers**: `LocalSandboxProvider` (filesystem) and `AioSandboxProvider` (Docker, in community/)
|
||
- **Virtual paths**: `/mnt/user-data/{workspace,uploads,outputs}` → thread-specific physical directories
|
||
- **Skills path**: `/mnt/skills` → `deer-flow/skills/` directory
|
||
- **Skills loading**: Recursively discovers nested `SKILL.md` files under `skills/{public,custom}` and preserves nested container paths
|
||
- **Conditional WeKnora LLMWiki**: a non-empty `llmwiki.weknora.api_base_url` enables the DeerFlow-native WeKnora BFF and UI; an empty address preserves the existing LLMWiki page. WeKnora remains independently deployed and owns its database, Redis, storage, parsers and models. DeerFlow keeps only the API/admin addresses, user/publication mappings, agent knowledge-base bindings and per-conversation selections; it stores and sends no WeKnora credential, so the endpoint must be restricted to a trusted network and allow DeerFlow traffic. Owners publish directly to the public LLMWiki without an approval queue for ordinary bases; the system-owned conversation-deposit base is always published, shared by all users, excluded from My Knowledge, immutable through user/admin detail operations, and recognized by both its reserved name and `system` owner. An existing same-name remote mapping is adopted in place as the system deposit so its mapping id and document links stay valid. Its user-facing knowledge-base description and generated document metadata use the `cmzs` brand; list/detail reads synchronize stale remote descriptions, while conversation-deposit search results, source previews and embedded-detail JSON rewrite legacy brand text to `cmzs`. Deposit requests must reference a thread accessible to the current account. Personal knowledge-base lists are scoped to the current DeerFlow account even for admins, while frontend LLMWiki query caches also include the active account id so switching accounts cannot reuse another account's list/detail data. Admins remain the only role that sees the WeKnora connection-config tab. Knowledge cards open a dedicated WeKnora-style detail route: Wiki-enabled bases default to Wiki, then documents and graph; standard bases hide Wiki and default to documents. Writable ordinary knowledge-base details expose a native multi-file upload action that submits files sequentially and reports aggregate success/failure; read-only public bases and the system deposit base do not. It supports file/URL/manual-Markdown imports, Wiki page browsing, document preview, real parsed chunks and Wiki graph nodes/edges. New chats use a compact WeKnora-style searchable/grouped multi-select inside the composer, while agent create/edit pages bind spaces through an inline panel. A single proxy service identity means DeerFlow users are isolated logically while their WeKnora entities share that identity's workspace. The read-only `llmwiki_search` tool reauthorizes the current DeerFlow user on every call and renders credential-free source previews; source previews prefer cleaned Wiki/LLMWiki text when a retrieved chunk maps to a Wiki page, with raw chunks as the fallback. The detail page first creates a signed iframe session through the normal DeerFlow API auth path, then points the iframe directly at the proxied WeKnora `/platform/...` route using the signed `deerflow_weknora_embed` cookie; global auth/CSRF bypass is limited to those signed iframe proxy requests, and HTTPS cookie attributes honor `X-Forwarded-Proto` behind reverse proxies. The iframe proxy keeps WeKnora's root-relative `/platform`, `/assets`, `/locales`, and `/api/v1/*` paths, but is registered last through `llmwiki.proxy_router` so it cannot intercept first-party Gateway APIs such as `POST /api/v1/auth/login/username`. See [docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md](docs/LLMWIKI_WEKNORA_INTEGRATION_ZH.md).
|
||
- **Enterprise Research workbench (temporarily disabled)**: the independent `/api/enterprise-research` router, its report executor, and its ORM table registration are intentionally inactive. The implementation remains in the repository, but startup does not create or alter `enterprise_research_tasks` or `enterprise_research_report_jobs`; this keeps the active Deep Research workflow independent of the paused workbench and avoids MySQL schema errors.
|
||
- **Embedded WeKnora detail chrome**: the proxy hides WeKnora's application sidebar (`.main > .aside_box`) and expands the detail outlet to the full iframe width, while preserving the Wiki page's own index/navigation panel. The host does not sandbox this trusted, permission-filtered detail frame because browser PDF viewers are disabled inside sandboxed ancestor frames. Preview and knowledge-base file redirects are followed by the Gateway so PDF bytes and protected Wiki images remain inside the authorized proxy instead of leaking to an unreachable object-storage URL.
|
||
- **Embedded WeKnora knowledge-base switching**: the original breadcrumb picker works inside the iframe, but lists only the signed-in user's own and published knowledge bases. Every switch is reauthorized through DeerFlow and receives a fresh short-lived iframe session, so it never exposes the upstream service-admin knowledge-base list.
|
||
- **Prefixed WeKnora deployment**: the frontend may be published under a separate prefix such as `/magentweb`, while embedded WeKnora detail traffic is served by the Gateway under `/deerflow`. The Gateway rewrites HTML root assets, WeKnora API URLs, redirects, and Vite dynamic dependency maps so page assets stay under `/deerflow`; the prefixed proxy routes must be registered before the unprefixed fallback. The signed embed cookie covers `/deerflow`, and auth/CSRF recognition includes all supported static directories.
|
||
- **File-write safety**: `str_replace` serializes read-modify-write per `(sandbox.id, path)` so isolated sandboxes keep concurrency even when virtual paths match
|
||
- **Tools**: `bash`, `ls`, `read_file`, `write_file`, `str_replace` (`bash` is disabled by default when using `LocalSandboxProvider`; use `AioSandboxProvider` for isolated shell access)
|
||
|
||
### Subagent System
|
||
|
||
Async task delegation with concurrent execution:
|
||
|
||
- **Built-in agents**: `general-purpose` (full toolset) and `bash` (command specialist, exposed only when shell access is available)
|
||
- **Concurrency**: Max 3 subagents per turn, 15-minute timeout
|
||
- **Execution**: Background thread pools with status tracking and SSE events
|
||
- **Flow**: Agent calls `task()` tool → executor runs subagent in background → polls for completion → returns result
|
||
|
||
### Memory System
|
||
|
||
LLM-powered persistent context retention across conversations:
|
||
|
||
- **Automatic extraction**: Analyzes conversations for user context, facts, and preferences
|
||
- **Structured storage**: User context (work, personal, top-of-mind), history, and confidence-scored facts
|
||
- **Debounced updates**: Batches updates to minimize LLM calls (configurable wait time)
|
||
- **System prompt injection**: Top facts + context injected into agent prompts
|
||
- **Storage**: JSON file with mtime-based cache invalidation
|
||
|
||
### Tool Ecosystem
|
||
|
||
| Category | Tools |
|
||
|----------|-------|
|
||
| **Sandbox** | `bash`, `ls`, `read_file`, `write_file`, `str_replace` |
|
||
| **Built-in** | `present_files`, `ask_clarification`, `view_image`, `task` (subagent) |
|
||
| **Community** | Tavily (web search), Jina AI (web fetch), Firecrawl (scraping), DuckDuckGo (image search) |
|
||
| **MCP** | Any Model Context Protocol server (stdio, SSE, HTTP transports) |
|
||
| **Skills** | Domain-specific workflows injected via system prompt |
|
||
|
||
#### Configurable HTTP web search
|
||
|
||
The config-defined `web_search` tool can call a deployment-specific JSON API
|
||
instead of a public search provider. Set its `endpoint`, static `payload`, SSL
|
||
verification, timeout, result limit, and `result_url_template` in `config.yaml`.
|
||
At runtime only `payload.query` is overwritten with the model's search text.
|
||
The API response must use `{"results": [{"recUuid": "", "content": "",
|
||
"title": "", "url": ""}]}`; missing or null result fields are normalized to empty strings. When
|
||
`recUuid` is present, `{recUuid}` in `result_url_template` replaces the original
|
||
result URL; otherwise the original URL is retained. Set `enabled: false` to keep
|
||
the tool completely out of the agent's registered tool list. `verify_ssl: false`
|
||
supports trusted internal HTTPS services with self-signed certificates.
|
||
|
||
### Gateway API
|
||
|
||
FastAPI application providing REST endpoints for frontend integration:
|
||
|
||
| Route | Purpose |
|
||
|-------|---------|
|
||
| `GET /api/models` | List available LLM models |
|
||
| `GET/PUT /api/mcp/config` | Manage MCP server configurations |
|
||
| `GET/PUT /api/skills` | List and manage skills |
|
||
| `POST /api/skills/install` | Install skill from `.skill` archive |
|
||
| `POST /api/skills/validate`, `POST /api/skills/install-upload` | Validate and upload `.zip`/`.skill` packages with owner-aware conflicts: owners may explicitly overwrite their own copy; other users' copies report the uploader and cannot be overwritten |
|
||
| `GET/POST /api/agents` | List visible agents or create a user-owned agent |
|
||
| `GET/PUT/DELETE /api/agents/{id}` | Read, update, or delete an id-addressed agent |
|
||
| `GET/POST/PUT/DELETE /api/roundtable-chains` | Manage user-owned roundtable business chains; each seat may optionally carry `position_id` for the separate position-collaboration workspace, plus `human_validation` to mark a Step 2 client-side pause checkpoint after that seat/stage completes (the original roundtable has no extra validation card) |
|
||
| `GET/PUT /api/business-mapping` | Read global business mappings. Updates to 3Q/6BF/7BF require an administrator; the 8BF public business-chain mapping is maintained exclusively by the trusted `lqq` account (derived from the authenticated email prefix). |
|
||
| `GET/POST/PUT /api/position-roundtable/sessions` | Standalone position-collaboration sessions: freeze task/intent/chain snapshots, index one direct-agent node per seat, build upstream handoff context (including bounded UTF-8 excerpts of text artifacts), and make archived sessions read-only without invoking the original roundtable coordinator or jobs. Built-in summary/action-plan completion mirrors visible report text into a Markdown artifact when a model omits its final `write_file` call. `POST /sessions/{id}/nodes/{nodeKey}/reject` accepts an optional reason, marks the rejected delivery for rework, and invalidates affected downstream deliveries so they are regenerated in order. The same router is also mounted at `/api/multi-agent/position-roundtable/*` for intranet gateways that only forward the established multi-agent namespace; the compatibility mount is intentionally omitted from OpenAPI. |
|
||
| `GET /api/memory` | Retrieve memory data |
|
||
| `POST /api/memory/reload` | Force memory reload |
|
||
| `GET /api/memory/config` | Memory configuration |
|
||
| `GET /api/memory/status` | Combined config + data |
|
||
| `POST /api/threads/{id}/uploads` | Upload files (auto-converts PDF/PPT/Excel/Word to Markdown, rejects directory paths) |
|
||
| `GET /api/threads/{id}/uploads/list` | List uploaded files |
|
||
| `DELETE /api/threads/{id}` | Delete DeerFlow-managed local thread data after LangGraph thread deletion; unexpected failures are logged server-side and return a generic 500 detail |
|
||
| `GET /api/artifact-library` | List the current user's generated conversation files across normal threads, with filename/conversation search, file-kind filtering, pagination, and an optional cache-bypassing refresh. The filesystem `outputs/` directories remain authoritative; uploads and image/audio/video files are excluded. |
|
||
| `GET /api/threads/{id}/artifacts/{path}` | Serve generated artifacts |
|
||
| `GET/POST /api/deep-research/*` | Durable Deep Research sessions and jobs. `quick`/`basic` run the compact report flow; `detailed` plans subtopics and writes independent sections; `deep` recursively investigates bounded follow-up branches. All modes reuse DeerFlow's configured material provider, LLM factory, hidden-thread artifacts, and SSE event log. Deep Research session/job/event/source/message writes serialise their rows in the writer transaction, rather than doing a post-commit ORM refresh or readback; this prevents MySQL read/write-splitting replica lag from failing “start writing”, dispatcher claim, report persistence, or follow-up history. Whole-document artifact/report rewrites likewise create their reversible version snapshots from writer-local ORM data and never refresh after commit, so replica lag cannot restore an otherwise successful rewrite. A report is written to the session and confirmed before its Job becomes `completed`; an unconfirmed write is retried and never exposed as a false completion. Retrying an already-created queued job re-nudges the dispatcher, recovering an interrupted HTTP response without creating a second report. If progress-event persistence, durable replay, or live fan-out is briefly unavailable, the worker/stream logs the degradation and continues or reconnects rather than failing the whole job. Whole-report rewrite accepts an optional `reportOutline` Markdown template; the durable job persists it and applies it as the chapter structure constraint. Optional report illustrations use a separately configured OpenAI-compatible image endpoint and are served only through the owning session's authenticated artifact route. |
|
||
| `/api/enterprise-research/*` | Temporarily disabled. Its router and persistence initialization are not registered, so these endpoints return 404 and no enterprise-research tables are created or migrated during startup. |
|
||
| `POST /api/ai-writing/sample/extract` | Extract text from Word/PDF/Markdown/TXT samples for AI-writing imitation mode |
|
||
| `POST /api/writing/export/docx` | Generate the shared formal Word download from Markdown. Uses the standard-library OOXML builder in `app/gateway/word_export.py`, removes manual/compound heading prefixes and Markdown horizontal rules, applies the prescribed A4/margin/heading/body/table/footer profile, and embeds the server-bundled `方正小标宋简体.ttf`; the client computer does not need that font installed. |
|
||
| `POST /cop/saveSuperiorTask` | TaskCOP compatibility API (DeerFlow Bearer/session authentication required): create a situation-overview task with a sequential four-digit id (`0001`, `0002`, ...) and return the legacy `{state, msg, data}` envelope |
|
||
| `POST /api/taskcop/import/tasks` | TaskCOP compatibility API (authenticated): import tasks from a JSON array or `tasks`/`records` envelope; every row's supplied `id`/`taskId` is authoritative and a duplicate id completely replaces the stored task |
|
||
| `DELETE /cop/tasks/{task_id}` | TaskCOP compatibility API (authenticated): delete one task and its task-scoped situation-report document |
|
||
| `GET /taskAnalyseSearch/cop-task-three-list` | TaskCOP compatibility API (authenticated): paginated task list for the situation-overview sentiment page (`content`, `taskStatus`, `taskDirection`, `startTime`, `endTime`) |
|
||
| `GET/POST/PUT /api/task-reports/by-task/{task_id}` | Shared task-scoped situation-report detail. Query, first save, and full update all use the legacy-compatible `{state, msg, data:[{id, taskId, createTime, updateTime, sessionId, contentJson, categoryType}]}` shape; `contentJson` is preserved as a JSON string. Saving a non-empty report automatically moves the matching TaskCOP task to status `25` (the legacy list's “已完成” state). |
|
||
| `POST /api/task-reports/import/by-task/{task_id}` | TaskCOP detail import (authenticated): accept a `sq-report-mock.json`-compatible `data`/`records` array and atomically replace the selected task's full report. The path task id overrides task ids inside the file. The agentfx built-in agent fills that file with the `task-report-build` convert skill in one call (search JSON → 14 categories; `task-report-{enemy,our,env,judge}` re-run a single dashboard page) then imports via `task-report-import`. |
|
||
| `POST /api/sentiment-agent/stream` | Authenticated BFF for the frontend sentiment-analysis virtual agent. Forwards the AG-UI JSON body to the `apiUrl` supplied from `runtime-config.js` (Basic Auth from the same payload). TLS certificate verification is off. Upstream SSE is piped through unchanged; connection/HTTP failures return the complete error in `detail`. |
|
||
|
||
The separately deployed TaskCOP Vue app uses `POST /api/v1/auth/login/username` directly (its `username`, `userId`, or `yUserId` URL parameter identifies the user), so this compatibility flow does not depend on the retired Consumer login service. It calls the Gateway address configured as `b1ConsumerUrl` directly; it does not use a Vite `/api` reverse proxy. An empty TaskCOP store remains empty; task data is introduced through normal creation or the JSON import API. The Gateway CORS middleware always merges local Vite origins `http://localhost|127.0.0.1:{5173,5174,3000,8080}`, including when `GATEWAY_CORS_ORIGINS` is set for another frontend. This keeps a Vite server that falls back from 5173 to 5174 (or the reverse) working after a backend restart. For a deployment, set `GATEWAY_CORS_ORIGINS` to the exact additional frontend origins that are allowed to call the Gateway.
|
||
|
||
### Workflow Studio interface resources
|
||
|
||
The Workflow Studio data-source API also registers reusable HTTP interface resources. Create them through
|
||
POST /api/workflows/data-sources with kind set to http; provide a non-secret baseUrl, an allowedMethods
|
||
allowlist, and optional request headers. The service encrypts the headers immediately and never returns
|
||
them through list, detail, resource-catalog, run-event, or canvas APIs. The canvas resource catalog
|
||
exposes only the name, description, HTTP address, and allowed methods. HTTP execution continues to
|
||
enforce the workflow target security policy, including the default SSRF restrictions.
|
||
|
||
The same catalog's agent and skill entries are display projections, not a second ownership store:
|
||
`/api/workflows/resources/agents` returns `agentId`, cached available `skills`, and `createdBy`; the
|
||
skills endpoint returns `skillId`, direct-call capability, and `createdBy`. Creator names are resolved
|
||
from the immutable owner id at read time (falling back to the id if an account was removed), while
|
||
built-in resources explicitly return `系统内置`.
|
||
|
||
### Conversation-driven workflow planning
|
||
|
||
The Workflow Studio chat composer does not silently start a template. A normal user
|
||
message first creates a durable planning session with
|
||
`POST /api/workflows/{workflow_id}/planning-sessions/stream` (SSE; the legacy
|
||
non-streaming `POST /api/workflows/{workflow_id}/planning-sessions` remains for
|
||
integrations). The stream relays safe DeerFlow controller milestones—catalog read,
|
||
controller parsing, role selection, graph assembly and validation—before returning
|
||
the final session, so the UI never has to fabricate a progress bar. The constrained planner builds
|
||
two or three validated candidate graphs from the current draft and the visible business
|
||
agents the user can use. It is a dedicated built-in **工作流总控** (`workflow-planner`),
|
||
not the multi-agent roundtable coordinator: it only reads the user requirement and
|
||
catalog metadata, selects permitted agent ids, and ranks the supported strategies. It
|
||
does not dispatch a roundtable, use research tools, or execute a formal run. The server
|
||
constructs the executable graph from those selected, visible resources; a controller
|
||
failure returns an explicit planning error rather than a keyword/template fallback. The
|
||
controller call has no callable tools and disables model thinking, so it adds one bounded
|
||
planning inference without starting research or worker-agent costs. A
|
||
browser selects one candidate, automatically applies the editor's DAG layout when
|
||
it loads, then may explicitly save an edited graph through
|
||
`PATCH /api/workflows/planning-sessions/{session_id}/proposals/{proposal_id}/graph`,
|
||
then creates the one formal run through the existing
|
||
`POST /api/workflows/{workflow_id}/runs` endpoint with `planningSessionId` and
|
||
`proposalId`. The gateway rejects unselected or invalid candidates and makes a repeated
|
||
confirmation return the already-created run, so a double click cannot execute a second
|
||
workflow. Planning sessions are persisted separately from runs (migration
|
||
`20260831_01`); their parent row is flushed before candidate rows so the foreign-key
|
||
transaction works consistently across SQLite, MySQL, and PostgreSQL. A candidate is
|
||
not consuming execution resources until it is confirmed.
|
||
|
||
The development Studio proxy must not gzip `/api` responses: compressed SSE can hold
|
||
small planning and run-event frames until the response ends. The backend already marks
|
||
both streams `no-transform` and `X-Accel-Buffering: no`; `apps/coze-studio` also sets
|
||
its Rsbuild development server `compress: false`. Agent/skill nodes use the bounded
|
||
`workflows.agent_recursion_limit` (default `250`), rather than a chat-sized 60-step
|
||
limit, because a model-to-tool research turn consumes about two LangGraph super-steps.
|
||
The composer may also send a configured `modelName`: the server validates it against
|
||
the safe model catalog, uses it for the workflow controller, and writes it into the
|
||
generated candidate agent nodes. For a direct single-agent task it overrides only that
|
||
projected node; confirmation still executes the immutable selected proposal snapshot.
|
||
The controller additionally creates a per-selected-agent task contract (`mission`,
|
||
`deliverable`, `scope`, `handoff`). The graph builder accepts contracts only for
|
||
catalog-visible, selected agent ids and injects each one into that worker's prompt, so
|
||
parallel branches receive complementary research assignments and sequential branches
|
||
have an explicit review/handoff boundary rather than all workers attempting the whole
|
||
user task. Missing controller fields receive a deterministic role-aware contract.
|
||
When an agent emits DeerFlow's native `ask_clarification` ToolMessage, the workflow
|
||
enters `awaiting_input` rather than completing or failing. The pause event references
|
||
the originating `nodeRunId` and `toolCallId`; the browser renders that original card
|
||
inside the agent message and `POST /runs/{run_id}/resume` supplies the user's reply as
|
||
the next turn of the same agent thread.
|
||
|
||
The current planner is deliberately constrained: it can rank supported strategies and
|
||
focus resources, but the server builds the graph from the already configured agent
|
||
nodes. It cannot invent a node type, external target, or credential from chat text.
|
||
For a parallel-research candidate, the configured research agents first flow into the
|
||
deterministic `evidence_normalizer` node. It emits a bounded Evidence Pack with the
|
||
originating node, role label, explicit gaps, and no invented verification claim before
|
||
the synthesis agent sees it. The candidate then runs `deep_research_write`: its
|
||
app-layer adapter converts exactly that Evidence Pack into selected, provenance-tagged
|
||
Deep Research source rows, starts or reattaches to the existing durable Deep Research
|
||
job, forwards report deltas as workflow `node.output.delta` events, and exposes the
|
||
final Markdown as a run artifact. Cancellation is forwarded to the same research job;
|
||
the node never calls the Deep Research HTTP API or launches a second writer. The
|
||
session key includes the topic, evidence snapshot and writing configuration, so a retry
|
||
reuses the same job while a changed upstream pack cannot return stale report prose. `multi_agent` Deep Research is
|
||
intentionally not nested yet because its separate plan-review pause would conflict
|
||
with workflow `human_input`. A completed `agent` or `skill` node in a terminal,
|
||
acyclic run can now receive targeted feedback through
|
||
`POST /api/workflows/runs/{run_id}/feedback`: the original remains immutable, a
|
||
revision run reuses completed nodes outside the target's downstream closure, and
|
||
only the target prompt receives the bounded feedback block before it and its
|
||
downstream nodes stream again. The public revision run retains the affected and
|
||
reused node ids in its feedback summary for the conversation/history UI. Running and looped workflows deliberately continue
|
||
to use `human_input` or a new full run; they are not modified in place. See
|
||
`docs/WORKFLOW_CONVERSATIONAL_MULTI_AGENT_ORCHESTRATION_ZH.md` for the roadmap.
|
||
|
||
### Deep Research modes
|
||
|
||
Deep Research is a separate Gateway feature rather than part of AI Writing. Its
|
||
`detailed` runner adapts GPT Researcher's detailed-report control flow (plan
|
||
subtopics, collect materials concurrently, write sections concurrently, then
|
||
assemble the report). Its `deep` runner adapts the recursive DeepResearchSkill
|
||
flow (branch, extract learnings, investigate follow-ups) with a deterministic
|
||
query budget derived from `deep_breadth` and `deep_depth`. The runners live in
|
||
`packages/harness/deerflow/agents/deep_research/runners/` and only use injected
|
||
DeerFlow adapters. `multi_agent` is a LangGraph 1.x editor/researcher/writer/
|
||
reviewer graph: it persists a plan-review pause, releases the worker lease, and
|
||
resumes from the approved or revised plan through the existing durable-job API.
|
||
Report prose uses the configured LangChain model's `astream()` path. Each
|
||
native model delta is fanned out as an in-process `report_delta` SSE frame so
|
||
the right-side `report.md` sandbox can write smoothly, while bounded,
|
||
persisted `report_chunk` events remain reconnect checkpoints. The durable log
|
||
is still authoritative: a browser reconnect, slow consumer, or another worker
|
||
falls back to `?after=<seq>` replay rather than starting a second model call.
|
||
Reasoning exposed by providers either through structured reasoning fields or
|
||
inline `<think>...</think>` text is separated from report deltas, checkpoints,
|
||
and the persisted Markdown projection. It may be forwarded as an ephemeral
|
||
live `report_thinking` frame for the message-list thought trace, but is never
|
||
persisted or written to the sandbox; only answer Markdown reaches the sandbox.
|
||
The short non-streaming control-plane calls used to plan queries, sections, and
|
||
reviews are independently capped at 45 seconds (maximum 120 seconds). If query
|
||
or section planning times out, the runner emits a visible recoverable warning
|
||
and searches the original research topic instead of leaving the job at the
|
||
planning step indefinitely. Report-prose streaming is intentionally not capped
|
||
by this control-plane limit.
|
||
Planning also persists `queries_planned` before each basic, detailed, or
|
||
multi-agent parallel search fan-out; the completed session stores its terminal
|
||
`lastJobId` in the usage snapshot so
|
||
the conversation workbench can replay the full planning/search/source/curation
|
||
trace after a page refresh. These are observable execution facts only, never
|
||
hidden model reasoning.
|
||
Detailed reports use stable introduction/section/conclusion targets; a
|
||
multi-agent review revision emits `report_reset` before the replacement draft.
|
||
Every session may also provide a bounded `custom_outline`. Markdown headings or
|
||
numbered entries become the deterministic section plan for `detailed` and
|
||
`multi_agent`; `basic` and `deep` receive the same outline as a guarded
|
||
report-writing constraint. The outline is frozen in the session snapshot and
|
||
never treated as a system instruction.
|
||
|
||
All owner-scoped Deep Research routes resolve the platform's asynchronous
|
||
`get_current_user` identity before querying or writing persistence, so session
|
||
and job records are always bound to a concrete user ID.
|
||
|
||
Report artifacts are server-named and written through the canonical sandbox
|
||
virtual path (`/mnt/user-data/outputs/<name>`), which the path layer maps to
|
||
the host filesystem. This keeps native Windows deployments compatible without
|
||
relaxing the sandbox traversal guard.
|
||
|
||
Deep Research reuses the DeerFlow Q&A retrieval stack instead of a separate
|
||
scraper. It first runs enabled research skills (`deep-search`, `web-research`,
|
||
and enabled knowledge-base skills) through the same `SubagentExecutor` runtime
|
||
used by Q&A, then uses a configured `web_search` provider when present. If this
|
||
offline deployment still has the shipped placeholder/disabled `web_search`
|
||
configuration, it falls back to DeerFlow's existing keyless DuckDuckGo provider
|
||
and, when that provider returns only irrelevant engine noise, a second real-web
|
||
search source. It persists only real returned sources. A missing intranet
|
||
endpoint therefore cannot by itself produce an empty Deep Research report. The capabilities API
|
||
reports `skill` only when its required Q&A tool is available, and reports
|
||
`web_search` when configured search or the online fallback can run.
|
||
Cancelling a queued research job finalizes its job row and immediately marks
|
||
the owning session `cancelled` with no active job, so it cannot retain a
|
||
per-user concurrency slot.
|
||
Saved `draft` sessions do not consume a concurrency slot; the two-job limit
|
||
only applies to sessions with a `running` or `awaiting_input` durable job.
|
||
|
||
Whole-report rewrite builds its requirement plan locally from the selected
|
||
style and instruction, publishes that plan immediately, and then opens only
|
||
one model stream for the actual Markdown writer. This avoids making users wait
|
||
for a separate model-thinking/planning pass before sandbox output begins.
|
||
`POST /api/deep-research/sessions/{id}/report-variants` is the structural
|
||
regeneration boundary: it copies the completed report and selected evidence
|
||
into a new session (including both saved `[来源:id]` and legacy streamed
|
||
`[[source:id]]` citation markers remapped to the cloned ids). The following
|
||
durable writer is explicitly queued with `operation=generate`: it does not send
|
||
the inherited report to the model as text to edit, and instead writes a fresh
|
||
report from the complete selected evidence set plus the newly confirmed
|
||
outline. Model reasoning remains a live SSE event before the first Markdown
|
||
token. New output uses display-ready `[来源:id]` citations; validation repairs
|
||
only unique near-complete `drs_src_*` ids (for example a provider-truncated
|
||
suffix) and still rejects genuinely unknown evidence ids.
|
||
The report-only model stream requests an 8192-token completion budget so a
|
||
normal multi-section report and its reference list do not stop at the common
|
||
4096-token provider default; malformed trailing citation syntax is rejected
|
||
rather than committed as a partial document.
|
||
The original report therefore remains independently available in history even
|
||
if the new report is cancelled or fails.
|
||
|
||
#### Optional report illustrations
|
||
|
||
DeerFlow's stock configuration includes image recognition but no image-generation
|
||
provider. Deep Research therefore keeps illustrations disabled until an operator
|
||
fills `config.yaml -> deep_research.image_generation` with an
|
||
image endpoint, API key, and model. `provider: openai_compatible` uses the
|
||
standard `/images/generations` protocol and requires `b64_json`; `provider:
|
||
dashscope_native` supports Qwen Image 3.0 and Wan 2.7 on DashScope / Token Plan
|
||
with the native multimodal-generation protocol. Native provider image URLs are
|
||
downloaded immediately, so both providers produce protected local artifacts.
|
||
When configured, the capabilities endpoint enables the page's “生成报告配图” switch.
|
||
A selected job writes at most four `research-image-*.png|jpg|webp` artifacts,
|
||
injects `deep-research://` links into the Markdown report, and emits image SSE
|
||
events. A failed/unconfigured image provider only emits a recoverable warning;
|
||
the text report still completes.
|
||
|
||
Completed reports also expose `GET /api/deep-research/sessions/{id}/messages`
|
||
and `POST /api/deep-research/sessions/{id}/chat`. Follow-up answers are grounded
|
||
only in that session's report, selected sources, and recent follow-up history;
|
||
the assistant returns persisted source ids for every accepted citation. Passing
|
||
`allow_new_research=true` is rejected rather than silently triggering another
|
||
search job. Owners can change the follow-up evidence set with
|
||
`PATCH /api/deep-research/sessions/{id}/sources/{source_id}` after a research
|
||
job stops. That selection is applied immediately to later follow-up prompts;
|
||
it never rewrites the completed report, and edits are rejected while the job is
|
||
running so they cannot race the automatic curation projection.
|
||
|
||
The Deep Research page is a conversation workspace: the original research
|
||
question, observable retrieval steps, completion summary, and report-grounded
|
||
follow-ups appear in chronological order. The report itself is written in the
|
||
adjacent sandbox as `/mnt/user-data/outputs/report.md`, with Word, Markdown,
|
||
and HTML export actions. Word export calls the shared Gateway
|
||
`POST /api/writing/export/docx` generator, embeds the licensed title font, and
|
||
includes the collected reference list without adding a backend Python package.
|
||
In chat-collection sessions the report configuration card
|
||
shows the unique material count inferred from the collector's visible tool
|
||
results before writing begins. The write lifecycle itself is rendered as
|
||
ordinary workspace message-list tool steps (not a second timeline component),
|
||
and remains in the transcript after completion; the generated digest is a
|
||
normal assistant message below those steps. As soon as the runner enters its
|
||
summary phase, that same assistant message is inserted in a streaming state
|
||
and receives each SSE `summary_delta`; completion only changes it in place to
|
||
the final report-complete wording rather than appending a delayed second
|
||
summary card.
|
||
|
||
Each completed chat-collection report also contributes a normal file card to
|
||
the message history, named from the report Markdown title rather than the
|
||
internal `report.md` artifact path. Historical sessions keep the sandbox closed
|
||
on replay, and opening that card selects the session's virtual report artifact
|
||
in the sandbox; its download action uses the existing Markdown/Word export
|
||
dialog. The per-user active-job limit and the default in-job retrieval/write
|
||
concurrency are both three. The Deep Research composer uses the installed
|
||
TDesign `Select`, populated from `/api/models`, for an optional per-research
|
||
model override. It is disabled during active collection/writing, is sent as
|
||
`model_name` to collector turns, and is applied to the `fast_model`,
|
||
`smart_model`, and `strategic_model` roles when a session or report job starts.
|
||
|
||
Completed-report Q&A uses `POST /api/deep-research/sessions/{id}/chat/stream`
|
||
for native SSE deltas; the earlier `/chat` JSON endpoint remains for API
|
||
compatibility. An explicit scoped edit such as “修改第一段”、
|
||
“重新生成一下第一段” or “重写第二节” uses the separate
|
||
`POST /api/deep-research/sessions/{id}/rewrite/stream` report-section writing
|
||
pipeline. That route requires a deterministic Markdown range and returns 422
|
||
instead of falling back to Q&A when no target is supplied. Ordinary follow-up
|
||
questions remain report/source-grounded and never mutate the report. The streamed
|
||
replacement candidate is shown in the message list while the sandbox keeps the
|
||
current report unchanged. The completed candidate creates a confirmation card;
|
||
only `POST /sessions/{id}/messages/{messageId}/rewrite-proposal` with
|
||
`{"action":"apply"}` substitutes the matching range and saves the assembled
|
||
Markdown. The proposal carries the source report hash, so a stale candidate is
|
||
rejected instead of overwriting a later edit; dismissing/cancelling leaves the
|
||
persisted report unchanged. This does not initiate new web research. Legacy
|
||
unmarked candidates from before this route split are inferred only when their
|
||
immediately preceding user message contains the same explicit scoped edit.
|
||
|
||
Position-roundtable history supports two modes. A non-empty `external_task_id` from the frontend route's `taskId` creates a row with `task_scoped = true`, so all sessions, nodes and persisted conversation snapshots for that task are shared across users; `user_id` is then creator audit metadata only. When the field is absent or empty, the original per-user personal-session behavior remains unchanged. Alembic revision `20260808_01` adds `task_scoped` with the safe default `false` for historical rows plus the task/update-time index.
|
||
|
||
### Roundtable concurrency guarantees
|
||
|
||
Position-roundtable node completion, rejection, thread binding and activation
|
||
are database transactions with row locking and retry-stable command ids. This
|
||
keeps simultaneous seat completion from leaving a downstream stage locked.
|
||
Task and intent snapshots are frozen when the chain becomes active, so a late
|
||
request from another user cannot alter the confirmed input. Background
|
||
roundtable jobs use a database unique active-job key and renewable worker
|
||
leases; cancelled jobs are finalized by the lease holder or recovered after an
|
||
expired lease, allowing users to start a new job after a worker failure.
|
||
|
||
The ordinary unit tests use SQLite. To verify deployed database semantics with
|
||
real PostgreSQL row locks and independent connections, set
|
||
`DEERFLOW_POSTGRES_CONCURRENCY_TEST_URL` and run:
|
||
|
||
```bash
|
||
PYTHONPATH=. uv run --no-sync pytest tests/test_roundtable_postgres_concurrency.py -q
|
||
```
|
||
|
||
AI writing supports a sample-imitation mode: the frontend can upload or paste a sample article, the Gateway extracts text for document files, and the `ai_writing` graph runs a `sample_analyzer` node before normal intent/research/outline/draft stages. The resulting style profile is injected into outline and draft prompts with explicit guardrails against copying source sentences, facts, names, or data.
|
||
|
||
### Custom Agents
|
||
|
||
Custom agents are indexed in the `agents` database table. The `id` column is the stable runtime identifier and filesystem directory name (`.deer-flow/agents/{id}/SOUL.md`); `name` is display-only, may be Chinese, and may be duplicated. Rows with `user_id = NULL` are built-in agents, user rows are private unless `published = true`, and users can only update or delete agents they created.
|
||
|
||
The built-in `forced-research-responder` (「强制检索输出助手」) is seeded at Gateway startup. Unlike an ordinary skill-enabled custom agent, it has a runtime-enforced per-turn collection gate: the model cannot finish a response until a retrieval/knowledge/MCP tool or a skill Python script returns usable information. It then answers directly in the user's requested chat format rather than generating an artifact. The agent may create or edit only temporary `.py` files below `/mnt/user-data/workspace/` and may execute `.py` skill scripts through `bash`; writes to `outputs`, Markdown/document creation, `present_files`, general shell commands, and skill mutation are denied. Administrators can edit its model and skill allowlist from agent management; the runtime safety policy remains fixed. Regression coverage: `tests/test_forced_research_agent.py`.
|
||
|
||
### AI Writing Intranet Retrieval
|
||
|
||
AI 写作的「知识库检索」会直连 `ai_writing.intranet_search_url` 内网 ES/知识库接口,并默认跳过代理环境变量和 HTTPS 证书校验,适配自签名或内部 CA 证书的内网服务。需要强校验时可在 `config.yaml` 里设置 `ai_writing.intranet_verify_ssl: true`。通用检索里的非内置技能在 `ai_writing.researcher_skill_agent: true` 时会用真实 agent 运行时执行技能:以 `ai-writing-researcher` 身份读取该技能的 `SKILL.md`,并按技能要求通过 `bash` 运行脚本,弱模型也会在本轮任务里收到明确的读文件和脚本执行指令。
|
||
|
||
### Installable Sentiment Analysis Skill
|
||
|
||
skill-packages/sentiment-analysis.skill is a separately installable skill package. It keeps the external AG-UI gateway address in config/sentiment-agent.json, resolves Basic Auth from deployment environment variables, and normalizes the upstream SSE stream into NDJSON (thinking_delta / answer_delta / done) for callers that need incremental rendering. It does not alter the existing frontend sentiment-analysis page. See skill-packages/sentiment-analysis/docs/调用与安装说明.md for installation and invocation.
|
||
|
||
### IM Channels
|
||
|
||
The IM bridge supports Feishu, Slack, and Telegram. Slack and Telegram still use the final `runs.wait()` response path, while Feishu now streams through `runs.stream(["messages-tuple", "values"])` and updates a single in-thread card in place.
|
||
|
||
For Feishu card updates, DeerFlow stores the running card's `message_id` per inbound message and patches that same card until the run finishes, preserving the existing `OK` / `DONE` reaction flow.
|
||
|
||
---
|
||
|
||
## Quick Start
|
||
|
||
### Prerequisites
|
||
|
||
- Python 3.12+
|
||
- [uv](https://docs.astral.sh/uv/) package manager
|
||
- API keys for your chosen LLM provider
|
||
|
||
### Installation
|
||
|
||
```bash
|
||
cd deer-flow
|
||
|
||
# Copy configuration files
|
||
cp config.example.yaml config.yaml
|
||
|
||
# Install backend dependencies
|
||
cd backend
|
||
make install
|
||
```
|
||
|
||
### Configuration
|
||
|
||
Edit `config.yaml` in the project root:
|
||
|
||
```yaml
|
||
models:
|
||
- name: gpt-4o
|
||
display_name: GPT-4o
|
||
use: langchain_openai:ChatOpenAI
|
||
model: gpt-4o
|
||
api_key: $OPENAI_API_KEY
|
||
supports_thinking: false
|
||
supports_vision: true
|
||
|
||
- name: gpt-5-responses
|
||
display_name: GPT-5 (Responses API)
|
||
use: langchain_openai:ChatOpenAI
|
||
model: gpt-5
|
||
api_key: $OPENAI_API_KEY
|
||
use_responses_api: true
|
||
output_version: responses/v1
|
||
supports_vision: true
|
||
```
|
||
|
||
Set your API keys:
|
||
|
||
```bash
|
||
export OPENAI_API_KEY="your-api-key-here"
|
||
```
|
||
|
||
For a LiteLLM OpenAI-compatible proxy, keep `supports_thinking: false` unless
|
||
the configured model has an explicit, proxy-supported thinking control. This
|
||
prevents DeerFlow's connection check from sending vendor-specific `thinking`
|
||
fields that a generic OpenAI-compatible gateway may reject.
|
||
|
||
### Running
|
||
|
||
**Full Application** (from project root):
|
||
|
||
```bash
|
||
make dev # Starts LangGraph + Gateway + Frontend + Nginx
|
||
```
|
||
|
||
Access at: http://localhost:2026
|
||
|
||
**Backend Only** (from backend directory):
|
||
|
||
```bash
|
||
# Terminal 1: LangGraph server
|
||
make dev
|
||
|
||
# Terminal 2: Gateway API
|
||
make gateway
|
||
```
|
||
|
||
Direct access: LangGraph at http://localhost:2024, Gateway at http://localhost:8001
|
||
|
||
---
|
||
|
||
## Project Structure
|
||
|
||
```
|
||
backend/
|
||
├── src/
|
||
│ ├── agents/ # Agent system
|
||
│ │ ├── lead_agent/ # Main agent (factory, prompts)
|
||
│ │ ├── middlewares/ # 9 middleware components
|
||
│ │ ├── memory/ # Memory extraction & storage
|
||
│ │ └── thread_state.py # ThreadState schema
|
||
│ ├── gateway/ # FastAPI Gateway API
|
||
│ │ ├── app.py # Application setup
|
||
│ │ └── routers/ # 6 route modules
|
||
│ ├── sandbox/ # Sandbox execution
|
||
│ │ ├── local/ # Local filesystem provider
|
||
│ │ ├── sandbox.py # Abstract interface
|
||
│ │ ├── tools.py # bash, ls, read/write/str_replace
|
||
│ │ └── middleware.py # Sandbox lifecycle
|
||
│ ├── subagents/ # Subagent delegation
|
||
│ │ ├── builtins/ # general-purpose, bash agents
|
||
│ │ ├── executor.py # Background execution engine
|
||
│ │ └── registry.py # Agent registry
|
||
│ ├── tools/builtins/ # Built-in tools
|
||
│ ├── mcp/ # MCP protocol integration
|
||
│ ├── models/ # Model factory
|
||
│ ├── skills/ # Skill discovery & loading
|
||
│ ├── config/ # Configuration system
|
||
│ ├── community/ # Community tools & providers
|
||
│ ├── reflection/ # Dynamic module loading
|
||
│ └── utils/ # Utilities
|
||
├── docs/ # Documentation
|
||
├── tests/ # Test suite
|
||
├── langgraph.json # LangGraph server configuration
|
||
├── pyproject.toml # Python dependencies
|
||
├── Makefile # Development commands
|
||
└── Dockerfile # Container build
|
||
```
|
||
|
||
---
|
||
|
||
## Configuration
|
||
|
||
### Main Configuration (`config.yaml`)
|
||
|
||
Place in project root. Config values starting with `$` resolve as environment variables.
|
||
|
||
Key sections:
|
||
- `models` - LLM configurations with class paths, API keys, thinking/vision flags
|
||
- `tools` - Tool definitions with module paths and groups
|
||
- `tool_groups` - Logical tool groupings
|
||
- `sandbox` - Execution environment provider
|
||
- `skills` - Skills directory paths plus optional prompt routing; `skills.es_query_routing`
|
||
can inject an editable system-prompt rule that sends flexible Elasticsearch Query DSL
|
||
requests to the `es_query` skill, and can be disabled without changing code. Each ordinary
|
||
Q&A agent build logs `ES query routing prompt registration` with `registered`, eligibility,
|
||
character count, and a short SHA-256 fingerprint so operators can verify the exact configured
|
||
block reached the final model system prompt without logging its contents
|
||
- `title` - Auto-title generation settings
|
||
- `summarization` - Context summarization settings
|
||
- `subagents` - Subagent system (enabled/disabled)
|
||
- `memory` - Memory system settings (enabled, storage, debounce, facts limits)
|
||
|
||
Provider note:
|
||
- `models[*].use` references provider classes by module path (for example `langchain_openai:ChatOpenAI`).
|
||
- If a provider module is missing, DeerFlow now returns an actionable error with install guidance (for example `uv add langchain-google-genai`).
|
||
|
||
### Browser CORS
|
||
|
||
The Gateway permits the local Vite development origins on ports 5173, 5174,
|
||
3000, and 8080 for both `localhost` and `127.0.0.1` with credentials, even
|
||
when `GATEWAY_CORS_ORIGINS` is set. For a separate deployed frontend, set
|
||
`GATEWAY_CORS_ORIGINS` to its exact origin (including scheme and port), for
|
||
example `http://47.88.25.99:7010`.
|
||
|
||
### Extensions Configuration (`extensions_config.json`)
|
||
|
||
MCP servers and skill states in a single file:
|
||
|
||
```json
|
||
{
|
||
"mcpServers": {
|
||
"github": {
|
||
"enabled": true,
|
||
"type": "stdio",
|
||
"command": "npx",
|
||
"args": ["-y", "@modelcontextprotocol/server-github"],
|
||
"env": {"GITHUB_TOKEN": "$GITHUB_TOKEN"}
|
||
},
|
||
"secure-http": {
|
||
"enabled": true,
|
||
"type": "http",
|
||
"url": "https://api.example.com/mcp",
|
||
"oauth": {
|
||
"enabled": true,
|
||
"token_url": "https://auth.example.com/oauth/token",
|
||
"grant_type": "client_credentials",
|
||
"client_id": "$MCP_OAUTH_CLIENT_ID",
|
||
"client_secret": "$MCP_OAUTH_CLIENT_SECRET"
|
||
}
|
||
}
|
||
},
|
||
"skills": {
|
||
"pdf-processing": {"enabled": true}
|
||
}
|
||
}
|
||
```
|
||
|
||
### Environment Variables
|
||
|
||
- `DEER_FLOW_CONFIG_PATH` - Override config.yaml location
|
||
- `DEER_FLOW_EXTENSIONS_CONFIG_PATH` - Override extensions_config.json location
|
||
- Model API keys: `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `DEEPSEEK_API_KEY`, etc.
|
||
- Tool API keys: `TAVILY_API_KEY`, `GITHUB_TOKEN`, etc.
|
||
|
||
### LangSmith Tracing
|
||
|
||
DeerFlow has built-in [LangSmith](https://smith.langchain.com) integration for observability. When enabled, all LLM calls, agent runs, tool executions, and middleware processing are traced and visible in the LangSmith dashboard.
|
||
|
||
**Setup:**
|
||
|
||
1. Sign up at [smith.langchain.com](https://smith.langchain.com) and create a project.
|
||
2. Add the following to your `.env` file in the project root:
|
||
|
||
```bash
|
||
LANGSMITH_TRACING=true
|
||
LANGSMITH_ENDPOINT=https://api.smith.langchain.com
|
||
LANGSMITH_API_KEY=lsv2_pt_xxxxxxxxxxxxxxxx
|
||
LANGSMITH_PROJECT=xxx
|
||
```
|
||
|
||
**Legacy variables:** The `LANGCHAIN_TRACING_V2`, `LANGCHAIN_API_KEY`, `LANGCHAIN_PROJECT`, and `LANGCHAIN_ENDPOINT` variables are also supported for backward compatibility. `LANGSMITH_*` variables take precedence when both are set.
|
||
|
||
### Langfuse Tracing
|
||
|
||
DeerFlow also supports [Langfuse](https://langfuse.com) observability for LangChain-compatible runs.
|
||
|
||
Add the following to your `.env` file:
|
||
|
||
```bash
|
||
LANGFUSE_TRACING=true
|
||
LANGFUSE_PUBLIC_KEY=pk-lf-xxxxxxxxxxxxxxxx
|
||
LANGFUSE_SECRET_KEY=sk-lf-xxxxxxxxxxxxxxxx
|
||
LANGFUSE_BASE_URL=https://cloud.langfuse.com
|
||
```
|
||
|
||
If you are using a self-hosted Langfuse deployment, set `LANGFUSE_BASE_URL` to your Langfuse host.
|
||
|
||
### Dual Provider Behavior
|
||
|
||
If both LangSmith and Langfuse are enabled, DeerFlow initializes and attaches both callbacks so the same run data is reported to both systems.
|
||
|
||
If a provider is explicitly enabled but required credentials are missing, or the provider callback cannot be initialized, DeerFlow raises an error when tracing is initialized during model creation instead of silently disabling tracing.
|
||
|
||
**Docker:** In `docker-compose.yaml`, tracing is disabled by default (`LANGSMITH_TRACING=false`). Set `LANGSMITH_TRACING=true` and/or `LANGFUSE_TRACING=true` in your `.env`, together with the required credentials, to enable tracing in containerized deployments.
|
||
|
||
---
|
||
|
||
## Development
|
||
|
||
### Commands
|
||
|
||
```bash
|
||
make install # Install dependencies
|
||
make dev # Run LangGraph server (port 2024)
|
||
make gateway # Run Gateway API (port 8001)
|
||
make lint # Run linter (ruff)
|
||
make format # Format code (ruff)
|
||
```
|
||
|
||
### Code Style
|
||
|
||
- **Linter/Formatter**: `ruff`
|
||
- **Line length**: 240 characters
|
||
- **Python**: 3.12+ with type hints
|
||
- **Quotes**: Double quotes
|
||
- **Indentation**: 4 spaces
|
||
|
||
### Testing
|
||
|
||
```bash
|
||
uv run pytest
|
||
```
|
||
|
||
### Runtime appearance themes
|
||
|
||
The admin appearance settings are persisted in `.deer-flow/system_settings.json`.
|
||
Besides the built-in themes, an administrator can create up to 12 runtime
|
||
themes through `PUT /api/system-settings/appearance`. Each theme stores a
|
||
safe six-digit `primary` colour, an optional primary-button `text_color`,
|
||
`light`/`dark` mode, name, enabled flag and sort order. The frontend derives
|
||
the rest of the visual tokens; the optional `overrides` object only exposes
|
||
ten interaction colours (`button_hover`, `navigation_gradient_start`,
|
||
`navigation_gradient_end`, `selected_background`, `border`, `muted`,
|
||
`muted_text`, `destructive`, `brand_title`, `primary_foreground`) for
|
||
exceptional brand requirements. The two navigation values override the
|
||
automatically derived top-bar gradient. A custom theme id
|
||
must use the `custom-theme-` prefix. Disabled or removed custom themes cannot
|
||
remain the system default and automatically fall back to `light`.
|
||
|
||
Compact top-bar shortcuts use the same appearance endpoint. In addition to
|
||
login and theme query parameters, each shortcut may include ordered keyless
|
||
`path` parameters. The frontend resolves each selected login field and appends
|
||
it as an encoded path segment (including inside a `#/…` hash route); password
|
||
remains an ordinary login query parameter with the key `password`.
|
||
|
||
### Citation display settings
|
||
|
||
Reference-mode entry points are controlled by `citation_display` in
|
||
`.deer-flow/system_settings.json`. `reference_mode_enabled` defaults to `false`;
|
||
when disabled, the frontend hides the "参考文献" composer button in both the
|
||
normal workspace and iframe chat, and hides the per-user RAG citation toggle
|
||
from appearance settings. Conversations with selected LLMWiki knowledge bases
|
||
still show their source documents in the reference panel even when the manual
|
||
entry is hidden. `GET /api/system-settings/citation-display` is
|
||
readable by logged-in users so the UI can decide whether to expose the entry;
|
||
`PUT /api/system-settings/citation-display` is admin-only and supports partial
|
||
updates for both `reference_mode_enabled` and `cleanup_words`.
|
||
|
||
### Ordinary Q&A Markdown format switch
|
||
|
||
Administrators can opt into the formal long-answer Markdown prompt from the
|
||
admin "问答提示词配置" panel. The value is persisted as
|
||
`prompt_prefix.ordinary_qa_markdown_format_enabled` in
|
||
`.deer-flow/system_settings.json` and defaults to `false`. When enabled, only
|
||
ordinary Q&A receives the `# 主标题` / `## 一、标题` / `### (一) 标题` /
|
||
`#### 1. 标题` rules; writing mode, notebook mode and scheduled runs remain
|
||
unaffected. The change takes effect when the next prompt is built.
|
||
|
||
---
|
||
|
||
## Technology Stack
|
||
|
||
- **LangGraph** (1.0.6+) - Agent framework and multi-agent orchestration
|
||
- **LangChain** (1.2.3+) - LLM abstractions and tool system
|
||
- **FastAPI** (0.115.0+) - Gateway REST API
|
||
- **langchain-mcp-adapters** - Model Context Protocol support
|
||
- **agent-sandbox** - Sandboxed code execution
|
||
- **markitdown** - Multi-format document conversion
|
||
- **tavily-python** / **firecrawl-py** - Web search and scraping
|
||
|
||
---
|
||
|
||
## Documentation
|
||
|
||
- [Configuration Guide](docs/CONFIGURATION.md)
|
||
- [Architecture Details](docs/ARCHITECTURE.md)
|
||
- [API Reference](docs/API.md)
|
||
- [Offline Docker wheelhouse deployment (ZH)](../README.md)
|
||
- [File Upload](docs/FILE_UPLOAD.md)
|
||
- [Path Examples](docs/PATH_EXAMPLES.md)
|
||
- [Context Summarization](docs/summarization.md)
|
||
- [Plan Mode](docs/plan_mode_usage.md)
|
||
- [Backend Run/Stop Guide (ZH)](docs/BACKEND_RUN_STOP_ZH.md)
|
||
- [Setup Guide](docs/SETUP.md)
|
||
|
||
---
|
||
|
||
## License
|
||
|
||
See the [LICENSE](../LICENSE) file in the project root.
|
||
|
||
## Contributing
|
||
|
||
See [CONTRIBUTING.md](CONTRIBUTING.md) for contribution guidelines.
|
||
|
||
## Position roundtable durable history
|
||
|
||
The position-roundtable workflow stores database-owned recovery snapshots for
|
||
Step-1 intent chat, every position-node chat, the summary/action-plan chats,
|
||
unsent composer drafts, and generated text deliverables. LangGraph checkpoints
|
||
remain the execution source, but SQL history continues to render after a
|
||
checkpoint or sandbox-output volume is replaced. Apply Alembic revision
|
||
`20260723_05` (or simply upgrade to `head`) before deploying this version.
|
||
|
||
Its live intent conversation still uses the shared `/api/intent/init` and
|
||
`/api/intent/stream` endpoints used by the original multi-agent roundtable.
|
||
Position roundtable identifies itself with `intent_agent_id: "position-roundtable-intent"`;
|
||
this is a separately seeded built-in agent, not an alias of `roundtable-intent`.
|
||
It preserves the same thread, SSE streaming, and file-reading lifecycle as the
|
||
generic Step-1 agent, but has its own **at-most-five user-assistance-round**
|
||
brake. A round is one AI turn plus the user reply, not one tool call: the
|
||
dedicated agent may emit several independent `ask_clarification` calls in the
|
||
same turn, which the frontend renders as several separate selection cards and
|
||
submits together. A topic-only request is not considered resolved: the agent
|
||
asks for every independent missing decision (such as scope, decision use, or
|
||
task constraints) as separate cards. For the dedicated agent, a fresh
|
||
`ask_clarification` emitted in a stream is a strict user-input pause and wins over any same-run premature
|
||
`[INTENT_READY]`; a historical clarification from the prior turn does not
|
||
block the user answer from completing. Its final
|
||
`[INTENT_READY]` JSON contains four string arrays: `coreGoals`, `riskWarnings`,
|
||
`keyPoints`, and `strategicSignificance`; it must contain completed conclusions,
|
||
not “待确认” or other unresolved placeholders. The multi-agent endpoint remains
|
||
backwards compatible when this optional field is omitted and continues to emit
|
||
`objective`, `constraints`, and `assumptions`.
|
||
`/api/intent/init` creates its empty thread directly through the in-process
|
||
thread store/checkpointer, matching multi-agent initialization and avoiding a
|
||
deployment-sensitive `127.0.0.1:<gateway-port>` loopback. If initial checkpoint
|
||
creation fails or is cancelled, bootstrap removes both the partial checkpoint and
|
||
its thread metadata row. Thread metadata creation issues no post-commit ORM
|
||
read-back on any code path: all returned fields are assigned before commit, and
|
||
skipping the extra read keeps a MySQL read/write-splitting endpoint from sending
|
||
it to a lagging replica (which fails the request with
|
||
`Could not refresh instance`). Only the
|
||
recovery snapshot layer is position-specific. Deployments must expose
|
||
either `/api/position-roundtable/*` or the compatibility mirror
|
||
`/api/multi-agent/position-roundtable/*`; updated clients probe both and only
|
||
retry when the first response is the framework/proxy-level `404 Not Found`.
|
||
|
||
## Position-roundtable role directory
|
||
|
||
The position-roundtable work views are persisted separately from organisation
|
||
positions/RBAC in `position_roles`. Gateway startup idempotently inserts the
|
||
four built-in ids (testing analysis, organisation test, solution design and
|
||
comprehensive test), but only when an id is missing: administrator changes to
|
||
the name, order, enabled state or role type are never overwritten.
|
||
|
||
`GET /api/position-roles` is consumed by the workbench and `PUT
|
||
/api/position-roles` is administrator-only. Exactly one enabled `intent` role
|
||
is required. It is the initial intent-identification workspace and owns the
|
||
summary/action-plan closing agents. Built-in rows retain stable ids so
|
||
historical chain nodes remain readable; they can be edited or disabled but not
|
||
deleted. Custom roles may be freely added or removed.
|
||
|
||
## Roundtable artifact submission status
|
||
|
||
The task-scoped roundtable workbench exposes a lightweight submission-status
|
||
API for generated deliverables:
|
||
|
||
- `GET /api/roundtable/tasks/{taskId}/artifacts?submitted=true` returns the
|
||
submitted artifact references for a task (or `items: []`).
|
||
- `PUT /api/roundtable/tasks/{taskId}/artifacts/submission` records
|
||
`submitted: true` or cancels it with `submitted: false`.
|
||
|
||
Both routes use the normal Gateway `Authorization: Bearer <token>` middleware.
|
||
The record is shared by `taskId`; the authenticated user is retained only in
|
||
audit columns. Its server-owned idempotency key is
|
||
`taskId + sourceKey + path + revision`. Repeating the same submitted payload
|
||
returns the existing record; cancellation updates its state and timestamp
|
||
without deleting the audit row. DeerFlow stores only the artifact reference
|
||
(`threadId` + sandbox `path`, plus the optional session/node/name/MIME fields),
|
||
never the Markdown body and never a task-system delivery. Continue to preview
|
||
the file through the existing thread-artifact endpoint.
|
||
|
||
## 普通知识库管理
|
||
|
||
在「普通知识库」的「我的知识库」或「公共 LLMWiki」中,知识库所有者和管理员可以点击卡片上的「编辑」,修改名称和描述(描述可清空)。保存后同步更新知识服务和本地记录,并刷新列表;系统「对话沉淀」库不提供编辑入口。
|
||
|
||
远端已经删除的知识库会在下次列表刷新时自动隐藏,包括个人、公共列表和知识库选择器;页面每 15 秒自动刷新,也可以手动点击「刷新」。远端服务请求失败时展示错误,不把连接故障当成删除,不自动删除本地映射或已有引用。
|
||
|
||
## DeerFlow-local LLMWiki Wiki vector index
|
||
|
||
LLMWiki can mirror processed WeKnora Wiki pages into DeerFlow and search only
|
||
processed Wiki content. A knowledge base uses exact NumPy cosine similarity
|
||
only after its whole Wiki has completed one coherent vectorization pass; until
|
||
then internal search calls WeKnora's Wiki-page search/list API and reads the
|
||
matched Wiki pages. It never falls back to raw documents or chunks. Configure
|
||
`llmwiki.local_wiki_index.embedding`, then enable the local index in
|
||
`config.yaml`. `auto_sync` defaults to true: the scheduler performs one
|
||
immediate incremental scan at startup and then scans every configured interval
|
||
(30 seconds by default). Wiki changes are content-hash checked, so unchanged
|
||
pages reuse their vectors instead of being encoded again.
|
||
`database_batch_size` caps each vector insert statement (default `50`) while
|
||
the page replacement remains atomic, preventing a large Wiki from flooding
|
||
the SQL connection in one statement.
|
||
The optional `embedding.dimensions` value is used only to validate the returned
|
||
vector size and build the index fingerprint; it is deliberately not sent in
|
||
the embedding request because fixed-dimension BGE/OpenAI-compatible endpoints
|
||
may reject that extra field with HTTP 400. Omit it to discover the dimension
|
||
from the first successful response.
|
||
|
||
Embedding failures log the upstream HTTP status and a bounded response body.
|
||
The client never sends a `dimensions` request parameter.
|
||
|
||
The migration creates `llmwiki_wiki_pages`, `llmwiki_wiki_vectors`, and
|
||
`llmwiki_wiki_sync_states`. Wiki Markdown and normalized float32 vectors are
|
||
stored in DeerFlow's configured SQL database. Multi-worker sync is protected by
|
||
database leases; searches use revision-checked per-knowledge-base LRU snapshots
|
||
after the full-library readiness gate succeeds. Knowledge-base owners can click
|
||
**手动向量化** in Wiki management (`POST
|
||
/api/llmwiki/knowledge-bases/{id}/vectorize`) and follow its status at
|
||
`GET /api/llmwiki/knowledge-bases/{id}/wiki-index-status`. The response includes
|
||
aggregate `progress` plus a `pages` list with every Wiki page's index state,
|
||
vector-section count, timestamps, and last error. The management page renders
|
||
the percentage and labels each article as 已向量化 / 待向量化 / 失败, so it is
|
||
clear exactly what is searchable. Existing chunk citations remain readable
|
||
while all new retrieval citations open Wiki pages.
|
||
|
||
The frontend labels these results simply as **Wiki** references; provider-brand
|
||
wording is intentionally hidden from user-facing labels and status text. Clicking any
|
||
part of a Wiki reference card opens one right-side drawer (rather than a second
|
||
detail dialog), renders the complete page as Markdown, and reads the remote
|
||
WeKnora Wiki page directly when the local index is disabled. Even if a Wiki
|
||
citation carries a legacy document id, the drawer does not mix in the uploaded
|
||
original-file preview. A single page can be downloaded as Markdown or Word from
|
||
the drawer. **Open in Wiki** navigates to the exact DeerFlow knowledge-base
|
||
mapping and Wiki slug. Wiki search-result cards use the same drawer interaction.
|
||
Exact Wiki links now open a lightweight native reader at
|
||
`knowledge/{mapping_id}/wiki/view?wikiSlug=...`; it fetches only the selected
|
||
page and never creates a WeKnora iframe session. The former
|
||
`knowledge/{mapping_id}?tab=wiki&wikiSlug=...` links redirect to the reader
|
||
before waiting for the LLMWiki runtime check. Drawer navigation stays in the
|
||
same SPA and passes the already-loaded page as route state for immediate paint.
|
||
The native reader also keeps a searchable, collapsible Wiki directory on the
|
||
left. It reconstructs hierarchy from `parent_slug`, `wiki_path`,
|
||
`category_path`, and the slug fallback, highlights and expands the current
|
||
page, and switches directory entries through the same lightweight reader route
|
||
without mounting the WeKnora iframe.
|
||
|
||
All generated SPA deep links preserve the deployment document path (for
|
||
example `/web/`) before the `#` route, including Q&A references and Wiki index
|
||
links. Protected Wiki images no longer render from temporary browser `blob:`
|
||
URLs. Public images use the published file proxy directly; private images first
|
||
obtain a 24-hour signed capability and then load through
|
||
`GET /api/public/llmwiki/files/{token}` with a real HTTP URL and inline
|
||
filename, allowing zoom, new-tab viewing, and browser Save As without exposing
|
||
the user's bearer token. The embedded WeKnora detail page follows the same
|
||
rule: late `img[data-protected-src]` placeholders keep their scoped HTTP proxy
|
||
URL on the DOM node, so WeKnora's click-to-enlarge viewer never receives an
|
||
already-revoked `blob:` URL.
|
||
|
||
## Leaderboard query and schema startup safety
|
||
|
||
The leaderboard excludes scheduled runs with indexed correlated `NOT EXISTS`
|
||
checks. Tool metrics no longer outer-join the full `runs` table merely to
|
||
recognize legacy scheduler messages. New databases receive dimension-aware
|
||
analytics indexes from ORM metadata; existing large MySQL deployments should
|
||
run `scripts/sql/leaderboard_indexes_mysql.sql` during a low-traffic window so
|
||
index creation never delays application startup.
|
||
|
||
Startup Alembic migrations now have one unambiguous head. The historical
|
||
duplicate `20260821_01` identifier is repaired by unique fixed-question
|
||
revisions and the `20260825_01` merge revision, which also fills the
|
||
report-structure column for databases stamped by either old variant.
|
||
|
||
Published knowledge bases also have a login-free, chrome-free reader at
|
||
`/#/embed/knowledge/{mapping_id}` for direct third-party iframe use. With no
|
||
`wikiSlug`, it selects the Wiki **索引** page first (recognized by title, page
|
||
type, or `index`/`home`/`readme` slug) and pins that page to the top of the
|
||
public directory. If WeKnora has pages but no explicit index page, the reader
|
||
creates a virtual `__index__` landing page from the public directory; empty
|
||
dynamic-index groups also fall back to those page summaries;
|
||
`?wikiSlug=...` deep-links one page. The reader supports
|
||
`theme=light|dark`, `bg`, and `text` embed appearance parameters. Index-article
|
||
catalogue sections (page-type counts, nested categories, page summaries) are
|
||
read from the published-only `/wiki/index` proxy because WeKnora renders that
|
||
part dynamically rather than storing it in the index page body; group and item
|
||
ordering therefore stays identical to WeKnora. Links generated as
|
||
knowledge-base-shell placeholders are matched
|
||
against the public page title, slug, or alias and rewritten to exact embedded
|
||
`wikiSlug` links. They retain WeKnora's brand-color medium text and dashed
|
||
underline (solid on hover), and switch articles inside the current reader
|
||
without a full-page reload. Data comes
|
||
only from the read-only `/api/public/llmwiki/knowledge-bases/{id}/...` namespace,
|
||
which rechecks `publication_status=published` on every request and denies the
|
||
system conversation-deposit knowledge base. Private or unpublished bases are
|
||
never exposed by this route. Anonymous consumers can discover the available
|
||
public Wiki libraries through `GET /api/public/llmwiki/knowledge-bases`. The
|
||
endpoint validates each published mapping against WeKnora, omits stale remote
|
||
mappings, and returns safe comprehensive metadata: description/type,
|
||
document/chunk/processing/share counts, non-secret configuration and
|
||
timestamps, Wiki page/type/status counts plus explicit/generated index mode and
|
||
default embed route, and aggregate local-vector state/counts/timestamps. It
|
||
never returns the WeKnora id, tenant/creator identity, storage credentials, API
|
||
credentials, raw vectors, or per-page indexing errors.
|
||
Protected `resource://` and supported object-storage images are fetched through
|
||
the current mapping instead of being exposed as upstream URLs. Authenticated
|
||
readers use `GET /api/llmwiki/knowledge-bases/{id}/files?file_path=...`; public
|
||
readers use the corresponding `/api/public/llmwiki/...` route, which requires a
|
||
published mapping. The embedded WeKnora guard applies the same mapping-scoped
|
||
proxy to document previews and Markdown images, including images inserted after
|
||
the drawer/page first renders.
|
||
|
||
Wiki management also provides **导出 Excel**. `GET
|
||
/api/llmwiki/knowledge-bases/{id}/wiki/export.xlsx` exports every processed Wiki
|
||
page, including its full directory path, current directory, parent-directory
|
||
path, one column per directory level, metadata, and complete Markdown content
|
||
(split across continuation columns when Excel's per-cell limit is reached).
|
||
|
||
Authenticated operations are available below `/api/llmwiki/wiki-index` and
|
||
`POST /api/llmwiki/wiki-vector-search`. The service-to-service endpoints below
|
||
`/api/external/llmwiki/wiki` require `X-API-Key`, are rate-limited, and only
|
||
expose mappings that are both internally published and explicitly marked
|
||
`external_search_enabled`; the system `对话沉淀` mapping is always denied.
|
||
|
||
A separate browser-friendly endpoint, `POST /api/knowledge/vector-search`, is
|
||
fully anonymous: it requires neither a DeerFlow token nor `X-API-Key`, has no
|
||
rate limit, and responds with `Access-Control-Allow-Origin: *`. It searches the
|
||
local vectors of every published knowledge base regardless of
|
||
`external_search_enabled`, while still excluding the system `对话沉淀` base.
|
||
The minimal request is `{"query":"如何部署"}`; optional fields are `top_k`
|
||
(default `10`, maximum `100`) and `include_page_content` (default `false`). Each
|
||
matched vector section is returned as its own ranked result, including the
|
||
matched text, score, DeerFlow and WeKnora knowledge/Wiki ids, page metadata,
|
||
raw source/chunk references, best-effort document/chunk ids, and a directly
|
||
usable `wiki_url`. Set `llmwiki.local_wiki_index.frontend_base_url` to the
|
||
deployed frontend origin/base path; its development default is
|
||
`http://127.0.0.1:5174`.
|
||
|
||
See `docs/LLMWIKI_DEERFLOW_LOCAL_VECTOR_INDEX_ZH.md` for configuration,
|
||
security boundaries, rollout, and API contracts.
|
||
|
||
## Document rewrite fallback
|
||
|
||
Whole-document rewrite commits the artifact/report before it attempts to store
|
||
the optional reversible before-image. If that version-snapshot write fails, the
|
||
rewritten content remains available and Gateway logs the exception; it never
|
||
restores the original solely because undo history is unavailable. The durable
|
||
result marks `versionSnapshotFailed`, allowing the sandbox to enable its normal
|
||
manual-save action without showing the user an error notification.
|
||
|
||
The optional rewrite-history hydration endpoints scope results to the
|
||
authenticated job owner. They deliberately return an empty list rather than
|
||
performing a second thread-owner check, so a legacy guest/SSO display alias
|
||
cannot make an otherwise usable sandbox report fail with a history-only 404.
|
||
|
||
Artifact and rewrite path resolution uses the persisted thread creator's
|
||
directory when metadata exists, after the route's normal ownership check. This
|
||
prevents a missing request context from incorrectly looking under
|
||
`users/default` and returning 404 for an existing sandbox file.
|
||
|
||
The final rewrite commit uses a deliberately short sibling temporary filename,
|
||
which keeps long Windows artifact paths below the path-length limit. It also
|
||
recreates a missing `outputs` parent before retrying, so transient
|
||
sandbox-directory recreation cannot discard a completed rewrite.
|
||
|
||
## Offline local-sandbox Office runtime
|
||
|
||
The pure DeerFlow offline image treats the backend container as the default
|
||
`LocalSandboxProvider` runtime. It bundles MarkItDown, python-docx, openpyxl,
|
||
python-pptx, LibreOffice Headless, Pandoc, Poppler, and CJK fonts. Modern OOXML
|
||
uploads are parsed directly; legacy DOC/XLS/PPT and ODT/ODS/ODP files are
|
||
normalized through an isolated LibreOffice profile before Markdown extraction.
|
||
Run `./OFFLINE_INSTALL.sh --office-check` from the offline bundle to generate
|
||
and re-parse Word, Excel, and PowerPoint samples without requiring a database.
|
||
|
||
## Skill knowledge distillation and assistant knowledge
|
||
|
||
The Gateway now exposes an administrator-only skill distillation pipeline at
|
||
`/api/skill-knowledge`. Each request can send up to 200 skills to one writable
|
||
WeKnora Wiki knowledge base; multi-base distillation is done by repeated
|
||
operations so the local binding remains `skill_name + target_id`. The scanner
|
||
never reads or uploads code/script files, strips Markdown code fences, blocks
|
||
credential-like files and content, safely inspects ZIP members, and converts
|
||
supported Office/PDF/spreadsheet files off the event loop. Jobs use database
|
||
claims and renewable leases, stage snapshots before activation, retain the
|
||
prior active version on failure, support per-item retry, and keep per-binding
|
||
contribution records for safe detach/delete. Direct assistant-base targets,
|
||
non-Wiki/FAQ/conversation/readonly targets, and direct Neo4j writes are rejected
|
||
for this product version.
|
||
|
||
`/api/assistant-knowledge` manages DeerFlow-native assistant knowledge bases.
|
||
There is one global full-library base named `知识梳理总库`, plus admin-created
|
||
custom assistant bases. Reads and search are available to authenticated users;
|
||
global initialization, custom-base creation, WeKnora export/import, source
|
||
reimport, and source deletion are admin-only. Assistant bases only accept
|
||
WeKnora export packages or imports triggered from a WeKnora mapping; direct
|
||
skill import and arbitrary local-file import return 410. Importing into a custom
|
||
assistant base also mirrors the same package into the global base, where
|
||
canonical entities and relations are aligned across all imported sources.
|
||
The database schema is introduced by Alembic revision `20260901_01`.
|
||
|
||
### Wiki 流转与本地小批量编码
|
||
|
||
所有登录用户都可在「技能 → 归纳到知识库」中多选自己可见的技能,并归纳到一个
|
||
自己可写的 WeKnora Wiki 知识库;用户只能看到和维护自己发起的归纳记录,管理员
|
||
接口仍可查看全量记录。技能归纳上传清理后的知识正文,等待 WeKnora 完成 Wiki 分析,再建立本地
|
||
Wiki 向量并通过落盘导出包自动沉淀到「知识梳理」。内置技能没有数据库所有权记录
|
||
也可以归纳。若解析结束却没有 Wiki,会提示检查生成模型和资料;文档解析完成
|
||
不等于 Wiki 生成成功,例如生成服务余额不足仍可能表现为文档已完成。
|
||
|
||
管理员可以在普通知识库直接生成向量数据包,每次操作有独立后台记录和下载入口。
|
||
助手知识库支持上传包、异步导入和导出;新建空的普通 Wiki 库可上传助手导出包。管理员可编辑助手知识库名称与描述,也可删除自建助手知识库及其 Wiki、向量、实体、关系、导入记录和落盘数据包;承担全库沉淀的「知识梳理总库」允许编辑但禁止删除。
|
||
包保留正文、别名/标签、目录分类、双向链接、实体关系及 float32 向量。
|
||
回流时通过 WeKnora API 重建目录/页面,向量进入 DeerFlow 本地 Wiki 索引,
|
||
**不直接写 WeKnora 的底层向量或图数据库**。普通与助手 Wiki 使用同一个阅读组件。
|
||
|
||
小批量 CPU 编码可配置如下,无需额外部署 HTTP 编码服务:
|
||
|
||
```yaml
|
||
llmwiki:
|
||
local_wiki_index:
|
||
enabled: true
|
||
auto_sync: true
|
||
sync_interval_seconds: 30
|
||
embedding:
|
||
provider: local
|
||
model: bge-embedding-m3
|
||
dimensions: 1024
|
||
batch_size: 2
|
||
local_threads: 2
|
||
local_process_isolation: true
|
||
# 离线部署时可指定已下载的 BAAI/bge-m3 模型目录:
|
||
# local_model_path: /opt/models/bge-m3
|
||
```
|
||
|
||
本地编码依赖随标准后端安装。将完整 BGE-M3 模型目录提前放到离线服务器,并通过
|
||
`local_model_path` 或 `DEERFLOW_BGE_M3_MODEL_PATH` 指向它。管理员也可在
|
||
「设置和更多 → 设置 → 本地编码器」中选择服务器目录,或从浏览器上传完整模型目录;
|
||
页面配置持久化到运行目录的 `system_settings.json`,上传文件落到
|
||
`.deer-flow/models/bge-m3`,均不进入源码仓库。系统强制 `local_files_only`,
|
||
**不会访问互联网或下载模型**。管理员可在同一设置页点击「一键启动编码器」异步加载;该操作不挂在应用启动流程,加载失败不会
|
||
中断 DeerFlow 主服务。默认由独立低优先级子进程承载模型内存,正常停止 DeerFlow
|
||
时会一并回收该子进程;编码进程异常退出只会将当前向量任务标记失败。模型目录应包含
|
||
`onnx/model.onnx`(官方权重约 2.3 GB)及对应 tokenizer/config 文件。权重不进入业务源码仓库。
|
||
浏览器目录上传采用磁盘暂存和 4 MB 分块复制,页面显示上传进度;经 Nginx 等反向
|
||
代理部署时仍需把请求体上限和读写超时配置到可承载完整模型目录。
|
||
相同模型指纹、相同内容的页面复用现有向量;短正文/摘要至少保留一个非空分段。
|
||
数据包导入不重新编码,模型指纹或维度不一致时明确拒绝,不能只按维度混用模型。
|
||
问答编码查询本身仍需可用的同一模型。WeKnora 的原始文档向量不能当作生成后
|
||
Wiki 正文的向量使用,两种文本需要分别建立各自的首次索引。
|
||
|
||
Wiki 页面新增、编辑、删除及文档上传/重解析都会立即触发一次增量检查;WeKnora
|
||
后台稍后才完成的 Wiki 抽取由 30 秒周期补扫接管。每次成功向量化后自动把该普通
|
||
知识库的最新完整快照同步到「知识梳理总库」,包括服务重启前已有库和新建库;来源
|
||
页面被删除时也会在总库中退役,不保留失效 Wiki。
|
||
|
||
「知识梳理总库」不是按来源简单堆叠:导入时会规范化全/半角、空白、标点及常见实体
|
||
类型,并利用实体别名对齐同一人物、组织、产品或概念;不同来源中同标题的实体 Wiki
|
||
只保留一个规范页面,正文按去重后的知识块合并,来源与实体/关系贡献仍分别留痕。
|
||
重复上传同一个 WeKnora 导出包时,以包内稳定知识库 ID 识别同一来源并执行快照替换,
|
||
不会再产生一批新的片段副本。已有历史重复数据会在总库首次读取、检索、导出或下次
|
||
同步时自动归并。助手知识检索排除目录、索引和摘要页;关键词只查标题与知识正文,
|
||
向量检索返回实际命中的正文片段,问答上下文和检索卡片均不再用摘要代替知识。
|
||
|
||
大包采用流式 JSONL 落盘、同盘临时 SQLite 分阶段读取及异步导入,避免一次性将
|
||
整个包加载进内存。部署需预留包、暂存索引和业务数据库空间,并配置代理上传大小
|
||
与超时。当前不宣称断点续传或 GB 级压力验收已完成。
|
||
|
||
## openGauss / GaussDB application database
|
||
|
||
The application ORM can use a PG-compatible openGauss/GaussDB database through
|
||
the official `opengauss-sqlalchemy` async dialect. Install the optional driver
|
||
set before enabling it:
|
||
|
||
```bash
|
||
uv sync --extra gauss
|
||
```
|
||
|
||
Configure the business database independently from the LangGraph checkpointer:
|
||
|
||
```yaml
|
||
database:
|
||
backend: gauss
|
||
gauss_url: $GAUSS_DATABASE_URL
|
||
|
||
checkpointer:
|
||
type: sqlite
|
||
connection_string: .deer-flow/data/deerflow.db
|
||
```
|
||
|
||
`GAUSS_DATABASE_URL` should normally be
|
||
`opengauss+asyncpg://user:password@host:26000/deerflow`. Plain
|
||
`opengauss://`, `gaussdb://`, and PostgreSQL-style URLs are accepted and
|
||
normalized to the openGauss dialect. Do not configure the LangGraph
|
||
PostgreSQL checkpointer against GaussDB: its migrations are PostgreSQL-specific.
|
||
For a production image, build with `--build-arg UV_EXTRAS=gauss`.
|