deerflow-code/offline-backend-20260512/backend/packages/harness/deerflow/agents/deep_research/UPSTREAM.md
2026-09-07 18:24:55 +08:00

3.7 KiB

GPT Researcher adaptation manifest

This directory implements Deep Research for DeerFlow. It is an algorithmic adaptation of selected GPT Researcher control flows, not an import of the upstream package at runtime. This boundary is intentional: GPT Researcher's LLM providers, search/scrape stack, MCP setup, global configuration and dependency graph must not bypass DeerFlow's model factory, material-security policy, user context or durable-job lifecycle.

Fixed upstream baseline

Field Value
Project assafelovic/gpt-researcher
Local review checkout F:\react01\deeflow-code\gpt-researcher
Commit 5d84d2f5553e70a2765a8ff3a0d2672d60437ce8
License Apache License 2.0 (F:\react01\deeflow-code\gpt-researcher\LICENSE)
Upstream repository https://github.com/assafelovic/gpt-researcher

No upstream executable source file is copied into this directory at present. If a future change copies an upstream file or a non-trivial portion of one, add its copyright/license notice next to that file, record the exact source path and commit below, and keep the Apache-2.0 notice in the distribution.

Adapted control-flow mapping

DeerFlow code Upstream reference Adaptation made
runners/basic.py gpt_researcher/agent.py, backend/report_type/basic_report/basic_report.py Planning → bounded material collection → de-duplication/context compression → cited report is expressed through AdapterBundle.
runners/detailed.py backend/report_type/detailed_report/detailed_report.py Subtopic planning, concurrent sub-research and section assembly are retained; all upstream GPTResearcher instances become bounded DeerFlow adapters.
runners/deep.py gpt_researcher/skills/deep_research.py Breadth/depth branching, learnings and follow-up questions are preserved with an explicit global budget, semaphores and cooperative cancellation.
runners/multi_agent.py multi_agents/agents/{orchestrator,editor,writer,fact_checker}.py Editor/researcher/writer/reviewer roles become a LangGraph 1.x StateGraph, with the review pause persisted as a durable job state.
adapters/context.py, adapters/material.py gpt_researcher/context/compression.py, gpt_researcher/skills/researcher.py Retains context/de-duplication intent while preserving DeerFlow provenance (rec_uuid, source id, source type) and enforcing the existing material-provider boundary.
runners/images.py, adapters/image.py gpt_researcher/skills/image_generator.py Retains report-image placement intent, while configured OpenAI-compatible and DashScope-native adapters both persist protected hidden-thread artifacts.

Required adaptation invariants

  1. Harness code may not import app.*; this is covered by tests/test_deep_research_harness_boundary.py.
  2. Research collection may only use MaterialProvider; upstream retrievers, scrapers and MCP clients are not imported.
  3. Model calls may only use DeerFlowCompletionBackend and create_chat_model(); no upstream provider configuration is accepted.
  4. Every job uses a frozen DeepResearchConfig snapshot, a user-scoped session and a durable event log.
  5. Upstream examples that run network calls, write unrestricted paths, or block for terminal input must be redesigned rather than copied.

Upgrade procedure

  1. Review the candidate upstream commit in the local checkout.
  2. Update the commit and mapping table in this manifest and in docs/gpt-researcher-deerflow-集成开发方案.md.
  3. Re-evaluate dependency, license and security implications before copying any source.
  4. Run the complete deep-research suite and the harness-boundary test.