Nan Gao 1f792d0f4b
feat(extensions): add middleware plugin foundation (#4636)
* feat(extensions): add middleware plugin foundation

* fix(extensions): stop config resolution from masking extension loading

`create_app()` resolved the configured plugin list inside the fail-open
guard around `load_extensions()`. CI has no `config.yaml` (gitignored and
never generated by the workflow), so `get_app_config()` raised
`FileNotFoundError` there and was swallowed as an extension failure --
`load_extensions()` never ran at all, and the four `create_app()` tests in
`test_extension_app_loading.py` passed locally but failed on every runner.

Resolve the plugin list before the guard. Only an absent `config.yaml` is
tolerated, mirroring `_resolve_trace_enabled_for_app_construction()`:
`create_app()` runs at import time, and lifespan still performs strict
config loading before serving. A `config.yaml` that exists but fails to
parse or validate now propagates instead of being reported as an extension
failure -- reporting it as the latter silently dropped a `required: true`
extension rather than failing the boot.

Make the tests config-independent with an autouse `stub_app_config`
fixture, following the existing pattern in `test_gateway_lifespan_shutdown.py`,
and cover both new branches of the config-resolution boundary.

* fix(extensions): bind the run's extension snapshot through subagent delegation

The lead-agent path resolves one immutable loaded-extension snapshot per run
and binds it through task-store allocation and graph construction, but the
subagent path re-read the process-wide singleton at execution time. In
production both are the same object, yet a `set_loaded_extensions()` between
the lead run's start and a subagent's execution (test teardown, a future
hot-reload path) would let one run mix two extension generations — exactly what
the documented invariant exists to prevent.

The graph-build binding is a ContextVar scoped to synchronous construction, so
it has already exited by the time a tool delegates; the snapshot has to travel
through runtime context instead. The run worker publishes it under the
host-internal `EXTENSION_SNAPSHOT_CONTEXT_KEY` (written after the caller merge,
popped when the run has none, so a caller-supplied value is never
authoritative), `task_tool` reads it back through the type-checking
`resolve_run_extensions()`, and `SubagentExecutor` binds it at construction.

Callers outside the Gateway run path — embedded `DeerFlowClient`, standalone
LangGraph Server — install no snapshot and keep the existing
`get_loaded_extensions()` fallback.

* refactor(extensions): defer the ordering table by call, not by a lying tuple

`CORE_ORDERING_CONSTRAINTS` was a `tuple` subclass that overrode only
`__iter__` and resolved into a class-level `_resolved` side channel. A tuple
cannot populate its own storage after construction, so the instance stayed the
empty tuple it was built as: `len()` was 0, `bool()` was False, `in` was always
False, indexing raised, slicing and `reversed()` came back empty, and it
compared unequal to the plain tuples tests substitute for it — all while
iteration yielded the real constraints. Only `assert_ordering` consumed it, and
only by iterating, so the split went unnoticed.

The sibling `_AnchorTable(dict)` uses the same idea soundly because dict is
mutable: `self.update()` fills the real storage, making every inherited
operation correct. That trick does not survive the port to an immutable type.

Replace it with `core_ordering_constraints()`, matching how `stack.py` defers
the same kind of table via `_anchors()`. The deferral is kept — it is about
dependency direction, not just cycles: `extensions/` is the layer the
middleware layer calls into, so a module-scope `agents.middlewares` import here
points the dependency backwards and closes a cycle as soon as any middleware
imports something under `extensions/` at module level. Resolution stays at
`assert_ordering` time, which already runs inside the middleware builder.

Tests pin both halves: the returned value is a plain tuple whose len/bool/
membership/indexing/reversal/equality agree with iteration, and a subprocess
probe asserts importing `extensions.ordering` does not load the middleware
layer while calling the function does.
2026-08-04 22:33:26 +08:00

94 lines
3.4 KiB
TOML

[project]
name = "deerflow-harness"
version = "2.1.0"
description = "DeerFlow agent harness framework"
requires-python = ">=3.12"
dependencies = [
"agent-client-protocol>=0.4.0",
"agent-sandbox>=0.0.30",
"croniter>=6.0.0",
# Exact pin by design (extension-system version contract): the host pins
# the contract version it implements, extensions declare ranges. A range
# here would let pip resolve a newer contract package than this harness
# implements, making newer extensions look supported at runtime.
"deerflow-extension-api==0.1.0",
"dotenv>=0.9.9",
"exa-py>=1.0.0",
"httpx>=0.28.0",
"kubernetes>=30.0.0",
# Lower bound reflects what the lockfile resolves and tests run against
# (langgraph 1.2.9 pulls langchain >=1.3 transitively).
"langchain>=1.3",
"langchain-anthropic>=1.4.1",
"langchain-deepseek>=1.0.1",
"langchain-mcp-adapters>=0.2.2",
"langchain-openai>=1.2.1",
"langfuse>=3.4.1",
"langgraph>=1.2.9,<1.3",
"langgraph-api>=0.8.1",
"langgraph-cli>=0.4.24",
"langgraph-runtime-inmem>=0.28.0",
"markdownify>=1.2.2",
"markitdown[all,xlsx]>=0.0.1a2",
"pydantic>=2.12.5",
"pyyaml>=6.0.3",
"readabilipy>=0.3.0",
"tavily-python>=0.7.17",
"firecrawl-py>=1.15.0",
"tiktoken>=0.8.0",
"ddgs>=9.10.0",
"duckdb>=1.4.4",
"langchain-google-genai>=4.2.1",
"langgraph-checkpoint-sqlite>=3.1.0,<3.2",
"langgraph-sdk>=0.1.51",
"sqlalchemy[asyncio]>=2.0,<3.0",
"aiosqlite>=0.19",
"alembic>=1.13",
"cryptography>=48.0.1",
"e2b-code-interpreter>=2.8.0",
]
[project.scripts]
deerflow = "deerflow.tui.cli:main"
[project.optional-dependencies]
# Terminal workbench (TUI). Kept optional so the core harness install stays lean;
# the `deerflow` console script degrades to headless help when textual is absent.
tui = ["textual>=0.80"]
# GroundRoute needs no extra packages (httpx is already a core dependency). This
# empty extra exists so the documented `uv add 'deerflow-harness[groundroute]'`
# install command resolves cleanly without an "unknown extra" warning.
groundroute = []
ollama = ["langchain-ollama>=0.3.0"]
postgres = [
"asyncpg>=0.29",
"langgraph-checkpoint-postgres>=3.1.0,<3.2",
"psycopg[binary]>=3.3.3",
"psycopg-pool>=3.3.0",
]
# Cross-process SSE stream bridge (stream_bridge.type: redis). Optional so
# single-process / memory-bridge installs do not pull redis. The Docker image
# always installs this extra because Docker defaults to the redis bridge.
redis = ["redis>=5.0.0"]
pymupdf = ["pymupdf4llm>=0.0.17"]
boxlite = ["boxlite>=0.9.7"]
# Tenki cloud sandbox provider (deerflow.community.tenki). Optional so a default
# install stays free of the Tenki SDK; only pulled in when the provider is used.
tenki = ["tenki-sandbox>=0.4.0"]
# Agent observability (Monocle). Optional so a default install stays free of the
# OpenTelemetry stack; only pulled in when MONOCLE_TRACING is used.
monocle = ["monocle_apptrace>=0.8.8"]
# Agentic browser control (browser_navigate/click/type/... tool group). Optional
# so the core harness install stays lean; import is lazy inside the private
# Playwright loop. After install, run `playwright install chromium` once.
browser = ["playwright>=1.40"]
# Optional Chinese tokenization for the FTS5 memory retrieval adapter.
memory-zh = ["jieba>=0.42.1"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel]
packages = ["deerflow"]