Nan Gao 1f792d0f4b
feat(extensions): add middleware plugin foundation (#4636)
* feat(extensions): add middleware plugin foundation

* fix(extensions): stop config resolution from masking extension loading

`create_app()` resolved the configured plugin list inside the fail-open
guard around `load_extensions()`. CI has no `config.yaml` (gitignored and
never generated by the workflow), so `get_app_config()` raised
`FileNotFoundError` there and was swallowed as an extension failure --
`load_extensions()` never ran at all, and the four `create_app()` tests in
`test_extension_app_loading.py` passed locally but failed on every runner.

Resolve the plugin list before the guard. Only an absent `config.yaml` is
tolerated, mirroring `_resolve_trace_enabled_for_app_construction()`:
`create_app()` runs at import time, and lifespan still performs strict
config loading before serving. A `config.yaml` that exists but fails to
parse or validate now propagates instead of being reported as an extension
failure -- reporting it as the latter silently dropped a `required: true`
extension rather than failing the boot.

Make the tests config-independent with an autouse `stub_app_config`
fixture, following the existing pattern in `test_gateway_lifespan_shutdown.py`,
and cover both new branches of the config-resolution boundary.

* fix(extensions): bind the run's extension snapshot through subagent delegation

The lead-agent path resolves one immutable loaded-extension snapshot per run
and binds it through task-store allocation and graph construction, but the
subagent path re-read the process-wide singleton at execution time. In
production both are the same object, yet a `set_loaded_extensions()` between
the lead run's start and a subagent's execution (test teardown, a future
hot-reload path) would let one run mix two extension generations — exactly what
the documented invariant exists to prevent.

The graph-build binding is a ContextVar scoped to synchronous construction, so
it has already exited by the time a tool delegates; the snapshot has to travel
through runtime context instead. The run worker publishes it under the
host-internal `EXTENSION_SNAPSHOT_CONTEXT_KEY` (written after the caller merge,
popped when the run has none, so a caller-supplied value is never
authoritative), `task_tool` reads it back through the type-checking
`resolve_run_extensions()`, and `SubagentExecutor` binds it at construction.

Callers outside the Gateway run path — embedded `DeerFlowClient`, standalone
LangGraph Server — install no snapshot and keep the existing
`get_loaded_extensions()` fallback.

* refactor(extensions): defer the ordering table by call, not by a lying tuple

`CORE_ORDERING_CONSTRAINTS` was a `tuple` subclass that overrode only
`__iter__` and resolved into a class-level `_resolved` side channel. A tuple
cannot populate its own storage after construction, so the instance stayed the
empty tuple it was built as: `len()` was 0, `bool()` was False, `in` was always
False, indexing raised, slicing and `reversed()` came back empty, and it
compared unequal to the plain tuples tests substitute for it — all while
iteration yielded the real constraints. Only `assert_ordering` consumed it, and
only by iterating, so the split went unnoticed.

The sibling `_AnchorTable(dict)` uses the same idea soundly because dict is
mutable: `self.update()` fills the real storage, making every inherited
operation correct. That trick does not survive the port to an immutable type.

Replace it with `core_ordering_constraints()`, matching how `stack.py` defers
the same kind of table via `_anchors()`. The deferral is kept — it is about
dependency direction, not just cycles: `extensions/` is the layer the
middleware layer calls into, so a module-scope `agents.middlewares` import here
points the dependency backwards and closes a cycle as soon as any middleware
imports something under `extensions/` at module level. Resolution stays at
`assert_ordering` time, which already runs inside the middleware builder.

Tests pin both halves: the returned value is a plain tuple whose len/bool/
membership/indexing/reversal/equality agree with iteration, and a subprocess
probe asserts importing `extensions.ordering` does not load the middleware
layer while calling the function does.
2026-08-04 22:33:26 +08:00

89 lines
4.0 KiB
Python

"""Declarative ordering invariants for the middleware stack.
Replaces hand-written index comparisons. Extension-contributed middlewares are
merged before validation runs, so a contribution cannot slip past an invariant,
and the failure names the extension responsible.
A broken invariant is the one hard failure in this system: unlike a missing
observation, it produces wrong behaviour without an error.
"""
from __future__ import annotations
from collections.abc import Mapping, Sequence
from dataclasses import dataclass
from functools import cache
from deerflow.extensions.isolation import IsolatedMiddleware
@dataclass(frozen=True)
class OrderingConstraint:
outer: type
inner: type
reason: str
def _indices_of(middlewares: Sequence[object], target: type) -> list[int]:
indices: list[int] = []
for index, middleware in enumerate(middlewares):
candidate = middleware.inner if isinstance(middleware, IsolatedMiddleware) else middleware
if isinstance(candidate, target):
indices.append(index)
return indices
def assert_ordering(
middlewares: Sequence[object],
provenance: Mapping[int, str],
constraints: Sequence[OrderingConstraint] | None = None,
) -> None:
"""Raise when a constraint is violated. No-op when both sides are absent."""
for constraint in constraints if constraints is not None else core_ordering_constraints():
outer_indices = _indices_of(middlewares, constraint.outer)
inner_indices = _indices_of(middlewares, constraint.inner)
if not outer_indices or not inner_indices:
continue
if max(outer_indices) < min(inner_indices):
continue
violating_indices = [index for index in outer_indices if index >= min(inner_indices)] + [index for index in inner_indices if index <= max(outer_indices)]
culprits = sorted({source for index in violating_indices if (source := provenance.get(index)) is not None})
blame = ", ".join(culprits) if culprits else "core middleware order"
raise RuntimeError(
f"Middleware ordering constraint violated: {constraint.outer.__name__} must be outer "
f"(lower index) of every {constraint.inner.__name__}, but found outer indices "
f"{outer_indices} vs inner indices {inner_indices}. Reason: {constraint.reason}. "
f"Contributed by: {blame}."
)
@cache
def core_ordering_constraints() -> tuple[OrderingConstraint, ...]:
"""The host's ordering invariants, resolved on first use.
Deferred deliberately, and the deferral is about dependency *direction*,
not just cycles: ``extensions/`` is the layer the middleware layer calls
into, so importing ``agents.middlewares`` at module scope here would point
the dependency backwards and close a cycle the moment any middleware
imports something under ``extensions/`` at module level. Resolution instead
happens at ``assert_ordering`` time, which already runs inside the
middleware builder — a forward reference within one layer.
Returns a plain tuple. The predecessor deferred by way of a ``tuple``
subclass overriding only ``__iter__``; because a tuple cannot populate its
own storage after construction, every operation reading that storage
(``len``, ``bool``, ``in``, indexing, slicing, ``reversed``, ``==``)
reported an empty sequence while iteration yielded the real constraints.
Deferring the call instead of faking the value keeps one answer.
"""
from deerflow.agents.middlewares.tool_error_handling_middleware import ToolErrorHandlingMiddleware
from deerflow.agents.middlewares.tool_progress_middleware import ToolProgressMiddleware
return (
OrderingConstraint(
outer=ToolProgressMiddleware,
inner=ToolErrorHandlingMiddleware,
reason=("ToolProgressMiddleware reads deerflow_tool_meta in _update_state_from_result, so its wrap_tool_call chain must enclose the ToolErrorHandlingMiddleware step that stamps it"),
),
)