Zheng Feng dcb2e687d5
feat(channels): add GitHub as a webhook-driven channel (#3754)
* feat(channels): add GitHub event-driven agents (#3754)

Add a webhook-driven GitHub channel with fail-closed webhook routing, deterministic per-agent PR/issue threads, mention-gated trigger fan-out, GitHub App token injection for sandboxed gh/git commands, and backend/AGENTS.md documentation.

* fix(llm-middleware): classify bare IndexError as transient

Upstream chat providers occasionally return 200 OK with an empty
generations list (observed against Volces "coding" on
ark.cn-beijing.volces.com). When that happens,
langchain_core.language_models.chat_models.ainvoke raises
``IndexError: list index out of range`` at
``llm_result.generations[0][0].message`` and kills the run.

Treat a bare IndexError reaching the middleware as a transient
upstream-payload glitch and route it through the existing
retry/backoff path instead of failing the whole agent run. The
retry budget and backoff schedule are unchanged.

Adds three regression tests covering the classifier and both the
recover-on-retry and exhausted-retries paths.

* fix(runtime): ignore stale LLM fallback markers from prior runs

When a run on a thread ends with the LLM-error-handling middleware emitting
a `deerflow_error_fallback`-marked AIMessage (e.g. after the IndexError
empty-generations classification fix lands), that message is persisted to
the thread's checkpoint as part of the messages channel. LangGraph replays
the full message history in `stream_mode="values"` chunks, so every
subsequent run on the same thread re-streams the stale fallback marker —
and the worker's chunk scanner faithfully picks it up, flipping
`RunStatus.success` to `RunStatus.error` for runs that themselves had
no LLM failure at all.

Snapshot the set of pre-existing message ids from the pre-run checkpoint
and thread it through `_extract_llm_error_fallback_message` /
`_try_extract_from_message` as a filter. Markers on history messages are
ignored; markers on fresh messages produced during this run still trip
the error path. Falls back to an empty set when the checkpointer is
absent or the snapshot can't be captured, preserving the prior behavior
on first-run / no-state paths.

Adds unit tests for the new filter (helper-level and `_collect_pre_existing_message_ids`)
plus an integration test exercising the full `run_agent` path with a stale
history checkpointer.

* fix(channels): make github channel fire-and-forget to avoid httpx.ReadTimeout on long runs

GitHub agent runs (clone -> edit -> test -> push -> PR) routinely exceed
the langgraph_sdk default 300s read deadline. The manager's runs.wait
call kept an HTTP stream open for the entire run lifetime, so the long
run blew up with httpx.ReadTimeout and the outer except branch then
released the dedupe key and emitted a false 'internal error' outbound.

The GitHub channel's outbound send is log-only by design: agents post to
the issue/PR via the gh CLI in the sandbox when they choose to comment
or create a PR. There is nothing for the manager to ferry back, so the
long-poll was pure overhead.

This change adds ChannelRunPolicy.fire_and_forget (default False) and
sets it True for the github channel. When fire_and_forget is True,
_handle_chat dispatches via client.runs.create (short POST, returns
once the run is pending) instead of client.runs.wait, and skips the
response-extraction + outbound-publish block. ConflictError on a busy
thread still trips the standard THREAD_BUSY_MESSAGE path so behavior on
the busy case is preserved for any future non-github fire-and-forget
channel.

Other (non-github) channels are unchanged: their policy defaults
fire_and_forget=False and they continue to dispatch via runs.wait.

Adds 6 regression tests in tests/test_channels.py::TestGithubFireAndForget:
- Default ChannelRunPolicy.fire_and_forget is False.
- The github policy registers fire_and_forget=True.
- github inbound calls runs.create, not runs.wait, with the right kwargs.
- github inbound publishes no outbound on success.
- ConflictError from runs.create still emits THREAD_BUSY_MESSAGE.
- Non-github channels (slack) still dispatch via runs.wait.

* test(lead-agent): accept user_id kwarg in skill-policy test stubs

The two GitHub-channel tests added in #3754 stubbed
_load_enabled_skills_for_tool_policy with a lambda that only accepted
`available_skills` and `app_config`, but the real function (and its call
site in agent.py) also passes `user_id`. This raised TypeError on every
run, failing backend-unit-tests.

Add `user_id=None` to match the three sibling stubs in the same file.

* refactor(gateway): disambiguate context-key set names

The two frozensets _INTERNAL_ONLY_CONTEXT_KEYS and _CONTEXT_ONLY_KEYS
shared a confusable "CONTEXT_ONLY" token in different orders, and the
first broke the _CONTEXT_<X>_KEYS pattern of its sibling
_CONTEXT_CONFIGURABLE_KEYS. Rename to make the distinct axes explicit:

  _CONTEXT_INTERNAL_CALLER_KEYS  - WHO: internal callers (scheduler) only
  _CONTEXT_RUNTIME_ONLY_KEYS     - WHERE: runtime context only, never configurable

Pure rename, no behavior change.
2026-07-04 22:56:24 +08:00

116 lines
4.7 KiB
Python

"""GitHub channel — webhook-driven IM channel for PR/issue comments.
Unlike other IM channels (Feishu, Slack, Telegram) which long-poll or use
WebSockets, GitHub delivers messages via HTTP push webhooks. This channel
therefore has a no-op ``start``/``stop`` — inbound messages arrive through
``POST /api/webhooks/github`` and are published to the bus by the webhook
route handler.
**The channel does not auto-post the agent's final response.** Each GitHub
agent (coder, reviewer, …) has the `gh` CLI in its sandbox and is expected to
decide for itself what — if anything — to post on the issue or PR, and to use
``gh issue comment`` / ``gh pr comment`` / ``gh pr create`` during the run.
The agent's final assistant message is logged at INFO for visibility in
``gateway.log`` but is **not** sent to GitHub.
Why log-only rather than auto-post:
- Two agents can bind the same event (e.g. coder + reviewer on a mention).
If both auto-posted their final messages, the user would see two replies
for every mention even when only one had useful work to do. Letting the
LLM call ``gh`` mid-run means silence is just "the LLM did not call gh."
- The agent often wants to post *intermediate* updates (an issue comment
linking the PR, a separate comment on a new sub-issue, …) — the
auto-post-the-final-message contract didn't model that and forced the
final message to play double duty.
- The dispatcher's per-agent ``_is_self_event`` gate already prevents the
comments the LLM posts via ``gh`` from looping the webhook back into a
new run for the same agent.
"""
from __future__ import annotations
import logging
from typing import Any
from app.channels.base import Channel
from app.channels.message_bus import MessageBus, OutboundMessage
logger = logging.getLogger(__name__)
class GitHubChannel(Channel):
"""Webhook-driven GitHub channel.
Inbound: ``POST /api/webhooks/github`` publishes ``InboundMessage`` to
the bus. Outbound: ``send`` is log-only (see module docstring) — agents
post to GitHub themselves via the ``gh`` CLI in their sandbox.
Configuration keys (in ``config.yaml`` under ``channels.github``):
- ``enabled`` (bool): set to ``true`` to activate.
- ``default_mention_login`` (str, optional): bot handle used by
``require_mention`` when the agent binding does not set one.
Falls back to ``"deerflow-bot"``.
"""
def __init__(self, bus: MessageBus, config: dict[str, Any]) -> None:
super().__init__(name="github", bus=bus, config=config)
# -- lifecycle ---------------------------------------------------------
async def start(self) -> None:
"""Register the outbound callback.
GitHub is push-based (webhooks), so no long-poll or socket
listener is needed. We only register for outbound replies so the
agent's final message gets logged.
"""
if self._running:
return
self.bus.subscribe_outbound(self._on_outbound)
self._running = True
logger.info("GitHubChannel started (webhook-driven, no polling)")
async def stop(self) -> None:
"""Unregister the outbound callback."""
if not self._running:
return
self.bus.unsubscribe_outbound(self._on_outbound)
self._running = False
logger.info("GitHubChannel stopped")
# -- outbound ----------------------------------------------------------
async def send(self, msg: OutboundMessage) -> None:
"""Log the agent's final message — do NOT post it to GitHub.
GitHub agents post to issues/PRs themselves via ``gh`` mid-run; the
final assistant message is logged for ``gateway.log`` visibility but
is not delivered to the platform. See the module docstring for why.
Metadata layout (read for logging context only):
- ``repo`` (str, e.g. ``"owner/name"``) — falls back to ``chat_id``
- ``number`` (int, issue or PR number)
- ``installation_id`` (int)
"""
gh = msg.metadata.get("github", {}) if isinstance(msg.metadata, dict) else {}
if not isinstance(gh, dict):
gh = {}
repo = gh.get("repo") or msg.chat_id
number = gh.get("number")
body = msg.text or ""
logger.info(
"[GitHubChannel] final message from agent for %s#%s (text_len=%d) — not posted; agents use `gh` directly",
repo,
number,
len(body),
)
# Mirror the body itself at DEBUG so operators can correlate without
# spamming INFO. Truncate to keep log lines bounded.
if body:
logger.debug("[GitHubChannel] final body (truncated to 2000 chars): %s", body[:2000])