feat(community): add Sofya web search provider (#5239)

* feat(community): add Sofya web search provider

Add a community provider backed by Sofya (https://sofya.co). Its search
endpoint returns the content of the result pages, not only their snippets,
and its fetch endpoint returns a page as markdown. Both are plain JSON over
HTTP, so this needs no extra Python package (uses httpx, already a
dependency).

Changes:
- backend/packages/harness/deerflow/community/sofya/__init__.py
- backend/packages/harness/deerflow/community/sofya/tools.py
  Implements web_search_tool and web_fetch_tool using httpx.
  API key is read from the config.yaml `api_key` field or the SOFYA_API_KEY
  env var. Follows the same interface and output shape as the existing
  ddg_search and serper providers, including the max_results parameter with
  config override and the structured "No results found" error.
- backend/tests/test_sofya_tools.py
  Unit tests covering API key resolution, config overrides, result mapping,
  time range, HTTP errors, empty results, and fetch failures.
- config.example.yaml: add commented-out Sofya web_search and web_fetch
  examples alongside the other providers
- .env.example: add SOFYA_API_KEY placeholder
- backend/docs/CONFIGURATION.md: list Sofya under web_search, web_fetch and
  the environment variables

* fix(sofya): honor caller max_results, validate search_depth, join time_range contract test

- Caller-supplied max_results now wins; config is used only when the
  argument is omitted, matching GroundRoute.
- search_depth is clamped to basic/snippets; an unsupported value logs a
  warning and falls back to basic.
- Sofya added to the shared time_range schema contract test.

* fix(sofya): cap per-result content so a search stays inline

An unbounded search payload (up to 20 read pages) crossed the tool output
budget middleware's externalize_min_chars threshold, which replaces the
result list with a file reference. Cap each result's content at
contents_max_characters (default 2000, 0 disables), matching Exa's config
key. Five capped results stay under the 12000 char threshold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5

* fix(sofya): list Sofya in the recency contract, coerce non-string content

_clip subscripted its input, so a non-string content or description from
the API raised TypeError instead of degrading. Coerce to text first, the
way _sofya_post and _response_results guard the shapes around it. Also add
Sofya to the Web Search Recency section in backend/AGENTS.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5

* fix(sofya): coerce web_fetch content, list sofya in the tools guide, add changelog

web_fetch sliced its content the same way web_search did before the last
push: a truthy non-string from the API passed the falsiness guard and then
raised TypeError. Reuse _clip, keeping the `or ""` so empty content still
reports "No content found".

Also add sofya to the community provider inventory in
packages/harness/deerflow/tools/AGENTS.md and an [Unreleased] changelog entry.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5

* docs(zh): add the missing InfoQuest and Firecrawl web_fetch tabs

The ZH web_fetch tab list named five providers where EN names seven. Both
tabs mirror their EN counterparts, so the two locales list the same
web_fetch providers again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
This commit is contained in:
Yusuf Gürdoğan 2026-09-08 05:50:03 +03:00 committed by GitHub
parent 9ce6fdcb22
commit 23bd76046a
No known key found for this signature in database
GPG Key ID: B5690EEEBB952194
15 changed files with 952 additions and 11 deletions

View File

@ -12,6 +12,10 @@ JINA_API_KEY=your-jina-api-key
# InfoQuest API Key
INFOQUEST_API_KEY=your-infoquest-api-key
# Sofya API Key (web search and fetch) - https://sofya.co
SOFYA_API_KEY=your-sofya-api-key
# Browser CORS allowlist for split-origin or port-forwarded deployments (comma-separated exact origins).
# Leave unset when using the unified nginx endpoint, e.g. http://localhost:2026.
# GATEWAY_CORS_ORIGINS=http://localhost:3000,http://127.0.0.1:3000

View File

@ -251,6 +251,9 @@ This section accumulates work toward the **2.1.0** milestone
verticals) and Tencent Cloud WSA - plus native recency filters
(day/week/month/year) shared across DDGS, Brave, Tavily, and SearXNG.
([#5023], [#5057], [#5099])
- **community:** New Sofya `web_search` and `web_fetch` provider - search
results carry the content of each page, capped per result so a default
search stays inline. ([#5239])
- **knowledge:** Opt-in read-only RAGFlow retrieval exposes a
`knowledge_search(query)` agent tool over configured RAGFlow datasets, with
a dataset-ID allowlist and credential/dataset-id redaction on error paths.
@ -2379,4 +2382,5 @@ with **180 merged pull requests** since the first 2.0 milestone tag.
[#5117]: https://github.com/bytedance/deer-flow/pull/5117
[#5133]: https://github.com/bytedance/deer-flow/pull/5133
[#5136]: https://github.com/bytedance/deer-flow/pull/5136
[#5239]: https://github.com/bytedance/deer-flow/pull/5239
[#5119]: https://github.com/bytedance/deer-flow/pull/5119

View File

@ -310,9 +310,10 @@ When using `make dev` from root, the frontend automatically connects through ngi
### Web Search Recency
DDG, Brave, Tavily, and SearXNG `web_search` share optional
DDG, Brave, Tavily, SearXNG, and Sofya `web_search` share optional
`time_range=day|week|month|year`; omission preserves request shape. DDG maps to
`d|w|m|y`, Brave to `pd|pw|pm|py`, and Tavily/SearXNG pass values unchanged.
`d|w|m|y`, Brave to `pd|pw|pm|py`, Tavily/SearXNG pass values unchanged, and
Sofya passes them unchanged as `freshness`.
For recency, DDGS 9.14.1 uses only enabled Brave, DuckDuckGo, and Yahoo engines
that honor `timelimit`: `auto`/`all` resolves to this set, incompatible configured
engines are removed, and an empty set falls back to it. Re-check on DDGS upgrades.

View File

@ -428,8 +428,8 @@ tools:
```
**Built-in Tools**:
- `web_search` - Search the web (DuckDuckGo, Tavily, Brave, Serply, Exa, InfoQuest, Tencent Cloud WSA, Firecrawl, fastCRW, GroundRoute)
- `web_fetch` - Fetch web pages (Jina AI, Crawl4AI, Exa, InfoQuest, Firecrawl, fastCRW, GroundRoute, Browserless)
- `web_search` - Search the web (DuckDuckGo, Tavily, Brave, Serply, Exa, InfoQuest, Tencent Cloud WSA, Firecrawl, fastCRW, GroundRoute, Sofya)
- `web_fetch` - Fetch web pages (Jina AI, Crawl4AI, Exa, InfoQuest, Firecrawl, fastCRW, GroundRoute, Browserless, Sofya)
- `web_capture` - Capture rendered webpage screenshots as artifacts (Browserless)
- `image_search` - Search for reference images (DuckDuckGo, InfoQuest, Serper, Brave)
- `ls` - List directory contents
@ -981,6 +981,7 @@ models:
- `SERPER_API_KEY` - Serper (Google Search/Images API) key for `web_search` and `image_search`
- `SERPLY_API_KEY` - [Serply](https://serply.io) key for `web_search` (Google Search, plus Google News and Google Scholar via `vertical`)
- `GROUNDROUTE_API_KEY` - GroundRoute meta-search API key for `web_search` and `web_fetch` (routes across Serper, Brave, Exa, Tavily, Firecrawl, Perplexity with gain-share pricing)
- `SOFYA_API_KEY` - [Sofya](https://sofya.co) key for `web_search` and `web_fetch`
- `BROWSERLESS_TOKEN` - Browserless Cloud token for `web_capture` (optional for self-hosted Browserless)
- `DEER_FLOW_PROJECT_ROOT` - Project root for relative runtime paths
- `DEER_FLOW_CONFIG_PATH` - Custom config file path

View File

@ -0,0 +1,3 @@
from .tools import web_fetch_tool, web_search_tool
__all__ = ["web_fetch_tool", "web_search_tool"]

View File

@ -0,0 +1,223 @@
"""
Web search and fetch tools powered by Sofya.
Sofya is a hosted web API for agents. Search returns the content of the result
pages, not only their snippets, and fetch returns a single page as clean
markdown. An API key is required. Sign up at https://sofya.co to get one.
"""
import json
import logging
import os
import httpx
from langchain.tools import tool
from deerflow.community.search_time_range import SearchTimeRange
from deerflow.config import get_app_config
logger = logging.getLogger(__name__)
_SOFYA_BASE_URL = "https://sofya.co/v1"
_SOFYA_SEARCH_PATH = "/search"
_SOFYA_FETCH_PATH = "/fetch"
_SOFYA_MAX_RESULTS = 20
_SOFYA_TIMEOUT = 60
_SOFYA_FETCH_MAX_CHARS = 4096
_DEFAULT_SEARCH_DEPTH = "basic"
_DEFAULT_CONTENTS_MAX_CHARACTERS = 2000
_SEARCH_DEPTHS = ("basic", "snippets")
_api_key_warned: set[str] = set()
def _get_api_key(tool_name: str) -> str | None:
config = get_app_config().get_tool_config(tool_name)
if config is not None:
api_key = config.model_extra.get("api_key")
if isinstance(api_key, str) and api_key.strip():
return api_key.strip()
env_key = os.getenv("SOFYA_API_KEY")
if isinstance(env_key, str) and env_key.strip():
return env_key.strip()
return None
def _coerce_max_results(value: object, default: int = 5, max_allowed: int = _SOFYA_MAX_RESULTS) -> int:
"""Coerce config/parameter input into a bounded positive result count."""
try:
count = int(value) # type: ignore[call-overload]
except (TypeError, ValueError):
return default
if count <= 0:
return default
return min(count, max_allowed)
def _coerce_content_limit(value: object, default: int = _DEFAULT_CONTENTS_MAX_CHARACTERS) -> int:
"""Coerce the per-result content limit. 0 means no limit; anything invalid falls back to the default."""
try:
limit = int(value) # type: ignore[call-overload]
except (TypeError, ValueError):
return default
return limit if limit >= 0 else default
def _resolve_search_depth(value: object) -> str:
"""Return a supported search depth, falling back to the default with a warning."""
if value is None:
return _DEFAULT_SEARCH_DEPTH
depth = str(value).strip().lower()
if depth in _SEARCH_DEPTHS:
return depth
logger.warning("Ignoring unsupported Sofya search_depth %r; using %r (supported: %s)", value, _DEFAULT_SEARCH_DEPTH, ", ".join(_SEARCH_DEPTHS))
return _DEFAULT_SEARCH_DEPTH
def _missing_key_message(tool_name: str) -> str:
if tool_name not in _api_key_warned:
_api_key_warned.add(tool_name)
logger.warning("Sofya API key is not set for '%s'. Set SOFYA_API_KEY in your environment or provide api_key in config.yaml. Sign up at https://sofya.co", tool_name)
return "SOFYA_API_KEY is not configured"
def _sofya_post(path: str, api_key: str, payload: dict) -> tuple[dict | None, str | None]:
"""Send a POST request to a Sofya endpoint.
Returns a ``(data, error)`` tuple: on success ``data`` is the parsed JSON
object and ``error`` is ``None``; on failure ``data`` is ``None`` and
``error`` is a message the caller can hand back to the model.
"""
headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
}
try:
with httpx.Client(timeout=_SOFYA_TIMEOUT) as client:
response = client.post(f"{_SOFYA_BASE_URL}{path}", headers=headers, json=payload)
response.raise_for_status()
data = response.json()
except httpx.HTTPStatusError as e:
logger.error("Sofya API returned HTTP %s: %s", e.response.status_code, (e.response.text or "")[:500])
return None, f"Sofya API error: HTTP {e.response.status_code}"
except Exception as e:
logger.error("Sofya request failed: %s: %s", type(e).__name__, str(e)[:500])
return None, str(e)[:500]
if not isinstance(data, dict):
logger.error("Sofya returned an unexpected payload type: %s", type(data).__name__)
return None, "Sofya returned an unexpected response format"
return data, None
def _clip(value: object, limit: int) -> str:
"""Coerce a result field to text and truncate it. A limit of 0 means no truncation."""
text = value if isinstance(value, str) else str(value)
return text if limit <= 0 else text[:limit]
def _response_results(data: dict) -> list[dict] | None:
"""Return the result dicts of a Sofya response, or None if malformed."""
results = data.get("results")
if results is None:
return []
if not isinstance(results, list):
logger.error("Sofya returned an unexpected 'results' payload type: %s", type(results).__name__)
return None
return [item for item in results if isinstance(item, dict)]
@tool("web_search", parse_docstring=True)
def web_search_tool(query: str, max_results: int | None = None, time_range: SearchTimeRange | None = None) -> str:
"""Search the web for information. Use this tool to find current information, news, articles, and facts from the internet.
Args:
query: Search keywords describing what you want to find. Be specific for better results.
max_results: Maximum number of results to return. If omitted, uses the configured value (default 5). Capped at 20.
time_range: Optional relative publication/update window. Use only when the request requires recent results.
"""
config = get_app_config().get_tool_config("web_search")
config_extra = (config.model_extra or {}) if config is not None else {}
# Honor the caller-supplied max_results; fall back to config only when omitted.
if max_results is None:
max_results = config_extra.get("max_results")
max_results = _coerce_max_results(max_results)
search_depth = _resolve_search_depth(config_extra.get("search_depth"))
content_limit = _coerce_content_limit(config_extra.get("contents_max_characters"))
api_key = _get_api_key("web_search")
if not api_key:
return json.dumps({"error": _missing_key_message("web_search"), "query": query}, ensure_ascii=False)
payload: dict[str, object] = {
"query": query,
"max_results": max_results,
"search_depth": search_depth,
}
if time_range is not None:
payload["freshness"] = time_range
data, error = _sofya_post(_SOFYA_SEARCH_PATH, api_key, payload)
if error is not None:
return json.dumps({"error": error, "query": query}, ensure_ascii=False)
results = _response_results(data)
if results is None:
return json.dumps({"error": "Sofya returned an unexpected response format", "query": query}, ensure_ascii=False)
if not results:
return json.dumps({"error": "No results found", "query": query}, ensure_ascii=False)
normalized_results = [
{
"title": r.get("title", ""),
"url": r.get("url", ""),
# Page content when the result was read, the search snippet otherwise.
# Capped so a normal search stays inline rather than being written to disk.
"content": _clip(r.get("content") or r.get("description") or "", content_limit),
}
for r in results[:max_results]
]
output = {
"query": query,
"total_results": len(normalized_results),
"results": normalized_results,
}
return json.dumps(output, indent=2, ensure_ascii=False)
@tool("web_fetch", parse_docstring=True)
def web_fetch_tool(url: str) -> str:
"""Fetch the contents of a web page at a given URL.
Only fetch EXACT URLs that have been provided directly by the user or have been returned in results from the web_search and web_fetch tools.
This tool can NOT access content that requires authentication, such as private Google Docs or pages behind login walls.
Do NOT add www. to URLs that do NOT have them.
URLs must include the schema: https://example.com is a valid URL while example.com is an invalid URL.
Args:
url: The URL to fetch the contents of.
"""
api_key = _get_api_key("web_fetch")
if not api_key:
return f"Error: {_missing_key_message('web_fetch')}"
data, error = _sofya_post(_SOFYA_FETCH_PATH, api_key, {"urls": [url]})
if error is not None:
return f"Error: {error}"
results = _response_results(data)
if results is None:
return "Error: Sofya returned an unexpected response format"
if not results:
return "Error: No results found"
result = results[0]
if not result.get("success", True):
return f"Error: {result.get('error') or 'Failed to fetch the URL'}"
content = _clip(result.get("content") or "", _SOFYA_FETCH_MAX_CHARS)
if not content:
return "Error: No content found"
title = result.get("title") or "Untitled"
return f"# {title}\n\n{content}"

View File

@ -24,6 +24,7 @@ Durable MCP task-management tools are added only while the process-local task su
**Community tools** (`packages/harness/deerflow/community/`): optional integrations, each in its own subpackage and wired through `config.yaml`. Documented examples:
- `tavily/` - Web search (5 results default) and web fetch (4KB limit)
- `sofya/` - Web search (5 results default, per-result content capped at 2000 chars) and web fetch (4KB limit)
- `jina_ai/` - Web fetch via Jina reader API with readability extraction
- `firecrawl/` - Web scraping via Firecrawl API
- `image_search/` - Image search via DuckDuckGo
@ -31,7 +32,7 @@ Durable MCP task-management tools are added only while the process-local task su
- `browser_automation/` - Agentic browser control (stateful `navigate → observe → click/type` loop) via Playwright, distinct from the read-only `web_fetch`/`web_capture` tools. Tools: `browser_navigate`, `browser_snapshot`, `browser_click`, `browser_type`, `browser_get_text`, `browser_back`, `browser_screenshot`, `browser_close` (config `group: browser`). A process-local `BrowserSessionManager` owns one private, loop-affine Playwright event-loop thread (same pattern as the BoxLite provider) so a per-thread browser session survives across turns regardless of the caller's loop (Gateway / TUI / test). Each action returns a fresh page snapshot whose interactive elements are addressed by a stable numeric `[ref]` index (stamped as `data-df-ref` during snapshot), so the model acts on what it just observed instead of holding stale handles or guessing selectors. URLs are SSRF-screened via the shared `validate_public_http_url` (opt-out `allow_private_addresses` only for intentional internal targets). CDP attachment cannot install the request guard on an existing Chrome context, so `cdp_url` fails closed unless the operator explicitly sets `allow_unguarded_cdp: true` for a trusted local browser. Browser REST/Live access also requires an exact non-NULL thread owner, rather than the general legacy shared-thread policy, because retained pages may contain authenticated state. Session admission is a hard `max_sessions` cap: pinned Live/operation sessions are never evicted, and a new thread is rejected when no unpinned session can be closed; one Live viewer owns a session at a time. Optional dependency: `cd backend && uv sync --extra browser && uv run playwright install chromium`; `scripts/detect_uv_extras.py` preserves the extra when `config.yaml` enables `browser_navigate`, and Gateway startup fails fast if configured browser control cannot import Playwright. Tests: `tests/test_browser_automation.py` (mocked tools + a real-Chromium integration test guarded by `importorskip`); `tests/manual_browser_live_check.py` is a manual DeepSeek-driven end-to-end check (not collected by pytest).
Live UI input dispatch is kept independent from JPEG capture: non-move actions start a rate-limited background refresh loop, so pointer, wheel, or keyboard input stays responsive while continuous gestures still produce frames throughout the interaction.
Additional providers also live here (`boxlite`, `brave`, `browserless`, `crawl4ai`, `ddg_search`, `e2b_sandbox`, `exa`, `fastcrw`, `groundroute`, `infoquest`, `searxng`, `serper`, `serply`, `tencent_wsa`, `tenki`); see each subpackage for specifics. `tencent_wsa` uses Tencent Cloud Web Search's service API key endpoint (`TENCENTCLOUD_WSA_APIKEY`), not Tencent Cloud SecretId/SecretKey signing. Its `max_results` is capped at 50; requests above 10 use Tencent's optional `Cnt` parameter, which needs a Tencent Cloud plan that supports it. E2B bootstrap is required. If it fails, the provider kills and closes the unusable remote sandbox. New sandbox creation raises an error. Warm-pool reclaim and remote discovery discard the sandbox and continue acquisition. E2B mounts remain optional.
Additional providers also live here (`boxlite`, `brave`, `browserless`, `crawl4ai`, `ddg_search`, `e2b_sandbox`, `exa`, `fastcrw`, `groundroute`, `infoquest`, `searxng`, `serper`, `serply`, `sofya`, `tencent_wsa`, `tenki`); see each subpackage for specifics. `tencent_wsa` uses Tencent Cloud Web Search's service API key endpoint (`TENCENTCLOUD_WSA_APIKEY`), not Tencent Cloud SecretId/SecretKey signing. Its `max_results` is capped at 50; requests above 10 use Tencent's optional `Cnt` parameter, which needs a Tencent Cloud plan that supports it. E2B bootstrap is required. If it fails, the provider kills and closes the unusable remote sandbox. New sandbox creation raises an error. Warm-pool reclaim and remote discovery discard the sandbox and continue acquisition. E2B mounts remain optional.
E2B output sync records remote file versions and actual host file metadata in a thread-local manifest. The manifest binds to the remote sandbox ID. A complete output listing removes entries for deleted files. This avoids repeat downloads when the host filesystem rounds modification times. A single release-time sync pass is bounded by aggregate ceilings (`_MAX_SYNC_TOTAL_BYTES`, `_MAX_SYNC_FILES`, `_SYNC_DEADLINE_SECONDS`) on top of the per-file `_MAX_DOWNLOAD_SIZE` cap, so a pathological outputs tree cannot make release download unboundedly; a truncated pass logs what it dropped and leaves the manifest un-pruned (only entries observed in that pass are reconciled), so files it never reached are retried on the next release rather than being forgotten.

View File

@ -302,6 +302,22 @@ class TestCheckWebSearch:
assert result.status == "ok"
assert "BRAVE_SEARCH_API_KEY set from config" in result.detail
def test_sofya_with_key_ok(self, tmp_path, monkeypatch):
monkeypatch.setenv("SOFYA_API_KEY", "test-key")
cfg = tmp_path / "config.yaml"
cfg.write_text("config_version: 5\ntools:\n - name: web_search\n use: deerflow.community.sofya.tools:web_search_tool\n")
result = doctor.check_web_search(cfg)
assert result.status == "ok"
assert "sofya" in result.detail
def test_sofya_without_key_warns(self, tmp_path, monkeypatch):
monkeypatch.delenv("SOFYA_API_KEY", raising=False)
cfg = tmp_path / "config.yaml"
cfg.write_text("config_version: 5\ntools:\n - name: web_search\n use: deerflow.community.sofya.tools:web_search_tool\n")
result = doctor.check_web_search(cfg)
assert result.status == "warn"
assert "SOFYA_API_KEY" in (result.fix or "")
def test_serper_with_key_ok(self, tmp_path, monkeypatch):
monkeypatch.setenv("SERPER_API_KEY", "test-key")
cfg = tmp_path / "config.yaml"
@ -423,6 +439,14 @@ class TestCheckWebFetch:
assert result.status == "warn"
assert "FIRECRAWL_API_KEY" in (result.fix or "")
def test_sofya_without_key_warns(self, tmp_path, monkeypatch):
monkeypatch.delenv("SOFYA_API_KEY", raising=False)
cfg = tmp_path / "config.yaml"
cfg.write_text("config_version: 5\ntools:\n - name: web_fetch\n use: deerflow.community.sofya.tools:web_fetch_tool\n")
result = doctor.check_web_fetch(cfg)
assert result.status == "warn"
assert "SOFYA_API_KEY" in (result.fix or "")
def test_no_fetch_tool_warns(self, tmp_path):
cfg = tmp_path / "config.yaml"
cfg.write_text("config_version: 5\ntools: []\n")

View File

@ -0,0 +1,593 @@
"""Unit tests for the Sofya community web search and fetch tools."""
import json
import logging
from unittest.mock import MagicMock, patch
import httpx
import pytest
@pytest.fixture(autouse=True)
def reset_api_key_warned():
"""Reset the module-level warning flag before each test."""
import deerflow.community.sofya.tools as sofya_mod
sofya_mod._api_key_warned = set()
yield
sofya_mod._api_key_warned = set()
@pytest.fixture
def mock_config_with_key():
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
tool_config = MagicMock()
tool_config.model_extra = {"api_key": "test-sofya-key", "max_results": 5}
mock.return_value.get_tool_config.return_value = tool_config
yield mock
@pytest.fixture
def mock_config_no_key():
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
tool_config = MagicMock()
tool_config.model_extra = {}
mock.return_value.get_tool_config.return_value = tool_config
yield mock
def _make_response(payload: object) -> MagicMock:
mock_resp = MagicMock()
mock_resp.json.return_value = payload
mock_resp.raise_for_status = MagicMock()
return mock_resp
def _make_search_response(results: list) -> MagicMock:
return _make_response({"query": "test", "results": results})
def _make_fetch_response(results: list) -> MagicMock:
return _make_response({"results": results})
class TestGetApiKey:
def test_returns_config_key_when_present(self):
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
tool_config = MagicMock()
tool_config.model_extra = {"api_key": "from-config"}
mock.return_value.get_tool_config.return_value = tool_config
from deerflow.community.sofya.tools import _get_api_key
assert _get_api_key("web_search") == "from-config"
def test_falls_back_to_env_when_config_key_whitespace(self):
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
tool_config = MagicMock()
tool_config.model_extra = {"api_key": " "}
mock.return_value.get_tool_config.return_value = tool_config
with patch.dict("os.environ", {"SOFYA_API_KEY": "env-key"}):
from deerflow.community.sofya.tools import _get_api_key
assert _get_api_key("web_search") == "env-key"
def test_uses_env_when_tool_is_not_configured(self):
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
mock.return_value.get_tool_config.return_value = None
with patch.dict("os.environ", {"SOFYA_API_KEY": "env-only"}):
from deerflow.community.sofya.tools import _get_api_key
assert _get_api_key("web_fetch") == "env-only"
def test_returns_none_when_no_key_anywhere(self):
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
mock.return_value.get_tool_config.return_value = None
with patch.dict("os.environ", {}, clear=True):
from deerflow.community.sofya.tools import _get_api_key
assert _get_api_key("web_search") is None
def test_returns_none_when_env_key_whitespace(self):
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
mock.return_value.get_tool_config.return_value = None
with patch.dict("os.environ", {"SOFYA_API_KEY": " "}):
from deerflow.community.sofya.tools import _get_api_key
assert _get_api_key("web_search") is None
def test_reads_config_for_requested_tool_name(self):
with patch("deerflow.community.sofya.tools.get_app_config") as mock:
tool_config = MagicMock()
tool_config.model_extra = {"api_key": "fetch-key"}
mock.return_value.get_tool_config.return_value = tool_config
from deerflow.community.sofya.tools import _get_api_key
assert _get_api_key("web_fetch") == "fetch-key"
mock.return_value.get_tool_config.assert_called_with("web_fetch")
class TestCoerceMaxResults:
def test_returns_value_when_valid_positive_int(self):
from deerflow.community.sofya.tools import _coerce_max_results
assert _coerce_max_results(3) == 3
def test_returns_value_for_numeric_string(self):
from deerflow.community.sofya.tools import _coerce_max_results
assert _coerce_max_results("7") == 7
def test_caps_value_at_default_maximum(self):
from deerflow.community.sofya.tools import _coerce_max_results
assert _coerce_max_results(999) == 20
def test_returns_default_for_non_numeric_string(self):
from deerflow.community.sofya.tools import _coerce_max_results
assert _coerce_max_results("oops") == 5
def test_returns_default_for_none(self):
from deerflow.community.sofya.tools import _coerce_max_results
assert _coerce_max_results(None) == 5
def test_returns_default_for_zero_or_negative(self):
from deerflow.community.sofya.tools import _coerce_max_results
assert _coerce_max_results(0) == 5
assert _coerce_max_results(-3) == 5
class TestMissingKeyMessage:
def test_warns_once_per_tool_name(self, caplog):
import deerflow.community.sofya.tools as sofya_mod
with caplog.at_level(logging.WARNING):
sofya_mod._missing_key_message("web_search")
sofya_mod._missing_key_message("web_search")
warnings = [r for r in caplog.records if r.levelno == logging.WARNING]
assert len(warnings) == 1
assert "web_search" in warnings[0].getMessage()
def test_warns_separately_for_each_tool(self, caplog):
import deerflow.community.sofya.tools as sofya_mod
with caplog.at_level(logging.WARNING):
sofya_mod._missing_key_message("web_search")
sofya_mod._missing_key_message("web_fetch")
warnings = [r for r in caplog.records if r.levelno == logging.WARNING]
assert len(warnings) == 2
class TestWebSearchTool:
def test_basic_search_returns_normalized_results(self, mock_config_with_key):
results = [
{"title": "Result 1", "url": "https://example.com/1", "content": "Page content 1", "description": "Snippet 1"},
{"title": "Result 2", "url": "https://example.com/2", "content": "Page content 2", "description": "Snippet 2"},
]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "python tutorial"}))
assert parsed["query"] == "python tutorial"
assert parsed["total_results"] == 2
assert parsed["results"][0] == {"title": "Result 1", "url": "https://example.com/1", "content": "Page content 1"}
assert mock_post.call_args.kwargs["headers"]["Authorization"] == "Bearer test-sofya-key"
assert mock_post.call_args.kwargs["json"]["query"] == "python tutorial"
assert mock_post.call_args.kwargs["json"]["search_depth"] == "basic"
assert "freshness" not in mock_post.call_args.kwargs["json"]
def test_falls_back_to_description_when_content_is_empty(self, mock_config_with_key):
results = [{"title": "Result", "url": "https://example.com", "content": "", "description": "Snippet"}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["results"][0]["content"] == "Snippet"
def test_time_range_is_sent_as_freshness(self, mock_config_with_key):
results = [{"title": "Result", "url": "https://example.com", "content": "Body"}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
web_search_tool.invoke({"query": "test", "time_range": "week"})
assert mock_post.call_args.kwargs["json"]["freshness"] == "week"
def test_search_depth_can_be_set_from_config(self, mock_config_with_key):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"search_depth": "snippets",
}
results = [{"title": "Result", "url": "https://example.com", "description": "Snippet"}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
web_search_tool.invoke({"query": "test"})
assert mock_post.call_args.kwargs["json"]["search_depth"] == "snippets"
def test_non_string_content_does_not_raise(self, mock_config_with_key):
results = [
{"title": "Numeric", "url": "https://example.com/1", "content": 12345},
{"title": "Listy", "url": "https://example.com/2", "content": None, "description": ["a", "b"]},
]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["results"][0]["content"] == "12345"
assert parsed["results"][1]["content"] == "['a', 'b']"
def test_result_content_is_capped_by_default(self, mock_config_with_key):
results = [{"title": "Result", "url": "https://example.com", "content": "x" * 9000}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert len(parsed["results"][0]["content"]) == 2000
def test_contents_max_characters_from_config(self, mock_config_with_key):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"contents_max_characters": 100,
}
results = [{"title": "Result", "url": "https://example.com", "content": "x" * 9000}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert len(parsed["results"][0]["content"]) == 100
def test_contents_max_characters_zero_disables_the_cap(self, mock_config_with_key):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"contents_max_characters": 0,
}
results = [{"title": "Result", "url": "https://example.com", "content": "x" * 9000}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert len(parsed["results"][0]["content"]) == 9000
def test_invalid_contents_max_characters_falls_back_to_default(self):
from deerflow.community.sofya.tools import _coerce_content_limit
assert _coerce_content_limit("oops") == 2000
assert _coerce_content_limit(None) == 2000
assert _coerce_content_limit(-5) == 2000
assert _coerce_content_limit(0) == 0
assert _coerce_content_limit("150") == 150
def test_default_search_stays_under_the_externalize_threshold(self, mock_config_with_key):
"""Five capped results must stay inline rather than being persisted to disk."""
results = [{"title": f"R{i}", "url": f"https://example.com/{i}", "content": "x" * 20000} for i in range(5)]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
output = web_search_tool.invoke({"query": "test"})
assert len(output) < 12000
def test_caller_max_results_wins_over_config(self, mock_config_with_key):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"max_results": 5,
}
results = [{"title": f"R{i}", "url": f"https://x.com/{i}", "content": f"C{i}"} for i in range(10)]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test", "max_results": 8}))
assert parsed["total_results"] == 8
assert mock_post.call_args.kwargs["json"]["max_results"] == 8
def test_unsupported_search_depth_falls_back_with_warning(self, mock_config_with_key, caplog):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"search_depth": "advanced",
}
results = [{"title": "Result", "url": "https://example.com", "content": "Body"}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
with caplog.at_level(logging.WARNING):
web_search_tool.invoke({"query": "test"})
assert mock_post.call_args.kwargs["json"]["search_depth"] == "basic"
assert any("search_depth" in r.getMessage() for r in caplog.records if r.levelno == logging.WARNING)
def test_search_depth_is_normalized(self, mock_config_with_key):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"search_depth": " Snippets ",
}
results = [{"title": "Result", "url": "https://example.com", "description": "Snippet"}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
web_search_tool.invoke({"query": "test"})
assert mock_post.call_args.kwargs["json"]["search_depth"] == "snippets"
def test_respects_max_results_from_config(self, mock_config_with_key):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"max_results": 3,
}
results = [{"title": f"R{i}", "url": f"https://x.com/{i}", "content": f"C{i}"} for i in range(10)]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["total_results"] == 3
assert mock_post.call_args.kwargs["json"]["max_results"] == 3
def test_config_max_results_is_capped(self, mock_config_with_key):
mock_config_with_key.return_value.get_tool_config.return_value.model_extra = {
"api_key": "test-key",
"max_results": 999,
}
results = [{"title": f"R{i}", "url": f"https://x.com/{i}", "content": f"C{i}"} for i in range(30)]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["total_results"] == 20
assert mock_post.call_args.kwargs["json"]["max_results"] == 20
def test_max_results_parameter_accepted(self, mock_config_no_key):
"""Tool accepts max_results as a call parameter when config does not override it."""
results = [{"title": f"R{i}", "url": f"https://x.com/{i}", "content": f"C{i}"} for i in range(10)]
with patch.dict("os.environ", {"SOFYA_API_KEY": "env-key"}):
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response(results)
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test", "max_results": 2}))
assert parsed["total_results"] == 2
def test_empty_results_return_error_json(self, mock_config_with_key):
"""An empty result list returns a structured error, matching ddg_search convention."""
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_search_response([])
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "no results"}))
assert parsed["error"] == "No results found"
assert parsed["query"] == "no results"
def test_unexpected_results_type_returns_error_json(self, mock_config_with_key):
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_response({"results": "nope"})
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["error"] == "Sofya returned an unexpected response format"
def test_http_error_returns_error_json(self, mock_config_with_key):
request = httpx.Request("POST", "https://sofya.co/v1/search")
response = httpx.Response(402, text="Insufficient credits", request=request)
mock_resp = MagicMock()
mock_resp.raise_for_status.side_effect = httpx.HTTPStatusError("error", request=request, response=response)
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = mock_resp
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["error"] == "Sofya API error: HTTP 402"
def test_network_error_returns_error_json(self, mock_config_with_key):
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.side_effect = httpx.ConnectError("boom")
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["error"] == "boom"
def test_missing_key_returns_error_json(self, mock_config_no_key):
with patch.dict("os.environ", {}, clear=True):
from deerflow.community.sofya.tools import web_search_tool
parsed = json.loads(web_search_tool.invoke({"query": "test"}))
assert parsed["error"] == "SOFYA_API_KEY is not configured"
assert parsed["query"] == "test"
class TestWebFetchTool:
def test_returns_title_and_content(self, mock_config_with_key):
results = [{"title": "Example Page", "url": "https://example.com", "content": "# Markdown body", "success": True}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_post = mock_client_cls.return_value.__enter__.return_value.post
mock_post.return_value = _make_fetch_response(results)
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "# Example Page\n\n# Markdown body"
assert mock_post.call_args.kwargs["json"] == {"urls": ["https://example.com"]}
def test_truncates_long_content(self, mock_config_with_key):
results = [{"title": "Long", "url": "https://example.com", "content": "x" * 9000, "success": True}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_fetch_response(results)
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert len(result) == len("# Long\n\n") + 4096
def test_non_string_content_does_not_raise(self, mock_config_with_key):
results = [{"title": "Numeric", "url": "https://example.com", "content": 12345, "success": True}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_fetch_response(results)
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "# Numeric\n\n12345"
def test_missing_content_still_reports_no_content(self, mock_config_with_key):
results = [{"title": "Empty", "url": "https://example.com", "content": None, "success": True}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_fetch_response(results)
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "Error: No content found"
def test_falls_back_to_untitled(self, mock_config_with_key):
results = [{"title": "", "url": "https://example.com", "content": "Body", "success": True}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_fetch_response(results)
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "# Untitled\n\nBody"
def test_failed_result_returns_its_error(self, mock_config_with_key):
results = [{"url": "https://example.com", "success": False, "error": "404 Not Found"}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_fetch_response(results)
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "Error: 404 Not Found"
def test_empty_content_returns_error(self, mock_config_with_key):
results = [{"title": "Empty", "url": "https://example.com", "content": "", "success": True}]
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_fetch_response(results)
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "Error: No content found"
def test_no_results_returns_error(self, mock_config_with_key):
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = _make_fetch_response([])
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "Error: No results found"
def test_http_error_returns_error_string(self, mock_config_with_key):
request = httpx.Request("POST", "https://sofya.co/v1/fetch")
response = httpx.Response(401, text="Invalid API key", request=request)
mock_resp = MagicMock()
mock_resp.raise_for_status.side_effect = httpx.HTTPStatusError("error", request=request, response=response)
with patch("deerflow.community.sofya.tools.httpx.Client") as mock_client_cls:
mock_client_cls.return_value.__enter__.return_value.post.return_value = mock_resp
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "Error: Sofya API error: HTTP 401"
def test_missing_key_returns_error_string(self, mock_config_no_key):
with patch.dict("os.environ", {}, clear=True):
from deerflow.community.sofya.tools import web_fetch_tool
result = web_fetch_tool.invoke({"url": "https://example.com"})
assert result == "Error: SOFYA_API_KEY is not configured"

View File

@ -6,6 +6,7 @@ from langchain_core.utils.function_calling import convert_to_openai_tool
from deerflow.community.brave.tools import web_search_tool as brave_web_search
from deerflow.community.ddg_search.tools import web_search_tool as ddg_web_search
from deerflow.community.searxng.tools import web_search_tool as searxng_web_search
from deerflow.community.sofya.tools import web_search_tool as sofya_web_search
from deerflow.community.tavily.tools import web_search_tool as tavily_web_search
EXPECTED_TIME_RANGES = {"day", "week", "month", "year"}
@ -13,8 +14,8 @@ EXPECTED_TIME_RANGES = {"day", "week", "month", "year"}
@pytest.mark.parametrize(
"tool_obj",
[ddg_web_search, brave_web_search, tavily_web_search, searxng_web_search],
ids=["ddg", "brave", "tavily", "searxng"],
[ddg_web_search, brave_web_search, tavily_web_search, searxng_web_search, sofya_web_search],
ids=["ddg", "brave", "tavily", "searxng", "sofya"],
)
def test_web_search_time_range_schema_is_consistent(tool_obj) -> None:
parameters = convert_to_openai_tool(tool_obj)["function"]["parameters"]

View File

@ -882,6 +882,19 @@ tools:
# # api_key: $CRW_API_KEY
# # base_url: https://fastcrw.com/api # default cloud; set to e.g. http://localhost:3000 for self-host
# Web search tool (uses Sofya, requires SOFYA_API_KEY)
# Sofya returns the content of the result pages, not only their snippets.
# Sign up at https://sofya.co to get a key.
# Note: set SOFYA_API_KEY in your environment before starting the app.
# Avoid putting literal API keys in config.yaml; use the $VAR form instead.
# - name: web_search
# group: web
# use: deerflow.community.sofya.tools:web_search_tool
# max_results: 5 # capped at 20 by the Sofya provider
# # search_depth: basic # basic reads the result pages, snippets returns snippets only
# contents_max_characters: 2000 # per-result content cap; 0 disables it
# # api_key: $SOFYA_API_KEY # Optional explicit env-var reference
# Web fetch tool (uses Browserless - headless Chrome, self-hosted or cloud)
# Browserless renders pages with a real headless Chrome, ideal for JavaScript-heavy
# sites and SPAs. Deploy your own: https://github.com/browserless/browserless
@ -1058,6 +1071,15 @@ tools:
# # base_url: https://fastcrw.com/api # default cloud; set to e.g. http://localhost:3000 for self-host
# # allow_private_addresses: false # SSRF guard: keep false in production. Set true ONLY for intentional internal targets.
# Web fetch tool (uses Sofya, requires SOFYA_API_KEY)
# Returns the page as markdown. Also handles PDF, DOCX and other documents.
# NOTE: Only one web_fetch provider can be active at a time.
# Comment out the Jina AI web_fetch entry above before enabling this one.
# - name: web_fetch
# group: web
# use: deerflow.community.sofya.tools:web_fetch_tool
# # api_key: $SOFYA_API_KEY
# Image search tool (uses DuckDuckGo)
# Use this to find reference images before image generation
- name: image_search

View File

@ -113,7 +113,7 @@ Community tools connect the agent to external services. They are configured in `
### Web search
<Tabs items={["DuckDuckGo (default)", "Tavily", "Brave", "Serply", "Exa", "InfoQuest", "Firecrawl", "GroundRoute"]}>
<Tabs items={["DuckDuckGo (default)", "Tavily", "Brave", "Serply", "Exa", "InfoQuest", "Firecrawl", "GroundRoute", "Sofya"]}>
<Tabs.Tab>
```yaml
tools:
@ -198,11 +198,19 @@ Requires a [GroundRoute](https://groundroute.ai) API key. Also supports `web_fet
Install: `cd backend && uv add 'deerflow-harness[groundroute]'`
</Tabs.Tab>
<Tabs.Tab>
```yaml
tools:
- use: deerflow.community.sofya.tools:web_search_tool
api_key: $SOFYA_API_KEY
```
Results include the content of each page, not only a snippet. Requires a [Sofya](https://sofya.co) API key. `max_results` is capped at 20, and each result's content is capped at `contents_max_characters` (default 2000, set 0 for no cap). Also supports `web_fetch`.
</Tabs.Tab>
</Tabs>
### Web fetch (page content extraction)
<Tabs items={["Jina AI (default)", "Browserless", "Exa", "InfoQuest", "Firecrawl", "GroundRoute"]}>
<Tabs items={["Jina AI (default)", "Browserless", "Exa", "InfoQuest", "Firecrawl", "GroundRoute", "Sofya"]}>
<Tabs.Tab>
```yaml
tools:
@ -251,6 +259,14 @@ tools:
api_key: $GROUNDROUTE_API_KEY
```
</Tabs.Tab>
<Tabs.Tab>
```yaml
tools:
- use: deerflow.community.sofya.tools:web_fetch_tool
api_key: $SOFYA_API_KEY
```
Returns the page as markdown. Requires a [Sofya](https://sofya.co) API key.
</Tabs.Tab>
</Tabs>
### Web capture (rendered screenshots)

View File

@ -110,7 +110,7 @@ tools:
### 网络搜索
<Tabs items={["DuckDuckGo默认", "Tavily", "Brave", "Serply", "Exa", "InfoQuest", "Firecrawl", "GroundRoute"]}>
<Tabs items={["DuckDuckGo默认", "Tavily", "Brave", "Serply", "Exa", "InfoQuest", "Firecrawl", "GroundRoute", "Sofya"]}>
<Tabs.Tab>
```yaml
tools:
@ -192,11 +192,19 @@ GroundRoute 是一个元搜索层:一个 API 接入六个搜索引擎Serper
安装:`cd backend && uv add 'deerflow-harness[groundroute]'`
</Tabs.Tab>
<Tabs.Tab>
```yaml
tools:
- use: deerflow.community.sofya.tools:web_search_tool
api_key: $SOFYA_API_KEY
```
搜索结果包含每个页面的正文内容,而不仅是摘要。需要 [Sofya](https://sofya.co) API Key`max_results` 上限为 20每条结果的正文长度上限由 `contents_max_characters` 控制(默认 2000设为 0 则不限制)。同时支持 `web_fetch`。
</Tabs.Tab>
</Tabs>
### 网页内容抓取
<Tabs items={["Jina AI默认", "Browserless", "Exa", "GroundRoute"]}>
<Tabs items={["Jina AI默认", "Browserless", "Exa", "InfoQuest", "Firecrawl", "GroundRoute", "Sofya"]}>
<Tabs.Tab>
```yaml
tools:
@ -225,11 +233,33 @@ tools:
</Tabs.Tab>
<Tabs.Tab>
```yaml
tools:
- use: deerflow.community.infoquest.tools:web_fetch_tool
api_key: $INFOQUEST_API_KEY
```
</Tabs.Tab>
<Tabs.Tab>
```yaml
tools:
- use: deerflow.community.firecrawl.tools:web_fetch_tool
api_key: $FIRECRAWL_API_KEY
```
</Tabs.Tab>
<Tabs.Tab>
```yaml
tools:
- use: deerflow.community.groundroute.tools:web_fetch_tool
api_key: $GROUNDROUTE_API_KEY
```
</Tabs.Tab>
<Tabs.Tab>
```yaml
tools:
- use: deerflow.community.sofya.tools:web_fetch_tool
api_key: $SOFYA_API_KEY
```
以 Markdown 形式返回页面内容。需要 [Sofya](https://sofya.co) API Key。
</Tabs.Tab>
</Tabs>
### 网页截图

View File

@ -497,6 +497,7 @@ def check_web_tool(config_path: Path, *, tool_name: str, label: str) -> CheckRes
"brave": "BRAVE_SEARCH_API_KEY",
"serper": "SERPER_API_KEY",
"serply": "SERPLY_API_KEY",
"sofya": "SOFYA_API_KEY",
"tencent_wsa": "TENCENTCLOUD_WSA_APIKEY",
},
"web_fetch": {
@ -504,6 +505,7 @@ def check_web_tool(config_path: Path, *, tool_name: str, label: str) -> CheckRes
"exa": "EXA_API_KEY",
"firecrawl": "FIRECRAWL_API_KEY",
"fastcrw": "CRW_API_KEY",
"sofya": "SOFYA_API_KEY",
},
"image_search": {
"brave": "BRAVE_SEARCH_API_KEY",

View File

@ -655,6 +655,14 @@ SEARCH_PROVIDERS: list[SearchProvider] = [
env_var="SERPLY_API_KEY",
extra_config={"max_results": 5},
),
SearchProvider(
name="sofya",
display_name="Sofya",
description="Search results include page content, API key required",
use="deerflow.community.sofya.tools:web_search_tool",
env_var="SOFYA_API_KEY",
extra_config={"max_results": 5, "contents_max_characters": 2000},
),
SearchProvider(
name="groundroute",
display_name="GroundRoute",
@ -716,6 +724,14 @@ WEB_FETCH_PROVIDERS: list[WebProvider] = [
env_var="CRW_API_KEY",
tool_name="web_fetch",
),
WebProvider(
name="sofya",
display_name="Sofya",
description="Page as markdown, API key required",
use="deerflow.community.sofya.tools:web_fetch_tool",
env_var="SOFYA_API_KEY",
tool_name="web_fetch",
),
WebProvider(
name="crawl4ai",
display_name="Crawl4AI",