mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-09-09 21:49:37 +00:00
* feat(community): add Sofya web search provider Add a community provider backed by Sofya (https://sofya.co). Its search endpoint returns the content of the result pages, not only their snippets, and its fetch endpoint returns a page as markdown. Both are plain JSON over HTTP, so this needs no extra Python package (uses httpx, already a dependency). Changes: - backend/packages/harness/deerflow/community/sofya/__init__.py - backend/packages/harness/deerflow/community/sofya/tools.py Implements web_search_tool and web_fetch_tool using httpx. API key is read from the config.yaml `api_key` field or the SOFYA_API_KEY env var. Follows the same interface and output shape as the existing ddg_search and serper providers, including the max_results parameter with config override and the structured "No results found" error. - backend/tests/test_sofya_tools.py Unit tests covering API key resolution, config overrides, result mapping, time range, HTTP errors, empty results, and fetch failures. - config.example.yaml: add commented-out Sofya web_search and web_fetch examples alongside the other providers - .env.example: add SOFYA_API_KEY placeholder - backend/docs/CONFIGURATION.md: list Sofya under web_search, web_fetch and the environment variables * fix(sofya): honor caller max_results, validate search_depth, join time_range contract test - Caller-supplied max_results now wins; config is used only when the argument is omitted, matching GroundRoute. - search_depth is clamped to basic/snippets; an unsupported value logs a warning and falls back to basic. - Sofya added to the shared time_range schema contract test. * fix(sofya): cap per-result content so a search stays inline An unbounded search payload (up to 20 read pages) crossed the tool output budget middleware's externalize_min_chars threshold, which replaces the result list with a file reference. Cap each result's content at contents_max_characters (default 2000, 0 disables), matching Exa's config key. Five capped results stay under the 12000 char threshold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5 * fix(sofya): list Sofya in the recency contract, coerce non-string content _clip subscripted its input, so a non-string content or description from the API raised TypeError instead of degrading. Coerce to text first, the way _sofya_post and _response_results guard the shapes around it. Also add Sofya to the Web Search Recency section in backend/AGENTS.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5 * fix(sofya): coerce web_fetch content, list sofya in the tools guide, add changelog web_fetch sliced its content the same way web_search did before the last push: a truthy non-string from the API passed the falsiness guard and then raised TypeError. Reuse _clip, keeping the `or ""` so empty content still reports "No content found". Also add sofya to the community provider inventory in packages/harness/deerflow/tools/AGENTS.md and an [Unreleased] changelog entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5 * docs(zh): add the missing InfoQuest and Firecrawl web_fetch tabs The ZH web_fetch tab list named five providers where EN names seven. Both tabs mirror their EN counterparts, so the two locales list the same web_fetch providers again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016TZhyPNCX2GYyBPkvTgJV5 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
109 lines
5.0 KiB
Plaintext
109 lines
5.0 KiB
Plaintext
# Serper API Key (Google Search) - https://serper.dev
|
|
SERPER_API_KEY=your-serper-api-key
|
|
|
|
# Serply API Key (Google Search, News and Scholar) - https://serply.io
|
|
SERPLY_API_KEY=your-serply-api-key
|
|
|
|
# TAVILY API Key
|
|
TAVILY_API_KEY=your-tavily-api-key
|
|
|
|
# Jina API Key
|
|
JINA_API_KEY=your-jina-api-key
|
|
|
|
# InfoQuest API Key
|
|
INFOQUEST_API_KEY=your-infoquest-api-key
|
|
|
|
# Sofya API Key (web search and fetch) - https://sofya.co
|
|
SOFYA_API_KEY=your-sofya-api-key
|
|
|
|
# Browser CORS allowlist for split-origin or port-forwarded deployments (comma-separated exact origins).
|
|
# Leave unset when using the unified nginx endpoint, e.g. http://localhost:2026.
|
|
# GATEWAY_CORS_ORIGINS=http://localhost:3000,http://127.0.0.1:3000
|
|
|
|
# Host interface the Docker stack publishes its entry port on. Defaults to
|
|
# 127.0.0.1 (loopback only), matching the local-trusted-environment deployment
|
|
# model documented in README.md -- the agent can execute commands.
|
|
# Set 0.0.0.0 only when the host is protected by your own TLS/auth front door
|
|
# or firewall, and complete first-run setup before it becomes reachable.
|
|
# BIND_HOST=0.0.0.0
|
|
# PORT=2026
|
|
|
|
# Optional:
|
|
# FIRECRAWL_API_KEY=your-firecrawl-api-key
|
|
# VOLCENGINE_API_KEY=your-volcengine-api-key
|
|
# OPENAI_API_KEY=your-openai-api-key
|
|
# OPENVIKING_API_KEY=your-openviking-user-api-key
|
|
# GEMINI_API_KEY=your-gemini-api-key
|
|
# DEEPSEEK_API_KEY=your-deepseek-api-key
|
|
# NOVITA_API_KEY=your-novita-api-key # OpenAI-compatible, see https://novita.ai
|
|
# MINIMAX_API_KEY=your-minimax-api-key # OpenAI-compatible, see https://platform.minimax.io
|
|
# STEPFUN_API_KEY=your-stepfun-api-key # OpenAI-compatible, see https://platform.stepfun.com
|
|
# VLLM_API_KEY=your-vllm-api-key # OpenAI-compatible
|
|
|
|
# E2B cloud sandbox API key — required only when using E2BSandboxProvider.
|
|
# Sign up at https://e2b.dev/dashboard
|
|
# E2B_API_KEY=your-e2b-api-key
|
|
# FEISHU_APP_ID=your-feishu-app-id
|
|
# FEISHU_APP_SECRET=your-feishu-app-secret
|
|
|
|
# SLACK_BOT_TOKEN=your-slack-bot-token
|
|
# SLACK_APP_TOKEN=your-slack-app-token
|
|
# TELEGRAM_BOT_TOKEN=your-telegram-bot-token
|
|
# DISCORD_BOT_TOKEN=your-discord-bot-token
|
|
|
|
# Enable LangSmith to monitor and debug your LLM calls, agent runs, and tool executions.
|
|
# LANGSMITH_TRACING=true
|
|
# LANGSMITH_ENDPOINT=https://api.smith.langchain.com
|
|
# LANGSMITH_API_KEY=your-langsmith-api-key
|
|
# LANGSMITH_PROJECT=your-langsmith-project
|
|
|
|
# GitHub API Token
|
|
# GITHUB_TOKEN=your-github-token
|
|
|
|
# Database (only needed when config.yaml has database.backend: postgres)
|
|
# DATABASE_URL=postgresql://deerflow:password@localhost:5432/deerflow
|
|
#
|
|
# WECOM_BOT_ID=your-wecom-bot-id
|
|
# WECOM_BOT_SECRET=your-wecom-bot-secret
|
|
# DINGTALK_CLIENT_ID=your-dingtalk-client-id
|
|
# DINGTALK_CLIENT_SECRET=your-dingtalk-client-secret
|
|
|
|
# Set to "false" to disable Swagger UI, ReDoc, and OpenAPI schema in production
|
|
# GATEWAY_ENABLE_DOCS=false
|
|
|
|
# Shared internal Gateway auth token for multi-worker deployments.
|
|
# `make up` generates and persists this automatically; set it manually only
|
|
# when you run Gateway workers outside the bundled deploy script.
|
|
# DEER_FLOW_INTERNAL_AUTH_TOKEN=your-shared-internal-token
|
|
|
|
# Provisioner API key for sandbox authentication (required when using provisioner/K8s sandbox mode).
|
|
# The same value must be set on the provisioner container and in config.yaml sandbox.provisioner_api_key.
|
|
# Generate: openssl rand -hex 32
|
|
# PROVISIONER_API_KEY=your-provisioner-api-key
|
|
|
|
# ── Frontend SSR → Gateway wiring ─────────────────────────────────────────────
|
|
# The Next.js server uses these to reach the Gateway during SSR (auth checks,
|
|
# /api/* rewrites). They default to localhost values that match `make dev` and
|
|
# `make start`, so most local users do not need to set them.
|
|
#
|
|
# Override only when the Gateway is not on localhost:8001 (e.g. when the
|
|
# frontend and gateway run on different hosts, in containers with a service
|
|
# alias, or behind a different port). docker-compose already sets these.
|
|
# DEER_FLOW_INTERNAL_GATEWAY_BASE_URL=http://localhost:8001
|
|
# DEER_FLOW_TRUSTED_ORIGINS=http://localhost:3000,http://localhost:2026
|
|
|
|
# ── Claude Code / Codex CLI subscription as a model provider (optional) ───────
|
|
# If you configure a ClaudeChatModel / Codex model provider (or an ACP agent)
|
|
# that reuses your CLI subscription login, prefer passing a token via env over
|
|
# bind-mounting your whole ~/.claude / ~/.codex into the container. The Gateway
|
|
# credential loader reads these first, so no directory mount is needed.
|
|
# CLAUDE_CODE_CREDENTIALS_PATH points at a single .credentials.json (Claude)
|
|
# rather than the whole dir. docker-compose.cli-auth.yaml is the opt-in
|
|
# directory-mount fallback for adapters that need the full CLI config.
|
|
# ACP adapters often take their own env API key (e.g. ANTHROPIC_API_KEY) and
|
|
# need no mount at all — check the adapter's docs. See SECURITY.md.
|
|
# CLAUDE_CODE_OAUTH_TOKEN=your-claude-code-oauth-token
|
|
# ANTHROPIC_AUTH_TOKEN=your-anthropic-auth-token
|
|
# CLAUDE_CODE_CREDENTIALS_PATH=/path/to/.claude/.credentials.json
|
|
# CODEX_AUTH_PATH=/path/to/codex/auth.json
|