mirror of
https://github.com/bytedance/deer-flow.git
synced 2026-07-28 17:06:05 +00:00
* feat(uploads): lazy-load historical files via list_uploaded_files tool
Replace per-turn injection of all historical upload metadata with on-demand
discovery via a new `list_uploaded_files` built-in tool, following the same
deferred-discovery pattern used by skills.
- Rename <uploaded_files> block to <current_uploads> (current-run files only)
- Add list_uploaded_files tool with include_outline: bool|list[str]
- Extract outline helpers to shared deerflow/utils/file_outline.py
- Update system prompt to reflect lazy-loading behaviour
- Historical file scan removed from UploadsMiddleware.before_agent()
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(uploads): clear uploaded_files state when no new files in current turn
When before_agent() returns None on empty turns, the LastValue
uploaded_files field retains the previous turn's filenames.
list_uploaded_files then incorrectly excludes those files as
"current-run" files, making them invisible until the next upload.
Fix: return {"uploaded_files": []} instead of None to explicitly
clear state. Add two-turn regression test covering the exact
scenario from review feedback.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: resolve CI lint errors and stale test assertion from merge
- Split long prompt line to fit 240-char limit
- Add missing `Any` import in list_uploaded_files_tool
- Remove unused `re` import in file_conversion (outline code moved)
- Remove unused `os` import in middleware test
- Fix test assertion: <uploaded_files> → <current_uploads> after main merge
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: resolve CI lint errors and stale test assertion from merge
- Split long prompt line to fit 240-char limit
- Add missing `Any` import in list_uploaded_files_tool
- Remove unused `re` import in file_conversion (outline code moved)
- Remove unused `os` import in middleware test
- Fix test assertion: <uploaded_files> → <current_uploads> after main merge
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: add current_uploads to input sanitization exempt tags
The lazy-loading PR renamed <uploaded_files> to <current_uploads>.
The anti-drift guard scans all framework XML blocks and requires each
to be either blocked or explicitly exempted. current_uploads wraps
trusted server-generated file metadata, not user input, so it belongs
in the exempt set.
Co-Authored-By: Claude <noreply@anthropic.com>
* test: regenerate replay golden after uploaded_files state change
before_agent now returns {"uploaded_files": []} instead of None,
adding uploaded_files to SSE values events. Regenerated via
DEERFLOW_WRITE_GOLDEN=1.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: review feedback — memory pipeline, stale tags, state clearing, nits
- Match both tags in memory stripping pipeline (uploaded_files|current_uploads)
- Remove stale uploaded_files from _BLOCKED_TAG_NAMES
- Clear uploaded_files on all before_agent early-return paths
- Fix ponytail: stray word in file_conversion re-export comment
- Remove dead total_omitted branch in _format_omitted_summary
- ruff format fixes
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: block current_uploads, sanitize only original user content
Per review feedback: instead of exempting <current_uploads> (which
allows user forgery), move it to _BLOCKED_TAG_NAMES and change
InputSanitizationMiddleware._process_request to scan only the
original user content (ORIGINAL_USER_CONTENT_KEY) when available.
Server-injected trusted blocks are no longer checked against the
blocked-tag denylist.
Co-Authored-By: Claude <noreply@anthropic.com>
* docs: clarify fallback reason in input sanitization comment
Co-Authored-By: Claude <noreply@anthropic.com>
* @
fix: third-round review feedback — state visibility, sanitization, regex, nits
- list_uploaded_files_tool: logger.warning instead of silent try/except
on runtime.state read failure (High)
- input_sanitization_middleware: _extract_text_from_content skips empty
text blocks to match message_content_to_text behaviour; rfind fallback
path logs warning for observability (Medium)
- memory pipeline regexes: backreference (?P<tag>)(?P=tag) in
message_processing.py and prompt.py (Low)
- file_conversion.py: re-export moved to top of file (Low)
- Tests: middleware→tool state bridge test; integrated forged-tag +
multimodal sanitization tests
PR #4174 — Follow-up issues: #4212, #4213, #4214
Co-Authored-By: Claude <noreply@anthropic.com>
@
* @
fix: 4th-round review — denylist, sanitization, scandir, nits
- Add "uploaded_files" back to _BLOCKED_TAG_NAMES (old tag still processed by
deermem; user forgery must be escaped) (consistency)
- Fix inaccurate rfind-fallback comment: UploadsMiddleware keeps string as
string, fallback is unreachable for strings (doc fix)
- Distinguish "empty string key" (upload without text) from "non-string key"
(caller forgery) so empty-text uploads never escape the server block (edge)
- Merge dual os.scandir(uploads_dir) calls into one list re-use (minor)
- Add comment on .md sibling skip known limitation: user-uploaded .md files
whose stem collides with a converted doc are hidden (boundary, no code change)
Co-Authored-By: Claude <noreply@anthropic.com>
@
* @
fix: tighten rfind-failure fallback — distinguish server blocks from user blocks
When _extract_text_from_content and message_content_to_text disagree on
multimodal list content and rfind fails, use content[0] (server-injected
<current_uploads> block) vs content[1:] (user blocks) to sanitize only
user blocks. Raw strings and non-standard dict blocks that
_extract_text_from_content misses are now also sanitized.
Non-distinguishable paths (< 2 text blocks, non-list content) still
degrade to full sanitization (safe — server block may be escaped but
user forgery never leaks). All fallback paths log via logger.warning.
Decision 18 / willem-bd 4th-round comment #3
Co-Authored-By: Claude <noreply@anthropic.com>
@
* @
fix: correct comments referencing text_blocks → content in rfind fallback
Co-Authored-By: Claude <noreply@anthropic.com>
@
* fix: 5th-round review — dead code, subagent gating, integration test, perf, consistency
- Delete unreachable ORIGINAL_USER_CONTENT_KEY guard in rfind fallback
branch (original_user_content guaranteed non-empty str at that point)
- Remove list_uploaded_files from BUILTIN_TOOLS; add include_upload_tool
param to get_available_tools(), default True; task_tool.py passes False
so subagents no longer receive a tool whose state exclusion is broken
- Add integration test exercising real create_agent graph (not mocked
runtime.state) to verify LangGraph propagates before_agent state writes
into ToolRuntime.state during same-turn tool calls
- Cache DirEntry.stat() st_size in candidates tuple to avoid second
per-file syscall in the rendering loop
- Make the upload-tag pre-check case-insensitive (content_str.lower())
to match _UPLOAD_BLOCK_RE re.IGNORECASE
PR #4174 — willem-bd 5th-round review items #1-#5
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(channels): pass files metadata through _human_input_message() for IM uploads
_human_input_message() was not passing additional_kwargs.files to the
downstream message. UploadsMiddleware read no files, wrote
uploaded_files=[], and list_uploaded_files reported same-run IM
attachments as historical files (fancyboi999 repro).
Fix: add files parameter to _human_input_message(), call site passes
files=uploaded. Regression test locks the contract.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(channels): remove legacy <uploaded_files> manual prepend to fix double-injection regression
Commit 8d86dbf6 added files= pass-through to UploadsMiddleware but
left the manual _format_uploaded_files_block() prepend in place.
Every IM attachment reached the model twice — once via the legacy
<uploaded_files> block and once via <current_uploads>.
This commit removes the manual prepend and the now-dead
_format_uploaded_files_block() function. UploadsMiddleware is the
sole upload-context producer for both IM and web paths.
Reported-by: fancyboi999 (PR review)
Co-Authored-By: Claude <noreply@anthropic.com>
* docs: update #4212 issue body to reflect completed fixes and narrowed remaining scope
* chore: remove temporary scratch file
* fix(middleware): neutralize user-derived values inside <current_uploads> block
Upload-derived filenames, paths, outline titles, and preview text are
interpolated verbatim inside the trusted <current_uploads> wrapper,
which InputSanitizationMiddleware exempts from sanitization. A crafted
filename or document heading containing blocked authority tags would
bypass the guardrail and enter model context as trusted framework data.
Fix: call neutralize_untrusted_tags() on all four user-derived values
inside _format_file_entry(), preserving the outer <current_uploads>
wrapper untouched.
Reported-by: fancyboi999 (P1 security review)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(middleware): neutralize extension labels in omitted-file summary
Files exceeding the 10-item context cap bypass _format_file_entry().
Their extensions, derived from user-controlled filenames via
_extension_label(), were interpolated verbatim into the trusted
<current_uploads> wrapper — another path for blocked authority tags
to escape the guardrail.
Fix: neutralize extension values inside _extension_label(), the
single extraction point for all extension labels.
Reported-by: fancyboi999 (P1 security review)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(tools): neutralize user-derived values in list_uploaded_files tool result
Apply neutralize_untrusted_tags() to every model-visible user-derived value
returned by list_uploaded_files: filename, virtual path, extension, outline
titles, outline preview lines, and omitted-file extension summary.
This closes the last remaining injection bypass in the upload lazy-loading
path - the <current_uploads> block and its omitted summary were already
neutralized (previous commits), but the list_uploaded_files tool produced
a second exit for the same attacker-controlled metadata that
ToolResultSanitizationMiddleware did not cover.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(tests): add missing include_upload_tool=False to task_tool mock assertions
PR #4174 added include_upload_tool parameter to get_available_tools().
task_tool.py correctly passes include_upload_tool=False for subagents
but 5 existing tests' assert_called_once_with expectations were not
updated, causing CI failures.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
218 lines
7.7 KiB
Python
218 lines
7.7 KiB
Python
"""File conversion utilities.
|
|
|
|
Converts document files (PDF, PPT, Excel, Word) to Markdown.
|
|
|
|
PDF conversion strategy (auto mode):
|
|
1. Try pymupdf4llm if installed — better heading detection, faster on most files.
|
|
2. If output is suspiciously short (< _MIN_CHARS_PER_PAGE chars/page, or < 200 chars
|
|
total when page count is unavailable), treat as image-based and fall back to MarkItDown.
|
|
3. If pymupdf4llm is not installed, use MarkItDown directly (existing behaviour).
|
|
|
|
Large files (> ASYNC_THRESHOLD_BYTES) are converted in a thread pool via
|
|
asyncio.to_thread() to avoid blocking the event loop (fixes #1569).
|
|
|
|
No FastAPI or HTTP dependencies — pure utility functions.
|
|
"""
|
|
|
|
import asyncio
|
|
import logging
|
|
from pathlib import Path
|
|
|
|
from deerflow.config.app_config import get_app_config
|
|
|
|
# Backward-compat re-exports — outline extraction moved to file_outline.py.
|
|
from deerflow.utils.file_outline import ( # noqa: F401
|
|
MAX_OUTLINE_ENTRIES,
|
|
extract_outline,
|
|
)
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
# File extensions that should be converted to markdown
|
|
CONVERTIBLE_EXTENSIONS = {
|
|
".pdf",
|
|
".ppt",
|
|
".pptx",
|
|
".xls",
|
|
".xlsx",
|
|
".doc",
|
|
".docx",
|
|
}
|
|
|
|
# Files larger than this threshold are converted in a background thread.
|
|
# Small files complete in < 1s synchronously; spawning a thread adds unnecessary
|
|
# scheduling overhead for them.
|
|
_ASYNC_THRESHOLD_BYTES = 1 * 1024 * 1024 # 1 MB
|
|
|
|
# If pymupdf4llm produces fewer characters *per page* than this threshold,
|
|
# the PDF is likely image-based or encrypted — fall back to MarkItDown.
|
|
# Rationale: normal text PDFs yield 200-2000 chars/page; image-based PDFs
|
|
# yield close to 0. 50 chars/page gives a wide safety margin.
|
|
# Falls back to absolute 200-char check when page count is unavailable.
|
|
_MIN_CHARS_PER_PAGE = 50
|
|
|
|
|
|
def _pymupdf_output_too_sparse(text: str, file_path: Path) -> bool:
|
|
"""Return True if pymupdf4llm output is suspiciously short (image-based PDF).
|
|
|
|
Uses chars-per-page rather than an absolute threshold so that both short
|
|
documents (few pages, few chars) and long documents (many pages, many chars)
|
|
are handled correctly.
|
|
"""
|
|
chars = len(text.strip())
|
|
doc = None
|
|
pages: int | None = None
|
|
try:
|
|
import pymupdf
|
|
|
|
doc = pymupdf.open(str(file_path))
|
|
pages = len(doc)
|
|
except Exception:
|
|
pass
|
|
finally:
|
|
if doc is not None:
|
|
try:
|
|
doc.close()
|
|
except Exception:
|
|
pass
|
|
if pages is not None and pages > 0:
|
|
return (chars / pages) < _MIN_CHARS_PER_PAGE
|
|
# Fallback: absolute threshold when page count is unavailable
|
|
return chars < 200
|
|
|
|
|
|
def _convert_pdf_with_pymupdf4llm(file_path: Path) -> str | None:
|
|
"""Attempt PDF conversion with pymupdf4llm.
|
|
|
|
Returns the markdown text, or None if pymupdf4llm is not installed or
|
|
if conversion fails (e.g. encrypted/corrupt PDF).
|
|
"""
|
|
try:
|
|
import pymupdf4llm
|
|
except ImportError:
|
|
return None
|
|
|
|
try:
|
|
return pymupdf4llm.to_markdown(str(file_path))
|
|
except Exception:
|
|
logger.exception("pymupdf4llm failed to convert %s; falling back to MarkItDown", file_path.name)
|
|
return None
|
|
|
|
|
|
def _convert_with_markitdown(file_path: Path) -> str:
|
|
"""Convert any supported file to markdown text using MarkItDown."""
|
|
from markitdown import MarkItDown
|
|
|
|
md = MarkItDown()
|
|
return md.convert(str(file_path)).text_content
|
|
|
|
|
|
def _do_convert(file_path: Path, pdf_converter: str) -> str:
|
|
"""Synchronous conversion — called directly or via asyncio.to_thread.
|
|
|
|
Args:
|
|
file_path: Path to the file.
|
|
pdf_converter: "auto" | "pymupdf4llm" | "markitdown"
|
|
"""
|
|
is_pdf = file_path.suffix.lower() == ".pdf"
|
|
|
|
if is_pdf and pdf_converter != "markitdown":
|
|
# Try pymupdf4llm first (auto or explicit)
|
|
pymupdf_text = _convert_pdf_with_pymupdf4llm(file_path)
|
|
|
|
if pymupdf_text is not None:
|
|
# pymupdf4llm is installed
|
|
if pdf_converter == "pymupdf4llm":
|
|
# Explicit — use as-is regardless of output length
|
|
return pymupdf_text
|
|
# auto mode: fall back if output looks like a failed parse.
|
|
# Use chars-per-page to distinguish image-based PDFs (near 0) from
|
|
# legitimately short documents.
|
|
if not _pymupdf_output_too_sparse(pymupdf_text, file_path):
|
|
return pymupdf_text
|
|
logger.warning(
|
|
"pymupdf4llm produced only %d chars for %s (likely image-based PDF); falling back to MarkItDown",
|
|
len(pymupdf_text.strip()),
|
|
file_path.name,
|
|
)
|
|
# pymupdf4llm not installed or fallback triggered → use MarkItDown
|
|
|
|
return _convert_with_markitdown(file_path)
|
|
|
|
|
|
async def convert_file_to_markdown(file_path: Path, output_path: Path | None = None) -> Path | None:
|
|
"""Convert a supported document file to Markdown.
|
|
|
|
PDF files are handled with a two-converter strategy (see module docstring).
|
|
Large files (> 1 MB) are offloaded to a thread pool to avoid blocking the
|
|
event loop.
|
|
|
|
Args:
|
|
file_path: Path to the file to convert.
|
|
output_path: Optional destination for the generated ``.md`` file.
|
|
When omitted, writes to ``file_path`` with a ``.md`` suffix.
|
|
Callers that track per-request filename uniqueness should pass a
|
|
pre-claimed path so companion markdown cannot clobber other uploads.
|
|
|
|
Returns:
|
|
Path to the generated .md file, or None if conversion failed.
|
|
"""
|
|
try:
|
|
pdf_converter = _get_pdf_converter()
|
|
file_size = file_path.stat().st_size
|
|
|
|
if file_size > _ASYNC_THRESHOLD_BYTES:
|
|
text = await asyncio.to_thread(_do_convert, file_path, pdf_converter)
|
|
else:
|
|
text = _do_convert(file_path, pdf_converter)
|
|
|
|
md_path = output_path if output_path is not None else file_path.with_suffix(".md")
|
|
md_path.write_text(text, encoding="utf-8")
|
|
|
|
logger.info("Converted %s to markdown: %s (%d chars)", file_path.name, md_path.name, len(text))
|
|
return md_path
|
|
except Exception as e:
|
|
logger.error("Failed to convert %s to markdown: %s", file_path.name, e)
|
|
return None
|
|
|
|
|
|
# Regex for bold-only lines that look like section headings.
|
|
# Targets SEC filing structural headings that pymupdf4llm renders as **bold**
|
|
# rather than # Markdown headings (because they use same font size as body text,
|
|
# distinguished only by bold+caps formatting).
|
|
#
|
|
# Pattern requires ALL of:
|
|
# 1. Entire line is a single **...** block (no surrounding prose)
|
|
# 2. Starts with a recognised structural keyword:
|
|
# - ITEM / PART / SECTION (with optional number/letter after)
|
|
# - SCHEDULE, EXHIBIT, APPENDIX, ANNEX, CHAPTER
|
|
|
|
_ALLOWED_PDF_CONVERTERS = {"auto", "pymupdf4llm", "markitdown"}
|
|
|
|
|
|
def _get_uploads_config_value(key: str, default: object) -> object:
|
|
"""Read a value from the uploads config, supporting dict and attribute access."""
|
|
cfg = get_app_config()
|
|
uploads_cfg = getattr(cfg, "uploads", None)
|
|
if isinstance(uploads_cfg, dict):
|
|
return uploads_cfg.get(key, default)
|
|
return getattr(uploads_cfg, key, default)
|
|
|
|
|
|
def _get_pdf_converter() -> str:
|
|
"""Read pdf_converter setting from app config, defaulting to 'auto'.
|
|
|
|
Normalizes the value to lowercase and validates it against the allowed set
|
|
so that values like 'AUTO' or 'MarkItDown' from config.yaml don't silently
|
|
fall through to unexpected behaviour.
|
|
"""
|
|
try:
|
|
raw = str(_get_uploads_config_value("pdf_converter", "auto")).strip().lower()
|
|
if raw not in _ALLOWED_PDF_CONVERTERS:
|
|
logger.warning("Invalid pdf_converter value %r; falling back to 'auto'", raw)
|
|
return "auto"
|
|
return raw
|
|
except Exception:
|
|
pass
|
|
return "auto"
|