Daoyuan Li b565e6c0f0
fix(workspace-changes): classify a symlink replacing a file distinctly from deleted (#4170)
* fix(workspace-changes): classify a symlink replacing a file distinctly from deleted

scan_workspace_roots() skipped every symlinked path entirely
(host_file.is_symlink() -> continue), so the path was completely absent
from a snapshot instead of being recorded as a metadata-only stub the
way binary/large/sensitive-looking files already are. When an agent run
replaces a tracked file with a symlink (e.g. rm config.txt && ln -s
/some/other/path config.txt), the after-snapshot never contained that
path at all, so compare_snapshots()'s _status() saw after_file=None and
reported plain "deleted" -- silently hiding that the path is still alive
on disk, now as a symlink that can point anywhere on the host, including
outside the workspace root.

Add a "symlink" classification mirroring the existing binary/large/
sensitive pattern: scan_workspace_roots() now records a symlink as a
metadata-only FileSnapshot stub (symlink=True, symlink_target from
os.readlink(), lstat'd without ever following the link) instead of
omitting it. _status() reports "symlink_created" whenever a symlink
newly occupies a path that was not already a symlink (brand new or
replacing a prior file), so the security-relevant fact surfaces
distinctly instead of collapsing into "deleted". A symlink genuinely
removed with nothing replacing it is unchanged: still "deleted".

Verified against a real POSIX symlink (WSL; native Windows symlink
creation needs elevated privilege) driving the unmodified
scan_workspace_roots()/compare_snapshots() functions, and via a
patch-file revert/reapply cycle on this same fix to confirm the added
regression tests fail before and pass after.

* fix(workspace-changes): count symlink_created in the changed-file badge

getChangedFileCount summed only created + modified + deleted, so a run
whose only change was a symlink replacing a file (reported as the new
symlink_created status, not deleted) produced a count of 0 and
WorkspaceChangeBadge hid the badge entirely -- the exact scenario this
PR targets, with the opposite of the intended result.

Add symlink_created to the frontend WorkspaceChangeSummary interface
(already emitted by the backend) and include it in the count. The
per-file StatusIcon/statusLabel "modified" fallthrough is unaffected
and left as a follow-up, per review.

Regression test reverts cleanly to reproduce a count of 0 pre-fix.

* fix(workspace-changes): rank symlink_created in file sort and complete the type contract

sortWorkspaceChanges's statusRank had no entry for the new
symlink_created status, so statusRank[left.status] - statusRank[right.status]
evaluated to NaN for any comparison involving a symlink-created file,
violating Array#sort's ordering contract instead of producing a
deterministic order. The frontend WorkspaceChangeStatus and
DiffUnavailableReason unions, and the WorkspaceFileChange symlink
fields, also stayed narrower than what the backend now emits, so
TypeScript's satisfies Record<...> guard on statusRank could not catch
the gap.

Widen WorkspaceChangeStatus to include "symlink_created" and
DiffUnavailableReason to include "symlink", add the matching
symlink/symlink_target_before/symlink_target_after fields to
WorkspaceFileChange, and give symlink_created a rank alongside
modified in statusRank -- restoring the satisfies guard's ability to
catch a future unranked status. unavailableLabel now has an explicit
"symlink" branch (new symlinkUnavailable i18n string, en-US + zh-CN)
instead of falling through to the generic label.

Also fixes the failing e2e-tests CI check: the existing
workspace-changes.spec.ts mock summary predates the symlink_created
field, so getChangedFileCount computed 1 + 1 + 0 + undefined = NaN and
the badge rendered "Edited NaN files" instead of "Edited 2 files".
Added symlink_created: 0 to the mock to match the real backend
contract.

New sortWorkspaceChanges unit tests revert cleanly against the
unranked statusRank to reproduce the NaN-driven misordering. pnpm test
(626 tests), pnpm check, and pnpm format are clean, and the
previously-failing e2e spec plus the full e2e suite (94 tests) pass.
2026-07-21 10:22:55 +08:00

324 lines
9.0 KiB
Python

from __future__ import annotations
import fnmatch
import hashlib
import os
from codecs import BOM_UTF16_BE, BOM_UTF16_LE, getincrementaldecoder
from pathlib import Path
from .types import (
DiffUnavailableReason,
FileSnapshot,
WorkspaceChangeLimits,
WorkspaceRoot,
WorkspaceSnapshot,
)
EXCLUDED_DIR_NAMES = {
".git",
".hg",
".svn",
".cache",
".next",
".venv",
"__pycache__",
"build",
"dist",
"node_modules",
}
BINARY_EXTENSIONS = {
".7z",
".avif",
".bmp",
".class",
".db",
".dll",
".dmg",
".doc",
".docx",
".exe",
".gif",
".gz",
".ico",
".jar",
".jpeg",
".jpg",
".mov",
".mp3",
".mp4",
".o",
".pdf",
".png",
".pyc",
".so",
".tar",
".webp",
".xls",
".xlsx",
".zip",
}
SENSITIVE_PATH_PATTERNS = (
".env",
".env.*",
"*api_key*",
"*apikey*",
"*.key",
"*.pem",
"*credential*",
"*password*",
"*private_key*",
"*secret*",
"*token*",
)
SAMPLE_BYTES = 4096
_UTF16_BOMS = (BOM_UTF16_LE, BOM_UTF16_BE)
def is_sensitive_workspace_path(path: str) -> bool:
normalized = path.lower()
parts = [part.lower() for part in Path(path).parts]
basename = parts[-1] if parts else normalized
for pattern in SENSITIVE_PATH_PATTERNS:
if fnmatch.fnmatch(basename, pattern) or fnmatch.fnmatch(normalized, pattern):
return True
if any(fnmatch.fnmatch(part, pattern) for part in parts):
return True
return False
def scan_workspace_roots(
roots: list[WorkspaceRoot],
*,
limits: WorkspaceChangeLimits | None = None,
include_text: bool = True,
text_paths: set[str] | None = None,
text_cache_dir: Path | None = None,
) -> WorkspaceSnapshot:
resolved_limits = limits or WorkspaceChangeLimits()
cache_dir = Path(text_cache_dir) if text_cache_dir is not None else None
if cache_dir is not None:
cache_dir.mkdir(parents=True, exist_ok=True)
files: dict[str, FileSnapshot] = {}
scanned = 0
truncated = False
for root in roots:
if not root.host_path.exists():
continue
for dirpath, dirnames, filenames in os.walk(root.host_path, followlinks=False):
dirnames[:] = [dirname for dirname in dirnames if dirname not in EXCLUDED_DIR_NAMES and not (Path(dirpath) / dirname).is_symlink()]
for filename in sorted(filenames):
if scanned >= resolved_limits.max_scanned_files:
truncated = True
return WorkspaceSnapshot(
files=files,
truncated=truncated,
text_cache_dir=str(cache_dir) if cache_dir is not None else None,
)
host_file = Path(dirpath) / filename
if host_file.is_symlink():
# A symlink must never be followed for stat/content purposes: its
# target can point anywhere on the host (including outside the
# scanned root), so it is recorded as a metadata-only stub -
# mirroring how binary/large/sensitive-looking files are handled
# below - instead of being silently omitted from the snapshot.
symlink_snapshot = _snapshot_symlink(root, host_file)
if symlink_snapshot is not None:
files[symlink_snapshot.path] = symlink_snapshot
scanned += 1
continue
if not host_file.is_file():
continue
snapshot = _snapshot_file(
root,
host_file,
limits=resolved_limits,
include_text=include_text,
text_paths=text_paths,
text_cache_dir=cache_dir,
)
if snapshot is not None:
files[snapshot.path] = snapshot
scanned += 1
return WorkspaceSnapshot(
files=files,
truncated=truncated,
text_cache_dir=str(cache_dir) if cache_dir is not None else None,
)
def _snapshot_file(
root: WorkspaceRoot,
host_file: Path,
*,
limits: WorkspaceChangeLimits,
include_text: bool,
text_paths: set[str] | None,
text_cache_dir: Path | None,
) -> FileSnapshot | None:
try:
stat = host_file.stat()
size = stat.st_size
mtime_ns = stat.st_mtime_ns
relative = host_file.relative_to(root.host_path).as_posix()
virtual_path = f"{root.virtual_prefix}/{relative}"
sensitive = is_sensitive_workspace_path(virtual_path)
except OSError:
return None
if sensitive:
return FileSnapshot(
path=virtual_path,
root=root.name,
size=size,
mtime_ns=mtime_ns,
sha256=None,
binary=False,
sensitive=True,
text=None,
content_unavailable_reason="sensitive",
)
try:
sample = host_file.read_bytes()[:SAMPLE_BYTES] if size <= SAMPLE_BYTES else _read_sample(host_file)
except OSError:
return None
binary = host_file.suffix.lower() in BINARY_EXTENSIONS or _looks_binary(sample)
sha256 = _sha256_file(host_file) if size <= limits.max_file_bytes_for_diff else None
text: str | None = None
text_path: str | None = None
reason: DiffUnavailableReason | None = None
should_include_text = include_text and (text_paths is None or virtual_path in text_paths)
if binary:
reason = "binary"
elif size > limits.max_file_bytes_for_diff:
reason = "large"
elif not should_include_text:
text = None
else:
try:
raw = host_file.read_bytes()
except OSError:
return None
decoded = _decode_text_bytes(raw)
if decoded is None:
binary = True
reason = "binary"
elif text_cache_dir is not None:
text_path = str(_cache_text_file(decoded, virtual_path, text_cache_dir))
else:
text = decoded
return FileSnapshot(
path=virtual_path,
root=root.name,
size=size,
mtime_ns=mtime_ns,
sha256=sha256,
binary=binary,
sensitive=sensitive,
text=text,
text_path=text_path,
content_unavailable_reason=reason,
)
def _snapshot_symlink(root: WorkspaceRoot, host_file: Path) -> FileSnapshot | None:
# Deliberately never follows the link (no read_bytes()/open() on the target):
# the target may point anywhere on the host, including outside the scanned
# root, so stat'ing or reading through it here would risk exposing arbitrary
# host file content/metadata as if it belonged to the workspace.
try:
stat = host_file.lstat()
size = stat.st_size
mtime_ns = stat.st_mtime_ns
relative = host_file.relative_to(root.host_path).as_posix()
virtual_path = f"{root.virtual_prefix}/{relative}"
sensitive = is_sensitive_workspace_path(virtual_path)
except OSError:
return None
try:
target = os.readlink(host_file)
except OSError:
target = None
return FileSnapshot(
path=virtual_path,
root=root.name,
size=size,
mtime_ns=mtime_ns,
sha256=None,
binary=False,
sensitive=sensitive,
text=None,
content_unavailable_reason="symlink",
symlink=True,
symlink_target=target,
)
def _cache_text_file(text: str, virtual_path: str, cache_dir: Path) -> Path:
cache_name = hashlib.sha256(virtual_path.encode("utf-8")).hexdigest()
target = cache_dir / cache_name
target.write_text(text, encoding="utf-8")
return target
def _read_sample(path: Path) -> bytes:
with path.open("rb") as file:
return file.read(SAMPLE_BYTES)
def _sha256_file(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as file:
for chunk in iter(lambda: file.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def _decode_text_bytes(data: bytes) -> str | None:
for encoding in ("utf-8-sig", "utf-8"):
try:
return data.decode(encoding)
except UnicodeDecodeError:
continue
if data.startswith(_UTF16_BOMS):
try:
return data.decode("utf-16")
except UnicodeDecodeError:
return None
return None
def _sample_decodes_as_text(sample: bytes, encoding: str) -> bool:
try:
decoder = getincrementaldecoder(encoding)()
decoder.decode(sample, final=False)
except UnicodeDecodeError:
return False
return True
def _looks_binary(sample: bytes) -> bool:
if sample.startswith(_UTF16_BOMS) and _sample_decodes_as_text(sample, "utf-16"):
return False
if b"\x00" in sample:
return True
if _sample_decodes_as_text(sample, "utf-8"):
return False
return True