simon ea9b70148e
fix(clarification): drop sibling tool calls before interrupt (#4908)
* fix(clarification): drop sibling tool calls before interrupt

- Rewrite the AIMessage in ClarificationMiddleware.after_model so a
  parallel bash/write_file cannot run before the user answers
- langchain return_direct only inspects the last ToolMessage; siblings
  both execute and can keep the agent loop alive
- Skip the rewrite when disable_clarification is set
- Prompt and tool docs: do not call other tools in the same turn

Fixes #4906

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(clarification): enhance sibling tool call handling in ClarificationMiddleware

- Update ClarificationMiddleware to ensure sibling tool calls are dropped when `ask_clarification` is invoked, preventing unintended execution before user input.
- Modify documentation to clarify that the `return_direct` router now inspects all client-side tool calls of the last AIMessage, ensuring proper routing behavior.
- Introduce a new integration test to validate that sibling tools do not execute when `ask_clarification` is present in the same turn.

This change addresses potential issues with tool execution order and improves the overall reliability of the middleware.

Fixes #4906

* fix(clarification): enhance tool call filtering in ClarificationMiddleware

- Update _filter_content_tool_use to handle Gemini-style function_call blocks by matching on name when no id is present, ensuring proper filtering of tool calls.
- Modify ClarificationMiddleware to maintain sibling tool call integrity by dropping unnecessary blocks, improving the clarity of the AIMessage content.
- Add a new test to validate the correct stripping of idless function call content blocks, ensuring that sibling tool calls do not execute prematurely.

This change improves the robustness of the middleware and addresses potential execution order issues.

Fixes #4906

* fix(clarification): drop siblings when ask_clarification is malformed

LangChain parks invalid args on invalid_tool_calls independently, so a
valid sibling would otherwise still execute before the user answers.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-08-24 07:37:35 +08:00

95 lines
4.9 KiB
Python

from typing import Literal, Required, TypedDict
from langchain.tools import tool
class ClarificationFormField(TypedDict, total=False):
"""One form field definition for a structured clarification card.
The model-visible schema documents the item shape; runtime validation
still happens defensively in ``ClarificationMiddleware`` because the
middleware intercepts the call before tool execution.
"""
name: Required[str]
label: str
type: Literal["text", "textarea", "number", "select", "multi_select", "checkbox", "date"]
required: bool
options: list[str]
placeholder: str
@tool("ask_clarification", parse_docstring=True, return_direct=True)
def ask_clarification_tool(
question: str,
clarification_type: Literal[
"missing_info",
"ambiguous_requirement",
"approach_choice",
"risk_confirmation",
"suggestion",
],
context: str | None = None,
options: list[str] | None = None,
fields: list[ClarificationFormField] | None = None,
) -> str:
"""Ask the user for clarification when you need more information to proceed.
Use this tool when you encounter situations where you cannot proceed without user input:
- **Missing information**: Required details not provided (e.g., file paths, URLs, specific requirements)
- **Ambiguous requirements**: Multiple valid interpretations exist
- **Approach choices**: Several valid approaches exist and you need user preference
- **Risky operations**: Destructive actions that need explicit confirmation (e.g., deleting files, modifying production)
- **Suggestions**: You have a recommendation but want user approval before proceeding
The execution will be interrupted and the question will be presented to the user.
Wait for the user's response before continuing.
When to use ask_clarification:
- You need information that wasn't provided in the user's request
- The requirement can be interpreted in multiple ways
- Multiple valid implementation approaches exist
- You're about to perform a potentially dangerous operation
- You have a recommendation but need user approval
Choosing the interaction shape:
- One open question -> just `question` (free text input)
- Pick exactly one option -> `options`
- Pick several options -> a single `fields` entry of type `multi_select`
- Collect several values at once (e.g. a set of parameters for one action) ->
`fields`, which renders a single structured form instead of several
sequential questions. Prefer one form over asking field-by-field.
Best practices:
- Ask ONE clarification at a time for clarity; a form with several fields
still counts as one clarification
- Be specific and clear in your question
- Don't make assumptions when clarification is needed
- For risky operations, ALWAYS ask for confirmation
- If a skill provides a predefined field template, pass it through `fields`
unchanged instead of redesigning it
- After calling this tool, execution will be interrupted automatically
- Do not call any other tool in the same turn as this one; sibling tool
calls are dropped so they cannot run before the user answers
Args:
question: The clarification question to ask the user. Be specific and clear.
clarification_type: The type of clarification needed (missing_info, ambiguous_requirement, approach_choice, risk_confirmation, suggestion).
context: Optional context explaining why clarification is needed. Helps the user understand the situation.
options: Optional list of choices (for approach_choice or suggestion types). Present clear options for the user to choose from.
fields: Optional form field definitions for collecting multiple values in one card; takes precedence over `options`.
Each field is an object with `name` (unique identifier, required; avoid JavaScript prototype names like
`constructor` or `toString`), `label` (display text, defaults to name), `type` (one of: text, textarea,
number, select, multi_select, checkbox, date; defaults to text), `required` (boolean, defaults to false),
`options` (list of strings, required for select/multi_select types), and `placeholder` (optional hint text).
A `checkbox` field is a boolean that defaults to "no"; set `required` on a checkbox only for
must-agree/consent semantics (the user has to tick it to submit). Keep forms bounded: at most 16 fields,
24 options per field, and 200 characters per name/label/option/placeholder — exceeding a limit degrades
the whole request to a plain-text question.
"""
# This is a placeholder implementation
# The actual logic is handled by ClarificationMiddleware which intercepts this tool call
# and interrupts execution to present the question to the user
return "Clarification request processed by middleware"