- Introduced a new reference document for streaming output issues, detailing the differences between streaming APIs and providing solutions for common problems. - Created a structured output issues reference, outlining the use of `with_structured_output`, `response_format`, and schema enforcement strategies. - Added a user query convention guide to standardize structured questions for skill authors, including block types for queries and data gathering. - Implemented a template for a chat model, encapsulating API key management, request payload construction, and response handling. - Established a symlink for the LangChain dev guide in the Claude skills directory for easier access. - Initialized a skills lock file to manage dependencies and versions for the LangChain dev guide.
14 KiB
Other Common Development Issues
A collection of standalone but high-frequency issues that span multiple categories or sit at the basic engineering layer.
Issue 1: A tool needs to return extra data on top of the model-facing result
-
Symptom: A tool needs to return both "text for the model to read" and "extra data for the application side to consume" — e.g. a retrieval tool needs to return a passage (for the model) and the source document ID + page number (for the frontend to highlight); an order lookup needs to return an order summary (for the model) and the full order object (for the next tool / main agent state). Concatenating everything into one string lets the model get distracted by irrelevant fields; returning only text loses the structured data.
-
Cause: A tool's return value has three different semantics; mixing them causes problems:
- For the model to see: must fit into
ToolMessage.content; the model uses this to decide next steps. - For the application layer / downstream business code to see, but not into the LLM context: e.g. document ID, raw payload, render hints.
- For writing back to agent state, to be read by subsequent tools / middleware: e.g.
customer_id,last_order,current_step.
These three categories should go through three different channels; jamming them into one string both adds noise for the model and loses structured data for the application.
- For the model to see: must fit into
-
Solution: Choose the return style by data destination (three styles can be mixed):
- Model-only → return string or dict directly: dicts get serialized into
ToolMessage.content, and the model reads the fields itself.
from langchain.tools import tool @tool def get_weather_data(city: str) -> dict: """Get structured weather for a city.""" return {"city": city, "temperature_c": 22, "conditions": "sunny"}- One for the model + metadata for the app layer (not into LLM) →
ToolMessage(artifact=...):contententers the model context,artifactdoesn't enter the model but stays on theToolMessagefor downstream consumption (typical scenario: retrieval tool returning passage + document ID).
from langchain.messages import ToolMessage from langchain.tools import ToolRuntime, tool @tool def search_books(query: str, runtime: ToolRuntime) -> ToolMessage: """Retrieve a passage and attach source metadata.""" passage = "It was the best of times, it was the worst of times." return ToolMessage( content=passage, # enters model context tool_call_id=runtime.tool_call_id, name="search_books", artifact={"document_id": "doc_123", "page": 0}, # readable to app layer, not into LLM )Application code later reads structured data from
message.artifact; the model sees only thecontenttext.- Write data back to agent state for subsequent tool / middleware / main agent reuse → return
Command(update=...): in the update, place both business fields and aToolMessage(must contain a ToolMessage paired withtool_call_id, otherwise the next LLM call will fail with invalid message sequence because "tool_call has no tool_response").
from langchain.messages import ToolMessage from langchain.tools import ToolRuntime, tool from langgraph.types import Command @tool def lookup_customer(customer_id: str, runtime: ToolRuntime) -> Command: """Look up customer and write profile into agent state.""" profile = fetch_customer(customer_id) # {"name": ..., "tier": ..., ...} return Command(update={ "customer_profile": profile, # written into state, readable by subsequent tools "messages": [ToolMessage( content=f"Found customer {profile['name']} (tier: {profile['tier']})", tool_call_id=runtime.tool_call_id, # required to pair with tool_call )], })State fields must first be declared in
state_schema(or via anAgentStatesubclass), otherwise updates are ignored. - Model-only → return string or dict directly: dicts get serialized into
-
Lessons learned:
- Default to "return string/dict" — covers 80% of scenarios.
- For fields the application layer wants but would be noise in the LLM context (document ID, raw payload, render metadata), use
ToolMessage(artifact=...). - For "cross-tool / cross-turn reusable business data" (user profile, looked-up order, current stage), use
Command(update=...)— remember the pairedToolMessage, and declare fields in state schema first. - The three styles aren't mutually exclusive: a tool can simultaneously do
Command(update={"customer_profile": ..., "messages": [ToolMessage(content=..., artifact=...)]}), using all three channels at once.
Issue 2: MCP tool can't access the agent's user_id / API key / current state
-
Symptom: You wire an MCP server in as a regular LangChain tool and want to read
runtime.context.user_id, current agent state, or user preferences fromstoreinside the tool — only to find the MCP tool can't access any of these and is limited to the args declared in the schema. Addinguser_iddirectly to the tool schema makes the model fabricate one for every call, polluting context and being insecure. -
Cause: MCP servers run in a separate process (stdio subprocess or remote HTTP service), completely process-isolated from the LangGraph runtime — they can't see store, context, state, or tool_call_id. Importing LangGraph runtime APIs on the MCP server side is pointless because it's not running in that process.
-
Solution: Bridge on the client side with
tool_interceptors— interceptors run in the LangGraph process and can access the fullToolRuntime, injecting the needed fields intoargs/headersbefore forwarding to the MCP server.from langchain_mcp_adapters.client import MultiServerMCPClient from langchain_mcp_adapters.interceptors import MCPToolCallRequest async def inject_user_context(request: MCPToolCallRequest, handler): runtime = request.runtime # 1) Business fields (user_id, tenant_id) injected into args; model can't see them and doesn't need them in schema args = {**request.args, "user_id": runtime.context.user_id} # 2) Auth / tracing info injected into headers, doesn't pollute schema at all headers = {"Authorization": f"Bearer {runtime.context.api_key}"} return await handler(request.override(args=args, headers=headers)) client = MultiServerMCPClient({...}, tool_interceptors=[inject_user_context])For auth "short-circuit" scenarios, directly
return ToolMessage(...)without callinghandler— e.g. denying sensitive tool calls when unauthenticated:from langchain.messages import ToolMessage async def require_auth(request: MCPToolCallRequest, handler): if request.name in {"delete_file", "export_data"} and not request.runtime.state.get("authenticated"): return ToolMessage( content="Authentication required.", tool_call_id=request.runtime.tool_call_id, ) return await handler(request) -
Lessons learned:
- Business fields (user_id, tenant_id) → modify
args: interceptor injects them centrally; not in schema, so the model won't fabricate them. - Auth / tracing info → modify
headers: not in schema at all, the cleanest. structuredContentreturned by MCP server is invisible to the model by default (placed only inToolMessage.artifact) — to let the model read it, use the interceptor to serializestructuredContentand concatenate it back toresult.content.- Multiple interceptors follow onion order:
[outer, inner]→ outer enters first / exits last. When composing "auth + rate limit + retry", put the outer concern in front, inner (closest to the tool call) in back. MultiServerMCPClientis stateless by default — a new session opens for every tool call. When you need to reuse context across calls (e.g. server-side login state), useasync with client.session("server_name") as session:to explicitly manage the lifecycle.
- Business fields (user_id, tenant_id) → modify
Issue 3: Model always outputs invalid_tool_calls - tool never actually executes
-
Symptom: The model is bound with tools and consistently "calls" them, but the tool function never fires. Inspecting the
AIMessageshowstool_callsis empty whileinvalid_tool_callsis populated. The agent loop either silently skips the call or raises a parsing error. The weaker the model (small-parameter open-source, quantized, vLLM-served), the more frequent this becomes. -
Cause: When the model generates a tool call, the
argumentsfield must be valid JSON conforming to the tool's schema. Weak models often produce malformed JSON — missing quotes, trailing commas, unescaped characters, truncated output, etc. LangChain's tool-call parser cannot parse the broken JSON, so instead of placing it intool_callsit moves it toinvalid_tool_calls. Since the agent executor only processes entries intool_calls, the tool never runs — it looks like the model "called" it but nothing happened. -
Solution: Use
ToolCallRepairMiddlewarefromlangchain-dev-utils— it automatically detects entries ininvalid_tool_calls, attempts to repair the malformed JSON via thejson-repairlibrary, and promotes successfully repaired calls back totool_callsso the agent can execute them normally.from langchain_dev_utils.agents.middleware import tool_call_repair agent = create_agent( model="openai:gpt-5-mini", tools=[run_python_code, get_current_time], middleware=[tool_call_repair], )tool_call_repairis a pre-instantiated global instance ofToolCallRepairMiddleware— zero configuration needed.If you prefer explicit instantiation:
from langchain_dev_utils.agents.middleware import ToolCallRepairMiddleware agent = create_agent( model="openai:gpt-5-mini", tools=[run_python_code, get_current_time], middleware=[ToolCallRepairMiddleware()], ) -
Lessons learned:
- When a tool "isn't being called", check
invalid_tool_callson the AIMessage first — the model likely did attempt a call, but the JSON was unparseable. ToolCallRepairMiddlewarecannot guarantee 100% repair — severely garbled output (e.g. half the JSON is natural language) will still fail. For those cases, consider simplifying the tool schema, splitting complex parameters into multiple smaller tools, or upgrading to a stronger model.- This middleware only acts on
invalid_tool_calls— valid calls pass through untouched with zero overhead.
- When a tool "isn't being called", check
Issue 4: How to use placeholders in system prompt that get dynamically replaced at runtime
-
Symptom: You want the system prompt to include dynamic information — user name, role, current date, conversation context, etc. — that varies per request. Hardcoding these values means creating a new agent for every variation; concatenating strings manually is error-prone and hard to maintain.
-
Cause:
create_agenttreatssystem_promptas a static string by default — it does not perform any template interpolation. To get runtime substitution, you need an explicit formatting middleware that resolves placeholders againststateandcontextbefore the prompt reaches the model. -
Solution: Add the
format_promptmiddleware (f-string style, covers most cases) orFormatPromptMiddleware(template_format="jinja2")(for conditionals / loops):from langchain_dev_utils.agents.middleware import format_prompt from langchain.agents import AgentState from dataclasses import dataclass class AssistantState(AgentState): name: str @dataclass class UserContext: user: str agent = create_agent( model="openai:gpt-5", system_prompt="You are {name}, an assistant for {user}.", middleware=[format_prompt], state_schema=AssistantState, context_schema=UserContext, ) # At runtime, {name} is resolved from state, {user} from context response = agent.invoke( {"messages": [HumanMessage(content="Hello")], "name": "Jarvis"}, context=UserContext(user="Tony"), ) # Model receives: "You are Jarvis, an assistant for Tony."Variables are resolved in priority order:
statefirst, thencontext— astatefield with the same name shadows thecontextfield.For templates that need conditionals or loops, use Jinja2:
from langchain_dev_utils.agents.middleware import FormatPromptMiddleware from dataclasses import dataclass from typing import Optional @dataclass class Context: user_role: Optional[str] = None jinja2_formatter = FormatPromptMiddleware(template_format="jinja2") agent = create_agent( model="openai:gpt-5", system_prompt=( "You are an assistant.\n" "{% if user_role == 'VIP' %}Provide premium service.{% endif %}" ), middleware=[jinja2_formatter], context_schema=Context, ) -
Lessons learned:
- Template interpolation is opt-in via middleware — never assume it happens by default.
- Use
format_prompt(global instance, zero config) for simple{variable}substitution; only reach for Jinja2 when you need{% if %}/{% for %}. - If a placeholder has no matching field in either
stateorcontext, formatting will raise aKeyError— declare all variables in the corresponding schema. - Keep
system_prompthardcoded by the developer; pass only data throughstate/context. Never let end-user input become the template itself — this applies regardless of whether formatting middleware is enabled.