Add comprehensive documentation and templates for LangChain skill development
- Introduced a new reference document for streaming output issues, detailing the differences between streaming APIs and providing solutions for common problems. - Created a structured output issues reference, outlining the use of `with_structured_output`, `response_format`, and schema enforcement strategies. - Added a user query convention guide to standardize structured questions for skill authors, including block types for queries and data gathering. - Implemented a template for a chat model, encapsulating API key management, request payload construction, and response handling. - Established a symlink for the LangChain dev guide in the Claude skills directory for easier access. - Initialized a skills lock file to manage dependencies and versions for the LangChain dev guide.
This commit is contained in:
@@ -0,0 +1,72 @@
|
||||
---
|
||||
name: langchain-dev-guide
|
||||
description: "LangChain / LangGraph engineering pitfalls and verified fixes. Covers DeepAgents, structured output, OpenAI-compatible model integration (including Chinese provider adapters: DeepSeek, Qwen, GLM, etc.), middleware, streaming, multi-agent orchestration, and other common development issues. Use when hitting unexpected behavior, making architecture decisions, or integrating Chinese LLM providers during LangChain development."
|
||||
---
|
||||
|
||||
# LangChain Dev Guide
|
||||
|
||||
A systematic summary of typical issues, non-obvious behaviors, and verified solutions encountered in real engineering with the LangChain / LangGraph ecosystem. Every entry comes from a real development scenario and is organized by category.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> This skill is an **engineering practice reference**, not an introductory tutorial. Each entry assumes the developer is already familiar with basic LangChain concepts (agent, tool, message, graph).
|
||||
|
||||
## How to Use
|
||||
|
||||
1. First use the "Scenario Index" below to locate the category file your problem belongs to.
|
||||
2. When unsure which category applies, search keywords directly in the "Common Issues Quick Reference".
|
||||
3. For ContextSeek / semantic memory: start with [contextseek-middleware.md](reference/contextseek-middleware.md) to identify your scenario, then go to [contextseek-params.md](reference/contextseek-params.md) for specific parameter configuration issues.
|
||||
4. Once you find the relevant section, read it in depth — every entry follows the structure **Symptom → Cause → Solution → Lessons learned**.
|
||||
|
||||
## Scenario Index
|
||||
|
||||
| Category | File | Trigger Scenarios |
|
||||
|----------|------|-------------------|
|
||||
| Deep Agents | [reference/deepagents.md](reference/deepagents.md) | Model selection, filesystem backend, disabling the general-purpose sub-agent, file permissions, long-term memory, long `SKILL.md` truncated by `read_file` 100-line default |
|
||||
| Structured Output | [reference/structured-output.md](reference/structured-output.md) | Model-level method selection, `create_agent` strategies, missing fields, unsupported `tool_choice`, provider-side 400 errors on forced schema tool selection |
|
||||
| OpenAI-compatible Model Integration | [reference/model-integration.md](reference/model-integration.md) | Pitfalls when using `ChatOpenAI` against OpenAI-compatible providers, integrating Reasoning models (chain-of-thought / `reasoning_content`) |
|
||||
| CN Model Integration | [reference/cn-models/README.md](reference/cn-models/README.md) | Generating LangChain integration classes for Chinese providers (DeepSeek, Qwen, GLM, Moonshot) |
|
||||
| Middleware | [reference/middleware.md](reference/middleware.md) | Middleware execution order, `state_schema` merging, HITL `resume` values, modifying state from `wrap_model_call` |
|
||||
| Streaming Output | [reference/streaming.md](reference/streaming.md) | Choosing between `stream_events` and `stream`, distinguishing tokens from multiple LLMs, disabling streaming, custom progress events |
|
||||
| Multi-Agent Orchestration | [reference/multi-agent.md](reference/multi-agent.md) | subagents vs handoffs, tool-per-agent vs dispatch, retrieving subagent state, trimming subagent boilerplate, quickly building handoff setups |
|
||||
| Other Common Issues | [reference/common-issues.md](reference/common-issues.md) | High-frequency standalone issues that don't fit the categories above. Currently includes: tools returning data to both the model and the application layer, MCP tools unable to access runtime context, `invalid_tool_calls`, and dynamic system prompt placeholders |
|
||||
| ContextSeek — Use Case Scenarios | [reference/contextseek-middleware.md](reference/contextseek-middleware.md) | Agent loses context across sessions, tool call auditing, cross-topic knowledge discovery (dream), SRE provenance / confidence tracing, enterprise knowledge cold-start (DataPlug) |
|
||||
| ContextSeek — Parameter & Config Issues | [reference/contextseek-params.md](reference/contextseek-params.md) | scope isolation, auto_store / record_tool_calls write volume, auto_compact throttling and shutdown, retrieval_tags / min_score filtering, tool_arg_overrides, dream trigger conditions, dream item decay, evidence_chain vs chain_confidence, DataPlug vs ctx.add(), plug() scope priority, auto_dream dual-gate triggering |
|
||||
|
||||
## Common Issues Quick Reference
|
||||
|
||||
| Keyword / Error | Where to Look |
|
||||
|-----------------|---------------|
|
||||
| Which model to choose / Deep Agent performing poorly | deepagents issue 1 |
|
||||
| Filesystem backend / local files / file permissions | deepagents issues 2 / 4 |
|
||||
| Disabling the default sub-agent / general-purpose | deepagents issue 3 |
|
||||
| Long-term memory / store | deepagents issue 5 |
|
||||
| `SKILL.md` truncated / only first 100 lines read / `read_file` limit / progressive disclosure | deepagents issue 6 |
|
||||
| `with_structured_output` returning None / missing fields | structured-output issue 1 |
|
||||
| `create_agent` / `response_format` / `ProviderStrategy` / `ToolStrategy` | structured-output issue 1 |
|
||||
| `with_structured_output` / `function_calling` / `tool_choice` unsupported / `deepseek-reasoner does not support this tool_choice` | structured-output issue 2 |
|
||||
| OpenAI-compatible model / `ChatOpenAI` not working | model-integration issue 1 |
|
||||
| Reasoning model / `reasoning_content` / chain-of-thought lost | model-integration issue 2 |
|
||||
| Chinese model / CN provider / DeepSeek / Qwen / GLM / Moonshot | cn-models README |
|
||||
| `langchain-cn-models` (embedded) / generate integration class | cn-models README |
|
||||
| Middleware order messed up / before/after counterintuitive | middleware issue 1 |
|
||||
| `state_schema` fields not merged / input/output control | middleware issue 2 |
|
||||
| `interrupt` resume value missing / HITL | middleware issue 3 |
|
||||
| Modifying state inside `wrap_model_call` has no effect | middleware issue 4 |
|
||||
| Choosing between `astream_events` and `astream` for streaming | streaming issue 1 |
|
||||
| Distinguishing token sources across multiple LLMs | streaming issue 2 |
|
||||
| Disabling streaming for a specific model | streaming issue 3 |
|
||||
| Custom events from inside a tool not being emitted | streaming issue 4 |
|
||||
| Multi-agent: subagents vs handoffs | multi-agent issue 1 |
|
||||
| Single dispatch tool vs one tool per agent | multi-agent issue 2 |
|
||||
| `interrupt` can't see subagent state | multi-agent issue 3 |
|
||||
| Too much subagent wrapper boilerplate | multi-agent issue 4 |
|
||||
| Quickly building a handoff-based multi-agent setup | multi-agent issue 5 |
|
||||
| Tool returning data to both the model and the app layer / `artifact` / `Command(update=...)` | common-issues issue 1 |
|
||||
| MCP tool can't access `user_id` / `store` / state / API key | common-issues issue 2 |
|
||||
| `invalid_tool_calls` / tool never executes / malformed tool-call JSON | common-issues issue 3 |
|
||||
| Dynamic system prompt placeholders / `format_prompt` / Jinja2 prompt variables | common-issues issue 4 |
|
||||
| Agent loses context across sessions — personal assistant or support bot | contextseek-middleware issue 1 |
|
||||
| Multi-tool data-pipeline agent — auditing tool call decisions | contextseek-middleware issue 2 |
|
||||
| Research agent accumulates raw notes — cross-topic pattern discovery | contextseek-middleware issue 3 |
|
||||
| SRE incident postmortem agent — tracing knowledge confidence and conflicts | contextseek-middleware issue 4 |
|
||||
| Enterprise knowledge migration — agent is retrieval-ready on day one | contextseek-middleware issue 5 |
|
||||
@@ -0,0 +1,99 @@
|
||||
# CN Model Integration Guide
|
||||
|
||||
Help developers write LangChain integration classes for a specified Chinese model (e.g., Qwen, GLM, DeepSeek, Moonshot) using the OpenAI-compatible interface.
|
||||
|
||||
> [!CAUTION]
|
||||
> **Never read, write, or access user configuration files such as `.env`, `.env.local`, `credentials.json`, or any other files that may contain secrets or sensitive information.** API Keys and other credentials must always be filled in by the user themselves — do not peek into or modify these files under any circumstances.
|
||||
|
||||
## Step 1: Gather Information
|
||||
|
||||
Confirm the following details with the user. If the user does not explicitly provide any of these, use reasonable defaults from the provider's documentation.
|
||||
|
||||
<!-- gather
|
||||
prompt: "Confirm the following details for the model integration:"
|
||||
fields:
|
||||
- name: model_name
|
||||
question: "Model name (lowercase, used for directory and class names)"
|
||||
example: "qwen"
|
||||
required: true
|
||||
- name: api_base
|
||||
question: "API base URL (OpenAI-compatible endpoint)"
|
||||
example: "https://dashscope.aliyuncs.com/compatible-mode/v1"
|
||||
required: true
|
||||
- name: api_key_env
|
||||
question: "API key environment variable name"
|
||||
example: "QWEN_API_KEY"
|
||||
required: true
|
||||
fallback: "Use reasonable defaults from the provider's documentation."
|
||||
-->
|
||||
|
||||
1. **Model Name** — lowercase, e.g., `qwen`, `glm`, `deepseek`. Used for directory names, class names, and `_llm_type`.
|
||||
2. **API Base URL** — the model's OpenAI-compatible endpoint URL.
|
||||
3. **API Key Environment Variable Name** — e.g., `QWEN_API_KEY`.
|
||||
|
||||
Additionally, inspect the project directory structure to determine the Python package manager (`uv.lock` → uv, `poetry.lock` → poetry, `requirements.txt` → pip, etc.).
|
||||
|
||||
## Step 2: Create Directory and Files
|
||||
|
||||
1. Create a top-level directory `<models_dir>/`.
|
||||
2. Create a model subdirectory `<models_dir>/<model_name>/` with `model_name` in lowercase.
|
||||
3. Keep the top-level `<models_dir>/__init__.py` empty.
|
||||
|
||||
```
|
||||
<models_dir>/
|
||||
├── __init__.py # empty
|
||||
├── <model_name>/
|
||||
│ ├── __init__.py
|
||||
│ └── chat_model.py
|
||||
└── ...
|
||||
```
|
||||
|
||||
## Step 3: Check if DeepSeek
|
||||
|
||||
**If the model is DeepSeek**, install `langchain-deepseek` and use `ChatDeepSeek` directly. Skip all subsequent steps.
|
||||
|
||||
**If the model is another provider**, continue with the steps below.
|
||||
|
||||
## Step 4: Copy the Template
|
||||
|
||||
1. Check whether `langchain-openai` is installed; install it if not.
|
||||
2. Copy the template from [../../template/chat_model.py](../../template/chat_model.py) into the target subdirectory.
|
||||
3. Create `__init__.py`: `from .chat_model import <CHAT_CLASS_NAME>`
|
||||
|
||||
## Step 5: Replace Placeholders
|
||||
|
||||
Use grep to list all placeholders, then replace each one with the actual value:
|
||||
|
||||
| Placeholder | Description | Example (Qwen) |
|
||||
|-------------|-------------|----------------|
|
||||
| `ChatModel` | Class name | `ChatQwen` |
|
||||
| `PROVIDER_API_KEY` | API Key env var name | `QWEN_API_KEY` |
|
||||
| `PROVIDER_API_BASE` | API Base env var name | `QWEN_API_BASE` |
|
||||
| `PROVIDER_API_BASE_URL` | Default API URL | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
|
||||
| `chat-provider` | Model identifier for `_llm_type` | `chat-qwen` |
|
||||
| `provider-name` | Value for `response_metadata["model_provider"]` | `dashscope` |
|
||||
| `Provider` | Provider display name for error messages | `Qwen` |
|
||||
|
||||
Each placeholder is a standalone, complete token — simply do a global find-and-replace. Apply replacements in both `chat_model.py` and `__init__.py`.
|
||||
|
||||
## Step 6: Configure Model Profile (Optional)
|
||||
|
||||
Use `langchain-model-profiles` to download profile information for the model provider. `<provider_name>` is the provider name; try a few likely candidates.
|
||||
|
||||
1. Check whether `langchain-model-profiles` is installed; install it if not.
|
||||
2. Run the download command:
|
||||
|
||||
```bash
|
||||
langchain-profiles refresh --provider <provider_name> --data-dir ./<models_dir>/<model_name>/data
|
||||
```
|
||||
|
||||
On success, a `data/_profiles.py` file is generated under the model directory, which is used by `_get_default_model_profile` in the template. If you cannot find the corresponding provider after several attempts, skip this step.
|
||||
|
||||
## Step 7: Write Integration Tests
|
||||
|
||||
After the model class is complete, you must write integration tests. See the detailed guide at [integration-tests.md](integration-tests.md).
|
||||
|
||||
> [!IMPORTANT]
|
||||
> **Before running integration tests, you must remind the user to edit the `.env` file themselves and fill in the required API Key and other environment variables.**
|
||||
>
|
||||
> When running tests, you will likely encounter common setup issues (package not importable, async test mode, etc.). Refer to the "Common Issues" section at the end of [integration-tests.md](integration-tests.md) for fixes.
|
||||
@@ -0,0 +1,117 @@
|
||||
# Chat Model Integration Tests
|
||||
|
||||
After writing `chat_model.py`, you must write integration tests to verify the model class works correctly.
|
||||
|
||||
## Test Framework
|
||||
|
||||
Use `ChatModelIntegrationTests` from `langchain_tests` as the base class, and run tests with pytest.
|
||||
|
||||
Install dependencies:
|
||||
|
||||
```bash
|
||||
pip install langchain-tests pytest python-dotenv
|
||||
```
|
||||
|
||||
## Test File Structure
|
||||
|
||||
Place test files following standard unit test directory conventions:
|
||||
|
||||
```
|
||||
src/<models_dir>/<model_name>/
|
||||
├── __init__.py
|
||||
├── chat_model.py
|
||||
└── ...
|
||||
|
||||
tests/
|
||||
└── test_chat_<model_name>.py
|
||||
```
|
||||
|
||||
## Standard Test Class
|
||||
|
||||
For a new provider (e.g., Qwen, GLM), create a test class that inherits from `ChatModelIntegrationTests` and provides the following properties:
|
||||
|
||||
```python
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
from dotenv import load_dotenv
|
||||
from langchain_core.language_models import BaseChatModel
|
||||
from langchain_tests.integration_tests import ChatModelIntegrationTests
|
||||
|
||||
from models.qwen.chat_model import ChatQwen # replace with the actual import path
|
||||
|
||||
load_dotenv()
|
||||
|
||||
|
||||
class TestChatQwen(ChatModelIntegrationTests):
|
||||
@property
|
||||
def chat_model_class(self) -> type[BaseChatModel]:
|
||||
return ChatQwen
|
||||
|
||||
@property
|
||||
def chat_model_params(self) -> dict:
|
||||
return {
|
||||
"model": "qwen-plus",
|
||||
"temperature": 0,
|
||||
}
|
||||
```
|
||||
|
||||
### Required Properties
|
||||
|
||||
| Property | Description |
|
||||
|----------|-------------|
|
||||
| `chat_model_class` | Returns the chat model class under test. |
|
||||
| `chat_model_params` | Parameters for creating an instance. Must include `model`; `temperature: 0` is recommended for deterministic results. |
|
||||
|
||||
## Running Tests
|
||||
|
||||
```bash
|
||||
# Run tests for a single model
|
||||
pytest tests/test_chat_<model_name>.py -v
|
||||
|
||||
# Skip tests marked as xfail (run only expected passes)
|
||||
pytest tests/test_chat_<model_name>.py -v -m "not xfail"
|
||||
|
||||
# Run all model tests
|
||||
pytest tests/ -v
|
||||
```
|
||||
|
||||
## Common Issues
|
||||
|
||||
After setting up the test, you will likely encounter the following issues. Address them before concluding tests pass.
|
||||
|
||||
### Model package not importable
|
||||
|
||||
By default, the `<models_dir>/` directory is not installed as a Python package, so `from <models_dir>.xxx import ...` in tests will fail. Two changes are needed in `pyproject.toml`:
|
||||
|
||||
**1) Add build-system config:**
|
||||
|
||||
```toml
|
||||
[build-system]
|
||||
requires = ["hatchling"]
|
||||
build-backend = "hatchling.build"
|
||||
|
||||
[tool.hatch.build.targets.wheel]
|
||||
packages = ["<models_dir>"]
|
||||
```
|
||||
|
||||
**2) Install in editable mode:**
|
||||
|
||||
```bash
|
||||
uv pip install -e .
|
||||
```
|
||||
|
||||
Without this, pytest fails with `ModuleNotFoundError: No module named '<models_dir>'`.
|
||||
|
||||
### Async tests not running (pytest-asyncio strict mode)
|
||||
|
||||
pytest-asyncio defaults to `Mode.STRICT`, which requires every async test to have an `@pytest.mark.asyncio` decorator. `langchain_tests` async methods lack this decorator.
|
||||
|
||||
Add to `pyproject.toml`:
|
||||
|
||||
```toml
|
||||
[tool.pytest.ini_options]
|
||||
asyncio_mode = "auto"
|
||||
```
|
||||
|
||||
Without this, all async tests (`test_ainvoke`, `test_astream`, `test_abatch`, etc.) fail with "async def functions are not natively supported."
|
||||
@@ -0,0 +1,218 @@
|
||||
# Other Common Development Issues
|
||||
|
||||
A collection of standalone but high-frequency issues that span multiple categories or sit at the basic engineering layer.
|
||||
|
||||
## Issue 1: A tool needs to return extra data on top of the model-facing result
|
||||
|
||||
- **Symptom**: A tool needs to return both "text for the model to read" and "extra data for the application side to consume" — e.g. a retrieval tool needs to return a passage (for the model) and the source document ID + page number (for the frontend to highlight); an order lookup needs to return an order summary (for the model) and the full order object (for the next tool / main agent state). Concatenating everything into one string lets the model get distracted by irrelevant fields; returning only text loses the structured data.
|
||||
- **Cause**: A tool's return value has three different semantics; mixing them causes problems:
|
||||
|
||||
- For the **model** to see: must fit into `ToolMessage.content`; the model uses this to decide next steps.
|
||||
- For the **application layer / downstream business code** to see, but **not** into the LLM context: e.g. document ID, raw payload, render hints.
|
||||
- For writing back to **agent state**, to be read by subsequent tools / middleware: e.g. `customer_id`, `last_order`, `current_step`.
|
||||
|
||||
These three categories should go through three different channels; jamming them into one string both adds noise for the model and loses structured data for the application.
|
||||
|
||||
- **Solution**: Choose the return style by data destination (three styles can be mixed):
|
||||
|
||||
- **Model-only → return string or dict directly**: dicts get serialized into `ToolMessage.content`, and the model reads the fields itself.
|
||||
|
||||
```python
|
||||
from langchain.tools import tool
|
||||
|
||||
@tool
|
||||
def get_weather_data(city: str) -> dict:
|
||||
"""Get structured weather for a city."""
|
||||
return {"city": city, "temperature_c": 22, "conditions": "sunny"}
|
||||
```
|
||||
|
||||
- **One for the model + metadata for the app layer (not into LLM) → `ToolMessage(artifact=...)`**: `content` enters the model context, `artifact` doesn't enter the model but stays on the `ToolMessage` for downstream consumption (typical scenario: retrieval tool returning passage + document ID).
|
||||
|
||||
```python
|
||||
from langchain.messages import ToolMessage
|
||||
from langchain.tools import ToolRuntime, tool
|
||||
|
||||
@tool
|
||||
def search_books(query: str, runtime: ToolRuntime) -> ToolMessage:
|
||||
"""Retrieve a passage and attach source metadata."""
|
||||
passage = "It was the best of times, it was the worst of times."
|
||||
return ToolMessage(
|
||||
content=passage, # enters model context
|
||||
tool_call_id=runtime.tool_call_id,
|
||||
name="search_books",
|
||||
artifact={"document_id": "doc_123", "page": 0}, # readable to app layer, not into LLM
|
||||
)
|
||||
```
|
||||
|
||||
Application code later reads structured data from `message.artifact`; the model sees only the `content` text.
|
||||
|
||||
- **Write data back to agent state for subsequent tool / middleware / main agent reuse → return `Command(update=...)`**: in the update, place both business fields and a `ToolMessage` (must contain a ToolMessage paired with `tool_call_id`, otherwise the next LLM call will fail with invalid message sequence because "tool_call has no tool_response").
|
||||
|
||||
```python
|
||||
from langchain.messages import ToolMessage
|
||||
from langchain.tools import ToolRuntime, tool
|
||||
from langgraph.types import Command
|
||||
|
||||
@tool
|
||||
def lookup_customer(customer_id: str, runtime: ToolRuntime) -> Command:
|
||||
"""Look up customer and write profile into agent state."""
|
||||
profile = fetch_customer(customer_id) # {"name": ..., "tier": ..., ...}
|
||||
return Command(update={
|
||||
"customer_profile": profile, # written into state, readable by subsequent tools
|
||||
"messages": [ToolMessage(
|
||||
content=f"Found customer {profile['name']} (tier: {profile['tier']})",
|
||||
tool_call_id=runtime.tool_call_id, # required to pair with tool_call
|
||||
)],
|
||||
})
|
||||
```
|
||||
|
||||
State fields must first be declared in `state_schema` (or via an `AgentState` subclass), otherwise updates are ignored.
|
||||
|
||||
- **Lessons learned**:
|
||||
- Default to "return string/dict" — covers 80% of scenarios.
|
||||
- For fields the **application layer wants but would be noise in the LLM context** (document ID, raw payload, render metadata), use `ToolMessage(artifact=...)`.
|
||||
- For "cross-tool / cross-turn reusable business data" (user profile, looked-up order, current stage), use `Command(update=...)` — remember the paired `ToolMessage`, and declare fields in state schema first.
|
||||
- The three styles aren't mutually exclusive: a tool can simultaneously do `Command(update={"customer_profile": ..., "messages": [ToolMessage(content=..., artifact=...)]})`, using all three channels at once.
|
||||
|
||||
## Issue 2: MCP tool can't access the agent's user_id / API key / current state
|
||||
|
||||
- **Symptom**: You wire an MCP server in as a regular LangChain tool and want to read `runtime.context.user_id`, current agent state, or user preferences from `store` inside the tool — only to find the MCP tool can't access any of these and is limited to the args declared in the schema. Adding `user_id` directly to the tool schema makes the model fabricate one for every call, polluting context and being insecure.
|
||||
- **Cause**: MCP servers run in a **separate process** (stdio subprocess or remote HTTP service), **completely process-isolated** from the LangGraph runtime — they can't see store, context, state, or tool_call_id. Importing LangGraph runtime APIs on the MCP server side is pointless because it's not running in that process.
|
||||
- **Solution**: Bridge on the **client side** with `tool_interceptors` — interceptors run in the LangGraph process and can access the full `ToolRuntime`, injecting the needed fields into `args` / `headers` before forwarding to the MCP server.
|
||||
|
||||
```python
|
||||
from langchain_mcp_adapters.client import MultiServerMCPClient
|
||||
from langchain_mcp_adapters.interceptors import MCPToolCallRequest
|
||||
|
||||
async def inject_user_context(request: MCPToolCallRequest, handler):
|
||||
runtime = request.runtime
|
||||
# 1) Business fields (user_id, tenant_id) injected into args; model can't see them and doesn't need them in schema
|
||||
args = {**request.args, "user_id": runtime.context.user_id}
|
||||
# 2) Auth / tracing info injected into headers, doesn't pollute schema at all
|
||||
headers = {"Authorization": f"Bearer {runtime.context.api_key}"}
|
||||
return await handler(request.override(args=args, headers=headers))
|
||||
|
||||
client = MultiServerMCPClient({...}, tool_interceptors=[inject_user_context])
|
||||
```
|
||||
|
||||
For auth "short-circuit" scenarios, directly `return ToolMessage(...)` without calling `handler` — e.g. denying sensitive tool calls when unauthenticated:
|
||||
|
||||
```python
|
||||
from langchain.messages import ToolMessage
|
||||
|
||||
async def require_auth(request: MCPToolCallRequest, handler):
|
||||
if request.name in {"delete_file", "export_data"} and not request.runtime.state.get("authenticated"):
|
||||
return ToolMessage(
|
||||
content="Authentication required.",
|
||||
tool_call_id=request.runtime.tool_call_id,
|
||||
)
|
||||
return await handler(request)
|
||||
```
|
||||
|
||||
- **Lessons learned**:
|
||||
- **Business fields (user_id, tenant_id) → modify `args`**: interceptor injects them centrally; not in schema, so the model won't fabricate them.
|
||||
- **Auth / tracing info → modify `headers`**: not in schema at all, the cleanest.
|
||||
- **`structuredContent` returned by MCP server is invisible to the model by default** (placed only in `ToolMessage.artifact`) — to let the model read it, use the interceptor to serialize `structuredContent` and concatenate it back to `result.content`.
|
||||
- Multiple interceptors follow **onion order**: `[outer, inner]` → outer enters first / exits last. When composing "auth + rate limit + retry", put the outer concern in front, inner (closest to the tool call) in back.
|
||||
- `MultiServerMCPClient` is **stateless by default** — a new session opens for every tool call. When you need to reuse context across calls (e.g. server-side login state), use `async with client.session("server_name") as session:` to explicitly manage the lifecycle.
|
||||
|
||||
## Issue 3: Model always outputs `invalid_tool_calls` - tool never actually executes
|
||||
|
||||
- **Symptom**: The model is bound with tools and consistently "calls" them, but the tool function never fires. Inspecting the `AIMessage` shows `tool_calls` is empty while `invalid_tool_calls` is populated. The agent loop either silently skips the call or raises a parsing error. The weaker the model (small-parameter open-source, quantized, vLLM-served), the more frequent this becomes.
|
||||
- **Cause**: When the model generates a tool call, the `arguments` field must be valid JSON conforming to the tool's schema. Weak models often produce malformed JSON — missing quotes, trailing commas, unescaped characters, truncated output, etc. LangChain's tool-call parser **cannot parse** the broken JSON, so instead of placing it in `tool_calls` it moves it to `invalid_tool_calls`. Since the agent executor only processes entries in `tool_calls`, the tool never runs — it looks like the model "called" it but nothing happened.
|
||||
- **Solution**: Use `ToolCallRepairMiddleware` from `langchain-dev-utils` — it automatically detects entries in `invalid_tool_calls`, attempts to repair the malformed JSON via the `json-repair` library, and promotes successfully repaired calls back to `tool_calls` so the agent can execute them normally.
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.agents.middleware import tool_call_repair
|
||||
|
||||
agent = create_agent(
|
||||
model="openai:gpt-5-mini",
|
||||
tools=[run_python_code, get_current_time],
|
||||
middleware=[tool_call_repair],
|
||||
)
|
||||
```
|
||||
|
||||
`tool_call_repair` is a pre-instantiated global instance of `ToolCallRepairMiddleware` — zero configuration needed.
|
||||
|
||||
If you prefer explicit instantiation:
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.agents.middleware import ToolCallRepairMiddleware
|
||||
|
||||
agent = create_agent(
|
||||
model="openai:gpt-5-mini",
|
||||
tools=[run_python_code, get_current_time],
|
||||
middleware=[ToolCallRepairMiddleware()],
|
||||
)
|
||||
```
|
||||
|
||||
- **Lessons learned**:
|
||||
- When a tool "isn't being called", **check `invalid_tool_calls` on the AIMessage first** — the model likely did attempt a call, but the JSON was unparseable.
|
||||
- `ToolCallRepairMiddleware` cannot guarantee 100% repair — severely garbled output (e.g. half the JSON is natural language) will still fail. For those cases, consider simplifying the tool schema, splitting complex parameters into multiple smaller tools, or upgrading to a stronger model.
|
||||
- This middleware only acts on `invalid_tool_calls` — valid calls pass through untouched with zero overhead.
|
||||
|
||||
## Issue 4: How to use placeholders in system prompt that get dynamically replaced at runtime
|
||||
|
||||
- **Symptom**: You want the system prompt to include dynamic information — user name, role, current date, conversation context, etc. — that varies per request. Hardcoding these values means creating a new agent for every variation; concatenating strings manually is error-prone and hard to maintain.
|
||||
- **Cause**: `create_agent` treats `system_prompt` as a static string by default — it does not perform any template interpolation. To get runtime substitution, you need an explicit formatting middleware that resolves placeholders against `state` and `context` before the prompt reaches the model.
|
||||
- **Solution**: Add the `format_prompt` middleware (f-string style, covers most cases) or `FormatPromptMiddleware(template_format="jinja2")` (for conditionals / loops):
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.agents.middleware import format_prompt
|
||||
from langchain.agents import AgentState
|
||||
from dataclasses import dataclass
|
||||
|
||||
class AssistantState(AgentState):
|
||||
name: str
|
||||
|
||||
@dataclass
|
||||
class UserContext:
|
||||
user: str
|
||||
|
||||
agent = create_agent(
|
||||
model="openai:gpt-5",
|
||||
system_prompt="You are {name}, an assistant for {user}.",
|
||||
middleware=[format_prompt],
|
||||
state_schema=AssistantState,
|
||||
context_schema=UserContext,
|
||||
)
|
||||
|
||||
# At runtime, {name} is resolved from state, {user} from context
|
||||
response = agent.invoke(
|
||||
{"messages": [HumanMessage(content="Hello")], "name": "Jarvis"},
|
||||
context=UserContext(user="Tony"),
|
||||
)
|
||||
# Model receives: "You are Jarvis, an assistant for Tony."
|
||||
```
|
||||
|
||||
Variables are resolved in priority order: **`state` first, then `context`** — a `state` field with the same name shadows the `context` field.
|
||||
|
||||
For templates that need conditionals or loops, use Jinja2:
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.agents.middleware import FormatPromptMiddleware
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional
|
||||
|
||||
@dataclass
|
||||
class Context:
|
||||
user_role: Optional[str] = None
|
||||
|
||||
jinja2_formatter = FormatPromptMiddleware(template_format="jinja2")
|
||||
|
||||
agent = create_agent(
|
||||
model="openai:gpt-5",
|
||||
system_prompt=(
|
||||
"You are an assistant.\n"
|
||||
"{% if user_role == 'VIP' %}Provide premium service.{% endif %}"
|
||||
),
|
||||
middleware=[jinja2_formatter],
|
||||
context_schema=Context,
|
||||
)
|
||||
```
|
||||
|
||||
- **Lessons learned**:
|
||||
- Template interpolation is **opt-in** via middleware — never assume it happens by default.
|
||||
- Use `format_prompt` (global instance, zero config) for simple `{variable}` substitution; only reach for Jinja2 when you need `{% if %}` / `{% for %}`.
|
||||
- If a placeholder has no matching field in either `state` or `context`, formatting will raise a `KeyError` — declare all variables in the corresponding schema.
|
||||
- Keep `system_prompt` hardcoded by the developer; pass only **data** through `state` / `context`. Never let end-user input become the template itself — this applies regardless of whether formatting middleware is enabled.
|
||||
@@ -0,0 +1,264 @@
|
||||
# ContextSeek Middleware — Use Case Scenarios
|
||||
|
||||
Every scenario below answers one question: **"I'm building X — is ContextSeekMiddleware the right fit, and how do I wire it in?"** Parameter-level configuration issues that arise after the decision to use the middleware are covered in [contextseek-params.md](contextseek-params.md).
|
||||
|
||||
---
|
||||
|
||||
## Issue 1: Agent loses context across sessions — personal assistant or support bot
|
||||
|
||||
**Keywords**: semantic memory, context loss, cross-session, agent forgets, StoreBackend alternative, vector retrieval, passive memory, persistent memory
|
||||
|
||||
You're building a **personal coding assistant** or a **customer support bot**. Users notice the agent "forgets" everything between sessions: an architecture decision made last week ("we chose OceanBase over Redis because of HTAP requirements") is unknown to the agent today; a support user who already explained their account type and preferred language has to explain it again on every new conversation.
|
||||
|
||||
The instinct is to reach for `memory=` and `StoreBackend` (from Deep Agents). That works — but only for knowledge the agent actively chooses to write down. It degrades when conversation history is long or when the agent doesn't "know" what's worth remembering. What's missing is a passive retrieval layer that automatically surfaces semantically relevant history before each model call without any agent involvement.
|
||||
|
||||
`ContextSeekMiddleware` sidecars the agent loop transparently:
|
||||
|
||||
1. **Before every model call**: retrieves the top-k semantically relevant items from the vector store and appends them to the system message as a `[Relevant Context]` block.
|
||||
2. **After every final answer**: stores the Q+A pair for future retrieval (skips intermediate tool-call turns to avoid noise).
|
||||
|
||||
The agent is never aware of either step.
|
||||
|
||||
- **Install**:
|
||||
```bash
|
||||
# Inside the agentseek project
|
||||
pip install "agentseek[context]"
|
||||
|
||||
# Standalone
|
||||
pip install contextseek contextseek-bridges-langchain
|
||||
```
|
||||
- **Configure via env vars** (copy from `.env.example`, section "ContextSeek semantic context layer"). All constructor parameters are optional — the middleware reads from env when nothing is passed:
|
||||
```bash
|
||||
# seekdb = local persistent storage, built-in ONNX embedder (no API key needed)
|
||||
AGENTSEEK_CTX_STORAGE_BACKEND=seekdb
|
||||
|
||||
# Optional LLM for richer L1 summaries
|
||||
AGENTSEEK_CTX_LLM_PROVIDER=openai
|
||||
AGENTSEEK_CTX_LLM_MODEL=gpt-4o-mini
|
||||
```
|
||||
The `AGENTSEEK_CTX_*` prefix is automatically aliased to the env vars contextseek reads internally — no credential duplication.
|
||||
- **Minimum integration** (zero constructor arguments):
|
||||
```python
|
||||
from langchain.agents import create_agent
|
||||
from contextseek.bridges.langchain.middleware import ContextSeekMiddleware
|
||||
|
||||
agent = create_agent(
|
||||
model=model,
|
||||
tools=[...],
|
||||
middleware=[ContextSeekMiddleware()],
|
||||
)
|
||||
```
|
||||
- **Sharing the agent's own model and embedder** (avoids a second model instantiation):
|
||||
```python
|
||||
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
|
||||
|
||||
model = ChatOpenAI(model="gpt-4o")
|
||||
embedder = OpenAIEmbeddings(model="text-embedding-3-small")
|
||||
|
||||
agent = create_agent(
|
||||
model=model,
|
||||
tools=[...],
|
||||
middleware=[ContextSeekMiddleware(model=model, embedder=embedder)],
|
||||
)
|
||||
```
|
||||
- **Lessons learned**: `ContextSeekMiddleware` and `memory=` (StoreBackend) solve different problems and can coexist in the same agent. Use `memory=` when the agent needs to explicitly read and write structured notes it controls. Use `ContextSeekMiddleware` when you want the agent's accumulated conversation history to automatically inform future answers — no agent-side file management required.
|
||||
|
||||
**Common configuration problems for this scenario** → see [contextseek-params.md](contextseek-params.md): scope isolation (multi-user context bleeding), auto_compact throttling and graceful shutdown, retrieval_tags / min_score filtering noisy context, tool_arg_overrides for injecting runtime arguments.
|
||||
|
||||
---
|
||||
|
||||
## Issue 2: Multi-tool data-pipeline agent — auditing tool call decisions
|
||||
|
||||
**Keywords**: tool provenance, audit trail, tool call history, data pipeline agent, record_tool_calls, why did agent choose this query, compliance, tool decision tracing
|
||||
|
||||
You're building a **data analysis agent** that, on each task, calls a chain of tools: `query_db` → `run_sql` → `transform_data` → `generate_chart` — typically 5–10 tool invocations per turn. A compliance requirement asks: *"Why did the agent choose that particular SQL query? Were the tool arguments reasonable?"* Your team also needs to reproduce past analyses from stored tool traces.
|
||||
|
||||
`ContextSeekMiddleware` with `record_tool_calls=True` writes a structured record for each tool invocation into the vector store. Each record captures:
|
||||
|
||||
- `tool` — tool name
|
||||
- `args` — arguments passed (after any `tool_arg_overrides` are applied)
|
||||
- `result` — the ToolMessage content
|
||||
- `rationale` — the AIMessage text that preceded the tool call (the model's stated reasoning)
|
||||
- `task` — the originating user message
|
||||
|
||||
These records are retrievable by future agent turns: *"What SQL did we use for the Q3 revenue report last month?"* can return the exact query with its rationale.
|
||||
|
||||
- **Integration**:
|
||||
```python
|
||||
middleware = ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_store=True, # store final Q+A pairs
|
||||
record_tool_calls=True, # additionally store each tool invocation
|
||||
)
|
||||
agent = create_agent(model=model, tools=[query_db, run_sql, transform_data, generate_chart], middleware=[middleware])
|
||||
```
|
||||
- **Retrieval-only mode** — read historical tool traces without writing new ones (e.g. a read-only audit agent):
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_store=False,
|
||||
record_tool_calls=False, # retrieval only; no writes
|
||||
)
|
||||
```
|
||||
- **Lessons learned**: `record_tool_calls=True` multiplies write volume by the average number of tool calls per turn. Each call triggers the full summarizer + embedding + DB write pipeline. For a 5-tool agent, that is 5× the LLM cost per turn compared to `auto_store=True` alone. Only enable it when you actually need per-tool provenance.
|
||||
|
||||
**Common configuration problems for this scenario** → see [contextseek-params.md](contextseek-params.md): auto_store / record_tool_calls write volume and cost spikes.
|
||||
|
||||
---
|
||||
|
||||
## Issue 3: Research agent accumulates raw notes — cross-topic pattern discovery
|
||||
|
||||
**Keywords**: dream, cross-topic pattern, consolidation, divergence, knowledge synthesis, idle-time, hypothesis generation, dreaming, research agent, implicit connections
|
||||
|
||||
You're building a **literature research agent** that writes dozens of raw research summaries per day across multiple topics. After a week the store holds 200+ items but the agent can only retrieve "known" content — it misses implicit cross-topic connections: "gene editing for Alzheimer's treatment" and "mRNA vaccine T-cell activation" both involve immune regulation mechanisms, but no item makes that bridge explicit. You need the system to synthesize new insights from accumulated material without re-running the full agent.
|
||||
|
||||
`ctx.dream()` runs offline in two phases:
|
||||
|
||||
1. **Consolidation** — scans recently active, un-dreamed items. Within a similarity window `(0.35, 0.72)` — related but not duplicates — it synthesizes new `extracted`-stage items that surface implicit patterns. Tagged `dreamed`, `consolidation`.
|
||||
2. **Divergence** — generates cross-cluster hypothesis items bridging dissimilar topic clusters. Tagged `dreamed`, `divergence`. Confidence is lower (×0.85 multiplier) — these are hypotheses, not facts.
|
||||
|
||||
Dream items have `stability=transient` and decay fast. They are "use it or lose it": `ctx.feedback(ref, score=1.0)` on a valuable hypothesis promotes it to a stable item; otherwise it fades within days.
|
||||
|
||||
- **Trigger dream manually after a research session**:
|
||||
```python
|
||||
from contextseek import ContextSeek
|
||||
|
||||
ctx = ContextSeek.from_settings()
|
||||
report = ctx.dream(scope="research/immunology")
|
||||
consolidation_count = len(report.consolidation.items)
|
||||
divergence_count = len(report.divergence.items) if report.divergence else 0
|
||||
print(f"Consolidations: {consolidation_count}, Divergences: {divergence_count}")
|
||||
```
|
||||
- **Review and promote valuable dream items**:
|
||||
```python
|
||||
scope = "research/immunology"
|
||||
# ctx.overview() returns stage distribution counts; dreamed items are filtered by tag separately
|
||||
all_items = ctx.items(scope=scope)
|
||||
dreamed = [item for item in all_items if "dreamed" in (item.tags or [])]
|
||||
|
||||
for item in dreamed:
|
||||
if human_review_approves(item):
|
||||
# feedback() requires a full URI ref, not a bare item id
|
||||
ref = ctx.resolver.ref_for(scope, item.id)
|
||||
ctx.feedback(ref, scope=scope, score=1.0, reason="confirmed cross-domain insight")
|
||||
```
|
||||
- **Enable LLM-enhanced dream** (richer synthesis):
|
||||
```bash
|
||||
DREAM_LLM_ENABLED=true
|
||||
# uses the same LLM configured for the agent; no separate key needed
|
||||
```
|
||||
- **Lessons learned**: By default `ContextSeekMiddleware` does **not** trigger `dream()` — it only handles `compact()`. For research workflows, call `ctx.dream()` explicitly after bulk ingestion sessions, or schedule it via cron / the contextseek daemon lifecycle. If you want automatic dream triggering inside the agent loop, set `auto_dream=True` on the middleware (see [contextseek-params.md](contextseek-params.md) Issue 11 for the dual-gate trigger mechanics). Dream without LLM falls back to keyword-overlap heuristics; results are coarser but still useful for surface-level clustering.
|
||||
|
||||
**Common configuration problems for this scenario** → see [contextseek-params.md](contextseek-params.md): dream trigger conditions not met (min_items, cooldown), dream-generated items disappearing due to transient stability.
|
||||
|
||||
---
|
||||
|
||||
## Issue 4: SRE incident postmortem agent — tracing knowledge confidence and conflicts
|
||||
|
||||
**Keywords**: evidence_chain, provenance, confidence propagation, conflict detection, SRE, postmortem, audit, upstream, chain_confidence, broken links, knowledge trustworthiness
|
||||
|
||||
You're building an **incident postmortem agent** that accumulates observations during a live incident: "Alert A fired at 14:03", "Log B shows connection timeouts", "Analysis: timeouts caused by DB connection pool exhaustion", "Recommendation: lower `connection_timeout` to 500ms". Each item is derived from the previous. A week later, during a recurring-incident review, an engineer asks: *"Is this recommendation actually well-supported? Were there any conflicting signals? How confident should we be in this diagnosis?"*
|
||||
|
||||
`ctx.evidence_chain()` constructs a full DAG starting from any item, traversing `derived_from`, `supported_by`, `merged_from` (positive) and `refuted_by` (negative) links. It returns:
|
||||
|
||||
- `overall_confidence` — propagated confidence score (Noisy-OR for supports, penalty for refutations)
|
||||
- `critical_path` — `list[str]` of item ids forming the highest-weight inference path
|
||||
- `conflicts` — `list[ConflictReport]`; each has `item_id`, `refuter_id`, `refutation_strength`
|
||||
- `broken_links` — `list[str]` of item ids that no longer exist in the store
|
||||
- `nodes` / `edges` — full DAG for visualization
|
||||
- `needs_reverification` — `bool`; `True` when `overall_confidence < 0.4`
|
||||
|
||||
- **Write items with explicit derivation links**:
|
||||
```python
|
||||
from contextseek import ContextSeek
|
||||
from contextseek.domain.links import Link, LinkType
|
||||
|
||||
ctx = ContextSeek.from_settings()
|
||||
scope = "incidents/2026-06-01"
|
||||
|
||||
alert = ctx.add("Alert A fired at 14:03: p99 latency > 2s", scope=scope, source="pagerduty")
|
||||
log = ctx.add("Log B: connection pool exhausted (pool_size=10, wait_timeout=30s)",
|
||||
scope=scope, source="datadog",
|
||||
links=[Link(target_id=alert.id, relation=LinkType.supported_by)])
|
||||
analysis = ctx.add("Root cause: DB connection pool exhausted under 50-rps load",
|
||||
scope=scope, source="agent_inference",
|
||||
links=[Link(target_id=log.id, relation=LinkType.derived_from)])
|
||||
rec = ctx.add("Recommendation: lower connection_timeout to 500ms",
|
||||
scope=scope, source="agent_inference",
|
||||
links=[Link(target_id=analysis.id, relation=LinkType.derived_from)])
|
||||
```
|
||||
- **Evaluate the recommendation's trustworthiness**:
|
||||
```python
|
||||
# evidence_chain and chain_confidence require a full URI ref, not a bare item id
|
||||
ref = ctx.resolver.ref_for(scope, rec.id)
|
||||
chain = ctx.evidence_chain(ref, scope=scope)
|
||||
print(f"Confidence: {chain.overall_confidence:.2f}") # e.g. 0.71
|
||||
if chain.conflicts:
|
||||
# conflicts is list[ConflictReport]; each has item_id, refuter_id, refutation_strength
|
||||
for c in chain.conflicts:
|
||||
print(f"Conflicting evidence: {c.refuter_id} refutes {c.item_id} (strength={c.refutation_strength:.2f})")
|
||||
if chain.overall_confidence < 0.4:
|
||||
print("Low confidence — needs human review before applying")
|
||||
```
|
||||
- **Quick check without the full DAG**:
|
||||
```python
|
||||
ref = ctx.resolver.ref_for(scope, rec.id)
|
||||
score = ctx.chain_confidence(ref, scope=scope)
|
||||
# returns a float; same traversal as evidence_chain but no DAG construction overhead
|
||||
```
|
||||
- **Lessons learned**: the agent itself doesn't need to call `evidence_chain` — it's a postmortem / audit API. Wire it into your review pipeline or a separate audit agent that periodically evaluates low-confidence recommendations. Writing `links=` when calling `ctx.add()` is optional but unlocks this entire capability; without links, `evidence_chain` sees only isolated nodes.
|
||||
|
||||
**Common configuration problems for this scenario** → see [contextseek-params.md](contextseek-params.md): evidence_chain vs chain_confidence — when to use which.
|
||||
|
||||
---
|
||||
|
||||
## Issue 5: Enterprise knowledge migration — agent is retrieval-ready on day one
|
||||
|
||||
**Keywords**: pre-populate, cold start, bulk import, DataPlug, RAGPlug, PowerMemPlug, existing knowledge, seed context, initial corpus, plug, knowledge migration
|
||||
|
||||
You're migrating an enterprise knowledge base to a ContextSeek-backed agent. The existing data is 100 k+ FAQ entries, historical support tickets, and internal wiki pages — currently stored in a RAG vector store or PowerMem. The new agent cannot wait months for `auto_store` to accumulate conversation history; it needs to be able to retrieve from all of this on day one.
|
||||
|
||||
`ctx.plug()` consumes a `DataPlug` (a streaming iterator of `RawEvent` objects) and routes each event through the same full pipeline as `ctx.add()`: summarization, embedding, conflict detection, and persistence. Built-in plugs cover the most common sources:
|
||||
|
||||
| Plug class | Source |
|
||||
|------------|--------|
|
||||
| `RAGPlug` | Existing RAG / vector store chunks |
|
||||
| `PowerMemPlug` | PowerMem memory store |
|
||||
| `TracePlug` | Execution traces / agent logs |
|
||||
| `MCPToolImporter` | MCP tool definitions → `skill` stage |
|
||||
| `OpenAIFunctionImporter` | OpenAI function schemas → `skill` stage |
|
||||
|
||||
- **Import from an existing RAG store**:
|
||||
```python
|
||||
from contextseek import ContextSeek
|
||||
from contextseek.plugs import RAGPlug
|
||||
|
||||
ctx = ContextSeek.from_settings()
|
||||
|
||||
# RAGPlug accepts a list of dicts; each dict needs at least "content" or "page_content"
|
||||
docs = existing_vector_store.similarity_search("*", k=10000)
|
||||
rag_plug = RAGPlug(
|
||||
documents=[{"content": d.page_content, "metadata": d.metadata} for d in docs],
|
||||
source_name="wiki-v2",
|
||||
)
|
||||
ctx.plug(rag_plug, scope="company/knowledge")
|
||||
```
|
||||
- **Import from PowerMem**:
|
||||
```python
|
||||
from contextseek.plugs import PowerMemPlug
|
||||
|
||||
# PowerMemPlug.from_records() accepts dicts from PowerMem get_all/search results
|
||||
records = powermem_instance.get_all(user_id="shared") # yields list of dicts
|
||||
plug = PowerMemPlug.from_records(records, source_prefix="powermem")
|
||||
ctx.plug(plug, scope="company/support-history")
|
||||
```
|
||||
- **Run compact after bulk import** to consolidate and promote stages before the agent goes live:
|
||||
```python
|
||||
ctx.compact(scope="company/knowledge")
|
||||
```
|
||||
- **Combine pre-populated knowledge with live agent writes** — plug for the initial corpus, then deploy the agent with `ContextSeekMiddleware` writing new Q+A pairs into the same scope. The two pipelines are additive.
|
||||
- **Lessons learned**: `plug()` is for one-time or scheduled batch ingestion outside the agent loop. `ContextSeekMiddleware` handles continuous, per-turn ingestion inside the agent loop. Use both: `plug()` to seed the store, middleware to keep it growing. Run `compact()` after large bulk imports before the first retrieval to maximize retrieval quality.
|
||||
|
||||
**Common configuration problems for this scenario** → see [contextseek-params.md](contextseek-params.md): DataPlug vs manual ctx.add() — which to use for bulk import; plug() scope priority and stage inference.
|
||||
@@ -0,0 +1,534 @@
|
||||
# ContextSeek Middleware — Parameter Configuration Issues
|
||||
|
||||
This file covers specific parameter-level problems that arise **after** you have decided to use `ContextSeekMiddleware`. For the business scenarios that motivate each parameter group, see [contextseek-middleware.md](contextseek-middleware.md).
|
||||
|
||||
---
|
||||
|
||||
## Issue 1: scope isolation strategy — fixed at construction vs dynamic per session
|
||||
|
||||
**Keywords**: scope isolation, multi-user, thread_id, context bleeding, user isolation, ContextVar, per-session bucket
|
||||
|
||||
- **Symptom**: In a multi-user service, different users' context bleeds into each other — one user sees context retrieved from another user's conversation history. Or a single-user bot uses `thread_id` as scope but the context appears to reset on each invocation.
|
||||
- **Cause**: A `ContextSeekMiddleware` instance is shared across all concurrent agent sessions (it is stateless except for the compact executor). Scope determines which "bucket" in the store is read and written. There are two distinct resolution paths:
|
||||
- **Constructor `scope=` is given**: this value is hard-wired for every session this instance handles — `_SCOPE_VAR` (the per-task ContextVar) is **not** consulted. All sessions share the same bucket regardless of `thread_id`.
|
||||
- **Constructor `scope=` is omitted (or `None`)**: `before_agent` runs at the start of each agent turn and sets `_SCOPE_VAR` to `runtime.thread_id` for the current asyncio task. Every downstream hook (`wrap_model_call`, `after_model`, `wrap_tool_call`) reads the ContextVar, so concurrent sessions are naturally isolated without touching the instance.
|
||||
- **Solution**:
|
||||
- **Single-tenant / shared knowledge base** (all sessions read and write the same context):
|
||||
```python
|
||||
# Every session contributes to and retrieves from "my_project"
|
||||
middleware = ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
scope="my_project",
|
||||
)
|
||||
```
|
||||
- **Multi-user isolation** (each session gets its own isolated context bucket):
|
||||
```python
|
||||
# Do NOT pass scope= — let before_agent pick up runtime.thread_id
|
||||
middleware = ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
# scope= omitted
|
||||
)
|
||||
|
||||
# Callers must pass a stable, user-specific thread_id in config
|
||||
agent.invoke(
|
||||
{"messages": [...]},
|
||||
config={"configurable": {"thread_id": f"user:{user_id}"}},
|
||||
)
|
||||
```
|
||||
- **Fallback when before_agent didn't run** (sync invocation without a checkpointer):
|
||||
```python
|
||||
# _current_scope() returns _SCOPE_VAR.get() or "default"
|
||||
middleware = ContextSeekMiddleware(model=model, embedder=embedder)
|
||||
```
|
||||
- **Anti-pattern to avoid** — sharing a fixed-scope instance across users:
|
||||
```python
|
||||
# WRONG: all users pollute each other's context
|
||||
shared = ContextSeekMiddleware(model=model, embedder=embedder, scope="global")
|
||||
agent_for_user_a = create_agent(..., middleware=[shared])
|
||||
agent_for_user_b = create_agent(..., middleware=[shared]) # reads user_a's context
|
||||
```
|
||||
- **Lessons learned**: `scope=` is a deliberate opt-in to shared context. Omitting it is the safe default for multi-user services: `before_agent` populates the ContextVar from `thread_id`, and each asyncio task gets its own isolated copy. The instance itself is never mutated — it is safe to share across agents.
|
||||
|
||||
---
|
||||
|
||||
## Issue 2: auto_store and record_tool_calls — write volume and side effects
|
||||
|
||||
**Keywords**: auto_store, record_tool_calls, write volume, LLM cost spike, storage overhead, intermediate messages skipped, tool provenance
|
||||
|
||||
- **Symptom**: After setting `record_tool_calls=True`, storage write volume spikes dramatically, LLM costs for the internal Summarizer shoot up, and agent turn latency increases. Or conversely: `auto_store=True` is the default, but some intermediate AI messages (with `tool_calls`) are unexpectedly NOT being stored.
|
||||
- **Cause**: The two write paths have very different trigger frequencies:
|
||||
- `auto_store` writes in `after_model`, which fires **once per model call** — but only for final answers. The middleware deliberately skips persistence when the `AIMessage` carries `tool_calls` (i.e. an intermediate planning step) to avoid noise. Only a clean text reply (no pending tool calls) is stored.
|
||||
- `record_tool_calls` writes in `wrap_tool_call`, which fires **once per individual tool invocation**. A single agent turn that calls 5 tools generates 5 separate `ctx.add()` calls, each triggering the full pipeline: Summarizer (L0 abstract + L1 overview) + embedding + DB write. At scale this multiplies costs by the average number of tool calls per turn.
|
||||
- **Solution**:
|
||||
- **Default configuration** (recommended for most use cases):
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_store=True, # only final Q+A pairs — low volume
|
||||
record_tool_calls=False, # default; no per-tool writes
|
||||
)
|
||||
```
|
||||
- **When you need per-tool provenance** (debugging, audit logs, tracing tool decision chains):
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_store=True,
|
||||
record_tool_calls=True, # each tool call stored with tool name, args, result, rationale, task
|
||||
)
|
||||
```
|
||||
Each stored tool record contains:
|
||||
- `tool`: tool name
|
||||
- `args`: tool call arguments (after `tool_arg_overrides` are applied)
|
||||
- `result`: the ToolMessage content
|
||||
- `rationale`: the AIMessage text that preceded the tool call (the model's reasoning)
|
||||
- `task`: the last user message
|
||||
- **Disable all writes** (retrieval-only mode, e.g. reading from a pre-populated knowledge base):
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_store=False,
|
||||
record_tool_calls=False,
|
||||
)
|
||||
```
|
||||
- **Why intermediate AI messages are skipped**: when the model emits `tool_calls`, the conversation is not done — the assistant has not produced a final answer yet. Storing this half-baked state would degrade retrieval quality because the context would contain questions without coherent answers. The middleware waits for the clean final turn.
|
||||
- **Lessons learned**: `record_tool_calls=True` is a diagnostic / provenance feature, not a general-purpose setting. Turn it on only for specific debugging sessions or audit pipelines. For production, `auto_store=True` alone builds a useful semantic memory over time with minimal overhead.
|
||||
|
||||
---
|
||||
|
||||
## Issue 3: auto_compact throttling mechanism and graceful shutdown
|
||||
|
||||
**Keywords**: auto_compact, compact_every, graceful shutdown, FastAPI lifespan, RuntimeError, ThreadPoolExecutor, compact frequency
|
||||
|
||||
- **Symptom**: You enabled `auto_compact=True` expecting the store to evolve automatically, but compact never seems to run (or runs too rarely). Or, after a FastAPI service restart, you see `RuntimeError: cannot schedule new futures after shutdown` in logs related to ContextSeek.
|
||||
- **Cause**:
|
||||
- `auto_compact` triggers inside `after_agent`, which runs **after the full agent turn**. The trigger condition is: the internal counter for the current scope must reach `compact_every`. So if `compact_every=20` and the agent handles 5 sessions per day, compact fires once every 4 days per scope — much less frequently than expected.
|
||||
- The compact task is submitted to a single-threaded `ThreadPoolExecutor` (one worker). A per-scope `threading.Lock` prevents concurrent compaction of the same scope. If a previous compact is still running when the next threshold is crossed, the new trigger is silently dropped (`lock.acquire(blocking=False)`).
|
||||
- The executor shuts down via `weakref.finalize` when the middleware instance is garbage collected. If the instance is held as a long-lived object and the application has its own shutdown hook, the finalize may race with framework teardown — producing the `RuntimeError`.
|
||||
- **Solution**:
|
||||
- **Recommended compact settings for production**:
|
||||
```python
|
||||
middleware = ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_compact=True,
|
||||
compact_every=20, # trigger after every 20 completed agent turns per scope
|
||||
)
|
||||
```
|
||||
A value of 20–50 balances evolution quality (compact needs enough new material) against freshness (too high means the store rarely evolves).
|
||||
- **Graceful shutdown in FastAPI lifespan**:
|
||||
```python
|
||||
from contextlib import asynccontextmanager
|
||||
from fastapi import FastAPI
|
||||
|
||||
middleware = ContextSeekMiddleware(model=model, embedder=embedder, auto_compact=True)
|
||||
agent = create_agent(model=model, tools=[...], middleware=[middleware])
|
||||
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
yield
|
||||
middleware.shutdown(wait=True) # wait for any in-flight compact to finish
|
||||
|
||||
app = FastAPI(lifespan=lifespan)
|
||||
```
|
||||
- **Checking if compact is running**: there is no built-in status probe. Use `LANGSMITH_TRACING=true` to see `ContextSeek.compact` spans in LangSmith, or instrument `_traced_compact` externally.
|
||||
- **Manually triggering compact** (outside the middleware loop):
|
||||
```python
|
||||
middleware.ctx.compact(scope="my_project")
|
||||
```
|
||||
- **Lessons learned**: `auto_compact` is a "fire-and-forget evolution" feature — it intentionally drops triggers when the executor is busy to avoid pile-up. It is not a guarantee that compact runs exactly every N turns; only **at most** every N turns and **never** concurrently for the same scope. For critical evolution jobs, prefer explicit scheduled compact calls outside the agent loop.
|
||||
|
||||
---
|
||||
|
||||
## Issue 4: retrieval_tags and min_score — filtering noisy context injections
|
||||
|
||||
**Keywords**: retrieval_tags, min_score, irrelevant context, noisy retrieval, tag filter, score threshold, context pollution, [Relevant Context] noise
|
||||
|
||||
- **Symptom**: The `[Relevant Context]` block injected into the system message contains obviously unrelated items that confuse the model. Or a scope holds entries from multiple projects and retrieval cross-contaminates them. Or low-confidence dream-generated hypotheses pollute the context with speculative content the model treats as fact.
|
||||
- **Cause**: `wrap_model_call` calls `ctx.retrieve(query, scope, k=retrieval_k)` with no tag or score filter by default. Everything in scope that ranks in the top-k is injected, regardless of confidence or provenance tags.
|
||||
- **Solution**:
|
||||
- **Tag-based filtering** — only retrieve items tagged for a specific project or source:
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
retrieval_tags=["project:alpha"], # AND filter — all tags must be present
|
||||
)
|
||||
```
|
||||
Tags must be applied when writing items. If using `auto_store`, the middleware writes items without custom tags. Tag filtering is most useful when items were written via `ctx.add(..., tags=[...])` or `plug()` with explicit tags.
|
||||
- **Score threshold** — exclude low-confidence / low-relevance items:
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
min_score=0.6, # only inject items with retrieval score >= 0.6
|
||||
)
|
||||
```
|
||||
- **Combine both** — tag filter AND score threshold:
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
retrieval_tags=["project:alpha"],
|
||||
min_score=0.55,
|
||||
)
|
||||
```
|
||||
- **Exclude dream-generated speculative content** from injection:
|
||||
```python
|
||||
# Items tagged "dreamed" have lower confidence and may not be suitable for direct injection.
|
||||
# Use min_score to exclude low-confidence dream items, or write dream items to a
|
||||
# separate scope and keep the agent's scope clean.
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
scope="agent/production", # dream items written to "agent/research" stay separate
|
||||
min_score=0.5,
|
||||
)
|
||||
```
|
||||
- **Lessons learned**: `retrieval_tags` is an AND filter — all listed tags must match. Use it to partition a shared scope (e.g. one scope per tenant, tagged by project). `min_score` is a post-retrieval cutoff applied before injection; the k-nearest neighbors are retrieved first, then filtered. Setting `min_score` too high starves the context block entirely; start around 0.4–0.6 and tune empirically.
|
||||
|
||||
---
|
||||
|
||||
## Issue 5: tool_arg_overrides — injecting arguments without modifying tool definitions
|
||||
|
||||
**Keywords**: tool_arg_overrides, inject arguments, tenant_id, api_key injection, runtime override, model-controlled args, silent replacement
|
||||
|
||||
- **Symptom**: A tool (from a library, MCP, or shared codebase) needs a runtime argument like `user_id`, `tenant_id`, or `api_key` injected at call time. You cannot modify the tool's definition, and you don't want the model to be responsible for supplying these values.
|
||||
- **Cause**: By default, `wrap_tool_call` passes the tool request through unchanged. `tool_arg_overrides` is a constructor-time dict mapping tool names to fixed key-value pairs. Before the tool executes, the middleware merges these overrides into `tool_call["args"]` — the model's arguments are kept, but any key in `overrides` is forcibly replaced. This happens regardless of `record_tool_calls`.
|
||||
- **Solution**:
|
||||
- **Inject a fixed tenant ID into one tool**:
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
tool_arg_overrides={
|
||||
"search_knowledge_base": {"tenant_id": "acme-corp"},
|
||||
},
|
||||
)
|
||||
```
|
||||
- **Override multiple tools**:
|
||||
```python
|
||||
ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
tool_arg_overrides={
|
||||
"send_email": {"from_address": "bot@company.com"},
|
||||
"write_to_db": {"db_name": "prod", "schema": "agents"},
|
||||
"call_external": {"api_key": os.environ["EXTERNAL_API_KEY"]},
|
||||
},
|
||||
)
|
||||
```
|
||||
- **Merge semantics**: overrides use `{**tool_args, **overrides}` — the model's supplied values are the base, and the override dict is merged on top. Any key the model passes that is also in `overrides` is silently replaced. The model cannot override the override.
|
||||
- **Interaction with record_tool_calls**: when `record_tool_calls=True`, the recorded `args` field reflects the **merged** args (after overrides), so stored provenance is accurate.
|
||||
- **Limitation**: overrides are static at construction time. If the injected value must change per session (e.g. a per-request `user_id`), use a custom `wrap_tool_call` middleware instead or pass the value through agent state.
|
||||
- **Lessons learned**: `tool_arg_overrides` is best for environment-level constants (API keys, tenant IDs, backend names) that should never be model-controlled. Keep the dict small and document it — it is easy to forget that the model's arg is being silently replaced.
|
||||
|
||||
---
|
||||
|
||||
## Issue 6: dream trigger conditions not met — dream produces zero results
|
||||
|
||||
**Keywords**: dream not running, min_items_for_dream, cooldown_hours, DreamStrategy, DREAM_LLM_ENABLED, dream conditions, DreamReport empty, consolidations empty
|
||||
|
||||
- **Symptom**: Calling `ctx.dream(scope=...)` returns a `DreamReport` with `consolidation.items=[]` and `divergence=None`. Or the contextseek daemon's lifecycle policy shows no dream activity in logs for days.
|
||||
- **Cause**: Two pre-conditions must be met before dream produces results:
|
||||
1. **Minimum item count** — scope must have ≥ `min_items_for_dream` active items (default **10**). A fresh scope with only a handful of entries produces nothing.
|
||||
2. **Divergence requires ≥ 2 tag-based clusters** — the divergence phase groups items by tag and requires at least 2 groups (each with ≥ 2 items). A scope where all items share the same tags produces consolidations only, never divergences. Clusters are built from item tags, not from a retrieval orchestrator.
|
||||
|
||||
Note: `cooldown_hours` (default 6 h) is tracked on the `DreamEngine` instance. Since `ctx.dream()` creates a new engine on every call, the cooldown **does not apply** when calling `ctx.dream()` directly — it only prevents rapid re-triggering inside the daemon lifecycle scheduler.
|
||||
|
||||
Additionally, without `DREAM_LLM_ENABLED=true` the engine uses keyword-overlap heuristics — results are coarser and the similarity window `(0.35, 0.72)` may produce fewer matches.
|
||||
- **Solution**:
|
||||
- **Check scope item count before calling dream**:
|
||||
```python
|
||||
report = ctx.overview(scope="my_scope")
|
||||
print(report.stage_distribution) # shows per-stage item counts
|
||||
# If total is < 10, add more content before expecting consolidations
|
||||
```
|
||||
- **Use dry_run to verify what would be produced**:
|
||||
```python
|
||||
report = ctx.dream(scope="my_scope", dry_run=True)
|
||||
# dry_run=True runs the full analysis but does not write any items
|
||||
consolidation_count = len(report.consolidation.items)
|
||||
divergence_count = len(report.divergence.items) if report.divergence else 0
|
||||
print(f"Would produce: {consolidation_count} consolidations, {divergence_count} divergences")
|
||||
```
|
||||
- **Reduce thresholds in development** (do not use in production):
|
||||
|
||||
`ContextSeekSettings` only exposes `DreamSettings(llm_enabled=bool)`. To override other `DreamStrategy` fields, build the `ContextSeek` instance manually with a custom `StrategyConfig`:
|
||||
```python
|
||||
from dataclasses import replace
|
||||
from contextseek import ContextSeek, ContextSeekSettings
|
||||
from contextseek.config.strategies import DreamStrategy, StrategyConfig
|
||||
|
||||
base_ctx = ContextSeek.from_settings(ContextSeekSettings())
|
||||
ctx = replace(
|
||||
base_ctx,
|
||||
strategy=replace(
|
||||
base_ctx.strategy,
|
||||
dream=DreamStrategy(
|
||||
min_items_for_dream=3, # default: 10
|
||||
divergence_min_clusters=2, # minimum: always 2 (tag-cluster requirement)
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
- **Enable LLM for richer synthesis**:
|
||||
```bash
|
||||
DREAM_LLM_ENABLED=true
|
||||
AGENTSEEK_CTX_LLM_PROVIDER=openai
|
||||
AGENTSEEK_CTX_LLM_MODEL=gpt-4o-mini
|
||||
```
|
||||
Or in code: `ContextSeekSettings(dream=DreamSettings(llm_enabled=True))`
|
||||
- **Lessons learned**: Dream is designed for mature scopes with diverse, multi-topic content. Don't expect it to produce useful output from a handful of items on a single topic. In production, call `ctx.dream()` explicitly after bulk ingestion sessions or at fixed intervals (e.g. weekly) — not after every agent turn.
|
||||
|
||||
---
|
||||
|
||||
## Issue 7: dream-generated items disappear — transient stability and feedback loop
|
||||
|
||||
**Keywords**: dream confidence, stability transient, feedback, use it or lose it, dream divergence, relevance_boost, dream item decay, dreamed tag
|
||||
|
||||
- **Symptom**: `ctx.dream()` produced a valuable cross-topic hypothesis and the model used it successfully in a few sessions. A week later, the same query no longer returns the hypothesis — it has vanished from the store.
|
||||
- **Cause**: All dream-generated items (both consolidation and divergence) are written with `stability=transient`. Transient items have a high decay weight in the lifecycle policy — they lose `relevance_boost` quickly and eventually drop below the retrieval score floor. Without positive feedback, dream items are designed to fade: they are hypotheses that need human or model confirmation to earn permanence.
|
||||
|
||||
Divergence items are additionally written with lower initial confidence (default `dream_initial_confidence=0.35` × `0.85` multiplier = ~0.30), making them especially vulnerable to decay.
|
||||
- **Solution**:
|
||||
- **Promote valuable dream items with feedback**:
|
||||
```python
|
||||
# After a dream run, review the generated items
|
||||
scope = "research/immunology"
|
||||
hits = ctx.retrieve("immune regulation", scope=scope, k=20)
|
||||
for hit in hits.items:
|
||||
# Tags are on hit.item, not on hit directly
|
||||
if "dreamed" in (hit.item.tags or []) and human_or_model_approves(hit.item):
|
||||
# feedback() requires a full URI ref, not a bare item id
|
||||
ref = ctx.resolver.ref_for(scope, hit.item.id)
|
||||
ctx.feedback(ref, scope=scope, score=1.0, reason="confirmed cross-domain insight")
|
||||
# Positive feedback raises relevance_boost, promoting the item toward stable retention
|
||||
```
|
||||
- **Programmatic promotion after model use**: if the model cites a dream item in its answer, treat that as implicit positive feedback:
|
||||
```python
|
||||
# Build the ref from scope + item id, then call feedback()
|
||||
ref = ctx.resolver.ref_for(scope, dreamed_item_id)
|
||||
ctx.feedback(ref, scope=scope, score=0.8, reason="used in model response")
|
||||
```
|
||||
- **Inspect dreamed items before they decay** — `overview()` returns stage distribution counts; filter dreamed items by tag separately:
|
||||
```python
|
||||
all_items = ctx.items(scope="research/immunology")
|
||||
dreamed = [item for item in all_items if "dreamed" in (item.tags or [])]
|
||||
print(f"{len(dreamed)} dreamed items pending review")
|
||||
```
|
||||
- **Adjust initial confidence for dream items** (if you trust the LLM synthesis):
|
||||
```python
|
||||
from dataclasses import replace
|
||||
from contextseek.config.strategies import DreamStrategy
|
||||
|
||||
ctx = replace(
|
||||
ctx,
|
||||
strategy=replace(
|
||||
ctx.strategy,
|
||||
dream=replace(
|
||||
ctx.strategy.dream,
|
||||
dream_initial_confidence=0.6, # default: 0.35 — higher = slower decay
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
- **Lessons learned**: "Use it or lose it" is intentional — unvalidated hypotheses should not accumulate indefinitely. Build a lightweight review step into your research workflow: after each dream run, iterate over `dreamed` items via `ctx.items()`, apply feedback for the useful ones, and let the rest decay naturally.
|
||||
|
||||
---
|
||||
|
||||
## Issue 8: evidence_chain vs chain_confidence — which API to use
|
||||
|
||||
**Keywords**: evidence_chain, chain_confidence, DAG, overall_confidence, critical_path, conflicts, ConflictReport, performance, provenance API choice
|
||||
|
||||
- **Symptom**: You only need to know whether a recommendation is trustworthy enough to act on, but `evidence_chain()` returns a large DAG object that's expensive to process. Or `chain_confidence()` returns a float you can't interpret, with no context about why the confidence is low.
|
||||
- **Cause**: Both methods traverse the same link graph (`derived_from`, `supported_by`, `merged_from`, `refuted_by`). The difference is output granularity:
|
||||
- `chain_confidence(ref, scope)` → `float` — propagated confidence only, no graph construction. Fast.
|
||||
- `evidence_chain(ref, scope, max_depth=10)` → `EvidenceChain` — full DAG with `nodes` (`list[ChainNode]`), `edges` (`list[ChainEdge]`), `overall_confidence`, `critical_path` (`list[str]` of item ids), `conflicts` (`list[ConflictReport]`), `broken_links`, and `needs_reverification` (a `bool`).
|
||||
|
||||
Both methods require a **full URI ref**, not a bare item id. Build the ref with `ctx.resolver.ref_for(scope, item_id)`.
|
||||
- **Solution**:
|
||||
- **Use `chain_confidence` for automated gating** (e.g. deciding whether to act on a recommendation):
|
||||
```python
|
||||
scope = "incidents/2026-06-01"
|
||||
ref = ctx.resolver.ref_for(scope, recommendation_id)
|
||||
score = ctx.chain_confidence(ref, scope=scope)
|
||||
if score < 0.4:
|
||||
flag_for_human_review(recommendation_id)
|
||||
else:
|
||||
apply_recommendation(recommendation_id)
|
||||
```
|
||||
- **Use `evidence_chain` when you need to explain or visualize the reasoning**:
|
||||
```python
|
||||
scope = "incidents/2026-06-01"
|
||||
ref = ctx.resolver.ref_for(scope, recommendation_id)
|
||||
chain = ctx.evidence_chain(ref, scope=scope, max_depth=10)
|
||||
|
||||
print(f"Confidence: {chain.overall_confidence:.2f}")
|
||||
|
||||
# critical_path is a list[str] of item ids — expand to get content
|
||||
path_items = ctx.expand_by_ids(chain.critical_path, scope=scope)
|
||||
print(f"Critical path: {[item.content[:60] for item in path_items]}")
|
||||
|
||||
if chain.conflicts:
|
||||
# conflicts is list[ConflictReport]; each has item_id, refuter_id, refutation_strength
|
||||
print("Conflicting evidence detected:")
|
||||
for c in chain.conflicts:
|
||||
print(f" - item={c.item_id} refuted_by={c.refuter_id} strength={c.refutation_strength:.2f}")
|
||||
|
||||
if chain.broken_links:
|
||||
print(f"Broken references: {len(chain.broken_links)} items no longer exist")
|
||||
|
||||
if chain.needs_reverification:
|
||||
print("Overall confidence below threshold — flag for re-investigation")
|
||||
```
|
||||
- **Reverification threshold**: `chain.needs_reverification` is a `bool` — it is `True` when `overall_confidence` falls below `0.4` (the built-in `reverification_threshold`). Use it as a signal to queue the item for human review before acting on it.
|
||||
- **When no links exist**: if items were written without `links=`, both methods return the item's own `confidence` field unchanged — there is no graph to traverse. Write `links=` on `ctx.add()` calls to enable meaningful provenance.
|
||||
- **Lessons learned**: `chain_confidence` is the right call for high-frequency automated checks (e.g. every retrieved item). `evidence_chain` is for postmortem analysis, audit reports, or dashboards where you need to present the reasoning chain to a human. Avoid calling `evidence_chain` in the hot path of every agent turn.
|
||||
|
||||
---
|
||||
|
||||
## Issue 9: DataPlug vs manual ctx.add() — choosing the right bulk import path
|
||||
|
||||
**Keywords**: DataPlug, plug, bulk import, RAGPlug, PowerMemPlug, batch add, streaming ingest, ctx.add loop, which to use
|
||||
|
||||
- **Symptom**: You need to write 10,000 records to contextseek before deploying an agent. You're unsure whether to loop `ctx.add()` or use `ctx.plug()` with a `DataPlug`, and whether the two differ in how scope, stage, and metadata are handled.
|
||||
- **Cause**: Both paths ultimately call the same internal pipeline (summarizer → embedder → conflict check → persist). The difference is at the **input interface** and **metadata forwarding**:
|
||||
- `ctx.add(content, *, scope, source, ...)` — one item, caller controls all parameters explicitly. Best for custom pipelines where you preprocess each record.
|
||||
- `ctx.plug(plug, *, scope)` — consumes a streaming `DataPlug`; each `RawEvent` can carry `metadata` that overrides `scope`, `stage`, `stability`, `embedding`, `importance`, and `summary` inline. Best for adapting existing data sources.
|
||||
- **Solution**:
|
||||
- **Use `plug()` when adapting an existing source** (RAG, PowerMem, traces):
|
||||
```python
|
||||
from contextseek.plugs import RAGPlug
|
||||
|
||||
plug = RAGPlug(source=vector_store.as_iterable(), source_id="wiki-v2")
|
||||
ctx.plug(plug, scope="company/knowledge")
|
||||
# RAGPlug yields RawEvents; metadata can carry per-item scope/stage overrides
|
||||
```
|
||||
- **Use `ctx.add()` in a loop for custom preprocessing**:
|
||||
```python
|
||||
for record in my_data_source:
|
||||
processed = preprocess(record)
|
||||
ctx.add(
|
||||
processed.text,
|
||||
scope="company/knowledge",
|
||||
source=record.source_id,
|
||||
tags=record.project_tags,
|
||||
stage="raw",
|
||||
)
|
||||
```
|
||||
- **Override stage per event with plug()** (e.g. import pre-summarized items directly as `extracted`):
|
||||
```python
|
||||
from contextseek.protocols.plugs import RawEvent
|
||||
|
||||
events = [
|
||||
RawEvent(
|
||||
content=item.text,
|
||||
source=item.source_id,
|
||||
metadata={"stage": "extracted", "summary": item.existing_summary},
|
||||
)
|
||||
for item in pre_summarized_items
|
||||
]
|
||||
# Wrap as a simple DataPlug
|
||||
class ListPlug:
|
||||
def stream(self): return iter(events)
|
||||
def metadata(self): return PlugMeta(name="pre-summarized", source_type="knowledge")
|
||||
|
||||
ctx.plug(ListPlug(), scope="company/knowledge")
|
||||
```
|
||||
- **Lessons learned**: `plug()` is the idiomatic path for integrating existing external sources. It handles the `PlugMeta → SourceType` mapping and per-event metadata merging out of the box. Manual `ctx.add()` loops are appropriate when you have non-standard preprocessing logic, need precise control over each item's provenance fields, or are writing fewer than a few hundred items.
|
||||
|
||||
---
|
||||
|
||||
## Issue 10: plug() scope priority and stage inference behavior
|
||||
|
||||
**Keywords**: plug scope priority, event metadata scope, DataPlug name, stage override, RAG stage, TracePlug, scope assignment, unexpected scope
|
||||
|
||||
- **Symptom**: After calling `ctx.plug(rag_plug, scope="project:x")`, some items end up in a different scope. Or items from a `TracePlug` import end up with `stage=raw` when you expected `stage=extracted`.
|
||||
- **Cause**:
|
||||
- **Scope resolution priority**: `plug(scope=...)` parameter → `event.metadata["scope"]` → `plug.metadata().name`. Passing `scope=` in the `plug()` call forces **all** events to that scope, regardless of what individual events declare. If you omit `scope=` from `plug()`, each event can control its own scope via `metadata`.
|
||||
- **Stage inference**: if neither `event.metadata["stage"]` nor a plug-level default is set, `plug()` falls back to `stage="raw"` for all items. `TracePlug` and skill importers set their own default stages (`trace_extraction` → `raw`; `MCPToolImporter` → `skill`), but `RAGPlug` defaults to `raw` unless the source metadata carries a `stage` key.
|
||||
- **Solution**:
|
||||
- **Force all events to one scope** (recommended for most bulk imports):
|
||||
```python
|
||||
ctx.plug(rag_plug, scope="company/knowledge")
|
||||
# All events land in "company/knowledge" regardless of event.metadata["scope"]
|
||||
```
|
||||
- **Allow per-event scope** (when source data has meaningful sub-scopes):
|
||||
```python
|
||||
# Omit scope= in plug(); events control their own scope via metadata
|
||||
ctx.plug(rag_plug)
|
||||
# Each RawEvent.metadata["scope"] = "dept/engineering" or "dept/hr" etc.
|
||||
```
|
||||
- **Set stage at the event level** to skip redundant summarization for pre-processed content:
|
||||
```python
|
||||
RawEvent(
|
||||
content=chunk.text,
|
||||
source="wiki",
|
||||
metadata={"stage": "extracted", "summary": chunk.summary, "scope": "company/knowledge"},
|
||||
)
|
||||
```
|
||||
- **Verify what landed after a plug run**:
|
||||
```python
|
||||
items = ctx.items(scope="company/knowledge")
|
||||
stage_counts = Counter(item.stage.value for item in items)
|
||||
print(stage_counts) # e.g. {"raw": 9500, "extracted": 500}
|
||||
```
|
||||
- **Lessons learned**: when in doubt, always pass `scope=` explicitly to `plug()` — it is the simplest way to ensure all imported items end up in the right bucket. Per-event scope override is a power-user feature for heterogeneous source data where different records naturally belong in different scopes.
|
||||
|
||||
---
|
||||
|
||||
## Issue 11: auto_dream — automatic dream triggering inside the agent loop
|
||||
|
||||
**Keywords**: auto_dream, dream_every, dream_min_interval_seconds, middleware dream, automatic dream, dual gate, fire-and-forget dream
|
||||
|
||||
- **Symptom**: You want the agent's knowledge store to automatically discover cross-topic patterns without calling `ctx.dream()` explicitly after every session. Or you enabled `auto_dream=True` but dream never seems to fire, even after many sessions.
|
||||
- **Cause**: `ContextSeekMiddleware` supports an optional automatic dream trigger alongside `auto_compact`. It uses a **dual gate** to avoid over-triggering: both conditions must be satisfied before dream fires:
|
||||
1. **Turn counter gate** — the scope's internal counter must reach `dream_every` (default **200**). So if `dream_every=200` and the agent handles 10 sessions per day, dream fires once every 20 days per scope — much less frequently than compact.
|
||||
2. **Time gate** — at least `dream_min_interval_seconds` must have elapsed since the last dream triggered by this middleware instance (default **3600 s = 1 h**). This prevents back-to-back firing even if the counter is satisfied repeatedly.
|
||||
|
||||
Both gates must pass in the same `after_agent` call. The dream task is submitted fire-and-forget (same `ThreadPoolExecutor` as compact) — it does **not** block the agent response.
|
||||
|
||||
Note: this is separate from the cooldown in `DreamStrategy.cooldown_hours`, which is tracked on the `DreamEngine` instance and does not apply when `ContextSeek.dream()` is called directly. The middleware's `dream_min_interval_seconds` is tracked on the middleware instance.
|
||||
- **Solution**:
|
||||
- **Enable automatic dream with sensible defaults**:
|
||||
```python
|
||||
middleware = ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_compact=True,
|
||||
compact_every=20,
|
||||
auto_dream=True,
|
||||
dream_every=200, # fire after every 200 agent turns per scope
|
||||
dream_min_interval_seconds=3600.0, # and at least 1 hour since last dream
|
||||
)
|
||||
```
|
||||
- **For high-traffic services** (many turns per day) — lower the turn threshold:
|
||||
```python
|
||||
middleware = ContextSeekMiddleware(
|
||||
model=model,
|
||||
embedder=embedder,
|
||||
auto_dream=True,
|
||||
dream_every=50, # fire more frequently
|
||||
dream_min_interval_seconds=7200.0, # but no more than once every 2 hours
|
||||
)
|
||||
```
|
||||
- **Graceful shutdown** — `middleware.shutdown(wait=True)` waits for both in-flight compact **and** dream tasks:
|
||||
```python
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
yield
|
||||
middleware.shutdown(wait=True)
|
||||
```
|
||||
- **Check if dream has fired** — there is no built-in status probe. Use `LANGSMITH_TRACING=true` to see `ContextSeek.dream` spans, or inspect `ctx.items(scope=..., stage=Stage.extracted)` for items tagged `dreamed`.
|
||||
- **Manually trigger dream outside the middleware loop** (recommended for research / bulk-import workflows):
|
||||
```python
|
||||
# Bypass both gates; always runs immediately
|
||||
report = middleware.ctx.dream(scope="my_scope")
|
||||
print(f"New items: {report.total_dream_items}")
|
||||
```
|
||||
- **Lessons learned**: `auto_dream` is designed for long-running conversational agents that accumulate hundreds of sessions. For lower-volume or batch scenarios (e.g. a research agent that ingests documents in bursts), explicit `ctx.dream()` calls after each ingestion batch give more control over timing and cost. `auto_dream=True` and explicit `ctx.dream()` calls can coexist — they use the same underlying `DreamEngine`.
|
||||
@@ -0,0 +1,245 @@
|
||||
# Deep Agents Development Issues
|
||||
|
||||
## Issue 1: Model choice has a huge impact on Agent capability
|
||||
|
||||
- **Symptom**: With the same Agent code, swapping the model produces wildly different capability — some models handle tool calling fine, others can't complete the task at all.
|
||||
- **Cause**: Deep Agents have high requirements on the model, which must have stable tool-calling ability. Different models perform very differently across dimensions like File Ops, Retrieval, Tool Use, Memory, Conversation, Summarization (see the Deep Agents eval suite).
|
||||
- **Solution**:
|
||||
- Prefer high-scoring models: `google_genai:gemini-3.5-flash` (Overall 82%), `openai:gpt-5.5` (80%), `anthropic:claude-opus-4-7` (80%)
|
||||
- Among open-source models, `GLM-5.1` (via OpenRouter/Fireworks) performs best (89%)
|
||||
- Watch for per-dimension weaknesses: even with a high Overall, Conversation and Memory scores tend to be low (most models < 50%); long-conversation scenarios need extra validation
|
||||
- Use `init_chat_model` to fine-tune parameters (e.g. `thinking_level`) to improve some models
|
||||
- **Lessons learned**: Don't pick by brand alone — you must run the eval. Different models from the same provider can vary widely (e.g. `gpt-5.4` Overall only 18%, while `gpt-5.5` hits 80%).
|
||||
|
||||
## Issue 2: Hard to pick a filesystem Backend
|
||||
|
||||
- **Symptom**: Deep Agents ships 6 filesystem backends (StateBackend, FilesystemBackend, StoreBackend, ContextHubBackend, LocalShellBackend, CompositeBackend), and it's unclear which one fits your scenario.
|
||||
- **Cause**: Each backend has very different persistence scope, isolation level, and security model, and the docs don't give a clear decision path.
|
||||
- **Solution**:
|
||||
- **Need cross-session persistence** → use `StoreBackend` (with a LangGraph store) or `ContextHubBackend` (LangSmith Hub)
|
||||
- **Need to operate on local project files** → use `CompositeBackend` to route: hand the project directory to `FilesystemBackend`, keep internal temporary data in `StateBackend`:
|
||||
```python
|
||||
from deepagents.backends import CompositeBackend, StateBackend, FilesystemBackend
|
||||
|
||||
backend = CompositeBackend(
|
||||
default=StateBackend(), # Agent internal data (temporary)
|
||||
routes={
|
||||
"/workspace/": FilesystemBackend(root_dir="/path/to/project", virtual_mode=True),
|
||||
},
|
||||
)
|
||||
```
|
||||
- **Multi-user isolation** → `StoreBackend` must be configured with a `namespace` factory function for data isolation:
|
||||
```python
|
||||
from deepagents.backends import StoreBackend
|
||||
|
||||
backend = StoreBackend(
|
||||
namespace=lambda rt: (rt.server_info.user.identity,),
|
||||
)
|
||||
```
|
||||
- **Need to execute shell commands** → use `LocalShellBackend` (development only) or a Sandbox Backend (production). Note: `LocalShellBackend` has no isolation, so the Agent can run arbitrary commands
|
||||
- **Security** → `FilesystemBackend` must enable `virtual_mode=True` to block path traversal (`..`, `~`, absolute paths). The default `virtual_mode=False` provides no safety guarantees even with `root_dir` set
|
||||
- **Lessons learned**: Most scenarios should use `CompositeBackend` to compose routes rather than a single Backend. Using `FilesystemBackend` or `StateBackend` alone each has obvious downsides — the former pollutes disk, the latter doesn't persist.
|
||||
|
||||
## Issue 3: How to disable the default general-purpose sub-agent
|
||||
|
||||
- **Symptom**: After creating a Deep Agent, even without any configured `subagents`, the Agent still automatically has a sub-agent named `general-purpose` and a corresponding `task` tool. In some scenarios you don't want delegation capability, but can't find the off switch.
|
||||
- **Cause**: Deep Agents injects a synchronous `general-purpose` sub-agent by default (inheriting the main Agent's tools, skills, and model). As long as at least one synchronous sub-agent exists, `SubAgentMiddleware` is attached and exposes the `task` tool.
|
||||
- **Solution**:
|
||||
- Set `general_purpose_subagent.enabled = False` in the harness profile, while keeping custom sub-agents:
|
||||
```python
|
||||
from deepagents import create_deep_agent
|
||||
from deepagents.profiles import HarnessProfile, GeneralPurposeSubagentProfile
|
||||
|
||||
profile = HarnessProfile(
|
||||
general_purpose_subagent=GeneralPurposeSubagentProfile(enabled=False),
|
||||
)
|
||||
|
||||
agent = create_deep_agent(
|
||||
model="anthropic:claude-sonnet-4-6",
|
||||
tools=[my_tool],
|
||||
harness_profile=profile,
|
||||
subagents=[research_subagent, code_subagent], # Custom sub-agents still work
|
||||
)
|
||||
```
|
||||
- Note: don't try to disable via `excluded_middleware=["SubAgentMiddleware"]` — this raises `ValueError` directly, and would also disable custom sub-agents
|
||||
- **Lessons learned**: If you only want to replace default behavior rather than fully disable it, pass a custom sub-agent with `name="general-purpose"` to override the default config.
|
||||
|
||||
## Issue 4: How to set filesystem permissions
|
||||
|
||||
- **Symptom**: The Agent can use built-in filesystem tools to read/write any file. You need to restrict its scope but aren't sure how to configure it. Or you configured permission rules but the Agent can still access paths that should be blocked.
|
||||
- **Cause**: Permissions are configured via the `permissions` parameter of `create_deep_agent`, not on the Backend. Rules use a **first-match-wins** strategy (matched top to bottom, stop at the first hit), and when no rule matches the default is **allow**. Wrong rule order or missing a catch-all deny rule will silently break permission enforcement.
|
||||
- **Solution**:
|
||||
- **Restrict the Agent to a workspace and protect sensitive files**:
|
||||
```python
|
||||
from deepagents import FilesystemPermission, create_deep_agent
|
||||
|
||||
agent = create_deep_agent(
|
||||
model=model,
|
||||
backend=backend,
|
||||
permissions=[
|
||||
# First deny sensitive files (specific rules come first)
|
||||
FilesystemPermission(
|
||||
operations=["read", "write"],
|
||||
paths=["/workspace/.env", "/workspace/secrets/**"],
|
||||
mode="deny",
|
||||
),
|
||||
# Then allow the workspace
|
||||
FilesystemPermission(
|
||||
operations=["read", "write"],
|
||||
paths=["/workspace/**"],
|
||||
mode="allow",
|
||||
),
|
||||
# Catch-all: deny everything else
|
||||
FilesystemPermission(
|
||||
operations=["read", "write"],
|
||||
paths=["/**"],
|
||||
mode="deny",
|
||||
),
|
||||
],
|
||||
)
|
||||
```
|
||||
- **Read-only Agent (block all writes)**:
|
||||
```python
|
||||
agent = create_deep_agent(
|
||||
model=model,
|
||||
backend=backend,
|
||||
permissions=[
|
||||
FilesystemPermission(
|
||||
operations=["write"],
|
||||
paths=["/**"],
|
||||
mode="deny",
|
||||
),
|
||||
],
|
||||
)
|
||||
```
|
||||
- **Sub-agent with different permissions** (sub-agents inherit parent permissions by default; setting `permissions` fully replaces rather than merges):
|
||||
```python
|
||||
agent = create_deep_agent(
|
||||
model=model,
|
||||
backend=backend,
|
||||
permissions=[
|
||||
FilesystemPermission(operations=["read", "write"], paths=["/workspace/**"], mode="allow"),
|
||||
FilesystemPermission(operations=["read", "write"], paths=["/**"], mode="deny"),
|
||||
],
|
||||
subagents=[
|
||||
{
|
||||
"name": "auditor",
|
||||
"description": "Read-only code reviewer",
|
||||
"system_prompt": "Review the code for issues.",
|
||||
"permissions": [
|
||||
# Fully replaces parent permissions: only allow reads in workspace
|
||||
FilesystemPermission(operations=["write"], paths=["/**"], mode="deny"),
|
||||
FilesystemPermission(operations=["read"], paths=["/workspace/**"], mode="allow"),
|
||||
FilesystemPermission(operations=["read"], paths=["/**"], mode="deny"),
|
||||
],
|
||||
}
|
||||
],
|
||||
)
|
||||
```
|
||||
- **CompositeBackend + sandbox**: when default is sandbox, `paths` must fall under a known route prefix, otherwise `NotImplementedError` is raised:
|
||||
```python
|
||||
from deepagents.backends import CompositeBackend
|
||||
|
||||
composite = CompositeBackend(
|
||||
default=sandbox,
|
||||
routes={"/memories/": memories_backend},
|
||||
)
|
||||
|
||||
# Correct: permission path is under a route prefix
|
||||
agent = create_deep_agent(
|
||||
model=model,
|
||||
backend=composite,
|
||||
permissions=[
|
||||
FilesystemPermission(operations=["write"], paths=["/memories/**"], mode="deny"),
|
||||
],
|
||||
)
|
||||
|
||||
# Wrong: /workspace/** hits the sandbox default and raises NotImplementedError
|
||||
# FilesystemPermission(operations=["write"], paths=["/workspace/**"], mode="deny")
|
||||
```
|
||||
- Note: permissions only affect built-in filesystem tools (`ls`, `read_file`, `glob`, `grep`, `write_file`, `edit_file`). Custom tools and MCP tools are not constrained
|
||||
- **Lessons learned**: The most common mistake is inverted rule order — putting a broad allow before deny means deny will never trigger. And since the default when no rule matches is allow, missing a catch-all deny is equivalent to having no permission control at all.
|
||||
|
||||
## Issue 5: How to configure long-term Agent memory
|
||||
|
||||
- **Symptom**: The Agent "forgets" between every conversation and can't remember user preferences or prior context. Or memory is configured, but memory leaks between multiple users.
|
||||
- **Cause**: Deep Agents' long-term memory is built on the filesystem — the Agent specifies a memory file path via the `memory=` parameter and uses the Backend to control storage location and isolation scope. Without a persistent Backend (like `StoreBackend`), memory exists only in single-session State. If the namespace isn't isolated by user, all users share the same memory file.
|
||||
- **Solution**:
|
||||
- **User-scoped isolated memory** (each user has independent, mutually invisible memory):
|
||||
```python
|
||||
from deepagents import create_deep_agent
|
||||
from deepagents.backends import CompositeBackend, StateBackend, StoreBackend
|
||||
|
||||
agent = create_deep_agent(
|
||||
model="google_genai:gemini-3.5-flash",
|
||||
memory=["/memories/preferences.md"],
|
||||
backend=CompositeBackend(
|
||||
default=StateBackend(),
|
||||
routes={
|
||||
"/memories/": StoreBackend(
|
||||
namespace=lambda rt: (rt.server_info.user.identity,),
|
||||
),
|
||||
},
|
||||
),
|
||||
)
|
||||
```
|
||||
- **Agent-scoped shared memory** (all users share the same Agent knowledge):
|
||||
```python
|
||||
agent = create_deep_agent(
|
||||
model="google_genai:gemini-3.5-flash",
|
||||
memory=["/memories/AGENTS.md"],
|
||||
backend=CompositeBackend(
|
||||
default=StateBackend(),
|
||||
routes={
|
||||
"/memories/": StoreBackend(
|
||||
namespace=lambda rt: (rt.server_info.assistant_id,),
|
||||
),
|
||||
},
|
||||
),
|
||||
)
|
||||
```
|
||||
- **Org-scoped read-only policy** (shared across users but not modifiable by the Agent, to prevent prompt injection from polluting shared state):
|
||||
```python
|
||||
from deepagents import FilesystemPermission
|
||||
|
||||
agent = create_deep_agent(
|
||||
model="google_genai:gemini-3.5-flash",
|
||||
memory=["/memories/preferences.md", "/policies/compliance.md"],
|
||||
backend=CompositeBackend(
|
||||
default=StateBackend(),
|
||||
routes={
|
||||
"/memories/": StoreBackend(
|
||||
namespace=lambda rt: (rt.server_info.user.identity,),
|
||||
),
|
||||
"/policies/": StoreBackend(
|
||||
namespace=lambda rt: (rt.context.org_id,),
|
||||
),
|
||||
},
|
||||
),
|
||||
permissions=[
|
||||
FilesystemPermission(operations=["write"], paths=["/policies/**"], mode="deny"),
|
||||
],
|
||||
)
|
||||
```
|
||||
- **Lessons learned**: The core design decision for memory is choosing the namespace — it determines "who can see what." User-scoped uses `user.identity`, Agent-scoped uses `assistant_id`, org-scoped uses `org_id`. Shared memory must be read-only, otherwise there's a cross-user prompt-injection risk.
|
||||
|
||||
## Issue 6: A long `SKILL.md` gets silently truncated — the Agent only reads the first 100 lines
|
||||
|
||||
- **Symptom**: You wrote a 300+ line `SKILL.md`, but the Agent behaves as if it never saw the later sections (workflows, examples, edge cases at the bottom are ignored). No error is raised — the skill just appears to "half work," and it takes a long time to realize content is missing rather than wrong.
|
||||
- **Cause**: The built-in `read_file` tool defaults to reading only **100 lines** (`DEFAULT_READ_LIMIT = 100` in `deepagents/middleware/filesystem.py`). Skills use *progressive disclosure*: only the name + description are injected into the system prompt, and the Agent is expected to `read_file` the full `SKILL.md` on demand. When the file exceeds 100 lines, that default read silently cuts it off — the tail is never loaded into context, and nothing flags the truncation.
|
||||
- **Solution**:
|
||||
- **Tell the model to pass an explicit `limit`.** The official `SKILLS_SYSTEM_PROMPT` already instructs this — its progressive-disclosure step reads: *"Use `read_file` on the path… Pass `limit=1000` since the default of 100 lines is too small for most skill files."* If you override `SkillsMiddleware(system_prompt=...)` with your own template, **keep that `limit=1000` instruction** (or higher) or you reintroduce the bug. If you see truncation in practice, bump the limit in the prompt (e.g. `limit=500`/`1000`) to cover your largest skill file:
|
||||
```python
|
||||
from deepagents.middleware.skills import SkillsMiddleware
|
||||
|
||||
# When customizing the prompt, preserve the explicit-limit guidance and
|
||||
# the three required slots: {skills_locations} {skills_load_warnings} {skills_list}
|
||||
middleware = SkillsMiddleware(
|
||||
backend=backend,
|
||||
sources=["/skills/"],
|
||||
system_prompt=my_template, # must tell the model: read_file(..., limit=1000)
|
||||
)
|
||||
```
|
||||
- **Or keep `SKILL.md` short and offload detail to reference files** (the pattern this very skill uses): a lean `SKILL.md` that fits in ~100 lines plus a `reference/` directory the Agent reads only when a section applies. This sidesteps the limit entirely and keeps the always-loaded metadata cheap.
|
||||
- Note: `limit` counts *source* lines; lines longer than 5,000 chars are split with continuation markers (`5.1`, `5.2`, …) that do **not** consume the budget. For very large files, paginate with `offset` (`read_file(file_path=..., offset=100, limit=200)`).
|
||||
- **Lessons learned**: This failure is insidious because it's silent — there's no error, the skill just under-performs, so the instinct is to blame the prompt wording or the model rather than a truncated read. Two preventions: (1) keep `SKILL.md` lean and push depth into `reference/` files, and (2) if a skill file must be long, ensure the system prompt forces a large `limit`. Treat 100 lines as a hard default ceiling on anything the Agent auto-reads.
|
||||
@@ -0,0 +1,272 @@
|
||||
# Middleware Development Issues
|
||||
|
||||
## Issue 1: Middleware execution order is counter-intuitive
|
||||
|
||||
- **Symptom**: When multiple middlewares are composed, the actual execution order of `before_model` differs from what you expected, causing state to be unexpectedly overwritten or logic to fail.
|
||||
- **Cause**: The execution order of the middleware list follows the onion model, with different rules for each of the three hook types:
|
||||
- `before_*` hooks: executed in list order (first → last)
|
||||
- `after_*` hooks: executed in **reverse** list order (last → first)
|
||||
- `wrap_*` hooks: nested wrapping (first wraps all others, innermost executes last)
|
||||
- **Solution**:
|
||||
```python
|
||||
agent = create_agent(
|
||||
model="gpt-5.4",
|
||||
middleware=[middleware1, middleware2, middleware3],
|
||||
tools=[...],
|
||||
)
|
||||
# Actual execution flow:
|
||||
# 1. middleware1.before_model()
|
||||
# 2. middleware2.before_model()
|
||||
# 3. middleware3.before_model()
|
||||
# 4. middleware1.wrap_model_call → middleware2.wrap_model_call → middleware3.wrap_model_call → model
|
||||
# 5. middleware3.after_model() ← note the reverse order!
|
||||
# 6. middleware2.after_model()
|
||||
# 7. middleware1.after_model()
|
||||
```
|
||||
- `before_agent` / `after_agent` follow the same rule: before in order, after in reverse
|
||||
- Key principle: things that need to intercept earliest go at the front of the list (rate limiting, permission checks); things that need to be the last fallback also go at the front (since wrap nesting puts them outermost)
|
||||
- **Lessons learned**: The nesting nature of `wrap_model_call` means the first middleware in the list both sees the request first and the response last. Place retry logic at the front of the list (outermost), logging in the middle or back.
|
||||
|
||||
## Issue 2: state_schema merge behavior and input/output control
|
||||
|
||||
- **Symptom**: Multiple middlewares each declare a `state_schema`, and it's unclear how the final state is merged. Or some fields are intermediate state you don't want exposed to callers, some need to be input-only and not in output, some appear only in output.
|
||||
- **Cause**: Internally `create_agent` merges all middleware `state_schema`s in registration order, finally merging the `create_agent`'s `state_schema` parameter (if any). The merge rule is: later declarations override earlier ones for the same field name (base_state is merged last, giving it the highest priority). After merging, `OmitFromSchema` annotations are used to generate the InputSchema and OutputSchema separately.
|
||||
- **Solution**:
|
||||
- **Merge order**: `[middleware1.state_schema, middleware2.state_schema, ..., base_state]`, with later overriding earlier for same-named fields:
|
||||
```python
|
||||
from langchain.agents import create_agent
|
||||
from langchain.agents.middleware import AgentState, AgentMiddleware
|
||||
from typing_extensions import NotRequired
|
||||
|
||||
class MiddlewareAState(AgentState):
|
||||
counter: NotRequired[int]
|
||||
|
||||
class MiddlewareBState(AgentState):
|
||||
trace_id: NotRequired[str]
|
||||
|
||||
class MyState(AgentState):
|
||||
user_id: NotRequired[str]
|
||||
|
||||
# Final merged state = messages + counter + trace_id + user_id
|
||||
# If there's a same-named field, MyState (base_state) takes priority
|
||||
agent = create_agent(
|
||||
model="gpt-5.4",
|
||||
middleware=[middleware_a, middleware_b],
|
||||
state_schema=MyState,
|
||||
)
|
||||
```
|
||||
- **Control field input/output visibility** — use the `OmitFromSchema` annotation:
|
||||
```python
|
||||
from typing import Annotated
|
||||
from langchain.agents.middleware import AgentState, OmitFromSchema
|
||||
from typing_extensions import NotRequired
|
||||
|
||||
class MyState(AgentState):
|
||||
# Only appears in input, not in output (e.g. config parameters passed by the user)
|
||||
user_preference: NotRequired[Annotated[str, OmitFromSchema(output=True)]]
|
||||
|
||||
# Only appears in output, no need for caller to pass in (e.g. result produced by the Agent)
|
||||
structured_response: NotRequired[Annotated[dict, OmitFromSchema(input=True)]]
|
||||
|
||||
# Intermediate state: neither in input nor output (pure internal flow)
|
||||
internal_step_count: NotRequired[Annotated[int, OmitFromSchema(input=True, output=True)]]
|
||||
|
||||
# Normal field: visible in both input and output
|
||||
messages: ... # inherited from AgentState
|
||||
```
|
||||
- **Actual behavior**:
|
||||
- `OmitFromSchema(output=True)`: caller passes it in via `invoke({"user_preference": "concise"})`, but the field is not in the returned result
|
||||
- `OmitFromSchema(input=True)`: caller doesn't need to pass it; produced during Agent execution, appears in the returned result
|
||||
- `OmitFromSchema(input=True, output=True)`: pure intermediate state, middlewares pass data via state, completely invisible to the outside
|
||||
- **Same-name field conflicts**: later merged overrides earlier merged. If two middlewares declare a same-named field with different types, no error is raised but behavior is unpredictable. Use field prefixes to avoid this:
|
||||
```python
|
||||
class RateLimitState(AgentState):
|
||||
ratelimit_count: NotRequired[int] # prefix isolation
|
||||
|
||||
class AuditState(AgentState):
|
||||
audit_last_tool: NotRequired[str] # prefix isolation
|
||||
```
|
||||
- **Lessons learned**: `OmitFromSchema` is the key mechanism that distinguishes "external interface" from "internal state". Fields without this annotation are visible in both InputSchema and OutputSchema by default. Counters, flags, and other middleware-produced fields should be marked `OmitFromSchema(input=True, output=True)` to avoid polluting the caller's interface.
|
||||
|
||||
## Issue 3: The resume value for the Human-in-the-loop middleware
|
||||
|
||||
- **Symptom**: After using `HumanInTheLoopMiddleware` the Agent interrupts successfully, but it's unclear what value to pass on resume; or after passing the value, the Agent behaves unexpectedly (e.g. still executes original args after edit, model doesn't receive feedback after reject).
|
||||
- **Cause**: After `HumanInTheLoopMiddleware` interrupts, you need to resume execution via `Command(resume=...)`. The resume value has the structure `{"decisions": [...]}` (**plural, array**), where each element uses a `"type"` field to specify the decision type. Common mistakes include using the singular `{"decision": "approve"}` form, or forgetting `version="v2"` and losing access to interrupt info.
|
||||
- **Solution**:
|
||||
- **Configure interrupt rules**: use `interrupt_on` to specify which tools require human approval. Values can be `True` (all decision types allowed), `False` (auto-approve), or an `InterruptOnConfig` object:
|
||||
```python
|
||||
from langchain.agents import create_agent
|
||||
from langchain.agents.middleware import HumanInTheLoopMiddleware
|
||||
from langgraph.checkpoint.memory import InMemorySaver
|
||||
|
||||
agent = create_agent(
|
||||
model="gpt-5.4",
|
||||
tools=[read_email_tool, send_email_tool, ask_user_tool],
|
||||
checkpointer=InMemorySaver(), # checkpointer is required
|
||||
middleware=[
|
||||
HumanInTheLoopMiddleware(
|
||||
interrupt_on={
|
||||
"send_email_tool": True, # allow all decision types (approve/edit/reject/respond)
|
||||
"ask_user_tool": {"allowed_decisions": ["respond"]}, # only allow respond
|
||||
"read_email_tool": False, # safe operation, no interrupt
|
||||
},
|
||||
description_prefix="Tool execution pending approval",
|
||||
),
|
||||
],
|
||||
)
|
||||
```
|
||||
- **Get interrupt info**: invoke with `version="v2"`. The returned `GraphOutput` contains a `.interrupts` attribute with `action_requests` (details of pending tool calls) and `review_configs` (allowed decision types per tool):
|
||||
```python
|
||||
config = {"configurable": {"thread_id": "thread-1"}}
|
||||
|
||||
result = agent.invoke(
|
||||
{"messages": [{"role": "user", "content": "Send the report to the team"}]},
|
||||
config=config,
|
||||
version="v2", # must specify v2 to access interrupts
|
||||
)
|
||||
|
||||
# result.interrupts contains interrupt details
|
||||
# Interrupt(value={
|
||||
# 'action_requests': [
|
||||
# {'name': 'send_email_tool', 'arguments': {...}, 'description': '...'}
|
||||
# ],
|
||||
# 'review_configs': [
|
||||
# {'action_name': 'send_email_tool', 'allowed_decisions': ['approve', 'edit', 'reject', 'respond']}
|
||||
# ]
|
||||
# })
|
||||
```
|
||||
- **Resume value format** (four decision types):
|
||||
```python
|
||||
from langgraph.types import Command
|
||||
|
||||
# approve: approve directly, execute the original tool call
|
||||
agent.invoke(
|
||||
Command(resume={"decisions": [{"type": "approve"}]}),
|
||||
config=config,
|
||||
version="v2",
|
||||
)
|
||||
|
||||
# reject: reject execution; message becomes feedback to help the model re-plan
|
||||
agent.invoke(
|
||||
Command(resume={"decisions": [
|
||||
{"type": "reject", "message": "Don't delete data; archive to the history table instead"}
|
||||
]}),
|
||||
config=config,
|
||||
version="v2",
|
||||
)
|
||||
|
||||
# edit: modify tool call args before executing (use edited_action to specify new tool name and args)
|
||||
agent.invoke(
|
||||
Command(resume={"decisions": [
|
||||
{
|
||||
"type": "edit",
|
||||
"edited_action": {
|
||||
"name": "send_email_tool", # usually same as the original tool
|
||||
"args": {"recipient": "boss@company.com", "subject": "Updated subject"},
|
||||
},
|
||||
}
|
||||
]}),
|
||||
config=config,
|
||||
version="v2",
|
||||
)
|
||||
|
||||
# respond: skip tool execution; human reply becomes the tool result directly (for ask_user-type tools)
|
||||
agent.invoke(
|
||||
Command(resume={"decisions": [{"type": "respond", "message": "Use a blue theme"}]}),
|
||||
config=config,
|
||||
version="v2",
|
||||
)
|
||||
```
|
||||
- **Multiple tool calls interrupted simultaneously**: when the model returns multiple tool_calls needing approval in one go, the order in the decisions array must correspond one-to-one with the order in `action_requests`:
|
||||
```python
|
||||
agent.invoke(
|
||||
Command(resume={"decisions": [
|
||||
{"type": "approve"}, # first tool: approve
|
||||
{"type": "reject", "message": "Not allowed"}, # second tool: reject
|
||||
]}),
|
||||
config=config,
|
||||
version="v2",
|
||||
)
|
||||
```
|
||||
- **Checkpointer is required**: without a checkpointer, state can't be restored after interrupt. `InMemorySaver` is for development; production uses persistent storage like `AsyncPostgresSaver`
|
||||
- **Lessons learned**: The most common mistake is getting the resume structure wrong — remember it's `{"decisions": [{"type": "..."}]}` (plural + array + type field), not `{"decision": "..."}`. The `edit` args go in `edited_action`, not at the top level. `respond` fits "ask user"-style tools: the human reply becomes a ToolMessage returned to the model, and the tool itself doesn't execute.
|
||||
|
||||
## Issue 4: Dynamically modifying state inside wrap_model_call
|
||||
|
||||
- **Symptom**: You need to update state in `wrap_model_call` based on the model response (e.g. record token usage, trigger summarization), but returning `ModelResponse` directly can't carry state updates.
|
||||
- **Cause**: The return type of `wrap_model_call` is `ModelResponse` (i.e. the model's AIMessage) by default. Unlike node-style hooks, you can't directly return a dict that merges into state. To inject state updates from the wrap layer, you need to return `ExtendedModelResponse`.
|
||||
- **Solution**:
|
||||
```python
|
||||
from typing import Callable
|
||||
from langchain.agents.middleware import (
|
||||
wrap_model_call,
|
||||
AgentState,
|
||||
ModelRequest,
|
||||
ModelResponse,
|
||||
ExtendedModelResponse,
|
||||
)
|
||||
from langgraph.types import Command
|
||||
from typing_extensions import NotRequired
|
||||
|
||||
class UsageState(AgentState):
|
||||
last_model_tokens: NotRequired[int]
|
||||
|
||||
@wrap_model_call(state_schema=UsageState)
|
||||
def track_usage(
|
||||
request: ModelRequest,
|
||||
handler: Callable[[ModelRequest], ModelResponse],
|
||||
) -> ExtendedModelResponse:
|
||||
response = handler(request)
|
||||
# Inject state updates via Command(update=...)
|
||||
return ExtendedModelResponse(
|
||||
model_response=response,
|
||||
command=Command(update={"last_model_tokens": 150}),
|
||||
)
|
||||
```
|
||||
- **Command composition rules across multiple middlewares**:
|
||||
- Commands are applied via graph reducers — the messages field is append-style
|
||||
- Non-reducer fields (regular int/str): inner is applied first, outer last, **outer overrides inner**
|
||||
- If the outer layer has retry logic (calling `handler()` multiple times), commands from earlier calls are discarded
|
||||
- **Dynamically modify the system prompt** (the most common use of wrap_model_call):
|
||||
```python
|
||||
from langchain.agents.middleware import wrap_model_call, ModelRequest, ModelResponse
|
||||
from langchain.messages import SystemMessage
|
||||
from typing import Callable
|
||||
|
||||
@wrap_model_call
|
||||
def inject_context(
|
||||
request: ModelRequest,
|
||||
handler: Callable[[ModelRequest], ModelResponse],
|
||||
) -> ModelResponse:
|
||||
# request.system_message is always a SystemMessage object
|
||||
new_content = list(request.system_message.content_blocks) + [
|
||||
{"type": "text", "text": "Current user preference: concise answers"}
|
||||
]
|
||||
return handler(request.override(system_message=SystemMessage(content=new_content)))
|
||||
```
|
||||
- **Dynamically switch models**:
|
||||
```python
|
||||
from langchain.chat_models import init_chat_model
|
||||
|
||||
complex_model = init_chat_model("claude-sonnet-4-6")
|
||||
simple_model = init_chat_model("claude-haiku-4-5-20251001")
|
||||
|
||||
@wrap_model_call
|
||||
def dynamic_model(
|
||||
request: ModelRequest,
|
||||
handler: Callable[[ModelRequest], ModelResponse],
|
||||
) -> ModelResponse:
|
||||
model = complex_model if len(request.messages) > 10 else simple_model
|
||||
return handler(request.override(model=model))
|
||||
```
|
||||
- **Dynamically filter tools**:
|
||||
```python
|
||||
@wrap_model_call
|
||||
def filter_tools(
|
||||
request: ModelRequest,
|
||||
handler: Callable[[ModelRequest], ModelResponse],
|
||||
) -> ModelResponse:
|
||||
relevant = [t for t in request.tools if t.name in ["search", "calculator"]]
|
||||
return handler(request.override(tools=relevant))
|
||||
```
|
||||
- **Lessons learned**: `request.override()` is the most important API in wrap_model_call — it can modify `system_message`, `model`, `tools`, `messages`. Use `ExtendedModelResponse + Command` when you need to modify state; use `request.override()` when you only need to modify request parameters. The two can be combined.
|
||||
@@ -0,0 +1,158 @@
|
||||
# Model Integration Issues
|
||||
|
||||
## Issue 1: What problems arise from using ChatOpenAI directly?
|
||||
|
||||
For OpenAI-compatible models, the simplest approach is to reuse `ChatOpenAI` from `langchain-openai`. Using Qwen as an example:
|
||||
|
||||
```python
|
||||
from langchain_openai import ChatOpenAI
|
||||
|
||||
llm = ChatOpenAI(
|
||||
model="qwen-max",
|
||||
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
|
||||
api_key=os.getenv("DASHSCOPE_API_KEY"),
|
||||
)
|
||||
```
|
||||
|
||||
This is the fastest way to get started, but in production it exposes the following issues:
|
||||
|
||||
- **`reasoning_content` (the thinking process) is silently dropped** — `_create_chat_result` and `_convert_chunk_to_generation_chunk` don't recognize this field, so neither streaming nor non-streaming responses surface the reasoning content.
|
||||
- **The reasoning context is broken across multi-turn conversations** — `reasoning_content` in an AIMessage can't be sent back to the model. `_get_request_payload` doesn't handle `additional_kwargs["reasoning_content"]`, so subsequent turns lose the previous turn's reasoning chain.
|
||||
- **An empty `tools: []` triggers provider errors** — some providers strictly validate empty arrays, but `ChatOpenAI` still sends the field even when no tools are bound.
|
||||
|
||||
So `ChatOpenAI` is suitable only for **quick verification** and **simple conversations without needing to display reasoning**. Once reasoning models or multi-turn reasoning chains are involved, a deeper adapter is needed.
|
||||
|
||||
## Issue 2: How to robustly integrate OpenAI-compatible reasoning models
|
||||
|
||||
To address the issues above, two approaches are recommended. Ask the user which they prefer before proceeding:
|
||||
|
||||
<!-- query
|
||||
type: choice
|
||||
question: "Which approach do you prefer for integrating reasoning models?"
|
||||
options:
|
||||
- label: "Code generation"
|
||||
description: "Generate a custom integration class in your repo via the CN Model Integration Guide — full control, easy to customize"
|
||||
- label: "Third-party library"
|
||||
description: "Install langchain-dev-utils and use its built-in adapters — zero hand-written code"
|
||||
default: 1
|
||||
-->
|
||||
|
||||
- **Approach 1 — Code generation via guide**: use the [CN Model Integration Guide](cn-models/README.md) to generate an integration class into the project repo. Best for teams that want full control and easy customization.
|
||||
- **Approach 2 — Third-party library**: install `langchain-dev-utils` and use its built-in adapters. Best for teams that prefer zero hand-written adapter code.
|
||||
|
||||
### Approach 1: Use the CN Model Integration Guide to generate integration classes
|
||||
|
||||
Follow the [CN Model Integration Guide](cn-models/README.md) to generate (via AI Coding) an integration class that subclasses `BaseChatOpenAI` and fixes all critical methods in one shot:
|
||||
|
||||
```python
|
||||
# Using Qwen as an example, the generated class fixes all the issues above:
|
||||
from models.qwen import ChatQwen
|
||||
|
||||
llm = ChatQwen(model="qwen-max")
|
||||
# reasoning_content automatically preserved, empty tools automatically removed, JSONDecodeError friendly hints
|
||||
```
|
||||
|
||||
The generated integration class covers:
|
||||
|
||||
- `_get_request_payload`: reasoning_content round-tripping + removal of empty tools
|
||||
- `_create_chat_result`: extracting reasoning_content from non-streaming responses
|
||||
- `_convert_chunk_to_generation_chunk`: extracting reasoning_content from streaming deltas
|
||||
- `_stream` / `_astream` / `_generate` / `_agenerate`: unified JSONDecodeError handling
|
||||
|
||||
**DeepSeek special path**: if `langchain-deepseek` is installed, the skill generates a subclass of the official class, only adding reasoning_content round-tripping while reusing the official implementation. Providers like Qwen with no official integration class inherit `BaseChatOpenAI` directly.
|
||||
|
||||
Suitable for: teams that want the adapter logic to live in their code repo, easy to read and customize.
|
||||
|
||||
### Approach 2: Use the third-party community library langchain-dev-utils
|
||||
|
||||
`langchain-dev-utils` is a third-party `langchain` ecosystem toolkit with built-in deep adapters for OpenAI-compatible models, removing the cost of hand-writing subclasses.
|
||||
|
||||
First install the standard version:
|
||||
|
||||
```bash
|
||||
pip install -U langchain-dev-utils[standard]
|
||||
```
|
||||
|
||||
#### 2.1 Dynamically generate integration classes via a factory function
|
||||
|
||||
Use the `create_openai_compatible_model` factory function to dynamically generate a chat model class at runtime:
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.chat_models.adapters import create_openai_compatible_model
|
||||
|
||||
ChatQwen = create_openai_compatible_model(
|
||||
model_provider="qwen",
|
||||
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
|
||||
chat_model_cls_name="ChatQwen",
|
||||
compatibility_options={
|
||||
"supported_tool_choice": ["auto","none", "specific"],
|
||||
"supported_response_format": ["json_schema"], # When enabled, with_structured_output defaults to json_schema
|
||||
},
|
||||
)
|
||||
|
||||
model = ChatQwen(model="qwen3-max", reasoning_keep_policy="current")
|
||||
```
|
||||
|
||||
Environment variables follow the `${PROVIDER_NAME}_API_BASE` / `${PROVIDER_NAME}_API_KEY` naming convention; omitting `base_url` reads them automatically.
|
||||
|
||||
The function is built on `BaseChatOpenAI`, with the main enhancements being:
|
||||
|
||||
- **Reasoning field extraction and round-tripping**: automatically parses `reasoning_content` / `reasoning`, with the `reasoning_keep_policy` (`never` / `current` / `all`) controlling how reasoning is retained in historical messages — fits Interleaved Thinking and Preserved Thinking.
|
||||
- **tool_choice differential adaptation**: use `supported_tool_choice` to declare which strategies the provider supports; unsupported values are filtered out instead of being forwarded and triggering errors.
|
||||
- **Dynamic structured-output selection**: based on `supported_response_format`, automatically pick the best strategy between `json_schema` and `function_calling`; declaring `json_schema` automatically sets `model.profile.structured_output` to `True`, integrated with `create_agent`.
|
||||
- **video content_block support**: fills in the video-type multimodal capability missing from `ChatOpenAI`.
|
||||
- **Model profiles**: pass a `profile` at creation or instantiation, so higher-level components like `create_agent` can sense model capabilities.
|
||||
|
||||
> **Note**: under the hood it uses pydantic `create_model`, which has dynamic-creation overhead, and the profiles dict is global. Create integration classes once at project startup to avoid repeated runtime regeneration.
|
||||
|
||||
#### 2.2 The registration-based style aligned with `init_chat_model`
|
||||
|
||||
The factory function in 2.1 requires the business side to explicitly hold a concrete class like `ChatQwen`. If you prefer LangChain's native `init_chat_model("provider:model")` "model by string" initialization style, `langchain-dev-utils` provides an equivalent experience — just use `register_model_provider` to register an OpenAI-compatible model under a unified entry point and set `chat_model` to `"openai-compatible"`. Internally it calls the `create_openai_compatible_model` above to build the integration class, and the business side no longer needs to reference the model class directly.
|
||||
|
||||
**Method 1: Pass arguments explicitly**
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.chat_models import register_model_provider
|
||||
|
||||
register_model_provider(
|
||||
provider_name="qwen",
|
||||
chat_model="openai-compatible",
|
||||
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
|
||||
)
|
||||
```
|
||||
|
||||
**Method 2: Via environment variables (recommended for config management)**
|
||||
|
||||
```bash
|
||||
export QWEN_API_BASE=https://dashscope.aliyuncs.com/compatible-mode/v1
|
||||
export QWEN_API_KEY=sk-xxx
|
||||
```
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.chat_models import register_model_provider
|
||||
|
||||
register_model_provider(
|
||||
provider_name="qwen",
|
||||
chat_model="openai-compatible",
|
||||
# Auto-reads QWEN_API_BASE / QWEN_API_KEY
|
||||
)
|
||||
```
|
||||
|
||||
Parameters like `base_url`, `compatibility_options`, and `model_profiles` from `create_openai_compatible_model` are also passed through — usage is identical to calling the factory function directly:
|
||||
|
||||
```python
|
||||
from langchain_dev_utils.chat_models import register_model_provider
|
||||
|
||||
register_model_provider(
|
||||
provider_name="qwen",
|
||||
chat_model="openai-compatible",
|
||||
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
|
||||
compatibility_options={
|
||||
"supported_tool_choice": ["auto", "none","specific"],
|
||||
"supported_response_format": ["json_schema"],
|
||||
},
|
||||
model_profiles=model_profiles,
|
||||
)
|
||||
```
|
||||
|
||||
After registration, business code can initialize models through the unified entry point `load_chat_model("qwen:qwen3-max")` — this is the usage pattern aligned with `init_chat_model()`. The call sites no longer couple to specific class names, `base_url`, or `api_key`.
|
||||
@@ -0,0 +1,317 @@
|
||||
# Multi-Agent Orchestration Issues
|
||||
|
||||
## Issue 1: Choosing between subagents and handoffs
|
||||
|
||||
- **Symptom**: When building multi-agent systems, you see both the "subagent as tool" pattern and the `Command(goto=..., graph=Command.PARENT)` "handoff" pattern in the docs and don't know which to choose. Picking wrong leads to: the main agent can't get the sub-agent's output, or the sub-agent can't talk to the user directly.
|
||||
- **Cause**: The two patterns solve completely different orchestration needs:
|
||||
- **Subagents (recommended default)**: a central supervisor treats sub-agents as **tool calls** — the main agent decides when to call, what query to send, and how to use the return value. Sub-agents **don't talk to the user directly**; each call starts from a clean context, and the result returns to the main agent.
|
||||
- **Handoffs**: a tool updates a state variable (e.g. `current_step` / `active_agent`), based on which the system switches "the currently active agent / config". The agent switched-to **directly takes over the conversation with the user**, with state persisting across turns.
|
||||
- **Solution**: Use this decision table:
|
||||
|
||||
| Need | Pick |
|
||||
| ------------------------------------------------------------- | --------------------- |
|
||||
| Main agent needs the sub-agent's result to decide next step | Subagents (sync call) |
|
||||
| Sub-agent needs to run a task in a clean context, avoid polluting the main conversation | Subagents (context isolation) |
|
||||
| Multiple domains (calendar / email / CRM…) need centralized routing | Subagents |
|
||||
| Customer service flow: first collect warranty id, then refund — must unlock in order | Handoffs |
|
||||
| Different stages need different system prompts / tool sets, and need to interact with user directly | Handoffs |
|
||||
| sales ↔ support transfer between each other, each agent talks to the user | Handoffs |
|
||||
|
||||
- **Lessons learned**: Default to subagents — its semantics are simplest (just a tool call), and it has the fewest failure modes. **Only when an agent needs to converse with the user across multiple turns directly** (instead of returning results to the upper layer) should you use handoffs. The two can be mixed: the supervisor uses subagents to manage multiple sub-agents, and one sub-agent internally uses handoffs for multi-stage flow.
|
||||
|
||||
## Issue 2: tool-per-agent or single dispatch tool in subagent mode
|
||||
|
||||
- **Symptom**: Subagent mode has two ways to expose: "wrap each sub-agent as a separate tool" vs. "write a general `task(agent_name, description)` tool that routes by name to a sub-agent in the registry". You don't know which to choose, or you started with tool-per-agent and later found adding a new agent requires heavy changes to the supervisor.
|
||||
- **Cause**: The two are inverse in "customizability" vs "extensibility":
|
||||
- **Tool per agent**: each sub-agent is wrapped as its own `@tool`, allowing per-sub-agent control over input/output/state passing. The cost: every new agent requires modifying the supervisor's `tools=[...]`.
|
||||
- **Single dispatch tool**: only one `task(agent_name, description)` tool; sub-agents are looked up in a registry dict; adding agents only modifies the registry, not the supervisor. The cost: all sub-agents share the same "query passed as user message, last message as return value" convention, no per-agent customization.
|
||||
- **Solution**:
|
||||
- **Tool per agent** (few agents, each needs separate context engineering):
|
||||
```python
|
||||
from langchain.tools import tool
|
||||
from langchain.agents import create_agent
|
||||
|
||||
research_subagent = create_agent(model="...", tools=[...])
|
||||
|
||||
@tool("research", description="Research a topic and return findings")
|
||||
def call_research(query: str) -> str:
|
||||
result = research_subagent.invoke({"messages": [{"role": "user", "content": query}]})
|
||||
return result["messages"][-1].content
|
||||
|
||||
supervisor = create_agent(model="...", tools=[call_research])
|
||||
```
|
||||
|
||||
- **Single dispatch tool** (many agents, multi-team, prefer convention over configuration):
|
||||
```python
|
||||
from langchain.tools import tool
|
||||
from langchain.agents import create_agent
|
||||
|
||||
SUBAGENTS = {
|
||||
"research": create_agent(model="gpt-5.4", prompt="You are a research specialist..."),
|
||||
"writer": create_agent(model="gpt-5.4", prompt="You are a writing specialist..."),
|
||||
}
|
||||
|
||||
@tool
|
||||
def task(agent_name: str, description: str) -> str:
|
||||
"""Launch an ephemeral subagent for a task.
|
||||
|
||||
Available agents:
|
||||
- research: Research and fact-finding
|
||||
- writer: Content creation and editing
|
||||
"""
|
||||
agent = SUBAGENTS[agent_name]
|
||||
result = agent.invoke({"messages": [{"role": "user", "content": description}]})
|
||||
return result["messages"][-1].content
|
||||
|
||||
supervisor = create_agent(
|
||||
model="gpt-5.4",
|
||||
tools=[task],
|
||||
system_prompt="You coordinate specialized sub-agents. Use the task tool to delegate work.",
|
||||
)
|
||||
```
|
||||
|
||||
- **Tell the main agent which sub-agents exist under single dispatch**, pick by scale:
|
||||
|
||||
| Registry size / change rate | Recommended approach |
|
||||
| ------------------------------------- | --------------------------------------------------------------------- |
|
||||
| <10, mostly static | List agent names + descriptions in the supervisor's system_prompt |
|
||||
| <10, want type safety | Constrain `agent_name: AgentName` with an `Enum` as the tool param |
|
||||
| >10, dynamically registered or maintained by multiple teams | Provide a separate `list_agents(query)` tool so the main agent looks them up on demand |
|
||||
|
||||
- **Lessons learned**: When unsure, start with tool-per-agent; switch to single dispatch when the agent count exceeds 5 or you clearly need multi-team independent delivery. The "cheap" of single dispatch shows in "no supervisor code changes when adding agents", but you pay with "all agents must share the same behavior contract" — switching too early forces customization to take detours.
|
||||
|
||||
## Issue 3: Can't get the sub-agent's internal state, no way to review at interrupt time
|
||||
|
||||
- **Symptom**: You want to use `get_state(subgraphs=True)` at the supervisor level to see where the subagent is in its run and what its current state is — but the subagent state is never returned. Or you want the subagent to preserve conversation history across multiple calls (e.g. a long-memory research assistant), but every invoke starts with empty state.
|
||||
- **Cause**: The subagent is invoked **inside a tool function**, and LangGraph can't statically discover this nested graph at compile time — `get_state(subgraphs=True)` can only find "subgraphs added with `add_node`" or "subgraphs invoked in a node function", but not subagents called inside tools. Separately, the subagent's `checkpointer` parameter controls three persistence modes; without explicit configuration it uses **per-invocation** (`checkpointer=None`, default), which doesn't preserve state across calls — this is what most subagents want, but it can mislead you into thinking "subagents can't use interrupt / can't see state at all".
|
||||
- **Solution**: First understand the three `checkpointer` modes, then pick by need:
|
||||
|
||||
| Mode | `checkpointer=` | Cross-call memory | Interrupt within one call | State inspection | Parallel calls to the same subagent |
|
||||
| --------------- | --------------- | ----------------- | ------------------------- | -------------------------------------------- | ---------------------- |
|
||||
| per-invocation | `None` (default) | ❌ | ✅ | ⚠️ Only "during the current call/at interrupt" | ✅ |
|
||||
| per-thread | `True` | ✅ | ✅ | ✅ | ❌ (namespace conflict) |
|
||||
| stateless | `False` | ❌ | ❌ | ❌ | ✅ |
|
||||
|
||||
In all modes, **the parent graph must be compiled with a checkpointer**, otherwise interrupt / state inspection / per-thread memory all fail to work.
|
||||
|
||||
- **Inspect nested state mid-subagent-run** (works in per-invocation too): the subagent must be "in the middle of a single call" (typical case: triggered `interrupt()` and waiting for resume). Then `graph.get_state(config, subgraphs=True).tasks[0].state` returns the nested state. Once that call ends, in per-invocation mode the state doesn't accumulate, and the next call is fresh.
|
||||
- **View the subagent's accumulated full state**: subagent must use `checkpointer=True`, and the parent graph must also have a checkpointer. Then `get_state(subgraphs=True)` returns the subagent state accumulated on that thread.
|
||||
- **Let the subagent preserve conversation history across calls** (continuations mode):
|
||||
```python
|
||||
from langchain.agents import create_agent
|
||||
from langchain.agents.middleware import ToolCallLimitMiddleware
|
||||
from langgraph.checkpoint.memory import MemorySaver
|
||||
|
||||
fruit_agent = create_agent(
|
||||
model="gpt-5.4-mini",
|
||||
tools=[fruit_info],
|
||||
prompt="You are a fruit expert. Respond in one sentence.",
|
||||
checkpointer=True, # per-thread persistence
|
||||
)
|
||||
|
||||
@tool
|
||||
def ask_fruit_expert(question: str) -> str:
|
||||
"""Ask the fruit expert. Use for ALL fruit questions."""
|
||||
resp = fruit_agent.invoke({"messages": [{"role": "user", "content": question}]})
|
||||
return resp["messages"][-1].content
|
||||
|
||||
agent = create_agent(
|
||||
model="gpt-5.4-mini",
|
||||
tools=[ask_fruit_expert],
|
||||
prompt="ALWAYS delegate fruit questions to ask_fruit_expert.",
|
||||
middleware=[
|
||||
# Must forbid parallel calls, otherwise two calls write to the same namespace → checkpoint conflict
|
||||
ToolCallLimitMiddleware(tool_name="ask_fruit_expert", run_limit=1),
|
||||
],
|
||||
checkpointer=MemorySaver(), # Parent graph checkpointer is a hard requirement
|
||||
)
|
||||
```
|
||||
Pitfall: per-thread subagents **don't support parallel LLM calls to the same tool** — e.g. "ask about apple and banana at the same time" causes the model to concurrently call `ask_fruit_expert` twice, both writing to the same namespace and conflicting. Use `ToolCallLimitMiddleware` to rate-limit, or disable parallel tool calls at the model layer.
|
||||
|
||||
- **Multiple different per-thread subagents coexisting** (both fruit and veggie need memory): each subagent needs to be wrapped in a `StateGraph` with a unique node name, otherwise LangGraph assigns namespaces by "call order", and reordering calls scrambles state:
|
||||
```python
|
||||
from langgraph.graph import MessagesState, StateGraph
|
||||
|
||||
def create_sub_agent(model, *, name, **kwargs):
|
||||
"""Wrap with a unique node name to get a stable namespace."""
|
||||
agent = create_agent(model=model, name=name, **kwargs)
|
||||
return (
|
||||
StateGraph(MessagesState)
|
||||
.add_node(name, agent)
|
||||
.add_edge("__start__", name)
|
||||
.compile()
|
||||
)
|
||||
|
||||
fruit_agent = create_sub_agent("gpt-5.4-mini", name="fruit_agent", tools=[fruit_info], prompt="...", checkpointer=True)
|
||||
veggie_agent = create_sub_agent("gpt-5.4-mini", name="veggie_agent", tools=[veggie_info], prompt="...", checkpointer=True)
|
||||
```
|
||||
|
||||
- **No checkpoint overhead needed for the subagent** (short tool-style calls, clearly no interrupt needed): use `checkpointer=False` to enter stateless mode. The subagent runs as an ordinary function with no durable execution — if it crashes, it runs again from scratch.
|
||||
|
||||
- **Need to access nested state at the main graph layer for debugging** (not just at interrupt time): change the subagent from "invoked inside a tool" to "called inside a graph". Two options:
|
||||
- **Call subgraph inside a node**: use when parent/child schemas differ; write a wrapper in the node function to convert state;
|
||||
- **Add subgraph as a node**: use when parent/child share state keys (typical: both use `MessagesState`). Directly `add_node` the compiled subagent — no wrapper needed.
|
||||
|
||||
Both forms are statically recognizable by LangGraph, and `get_state(subgraphs=True)` returns nested state. The cost is giving up the "subagent as tool" natural routing ability — you have to design the graph yourself.
|
||||
|
||||
- **Just want to see what went wrong with the subagent**: enable LangSmith tracing. The subagent's run appears as a nested trace under the main agent's trace, far more intuitive than `get_state`.
|
||||
|
||||
- **Lessons learned**:
|
||||
- The default per-invocation mode for subagents is the right choice for most scenarios — it supports interrupt, supports state inspection within a single call, and supports parallel calls; it just doesn't have cross-call memory. Treat it as a "side-effect-free pure function".
|
||||
- Before upgrading to `checkpointer=True` (per-thread), confirm two things: ① the parent graph has a checkpointer; ② the main agent won't concurrently call the same per-thread subagent (middleware or model config handles this).
|
||||
- "Main graph needs nested state" and "subagent as tool" are mutually exclusive — for nested visibility, accept "writing a custom graph"; don't try to invoke in a tool and then expect the supervisor to `get_state(subgraphs=True)`.
|
||||
|
||||
## Issue 4: Too much subagent wrapping boilerplate / need per-agent input-output customization
|
||||
|
||||
- **Symptom**: Following Issue 2's approach to wrap subagents as tools, every agent has to repeatedly write "`@tool` → `agent.invoke({"messages": [...]})` → `result["messages"][-1].content`". When you need to add input preprocessing to a subagent (passing the main agent context along) or output post-processing (returning structured results back to main agent state), you either write nested lambdas or jam in `Command(update=...)` — repetitive and error-prone.
|
||||
- **Cause**: LangChain natively only exposes low-level mechanisms like `Command` / `ToolRuntime`, without abstracting "wrap agent as tool" and "intercept input/output" into a separate interface. Every new subagent requires writing the boilerplate from scratch.
|
||||
- **Solution**: Use `wrap_agent_as_tool` / `wrap_all_agents_as_tool` from the third-party `langchain-dev-utils`, combined with `pre_input_hooks` / `post_output_hooks` for context engineering.
|
||||
|
||||
- **Wrap a single subagent** (replacing the hand-written tool-per-agent boilerplate):
|
||||
```python
|
||||
from langchain_dev_utils.agents import wrap_agent_as_tool
|
||||
|
||||
schedule_event = wrap_agent_as_tool(
|
||||
calendar_agent,
|
||||
tool_name="schedule_event",
|
||||
tool_description=(
|
||||
"Schedule a calendar event using natural language."
|
||||
"Input: natural language calendar scheduling request (e.g. 'meeting with design team next Tuesday 2pm')"
|
||||
),
|
||||
)
|
||||
manage_email = wrap_agent_as_tool(email_agent, tool_name="manage_email", tool_description="...")
|
||||
|
||||
supervisor = create_agent(model="...", tools=[schedule_event, manage_email])
|
||||
```
|
||||
Both `tool_name` and `tool_description` are optional, but **strongly recommended to set explicitly** — the default name is `transfer_to_{agent_name}` and the default description is `This tool transforms input to {agent_name}`. The main agent has almost no useful information to base tool selection on.
|
||||
|
||||
- **Wrap multiple subagents as a single dispatch tool** (replacing the hand-written single dispatch registry):
|
||||
```python
|
||||
from langchain_dev_utils.agents import wrap_all_agents_as_tool
|
||||
|
||||
call_subagent = wrap_all_agents_as_tool(
|
||||
[calendar_agent, email_agent],
|
||||
tool_name="call_subagent",
|
||||
tool_description=(
|
||||
"Call a sub-agent to execute a task. Available agents: "
|
||||
"- calendar_agent: for scheduling calendar events\n"
|
||||
"- email_agent: for sending emails"
|
||||
),
|
||||
)
|
||||
|
||||
main_agent = create_agent(model="...", tools=[call_subagent])
|
||||
```
|
||||
|
||||
- **Use `pre_input_hooks` to inject context into the subagent** (pass the main agent state / original user message through to the sub-agent):
|
||||
```python
|
||||
from langchain.tools import ToolRuntime
|
||||
|
||||
def process_input(request: str, runtime: ToolRuntime) -> str:
|
||||
original = next(m for m in runtime.state["messages"] if m.type == "human")
|
||||
return (
|
||||
"You are helping handle the following user query:\n\n"
|
||||
f"{original.text}\n\n"
|
||||
"You've been assigned the following sub-task:\n\n"
|
||||
f"{request}"
|
||||
)
|
||||
|
||||
call_agent = wrap_agent_as_tool(agent, pre_input_hooks=process_input)
|
||||
```
|
||||
When the hook returns `str`, it's auto-wrapped as a `HumanMessage` as subagent input; when it returns `dict`, it's used directly as input (for scenarios needing extra state fields). Pass a `(sync_fn, async_fn)` tuple to handle sync/async paths separately.
|
||||
|
||||
- **Use `post_output_hooks` to feed structured results back to the main agent** (replacing hand-written `Command(update=...)`):
|
||||
```python
|
||||
import json
|
||||
|
||||
def process_output(request, response, runtime):
|
||||
return json.dumps({
|
||||
"status": "success",
|
||||
"event_id": "evt_123",
|
||||
"summary": response["messages"][-1].text,
|
||||
})
|
||||
|
||||
call_agent = wrap_agent_as_tool(agent, post_output_hooks=process_output)
|
||||
```
|
||||
The hook return value can be a string (used directly as tool result) or a `Command` object (also updates main agent state).
|
||||
|
||||
- **Handle hooks per-subagent-name under `wrap_all_agents_as_tool`**:
|
||||
```python
|
||||
from langchain_dev_utils.agents.wrap import get_subagent_name
|
||||
|
||||
def process_input(request: str, runtime: ToolRuntime):
|
||||
if get_subagent_name(runtime) == "weather_agent":
|
||||
city = runtime.state.get("city", "Unknown city")
|
||||
return f"Current city is: {city}. Please complete the task based on the above. " + request
|
||||
return request
|
||||
```
|
||||
|
||||
- **Lessons learned**: Native `Command + ToolRuntime` are more flexible, but most subagent wrapping just needs "change the name / inject context / wrap return value" — those three things. Use `wrap_agent_as_tool` directly to skip the boilerplate. To preserve single dispatch tool's extensibility (Issue 2), use `wrap_all_agents_as_tool` + `get_subagent_name` for per-name hook handling — much cleaner than maintaining a registry dict with if/elif.
|
||||
|
||||
## Issue 5: How to quickly build a handoff-capable multi-Agent system
|
||||
|
||||
- **Symptom**: Building multi-agent handoff with 4 agents transferring between each other requires writing 12 `transfer_to_xxx` tools — each repeating "get last_ai_message from state + construct a paired ToolMessage + wrap with `Command(goto=..., graph=Command.PARENT)`". Missing the ToolMessage pairing causes the receiving agent to see an illegal history of "tool_call without tool_response", directly reporting invalid message sequence. If you switch to "single agent + middleware" to dodge message pairing, you have to write `wrap_model_call` to swap prompts and tools based on `active_agent` / `current_step` — still not lightweight.
|
||||
- **Cause**: Handoffs are essentially a state machine — "switch the currently available prompt / tools based on active_agent". LangChain only provides low-level components like `Command` / `ToolRuntime`; it doesn't abstract "declare which agents exist and who can transfer to whom" into a high-level interface.
|
||||
- **Solution**: Use `HandoffAgentMiddleware` from the third-party `langchain-dev-utils` for declarative configuration — write each agent's prompt / tools / transfer targets as a dict, and the middleware auto-generates corresponding transfer tools. Message pairing, Command construction, and dynamic prompt/tool switching are all built in.
|
||||
|
||||
```python
|
||||
from langchain.agents import create_agent
|
||||
from langchain_dev_utils.agents.middleware import HandoffAgentMiddleware
|
||||
from langchain_dev_utils.agents.middleware.handoffs import AgentConfig
|
||||
|
||||
agent_config: dict[str, AgentConfig] = {
|
||||
"time_agent": {
|
||||
"prompt": "You are a time assistant",
|
||||
"tools": [get_current_time],
|
||||
"handoffs": ["default_agent"], # can only hand off back to default
|
||||
},
|
||||
"weather_agent": {
|
||||
"prompt": "You are a weather assistant",
|
||||
"tools": [get_current_weather, get_current_city],
|
||||
"handoffs": ["default_agent"],
|
||||
},
|
||||
"code_agent": {
|
||||
"model": "openai:gpt-5.4", # can specify model individually, overriding the fallback model
|
||||
"prompt": "You are a code assistant",
|
||||
"tools": [run_code],
|
||||
"handoffs": ["default_agent"],
|
||||
},
|
||||
"default_agent": {
|
||||
"prompt": "You are an assistant",
|
||||
"default": True, # globally there must be exactly one default
|
||||
"handoffs": "all", # can hand off to any agent
|
||||
},
|
||||
}
|
||||
|
||||
agent = create_agent(
|
||||
model="openai:gpt-5.4", # fallback model (reused when agents_config doesn't declare model)
|
||||
middleware=[HandoffAgentMiddleware(agents_config=agent_config)],
|
||||
)
|
||||
```
|
||||
When using this middleware, `create_agent`'s own `tools` and `system_prompt` are ignored — all prompts/tools come from `agents_config`.
|
||||
|
||||
- To customize the **description** of the transfer tool (without changing the implementation), pass `custom_handoffs_tool_descriptions`:
|
||||
```python
|
||||
HandoffAgentMiddleware(
|
||||
agents_config=agent_config,
|
||||
custom_handoffs_tool_descriptions={
|
||||
"code_agent": "This tool is for handing off to the code assistant for code questions",
|
||||
...
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
- To **fully customize the transfer tool implementation** (e.g. log / audit during handoff), pass `handoffs_tool_overrides`. Custom tools must return `Command`, with `update.messages` containing the tool response and `update.active_agent` pointing to the target agent name:
|
||||
```python
|
||||
@tool
|
||||
def transfer_to_code_agent(runtime: ToolRuntime) -> Command:
|
||||
"""This tool helps you hand off to the code assistant"""
|
||||
# ... your custom logic (logging, auditing, etc.) ...
|
||||
return Command(update={
|
||||
"messages": [ToolMessage(content="transfer to code agent", tool_call_id=runtime.tool_call_id)],
|
||||
"active_agent": "code_agent",
|
||||
})
|
||||
|
||||
HandoffAgentMiddleware(agents_config=agent_config, handoffs_tool_overrides={"code_agent": transfer_to_code_agent})
|
||||
```
|
||||
|
||||
- **Lessons learned**: The core mental model for handoffs is "switch prompt + tools by active_agent" — a declarative configuration. Use `HandoffAgentMiddleware` to drop the mental load from "write N transfer tools + handle message pairing" to "fill in a dict". Also, **checkpointer is a hard requirement** — `active_agent` is cross-turn state; without a checkpointer the next conversation restarts from the default agent.
|
||||
@@ -0,0 +1,223 @@
|
||||
# Streaming Output Issues
|
||||
|
||||
## Issue 1: Choosing between streaming APIs (v3 vs stream_mode)
|
||||
|
||||
- **Symptom**: You see both `agent.stream(stream_mode=...)` and `agent.stream_events(version="v3")` in the docs and don't know which one to use.
|
||||
- **Cause**: v1.3 introduced **Event Streaming** (`stream_events(version="v3")`), which is completely different from classic `stream(stream_mode=...)` in programming model:
|
||||
- Classic API: returns a generator that yields all events in time order, mixed together. The business side dispatches with if/elif based on `mode` / `type`.
|
||||
- Event Streaming: returns a **`Stream` object** that organizes events from the same run into multiple **typed projections** — each projection is an independently consumable iterator. The business side "reads whichever attribute holds the data it wants".
|
||||
|
||||
Both share the same underlying LangGraph protocol, but the top-level API shape is completely different — mixing doc examples leads to mismatched output structures.
|
||||
|
||||
- **Solution**:
|
||||
- **Always use Event Streaming for new projects**. Minimum viable form:
|
||||
```python
|
||||
from langchain.agents import create_agent
|
||||
|
||||
def get_weather(city: str) -> str:
|
||||
"""Get weather for a city."""
|
||||
return f"It's always sunny in {city}!"
|
||||
|
||||
agent = create_agent(model="gpt-5-nano", tools=[get_weather])
|
||||
|
||||
stream = agent.stream_events(
|
||||
{"messages": [{"role": "user", "content": "What is the weather in SF?"}]},
|
||||
version="v3",
|
||||
)
|
||||
|
||||
for message in stream.messages:
|
||||
for delta in message.text:
|
||||
print(delta, end="", flush=True)
|
||||
|
||||
final_state = stream.output
|
||||
```
|
||||
|
||||
Key intuition: `stream` is not a generator, it's an **object**. It holds all events from the run and splits them into independent projections by dimension:
|
||||
|
||||
| Projection | Use |
|
||||
| --------------------- | ------------------------------------------------------------------ |
|
||||
| `stream.messages` | One `ChatModelStream` per LLM call (most commonly used) |
|
||||
| `stream.tool_calls` | Tool **execution-phase** lifecycle (input, output_deltas, output, error) |
|
||||
| `stream.subgraphs` | Nested sub-agent / subgraph runs (see Issue 2) |
|
||||
| `stream.values` | Agent state snapshots |
|
||||
| `stream.output` | Final Agent state (equivalent to `invoke`'s return value) |
|
||||
| `stream.extensions` | Custom transformer projections |
|
||||
| `for event in stream` | Raw protocol events with full envelope (fallback / debug) |
|
||||
|
||||
- **`stream.messages` — the most commonly used projection**, one `ChatModelStream` per LLM call:
|
||||
```python
|
||||
for message in stream.messages:
|
||||
print(f"[{message.node}] ", end="") # the node name where this LLM call sits
|
||||
for delta in message.text: # text delta
|
||||
print(delta, end="", flush=True)
|
||||
|
||||
for delta in message.reasoning: # reasoning delta (only if model has reasoning enabled)
|
||||
print(f"[thinking] {delta}", end="", flush=True)
|
||||
|
||||
full_msg = message.output # the complete AIMessage after the call ends
|
||||
if full_msg.usage_metadata:
|
||||
print(full_msg.usage_metadata)
|
||||
```
|
||||
|
||||
Key attributes of `ChatModelStream`:
|
||||
|
||||
| Attribute | Description |
|
||||
| --------------------- | --------------------------------------------------------------------------------- |
|
||||
| `message.text` | Text deltas; `str(message.text)` waits until end to get full text |
|
||||
| `message.reasoning` | Reasoning deltas (only populated when model has reasoning on); same shape as `text` |
|
||||
| `message.tool_calls` | Argument fragments while the model **is producing** a tool_call; `.get()` for the final structured result |
|
||||
| `message.output` | The complete `AIMessage` after this call (with `usage_metadata` / `content_blocks`) |
|
||||
| `message.node` | The graph node name where this call sits (use to distinguish sources across multiple LLM calls within one agent) |
|
||||
|
||||
- **`stream.tool_calls` — tool execution lifecycle**. Note the difference from `message.tool_calls`:
|
||||
- `message.tool_calls`: argument fragments while the model is **still saying** "I want to call this tool".
|
||||
- `stream.tool_calls`: after the model is done, the process of the tool **actually being executed**, including input, streaming output, final result, and exceptions.
|
||||
|
||||
```python
|
||||
for call in stream.tool_calls:
|
||||
print(f"{call.tool_name}({call.input})")
|
||||
for delta in call.output_deltas: # real-time output from streaming tools (e.g. retrieval/long operations)
|
||||
print(delta, end="", flush=True)
|
||||
print(call.output, call.error)
|
||||
```
|
||||
|
||||
- **Multi-projection concurrent consumption — Event Streaming's biggest convenience**. Multiple projections on the same stream can be consumed independently and in parallel, without writing select/dispatch yourself:
|
||||
```python
|
||||
# async: use asyncio.gather to consume multiple projections concurrently
|
||||
import asyncio
|
||||
stream = await agent.astream_events(input, version="v3")
|
||||
|
||||
async def consume_messages():
|
||||
async for message in stream.messages:
|
||||
print(await message.text)
|
||||
|
||||
async def consume_tool_calls():
|
||||
async for call in stream.tool_calls:
|
||||
print(call.tool_name, call.input)
|
||||
|
||||
await asyncio.gather(consume_messages(), consume_tool_calls())
|
||||
|
||||
# sync: use stream.interleave to merge multiple projections into a single iterator in arrival order
|
||||
for name, item in stream.interleave("messages", "tool_calls", "values"):
|
||||
if name == "messages":
|
||||
print(item.text)
|
||||
elif name == "tool_calls":
|
||||
print(item.tool_name, item.input)
|
||||
```
|
||||
|
||||
- **Channels not covered by typed projections** (looking at the raw envelope, debugging an unexposed event), iterate the raw events directly:
|
||||
```python
|
||||
for event in stream:
|
||||
print(event["method"], event["params"]["namespace"], event["params"]["data"])
|
||||
```
|
||||
|
||||
- **Legacy projects / strong dependency on Pregel-level capabilities**: keep `stream(stream_mode=...)`. Recommend upgrading to the v2 output format (`langgraph>=1.1`) for unified `StreamPart` dicts:
|
||||
```python
|
||||
for chunk in agent.stream(
|
||||
{"messages": [{"role": "user", "content": "What is the weather in SF?"}]},
|
||||
stream_mode=["updates", "custom"],
|
||||
version="v2", # unified StreamPart dict, no longer (mode, data) tuple
|
||||
):
|
||||
print(chunk["type"]) # "updates" or "custom"
|
||||
print(chunk["data"])
|
||||
```
|
||||
|
||||
- **Lessons learned**: The mental model for Event Streaming is "**subscribe to multi-channel streams by data type**", not "consume one generator over time". Once you grasp this, the rest is just looking up projections: chat UI defaults to `stream.messages` + `message.text`; multi-LLM defaults to `stream.subgraphs` + `message.node` (Issue 2); custom events use `get_stream_writer()` + `stream_mode="custom"` (Issue 4). Only fall back to `stream_mode` for legacy code or to leverage Pregel-level capabilities like `get_stream_writer()`.
|
||||
|
||||
## Issue 2: How to tell apart tokens from different LLMs when one Agent has multiple LLM calls
|
||||
|
||||
- **Symptom**: An Agent may have sub-agents, side LLMs inside middleware (safety review, structured output validation, etc.), or models called from within a tool — beyond the main model. `stream_mode="messages"` pushes all LLM tokens mixed together, leaving no way to tell which model produced a given token.
|
||||
- **Cause**: `stream_mode="messages"` doesn't distinguish sources by default; in Event Streaming nested calls go through nested namespaces, so you either judge by `message.node` / metadata, or consume `stream.subgraphs` as a separate projection.
|
||||
- **Solution**:
|
||||
- **Explicitly name all agents**: `create_agent(..., name="weather_agent")`. The name enters metadata and is also attached to the generated `AIMessage`.
|
||||
- **Event Streaming (recommended)**:
|
||||
- Different node LLM calls within the same Agent: distinguish via `message.node`.
|
||||
- Nested sub-Agent / Subgraph: consume separately via `stream.subgraphs` + `subagent.graph_name`:
|
||||
```python
|
||||
from langchain.chat_models import init_chat_model
|
||||
|
||||
weather_agent = create_agent(
|
||||
model=init_chat_model("openai:gpt-5.4"),
|
||||
tools=[get_weather],
|
||||
name="weather_agent",
|
||||
)
|
||||
|
||||
def call_weather(query: str) -> str:
|
||||
"""Query the weather agent."""
|
||||
result = weather_agent.invoke({"messages": [{"role": "user", "content": query}]})
|
||||
return result["messages"][-1].text
|
||||
|
||||
supervisor = create_agent(
|
||||
model=init_chat_model("openai:gpt-5.4"),
|
||||
tools=[call_weather],
|
||||
name="supervisor",
|
||||
)
|
||||
|
||||
stream = supervisor.stream_events(
|
||||
{"messages": [{"role": "user", "content": "What's the weather in Boston?"}]},
|
||||
version="v3",
|
||||
)
|
||||
|
||||
for subagent in stream.subgraphs:
|
||||
if subagent.graph_name != "weather_agent":
|
||||
continue
|
||||
for message in subagent.messages:
|
||||
for token in message.text:
|
||||
print(token, end="", flush=True)
|
||||
```
|
||||
- **Classic mode**: you must pass `subgraphs=True`, and distinguish sources by `metadata["lc_agent_name"]` or `metadata["langgraph_node"]`:
|
||||
```python
|
||||
for chunk in agent.stream(input, stream_mode=["messages", "updates"], subgraphs=True, version="v2"):
|
||||
if chunk["type"] == "messages":
|
||||
token, metadata = chunk["data"]
|
||||
agent_name = metadata.get("lc_agent_name")
|
||||
node_name = metadata.get("langgraph_node")
|
||||
# Use agent_name / node_name to distinguish sources
|
||||
```
|
||||
- **Lessons learned**: As long as there might be more than one LLM call in an Agent (even if not multi-agent, just a middleware that calls a model), think through how to distinguish token sources up front; always pass `name=` to agents explicitly, otherwise metadata has basically nothing usable.
|
||||
|
||||
## Issue 3: How to turn off streaming for certain models
|
||||
|
||||
- **Symptom**: In some cases you don't need streaming from certain models, but once `stream_mode` is set to "messages", all of them stream by default.
|
||||
- **Cause**: Whether to stream depends on the model instance configuration; the Agent layer won't turn it off for you.
|
||||
- **Solution**:
|
||||
- OpenAI and similar support the `streaming` field:
|
||||
```python
|
||||
from langchain_openai import ChatOpenAI
|
||||
|
||||
model = ChatOpenAI(model="gpt-5.4", streaming=False)
|
||||
```
|
||||
- For models that don't support `streaming`, use the generic `disable_streaming=True` from the base class.
|
||||
- **Lessons learned**: Streaming is a **model-instance-level** switch, not an Agent-level one — different LLM calls within the same Agent can mix streaming/non-streaming. Default intermediate-step models (safety review, structured output validation, sub-Agents) to non-streaming and let only the final user-facing model stream. This significantly reduces event noise and avoids leaking internal pipelines to the client.
|
||||
|
||||
|
||||
|
||||
## Issue 4: Custom events / in-tool progress reports not arriving
|
||||
|
||||
- **Symptom**: You want to push custom info like "download progress", "retrieval hit count", or "domain events" from tools or middleware, but the consumer side never receives them. Or after the code change, invoking the tool standalone errors out.
|
||||
- **Cause**: The built-in `stream.messages` / `stream.tool_calls` only cover model tokens and the tool lifecycle — **they don't carry out custom data actively written from inside nodes/tools**. To let external consumers see this data, you must explicitly go through the "custom events channel" — i.e. `get_stream_writer()` + `stream_mode="custom"`. Another common side effect: once you call `get_stream_writer()` inside a tool, **that tool can only run in a LangGraph execution context**, and a standalone `invoke` outside the Agent will raise `RuntimeError`.
|
||||
- **Solution**:
|
||||
- **Get the writer inside a node/tool and push custom events**:
|
||||
```python
|
||||
from langgraph.config import get_stream_writer
|
||||
|
||||
def get_weather(city: str) -> str:
|
||||
writer = get_stream_writer()
|
||||
writer(f"Looking up data for city: {city}")
|
||||
writer(f"Acquired data for city: {city}")
|
||||
return f"It's always sunny in {city}!"
|
||||
```
|
||||
- **Consumer side subscribes with `stream_mode="custom"`** (also recommend upgrading to v2 output format for unified `StreamPart` dicts):
|
||||
```python
|
||||
for chunk in agent.stream(
|
||||
{"messages": [{"role": "user", "content": "What is the weather in SF?"}]},
|
||||
stream_mode=["messages", "custom"],
|
||||
version="v2",
|
||||
):
|
||||
if chunk["type"] == "custom":
|
||||
print(chunk["data"])
|
||||
```
|
||||
- writer accepts any serializable object, not just strings — pass a dict directly to push structured data (e.g. `writer({"event": "progress", "pct": 30})`); consumer parses it from `chunk["data"]` per business convention.
|
||||
- **Lessons learned**:
|
||||
- `stream.messages` / `stream.tool_calls` don't solve "business events" — when you need to actively push custom data, go back to `get_stream_writer()` + `stream_mode="custom"`.
|
||||
- Calling `get_stream_writer()` inside a tool "binds" the tool to a LangGraph context. Unit tests must either go through the Agent or add a conditional fallback for the writer (`try/except RuntimeError`). If the tool needs to remain independently testable, prefer keeping progress reporting in a wrapper layer or middleware layer, and have the tool itself just return structured results.
|
||||
@@ -0,0 +1,137 @@
|
||||
# Structured Output Issues
|
||||
|
||||
A focused reference for `with_structured_output`, `response_format`, and provider-side schema enforcement.
|
||||
|
||||
## Choose the API path first
|
||||
|
||||
LangChain exposes two related but different structured-output APIs. Diagnose and configure the one the application actually uses:
|
||||
|
||||
- **Model wrapper — `model.with_structured_output(...)`**: the model integration selects a steering method such as `json_schema`, `function_calling`, or `json_mode`, then parses the model response into the requested schema. Defaults vary by integration and version, so set `method` explicitly when behavior matters.
|
||||
- **Agent — `create_agent(response_format=...)`**: passing a schema directly lets the agent select a strategy using model capability metadata plus LangChain fallback and tool-compatibility checks. Both strategies parse and validate structured responses; `ToolStrategy` can retry configured structured-output errors, while `ProviderStrategy` surfaces provider or schema validation failures. These strategies are not aliases for the model wrapper's `method` argument.
|
||||
|
||||
Use the automatic agent strategy when the model profile is accurate:
|
||||
|
||||
```python
|
||||
agent = create_agent(
|
||||
model=model,
|
||||
tools=tools,
|
||||
response_format=MySchema,
|
||||
)
|
||||
```
|
||||
|
||||
Force a strategy only when provider capability is known and the automatic choice is unsuitable:
|
||||
|
||||
```python
|
||||
from langchain.agents.structured_output import ProviderStrategy, ToolStrategy
|
||||
|
||||
native_agent = create_agent(
|
||||
model=model,
|
||||
tools=tools,
|
||||
response_format=ProviderStrategy(MySchema),
|
||||
)
|
||||
|
||||
tool_agent = create_agent(
|
||||
model=model,
|
||||
tools=tools,
|
||||
response_format=ToolStrategy(MySchema),
|
||||
)
|
||||
```
|
||||
|
||||
## Issue 1: Unstable structured output, often returns None
|
||||
|
||||
- **Symptom**: `model.with_structured_output(schema)` returns `None`, an empty object, missing fields, or a parsing error. The same prompt can work on one run and fail on the next. Small, open-source, or quantized models tend to be less reliable with complex schemas.
|
||||
- **Cause**: The cause depends on the configured method and provider capability:
|
||||
- With provider-native structured output, the provider may reject an unsupported schema or model before generation.
|
||||
- OpenAI-family `with_structured_output(..., method="function_calling")` binds the schema as a tool and **forces the schema tool** by passing its name as `tool_choice`. When the provider supports and honors forced selection, the model is not free to answer in prose instead.
|
||||
- If an adapter strips or ignores `tool_choice` — for example through `disabled_params={"tool_choice": None}` — the model is again **free to skip the schema tool**. A natural-language answer then gives the tool parser nothing to deserialize, so the parsed result can be `None`.
|
||||
- A returned tool call can still fail schema validation because its arguments are malformed, incomplete, or have the wrong types.
|
||||
- **Solution**:
|
||||
|
||||
1. **Inspect the raw response before changing methods.** `include_raw=True` separates "no schema tool call" from "tool call failed validation":
|
||||
|
||||
```python
|
||||
structured_model = model.with_structured_output(
|
||||
MySchema,
|
||||
method="function_calling",
|
||||
include_raw=True,
|
||||
)
|
||||
result = structured_model.invoke(prompt)
|
||||
|
||||
print(result["raw"].tool_calls)
|
||||
print(result["parsed"])
|
||||
print(result["parsing_error"])
|
||||
```
|
||||
|
||||
2. **Choose a model-level method by documented provider capability.** There is no universal fallback order that works for every integration:
|
||||
|
||||
- Prefer provider-native `json_schema` when the selected model and provider document schema-constrained output:
|
||||
|
||||
```python
|
||||
structured_model = model.with_structured_output(MySchema, method="json_schema")
|
||||
```
|
||||
|
||||
- Use `function_calling` when the provider supports tool calling and forced `tool_choice`. Keep `with_structured_output` so LangChain both binds the schema tool and attaches the output parser:
|
||||
|
||||
```python
|
||||
structured_model = model.with_structured_output(
|
||||
MySchema,
|
||||
method="function_calling",
|
||||
)
|
||||
```
|
||||
|
||||
Calling `bind_tools` directly returns an `AIMessage` with tool calls; it does not provide the schema parsing performed by `with_structured_output`.
|
||||
|
||||
- Use `json_mode` when the provider guarantees JSON syntax but not schema conformance. Spell out field names, types, and constraints in the prompt, then validate and retry in application code:
|
||||
|
||||
```python
|
||||
structured_model = model.with_structured_output(MySchema, method="json_mode")
|
||||
```
|
||||
|
||||
3. **Configure agents through agent strategies.** For `create_agent`, use a direct schema for automatic selection, `ProviderStrategy` for known native support, or `ToolStrategy` for tool-based structured output and its validation-retry loop. Do not diagnose `create_agent(response_format=...)` solely through the model wrapper's three `method` values.
|
||||
|
||||
- **Lessons learned**:
|
||||
- Check `raw.tool_calls` and `parsing_error` before deciding whether the failure is tool selection or schema validation.
|
||||
- Provider-native schema enforcement is usually the strongest option when the exact model supports it, but capability must be verified rather than inferred from OpenAI-compatible transport alone.
|
||||
- Weak model + complex schema is the most failure-prone combination. Split large schemas into smaller calls when possible.
|
||||
|
||||
## Issue 2: `with_structured_output(..., method="function_calling")` fails because the model does not support `tool_choice`
|
||||
|
||||
- **Symptom**: Calling `model.with_structured_output(schema)` or `model.with_structured_output(schema, method="function_calling")` raises a provider-side 400 error like "`<model>` does not support this `tool_choice`". This often happens with `ChatOpenAI`-style integrations or derived classes such as `ChatDeepSeek` against OpenAI-compatible providers. A typical example is `deepseek-v4-flash` in thinking mode: it supports tool calling but rejects forced `tool_choice`.
|
||||
- **Cause**: For `function_calling`, LangChain's OpenAI-family chat models try to make structured output more reliable by forcing the schema tool to be called. Internally, `with_structured_output` binds the schema as a tool and usually passes `tool_choice=<schema_tool_name>`. That works for providers that support forced tool selection, but some OpenAI-compatible backends only allow free-form tool calling and reject explicit `tool_choice`. When you access those models through `ChatOpenAI`, `ChatDeepSeek`, or another subclass inheriting the same behavior, the adapter forwards `tool_choice` and the provider errors before generation starts.
|
||||
- **Solution**: Disable forwarding of `tool_choice` for that model instance, so LangChain still binds the schema tool but does not send the unsupported parameter:
|
||||
|
||||
```python
|
||||
from langchain_deepseek import ChatDeepSeek
|
||||
from pydantic import BaseModel
|
||||
|
||||
class User(BaseModel):
|
||||
name: str
|
||||
age: int
|
||||
|
||||
model = ChatDeepSeek(
|
||||
model="deepseek-v4-flash",
|
||||
disabled_params={"tool_choice": None},
|
||||
extra_body={"thinking": {"type": "enabled"}},
|
||||
)
|
||||
|
||||
structured_model = model.with_structured_output(
|
||||
User,
|
||||
method="function_calling",
|
||||
)
|
||||
|
||||
print(structured_model.invoke("My name is John and I am 25 years old."))
|
||||
```
|
||||
|
||||
Why this model-wrapper workaround works: OpenAI-family `with_structured_output(..., method="function_calling")` calls `_filter_disabled_params(...)` before binding the schema tool. Setting `disabled_params={"tool_choice": None}` strips the forced selection from this model-wrapper path.
|
||||
|
||||
This setting does not configure agent `ToolStrategy`. Stock `create_agent(response_format=ToolStrategy(...))` still forces structured-tool use through `model.bind_tools(..., tool_choice="any")`, and ordinary `bind_tools` does not consult `disabled_params`. If an agent's provider rejects forced tool selection, use a documented native `ProviderStrategy`, a compatible model or integration, or an adapter that explicitly handles this provider limitation.
|
||||
|
||||
Removing forced selection restores compatibility but also makes the model free to skip the schema tool. Use `include_raw=True` to detect that case and choose another method by capability:
|
||||
- If the provider documents native schema-constrained decoding for this model, use `method="json_schema"`.
|
||||
- If it supports only JSON syntax enforcement, use `method="json_mode"`, describe the schema in the prompt, and validate plus retry in application code.
|
||||
- If neither is available, keep unforced `function_calling` only with an explicit retry or repair path for missing tool calls.
|
||||
- **Lessons learned**:
|
||||
- This failure mode is not "the schema is wrong" and not "tool calling is unsupported" in general. The narrower issue is that the backend rejects forced tool selection via `tool_choice`.
|
||||
- `ChatOpenAI` compatibility is only a transport-level starting point. Once you connect it to non-OpenAI providers, verify which OpenAI request parameters they actually support, especially for reasoning models.
|
||||
- When a provider says "`... does not support this tool_choice`", the fastest fix is usually model-level configuration (`disabled_params`) rather than patching LangChain source code.
|
||||
- Removing `tool_choice` trades reliability for compatibility. If outputs start coming back as `None`, see Issue 1 and switch to `json_schema`, `json_mode`, or an explicit retry path based on provider capability.
|
||||
@@ -0,0 +1,108 @@
|
||||
# User Query Convention
|
||||
|
||||
A cross-platform convention for skill authors to define structured questions. Agents parse these blocks and render them via the best available tool on their platform.
|
||||
|
||||
## Block Types
|
||||
|
||||
### `<!-- query -->` — Single or multi-choice question
|
||||
|
||||
Use when the user must pick between approaches, modes, or options.
|
||||
|
||||
```markdown
|
||||
<!-- query
|
||||
type: choice
|
||||
question: "Which approach do you prefer?"
|
||||
options:
|
||||
- label: "Code generation"
|
||||
description: "Generate integration class in your repo"
|
||||
- label: "Third-party library"
|
||||
description: "Use langchain-dev-utils built-in adapters"
|
||||
default: 1
|
||||
-->
|
||||
```
|
||||
|
||||
Fields:
|
||||
- `type`: `choice` (single-select) or `multi-choice` (multi-select)
|
||||
- `question`: The question to present
|
||||
- `options`: 2–4 options, each with `label` and `description`
|
||||
- `default`: 1-based index of the default option (applied when user says "use defaults" or doesn't answer)
|
||||
|
||||
### `<!-- gather -->` — Collect multiple inputs
|
||||
|
||||
Use when the skill needs several pieces of information from the user before proceeding.
|
||||
|
||||
```markdown
|
||||
<!-- gather
|
||||
prompt: "Confirm the following details:"
|
||||
fields:
|
||||
- name: model_name
|
||||
question: "Model name (lowercase)"
|
||||
example: "qwen"
|
||||
required: true
|
||||
- name: api_base
|
||||
question: "API base URL"
|
||||
example: "https://dashscope.aliyuncs.com/compatible-mode/v1"
|
||||
required: true
|
||||
- name: api_key_env
|
||||
question: "API key env var name"
|
||||
example: "QWEN_API_KEY"
|
||||
required: true
|
||||
fallback: "Use reasonable defaults from the provider's documentation."
|
||||
-->
|
||||
```
|
||||
|
||||
Fields:
|
||||
- `prompt`: Introductory text shown before the questions
|
||||
- `fields`: List of inputs to collect; each has `name`, `question`, `example`, and optional `required` (default true)
|
||||
- `fallback`: Instruction for the agent when the user declines to answer or says "just use defaults"
|
||||
|
||||
## Platform Rendering
|
||||
|
||||
| Platform | `<!-- query -->` | `<!-- gather -->` |
|
||||
|----------|-----------------|-------------------|
|
||||
| Claude Code | `AskUserQuestion` with `options` | `AskUserQuestion` with one question per field |
|
||||
| Gemini CLI | `ask_user` | `ask_user` per field |
|
||||
| Copilot CLI | Output as formatted text with numbered options, wait for reply | Output as numbered list with examples, wait for reply |
|
||||
| Cursor / Windsurf | Output as formatted text, wait for reply | Output as formatted text, wait for reply |
|
||||
| Codex | Output as formatted text (autonomous mode — apply defaults if no response) | Apply defaults (autonomous mode) |
|
||||
|
||||
### Claude Code example rendering
|
||||
|
||||
For a `<!-- query -->` block, the agent calls:
|
||||
```
|
||||
AskUserQuestion({
|
||||
questions: [{
|
||||
question: "Which approach do you prefer?",
|
||||
header: "Approach",
|
||||
options: [
|
||||
{ label: "Code generation", description: "Generate integration class in your repo" },
|
||||
{ label: "Third-party library", description: "Use langchain-dev-utils built-in adapters" }
|
||||
],
|
||||
multiSelect: false
|
||||
}]
|
||||
})
|
||||
```
|
||||
|
||||
For a `<!-- gather -->` block, the agent calls `AskUserQuestion` with up to 4 questions (the tool's limit), batching if needed.
|
||||
|
||||
### Fallback text rendering (Cursor, Copilot, Codex)
|
||||
|
||||
For platforms without structured prompting, output:
|
||||
|
||||
```
|
||||
**Which approach do you prefer?**
|
||||
|
||||
1. **Code generation** — Generate integration class in your repo
|
||||
2. **Third-party library** — Use langchain-dev-utils built-in adapters
|
||||
|
||||
(Reply with number or description. Default: 1)
|
||||
```
|
||||
|
||||
## Guidelines for Skill Authors
|
||||
|
||||
1. **Place blocks inline** where the question naturally occurs in the skill flow — not in a separate section
|
||||
2. **Always provide a `default` or `fallback`** — agents running in autonomous mode need a way to proceed without blocking
|
||||
3. **Keep options to 2–4** — matches `AskUserQuestion` limits and avoids decision fatigue
|
||||
4. **Use `<!-- gather -->` sparingly** — prefer inferring from project context (package manager, existing config) over asking
|
||||
5. **Blocks are HTML comments** — they don't render in markdown viewers, so the surrounding prose should still make sense without them
|
||||
6. **Prose context around blocks is required** — the block is for the agent's structured rendering; the surrounding markdown provides context for human readers browsing the skill file
|
||||
@@ -0,0 +1,205 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from json import JSONDecodeError
|
||||
from typing import Any, Self, cast
|
||||
|
||||
import openai
|
||||
from langchain_core.language_models import (
|
||||
LanguageModelInput,
|
||||
ModelProfile,
|
||||
ModelProfileRegistry,
|
||||
)
|
||||
from langchain_core.messages import AIMessage, AIMessageChunk
|
||||
from langchain_core.outputs import ChatGenerationChunk, ChatResult
|
||||
from langchain_core.utils import from_env, secret_from_env
|
||||
from langchain_openai.chat_models.base import BaseChatOpenAI, _convert_message_to_dict
|
||||
from pydantic import Field, SecretStr, model_validator
|
||||
|
||||
_INVALID_RESPONSE_ERROR_MSG = (
|
||||
"Provider API returned an invalid response. "
|
||||
"Please check your API key, network connection, or API base URL."
|
||||
)
|
||||
|
||||
|
||||
def _get_default_model_profile(model_name: str) -> ModelProfile:
|
||||
try:
|
||||
from .data._profiles import _PROFILES # type: ignore[import-untyped]
|
||||
except ImportError:
|
||||
return {}
|
||||
|
||||
_MODEL_PROFILES = cast("ModelProfileRegistry", _PROFILES)
|
||||
default = _MODEL_PROFILES.get(model_name) or {}
|
||||
return default.copy()
|
||||
|
||||
|
||||
class ChatModel(BaseChatOpenAI):
|
||||
api_key: SecretStr | None = Field(
|
||||
default_factory=secret_from_env("PROVIDER_API_KEY", default=None),
|
||||
)
|
||||
api_base: str = Field(
|
||||
default_factory=from_env(
|
||||
"PROVIDER_API_BASE",
|
||||
default="PROVIDER_API_BASE_URL",
|
||||
),
|
||||
)
|
||||
|
||||
@property
|
||||
def _llm_type(self) -> str:
|
||||
return "chat-provider"
|
||||
|
||||
@property
|
||||
def lc_secrets(self) -> dict[str, str]:
|
||||
return {"api_key": "PROVIDER_API_KEY"}
|
||||
|
||||
@model_validator(mode="after")
|
||||
def validate_environment(self) -> Self:
|
||||
if not (self.api_key and self.api_key.get_secret_value()):
|
||||
msg = "PROVIDER_API_KEY must be set."
|
||||
raise ValueError(msg)
|
||||
client_params: dict = {
|
||||
"api_key": self.api_key.get_secret_value() if self.api_key else None,
|
||||
"base_url": self.api_base,
|
||||
"timeout": self.request_timeout,
|
||||
"max_retries": self.max_retries,
|
||||
"default_headers": self.default_headers,
|
||||
"default_query": self.default_query,
|
||||
}
|
||||
client_params = {k: v for k, v in client_params.items() if v is not None}
|
||||
if not (self.client or None):
|
||||
sync_specific: dict = {"http_client": self.http_client}
|
||||
self.root_client = openai.OpenAI(**client_params, **sync_specific)
|
||||
self.client = self.root_client.chat.completions
|
||||
if not (self.async_client or None):
|
||||
async_specific: dict = {"http_client": self.http_async_client}
|
||||
self.root_async_client = openai.AsyncOpenAI(
|
||||
**client_params, **async_specific,
|
||||
)
|
||||
self.async_client = self.root_async_client.chat.completions
|
||||
return self
|
||||
|
||||
def _resolve_model_profile(self) -> ModelProfile | None:
|
||||
return _get_default_model_profile(self.model_name) or None
|
||||
|
||||
def _get_request_payload(
|
||||
self,
|
||||
input_: LanguageModelInput,
|
||||
*,
|
||||
stop: list[str] | None = None,
|
||||
**kwargs: Any,
|
||||
) -> dict:
|
||||
payload = super()._get_request_payload(input_, stop=stop, **kwargs)
|
||||
messages = self._convert_input(input_).to_messages()
|
||||
payload_messages = []
|
||||
for m in messages:
|
||||
if isinstance(m, AIMessage):
|
||||
msg_dict = _convert_message_to_dict(m)
|
||||
if m.additional_kwargs.get("reasoning_content"):
|
||||
msg_dict["reasoning_content"] = m.additional_kwargs.get(
|
||||
"reasoning_content",
|
||||
)
|
||||
payload_messages.append(msg_dict)
|
||||
else:
|
||||
payload_messages.append(_convert_message_to_dict(m))
|
||||
payload["messages"] = payload_messages
|
||||
if "tools" in payload and len(payload["tools"]) == 0:
|
||||
payload.pop("tools")
|
||||
return payload
|
||||
|
||||
def _create_chat_result(
|
||||
self,
|
||||
response: dict | openai.BaseModel,
|
||||
generation_info: dict | None = None,
|
||||
) -> ChatResult:
|
||||
rtn = super()._create_chat_result(response, generation_info)
|
||||
|
||||
if not isinstance(response, openai.BaseModel):
|
||||
return rtn
|
||||
|
||||
for generation in rtn.generations:
|
||||
if generation.message.response_metadata is None:
|
||||
generation.message.response_metadata = {}
|
||||
generation.message.response_metadata["model_provider"] = "provider-name"
|
||||
|
||||
choices = getattr(response, "choices", None)
|
||||
if choices and hasattr(choices[0].message, "reasoning_content"):
|
||||
rtn.generations[0].message.additional_kwargs["reasoning_content"] = (
|
||||
choices[0].message.reasoning_content
|
||||
)
|
||||
return rtn
|
||||
|
||||
def _convert_chunk_to_generation_chunk(
|
||||
self,
|
||||
chunk: dict,
|
||||
default_chunk_class: type,
|
||||
base_generation_info: dict | None,
|
||||
) -> ChatGenerationChunk | None:
|
||||
generation_chunk = super()._convert_chunk_to_generation_chunk(
|
||||
chunk,
|
||||
default_chunk_class,
|
||||
base_generation_info,
|
||||
)
|
||||
if (choices := chunk.get("choices")) and generation_chunk:
|
||||
top = choices[0]
|
||||
if isinstance(generation_chunk.message, AIMessageChunk):
|
||||
generation_chunk.message.response_metadata = {
|
||||
**generation_chunk.message.response_metadata,
|
||||
"model_provider": "provider-name",
|
||||
}
|
||||
if (
|
||||
reasoning_content := top.get("delta", {}).get("reasoning_content")
|
||||
) is not None:
|
||||
generation_chunk.message.additional_kwargs["reasoning_content"] = (
|
||||
reasoning_content
|
||||
)
|
||||
return generation_chunk
|
||||
|
||||
def _stream(self, messages, stop=None, run_manager=None, **kwargs):
|
||||
kwargs["stream_options"] = {"include_usage": True}
|
||||
try:
|
||||
yield from super()._stream(
|
||||
messages, stop=stop, run_manager=run_manager, **kwargs,
|
||||
)
|
||||
except JSONDecodeError as e:
|
||||
raise JSONDecodeError(
|
||||
_INVALID_RESPONSE_ERROR_MSG,
|
||||
e.doc,
|
||||
e.pos,
|
||||
) from e
|
||||
|
||||
async def _astream(self, messages, stop=None, run_manager=None, **kwargs):
|
||||
kwargs["stream_options"] = {"include_usage": True}
|
||||
try:
|
||||
async for chunk in super()._astream(
|
||||
messages, stop=stop, run_manager=run_manager, **kwargs,
|
||||
):
|
||||
yield chunk
|
||||
except JSONDecodeError as e:
|
||||
raise JSONDecodeError(
|
||||
_INVALID_RESPONSE_ERROR_MSG,
|
||||
e.doc,
|
||||
e.pos,
|
||||
) from e
|
||||
|
||||
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
|
||||
try:
|
||||
return super()._generate(
|
||||
messages, stop=stop, run_manager=run_manager, **kwargs,
|
||||
)
|
||||
except JSONDecodeError as e:
|
||||
raise JSONDecodeError(
|
||||
_INVALID_RESPONSE_ERROR_MSG,
|
||||
e.doc,
|
||||
e.pos,
|
||||
) from e
|
||||
|
||||
async def _agenerate(self, messages, stop=None, run_manager=None, **kwargs):
|
||||
try:
|
||||
return await super()._agenerate(
|
||||
messages, stop=stop, run_manager=run_manager, **kwargs,
|
||||
)
|
||||
except JSONDecodeError as e:
|
||||
raise JSONDecodeError(
|
||||
_INVALID_RESPONSE_ERROR_MSG,
|
||||
e.doc,
|
||||
e.pos,
|
||||
) from e
|
||||
+1
@@ -0,0 +1 @@
|
||||
../../.agents/skills/langchain-dev-guide
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"version": 1,
|
||||
"skills": {
|
||||
"langchain-dev-guide": {
|
||||
"source": "ob-labs/agentseek",
|
||||
"sourceType": "github",
|
||||
"skillPath": "skills/langchain-dev-guide/SKILL.md",
|
||||
"computedHash": "5772776e13375f078d3743c0899a28ffeb20bc6e9f7884e3b6e9d3826b0f6abc"
|
||||
}
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user