LangChain Agents Documentation: create_agent, Middleware, and What Changed Since 0.3

TLDR: LangChain's official agent documentation now lives at docs.langchain.com, and the current entry point is create_agent from langchain.agents, which builds a LangGraph graph you customize with middleware. Legacy AgentExecutor and initialize_agent moved to langchain-classic. This map, checked on September 28, 2026 against langchain 1.4.3, shows which official page answers which question, what changed since 0.3, and a working search agent.
A search for LangChain's agent docs returns three generations of advice: old python.langchain.com pages, LangGraph create_react_agent tutorials, and blog posts built on initialize_agent. The last group fails on a current install, because langchain.agents in 1.x exports only create_agent and AgentState (source). This page does not replace the official documentation. It routes each question to the right official page, records what changed between versions, and assembles the documented pieces into one agent you can run. API names were checked against the docs, API reference, and GitHub source.
If you want the concepts before the API, our guide to how the agent loop works covers goals, context, actions, and evaluation without tying them to a framework.
Where does the official LangChain agents documentation live?
The canonical page is Agents in the LangChain Python docs. It defines an agent as "a model calling tools in a loop until a given task is complete" and documents create_agent. Old links still resolve as redirects: python.langchain.com and its agent pages land on the LangChain overview, and the retired LangGraph agents page forwards to the new Agents page. Parameter-level detail lives in the create_agent API reference. TypeScript has a parallel track where the same harness is createAgent, imported from langchain (JavaScript Agents page).
| Your question | Official page | What it covers |
|---|---|---|
| How do I create and run an agent? | Agents | Model, tools, system prompt, state, invocation, streaming |
| How do I define a tool? | Tools | The tool decorator, schemas, runtime context, return types, MCP tools |
| How do I customize the loop? | Middleware overview, prebuilt, custom | Hooks, retries, call limits, human approval, PII handling |
| Which model string do I pass? | Models | Provider and model identifiers, parameters, dynamic model selection |
| How do I get typed output? | Structured output | Tool and provider strategies for response formats |
| How do I keep conversation history? | Short-term memory | Checkpointers, thread IDs, custom state |
| How do I connect an MCP server? | Model Context Protocol | Transports, authentication, tool results (beta) |
| How do I upgrade old code? | LangChain v1 migration guide | Moving to create_agent, hooks to middleware, namespace moves |
| What shipped, and when? | Changelog, release policy | Release notes, support windows, deprecation rules |
| How do I test without API calls? | Unit testing | Fake chat models and in-memory checkpointers |
If you work from an editor, the docs also run a public MCP server at https://docs.langchain.com/mcp that needs no API key, so a coding agent can search them directly.
Which entry point should you use: create_agent, Deep Agents, or LangGraph?
LangChain's frameworks, runtimes, and harnesses page separates three layers: LangChain is the framework, LangGraph the runtime, and Deep Agents the harness. The LangChain overview recommends Deep Agents for a batteries-included start, create_agent for a harness you configure yourself, and LangGraph for workflows that mix deterministic and agentic steps.
| Entry point | Import | Reach for it when | Status, September 2026 |
|---|---|---|---|
| create_agent | from langchain.agents import create_agent | You want the standard tool loop and will add capabilities through middleware | Stable; the 1.x line has long-term support |
| create_deep_agent | from deepagents import create_deep_agent | Long, multi-step tasks that need planning, files, subagents, and summarization built in | Pre-1.0 (0.7.19); minor releases can break; needs Python 3.11 or later |
| StateGraph | from langgraph.graph import StateGraph | Custom topology: routing, fan-out, or deterministic steps around agent calls | Stable; the 1.x line has long-term support |
| AgentExecutor, initialize_agent | from langchain_classic.agents import ... | Only while maintaining code you cannot migrate yet | Legacy namespace; initialize_agent and AgentType are deprecated, with removal targeted for 2.0.0 |
The layers nest rather than compete. create_agent returns a compiled LangGraph graph, and the middleware overview notes that hooks run inside that graph, so you can drop a whole agent into a larger StateGraph as a node and keep its retries and approvals. Deep Agents builds on LangChain's agent building blocks and the LangGraph runtime (overview, PyPI), and as a pre-1.0 package it churns: 0.7.0 (July 2026) stopped including to-do planning by default. If you are still deciding whether the task needs an agent loop at all, our guide to AI agent architecture patterns covers when a fixed workflow fits better.
What changed between LangChain 0.3 and 1.4?
LangChain 1.0 shipped in October 2025 and reorganized the package around agents. The v1 migration guide and the langchain-classic source account for most of what breaks old tutorials.
| Old pattern | Status in langchain 1.4.3 | Use instead |
|---|---|---|
from langchain.agents import initialize_agent, AgentType | Import fails; both moved to langchain_classic.agents, deprecated since 0.1.0 | create_agent |
AgentExecutor(..., max_iterations=15) | Lives in langchain_classic.agents; 15 is its default cap | Call-limit middleware (see below) |
langgraph.prebuilt.create_react_agent | Deprecated in LangGraph v1 | create_agent, with prompt= renamed system_prompt= |
pre_model_hook, post_model_hook | Replaced | Middleware with before_model and after_model |
tools=ToolNode([...], handle_tool_errors=...) | A ToolNode is no longer accepted as tools | wrap_tool_call middleware or ToolErrorMiddleware |
response_format=("prompt", Schema) | Prompted output removed | ToolStrategy or ProviderStrategy |
Per-run data in config["configurable"] | Superseded for static context; thread IDs stay in config | context= on invoke, context_schema= on the agent |
langchain.chains, langchain.retrievers, langchain.hub | Moved | The same modules under langchain_classic |
MultiServerMCPClient from langchain-mcp-adapters | Replaced in 1.4.0 | MCPAdapter from langchain.mcp (beta) |
Three smaller changes bite in code review. Custom state must be a TypedDict that extends AgentState, since Pydantic models and dataclasses are no longer supported. Code that filters streamed events by node name must match "model" instead of "agent". And message text is now a property, message.text; calling it as a method is deprecated in langchain-core 1.0.
The minor releases since 1.0, from the changelog:
- 1.1 (November 2025): model profiles,
ModelRetryMiddleware, andSystemMessagesupport for the system prompt. - 1.2 (December 2025): provider-specific tool options through a tool
extrasattribute, and strict schema adherence for structured output. - 1.3 and LangGraph 1.2 (May 2026):
version="v3"event streaming, plus LangGraph per-node timeouts and node-level error handlers. - 1.4 (September 2026): MCP support inside LangChain as the beta
langchain.mcpnamespace, built on FastMCP. Patch 1.4.3 shipped on September 28, 2026 and requires langgraph 1.2.11 or later, below 1.3.0 (PyPI).
Support windows matter if you still run 0.3 code. Per the release policy, LangChain 0.3 and LangGraph 0.4 are in maintenance mode until December 2026, receiving security patches and critical fixes only. LangChain 1.0 and LangGraph 1.0 are long-term support lines: active until 2.0 ships, then in maintenance for at least a year, and deprecated features keep working throughout 1.x. Patch releases can land up to a few times a week, so pin exact versions and rerun your tests before each upgrade.
How do tools work in create_agent?
The Tools page covers the three shapes create_agent accepts in tools=: plain Python callables with type hints and a docstring, functions decorated with @tool, and dicts describing a provider's built-in tools. Type hints are required because they become the input schema, and the docstring becomes the description the model reads when deciding whether to call the tool. The docs recommend snake_case names for compatibility across providers and reserve two argument names, config and runtime; a tool that needs agent state or per-run context takes a ToolRuntime parameter instead. A tool can return a string, an object for the model to inspect, or a Command that writes to state. For long tool lists, the same page covers dynamic tool selection.
create_agent runs tools through LangGraph's ToolNode, whose default handler turns invalid-argument errors into a message the model can correct but re-raises any exception thrown inside your tool (source). A timeout in your HTTP call therefore ends the whole run unless middleware catches it. The example further down handles that explicitly.
Tools from MCP servers
Install langchain[mcp] (1.4.0 or later), open an MCPAdapter on a server, and pass the result of list_tools() to create_agent. The namespace is in beta and raises a LangChainBetaWarning on import. You.com runs a hosted MCP server whose keyless free profile exposes you-search and you-discover at 100 queries per day, according to the MCP server docs. Adding tools=you-search narrows it to one tool, because enabled tools are the intersection of the profile and the allowlist.
import asyncio
from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
# Keyless free profile, narrowed to one tool: enabled tools are the
# intersection of the profile ceiling and the ?tools= allowlist.
YOU_MCP_URL = "https://api.you.com/mcp?profile=free&tools=you-search"
async def main() -> None:
async with MCPAdapter(YOU_MCP_URL) as adapter:
tools = await adapter.list_tools()
agent = create_agent("openai:gpt-5.5", tools)
result = await agent.ainvoke(
{"messages": [{"role": "user", "content": "Summarize this week's LangChain releases."}]}
)
print(result["messages"][-1].text)
if __name__ == "__main__":
asyncio.run(main())
Adapted tools keep their server names, so the model sees you-search. For keyed access to the server's four default tools, pass a FastMCP client that sends your key as a bearer token, MCPAdapter(Client("https://api.you.com/mcp", auth=os.environ["YDC_API_KEY"])), the pattern on LangChain's MCP authentication page. If MCP is new to you, start with what the Model Context Protocol is.
How does middleware change the agent loop?
Middleware is the customization layer in LangChain 1.x. The v1 release notes list six hooks: before_agent, before_model, wrap_model_call, wrap_tool_call, after_model, and after_agent. You can subclass AgentMiddleware or use the matching decorators, and the prebuilt catalog covers the common cases. These are the ones to know first:
| Middleware | What it does | Detail that matters |
|---|---|---|
| ToolErrorMiddleware | Turns selected tool exceptions into error messages the model can see | Needs langchain 1.3.14 or later; a handler that returns nothing lets the exception halt the run |
| ToolRetryMiddleware | Retries failed tool calls with exponential backoff | Exceptions that do not match its retry filter are re-raised immediately |
| ModelCallLimitMiddleware | Caps model calls per run or per thread | Thread limits need a checkpointer; the default exit ends the run gracefully |
| ToolCallLimitMiddleware | Caps tool calls per run or per thread, globally or for one named tool | By default, extra calls are blocked with an error message and the model decides how to finish |
| SummarizationMiddleware | Summarizes history as it nears the context limit | Trigger points can come from model profiles |
| HumanInTheLoopMiddleware | Pauses for approval before selected tool calls | Matches on each tool's name |
Order in the middleware list is behavior, not style. Earlier entries wrap later ones, so the docs put ToolErrorMiddleware before ToolRetryMiddleware and set on_failure="error" on the retry layer. Transient failures get retried first, and only an exhausted retry reaches the error handler.
One default changed quietly between generations. The legacy AgentExecutor stopped after 15 iterations unless told otherwise (source). create_agent compiles its graph with a recursion limit of 9,999 steps (source), so no small built-in cap stops a looping agent. Put the cap in middleware: ModelCallLimitMiddleware bounds model calls per invocation, and ToolCallLimitMiddleware bounds calls to one expensive tool.
A working example: a search agent built from the documented pieces
The official pages show each piece with stub tools; here they are assembled around a real one. The tool calls the You.com Web Search API over REST: a POST to https://ydc-index.io/v1/search with an X-API-Key header, as documented in the Web Search API guide. It requests highlights, the query-relevant passages the guide recommends for grounding an agent. You need Python 3.10 or later, a You.com API key, and an OpenAI key for the model string LangChain's own examples use.
python -m pip install "langchain==1.4.3" "langchain-openai==1.6.6"
export YDC_API_KEY="your-you.com-api-key"
export OPENAI_API_KEY="your-openai-api-key"
import json
import os
import urllib.error
import urllib.request
from typing import Optional
from langchain.agents import create_agent
from langchain.agents.middleware import (
ModelCallLimitMiddleware,
ToolCallLimitMiddleware,
ToolCallRequest,
ToolErrorMiddleware,
ToolRetryMiddleware,
)
from langchain.tools import tool
from langgraph.checkpoint.memory import InMemorySaver
SEARCH_URL = "https://ydc-index.io/v1/search"
TRANSIENT_STATUS = {429, 500, 502, 503, 504}
class TransientSearchError(Exception):
"""Rate limits, server errors, and network failures: safe to retry."""
class PermanentSearchError(Exception):
"""Missing key, bad key, no credits, missing scope, bad parameters: fix config."""
def search_you(query: str, count: int = 5, timeout: float = 15.0) -> dict:
key = os.environ.get("YDC_API_KEY", "").strip()
if not key:
raise PermanentSearchError("YDC_API_KEY is not set")
body = json.dumps({
"query": query,
"count": count,
"extraction": {"extraction_mode": "highlights"},
}).encode("utf-8")
request = urllib.request.Request(
SEARCH_URL,
data=body,
method="POST",
headers={"X-API-Key": key, "Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
return json.load(response)
except urllib.error.HTTPError as err:
if err.code in TRANSIENT_STATUS:
raise TransientSearchError(f"HTTP {err.code}") from err
raise PermanentSearchError(f"HTTP {err.code}") from err
except (urllib.error.URLError, TimeoutError) as err:
raise TransientSearchError(type(err).__name__) from err
def format_results(payload: dict, max_passages: int = 2) -> str:
web = (payload.get("results") or {}).get("web") or []
if not web:
return "NO_RESULTS: rephrase once, or answer without search and say so."
blocks = []
for hit in web:
contents = hit.get("contents") or {}
passages = (contents.get("highlights") or hit.get("snippets")
or [hit.get("description") or ""])
text = " ".join(p for p in passages[:max_passages] if p)
blocks.append(f"{hit.get('title') or 'Untitled'}\n{hit.get('url', '')}\n{text}")
return "\n\n".join(blocks)
@tool
def web_search(query: str) -> str:
"""Search the live web for facts that may have changed after training.
Returns titles, URLs, and query-relevant passages. Cite the URLs you use."""
return format_results(search_you(query))
def on_search_error(exc: Exception, request: ToolCallRequest) -> Optional[str]:
if isinstance(exc, TransientSearchError):
return ("web_search is temporarily unavailable. Answer from results you "
"already have and say which claims you could not verify.")
return None # PermanentSearchError and anything unexpected halt the run
def build_agent():
return create_agent(
model="openai:gpt-5.5",
tools=[web_search],
system_prompt=(
"Search before answering anything time-sensitive, at most three "
"times per question, and cite the URL behind every searched claim."
),
middleware=[
ModelCallLimitMiddleware(run_limit=6),
ToolCallLimitMiddleware(tool_name="web_search", run_limit=3),
ToolErrorMiddleware(on_search_error, tools=["web_search"]),
ToolRetryMiddleware(
max_retries=2,
retry_on=(TransientSearchError,),
on_failure="error",
tools=["web_search"],
),
],
checkpointer=InMemorySaver(),
)
if __name__ == "__main__":
agent = build_agent()
config = {"configurable": {"thread_id": "docs-map-demo"}}
question = "What changed in the most recent LangChain 1.4 patch release?"
result = agent.invoke(
{"messages": [{"role": "user", "content": question}]},
config=config,
)
print(result["messages"][-1].text)
What each piece does:
- The error classes follow documented status codes. You.com answers 429 when you exceed the rate limit and asks clients to back off and retry; 401, 402, 403, and 422 mean a bad key, no credits, a missing scope, or invalid parameters (error reference). Retrying those wastes calls, so they are permanent.
- Retry runs inside error handling. Transient failures get two retries with exponential backoff starting at one second; if they still fail,
on_search_errortells the model search is down so it can answer with a caveat. The middleware does not read theRetry-Afterheader the rate limit docs ask clients to honor, so raiseinitial_delayif 429s persist. Permanent errors halt the run, so monitoring sees a configuration problem instead of a quietly degraded answer. - Two limits do the job max_iterations used to do: at most three searches and six model calls per question. A fourth search request is blocked with an error message, and the model decides how to finish.
- The tool reads the web results only, keeps every URL, and caps passages at two per result so search output does not crowd the context window.
The cost is easy to bound. The Web Search API costs $5.00 per 1,000 calls as of September 2026, and highlights are included in that base price, per the billing page. Each search costs half a cent, so the three-search cap holds search spend to $0.015 per question before model tokens. New accounts start with $100 in free credits. Self-serve keys default to 10 requests per second on this endpoint (rate limits); concurrent users can hit that, which is what the 429 retry path handles.
If you would rather not write the tool, LangChain's own tools index lists You.com Search and links to You.com's LangChain integration page for the langchain-youdotcom package. Its YouSearchTool registers under the name you_search, so change tool_name in the limit middleware if you swap it in. Our LangChain web search tool guide covers that package's tools, retriever, and parameters.
What should you test before shipping a LangChain agent?
Test the tool first, without a network. Fixtures shaped like the documented response make parsing and error classification deterministic. Save the example as search_agent.py; this test file adds nothing beyond the standard library:
import io
import unittest
import urllib.error
from unittest import mock
import search_agent as sa
FIXTURE = {
"results": {"web": [{
"url": "https://example.com/release-notes",
"title": "Release notes",
"description": "Fallback text",
"contents": {"highlights": ["First passage.", "Second passage.", "Third."]},
}]},
"metadata": {"query": "q", "search_uuid": "0000", "latency": 0.3},
}
def http_error(code):
return urllib.error.HTTPError(sa.SEARCH_URL, code, "error", {}, io.BytesIO(b"{}"))
class SearchToolTests(unittest.TestCase):
def test_keeps_urls_and_caps_passages(self):
text = sa.format_results(FIXTURE)
self.assertIn("https://example.com/release-notes", text)
self.assertIn("First passage. Second passage.", text)
self.assertNotIn("Third.", text)
def test_empty_results_return_sentinel(self):
self.assertTrue(sa.format_results({"results": {"web": []}}).startswith("NO_RESULTS"))
def test_rate_limit_is_transient(self):
with mock.patch.dict("os.environ", {"YDC_API_KEY": "test"}), \
mock.patch("urllib.request.urlopen", side_effect=http_error(429)):
with self.assertRaises(sa.TransientSearchError):
sa.search_you("q")
def test_bad_key_is_permanent(self):
with mock.patch.dict("os.environ", {"YDC_API_KEY": "test"}), \
mock.patch("urllib.request.urlopen", side_effect=http_error(401)):
with self.assertRaises(sa.PermanentSearchError):
sa.search_you("q")
if __name__ == "__main__":
unittest.main()
Then test the agent with a scripted model. The unit testing page pairs GenericFakeChatModel with an InMemorySaver checkpointer to replay exact responses, tool calls included, with no API key. One catch: create_agent calls the model's bind_tools whenever tools are registered, and GenericFakeChatModel inherits the base implementation, which raises NotImplementedError (source). For an agent with tools, subclass the fake and have bind_tools return self. Then check that:
- a scripted fourth search is blocked and the run still ends with an answer;
- a persistent 429 produces the "temporarily unavailable" tool message after the retries, not a crash;
- a 401 halts the run and surfaces the permanent error in your logs;
- CI promotes
LangChainDeprecationWarningto an error, so calls to deprecated APIs such asinitialize_agentfail loudly, and handlesLangChainBetaWarningdeliberately if you uselangchain.mcp.
Search results are untrusted input. A page can carry instructions aimed at your agent, so keep tool output clearly separated from your system prompt; our guide to prompt injection and intent hijacking covers the defenses. To score the finished agent on real tasks, see how to run an AI agent evaluation.
Related Guides
LI Test
LI Test
Share Article:
