Blog
 / 

LangChain Agents Documentation: create_agent, Middleware, and What Changed Since 0.3

LangChain Agents Documentation: create_agent, Middleware, and What Changed Since 0.3

TLDR: LangChain's official agent documentation now lives at docs.langchain.com, and the current entry point is create_agent from langchain.agents, which builds a LangGraph graph you customize with middleware. Legacy AgentExecutor and initialize_agent moved to langchain-classic. This map, checked on September 28, 2026 against langchain 1.4.3, shows which official page answers which question, what changed since 0.3, and a working search agent.

A search for LangChain's agent docs returns three generations of advice: old python.langchain.com pages, LangGraph create_react_agent tutorials, and blog posts built on initialize_agent. The last group fails on a current install, because langchain.agents in 1.x exports only create_agent and AgentState (source). This page does not replace the official documentation. It routes each question to the right official page, records what changed between versions, and assembles the documented pieces into one agent you can run. API names were checked against the docs, API reference, and GitHub source.

If you want the concepts before the API, our guide to how the agent loop works covers goals, context, actions, and evaluation without tying them to a framework.

Where does the official LangChain agents documentation live?

The canonical page is Agents in the LangChain Python docs. It defines an agent as "a model calling tools in a loop until a given task is complete" and documents create_agent. Old links still resolve as redirects: python.langchain.com and its agent pages land on the LangChain overview, and the retired LangGraph agents page forwards to the new Agents page. Parameter-level detail lives in the create_agent API reference. TypeScript has a parallel track where the same harness is createAgent, imported from langchain (JavaScript Agents page).

Your questionOfficial pageWhat it covers
How do I create and run an agent?AgentsModel, tools, system prompt, state, invocation, streaming
How do I define a tool?ToolsThe tool decorator, schemas, runtime context, return types, MCP tools
How do I customize the loop?Middleware overview, prebuilt, customHooks, retries, call limits, human approval, PII handling
Which model string do I pass?ModelsProvider and model identifiers, parameters, dynamic model selection
How do I get typed output?Structured outputTool and provider strategies for response formats
How do I keep conversation history?Short-term memoryCheckpointers, thread IDs, custom state
How do I connect an MCP server?Model Context ProtocolTransports, authentication, tool results (beta)
How do I upgrade old code?LangChain v1 migration guideMoving to create_agent, hooks to middleware, namespace moves
What shipped, and when?Changelog, release policyRelease notes, support windows, deprecation rules
How do I test without API calls?Unit testingFake chat models and in-memory checkpointers

If you work from an editor, the docs also run a public MCP server at https://docs.langchain.com/mcp that needs no API key, so a coding agent can search them directly.

Which entry point should you use: create_agent, Deep Agents, or LangGraph?

LangChain's frameworks, runtimes, and harnesses page separates three layers: LangChain is the framework, LangGraph the runtime, and Deep Agents the harness. The LangChain overview recommends Deep Agents for a batteries-included start, create_agent for a harness you configure yourself, and LangGraph for workflows that mix deterministic and agentic steps.

Entry pointImportReach for it whenStatus, September 2026
create_agentfrom langchain.agents import create_agentYou want the standard tool loop and will add capabilities through middlewareStable; the 1.x line has long-term support
create_deep_agentfrom deepagents import create_deep_agentLong, multi-step tasks that need planning, files, subagents, and summarization built inPre-1.0 (0.7.19); minor releases can break; needs Python 3.11 or later
StateGraphfrom langgraph.graph import StateGraphCustom topology: routing, fan-out, or deterministic steps around agent callsStable; the 1.x line has long-term support
AgentExecutor, initialize_agentfrom langchain_classic.agents import ...Only while maintaining code you cannot migrate yetLegacy namespace; initialize_agent and AgentType are deprecated, with removal targeted for 2.0.0

The layers nest rather than compete. create_agent returns a compiled LangGraph graph, and the middleware overview notes that hooks run inside that graph, so you can drop a whole agent into a larger StateGraph as a node and keep its retries and approvals. Deep Agents builds on LangChain's agent building blocks and the LangGraph runtime (overview, PyPI), and as a pre-1.0 package it churns: 0.7.0 (July 2026) stopped including to-do planning by default. If you are still deciding whether the task needs an agent loop at all, our guide to AI agent architecture patterns covers when a fixed workflow fits better.

What changed between LangChain 0.3 and 1.4?

LangChain 1.0 shipped in October 2025 and reorganized the package around agents. The v1 migration guide and the langchain-classic source account for most of what breaks old tutorials.

Old patternStatus in langchain 1.4.3Use instead
from langchain.agents import initialize_agent, AgentTypeImport fails; both moved to langchain_classic.agents, deprecated since 0.1.0create_agent
AgentExecutor(..., max_iterations=15)Lives in langchain_classic.agents; 15 is its default capCall-limit middleware (see below)
langgraph.prebuilt.create_react_agentDeprecated in LangGraph v1create_agent, with prompt= renamed system_prompt=
pre_model_hook, post_model_hookReplacedMiddleware with before_model and after_model
tools=ToolNode([...], handle_tool_errors=...)A ToolNode is no longer accepted as toolswrap_tool_call middleware or ToolErrorMiddleware
response_format=("prompt", Schema)Prompted output removedToolStrategy or ProviderStrategy
Per-run data in config["configurable"]Superseded for static context; thread IDs stay in configcontext= on invoke, context_schema= on the agent
langchain.chains, langchain.retrievers, langchain.hubMovedThe same modules under langchain_classic
MultiServerMCPClient from langchain-mcp-adaptersReplaced in 1.4.0MCPAdapter from langchain.mcp (beta)

Three smaller changes bite in code review. Custom state must be a TypedDict that extends AgentState, since Pydantic models and dataclasses are no longer supported. Code that filters streamed events by node name must match "model" instead of "agent". And message text is now a property, message.text; calling it as a method is deprecated in langchain-core 1.0.

The minor releases since 1.0, from the changelog:

  • 1.1 (November 2025): model profiles, ModelRetryMiddleware, and SystemMessage support for the system prompt.
  • 1.2 (December 2025): provider-specific tool options through a tool extras attribute, and strict schema adherence for structured output.
  • 1.3 and LangGraph 1.2 (May 2026): version="v3" event streaming, plus LangGraph per-node timeouts and node-level error handlers.
  • 1.4 (September 2026): MCP support inside LangChain as the beta langchain.mcp namespace, built on FastMCP. Patch 1.4.3 shipped on September 28, 2026 and requires langgraph 1.2.11 or later, below 1.3.0 (PyPI).

Support windows matter if you still run 0.3 code. Per the release policy, LangChain 0.3 and LangGraph 0.4 are in maintenance mode until December 2026, receiving security patches and critical fixes only. LangChain 1.0 and LangGraph 1.0 are long-term support lines: active until 2.0 ships, then in maintenance for at least a year, and deprecated features keep working throughout 1.x. Patch releases can land up to a few times a week, so pin exact versions and rerun your tests before each upgrade.

How do tools work in create_agent?

The Tools page covers the three shapes create_agent accepts in tools=: plain Python callables with type hints and a docstring, functions decorated with @tool, and dicts describing a provider's built-in tools. Type hints are required because they become the input schema, and the docstring becomes the description the model reads when deciding whether to call the tool. The docs recommend snake_case names for compatibility across providers and reserve two argument names, config and runtime; a tool that needs agent state or per-run context takes a ToolRuntime parameter instead. A tool can return a string, an object for the model to inspect, or a Command that writes to state. For long tool lists, the same page covers dynamic tool selection.

create_agent runs tools through LangGraph's ToolNode, whose default handler turns invalid-argument errors into a message the model can correct but re-raises any exception thrown inside your tool (source). A timeout in your HTTP call therefore ends the whole run unless middleware catches it. The example further down handles that explicitly.

Tools from MCP servers

Install langchain[mcp] (1.4.0 or later), open an MCPAdapter on a server, and pass the result of list_tools() to create_agent. The namespace is in beta and raises a LangChainBetaWarning on import. You.com runs a hosted MCP server whose keyless free profile exposes you-search and you-discover at 100 queries per day, according to the MCP server docs. Adding tools=you-search narrows it to one tool, because enabled tools are the intersection of the profile and the allowlist.

import asyncio

from langchain.agents import create_agent
from langchain.mcp import MCPAdapter

# Keyless free profile, narrowed to one tool: enabled tools are the
# intersection of the profile ceiling and the ?tools= allowlist.
YOU_MCP_URL = "https://api.you.com/mcp?profile=free&tools=you-search"


async def main() -> None:
    async with MCPAdapter(YOU_MCP_URL) as adapter:
        tools = await adapter.list_tools()
        agent = create_agent("openai:gpt-5.5", tools)
        result = await agent.ainvoke(
            {"messages": [{"role": "user", "content": "Summarize this week's LangChain releases."}]}
        )
        print(result["messages"][-1].text)


if __name__ == "__main__":
    asyncio.run(main())

Adapted tools keep their server names, so the model sees you-search. For keyed access to the server's four default tools, pass a FastMCP client that sends your key as a bearer token, MCPAdapter(Client("https://api.you.com/mcp", auth=os.environ["YDC_API_KEY"])), the pattern on LangChain's MCP authentication page. If MCP is new to you, start with what the Model Context Protocol is.

How does middleware change the agent loop?

Middleware is the customization layer in LangChain 1.x. The v1 release notes list six hooks: before_agent, before_model, wrap_model_call, wrap_tool_call, after_model, and after_agent. You can subclass AgentMiddleware or use the matching decorators, and the prebuilt catalog covers the common cases. These are the ones to know first:

MiddlewareWhat it doesDetail that matters
ToolErrorMiddlewareTurns selected tool exceptions into error messages the model can seeNeeds langchain 1.3.14 or later; a handler that returns nothing lets the exception halt the run
ToolRetryMiddlewareRetries failed tool calls with exponential backoffExceptions that do not match its retry filter are re-raised immediately
ModelCallLimitMiddlewareCaps model calls per run or per threadThread limits need a checkpointer; the default exit ends the run gracefully
ToolCallLimitMiddlewareCaps tool calls per run or per thread, globally or for one named toolBy default, extra calls are blocked with an error message and the model decides how to finish
SummarizationMiddlewareSummarizes history as it nears the context limitTrigger points can come from model profiles
HumanInTheLoopMiddlewarePauses for approval before selected tool callsMatches on each tool's name

Order in the middleware list is behavior, not style. Earlier entries wrap later ones, so the docs put ToolErrorMiddleware before ToolRetryMiddleware and set on_failure="error" on the retry layer. Transient failures get retried first, and only an exhausted retry reaches the error handler.

One default changed quietly between generations. The legacy AgentExecutor stopped after 15 iterations unless told otherwise (source). create_agent compiles its graph with a recursion limit of 9,999 steps (source), so no small built-in cap stops a looping agent. Put the cap in middleware: ModelCallLimitMiddleware bounds model calls per invocation, and ToolCallLimitMiddleware bounds calls to one expensive tool.

A working example: a search agent built from the documented pieces

The official pages show each piece with stub tools; here they are assembled around a real one. The tool calls the You.com Web Search API over REST: a POST to https://ydc-index.io/v1/search with an X-API-Key header, as documented in the Web Search API guide. It requests highlights, the query-relevant passages the guide recommends for grounding an agent. You need Python 3.10 or later, a You.com API key, and an OpenAI key for the model string LangChain's own examples use.

python -m pip install "langchain==1.4.3" "langchain-openai==1.6.6"
export YDC_API_KEY="your-you.com-api-key"
export OPENAI_API_KEY="your-openai-api-key"
import json
import os
import urllib.error
import urllib.request
from typing import Optional

from langchain.agents import create_agent
from langchain.agents.middleware import (
    ModelCallLimitMiddleware,
    ToolCallLimitMiddleware,
    ToolCallRequest,
    ToolErrorMiddleware,
    ToolRetryMiddleware,
)
from langchain.tools import tool
from langgraph.checkpoint.memory import InMemorySaver

SEARCH_URL = "https://ydc-index.io/v1/search"
TRANSIENT_STATUS = {429, 500, 502, 503, 504}


class TransientSearchError(Exception):
    """Rate limits, server errors, and network failures: safe to retry."""


class PermanentSearchError(Exception):
    """Missing key, bad key, no credits, missing scope, bad parameters: fix config."""


def search_you(query: str, count: int = 5, timeout: float = 15.0) -> dict:
    key = os.environ.get("YDC_API_KEY", "").strip()
    if not key:
        raise PermanentSearchError("YDC_API_KEY is not set")
    body = json.dumps({
        "query": query,
        "count": count,
        "extraction": {"extraction_mode": "highlights"},
    }).encode("utf-8")
    request = urllib.request.Request(
        SEARCH_URL,
        data=body,
        method="POST",
        headers={"X-API-Key": key, "Content-Type": "application/json"},
    )
    try:
        with urllib.request.urlopen(request, timeout=timeout) as response:
            return json.load(response)
    except urllib.error.HTTPError as err:
        if err.code in TRANSIENT_STATUS:
            raise TransientSearchError(f"HTTP {err.code}") from err
        raise PermanentSearchError(f"HTTP {err.code}") from err
    except (urllib.error.URLError, TimeoutError) as err:
        raise TransientSearchError(type(err).__name__) from err


def format_results(payload: dict, max_passages: int = 2) -> str:
    web = (payload.get("results") or {}).get("web") or []
    if not web:
        return "NO_RESULTS: rephrase once, or answer without search and say so."
    blocks = []
    for hit in web:
        contents = hit.get("contents") or {}
        passages = (contents.get("highlights") or hit.get("snippets")
                    or [hit.get("description") or ""])
        text = " ".join(p for p in passages[:max_passages] if p)
        blocks.append(f"{hit.get('title') or 'Untitled'}\n{hit.get('url', '')}\n{text}")
    return "\n\n".join(blocks)


@tool
def web_search(query: str) -> str:
    """Search the live web for facts that may have changed after training.
    Returns titles, URLs, and query-relevant passages. Cite the URLs you use."""
    return format_results(search_you(query))


def on_search_error(exc: Exception, request: ToolCallRequest) -> Optional[str]:
    if isinstance(exc, TransientSearchError):
        return ("web_search is temporarily unavailable. Answer from results you "
                "already have and say which claims you could not verify.")
    return None  # PermanentSearchError and anything unexpected halt the run


def build_agent():
    return create_agent(
        model="openai:gpt-5.5",
        tools=[web_search],
        system_prompt=(
            "Search before answering anything time-sensitive, at most three "
            "times per question, and cite the URL behind every searched claim."
        ),
        middleware=[
            ModelCallLimitMiddleware(run_limit=6),
            ToolCallLimitMiddleware(tool_name="web_search", run_limit=3),
            ToolErrorMiddleware(on_search_error, tools=["web_search"]),
            ToolRetryMiddleware(
                max_retries=2,
                retry_on=(TransientSearchError,),
                on_failure="error",
                tools=["web_search"],
            ),
        ],
        checkpointer=InMemorySaver(),
    )


if __name__ == "__main__":
    agent = build_agent()
    config = {"configurable": {"thread_id": "docs-map-demo"}}
    question = "What changed in the most recent LangChain 1.4 patch release?"
    result = agent.invoke(
        {"messages": [{"role": "user", "content": question}]},
        config=config,
    )
    print(result["messages"][-1].text)

What each piece does:

  • The error classes follow documented status codes. You.com answers 429 when you exceed the rate limit and asks clients to back off and retry; 401, 402, 403, and 422 mean a bad key, no credits, a missing scope, or invalid parameters (error reference). Retrying those wastes calls, so they are permanent.
  • Retry runs inside error handling. Transient failures get two retries with exponential backoff starting at one second; if they still fail, on_search_error tells the model search is down so it can answer with a caveat. The middleware does not read the Retry-After header the rate limit docs ask clients to honor, so raise initial_delay if 429s persist. Permanent errors halt the run, so monitoring sees a configuration problem instead of a quietly degraded answer.
  • Two limits do the job max_iterations used to do: at most three searches and six model calls per question. A fourth search request is blocked with an error message, and the model decides how to finish.
  • The tool reads the web results only, keeps every URL, and caps passages at two per result so search output does not crowd the context window.

The cost is easy to bound. The Web Search API costs $5.00 per 1,000 calls as of September 2026, and highlights are included in that base price, per the billing page. Each search costs half a cent, so the three-search cap holds search spend to $0.015 per question before model tokens. New accounts start with $100 in free credits. Self-serve keys default to 10 requests per second on this endpoint (rate limits); concurrent users can hit that, which is what the 429 retry path handles.

If you would rather not write the tool, LangChain's own tools index lists You.com Search and links to You.com's LangChain integration page for the langchain-youdotcom package. Its YouSearchTool registers under the name you_search, so change tool_name in the limit middleware if you swap it in. Our LangChain web search tool guide covers that package's tools, retriever, and parameters.

What should you test before shipping a LangChain agent?

Test the tool first, without a network. Fixtures shaped like the documented response make parsing and error classification deterministic. Save the example as search_agent.py; this test file adds nothing beyond the standard library:

import io
import unittest
import urllib.error
from unittest import mock

import search_agent as sa

FIXTURE = {
    "results": {"web": [{
        "url": "https://example.com/release-notes",
        "title": "Release notes",
        "description": "Fallback text",
        "contents": {"highlights": ["First passage.", "Second passage.", "Third."]},
    }]},
    "metadata": {"query": "q", "search_uuid": "0000", "latency": 0.3},
}


def http_error(code):
    return urllib.error.HTTPError(sa.SEARCH_URL, code, "error", {}, io.BytesIO(b"{}"))


class SearchToolTests(unittest.TestCase):
    def test_keeps_urls_and_caps_passages(self):
        text = sa.format_results(FIXTURE)
        self.assertIn("https://example.com/release-notes", text)
        self.assertIn("First passage. Second passage.", text)
        self.assertNotIn("Third.", text)

    def test_empty_results_return_sentinel(self):
        self.assertTrue(sa.format_results({"results": {"web": []}}).startswith("NO_RESULTS"))

    def test_rate_limit_is_transient(self):
        with mock.patch.dict("os.environ", {"YDC_API_KEY": "test"}), \
                mock.patch("urllib.request.urlopen", side_effect=http_error(429)):
            with self.assertRaises(sa.TransientSearchError):
                sa.search_you("q")

    def test_bad_key_is_permanent(self):
        with mock.patch.dict("os.environ", {"YDC_API_KEY": "test"}), \
                mock.patch("urllib.request.urlopen", side_effect=http_error(401)):
            with self.assertRaises(sa.PermanentSearchError):
                sa.search_you("q")


if __name__ == "__main__":
    unittest.main()

Then test the agent with a scripted model. The unit testing page pairs GenericFakeChatModel with an InMemorySaver checkpointer to replay exact responses, tool calls included, with no API key. One catch: create_agent calls the model's bind_tools whenever tools are registered, and GenericFakeChatModel inherits the base implementation, which raises NotImplementedError (source). For an agent with tools, subclass the fake and have bind_tools return self. Then check that:

  • a scripted fourth search is blocked and the run still ends with an answer;
  • a persistent 429 produces the "temporarily unavailable" tool message after the retries, not a crash;
  • a 401 halts the run and surfaces the permanent error in your logs;
  • CI promotes LangChainDeprecationWarning to an error, so calls to deprecated APIs such as initialize_agent fail loudly, and handles LangChainBetaWarning deliberately if you use langchain.mcp.

Search results are untrusted input. A page can carry instructions aimed at your agent, so keep tool output clearly separated from your system prompt; our guide to prompt injection and intent hijacking covers the defenses. To score the finished agent on real tasks, see how to run an AI agent evaluation.

Related Guides

    Share Article:

  1. LI Test

  2. LI Test

Related resources.

No items found.
No items found.