Modern AI Engineering

Lesson 12.5 · 27 min

LangChain: Composable Components for LLM Applications

Every LLM app needs the same plumbing (prompts, model calls, parsing, memory, retrieval, tools), so why write it from scratch every time?

In short: LangChain is an open-source framework (Python and JavaScript) that gives standard building blocks for LLM apps: models, prompt templates, output parsers, retrievers, memory helpers and tools, all sharing one interface so they can be snapped together into chains with the | operator. Under the hood a chain is just function composition: each component's output becomes the next one's input. We will see real LangChain code and then rebuild the core idea in 30 lines of plain Python.

What is LangChain?

LangChain is an open-source framework for building applications on top of large language models. It was started by Harrison Chase in late 2022 and is available for Python and JavaScript/TypeScript. It does not provide a model of its own. Instead it provides standard components for everything around the model, and integrations with hundreds of model providers, vector databases and tools.

The project is split into packages. langchain-core holds the base interfaces (prompts, messages, runnables, parsers). Integration packages such as langchain-openai or langchain-anthropic connect specific providers. The langchain package adds higher-level pieces such as agents. Two sibling projects come from the same company: LangGraph for graph-based agents (next lesson) and LangSmith for tracing and evaluation. With LangChain 1.0 (released in late 2025) its agent API runs on top of LangGraph.

Think of LEGO bricks LEGO bricks come in many shapes, but every brick has the same studs, so any brick connects to any other. LangChain components are like that: a prompt, a model and a parser look different inside, but they all expose the same “studs” (an invoke method), so we can click them together in any order that makes sense.

Running example: a shop assistant that answers product questions using our product docs and returns structured data our website can display.

Why do we need LangChain?

Calling one model once is easy with the provider's own SDK. Real apps quickly need more, and that “more” is the same in almost every app:

  • Prompt assembly: filling variables into templates with system and user messages.
  • Provider differences: OpenAI, Anthropic, Google and local models all have slightly different APIs and message formats.
  • Parsing: turning free text into JSON or typed objects the rest of the code can use.
  • Memory: carrying conversation history between calls.
  • Retrieval: finding relevant documents and inserting them into the prompt (RAG).
  • Tools: describing functions to the model and running the ones it asks for.

Without a framework each team writes this glue again, slightly differently. LangChain's pitch is: use tested components with one common interface, swap a provider by changing one line, and get streaming, batching and tracing for free.

The core idea: everything is a Runnable

The heart of modern LangChain is one interface called a Runnable. Anything that is a Runnable has the same methods: invoke(input) for one input, batch([inputs]) for many, and stream(input) to get output piece by piece (plus async versions). Prompts, models, parsers and retrievers are all Runnables.

Because they share an interface, two Runnables can be joined with the | (pipe) operator. a | b creates a new Runnable that runs a and passes its output to b. This syntax is called the LangChain Expression Language (LCEL). In maths terms, a chain is function composition: chain(x) = parser(model(prompt(x))).

LLM and prompt template

A chat model component wraps a provider's API. We create it once with settings (model name, temperature) and call invoke with messages. Because all chat models share the interface, swapping ChatOpenAI for ChatAnthropic is a one-line change.

A prompt template is a reusable prompt with blanks. ChatPromptTemplate.from_messages takes a list of (role, text) pairs where the text may contain {variables}. Invoking the template with a dictionary fills the blanks and returns ready-to-send messages. Templates keep prompts out of string concatenation code and make them easy to version and test.

real LangChain: prompt + model (needs langchain-openai and an API key)

from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
prompt = ChatPromptTemplate.from_messages([
("system", "You are a helpful shop assistant. Answer only from the context."),
("human", "Context: {context}\n\nQuestion: {question}"),
])
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
messages = prompt.invoke({"context": "The Steelbrew kettle holds 1.7 L.",
"question": "How much water does the kettle hold?"})
reply = llm.invoke(messages)   # an AIMessage
print(reply.content)

What is a chain? And output parsers

A chain is a sequence of components where each output feeds the next input. With LCEL we write chain = prompt | llm | parser and then call chain.invoke(...) once. Chains are themselves Runnables, so a chain can be a piece of a bigger chain.

An output parser is the last link that turns the model's message into something our code can use. StrOutputParser returns plain text. JsonOutputParser extracts JSON. For typed results, many chat models support llm.with_structured_output(MySchema), which uses the provider's structured-output or tool-calling features so the model returns data matching a Pydantic schema.

Pause and think: In the mini version, what would (prompt | FakeLLM()).invoke({...}) return compared with the full chain?

A raw string like Sure! {"name": ...} instead of a dict, because the output parser link is missing. Each link changes the type: dict → prompt string → model text → parsed dict.

Memory

Models are stateless: each call knows only what is in its prompt. Memory in LangChain means storing past messages and adding them to the next prompt. A prompt template can include a MessagesPlaceholder where the history goes. RunnableWithMessageHistory wraps a chain so that, for each session id, it loads the history before the call and saves the new messages after.

Older LangChain versions had many Memory classes (such as ConversationBufferMemory); these are now legacy. For agents, current LangChain relies on LangGraph's checkpointers, which save the whole conversation state per thread. Whatever the API, the concept is the same: memory is the harness re-sending relevant history, often trimmed or summarised so it fits the context window.

Memory is not free Every remembered message is re-sent and paid for on every call. Long chats need trimming (keep the last N messages) or summarisation, or costs and latency climb and early instructions get lost.

Retrieval and RAG

Retrieval-augmented generation (RAG) means finding relevant documents and putting them in the prompt so the model answers from our data. LangChain provides the pieces as components: document loaders (read PDFs, web pages), text splitters (cut documents into chunks), embedding models (turn text into vectors), vector stores (store and search vectors) and retrievers (anything that takes a query and returns documents).

real LangChain RAG chain (not run here)

from langchain_core.vectorstores import InMemoryVectorStore
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
from langchain_openai import OpenAIEmbeddings
store = InMemoryVectorStore.from_texts(
["The Steelbrew kettle holds 1.7 L.", "Returns are accepted within 30 days."],
embedding=OpenAIEmbeddings())
retriever = store.as_retriever(search_kwargs={"k": 1})
format_docs = lambda docs: "\n".join(d.page_content for d in docs)
rag_chain = ({"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt | llm | StrOutputParser())     # prompt, llm from earlier
print(rag_chain.invoke("Can I return the kettle after two weeks?"))

Tools and agents

A tool is a Python function the model may ask to call. In LangChain we decorate a function with @tool; its name, docstring and type hints become the description and argument schema sent to the model. When a model supports tool calling, llm.bind_tools([...]) lets it reply with structured tool-call requests instead of text.

An agent is a model in a loop that keeps calling tools until it can answer. In LangChain 1.0 the main entry point is create_agent, which builds this loop (on LangGraph under the hood): the model decides, the tool runs, the result goes back, repeat.

real LangChain agent (not run here)

from langchain.agents import create_agent
from langchain_core.tools import tool
@tool
def get_order_status(order_id: str) -> str:
"""Look up the shipping status of an order by its id."""
return {"42": "shipped", "7": "packing"}.get(order_id, "unknown")
agent = create_agent(model="openai:gpt-4o-mini", tools=[get_order_status],
system_prompt="You are a polite shop support agent.")
result = agent.invoke({"messages": [{"role": "user", "content": "Where is order 42?"}]})
print(result["messages"][-1].content)

Going one level deeper

Most chain bugs are type bugs: one link hands the next link something it did not expect. The mini version from this lesson is small enough to trace by hand, so let us follow one value through it.

The value at each link of retriever | prompt | FakeLLM() | JsonParser() for the input "kettle"
LinkReceivesReturns
retrieverthe string kettlea dict with question and context
promptthat dictone filled-in prompt string
FakeLLMthe prompt stringchatty text with JSON inside
JsonParserthe chatty texta Python dict

Now let us break it on purpose, in our heads.

Three ways to break the mini chain
ChangeWhat happensWhy
Swap the last two links, so the parser comes before the modelThe parser fails: the prompt string has no { to findThe parser was built for model text and received a prompt
The retriever returns only the question keyThe template fails on the missing contextThe template has a blank that nothing filled
Drop the parserNo error at all. The caller gets a string and fails later, when it reads eta_daysThe chain ran fine; the type at its end is wrong

Finding the broken link

  1. Run the first link alone: Invoke the retriever by itself and look at the value and its type.
  2. Add one link at a time: A chain is itself a Runnable, so (retriever | prompt) can be invoked too. Print what comes out.
  3. Stop at the first surprise: The first link whose output is not what the next link expects is the bug. Everything after it is only a symptom.

The third row of the table is the dangerous one. A failure that raises an error is found at once. A wrong type that flows on quietly is found by a user.

Practice: try it yourself

We will build memory the way this lesson describes it: a wrapper around a chain that loads the stored history for a session before the call and saves the new messages after it. The chain is plain function composition, and the model is a scripted stand-in.

practice_session_memory.py

# Memory as a wrapper: load history before the call, save new messages after.
def pipe(*steps):
# Function composition: each step's output is the next step's input
def chain(x):
for step in steps:
x = step(x)
return x
return chain
def prompt(variables):
lines = ["system: You are a shop assistant."] + variables["history"]
return lines + ["human: " + variables["question"]]
def fake_llm(messages):            # stands in for a chat model call
seen = " ".join(messages).lower()
product = "kettle" if "kettle" in seen else "unknown product"
return f"ai: ({len(messages)} messages seen) You mean the {product}."
chain = pipe(prompt, fake_llm)
store = {}                         # session id -> list of past messages
def with_history(chain, keep_last=4):
def invoke(question, session_id):
history = store.setdefault(session_id, [])
reply = chain({"history": history[-keep_last:], "question": question})
history += ["human: " + question, reply]      # save after the call
return reply
return invoke
chat = with_history(chain)
print(chat("Do you sell a kettle?", "alice"))
print(chat("How much water does it hold?", "alice"))
print(chat("How much water does it hold?", "bob"))
print("stored:", {sid: len(msgs) for sid, msgs in store.items()})

Output:

ai: (2 messages seen) You mean the kettle.
ai: (4 messages seen) You mean the kettle.
ai: (2 messages seen) You mean the unknown product.
stored: {'alice': 4, 'bob': 2}

Now change it:

  • Set keep_last=1. Predict how many messages the model sees on Alice's second question, and whether it still names the product.
  • Add two more questions from Alice. Predict the “messages seen” number for each. Where does it stop growing, and why?
  • Give Bob the session id "alice". Predict his reply. Why would this be a serious bug in a real shop?

Pause and think: Bob asked the same question as Alice and got “unknown product”. The model function is the same for both. What is different?

The input. Memory is stored per session id, and Bob's session has no earlier messages, so his prompt held only the system line and his question. The model has no state of its own. It seemed to remember the kettle for Alice only because the wrapper sent her earlier messages again. Change what the wrapper sends, and we change what the model appears to know.

Pause and think: with_history saves the new messages after the chain returns. Suppose the chain raises an error halfway. What ends up in the store, and is that the behaviour we want?

Nothing new is saved, because the line that extends the history is never reached. That is usually what we want: a question with no answer would leave the history unbalanced and could confuse the next call. The price is that the failed question is forgotten, so the app should show the error and let the user ask again.

A complete flow, and how LangChain compares

What happens when our shop assistant answers one question

  1. Input: The user asks “Is the kettle in stock and when would it arrive?”. The chain receives the question (and a session id if memory is used).
  2. Retrieve: The retriever embeds the question and fetches the top product-doc chunks.
  3. Fill the prompt: The template combines system rules, history, context and the question into messages.
  4. Call the model: The chat model component sends the messages to the provider and gets an AI message back (possibly with tool calls if tools are bound).
  5. Parse: The output parser or structured-output wrapper turns the reply into a typed object.
  6. Return and trace: The website gets the object; if LangSmith tracing is enabled, every step's inputs, outputs and timing are recorded.

Common pitfalls Copying old tutorials: many use classes that were deprecated (old chain and memory classes). Check the version you install. Hiding prompts: always print or trace the final prompt the model saw; most bugs are visible there. Over-abstracting: for a single model call, a framework may add more complexity than it removes.

Key takeaways

  • LangChain provides standard components around LLMs: models, prompts, parsers, retrievers, memory helpers and tools.
  • Everything is a Runnable with invoke, batch and stream, so components snap together with |.
  • A chain is function composition: each link's output type must match the next link's input.
  • Output parsers and structured output turn model text into data your code can trust.
  • RAG, memory and agents are just more components in the pipeline; agents now run on LangGraph.
  • For complex control flow use LangGraph; for one call, the provider SDK may be enough.

Key terms

  • LangChain: An open-source framework of standard, composable components for building LLM applications.
  • Runnable: LangChain's common interface with invoke, batch and stream; prompts, models, parsers and chains are all Runnables.
  • LCEL: LangChain Expression Language: composing Runnables with the | operator.
  • Prompt template: A reusable prompt with named blanks that are filled from a dictionary at run time.
  • Output parser: A component that turns a model's reply into a string, JSON or typed object.
  • Retriever: A component that takes a query and returns relevant documents, often from a vector store.

← 12.4 Defining Done: Why Exit Criteria Shape Agent Quality · 12.6 LangGraph: Graph-Based Agent Orchestration Explained →