← Back to Agentic AI map
Lesson 2.5 · Building Agents with LangChain

Output Parsers

Turn the model’s text into a Recipe object your code can use - with PydanticOutputParser at the end of an LCEL chain.

parsers

What you will be able to do

  • Explain why an application needs data, not prose
  • Describe the expected structure with a Pydantic class
  • Create a PydanticOutputParser and put its format instructions in the prompt
  • Build prompt | llm | parser and use the Recipe object it returns
  • Tell parsing apart from validation
  • Remember that a valid structure is not the same as a true answer

The idea, in plain English

A model returns text. "Chocolate cake requires flour, sugar, eggs, butter and cocoa powder" is fine for a person, but your application wants a name and a list it can loop over. An output parser converts the model’s response into a format your program can use.

In this lesson the parser is PydanticOutputParser. You describe the result as a Pydantic class - class Recipe(BaseModel): name: str; ingredients: list[str] - and the parser does two jobs: it writes format instructions for the prompt, and it turns the model’s reply into a validated Recipe object. Put it last in the chain from Lesson 2.4: prompt | llm | parser.

This is Module 1’s structured output again - schema, parsing, validation - with less manual code. And the same limit applies: a schema checks the shape of the answer, not whether it is true.

Every output here comes from real runs against llama3 (the code says llama3.1).

Worked example: Parse a recipe into an ingredients list.

request flowprompt | llm | parserstep 1 / 4

1 - Fill the prompt

The template fills {dish} and {format_instructions}. The instructions come from parser.get_format_instructions() and contain the JSON schema of Recipe.

dish
vegetable pasta
format_instructions
Recipe schema
input_variables
dish, format_instructions
output
messages

One call to chain.invoke(), from the dish name to a Recipe object.

The problem: text is not data

Ask for ingredients in plain text and you have to extract them yourself. text.split(",") works for "pasta, tomato, onion and garlic" - until the model answers with a numbered list instead, or adds a sentence before it. Every new format needs new extraction code.

With a schema you say what you want once, and the same code reads every answer. The application gets predictable data instead of free-form text.

Pydantic describes the structure

Pydantic lets you define the expected structure as a Python class. Recipe has a name, which is a string, and ingredients, which is a list of strings. Recipe(name="Pasta", ingredients=["pasta", "tomato", "garlic"]) is valid; Recipe(name="Pasta", ingredients="tomato") is not.

PydanticOutputParser(pydantic_object=Recipe) connects that class to LangChain. From then on the parser knows the expected structure - and can explain it to the model.

Format instructions

parser.get_format_instructions() returns text: a sentence saying the output must be a JSON instance of a schema, a small example, and the Recipe schema itself. The model needs this to know what to produce.

Put a {format_instructions} placeholder in the template and pass the instructions as a value, as the example does. Do not paste them into the template text: the schema is full of braces, and Lesson 2.3 showed what braces in a template mean.

What the model actually returned

The model did not return bare JSON. It wrote "Here is the recipe for vegetable pasta in the format you requested:", then the JSON in a markdown code fence, then a paragraph explaining the schema. The parser extracted the JSON anyway - that is part of its job.

The content was also not the tidy list from the lesson’s sketch. Ingredients came back with quantities and preparation: "8 oz pasta of your choice", "2 cloves garlic, minced". The schema says ingredients is a list of strings; it says nothing about what the strings contain. In eight runs every reply parsed, and every list had quantities.

Parsing vs validation

Parsing asks: can I convert this text into data? Validation asks: does that data match what my application expects? The parser does both, in that order.

Valid JSON is not valid application data. {"name": "Priya", "age": null}, {"name": "Priya"} and {"name": "Priya", "age": "twenty-nine"} all parse. Against a class with age: int, Pydantic rejects all three.

class Person(BaseModel): name: str, age: int
{"name": "Priya", "age": 29}Valid JSON, valid data.
{"name": "Priya", "age": null}Valid JSON. Rejected - age must be an integer.
{"name": "Priya"}Valid JSON. Rejected - age is missing.
{"name": "Priya", "age": "twenty-nine"}Valid JSON. Rejected - not an integer.

Tip: Pydantic is not maximally strict by default: it converts "29" to 29 and silently drops fields the class does not define.

A schema does not make the answer true

Ask for a Person(name: str, age: int) from "Priya lives in Pune." - a sentence with no age. A required integer field puts pressure on the model to invent one.

Through the parser, llama3 wrote "age": null and the parse failed: honest model, unusable schema. When the shape was enforced by Ollama instead (Lesson 1.3’s format option), the same request returned age=-1 - valid, and meaningless. With age: int | None the parser returned age=None, the honest answer. Allow null for facts that may be missing.

Watch out: Constraining the format does not make the data true. Validation checks shape and types, never facts.

Module 1 and Module 2

In Module 1 you did each step by hand: read response["message"]["content"], call json.loads(), then check the fields yourself. The parser wraps those steps into one reusable component driven by one class.

The same job, two ways
Describe the shapeModule 1: an example in the prompt. Module 2: a Pydantic class.
Tell the modelModule 1: hand-written. Module 2: get_format_instructions().
ParseModule 1: json.loads(). Module 2: the parser.
ValidateModule 1: your own checks. Module 2: Pydantic, from the class.
UseModule 1: data["name"]. Module 2: result.name.

Where this is used

The pattern is always the same: unstructured text in, a model, structured output, then your application. Only the class changes.

One pattern, different classes
Resume parsingname, skills: list[str], experience_years: int
Support ticketscategory, priority, requires_human: bool
Course taggingtopic, category, difficulty
Recipe extractionname, ingredients: list[str]

Step-by-step code

Define the structure
from pydantic import BaseModel class Recipe(BaseModel): name: str ingredients: list[str] print(Recipe(name="Pasta", ingredients=["pasta", "tomato", "garlic"])) # name='Pasta' ingredients=['pasta', 'tomato', 'garlic']
The complete example
from pydantic import BaseModel from langchain_ollama import ChatOllama from langchain_core.prompts import ChatPromptTemplate from langchain_core.output_parsers import PydanticOutputParser class Recipe(BaseModel): name: str ingredients: list[str] llm = ChatOllama(model="llama3.1") parser = PydanticOutputParser(pydantic_object=Recipe) prompt = ChatPromptTemplate.from_template( """ Create a recipe for {dish}. {format_instructions} """ ) chain = prompt | llm | parser result = chain.invoke({ "dish": "vegetable pasta", "format_instructions": parser.get_format_instructions(), }) print(result)
Output - from a real run
name='Vegetable Pasta' ingredients=['8 oz pasta of your choice (e.g. spaghetti, linguine, or fettuccine)', '1 red bell pepper, sliced', '1 yellow bell pepper, sliced', '1 small onion, thinly sliced', '2 cloves garlic, minced', '1 cup cherry tomatoes, halved', '1 cup broccoli florets', '1 cup sliced mushrooms (button or cremini)', '1 tsp olive oil', 'Salt and pepper to taste', 'Grated Parmesan cheese (optional)'] Eight runs: every one parsed. Names varied ('Vegetable Pasta', 'Vegetable Pasta Recipe'), 7 to 11 ingredients, always with quantities.
What the parser received - the raw reply
Here is the recipe for vegetable pasta in the format you requested: ``` { "name": "Vegetable Pasta", "ingredients": [ "8 oz pasta of your choice", "2 cups mixed vegetables (such as bell peppers, zucchini, cherry tomatoes, and onions)", "2 cloves garlic, minced", ... ] } ``` This JSON instance conforms to the provided schema. ...
Use the data
print(result.name) # Vegetable Pasta for ingredient in result.ingredients: print(ingredient) # 8 oz pasta of your choice (e.g. spaghetti, linguine, or fettuccine) # 1 red bell pepper, sliced # ...
Validation - a string where a list belongs
from langchain_core.exceptions import OutputParserException try: parser.parse('{"name": "Pasta", "ingredients": "tomato"}') except OutputParserException as error: print(error) # Failed to parse Recipe from completion {"name": "Pasta", "ingredients": "tomato"}. # Got: 1 validation error for Recipe # ...
Missing information: int vs int | None
class Person(BaseModel): name: str age: int class PersonMaybeAge(BaseModel): name: str age: int | None # Same chain shape: prompt | llm | parser, for "Priya lives in Pune." # Person: OutputParserException - the model wrote "age": null, which is not an int # PersonMaybeAge: name='Priya' age=None
Module 1 vs Module 2
# Module 1 - by hand raw = response["message"]["content"] # a string raw["name"] # TypeError: string indices must be integers, not 'str' data = json.loads(raw) # parse data["name"] # ...then validate it yourself # Module 2 - the parser does both result = (prompt | llm | parser).invoke(values) result.name

Tip: Catch OutputParserException where you call the chain. Its message includes the model’s raw output - exactly what you need to see when parsing fails.

Output parsers at a glance

BaseModel

Define the structure: fields and types.

class Recipe(BaseModel): name: str
PydanticOutputParser

Format instructions before, parse and validate after.

PydanticOutputParser(pydantic_object=Recipe)
get_format_instructions

The schema as prompt text.

"format_instructions": parser.get_format_instructions()
The chain

The parser goes last.

prompt | llm | parser
OutputParserException

Raised when parsing or validation fails.

from langchain_core.exceptions import OutputParserException
int | None

Let the model say "not given".

age: int | None

Try it yourself

The code does not change. Swap the content string and the program does something else entirely.

Another dish

“Run the chain for "chocolate cake" and loop over result.ingredients.”

Add a field

“Add cooking_time: int to Recipe and run it again.”

Break it

“Call parser.parse() on JSON with a missing field and on prose with no JSON.”

Missing facts

“Extract a Person from "Priya lives in Pune." with age: int, then with age: int | None.”

What usually goes wrong

Treating the response content as data

The model’s content is a string. Parse it first - or let the parser do it.

✗ raw = response["message"]["content"]
raw["name"]
✓ result = chain.invoke(values)
result.name
Assuming valid JSON means correct data

{"name": "Priya", "age": "twenty-nine"} is valid JSON. With age: int, it is not valid data.

Forgetting format_instructions

The template has a {format_instructions} placeholder; invoke() must supply it, or the call fails with a missing-variable error.

✗ chain.invoke({"dish": "vegetable pasta"})
✓ chain.invoke({"dish": "vegetable pasta", "format_instructions": parser.get_format_instructions()})
Requiring fields the input cannot fill

Either the parse fails or the model invents a value. Use int | None when the answer may not be there.

✗ age: int
✓ age: int | None

Key points

  • Applications need predictable data, not free-form text.
  • A Pydantic class defines the structure: Recipe has name and ingredients.
  • PydanticOutputParser writes format instructions and turns the reply into a Recipe.
  • chain = prompt | llm | parser - the parser goes last.
  • Parsing: can I read it? Validation: does it fit my class?
  • Valid JSON is not valid application data.
  • Constraining the format does not make the data true.

Quick check before you move on

What problem does an output parser solve?
It converts the model’s text into structured data the application can use.
What is PydanticOutputParser?
A LangChain parser that uses a Pydantic class to describe, parse and validate the expected response.
Why use Pydantic?
To define the expected structure and types, and to validate data against them.
What does parser.get_format_instructions() return?
Text with the JSON schema and an explanation, to put in the prompt.
What is the difference between parsing and validation?
Parsing converts text into data. Validation checks that the data matches what the application expects.
Does structured output guarantee factual correctness?
No. It enforces structure and types, not truth.

Quiz

  1. 1.

    Why not extract ingredients with text.split(",")?

  2. 2.

    The model wrapped its JSON in a sentence and a code fence. What did the parser do?

  3. 3.

    The model returns {"name": "Pasta", "ingredients": "tomato"}. What happens?

  4. 4.

    Our recipes came back with "8 oz pasta" instead of "pasta". Did validation fail? Why not?

Interview questions

What is an output parser?

A component that converts the LLM’s response into a structured format the application can use.

Does structured output guarantee factual correctness?

No. It enforces structure and types. A required field the input cannot support will fail or be filled with something plausible - so allow nulls and verify important values.

Why is structured output useful for applications?

Applications need predictable data - fields they can read, loop over and store - rather than free-form text.

Comments

Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.

Loading comments...