Output Parsers
Turn the model’s text into a Recipe object your code can use - with PydanticOutputParser at the end of an LCEL chain.
What you will be able to do
- Explain why an application needs data, not prose
- Describe the expected structure with a Pydantic class
- Create a PydanticOutputParser and put its format instructions in the prompt
- Build prompt | llm | parser and use the Recipe object it returns
- Tell parsing apart from validation
- Remember that a valid structure is not the same as a true answer
The idea, in plain English
A model returns text. "Chocolate cake requires flour, sugar, eggs, butter and cocoa powder" is fine for a person, but your application wants a name and a list it can loop over. An output parser converts the model’s response into a format your program can use.
In this lesson the parser is PydanticOutputParser. You describe the result as a Pydantic class - class Recipe(BaseModel): name: str; ingredients: list[str] - and the parser does two jobs: it writes format instructions for the prompt, and it turns the model’s reply into a validated Recipe object. Put it last in the chain from Lesson 2.4: prompt | llm | parser.
This is Module 1’s structured output again - schema, parsing, validation - with less manual code. And the same limit applies: a schema checks the shape of the answer, not whether it is true.
Every output here comes from real runs against llama3 (the code says llama3.1).
Worked example: Parse a recipe into an ingredients list.
1 - Fill the prompt
The template fills {dish} and {format_instructions}. The instructions come from parser.get_format_instructions() and contain the JSON schema of Recipe.
One call to chain.invoke(), from the dish name to a Recipe object.
The problem: text is not data
Ask for ingredients in plain text and you have to extract them yourself. text.split(",") works for "pasta, tomato, onion and garlic" - until the model answers with a numbered list instead, or adds a sentence before it. Every new format needs new extraction code.
With a schema you say what you want once, and the same code reads every answer. The application gets predictable data instead of free-form text.
Pydantic describes the structure
Pydantic lets you define the expected structure as a Python class. Recipe has a name, which is a string, and ingredients, which is a list of strings. Recipe(name="Pasta", ingredients=["pasta", "tomato", "garlic"]) is valid; Recipe(name="Pasta", ingredients="tomato") is not.
PydanticOutputParser(pydantic_object=Recipe) connects that class to LangChain. From then on the parser knows the expected structure - and can explain it to the model.
Format instructions
parser.get_format_instructions() returns text: a sentence saying the output must be a JSON instance of a schema, a small example, and the Recipe schema itself. The model needs this to know what to produce.
Put a {format_instructions} placeholder in the template and pass the instructions as a value, as the example does. Do not paste them into the template text: the schema is full of braces, and Lesson 2.3 showed what braces in a template mean.
What the model actually returned
The model did not return bare JSON. It wrote "Here is the recipe for vegetable pasta in the format you requested:", then the JSON in a markdown code fence, then a paragraph explaining the schema. The parser extracted the JSON anyway - that is part of its job.
The content was also not the tidy list from the lesson’s sketch. Ingredients came back with quantities and preparation: "8 oz pasta of your choice", "2 cloves garlic, minced". The schema says ingredients is a list of strings; it says nothing about what the strings contain. In eight runs every reply parsed, and every list had quantities.
Parsing vs validation
Parsing asks: can I convert this text into data? Validation asks: does that data match what my application expects? The parser does both, in that order.
Valid JSON is not valid application data. {"name": "Priya", "age": null}, {"name": "Priya"} and {"name": "Priya", "age": "twenty-nine"} all parse. Against a class with age: int, Pydantic rejects all three.
{"name": "Priya", "age": 29}Valid JSON, valid data.{"name": "Priya", "age": null}Valid JSON. Rejected - age must be an integer.{"name": "Priya"}Valid JSON. Rejected - age is missing.{"name": "Priya", "age": "twenty-nine"}Valid JSON. Rejected - not an integer.Tip: Pydantic is not maximally strict by default: it converts "29" to 29 and silently drops fields the class does not define.
A schema does not make the answer true
Ask for a Person(name: str, age: int) from "Priya lives in Pune." - a sentence with no age. A required integer field puts pressure on the model to invent one.
Through the parser, llama3 wrote "age": null and the parse failed: honest model, unusable schema. When the shape was enforced by Ollama instead (Lesson 1.3’s format option), the same request returned age=-1 - valid, and meaningless. With age: int | None the parser returned age=None, the honest answer. Allow null for facts that may be missing.
Watch out: Constraining the format does not make the data true. Validation checks shape and types, never facts.
Module 1 and Module 2
In Module 1 you did each step by hand: read response["message"]["content"], call json.loads(), then check the fields yourself. The parser wraps those steps into one reusable component driven by one class.
Describe the shapeModule 1: an example in the prompt. Module 2: a Pydantic class.Tell the modelModule 1: hand-written. Module 2: get_format_instructions().ParseModule 1: json.loads(). Module 2: the parser.ValidateModule 1: your own checks. Module 2: Pydantic, from the class.UseModule 1: data["name"]. Module 2: result.name.Where this is used
The pattern is always the same: unstructured text in, a model, structured output, then your application. Only the class changes.
Resume parsingname, skills: list[str], experience_years: intSupport ticketscategory, priority, requires_human: boolCourse taggingtopic, category, difficultyRecipe extractionname, ingredients: list[str]Step-by-step code
from pydantic import BaseModel
class Recipe(BaseModel):
name: str
ingredients: list[str]
print(Recipe(name="Pasta", ingredients=["pasta", "tomato", "garlic"]))
# name='Pasta' ingredients=['pasta', 'tomato', 'garlic']from pydantic import BaseModel
from langchain_ollama import ChatOllama
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import PydanticOutputParser
class Recipe(BaseModel):
name: str
ingredients: list[str]
llm = ChatOllama(model="llama3.1")
parser = PydanticOutputParser(pydantic_object=Recipe)
prompt = ChatPromptTemplate.from_template(
"""
Create a recipe for {dish}.
{format_instructions}
"""
)
chain = prompt | llm | parser
result = chain.invoke({
"dish": "vegetable pasta",
"format_instructions": parser.get_format_instructions(),
})
print(result)name='Vegetable Pasta' ingredients=['8 oz pasta of your choice (e.g. spaghetti, linguine, or fettuccine)', '1 red bell pepper, sliced', '1 yellow bell pepper, sliced', '1 small onion, thinly sliced', '2 cloves garlic, minced', '1 cup cherry tomatoes, halved', '1 cup broccoli florets', '1 cup sliced mushrooms (button or cremini)', '1 tsp olive oil', 'Salt and pepper to taste', 'Grated Parmesan cheese (optional)']
Eight runs: every one parsed. Names varied ('Vegetable Pasta', 'Vegetable Pasta Recipe'),
7 to 11 ingredients, always with quantities.Here is the recipe for vegetable pasta in the format you requested:
```
{
"name": "Vegetable Pasta",
"ingredients": [
"8 oz pasta of your choice",
"2 cups mixed vegetables (such as bell peppers, zucchini, cherry tomatoes, and onions)",
"2 cloves garlic, minced",
...
]
}
```
This JSON instance conforms to the provided schema. ...print(result.name)
# Vegetable Pasta
for ingredient in result.ingredients:
print(ingredient)
# 8 oz pasta of your choice (e.g. spaghetti, linguine, or fettuccine)
# 1 red bell pepper, sliced
# ...from langchain_core.exceptions import OutputParserException
try:
parser.parse('{"name": "Pasta", "ingredients": "tomato"}')
except OutputParserException as error:
print(error)
# Failed to parse Recipe from completion {"name": "Pasta", "ingredients": "tomato"}.
# Got: 1 validation error for Recipe
# ...class Person(BaseModel):
name: str
age: int
class PersonMaybeAge(BaseModel):
name: str
age: int | None
# Same chain shape: prompt | llm | parser, for "Priya lives in Pune."
# Person: OutputParserException - the model wrote "age": null, which is not an int
# PersonMaybeAge: name='Priya' age=None# Module 1 - by hand
raw = response["message"]["content"] # a string
raw["name"] # TypeError: string indices must be integers, not 'str'
data = json.loads(raw) # parse
data["name"] # ...then validate it yourself
# Module 2 - the parser does both
result = (prompt | llm | parser).invoke(values)
result.nameTip: Catch OutputParserException where you call the chain. Its message includes the model’s raw output - exactly what you need to see when parsing fails.
Output parsers at a glance
BaseModelDefine the structure: fields and types.
class Recipe(BaseModel): name: str
PydanticOutputParserFormat instructions before, parse and validate after.
PydanticOutputParser(pydantic_object=Recipe)
get_format_instructionsThe schema as prompt text.
"format_instructions": parser.get_format_instructions()
The chainThe parser goes last.
prompt | llm | parser
OutputParserExceptionRaised when parsing or validation fails.
from langchain_core.exceptions import OutputParserException
int | NoneLet the model say "not given".
age: int | None
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Run the chain for "chocolate cake" and loop over result.ingredients.”
“Add cooking_time: int to Recipe and run it again.”
“Call parser.parse() on JSON with a missing field and on prose with no JSON.”
“Extract a Person from "Priya lives in Pune." with age: int, then with age: int | None.”
What usually goes wrong
The model’s content is a string. Parse it first - or let the parser do it.
✗ raw = response["message"]["content"]
raw["name"]✓ result = chain.invoke(values)
result.name{"name": "Priya", "age": "twenty-nine"} is valid JSON. With age: int, it is not valid data.
The template has a {format_instructions} placeholder; invoke() must supply it, or the call fails with a missing-variable error.
✗ chain.invoke({"dish": "vegetable pasta"})✓ chain.invoke({"dish": "vegetable pasta", "format_instructions": parser.get_format_instructions()})Either the parse fails or the model invents a value. Use int | None when the answer may not be there.
✗ age: int✓ age: int | NoneKey points
- Applications need predictable data, not free-form text.
- A Pydantic class defines the structure: Recipe has name and ingredients.
- PydanticOutputParser writes format instructions and turns the reply into a Recipe.
- chain = prompt | llm | parser - the parser goes last.
- Parsing: can I read it? Validation: does it fit my class?
- Valid JSON is not valid application data.
- Constraining the format does not make the data true.
Quick check before you move on
Quiz
- 1.
Why not extract ingredients with text.split(",")?
- 2.
The model wrapped its JSON in a sentence and a code fence. What did the parser do?
- 3.
The model returns {"name": "Pasta", "ingredients": "tomato"}. What happens?
- 4.
Our recipes came back with "8 oz pasta" instead of "pasta". Did validation fail? Why not?
Interview questions
What is an output parser?
A component that converts the LLM’s response into a structured format the application can use.
Does structured output guarantee factual correctness?
No. It enforces structure and types. A required field the input cannot support will fail or be filled with something plausible - so allow nulls and verify important values.
Why is structured output useful for applications?
Applications need predictable data - fields they can read, loop over and store - rather than free-form text.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned