Structured Output
Get the answer back as JSON your program can use - then parse it, validate it, and retry only when retrying can help.
What you will be able to do
- Explain why plain text is not enough when a program has to use the answer
- Ask a model for JSON in an exact shape, and turn it into a Python dict with json.loads()
- Explain why temperature 0 suits structured output - and why it makes a plain retry useless
- Tell invalid JSON from valid JSON that carries the wrong data, and handle each differently
- Retry with feedback, so the model can correct its own formatting
- Let Ollama enforce the format with format="json" or a schema - without making it invent values
The idea, in plain English
So far the model has written text for a person to read. "Priya is 29 years old and lives in Pune" is easy for you to understand, but a program that needs a name and an age would have to pick the sentence apart itself. What it wants instead is data: {"name": "Priya", "age": 29}. Getting that from a model is called structured output.
JSON is the usual format - a way of writing data as text, with braces, quotes, and brackets, that almost every language can read. You ask for JSON in the prompt, showing the exact shape you want, and turn the reply into a Python dictionary with json.loads(). Temperature goes to 0, because here you want the format right every time, not variety.
Then you stop trusting it. Two different things can go wrong, and they need different fixes. The reply might not be JSON at all - an explanation around it, a markdown fence, a trailing comma - so json.loads() fails. That is a formatting slip the model can fix, so you tell it what went wrong and ask again. Or the reply is perfectly valid JSON with the wrong data in it - an age of null because the sentence never gave one. No retry can fix that, so you validate the fields and reject it.
Every example here was run against a local model, and the outputs shown are real. They will vary a little on your machine - which is exactly why the parsing and validation matter.
Worked example: Extract a name and an age from a sentence as JSON.
From JSON text to Python data
response["message"]["content"] is always a string, even when it looks like JSON. json.loads() parses that string and returns Python values: an object becomes a dict, an array a list, true and false become True and False, and null becomes None.
Once it is a dict, the rest of your program can use it like any other data - data["name"], data["age"] - which is the whole point.
{ "name": "Priya" }dict["Pune", "Mumbai"]list"Priya"str29 / 4.5int / floattrue / falseTrue / FalsenullNoneThe prompt is a contract
Models write prose by default. Ask loosely - "Extract the name and age as JSON" - and the reply in our run began "Here is the extracted information in JSON format:", with the JSON inside a markdown code fence. A person reads that fine. json.loads() fails on the very first character.
So the prompt spells out the contract: ONLY valid JSON, this exact shape, no explanation, no markdown. It also says what to do when information is missing - use null - because otherwise the model has to guess, and a guessed age looks exactly like a real one.
Temperature 0 - and the catch
At temperature 0 the model takes the most likely token every time, so the same prompt gives the same output. We ran the extraction five times and got five identical replies. That is what you want for a format your code depends on.
It has a consequence people miss. If the output at temperature 0 is broken, sending the same prompt again produces the same broken output. We retried the loose prompt three times and it failed three times. A retry only helps if something changes between attempts - which is why the better version below sends the error back to the model.
Watch out: Retrying the identical request at temperature 0 repeats the identical failure. Either change the request - add feedback - or do not retry.
Two problems, two fixes
Problem one: the reply is not valid JSON, and json.loads() raises json.JSONDecodeError. The model made a formatting slip, so it can fix it: add its reply and the error to the conversation, and ask for only the JSON. In our run that fixed the loose prompt on the second call.
Problem two: the reply is valid JSON, but the data is wrong. {"name": "Priya", "age": null} parses cleanly. So would {"name": "Priya"} with no age, or an age of "twenty-nine". json.loads() cannot catch any of these, so you check the fields yourself. And you do not retry - "Priya lives in Pune" has no age in it, and asking again cannot invent one.
Prose or a markdown fence around the JSONNot JSON. Retry, telling the model what was wrong.A trailing comma or a missing quoteNot JSON. Retry with feedback.A required field is missingValid JSON, wrong data. Validate and reject.A field is null because the input did not sayValid JSON, honest answer. Reject it, or let the app handle None.A field has the wrong type - "twenty-nine"Valid JSON, wrong data. Validate and reject.Two people in the sentence, one in the JSONValid and silent. Only a better shape - a list of people - fixes it.Letting Ollama enforce the format
Ollama can do part of this for you. Pass format="json" and the output is constrained to valid JSON - in our run it fixed the loose prompt with no retry at all. Pass a JSON schema as format and the output follows that exact shape.
A schema needs care. Our first schema required age to be an integer, and for "Priya lives in Pune" the model returned "age": 0 - an invented value, perfectly valid, silently wrong. Allowing null in the schema fixed it: the model returned null. Constraining the format does not make the data true, so validation stays.
Building this by hand first is the point of Module 1. In Module 2, LangChain output parsers and Pydantic models (Lesson 2.5) wrap the same ideas - a schema, parsing, validation - in less code.
Where structured output is used
Any time a program has to act on what the model says rather than show it to a person: extracting fields from documents, classifying messages, filling in a form, routing a request. Every one follows the same shape as this lesson.
Resume parser{"name": "John", "skills": ["React", "Node.js"], "experience_years": 5}Support ticket triage{"category": "payment", "priority": "high", "requires_human": true}Course tagging{"topic": "React", "category": "Frontend", "difficulty": "Beginner"}Agent tool calls (Lesson 1.5 on)The model says which tool to run, in a format your code parses.Step-by-step code
import json
raw = '{"name": "Priya", "age": 29}' # a string - not a dict yet
data = json.loads(raw) # now a dict
print(data) # {'name': 'Priya', 'age': 29}
print(data["name"]) # Priya
print(data["age"]) # 29sentence
-> prompt asking for an exact JSON shape
-> ollama.chat(..., options={"temperature": 0})
-> response["message"]["content"] a string
-> json.loads() a dict (or JSONDecodeError)
-> validate the fields usable (or reject)
-> your applicationimport ollama
import json
def extract_json(sentence, retries=2):
prompt = f"""Extract the person's name and age from this sentence.
Respond with ONLY valid JSON in this exact format: {{"name": "...", "age": ...}}
No explanation, no markdown, just the JSON.
Sentence: {sentence}"""
for attempt in range(retries + 1):
response = ollama.chat(
model="llama3.1",
messages=[{"role": "user", "content": prompt}],
options={"temperature": 0}
)
raw = response["message"]["content"].strip()
try:
return json.loads(raw)
except json.JSONDecodeError:
if attempt == retries:
raise ValueError(f"Model never returned valid JSON. Last output: {raw}")
result = extract_json("Priya is 29 years old and lives in Pune.")
print(result) # {'name': 'Priya', 'age': 29}Loose prompt ("...as JSON"), temperature 0:
Here is the extracted information in JSON format:
```
{
"name": "Priya",
"age": 29
}
```
-> json.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
The same request retried three times: fail, fail, fail
(temperature 0 gives the same output every time)
"Priya lives in Pune." -> {"name": "Priya", "age": null}
"Priya is 29 and her brother Arjun is 34." -> {"name": "Priya", "age": 29}
both valid JSON - one has no age, one dropped a personimport json
import ollama
MODEL = "llama3.1"
PROMPT = """Extract the person's name and age from this sentence.
Respond with ONLY valid JSON in this exact format: {{"name": "...", "age": ...}}
Use null for anything the sentence does not say.
No explanation, no markdown, just the JSON.
Sentence: {sentence}"""
def validate_person(data):
"""Raise ValueError unless data is a person with a name and a whole-number age."""
if not isinstance(data, dict):
raise ValueError(f"expected a JSON object, got {type(data).__name__}")
missing = [field for field in ("name", "age") if field not in data]
if missing:
raise ValueError(f"missing fields: {missing}")
if not isinstance(data["name"], str) or not data["name"].strip():
raise ValueError("name must be a non-empty string")
if isinstance(data["age"], bool) or not isinstance(data["age"], int):
raise ValueError(f"age must be a whole number, got {data['age']!r}")
return data
def extract_person(sentence, retries=2):
messages = [{"role": "user", "content": PROMPT.format(sentence=sentence)}]
for attempt in range(retries + 1):
response = ollama.chat(
model=MODEL,
messages=messages,
options={"temperature": 0},
)
raw = response["message"]["content"].strip()
# Problem 1 - not JSON at all. The model can fix this, so retry with feedback.
try:
data = json.loads(raw)
except json.JSONDecodeError as error:
if attempt == retries:
raise ValueError(f"No valid JSON after {retries + 1} attempts. Last output: {raw}") from error
messages.append({"role": "assistant", "content": raw})
messages.append({
"role": "user",
"content": f"That was not valid JSON ({error.msg}). Reply with only the JSON object.",
})
continue
# Problem 2 - valid JSON, wrong data. Retrying cannot invent a missing age,
# so reject it straight away and let the caller decide.
return validate_person(data)
print(extract_person("Priya is 29 years old and lives in Pune."))
try:
print(extract_person("Priya lives in Pune."))
except ValueError as error:
print("Rejected:", error){'name': 'Priya', 'age': 29}
Rejected: age must be a whole number, got None
With the loose prompt instead: the first reply failed to parse,
and the retry with feedback succeeded on the second call.import json
import ollama
MODEL = "llama3.1"
PERSON_SCHEMA = {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": ["integer", "null"]}, # allow null, or the model invents an age
},
"required": ["name", "age"],
}
for sentence in ["Priya is 29 years old and lives in Pune.", "Priya lives in Pune."]:
response = ollama.chat(
model=MODEL,
messages=[{
"role": "user",
"content": f"Extract the person's name and age. Use null for anything not stated.\n\nSentence: {sentence}",
}],
format=PERSON_SCHEMA, # Ollama constrains the output to this shape
options={"temperature": 0},
)
print(json.loads(response["message"]["content"])){'name': 'Priya', 'age': 29}
{'name': 'Priya', 'age': None}
With "age": {"type": "integer"} - null not allowed:
{'name': 'Priya', 'age': 29}
{'name': 'Priya', 'age': 0} <- invented: valid, and wrongWatch out: Never trust the shape of what comes back. Valid JSON only means json.loads() did not fail - not that the fields you need are there, or true.
Tip: json.JSONDecodeError is a subclass of ValueError, so one except ValueError would catch a parse failure and your own validation errors alike. Keep them separate when they need different handling, as extract_person does.
Structured output toolkit
json.loads(text)Parse a JSON string into Python values.
data = json.loads(raw)
json.JSONDecodeErrorRaised when the text is not valid JSON. A subclass of ValueError.
except json.JSONDecodeError as error:
error.msgThe short reason - useful to send back as feedback.
"Expecting value"
temperature 0Same prompt, same output: a consistent format, and identical failures.
options={"temperature": 0}format="json"Ollama constrains the reply to valid JSON.
ollama.chat(..., format="json")
format=schemaOllama constrains the reply to a JSON schema. Allow null for optional facts.
"age": {"type": ["integer", "null"]}isinstanceCheck a field has the type you expect.
isinstance(data["age"], int)
Try it yourself
The code does not change. Swap the content string and the program does something else entirely.
“Arjun, 34, works as a data engineer in Bengaluru.”
“Priya lives in Pune.”
“Priya is twenty-nine and lives in Pune.”
“Priya is 29 and her brother Arjun is 34.”
What usually goes wrong
Any extra sentence or markdown fence makes the whole reply unparseable.
✗ Give me the answer as JSON and explain it.✓ Respond with ONLY valid JSON. No explanation, no markdown.The content is always a string. Indexing it with a key raises a TypeError.
✗ raw = response["message"]["content"]
raw["name"]✓ data = json.loads(response["message"]["content"])
data["name"]One stray word and it raises. Catch json.JSONDecodeError.
✗ data = json.loads(raw)✓ try:
data = json.loads(raw)
except json.JSONDecodeError as error:
...Same input, same output: the retry reproduces the failure. Add the error to the conversation so the next attempt is different.
✗ for attempt in range(3):
raw = ollama.chat(... same messages ...)✓ messages.append({"role": "assistant", "content": raw})
messages.append({"role": "user", "content": f"That was not valid JSON ({error.msg}). Reply with only the JSON object."})A missing key, a null, or "twenty-nine" all parse cleanly. Validate the fields you use.
✗ age = json.loads(raw)["age"] + 1✓ data = validate_person(json.loads(raw))Require an integer age and the model will invent one - 0 in our run - when the input has none. Allow null for anything the input might not say.
✗ "age": {"type": "integer"}✓ "age": {"type": ["integer", "null"]}Key points
- Applications need data, not prose - structured output means asking for an exact JSON shape.
- The reply is a string; json.loads() turns it into a dict, and null into None.
- Tell the model the exact shape, and what to use for missing information.
- Temperature 0 makes the format consistent - and makes identical retries fail identically.
- Invalid JSON: retry with feedback. Valid JSON with wrong data: validate and reject.
- Validation checks that required fields exist and have the right type.
- format="json" or a schema lets Ollama enforce the shape - but a strict schema can make the model invent values.
Quick check before you move on
Quiz
- 1.
Why do we set temperature=0 here?
- 2.
What does the retry loop protect against?
- 3.
What Python function turns a JSON string into a Python dictionary?
- 4.
The model returns {"name": "Priya", "age": null} for "Priya lives in Pune." Should you retry?
- 5.
A schema requires age to be an integer. What can go wrong when the input has no age?
Interview questions
How would you handle structured output from an LLM?
Ask for an explicitly defined JSON structure at temperature 0, parse it with json.loads(), retry parse failures with the error fed back to the model, and validate the required fields and types before the data reaches the application. Where the runtime supports it, constrain the output with JSON mode or a schema - and still validate.
What is the difference between invalid JSON and invalid data?
Invalid JSON fails to parse - a formatting slip the model can correct on a retry with feedback. Invalid data parses fine but is missing fields, null, or the wrong type - usually because the input lacks the information, so retrying does not help and validation must catch it.
Does a JSON schema or JSON mode make validation unnecessary?
No. It guarantees the shape, not the truth. A schema that requires a value can push the model to invent one when the input does not contain it.
Comments
Sign in to leave a comment. Your name and photo come from Google; nothing else is shared.
Loading comments...
AI
System Design
Backend
- GraphQL8 modules · 69 lessons planned
- Core Python13 modules · 75 lessons planned
- FastAPI5 sections · 20 lessons
- Node.js14 modules · 206 lessons planned
- Node.js Performance7 chapters · 36 topics
- Event Loop Lifecycle6 phases · 3 scenarios
- Docker & Containerization11 modules · 144 lessons planned
- AWS for Developers14 modules · 219 lessons planned
- CI/CD & DevOps Automation10 modules · 134 lessons planned