Chaotic scribbled text flowing into a geometric gate structure and emerging on the other side as neat, orderly brick-like blocks Tech
AI-generated, Working Theory
Tech · ◉ Evergreen

The model can say anything. Your schema decides what ships.

by · ·5 min·Working Theory

A prompt is a request; you need a contract. The model's output is not your result — it's a proposal, and a proposal gets checked before it ships.

Here is the whole problem with building on a language model, compressed into one line: it will do what you asked almost every time, and “almost” is not a type you can pass to the next function.

You write a prompt. “Return a JSON object with title, priority (1–5), and tags (array of strings).” You run it a hundred times and it works a hundred times, so you wire it into the pipeline and move on. Then, on request 3,412, the model gets chatty and prefixes the JSON with “Sure! Here’s your object:”. Or it decides priority should be "high" this time. Or it invents a tags field that’s a comma-separated string instead of an array, because that’s how it saw the world for one unlucky token. Your parser throws, or worse, it doesn’t — and the garbage flows downstream where it’s much more expensive to notice.

The instinct is to fix the prompt. Add “ONLY return valid JSON.” Add “do not include any preamble.” Add three exclamation points. This is persuasion, and persuasion is the wrong layer. A prompt is a request; you need a contract. The place a contract lives is not in the model’s good intentions — it’s in a validation layer the output has to pass through before it counts as real.

The mental shift is small and load-bearing: the model’s output is not your result. It’s a proposal. A proposal gets checked before it ships.

model free text schema gate valid ✓ ships UI · DB · next agent invalid ✗ repair & retry, or reject
The prompt persuades; the gate enforces. Nothing downstream sees output the schema didn't approve. Original diagram · Working Theory

In practice the gate has two places to stand, and mature systems use both. The first is at generation time: JSON mode, tool/function calling, or constrained decoding, where the runtime is only allowed to emit tokens that keep the output valid against a grammar or schema. This is the strongest guarantee, because it makes the malformed output nearly impossible rather than merely discouraged. The second is after the fact: you parse the string, validate it against the schema, and — this is the part people skip — you design the failure path. On a validation miss, you either run a bounded repair loop (hand the model the error and ask it to fix its own output, with a hard retry budget) or you reject and degrade gracefully. What you never do is let unvalidated output through because “it’s usually fine.”

And “schema” is bigger than shape. Once you have a gate, it’s the natural home for every other thing you don’t trust the model to get right on its own: value ranges (priority really is 1–5), allow-lists (this agent may call these tools and no others), refusal detection, and content checks on what’s about to hit a user or a database. The schema isn’t just deciding whether the JSON parses. It’s deciding what is allowed to ship to the next stage — and in an agentic system, the next stage might be an action in the real world, which is exactly where you want a gate and not a vibe.

There’s a real cost, so name it. Constrained decoding, pushed too hard, can flatten the quality of the answer — clamp the grammar too tight and you sometimes get valid-but-worse. Validation adds a hop of latency and forces you to own a failure path you’d rather pretend won’t fire. Those are the right costs to pay. A rejected response you can retry beats a malformed one you shipped; a slightly slower answer beats a corrupt row.

The one-line version to keep: the prompt is where you ask, and the schema is where you insist. Build products where the insisting is done by code, at the boundary, every time — not by a sentence in the prompt that the model is free to have a bad day about.

To look up: structured/constrained decoding and grammar-constrained generation; JSON schema validation as a runtime contract; the “LLM output as proposal, not result” framing; and, adjacent, guardrails and least-privilege tool allow-lists for agents.

Sources

  • Structured/constrained decoding
  • JSON schema validation as a runtime contract
  • least-privilege tool allow-lists for agents

Liked this? Get the next one in Working Theory.

Going weekly in August (it's in beta now). One genuinely interesting read on building, the brain, and the science most people missed.

Subscribe →
Got a reaction, a counter-example, or something I missed? Reply by email — I read everything.
◉ join in

Where have you hit this — in a product you use, or one you're building?

Threads open here soon. For now, the conversation lives two clicks away — discuss on GitHub, or just reply by email. I read and answer everything.