Research ยท Trustworthy AI for command and control
Constrained decoding explained: making language models produce valid structured output
Constrained decoding guarantees the shape of a language model's output. It does not guarantee that the content is true. In our evaluation it took structural validity from 60.8% to 100%, and left the question of accuracy for other parts of the system.
What goes wrong without it
A language model asked for a record in a fixed format usually produces one. Some outputs contain a bracket that never closes, omit a required field or carry a value outside the allowed options. In an automated pipeline any of those means a rejected message. In our test, the same model and prompt without a constraint produced a structurally valid record 60.8% of the time. Four outputs in ten could not be processed at all.
How it works
A language model writes one token at a time. At each step it scores every token in its vocabulary and one is chosen. Constrained decoding adds a check between the scoring and the choice.
- The required format is written down as a JSON Schema or a grammar, and compiled into a state machine that tracks where in the format the output currently is.
- At each step the state machine says which tokens could legally come next.
- Every other token is removed from consideration, and the model chooses among the rest.
- The chosen token moves the state machine forward, and the cycle repeats until the format is complete.
The result cannot be malformed, because a malformed choice was never on offer. Willard and Louf describe the general technique in Efficient Guided Generation for Large Language Models.
What it guarantees, and what it does not
| It guarantees | It does not guarantee |
|---|---|
| The output parses. | The values are correct. |
| Every required field is present. | The values came from the input and were not invented. |
| Each value has the allowed type and is one of the allowed options. | The model was confident, or should have been. |
The last row matters most. When the model is unsure between two allowed values, the constraint still makes it pick one, and the record looks exactly as clean as a confident one. That is the reason Gridnorth follows extraction with a per-field confidence check and an operator confirmation step. See how it works.
What we measured
- Structurally valid records: 100% with the constraint, 60.8% without.
- Field accuracy across all fields: 84.4%. The constraint made records processable. It did not make every value right.
- Median extraction time: 7.7 s on a single NVIDIA T4 GPU. We report end-to-end time and have not separated out the cost of the constraint.
The results page gives the method and the limits, and test integrity explains the precautions.
Last reviewed September 29, 2026