Gridnorth
Laboratory evaluation: results supporting TRL 3
Bottom line: on a held-out synthetic test set of 120 inputs, constrained extraction produced a structurally valid record every time, field accuracy was 84.4%, and flagging caught 64.4% of errors. Gaps remain in radio-style input and in flagging coverage. Conditions, findings and open items follow.
Scope and status
This page reports one laboratory evaluation of each critical Gridnorth function, run in September 2026. The results support Technology Readiness Level (TRL) 3: proof of concept demonstrated in the laboratory on representative data. TRL is a maturity rating for the technology. The data behind it are laboratory results on synthetic messages, not field data.
Test conditions
- Model: Mistral-7B-Instruct, quantized to 4 bits, running locally with no external service.
- Messages: three types built on public formats: an adapted 9-line MEDEVAC request, a SALUTE contact report and a logistics status report. These are stand-ins. They are not APP-11 definitions.
- Test set: 120 inputs, half English and half French, generated from a fixed seed and independent of the prompt. Reference answers came from the generator and never from the model. An automated check confirmed that no test value appeared in any prompt. Inputs varied by language, style, layout and mid-message corrections.
- Input styles: typed-style text, and radio-style text, which is written the way messages are spoken over the radio. No audio was used.
- Baseline: the same model and prompt without constrained decoding.
- Flagging threshold: set on half the items and reported on the other half.
- Encoding: EXIficient, an implementation of the W3C EXI standard, built from source at pinned versions.
A development run on a separate test set exposed two defects, in callsign formatting and in a confidence measure. Both were corrected before this run, and only this run is reported. Test integrity in the evaluation gives the reasoning.
Findings
Structure
| Measure | Result |
|---|---|
| Structurally valid records, constrained decoding | 100% |
| Structurally valid records, same model unconstrained | 60.8% |
Extraction accuracy
| Measure | Result |
|---|---|
| Field accuracy, all fields | 84.4% |
| Field accuracy, typed-style text | 94.3% |
| Field accuracy, radio-style text | 74.5% |
| Field accuracy, English | 86.4% |
| Field accuracy, French | 82.3% |
Operator confirmation
| Measure | Result |
|---|---|
| Fields flagged for confirmation | 15.3% |
| Errors caught by flagging | 64.4% |
| Critical fields wrong and not flagged | 2.9% |
Encoding
| Measure | Result |
|---|---|
| Mean message size, plain XML text | 328 bytes |
| Mean message size, gzip | 202 bytes |
| Mean message size, EXI without the schema | 190 bytes |
| Mean message size, schema-informed EXI | 19 bytes |
| EXI round trip | 100% lossless |
Timing
| Measure | Result |
|---|---|
| Median extraction time | 7.7 s on a single NVIDIA T4 graphics processing unit (GPU), a data-centre part |
Assessment
Constraining the model's output is what makes it usable for machine-to-machine messaging: without the constraint, four in ten outputs could not be processed at all. Typed-style text is above 94% field accuracy. The gap between English and French is about four points.
Limitations and open items
- Critical fields in radio-style text. Grid references, times and callsigns read out digit by digit are the main source of error. The next stage moves their parsing to deterministic code.
- Flagging coverage. Flagging catches about two thirds of errors. The target is at least 80%.
- Speech. All inputs were text. Speech from real speakers has not been tested.
- Schemas. The schemas are representative. Results on APP-11 schemas will differ, and the EXI reduction is likely to be smaller on messages with more free text.
- Hardware. Timing was measured on a data-centre GPU, not on edge hardware.
Last reviewed September 29, 2026