Research ยท Trustworthy AI for command and control

Test integrity in the Gridnorth evaluation

Language-model results are easy to flatter. This article lists the precautions taken in the laboratory evaluation and the weaknesses that remain.

What we did

  1. Generated the test set from a fixed seed. The 120 inputs do not depend on the prompt, and can be regenerated exactly.
  2. Took reference answers from the generator, never from the model. Whatever the model says is judged against what the generator put in.
  3. Checked for leakage automatically. A script confirmed that no test value appeared in any prompt.
  4. Kept development and test data separate. We built and tuned on one set and reported on another.
  5. Tuned the flagging threshold on half the items and reported on the other half. A threshold chosen on the data it is scored on flatters the result.
  6. Used a baseline. The same model and prompt, without constrained decoding.
  7. Pinned versions and built the EXI implementation from source, so the result can be reproduced.
  8. Reported only the final run. Our development run exposed two defects, in callsign formatting and in a confidence measure. We fixed both before the final run, and we say so.

What these precautions do not cover

  • The data is synthetic and we wrote the generator. If our generator writes messages more tidily than people do, the results are optimistic. Real messages are the test that matters.
  • All inputs were text. Speech recognition adds its own errors and has not been tested.
  • One model, one run. We have not measured how much the numbers move between runs or between models.
  • The schemas are stand-ins, not APP-11 definitions.
  • The sample is small. With half the test set in each language, the gap of about four points between English and French is an indication, not a finding.

What comes next

Test data written by people outside the team, real speech from real speakers, real message schemas, and enough repeat runs to put an error bar on each number. The results page will be updated as each arrives, and the old figures will stay visible for comparison.

Last reviewed September 29, 2026