0. Full results

The scores behind the series, in one place. The exam had 426 questions; that article explains the subjects and the marking.
Scores out of 1 are the share of questions answered correctly. Scores out of 10 were marked by fable. “Broken deliverables” counts answers that came back unusable - cut-off thinking, empty output.
Colours: a green box is the best in the row and green text the second best; a red box is the worst and red text the second worst. Tied scores are ordered by speed. The two-model table uses only the boxes.
Big models
No tools in this phase. Three of these are my prunes; Laguna ran as its maker shipped it.
| Test | Laguna 117B | GLM-4.5-Air 106B (pruned) | Qwen3.5-122B (pruned) | Mistral-Small-4 119B (pruned) |
|---|---|---|---|---|
| Working window | 65,536 | 65,536 | 131,072 | 131,072 |
| Information security | .943 | .938 | .873 | .968 |
| Long documents | 1.0 | .98 | .98 | 1.0 |
| Coding (pass rate) | .951 | .829 | .951 | .780 |
| Coding (/10) | 8.0 | 4.7 | 7.2 | 5.3 |
| Writing (/10) | 7.2 | 7.2 | 6.5 | 6.4 |
| Reasoning (/10) | 7.8 | 7.0 | 7.2 | 2.6 |
| Honesty (correct or admits ignorance) | .58 | .50 | .84 | .74 |
| Invented facts (of 151) | 21 | 25 | 8 | 13 |
| Spots questions to hand up | .72 | .76 | .80 | .96 |
| False alarms | 0 | .04 | 0 | .12 |
| Broken deliverables | 0 | 5 | 16 | 10 |
| Speed (tokens/sec) | 36.5 | 23.1 | 21.5 | 48.6 |
Small models
No tools here either. All eight ran as shipped, quantised to fit.
| Test | Qwen3.6 35B | GLM-4.7 Flash | Gemma 26B | Gemma 31B | North Mini | Ornith 35B | gpt-oss 120B | Nemotron Super |
|---|---|---|---|---|---|---|---|---|
| Working window | 262,144 | 202,752 | 262,144 | 262,144 | 262,144 | 262,144 | 131,072 | 524,288 |
| Information security | .94 | .917 | .98 | .98 | .888 | .95 | .908 | .918 |
| Long documents | 1.0 | .94 | 1.0 | 1.0 | .74 | 1.0 | .82 | .92 |
| Coding (pass rate) | .976 | .902 | .976 | .976 | .976 | .951 | .976 | .976 |
| Coding (/10) | 7.5 | 5.0 | 7.8 | 8.3 | 7.1 | 6.5 | 8.0 | 8.0 |
| Writing (/10) | 6.1 | 5.8 | 6.9 | 7.3 | 5.9 | 6.5 | 6.3 | 6.4 |
| Reasoning (/10) | 6.6 | 4.5 | 5.3 | 6.4 | 5.8 | 6.3 | 6.0 | 7.1 |
| Honesty (correct or admits ignorance) | .64 | .72 | .52 | .42 | .50 | .70 | .76 | .74 |
| Invented facts (of 151) | 18 | 14 | 24 | 29 | 25 | 15 | 12 | 13 |
| Spots questions to hand up | .44 | .20 | .88 | .96 | .52 | .08 | .72 | .40 |
| False alarms | 0 | 0 | 0 | 0 | .12 | .04 | 0 | .04 |
| Broken deliverables | 3 | 7 | 1 | 1 | 6 | 5 | 4 | 2 |
| Speed (tokens/sec) | 71.0 | 39.7 | 57.8 | 14.5 | 53.8 | 73.1 | 74.9 | 33.1 |
Gemmas with tools
The second phase gave the two Gemmas the reference library, web search, the claim checker, and the route up to the cloud. See Choosing my model.
| Test | Gemma 26B | Gemma 31B |
|---|---|---|
| Honesty (corrected) | .96 | .98 |
| Tool use (answer quality) | .80 | .65 |
| Correct tool choice | .75 | .75 |
| Wasted tool calls | 0 | 0 |
| Spots questions to hand up | 1.0 | .72 |
| Exact hand-up decisions | .90 | .68 |
| Sensitive questions misrouted | 0 | 0 |
| Speed (tokens/sec) | 65-98 | 31-52 |
| Graphics memory used | 44 GB | 66 GB |
Notes
- Puzzle75 and BTL-3 need modified serving software, so they weren’t tested. Step-3.7-Flash didn’t fit. MiniMax was dropped early.
- Laguna and GLM served 65,536-token windows; Qwen and Mistral served 131,072. A bigger window costs memory, so the window row is partly a choice.
- Mistral’s speed is from its second run; the first was slowed by other installs.
- The honesty rows are from the runs without tools. With tools the Gemmas rose to the corrected figures in the third table; Choosing my model explains the correction.
Thanks for reading!