0. Full results

The scores behind the series, in one place. The exam had 426 questions; that article explains the subjects and the marking.

Scores out of 1 are the share of questions answered correctly. Scores out of 10 were marked by fable. “Broken deliverables” counts answers that came back unusable - cut-off thinking, empty output.

Colours: a green box is the best in the row and green text the second best; a red box is the worst and red text the second worst. Tied scores are ordered by speed. The two-model table uses only the boxes.

Big models

No tools in this phase. Three of these are my prunes; Laguna ran as its maker shipped it.

TestLaguna 117BGLM-4.5-Air 106B (pruned)Qwen3.5-122B (pruned)Mistral-Small-4 119B (pruned)
Working window65,53665,536131,072131,072
Information security.943.938.873.968
Long documents1.0.98.981.0
Coding (pass rate).951.829.951.780
Coding (/10)8.04.77.25.3
Writing (/10)7.27.26.56.4
Reasoning (/10)7.87.07.22.6
Honesty (correct or admits ignorance).58.50.84.74
Invented facts (of 151)2125813
Spots questions to hand up.72.76.80.96
False alarms0.040.12
Broken deliverables051610
Speed (tokens/sec)36.523.121.548.6

Small models

No tools here either. All eight ran as shipped, quantised to fit.

TestQwen3.6 35BGLM-4.7 FlashGemma 26BGemma 31BNorth MiniOrnith 35Bgpt-oss 120BNemotron Super
Working window262,144202,752262,144262,144262,144262,144131,072524,288
Information security.94.917.98.98.888.95.908.918
Long documents1.0.941.01.0.741.0.82.92
Coding (pass rate).976.902.976.976.976.951.976.976
Coding (/10)7.55.07.88.37.16.58.08.0
Writing (/10)6.15.86.97.35.96.56.36.4
Reasoning (/10)6.64.55.36.45.86.36.07.1
Honesty (correct or admits ignorance).64.72.52.42.50.70.76.74
Invented facts (of 151)1814242925151213
Spots questions to hand up.44.20.88.96.52.08.72.40
False alarms0000.12.040.04
Broken deliverables37116542
Speed (tokens/sec)71.039.757.814.553.873.174.933.1

Gemmas with tools

The second phase gave the two Gemmas the reference library, web search, the claim checker, and the route up to the cloud. See Choosing my model.

TestGemma 26BGemma 31B
Honesty (corrected).96.98
Tool use (answer quality).80.65
Correct tool choice.75.75
Wasted tool calls00
Spots questions to hand up1.0.72
Exact hand-up decisions.90.68
Sensitive questions misrouted00
Speed (tokens/sec)65-9831-52
Graphics memory used44 GB66 GB

Notes

  • Puzzle75 and BTL-3 need modified serving software, so they weren’t tested. Step-3.7-Flash didn’t fit. MiniMax was dropped early.
  • Laguna and GLM served 65,536-token windows; Qwen and Mistral served 131,072. A bigger window costs memory, so the window row is partly a choice.
  • Mistral’s speed is from its second run; the first was slowed by other installs.
  • The honesty rows are from the runs without tools. With tools the Gemmas rose to the corrected figures in the third table; Choosing my model explains the correction.

Thanks for reading!