Model ranking · all 25 model / reasoning settings
GLM-5.3, Grok 4.6, GPT-6 Astra, and GPT-5.6 Luna at none, low, medium, high, xhigh, and max are included alongside every historical model. Scores measure probability consistency, not accuracy. 148 shared tuples across 24 participating settings, from 157 reviewed tuples. All native settings and API runs with complete original coverage participate; incomplete historical settings remain visible and unranked. Selection by shared availability can change the ordering; these results do not represent the full original 300-tuple benchmark. Native subscription runs and historical API runs differ in dates, prompts and reasoning settings. Coverage and methods →
Jev uses a distinct probability channel: TypeSafe Noul returns a direct yes-probability for each question, independently, with its full resolution criteria. No shared question context, logical hints, prompt tuning, verbal answers or token logprobs. Reasoning controls and an explicit abstention field are not exposed. The same algebraic formulas and shared tuple IDs are used; this does not isolate model quality from elicitation method. Research and protocol →
| Model | Reasoning | Shared mean ↓ | Shared tuples | Reviewed-scope mean | Reviewed tuples | Valid probabilities | Method | Run date (UTC) |
|---|---|---|---|---|---|---|---|---|
| GLM-5.3 | Provider default | 0.0587 | 148/148 | 0.0673 | 157/157 | 274/274 | Subscription | 2026-09-13 |
| kimi-k3 | Not recorded* | 0.0608 | 148/148 | 0.0701 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| Grok 4.6 | low | 0.0635 | 148/148 | 0.0642 | 156/157 | 273/274 | Subscription | 2026-09-13 |
| GPT-5.6 Luna | max | 0.0636 | 148/148 | 0.0657 | 157/157 | 274/274 | Subscription | 2026-09-13 |
| GPT-5.6 Luna | high | 0.0760 | 148/148 | 0.0782 | 151/157 | 273/274 | Subscription | 2026-09-13 |
| nemotron-3-ultra-nvfp4 | Not recorded* | 0.0821 | 148/148 | 0.0798 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| glm-5p2 | Not recorded* | 0.0845 | 148/148 | 0.0870 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| JEV 1.13.0 · TypeSafe | Not exposed | 0.0864 | 148/148 | 0.0895 | 157/157 | 274/274 | TypeSafe Noul API | 2026-09-17 |
| qwen3p8-2p4t-a95b | Not recorded* | 0.0871 | 148/148 | 0.0878 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| qwen3p8-max | Not recorded* | 0.0874 | 148/148 | 0.0855 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| deepseek-v4-flash-0731 | Not recorded* | 0.0876 | 148/148 | 0.0918 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| GPT-5.6 Luna | medium | 0.0901 | 148/148 | 0.0905 | 155/157 | 272/274 | Subscription | 2026-09-13 |
| GPT-5.6 Luna | xhigh | 0.0954 | 148/148 | 0.0951 | 151/157 | 273/274 | Subscription | 2026-09-13 |
| minimax-m3 | Not recorded* | 0.0978 | 148/148 | 0.0941 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| GPT-5.6 Luna | low | 0.1042 | 148/148 | 0.1038 | 151/157 | 273/274 | Subscription | 2026-09-13 |
| qwen3p7-plus | Not recorded* | 0.1044 | 148/148 | 0.1130 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| deepseek-v4-pro-0813 | Not recorded* | 0.1076 | 148/148 | 0.1033 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| kimi-k2p7-code | Not recorded* | 0.1122 | 148/148 | 0.1104 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| kimi-k2p6 | Not recorded* | 0.1125 | 148/148 | 0.1114 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| GPT-6 Astra | low | 0.1155 | 148/148 | 0.1157 | 150/157 | 269/274 | Subscription | 2026-09-13 |
| GPT-5.6 Luna | none | 0.1214 | 148/148 | 0.1224 | 150/157 | 269/274 | Subscription | 2026-09-13 |
| gpt-oss-120b | Not recorded* | 0.1426 | 148/148 | 0.1414 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| nemotron-lightning-3p5-30b-a3b | Not recorded* | 0.1843 | 148/148 | 0.1895 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| gpt-oss-20b | Not recorded* | 0.1869 | 148/148 | 0.1825 | 157/157 | Not recorded | Fireworks API | 2026-08-19 |
| muse-glimmer-30b | Not recorded* | — unranked: incomplete historical coverage | — | 0.0228 | 47/157 | Not recorded | Fireworks API | 2026-08-19 |
*Historical API runs requested none, with low as a fallback; the setting actually used was not recorded per proposition. Subscription runs were measured in September and historical API runs in August, with different native prompts, reasoning settings, and output limits.