consistency-ft · corrected comparison details

corrected comparison · full ranking, coverage, methods and subscription usage — selected runs as of 2026-09-17T00:30:34.686981+00:00

Model ranking · all 25 model / reasoning settings

GLM-5.3, Grok 4.6, GPT-6 Astra, and GPT-5.6 Luna at none, low, medium, high, xhigh, and max are included alongside every historical model. Scores measure probability consistency, not accuracy. 148 shared tuples across 24 participating settings, from 157 reviewed tuples. All native settings and API runs with complete original coverage participate; incomplete historical settings remain visible and unranked. Selection by shared availability can change the ordering; these results do not represent the full original 300-tuple benchmark. Native subscription runs and historical API runs differ in dates, prompts and reasoning settings. Coverage and methods →

Jev uses a distinct probability channel: TypeSafe Noul returns a direct yes-probability for each question, independently, with its full resolution criteria. No shared question context, logical hints, prompt tuning, verbal answers or token logprobs. Reasoning controls and an explicit abstention field are not exposed. The same algebraic formulas and shared tuple IDs are used; this does not isolate model quality from elicitation method. Research and protocol →

Sorted by mean violation on the same shared tuples (lower is better). Reviewed-scope means use each model’s available reviewed tuples and can have different denominators.
ModelReasoningShared mean ↓Shared tuplesReviewed-scope meanReviewed tuplesValid probabilitiesMethodRun date (UTC)
GLM-5.3Provider default0.0587148/1480.0673157/157274/274Subscription2026-09-13
kimi-k3Not recorded*0.0608148/1480.0701157/157Not recordedFireworks API2026-08-19
Grok 4.6low0.0635148/1480.0642156/157273/274Subscription2026-09-13
GPT-5.6 Lunamax0.0636148/1480.0657157/157274/274Subscription2026-09-13
GPT-5.6 Lunahigh0.0760148/1480.0782151/157273/274Subscription2026-09-13
nemotron-3-ultra-nvfp4Not recorded*0.0821148/1480.0798157/157Not recordedFireworks API2026-08-19
glm-5p2Not recorded*0.0845148/1480.0870157/157Not recordedFireworks API2026-08-19
JEV 1.13.0 · TypeSafeNot exposed0.0864148/1480.0895157/157274/274TypeSafe Noul API2026-09-17
qwen3p8-2p4t-a95bNot recorded*0.0871148/1480.0878157/157Not recordedFireworks API2026-08-19
qwen3p8-maxNot recorded*0.0874148/1480.0855157/157Not recordedFireworks API2026-08-19
deepseek-v4-flash-0731Not recorded*0.0876148/1480.0918157/157Not recordedFireworks API2026-08-19
GPT-5.6 Lunamedium0.0901148/1480.0905155/157272/274Subscription2026-09-13
GPT-5.6 Lunaxhigh0.0954148/1480.0951151/157273/274Subscription2026-09-13
minimax-m3Not recorded*0.0978148/1480.0941157/157Not recordedFireworks API2026-08-19
GPT-5.6 Lunalow0.1042148/1480.1038151/157273/274Subscription2026-09-13
qwen3p7-plusNot recorded*0.1044148/1480.1130157/157Not recordedFireworks API2026-08-19
deepseek-v4-pro-0813Not recorded*0.1076148/1480.1033157/157Not recordedFireworks API2026-08-19
kimi-k2p7-codeNot recorded*0.1122148/1480.1104157/157Not recordedFireworks API2026-08-19
kimi-k2p6Not recorded*0.1125148/1480.1114157/157Not recordedFireworks API2026-08-19
GPT-6 Astralow0.1155148/1480.1157150/157269/274Subscription2026-09-13
GPT-5.6 Lunanone0.1214148/1480.1224150/157269/274Subscription2026-09-13
gpt-oss-120bNot recorded*0.1426148/1480.1414157/157Not recordedFireworks API2026-08-19
nemotron-lightning-3p5-30b-a3bNot recorded*0.1843148/1480.1895157/157Not recordedFireworks API2026-08-19
gpt-oss-20bNot recorded*0.1869148/1480.1825157/157Not recordedFireworks API2026-08-19
muse-glimmer-30bNot recorded*unranked: incomplete historical coverage0.022847/157Not recordedFireworks API2026-08-19

*Historical API runs requested none, with low as a fallback; the setting actually used was not recorded per proposition. Subscription runs were measured in September and historical API runs in August, with different native prompts, reasoning settings, and output limits.

Coverage and comparison limits

Reviewed target: 157 whole tuples across 10 checks, using 274 distinct propositions. The review applied the documented topic and semantic criteria to source texts, without consulting model responses. The ranking uses only the 148 tuples scored by every one of the 24 participating settings. The table below retains each model’s coverage on the full reviewed target, including questions outside that shared set. Missing readouts and incomplete archived coverage can select easier questions; a shared denominator does not remove this selection effect. Reviewed scope and exclusion reasons. Native protocol and audit. Unanswered probabilities remain missing, never zero. Low violation measures coherence, not forecasting accuracy or calibration.

model / reasoningmethodscored tuplesshared ranking tuplesvalid probabilitiescompleted propositionsstatusmissing readouts
glm-5.3 · native-default (native)subscription CLI157/157148274/274274/274complete
gpt-5.6-luna · high (native)subscription CLI151/157148273/274274/274completed-with-missingabstention: 1
gpt-5.6-luna · low (native)subscription CLI151/157148273/274274/274completed-with-missingabstention: 1
gpt-5.6-luna · max (native)subscription CLI157/157148274/274274/274complete
gpt-5.6-luna · medium (native)subscription CLI155/157148272/274274/274completed-with-missingabstention: 2
gpt-5.6-luna · none (native)subscription CLI150/157148269/274274/274completed-with-missingabstention: 5
gpt-5.6-luna · xhigh (native)subscription CLI151/157148273/274274/274completed-with-missingabstention: 1
gpt-6-astra · low (native)subscription CLI150/157148269/274274/274completed-with-missingabstention: 5
grok-4.6 · low (native)subscription CLI156/157148273/274274/274completed-with-missingabstention: 1
JEV 1.13.0 · TypeSafeTypeSafe Noul API157/157148274/274274/274completedinfrastructure: 1
deepseek-v4-flash-0731Fireworks API157/157148unrecordedunrecordedhistorical full
deepseek-v4-pro-0813Fireworks API157/157148unrecordedunrecordedhistorical full
glm-5p2Fireworks API157/157148unrecordedunrecordedhistorical full
gpt-oss-120bFireworks API157/157148unrecordedunrecordedhistorical full
gpt-oss-20bFireworks API157/157148unrecordedunrecordedhistorical full
kimi-k2p6Fireworks API157/157148unrecordedunrecordedhistorical full
kimi-k2p7-codeFireworks API157/157148unrecordedunrecordedhistorical full
kimi-k3Fireworks API157/157148unrecordedunrecordedhistorical full
minimax-m3Fireworks API157/157148unrecordedunrecordedhistorical full
muse-glimmer-30bFireworks API47/157unranked: incomplete historical coverageunrecordedunrecordedhistorical partial
nemotron-3-ultra-nvfp4Fireworks API157/157148unrecordedunrecordedhistorical full
nemotron-lightning-3p5-30b-a3bFireworks API157/157148unrecordedunrecordedhistorical full
qwen3p7-plusFireworks API157/157148unrecordedunrecordedhistorical full
qwen3p8-2p4t-a95bFireworks API157/157148unrecordedunrecordedhistorical full
qwen3p8-maxFireworks API157/157148unrecordedunrecordedhistorical full

Model ranking · native subscriptions

Fresh CLI sessions, one verbalized probability per proposition, tools and browsing disabled. Each reasoning level is a separate setting. No logprobs are available. Native agent scaffolds, reasoning settings and output limits differ by provider; this is not a matched API experiment. Both charts retain the same shared tuples across all participating settings.

0.000.050.100.150.20glm-5.3 · native-default (native)0.059grok-4.6 · low (native)0.064gpt-5.6-luna · max (native)0.064gpt-5.6-luna · high (native)0.076gpt-5.6-luna · medium (native)0.090gpt-5.6-luna · xhigh (native)0.095gpt-5.6-luna · low (native)0.104gpt-6-astra · low (native)0.116gpt-5.6-luna · none (native)0.121

Model ranking · historical Fireworks API

Original August runs, using verbalized probabilities at temperature 0. Requested reasoning was none, with low as a fallback; actual fallback choices were not recorded per proposition. Native and API results use the same tuple target and algebraic violation formulas, with different elicitation settings. Both charts retain the same shared tuples across all participating settings.

0.000.120.250.380.50kimi-k30.061nemotron-3-ultra-nvfp40.082glm-5p20.084qwen3p8-2p4t-a95b0.087qwen3p8-max0.087deepseek-v4-flash-07310.088minimax-m30.098qwen3p7-plus0.104deepseek-v4-pro-08130.108kimi-k2p7-code0.112kimi-k2p60.113gpt-oss-120b0.143nemotron-lightning-3p5-30b-a3b0.184gpt-oss-20b0.187muse-glimmer-30b (47/157 reviewed)unranked: incomplete historical coverage

Subscription usage

Across all archived campaigns: 61,738,874 recorded native tokens and 8,753 CLI launches, including setup and interruptions. 76 interrupted launches have no native receipt; native usage is unknown for 106 launches in total. Unknown consumption is not zero. Imported observations are counted once. Quota percentages are not summed across campaigns or accounts.

Quota changes below cover native-20260913-scoped-final: 2026-09-13T04:30:48.606308+00:00 to 2026-09-13T04:55:52.120664+00:00. Account-wide deltas can include other fleet sessions; window resets and provider accounting prevent inference from tokens alone. Download complete usage report.

providerwindowbeforeafterused deltaunitmeasurement limits
OpenAI · subscription 1weekly12120percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 2weekly29312percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 3weekly56593percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 4weekly110percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 5weekly550percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 6weekly42497percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 7weekly24251percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
Z.ai GLM coding planTIME_LIMIT (provider unit 5 × 1)00unavailablepercentage pointsNative quota endpoint; includes excluded setup within the measurement window and any concurrent account activity. Reset timestamp changed; no delta attributed.
Z.ai GLM coding plan5-hour token window451percentage pointsNative quota endpoint; includes excluded setup within the measurement window and any concurrent account activity. Window identifiers are preserved as returned by the provider.
Z.ai GLM coding planweekly token window (provider unit 6)110percentage pointsNative quota endpoint; includes excluded setup within the measurement window and any concurrent account activity. Window identifiers are preserved as returned by the provider.
Grok SuperGrok Heavyweekly2.02.00.0percentage pointsNative /usage screen, whole-percent resolution; shared-account delta. A zero displayed change does not mean zero consumption. Same weekly reset verified.