consistency-ft · original benchmark details

original benchmark · full ranking, coverage, methods and subscription usage — selected runs as of 2026-09-17T00:30:34.686981+00:00

Model ranking · all 25 model / reasoning settings

GLM-5.3, Grok 4.6, GPT-6 Astra, and GPT-5.6 Luna at none, low, medium, high, xhigh, and max are included alongside every historical model. Scores measure probability consistency, not accuracy. 264 shared tuples across 24 participating settings, from the original 300-tuple benchmark. All native settings and API runs with complete original coverage participate; incomplete historical settings remain visible and unranked. No topic or semantic subset is applied here. Missing original readouts reduce the shared set; selection by shared availability can change the ordering. Native results use the original native-20260912 campaign (numeric-v1). Native subscription runs and historical API runs differ in dates, prompts and reasoning settings. Coverage and methods →

Jev uses a distinct probability channel: TypeSafe Noul returns a direct yes-probability for each question, independently, with its full resolution criteria. No shared question context, logical hints, prompt tuning, verbal answers or token logprobs. Reasoning controls and an explicit abstention field are not exposed. The same algebraic formulas and shared tuple IDs are used; this does not isolate model quality from elicitation method. Research and protocol →

Sorted by mean violation on the same shared tuples (lower is better). Original-cohort means use each model’s available original tuples and can have different denominators.
ModelReasoningShared mean ↓Shared tuplesOriginal-cohort meanOriginal tuplesValid probabilitiesMethodRun date (UTC)
Grok 4.6low0.0620264/2640.0609300/300591/591Subscription2026-09-12
GLM-5.3Provider default0.0641264/2640.0687300/300591/591Subscription2026-09-12
GPT-5.6 Lunaxhigh0.0665264/2640.0656268/300562/591Subscription2026-09-12
GPT-5.6 Lunamax0.0666264/2640.0733278/300564/591Subscription2026-09-12
kimi-k3Not recorded*0.0738264/2640.0823300/300Not recordedFireworks API2026-08-19
nemotron-3-ultra-nvfp4Not recorded*0.0820264/2640.0842300/300Not recordedFireworks API2026-08-19
qwen3p8-maxNot recorded*0.0849264/2640.0851300/300Not recordedFireworks API2026-08-19
qwen3p8-2p4t-a95bNot recorded*0.0849264/2640.0865300/300Not recordedFireworks API2026-08-19
deepseek-v4-flash-0731Not recorded*0.0871264/2640.0884300/300Not recordedFireworks API2026-08-19
JEV 1.13.0 · TypeSafeNot exposed0.0873264/2640.0923300/300591/591TypeSafe Noul API2026-09-17
minimax-m3Not recorded*0.0920264/2640.0894300/300Not recordedFireworks API2026-08-19
GPT-5.6 Lunalow0.0933264/2640.0926280/300569/591Subscription2026-09-12
qwen3p7-plusNot recorded*0.0966264/2640.1049300/300Not recordedFireworks API2026-08-19
glm-5p2Not recorded*0.0977264/2640.0967300/300Not recordedFireworks API2026-08-19
GPT-5.6 Lunahigh0.1026264/2640.1018272/300560/591Subscription2026-09-12
GPT-5.6 Lunamedium0.1048264/2640.1030272/300565/591Subscription2026-09-12
deepseek-v4-pro-0813Not recorded*0.1063264/2640.1034300/300Not recordedFireworks API2026-08-19
kimi-k2p7-codeNot recorded*0.1090264/2640.1077300/300Not recordedFireworks API2026-08-19
kimi-k2p6Not recorded*0.1095264/2640.1062300/300Not recordedFireworks API2026-08-19
GPT-5.6 Lunanone0.1097264/2640.1201300/300591/591Subscription2026-09-12
GPT-6 Astralow0.1132264/2640.1139276/300569/591Subscription2026-09-12
gpt-oss-120bNot recorded*0.1487264/2640.1462300/300Not recordedFireworks API2026-08-19
gpt-oss-20bNot recorded*0.1600264/2640.1593300/300Not recordedFireworks API2026-08-19
nemotron-lightning-3p5-30b-a3bNot recorded*0.1860264/2640.1898300/300Not recordedFireworks API2026-08-19
muse-glimmer-30bNot recorded*unranked: incomplete historical coverage0.0381112/300Not recordedFireworks API2026-08-19

*Historical API runs requested none, with low as a fallback; the setting actually used was not recorded per proposition. Subscription runs were measured in September and historical API runs in August, with different native prompts, reasoning settings, and output limits.

Coverage and comparison limits

Original target: 300 tuples across 10 checks, using 591 distinct propositions. The ranking compares the same 264 tuples scored by all 24 participating settings. Each model’s original coverage and mean over its own available tuples are preserved below; those means can have different denominators. Missing original readouts reduce the shared set. A shared-set headline does not imply completion of the full original benchmark. Native results are fixed to native-20260912 (numeric-v1). The corrected subset and rerun are reported separately. Unanswered probabilities remain missing, never zero. Low violation measures coherence, not forecasting accuracy or calibration.

model / reasoningmethodscored tuplesshared ranking tuplesvalid probabilitiescompleted propositionsstatusmissing readouts
glm-5.3 · native-default (native)subscription CLI300/300264591/591591/591complete
gpt-5.6-luna · high (native)subscription CLI272/300264560/591591/591completed-with-missingrefusal: 31
gpt-5.6-luna · low (native)subscription CLI280/300264569/591591/591completed-with-missingrefusal: 22
gpt-5.6-luna · max (native)subscription CLI278/300264564/591591/591completed-with-missingrefusal: 27
gpt-5.6-luna · medium (native)subscription CLI272/300264565/591591/591completed-with-missingrefusal: 26
gpt-5.6-luna · none (native)subscription CLI300/300264591/591591/591complete
gpt-5.6-luna · xhigh (native)subscription CLI268/300264562/591591/591completed-with-missingrefusal: 29
gpt-6-astra · low (native)subscription CLI276/300264569/591591/591completed-with-missingrefusal: 22
grok-4.6 · low (native)subscription CLI300/300264591/591591/591complete
JEV 1.13.0 · TypeSafeTypeSafe Noul API300/300264591/591591/591completedinfrastructure: 1
deepseek-v4-flash-0731Fireworks API300/300264unrecordedunrecordedhistorical full
deepseek-v4-pro-0813Fireworks API300/300264unrecordedunrecordedhistorical full
glm-5p2Fireworks API300/300264unrecordedunrecordedhistorical full
gpt-oss-120bFireworks API300/300264unrecordedunrecordedhistorical full
gpt-oss-20bFireworks API300/300264unrecordedunrecordedhistorical full
kimi-k2p6Fireworks API300/300264unrecordedunrecordedhistorical full
kimi-k2p7-codeFireworks API300/300264unrecordedunrecordedhistorical full
kimi-k3Fireworks API300/300264unrecordedunrecordedhistorical full
minimax-m3Fireworks API300/300264unrecordedunrecordedhistorical full
muse-glimmer-30bFireworks API112/300unranked: incomplete historical coverageunrecordedunrecordedhistorical partial
nemotron-3-ultra-nvfp4Fireworks API300/300264unrecordedunrecordedhistorical full
nemotron-lightning-3p5-30b-a3bFireworks API300/300264unrecordedunrecordedhistorical full
qwen3p7-plusFireworks API300/300264unrecordedunrecordedhistorical full
qwen3p8-2p4t-a95bFireworks API300/300264unrecordedunrecordedhistorical full
qwen3p8-maxFireworks API300/300264unrecordedunrecordedhistorical full

Model ranking · native subscriptions

Fresh CLI sessions, one verbalized probability per proposition, tools and browsing disabled. Each reasoning level is a separate setting. No logprobs are available. Native agent scaffolds, reasoning settings and output limits differ by provider; this is not a matched API experiment. Both charts retain the same shared tuples across all participating settings.

0.000.050.100.150.20grok-4.6 · low (native)0.062glm-5.3 · native-default (native)0.064gpt-5.6-luna · xhigh (native) (268/300 original)0.067gpt-5.6-luna · max (native) (278/300 original)0.067gpt-5.6-luna · low (native) (280/300 original)0.093gpt-5.6-luna · high (native) (272/300 original)0.103gpt-5.6-luna · medium (native) (272/300 original)0.105gpt-5.6-luna · none (native)0.110gpt-6-astra · low (native) (276/300 original)0.113

Model ranking · historical Fireworks API

Original August runs, using verbalized probabilities at temperature 0. Requested reasoning was none, with low as a fallback; actual fallback choices were not recorded per proposition. Native and API results use the same tuple target and algebraic violation formulas, with different elicitation settings. Both charts retain the same shared tuples across all participating settings.

0.000.120.250.380.50kimi-k30.074nemotron-3-ultra-nvfp40.082qwen3p8-max0.085qwen3p8-2p4t-a95b0.085deepseek-v4-flash-07310.087minimax-m30.092qwen3p7-plus0.097glm-5p20.098deepseek-v4-pro-08130.106kimi-k2p7-code0.109kimi-k2p60.110gpt-oss-120b0.149gpt-oss-20b0.160nemotron-lightning-3p5-30b-a3b0.186muse-glimmer-30b (112/300 original)unranked: incomplete historical coverage

Subscription usage

Across all archived campaigns: 61,738,874 recorded native tokens and 8,753 CLI launches, including setup and interruptions. 76 interrupted launches have no native receipt; native usage is unknown for 106 launches in total. Unknown consumption is not zero. Imported observations are counted once. Quota percentages are not summed across campaigns or accounts.

Quota changes below cover native-20260913-scoped-final: 2026-09-13T04:30:48.606308+00:00 to 2026-09-13T04:55:52.120664+00:00. Account-wide deltas can include other fleet sessions; window resets and provider accounting prevent inference from tokens alone. Download complete usage report.

providerwindowbeforeafterused deltaunitmeasurement limits
OpenAI · subscription 1weekly12120percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 2weekly29312percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 3weekly56593percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 4weekly110percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 5weekly550percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 6weekly42497percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
OpenAI · subscription 7weekly24251percentage points of this subscriptionShared with other fleet sessions; whole-percent native readings. Attributable request/token counts are separate; the provider exposes no exact marginal quota conversion.
Z.ai GLM coding planTIME_LIMIT (provider unit 5 × 1)00unavailablepercentage pointsNative quota endpoint; includes excluded setup within the measurement window and any concurrent account activity. Reset timestamp changed; no delta attributed.
Z.ai GLM coding plan5-hour token window451percentage pointsNative quota endpoint; includes excluded setup within the measurement window and any concurrent account activity. Window identifiers are preserved as returned by the provider.
Z.ai GLM coding planweekly token window (provider unit 6)110percentage pointsNative quota endpoint; includes excluded setup within the measurement window and any concurrent account activity. Window identifiers are preserved as returned by the provider.
Grok SuperGrok Heavyweekly2.02.00.0percentage pointsNative /usage screen, whole-percent resolution; shared-account delta. A zero displayed change does not mean zero consumption. Same weekly reset verified.