GPT-6 Astra and GPT-5.6 Luna, frozen, trade a combinatorial prediction market over Turtel et al.'s 1,265 released test questions. Each session forecasts eight questions at once and must emit valid trades against the bayes-market engine. Nothing is trained. The question is whether repeated, coherence-enforced trading beats the model's first pass.
Brier score of the market's marginal probabilities against the known resolutions, all 1,265 questions, after each pass. The table also gives the joint log loss, −log P(ω)/n under the whole junction tree by the chain rule, which is what the LMSR settles on and the only number that credits the dependencies the model builds. Dashed lines are Turtel's published numbers on the same questions: his base model one-shot, o1, his RL-trained ReMax, and the Polymarket price at prediction time.
Contamination caveat. The test questions closed in February and March 2025, inside the training window of both models. The prompt orders the model to answer as of each question's date and to use no later knowledge, and we log everything, but the absolute Brier of a frontier model here is not a clean forecasting number. The within-run comparison, pass 1 against pass 3, and the trade-validity rate are the clean results; Turtel's rows are context.
Total KL divergence moved across a session's trades, in nats, by pass. Pass 1 is uncapped; later passes are budgeted at 0.3 nats per question.
Market probability of twenty random questions after each pass, coloured by resolution. Convergence toward the resolved side is the mechanism working; oscillation is herding.
One randomly chosen session from the latest pass: the market state it saw, its rationale, and the trades that were executed or rejected. These traces are the raw material for later distillation into a 14B model.
no sessions yet