Better elicitation inputs
Our probability readouts and ten checks could audit judgments before integration. The roadmap also proposes sliders and betting-odds interfaces for experts.
Where could belief-consistency research fit into Open Agency Architecture? At the point where individual judgments become a model of the world.
The Open Agency roadmap breaks collective decision-making into modular problems. Consistency research could help check the forecasts that connect possible actions to their consequences.
Define feasible actions and constraints with formal descriptions or templates.
Elicit expert predictions and represent uncertainty in a world model.
Combine stakeholder preferences with an explicit choice mechanism.
Our probability readouts and ten checks could audit judgments before integration. The roadmap also proposes sliders and betting-odds interfaces for experts.
Compare judgments with a program's implied probabilities. Squiggle-to-PRISM/JANI translation, discretization and semantic fidelity need separate validation.
Test whether better inputs improve sequential forecasts and decisions. Verified RL, preference elicitation and Nash bargaining are further components.
A finite audit has a finite scope. Passing local checks does not establish global world-model validity. Project 7's cross-scale models and continuous-time coalgebras remain outside what this repository implements.
A project either succeeds or fails, under one shared definition. Its success and failure probabilities should add to 100%. Move the two judgments to see the gap.
Sell both claims for $1.20 and pay $1.00 in either outcome: $0.20 guaranteed profit in this toy contract.
A coherent forecast can still be confidently wrong.
The toy contract assumes frictionless, bounded stakes and $1/$0 settlement. Its Negation score is |p + q − 1|. The repair minimizes the sum of Bernoulli KL distances to the two inputs, clipped at 10⁻⁶. It is a numerical illustration, not a trained-model result. Metric source ↗
The repository measures inconsistencies in elicited forecasts. We still need to test whether reducing them improves predictions or collective decisions.
Negation, paraphrase, implication, conjunctions, conditionals and total probability.
Check definitions ↗Natural-language tuples from the Paleka corpus. Exact-logic data construction is proposed.
Tuple corpus ↗A bounded-stake adversary measures guaranteed profit over a check's allowed outcomes.
LP implementation ↗Episode splits support a future transfer test. They are not a trained checkpoint.
Episode builder ↗The live model ranking uses algebraic violation, a different metric from the LP. Readout, reasoning settings and missing responses affect comparisons. Lower violation does not establish calibration or forecasting accuracy.
A KL anchor penalizes changes to reference beliefs while consistency pressure reduces contradictions. It can work on unresolved questions, but it can also preserve errors. Repeatedly moving the reference may allow drift.
Here KL compares Bernoulli distributions over reported beliefs, rather than token policies. The full note specifies the candidate loss and limitations. Neither the roadmap nor this page establishes that the loss improves OAA policy selection.
Start with a controlled funding decision whose true consequences are known. Keep stakeholder utilities and the choice rule fixed, and vary only the forecasts.
A tiny Bayesian world for each case, with explicit action effects and an exact-posterior oracle.
Preserve the original judgments, including contradictions and missing answers.
Fit one coherent finite joint distribution while limiting changes to the original inputs.
A coherent but uninformative comparison that exposes trivial improvements.
Lower decision regret, without worse predictive scores or hidden missing responses.
Lower inconsistency on fitted relations is expected by construction. The 24-case pilot could establish feasibility, not real-world effectiveness.
Shared evidence is part of the experiment. Experts with different information can reasonably disagree. Outside this controlled setting, match event definitions, evidence and timestamps before calling judgments inconsistent. Agreement among related models is not independent evidence.
The Open Agency roadmap proposes an iterative, modular, human-in-the-loop approach to institutional decision-making. This page illustrates one possible connection at its world-model input.
Futarchy's markets, OAA's proposed Nash bargaining, and consistency metrics address different pieces of the problem. A coherent belief system can still be mistaken—or deliberately fabricated.
Published by Deger Turan. Written with Tantum Collins, davidad, and Charbel-Raphael Segerie. Reviewed by Andrew Critch, Ben Goldhaber, Ozzie Gooen, and Evan Miyazono.
Detailed mappings, evidence links, distinctions and the proposed pilot protocol. The original repository review is pinned to commit 46eed97.