consistency-ft
A research connection

From consistent beliefs to shared decisions.

Where could belief-consistency research fit into Open Agency Architecture? At the point where individual judgments become a model of the world.

CONSISTENCY-FT × OPEN AGENCY
Illustrated research note · 13 September 2026

From separate judgments to a shared world model Three sources of probability judgments connect to an audit of their relations. The proposed audit feeds a shared world model, which informs decisions alongside separately supplied human values. A POSSIBLE CONNECTION Judgment AP(success)Judgment BP(failure)Judgment CP(A ∩ B) Check the relations PROPOSED AUDIT model Shared decisionsforecasts + human values values JUDGMENTS → MODEL → CHOICE
Our proposed entry point: audit the inputs to a shared model.
Policy verification and human values remain separate problems.
01 / Where the work meets

A small component.
A consequential place.

The Open Agency roadmap breaks collective decision-making into modular problems. Consistency research could help check the forecasts that connect possible actions to their consequences.

OAA PROJECTS 1–2

What could we do?

Define feasible actions and constraints with formal descriptions or templates.

OAA PROJECTS 3, 4 & 6

What might happen?

Elicit expert predictions and represent uncertainty in a world model.

OAA PROJECTS 8–9

What should we choose?

Combine stakeholder preferences with an explicit choice mechanism.

↑ Human preferences are a separate input.
Proposed consistency layer: preserve judgments → check relations → clarify or repair → retain both versions.
A simplified path through the roadmap, not a complete OAA implementation. Its minimal preference prototype can also evaluate actions directly without a predictive world model.
Direct connection · project 3

Better elicitation inputs

Our probability readouts and ten checks could audit judgments before integration. The roadmap also proposes sliders and betting-odds interfaces for experts.

Next boundary · projects 4 & 6

From numbers to models

Compare judgments with a program's implied probabilities. Squiggle-to-PRISM/JANI translation, discretization and semantic fidelity need separate validation.

Later questions · projects 5, 8 & 9

Updating and choosing

Test whether better inputs improve sequential forecasts and decisions. Verified RL, preference elicitation and Nash bargaining are further components.

A finite audit has a finite scope. Passing local checks does not establish global world-model validity. Project 7's cross-scale models and continuous-time coalgebras remain outside what this repository implements.

02 / A probability playground

Two answers.
One logical relation.

A project either succeeds or fails, under one shared definition. Its success and failure probabilities should add to 100%. Move the two judgments to see the gap.

Illustrative inputs · not model results
Total probability120%
Negation violation0.20

Sell both claims for $1.20 and pay $1.00 in either outcome: $0.20 guaranteed profit in this toy contract.

Coherent complement probabilities lie on a diagonal lineSuccess probability is the horizontal axis; failure probability is the vertical axis. The blue line contains all pairs summing to one. The amber point marks the entered judgments. P(success) + P(failure) = 1 00.5100.51Probability of successProbability of failure
Success and failure are exact complements. The line describes logical compatibility; it tells us nothing about which forecast is accurate.

A coherent forecast can still be confidently wrong.

The toy contract assumes frictionless, bounded stakes and $1/$0 settlement. Its Negation score is |p + q − 1|. The repair minimizes the sum of Bernoulli KL distances to the two inputs, clipped at 10⁻⁶. It is a numerical illustration, not a trained-model result. Metric source ↗

03 / Evidence and open work

We have an audit.
The bridge is a hypothesis.

The repository measures inconsistencies in elicited forecasts. We still need to test whether reducing them improves predictions or collective decisions.

Implemented
10

Relation checks

Negation, paraphrase, implication, conjunctions, conditionals and total probability.

Check definitions ↗
Available
3,171

Benchmark tuples

Natural-language tuples from the Paleka corpus. Exact-logic data construction is proposed.

Tuple corpus ↗
Implemented
LP

Exploitability metric

A bounded-stake adversary measures guaranteed profit over a check's allowed outcomes.

LP implementation ↗
Training prepared
4

Held-out check families

Episode splits support a future transfer test. They are not a trained checkpoint.

Episode builder ↗

The live model ranking uses algebraic violation, a different metric from the LP. Readout, reasoning settings and missing responses affect comparisons. Lower violation does not establish calibration or forecasting accuracy.

Proposed research direction

Correct the contradiction.
Keep the information.

A KL anchor penalizes changes to reference beliefs while consistency pressure reduces contradictions. It can work on unresolved questions, but it can also preserve errors. Repeatedly moving the reference may allow drift.

The proposed loss balances consistency pressure with a belief anchorArrows from consistency pressure and reference beliefs meet at a revised belief. This is a tradeoff to test, not an established improvement.qcoherencepressurereferencebeliefs
Exploit(q, S) + λ · mean KL(q ∥ pref)

Here KL compares Bernoulli distributions over reported beliefs, rather than token policies. The full note specifies the candidate loss and limitations. Neither the roadmap nor this page establishes that the loss improves OAA policy selection.

04 / The next test

Does a better input
make a better decision?

Start with a controlled funding decision whose true consequences are known. Keep stakeholder utilities and the choice rule fixed, and vary only the forecasts.

Proposed · not yet run

24 synthetic cases

A tiny Bayesian world for each case, with explicit action effects and an exact-posterior oracle.

Twenty-four proposed synthetic casesFour rows of six outlined circles represent the proposed cases. These are planned cases, not completed results.
2funding choices3model settings1fixed rule
Three arms · identical recorded inputs

Raw forecasts

Preserve the original judgments, including contradictions and missing answers.

KL-minimal repair

Fit one coherent finite joint distribution while limiting changes to the original inputs.

Uniform-joint baseline

A coherent but uninformative comparison that exposes trivial improvements.

What would count as progress?

Lower decision regret, without worse predictive scores or hidden missing responses.

Lower inconsistency on fitted relations is expected by construction. The 24-case pilot could establish feasibility, not real-world effectiveness.

  1. Freeze the worldPublish seeds, action semantics, logical relations and synthetic utilities.
  2. Record judgmentsUse the same evidence, independent contexts and fixed model/effort settings.
  3. Compare the armsKeep the decision expression fixed. Mix repaired joints before deriving conditionals.
  4. Measure consequencesTrack regret, posterior error, expected Brier loss, coverage, repair size and cost.

Shared evidence is part of the experiment. Experts with different information can reasonably disagree. Outside this controlled setting, match event definitions, evidence and timestamps before calling judgments inconsistent. Agreement among related models is not independent evidence.

05 / Follow the sources

An entry point into
a larger research program.

The Open Agency roadmap proposes an iterative, modular, human-in-the-loop approach to institutional decision-making. This page illustrates one possible connection at its world-model input.

Futarchy's markets, OAA's proposed Nash bargaining, and consistency metrics address different pieces of the problem. A coherent belief system can still be mistaken—or deliberately fabricated.

Explore the futarchy bridge →