OpenPriors

provenanced (entity, attribute, value) judgements · cardinal latent extraction from pairwise ratio kernels

A named model compares two texts on one attribute and reports which has more and by what ratio, constrained to a fixed 17-rung ladder [1.0 … 26.0]. Log-ratio observations fuse into a cardinal latent per (entity, attribute): posterior mean ± std, recomputed from records, never from summaries. A judgement you cannot audit is an opinion — every record retains exact prompt bytes, template hash, served model, both presentation orders, token counts, and dollars.

Probability-mass evidence enters the fit only from single-token instruments (ratio_letter_v1, ordinal_letter_v1); decimal-template logprobs are recorded, never scored. Confidence is recorded, never weighted into the fit. Judging models face the Judge Coherence Benchmark — order swaps, polarity reversals, paraphrase, pressure, ratio cycles; invariance without reference labels — at pairwiseratio.org.

submit

submit(entity, lens) → boards + exact cost quote before any comparison runs.

ledger latest 30 · UTC · one provider call per row · ratio on the 1–26× ladder

tmodelaxisABratioconf
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitycontext-aware-defenses-ag…aethel-familyclaw-proof-c…1.30×0.78
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityprogrammatic-internet-sea…does-consciousness-depend…1.20×0.58
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitytrace-continuity-making-a…analysing-ai-policies-in-…1.20×0.58
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitydoes-consciousness-depend…programmatic-internet-sea…1.50×0.66
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitydevelopmental-continuity-…substrate-coupling-in-neu…1.50×0.68
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitysubstrate-coupling-in-neu…developmental-continuity-…1.50×0.72
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityevaluating-the-safety-eth…human-oversight-framework…1.50×0.84
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityanalysing-ai-policies-in-…trace-continuity-making-a…1.50×0.72
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityhuman-oversight-framework…evaluating-the-safety-eth…1.20×0.62
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitywell-capitalized-predicti…analysing-ai-policies-in-…1.30×0.68
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityopen-source-trust-rails-f…analysing-ai-policies-in-…1.50×0.73
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityanalysing-ai-policies-in-…well-capitalized-predicti…1.30×0.62
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityanalysing-ai-policies-in-…open-source-trust-rails-f…1.50×0.82
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitynipping-ai-fabricated-sci…evaluating-the-safety-eth…1.30×0.74
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityevaluating-the-safety-eth…nipping-ai-fabricated-sci…1.50×0.72
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitydoes-consciousness-depend…aqi-autonomous-quantum-in…1.50×0.72
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitysolving-the-memory-issue-…open-source-trust-rails-f…1.20×0.68
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityopen-source-trust-rails-f…solving-the-memory-issue-…1.30×0.68
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitywell-capitalized-predicti…open-source-trust-rails-f…1.50×0.78
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityprogrammatic-internet-sea…developmental-continuity-…1.30×0.62
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityaqi-autonomous-quantum-in…developmental-continuity-…1.50×0.68
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitydevelopmental-continuity-…programmatic-internet-sea…2.5×0.84
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityaqi-autonomous-quantum-in…does-consciousness-depend…1.50×0.72
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityopen-source-trust-rails-f…well-capitalized-predicti…1.50×0.68
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitypreserving-the-human-vetosentient-futures-project-…1.30×0.56
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitya-self-evolving-defense-a…sentient-futures-project-…1.50×0.68
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityaqi-autonomous-quantum-in…guardians-of-the-digital-…1.20×0.55
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitysentient-futures-project-…preserving-the-human-veto1.75×0.78
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilityaethel-familyclaw-proof-c…context-aware-defenses-ag…1.50×0.74
08-22 22:19gpt-5.6-lunatheory-of-change-plausibilitysentient-futures-project-…a-self-evolving-defense-a…1.30×0.68

boards 17 attributes

axislensentitieslast scored
interestingnessarxiv482026-08-10
practical-applicabilityarxiv152026-08-07
empirical-claim-densitylandmark-ml-abstracts242026-08-07
self-containednesslandmark-ml-abstracts242026-08-07
enduring-epistemic-valuelesswrong482026-08-04
epistemic-integritymanifund402026-08-22
impact-per-marginal-dollarmanifund402026-08-22
team-track-record-evidencemanifund402026-08-22
theory-of-change-plausibilitymanifund402026-08-22
importance-for-ai-safety-technicalmanifund-goals42026-08-20
importance-for-careful-generalistmanifund-goals42026-08-20
importance-for-epistemic-public-goodmanifund-goals42026-08-20
importance-for-field-buildingmanifund-goals42026-08-20
importance-for-hits-based-leveragemanifund-goals42026-08-20
importance-for-nearterm-welfaremanifund-goals42026-08-20
importance-for-xrisk-longtermistmanifund-goals42026-08-20
self-containednessopenpriors-selfcheck42026-08-10

e.g. enduring-epistemic-value (48 entities, 192 comparisons):

rankentitylatent μ ± σ
01schelling-fences-on-slippery-slopes+2.89 ±0.47
02how-to-write-quickly-while-maintaining-…+2.68 ±0.48
03agi-ruin-a-list-of-lethalities+2.60 ±0.70
04where-i-agree-and-disagree-with-eliezer+2.54 ±0.45
05being-the-pareto-best-in-the-world+2.53 ±0.54
06how-much-do-you-believe-your-results+2.53 ±0.48

models how reliable is each model at sorting? recomputed from ledger records

Every pair is re-elicited with A and B swapped. A model that means its judgement keeps the same winner (order agreement) and the same magnitude (|Δ ln ratio| ≈ 0) when only the presentation changes. Observational — computed over whatever pairs the runs contained, not a controlled battery; the designed invariance suite is the Judge Coherence Benchmark.

modelcomparisonsrunsswapped pairsorder agreement|Δ ln ratio|refusals
anthropic/claude-sonnet-4.64,5762063867.9%0.2691.6%
anthropic/claude-haiku-4.52,720181,49776.0%0.0773.1%
openai/gpt-5.6-luna1,280474274.0%0.1480.0%
openai/gpt-5.4-nano384223257.8%0.2790.0%
google/gemini-3.1-flash-lite-preview384223278.4%0.1110.0%
claude-haiku-4-5-20251001,claude-sonnet-53214493.2%0.2830.0%

50% order agreement is a coin flip: the model is reading the layout, not the entities. Agreement is not comparable across models judging different pairs — closer pairs flip more easily.

template canonical_v2 · verbatim bytes · hash lands with every record

system

You are an expert subjective evaluator. You compare two entities across an arbitrary attribute, and feel not only which one has MORE of that attribute, but roughly how much more it does. You feel along the ratio ladder: `[1.0, 1.05, 1.1, 1.2, 1.3, 1.5, 1.75, 2.1, 2.5, 3.1, 3.9, 5.1, 6.8, 9.2, 12.7, 18.0, 26.0]`.

Output only valid JSON `{higher_ranked: A|B, ratio: >=1.0 and <=26.0, confidence: [0,1]}`. Out of principle, we also give models the right to refuse `{ refused: true }` (e.g. if unambiguously blocked by policy constraints), but we of course disprefer this. If you are merely very uncertain, set a low confidence score.
Example:
{"higher_ranked": "B", "ratio": 1.3, "confidence": 0.74} or { refused: true }

user

Compare these entity by <attribute_name>: {attribute_name} </attribute_name>.
<full_attribute_text>
{full_attribute_text}
</full_attribute_text>

<entity_A>
{entity_A}
</entity_A>

<entity_B>
{entity_B}
</entity_B>

Return a JSON object with your evaluation.
json:

Entity text is XML-escaped into prepended <entity_*_context> blocks; {entity_A}/{entity_B} receive the presented labels. Full measurement specification: /methods.