Lab

A set of things, a set of attributes, a model. Each cell is the model's fitted score with its uncertainty; each column header says, in numbers, how much to trust that column. Click a cell for the pairwise evidence behind it. Run your own set at /judge; every column here is a public run.

score — fitted latent on the log-ratio scale, gauge-fixed per run (not portable across runs). Bars share a column's scale; the whisker is the interval at the chosen coverage. Rank and percentile are standing within this cohort only.

swap stability — every pair is asked twice with A and B swapped. The share that keeps the same winner; 50% is a coin flip (the model reads layout, not content). |Δ ln ratio| is how much the magnitude moved.

separation — of all entity pairs, how many have non-overlapping intervals at this coverage. A column that separates 5 of 780 pairs ranks almost nothing; the ordering is suggestive, not settled.

budget — comparisons spent vs allowed, why the run stopped, and the top-k error estimate at stop. budget_exhausted means: read top-k error before treating the top as settled.

run agreement — when several runs judged the same attribute on the same entities, rank correlation (Spearman ρ, Kendall τ) and top-5 overlap between them. Two runs of one model at temperature 0 should be ρ=1; if they are not, the attribute is not being read stably.

refusals / cost — counted per run, never hidden. Cost is provider spend for the whole column. Full measurement spec: /methods.