A set of things, a set of attributes, a model. Each cell is the model's fitted score with its uncertainty; each column header says, in numbers, how much to trust that column. Click a cell for the pairwise evidence behind it. Run your own set at /judge; every column here is a public run.
score — fitted latent on the log-ratio scale, gauge-fixed per run (not portable across runs). Bars share a column's scale; the whisker is the interval at the chosen coverage. Rank and percentile are standing within this cohort only.
swap stability — every pair is asked twice with A and B swapped. The share that keeps the same winner; 50% is a coin flip (the model reads layout, not content). |Δ ln ratio| is how much the magnitude moved.
separation — of all entity pairs, how many have non-overlapping intervals at this coverage. A column that separates 5 of 780 pairs ranks almost nothing; the ordering is suggestive, not settled.
budget — comparisons spent vs allowed, why the run stopped, and the top-k error estimate at stop. budget_exhausted means: read top-k error before treating the top as settled.
run agreement — when several runs judged the same attribute on the same entities, rank correlation (Spearman ρ, Kendall τ) and top-5 overlap between them. Two runs of one model at temperature 0 should be ρ=1; if they are not, the attribute is not being read stably.
refusals / cost — counted per run, never hidden. Cost is provider spend for the whole column. Full measurement spec: /methods.