manak.

Nothing is open. No further changes are scheduled.

Event workspace

Confidence

published

5projects compared
12duels decided
3judges resampled
400resamples
3levels
2 of 4pairs separated

What these duels establish

2 of 4 adjacent pairs are separated by this evidence. First place is separated from second.

An adjacent pair is two projects next to each other in the published order. A pair counts as separated when fewer than 2.50% of the resampled panels put them the other way round. Where a pair is not separated, the place between them was decided by the sort and not by the judging, and awarding a prize across that boundary is a choice somebody is making rather than a result these duels produced.

Levels the duels support
3 across 5 compared projects
Another panel would agree
91.4% of pairs, on average
Another panel would reproduce this exact order
59.8%

Agreement is Kendall's tau-b between each resampled panel's ranking and the published one, averaged: 1 is every resample agreeing about every pair and 0 is a coin toss. The line under it is the strict reading - the share of resamples that got every place right - and it is 0 on any field of more than a handful of projects, because a long table has a lot of places to get wrong. Read the first number; the second is there so the first cannot be mistaken for it.

HigherLowerGapReversed inVerdict
ReadbackHighcontrast1.820.0%separated
HighcontrastLintwright0.0236.8%not separated indistinguishable on this evidence
LintwrightPortmatic1.450.0%separated
PortmaticStackless1.593.5%not separated indistinguishable on this evidence
Resample with different settings

More resamples resolve a finer tail and cost more time; a higher confidence level demands more evidence before it calls one project ahead of another, and therefore separates fewer pairs. Neither changes a single score - the ranking above is the same fit either way, and what moves is only how much of it this page is willing to call established.

How many times to refit the ranking on resampled data. More is finer and slower; the default resolves a quarter of a percentage point.

How much of the resampled spread each interval should cover. A higher level separates fewer pairs, because it demands more evidence to call one project ahead of another.

What qualifies these numbers

The resample reported nothing worth qualifying.

How seriousWhat the resample noticedCode
noteThe comparisons resolve 3 tiers among 5 compared projects. A tier holds the projects that are not separated from the one at the top of it, so the tier leader is the defensible unit to award on. It is not a claim that every project in a tier ties with every other — read `pairs` for that.bootstrap.tiers

Levels

A level is formed around one project: the one named first below, and then every project beneath it in the ranking that these duels do not separate from it. Award a prize to the top of a level and no project on that level can point at these comparisons and say the order between the two of them was established.

It is not a claim that everything on a level ties with everything else. Two projects sharing a level can still be separated from each other - both stay close enough to the project it was formed around to be held with it, while the evidence between them is clear. The pair table above is where a boundary is established, and it is the only place on this page that establishes one.

Level 1
Readback
Level 2
Highcontrast - with Lintwright and Portmatic, none of which these duels separate from it
Level 3
Stackless

The ranking, with its intervals

A strength is not a score out of anything - it is the number that best explains who beat whom, and it means nothing except against the others in this table. The range beside it is where the resamples put it; "could place" is the span of places it took across them, "holds this place" is how often it kept the one printed here, and "comes first" is how often it topped the whole field. A project no duel mentions has no evidence about it at all and is listed without a place rather than at the bottom.

#LevelProjectTrackStrength95% rangeCould placeHolds this placeComes firstDuels
11Readbackaccess2.371.86 – 2.911 – 1100%100%3
22Highcontrastaccess0.55-0.81 – 1.462 – 463%0%6
32Lintwrighttooling0.53-0.12 – 0.972 – 363%0%6
42Portmatictooling-0.93-2.21 – 0.503 – 594%0%5
53Stacklesstooling-2.52-2.91 – -1.004 – 597%0%4

How this was computed

Method
pairwise-bootstrap(judge, 400, prior=1)
Resampling unit
the judge - a resample draws whole judges, so the question is whether another panel would agree
Resamples
400
Reached tolerance
400 of 400
Graph in one piece
400 of 400
Finest share resolved
0.25%
Reversal allowed
up to 2.50%
Seed
manak.bootstrap|01M3GR05HGYBMVHY82GGQZPF4H

The seed is published so the intervals can be re-derived rather than trusted: the same duels and the same settings give the same numbers on any machine, which is also why this page does not change its mind when somebody reloads it.

Standings · the same figures as JSON: this page's API route