Event workspace
Confidence
published
What these duels establish
2 of 4 adjacent pairs are separated by this evidence. First place is separated from second.
An adjacent pair is two projects next to each other in the published order. A pair counts as separated when fewer than 2.50% of the resampled panels put them the other way round. Where a pair is not separated, the place between them was decided by the sort and not by the judging, and awarding a prize across that boundary is a choice somebody is making rather than a result these duels produced.
- Levels the duels support
- 3 across 5 compared projects
- Another panel would agree
- 91.4% of pairs, on average
- Another panel would reproduce this exact order
- 59.8%
Agreement is Kendall's tau-b between each resampled panel's ranking and the published one, averaged: 1 is every resample agreeing about every pair and 0 is a coin toss. The line under it is the strict reading - the share of resamples that got every place right - and it is 0 on any field of more than a handful of projects, because a long table has a lot of places to get wrong. Read the first number; the second is there so the first cannot be mistaken for it.
| Higher | Lower | Gap | Reversed in | Verdict |
|---|---|---|---|---|
| Readback | Highcontrast | 1.82 | 0.0% | separated |
| Highcontrast | Lintwright | 0.02 | 36.8% | not separated indistinguishable on this evidence |
| Lintwright | Portmatic | 1.45 | 0.0% | separated |
| Portmatic | Stackless | 1.59 | 3.5% | not separated indistinguishable on this evidence |
Resample with different settings
More resamples resolve a finer tail and cost more time; a higher confidence level demands more evidence before it calls one project ahead of another, and therefore separates fewer pairs. Neither changes a single score - the ranking above is the same fit either way, and what moves is only how much of it this page is willing to call established.
What qualifies these numbers
The resample reported nothing worth qualifying.
| How serious | What the resample noticed | Code |
|---|---|---|
| note | The comparisons resolve 3 tiers among 5 compared projects. A tier holds the projects that are not separated from the one at the top of it, so the tier leader is the defensible unit to award on. It is not a claim that every project in a tier ties with every other — read `pairs` for that. | bootstrap.tiers |
Levels
A level is formed around one project: the one named first below, and then every project beneath it in the ranking that these duels do not separate from it. Award a prize to the top of a level and no project on that level can point at these comparisons and say the order between the two of them was established.
It is not a claim that everything on a level ties with everything else. Two projects sharing a level can still be separated from each other - both stay close enough to the project it was formed around to be held with it, while the evidence between them is clear. The pair table above is where a boundary is established, and it is the only place on this page that establishes one.
- Level 1
- Readback
- Level 2
- Highcontrast - with Lintwright and Portmatic, none of which these duels separate from it
- Level 3
- Stackless
The ranking, with its intervals
A strength is not a score out of anything - it is the number that best explains who beat whom, and it means nothing except against the others in this table. The range beside it is where the resamples put it; "could place" is the span of places it took across them, "holds this place" is how often it kept the one printed here, and "comes first" is how often it topped the whole field. A project no duel mentions has no evidence about it at all and is listed without a place rather than at the bottom.
| # | Level | Project | Track | Strength | 95% range | Could place | Holds this place | Comes first | Duels |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 1 | Readback | access | 2.37 | 1.86 – 2.91 | 1 – 1 | 100% | 100% | 3 |
| 2 | 2 | Highcontrast | access | 0.55 | -0.81 – 1.46 | 2 – 4 | 63% | 0% | 6 |
| 3 | 2 | Lintwright | tooling | 0.53 | -0.12 – 0.97 | 2 – 3 | 63% | 0% | 6 |
| 4 | 2 | Portmatic | tooling | -0.93 | -2.21 – 0.50 | 3 – 5 | 94% | 0% | 5 |
| 5 | 3 | Stackless | tooling | -2.52 | -2.91 – -1.00 | 4 – 5 | 97% | 0% | 4 |
How this was computed
- Method
- pairwise-bootstrap(judge, 400, prior=1)
- Resampling unit
- the judge - a resample draws whole judges, so the question is whether another panel would agree
- Resamples
- 400
- Reached tolerance
- 400 of 400
- Graph in one piece
- 400 of 400
- Finest share resolved
- 0.25%
- Reversal allowed
- up to 2.50%
- Seed
manak.bootstrap|01M3GR05HGYBMVHY82GGQZPF4H
The seed is published so the intervals can be re-derived rather than trusted: the same duels and the same settings give the same numbers on any machine, which is also why this page does not change its mind when somebody reloads it.
Standings · the same figures as JSON: this page's API route