Correction · 18 Jul 2026
Majority voting withdrawn from the annotation regime
It only works if errors are independent, and on ambiguous items they are not.
Specifics
The numbers, and where they come from
| Quantity | Value | Basis |
|---|---|---|
| Items resolved under v1 | ~6,000 | Counted |
| Triple-annotated share | 20% | Stated |
| Seeded gold share | 5% | Stated |
| Added programme hours | ~60 | Estimated |
The error
Why the original design was wrong
Regime v1 resolved disagreement between annotators by majority vote. That is only sound when annotator errors are independent of one another. On the items that actually matter, the ambiguous ones, they are not: annotators tend to fail in the same direction, because the ambiguity has a shape and most readers resolve it the same wrong way.
The consequence is that majority voting is most confident exactly where it is least reliable. A three-to-nothing vote on an ambiguous item looks like strong agreement and is in fact correlated error.
The replacement
Regime v2
- 01
Triple annotation on 20%
A fifth of all items go to three annotators, with Fleiss’ kappa reported per batch rather than in aggregate, so a bad batch cannot be averaged away.
- 02
Expert adjudication
Disagreements, and every item marked unsure, go to an adjudicator rather than to a vote. Slower and more expensive, and the only version that is sound.
- 03
Five per cent gold
Seeded items with known answers, scored continuously, so annotator drift is visible while it is happening rather than at the end.
- 04
Provenance per item
Every entry records whether it was WordNet-seeded, model-generated or human-verified, plus its agreement score. Soundness with respect to a knowledge base nobody checked is not soundness.
Status
What this changes for work already done
Roughly 6,000 items were resolved under v1. They are being re-adjudicated rather than accepted, and until that finishes those items carry a provenance flag marking them as v1-resolved. The flag is visible in the data, not just in this note.
Cost of the correction
Adjudication is slower than voting. This adds an estimated 60 hours to the programme and pushes the knowledge base milestone later. That is the correct trade and we are recording the cost rather than absorbing it quietly.