Correction · 18 Jul 2026

Majority voting withdrawn from the annotation regime

It only works if errors are independent, and on ambiguous items they are not.

Type
Correction
Regime
v2
Replaced
Majority vote
Now
Adjudication

Specifics

The numbers, and where they come from

Figures on this page, with the basis of each
QuantityValueBasis
Items resolved under v1~6,000Counted
Triple-annotated share20%Stated
Seeded gold share5%Stated
Added programme hours~60Estimated
vote_resultwithdrawn from the regimeannotatorsthree, on 20% of itemsadjudicated_bynamed expert, not a tallykappa_batchFleiss' kappa, per batchprovenanceseeded, generated or verified
What an item now carries after regime v2. The vote field is gone rather than left empty.

The error

Why the original design was wrong

Regime v1 resolved disagreement between annotators by majority vote. That is only sound when annotator errors are independent of one another. On the items that actually matter, the ambiguous ones, they are not: annotators tend to fail in the same direction, because the ambiguity has a shape and most readers resolve it the same wrong way.

The consequence is that majority voting is most confident exactly where it is least reliable. A three-to-nothing vote on an ambiguous item looks like strong agreement and is in fact correlated error.

The replacement

Regime v2

  • 01

    Triple annotation on 20%

    A fifth of all items go to three annotators, with Fleiss’ kappa reported per batch rather than in aggregate, so a bad batch cannot be averaged away.

  • 02

    Expert adjudication

    Disagreements, and every item marked unsure, go to an adjudicator rather than to a vote. Slower and more expensive, and the only version that is sound.

  • 03

    Five per cent gold

    Seeded items with known answers, scored continuously, so annotator drift is visible while it is happening rather than at the end.

  • 04

    Provenance per item

    Every entry records whether it was WordNet-seeded, model-generated or human-verified, plus its agreement score. Soundness with respect to a knowledge base nobody checked is not soundness.

Status

What this changes for work already done

Roughly 6,000 items were resolved under v1. They are being re-adjudicated rather than accepted, and until that finishes those items carry a provenance flag marking them as v1-resolved. The flag is visible in the data, not just in this note.

Cost of the correction

Adjudication is slower than voting. This adds an estimated 60 hours to the programme and pushes the knowledge base milestone later. That is the correct trade and we are recording the cost rather than absorbing it quietly.