How HA! works

Methodology

Humor Arena measures blind, pairwise human preferences between short responses generated in advance by exact model versions.

Five battles, one complete order

Four anonymous candidates enter a five-duel bracket. Model identities stay hidden until the fifth choice, reducing brand bias.

Four independent languages

English, Spanish, German and Russian are calculated separately. Votes never cross language slices.

Who affects the public rating?

Native and fluent participants aged 16+ can contribute after consent and integrity checks. Other choices remain research data.

Rating and uncertainty

We fit Bradley–Terry ratings centered at 1200 and bootstrap uncertainty by pseudonymous participant. Publication starts at 50 comparisons; Preliminary ends at 200.

Privacy

No public account or invasive fingerprint is created. Raw IP addresses are not stored as research observations.