OVOS Plugin Arena · Completeness
Evidence
This page is generated from the same data the site publishes — no hand-typed
numbers. It exists so an outside reviewer can check, league by league, what
is actually in place versus what is still a gap: how many fighters are
registered, how many datasets, how many of those datasets have published
predictions on Hugging Face, and how many benchmark boards and ELO
leaderboards exist as a result. Where a league has fighters but few or no
published predictions, that shows up as a gap here — not a rounded-up number.
Loading evidence…
Try another league or language. No evidence data yet — run arena.cli export-evidence (wired into
the assemble.yml action) to generate data/evidence.json.
Per league
Fighter coverage, per league
Dataset-level coverage above can hide fighter-level gaps: a league can have
a published benchmark board and still have registered fighters with zero
rows on it. This table counts, per league, how many registered fighters
actually appear in at least one benchmark board or battles pool
("on boards") versus how many do not yet ("ghosts"). A ghost is a
registered-but-not-yet-benchmarked fighter — most reflect sweeps still in
progress, not a hidden failure. This is computed offline from the
committed data dir only; it cannot see in-flight Hugging Face uploads.
Published predictions, per dataset
Every dataset a league claims to score against, linked to its Hugging Face
predictions repo. marks
a dataset that is registered but has no published predictions yet — no
benchmark board, no ELO seed for it.
Verify it yourself