OVOS Plugin Arena · Completeness
Evidence
This page is for grant reviewers and auditors who need to check the arena's completeness rather than watch a battle or read a leaderboard.
This page is generated from the same data the site publishes — no hand-typed numbers. It exists so an outside reviewer can check, league by league, what is actually in place versus what is still a gap: how many fighters are registered, how many datasets, how many of those datasets have publishedpredictions on Hugging Face, and how many benchmark boards and ELO leaderboards exist as a result. Where a league has fighters but few or no published predictions, that shows up as a gap here — not a rounded-up number.
Loading evidence…
Try another league or language. No evidence data yet — run arena.cli export-evidence (wired into the assemble.yml action) to generate data/evidence.json.
Per league
Fighter coverage, per league
Dataset-level coverage above can hide fighter-level gaps: a league can have a published benchmark board and still have registered fighters with zero rows on it. This table counts, per league, how many registered fighters actually appear in at least one benchmark board or battles pool ("on boards") versus how many do not yet ("ghosts"). A ghost is a registered-but-not-yet-benchmarked fighter — most reflect sweeps still in progress, not a hidden failure. This is computed offline from the committed data dir only; it cannot see in-flight Hugging Face uploads.
Published predictions, per dataset
Every dataset a league claims to score against, linked to its Hugging Face predictions repo. marks a dataset that is registered but has no published predictions yet — no benchmark board, no ELO seed for it.
Verify it yourself