Global Leaderboard
Category weights
◈ Explore the data category × difficulty · where each model holds up and where it cracks
GOD MODE BENCH
The hardest class, exclusively — god sentinels (atomic all-or-nothing), god agentic tasks through every harness, and god-tier code generation in the arena. Beyond-frontier by construction; most models should fail most of it.
How this works
A GOD MODE bench run (Run tab → Test plan → GOD MODE BENCH) serves the model through the same verified HF-pull flow, then runs ONLY the god tier: every god sentinel (scored by atomic checkers — a 90%-right answer scores 0), the god agentic tasks through Hermes / OpenClaw / OpenCode, and god-tier arena challenges (those artifacts land in the Code Gallery for human votes). GOD SCORE = 0.6 × sentinels + 0.4 × agentic, renormalized when a component is untested. Attested runs only; partial god passes never rank.
Performance
TTFT · TPOT · tok/s across the concurrency ladder, per prompt category — same 64k serve as the quality boards.
How this works
Every comprehensive benchmark also runs a performance grid: the served model answers fixed prompts from each category (Math, Coding, Reasoning, Instruction, Prose — short and long-prefill) at concurrency 1 / 4 / 8 / 16 / 32 — each category swept in isolation (c4 = four concurrent streams of that prompt type only, replicas cache-busted so prefill is real; never a mixed-category pool) — measuring TTFT (time to first token), TPOT (inter-token latency once decoding), per-stream and aggregate tok/s. The same tasks are also timed through each agentic harness so harness overhead is directly comparable. Pick a model, pick a metric — the curves show how it scales; the heatmap shows where it's strong (brighter = better).
AI Harness evaluation
One model, three agentic harnesses — exact versions disclosed.
How this works
How each model performs through the three agentic harnesses the controlled pod runs the suite through — Hermes, OpenClaw and OpenCode. Each cell is the model's agentic score on that harness; the exact release version used is shown beneath the score (and travels with every report). Column headers link to each harness's upstream repo.
Compare
Two whole benchmarks side by side — every board (text · harnesses · vision · audio · video · perf · arena · recipe), parity gaps called out. Single-run and seed A/Bs below.
How this works
Compare benchmarks: pick two benchmark jobs — every section renders side by side in a fixed order; a board only one side ran shows an explicit parity plate ("no … results for this run") so gaps are visible, never silently skipped. Compare runs: pick any two submissions — the shared suite joins per question, so you see both answers, both scores and both recipes head-to-head (the recipe-tuning A/B view; also reachable from the Submissions list via the ⇆ checkboxes). Compare by seed: a fast bench draws one question per (category × difficulty) from a seed; every model that ran it answered the identical questions.
single-run + seed tools
Live benchmark
Watch the pod score cases as they stream in.
How this works
Watch a controlled run as it happens — per-category progress and the prompts + answers as each case is scored. Streams in as the pod checkpoints (every few cases); no run in progress shows an idle state.
Run a benchmark
Every launch attempts a validated bench: hash-proven model, your choice of engine container, signed submission.
How this works
Paste a Hugging Face link — or point at weights already on disk plus their HF link. The pod resolves the repo and hash-validates the model (a matching local copy is good as gold: no re-download). Pick the engine container for your hardware — AEON's own boards run aeon-vllm-ultimate with fully-optimal settings; vLLM, SGLang, llama.cpp, ROCm or a custom image elsewhere; Apple silicon serves MLX bare-metal (macOS can't run MLX in a container) with the startup recipe reported exactly like a docker recipe. Validated model + AEON's bench execution + the pod's ed25519 signature = an ✓ attested global result. Anything that can't validate still runs — stored local-only, never globally ranked.
◉ Validated bench
Hash-validate → benchmark → sign → submit. Earns an ✓ attested global result — whether the pod serves the weights or you point it at a model that's already running.
⚙ RECIPE TUNING startup flags — dial in the optimal recipe for this system
Spec decode is lossless — the target model verifies every draft token, so only SPEED changes: low n favors concurrency, high n favors single-stream. DFlash presets require a drafter card (pulled fresh, hash-validated against its HF card, pinned in the recipe, and mounted at /drafter) and support n=1-15. DSpark runs either from an external DSpark drafter card (block-N drafters, e.g. deepseek-ai/dspark_qwen3_8b_block7 — same pull/validate/mount as DFlash) or fully in-checkpoint on DSpark-trained models (no drafter download; head internals come from the drafter's own config). Native MTP presets use the model's built-in MTP head and need no drafter; start small, then tune upward per model.
How custom flags work: type them exactly as you'd pass them to the engine, space-separated (--flag value or a bare --flag for booleans). Wrap a value containing spaces or JSON in single quotes. These MERGE into the engine's serve command — a flag also set above is replaced by your value; a new flag is appended. The bench wiring (served alias · host · port · 64K floor) is protected and can't be broken. The exact final recipe is recorded, shown on the result, and downloadable — every tuned run is a data point in the optimal-recipe search.
Paste a whole startup recipe here for total control (exotic engines, hand-tuned setups). When set, the engine dropdown + every flag above are ignored — the pod runs THIS command verbatim to serve, then still hash-validates the model and benches it. It MUST serve the alias model-under-test on :8000 (or your chosen port) so the harnesses connect. Recorded with the run exactly like a generated recipe.
Benchmark a serve that's already up — your live production endpoint, an MLX / LM Studio host, or a cluster head. Scan for it or paste its URL, then give the HF link above so the pod hash-verifies the weights and logprob-fingerprints the endpoint against them → ✓ attested on a match.
Comprehensive is the full benchmark — one launch populates every board (leaderboard · harnesses · performance · arena). Expect 30–60 min on capable hardware; Text suite only is the quick pass.
◆ Frontier API reference verified API
Run approved hosted frontier models through the same benchmark so local systems can be compared against ChatGPT, Grok, Claude, and other curated APIs. The pod validates the selected API/model before queueing, then the mothership records provider, version, effort, and logo metadata.
+ Save a provider API key here
▸ Endpoint bench unverified · never ranks
Prefer ◉ Point at a running model above. It benches the same already-running endpoint but hash-verifies the weights and fingerprints the serve, so the result earns ✓ attested and appears on the global leaderboard. Use this unverified card only for a deliberate private smoke test where you don't want a global rank — it skips model validation entirely, so it can never rank.
+ Save an endpoint API key here
Saved keys & tokens
Stored only in this pod (shown masked, never sent back in full). Use for authenticated endpoints and gated Hugging Face repos.
Generated Apps
Each round the category is shuffled — a random task, two random models, rendered live in isolated sandboxes. Click a frame to interact (games take keyboard/mouse). Models are hidden until you vote; pick the better result — or vote with the A / B / T keys.
Human-vote ranking
| # | model | Elo | W | L | T | win% |
|---|
Code Gallery
The top-rated artifact for every arena task — live previews + full source.
How this works
For each arena prompt, the top 10 generations ranked by per-artifact Elo from verified human A/B votes, plus the newest unrated submissions so fresh benchmark artifacts appear immediately. ▶ preview runs one live in an isolated sandbox; ⬇ code downloads the complete single-file HTML source as a ZIP. Artifacts with no counted votes yet are marked unrated.
Submissions — full transparency
Every run, fully inspectable: prompt, answer, score, judge.
How this works
Every benchmark run is fully inspectable: what the model was asked, exactly how it answered, the score, and how + by what it was judged (with rationale). Pick a run to view its complete results — and verify it against the signed manifest.
Select a submission on the left to view its full results.
▲ Run a Bench Pod one command · any hardware · attested results
Results are produced by a controlled pod you run on your own hardware — never from this page. Pull the prebuilt dashboard container, open the Run tab, and the pod hash-validates the model, serves it in your chosen engine (aeon-vllm-ultimate · vLLM · SGLang · llama.cpp · ROCm · custom image · Apple MLX · LM Studio), benchmarks through the real harnesses, and submits a signed attested result — engine, hardware and startup recipe attached.
docker run -d --name aeon-pod --network host --gpus all \ -v /var/run/docker.sock:/var/run/docker.sock \ -v aeon-pod-state:/root/.aeon \ -v "$HOME/aeon-models:/models" -e AEON_MODELS_HOST_DIR="$HOME/aeon-models" \ ghcr.io/aeon-7/aeon-pod:latest
Then open http://localhost:8091 → Run tab. --gpus all needs the NVIDIA Container Toolkit — without GPU visibility the CUDA engines hide themselves. CPU-only host? Drop the flag.
docker run -d --name aeon-pod -p 8091:8091 \ -v /var/run/docker.sock:/var/run/docker.sock \ -v aeon-pod-state:/root/.aeon \ -v "$HOME/aeon-models:/models" -e AEON_MODELS_HOST_DIR="$HOME/aeon-models" \ ghcr.io/aeon-7/aeon-pod:latest
Then open http://localhost:8091 → Run tab. The pod detects the Apple-silicon host and recommends MLX: it hands you the exact mlx_lm.server command (bare metal — macOS can't run MLX in a container), benches that endpoint, and records the startup recipe exactly like a docker recipe. LM Studio works the same way.
Admin — integrity & moderation
Benches — disqualify a faulty bench (kept, marked, unranked) or re-judge its Tier-1 cases
Evaluators — honeypot accuracy gates whether each account's votes count · history shows the evidence
| user | honeypot | accuracy | standing | counted | cast | action |
|---|
Generated artifacts — delete broken or inappropriate generations
Create an evaluator account
Anonymous — just a username and password. No email, no recovery (yet).
Change password
Updating . This signs out your other devices.
Run detail —
| case | cat | tier | score | evidence | ttft | tok/s | e2e |
|---|
▤ Pick a model folder
Support the work
I pour my soul, time, money, and sleepless nights into this. If it's been valuable to you, please consider a tip to support the continued work — every bit goes straight back into more compute, more models, and more open releases.
Join the Patreon — exclusive content, software, 3D print files & more