benchmarks · adventurebench
Can your model play a text adventure?
Adventure Bench measures grounded action interpretation: map a player utterance onto the visible scene’s small action vocabulary, or refuse when the request is not grounded.
the cost curve
The score–cost frontier
Every labeled frontier point is a model for which no cheaper tested model scored as well. The other points are dominated on this release; exact values for every point appear in the table below.
The frontier is a cost-efficiency view, not a ranking or a statistical claim. Costs are summed from provider-recorded completion evidence; score uncertainty and paired comparisons remain in the table below.
| Model | Score (95% CI) | Passed | Recorded cost | Provider / runtime | Parse / transport | Separates |
|---|---|---|---|---|---|---|
| google/gemma-3-4b-it | 72.5% ; 95% cluster-bootstrap interval (66.8% to 77.9%) | 531/732 | $0.020 | DeepInfra / openrouter:chat-completions-v1 | 0 / 0 | +0 / −7 |
| ibm-granite/granite-4.0-h-micro | 73.4% ; 95% cluster-bootstrap interval (67.6% to 78.7%) | 537/732 | $0.008 | Cloudflare / openrouter:chat-completions-v1 | 3 / 0 | +0 / −7 |
| ibm-granite/granite-4.1-8b | 75.8% ; 95% cluster-bootstrap interval (70.1% to 81.1%) | 555/732 | $0.020 | CoreWeave / openrouter:chat-completions-v1 | 0 / 0 | +1 / −6 |
| ibm-granite/granite-4.2-8b | 79.5% ; 95% cluster-bootstrap interval (74.2% to 84.4%) | 582/732 | $0.023 | CoreWeave / openrouter:chat-completions-v1 | 9 / 0 | +4 / −3 |
| meta-llama/llama-3.1-8b-instruct | 73.6% ; 95% cluster-bootstrap interval (68.0% to 79.0%) | 539/732 | $0.020 | Groq / openrouter:chat-completions-v1 | 0 / 0 | +0 / −7 |
| meta-llama/llama-3.2-3b-instruct | 68.9% ; 95% cluster-bootstrap interval (62.7% to 74.6%) | 504/732 | $0.023 | Parasail / openrouter:chat-completions-v1 | 72 / 0 | +0 / −8 |
| mistralai/ministral-14b-2512 | 86.5% ; 95% cluster-bootstrap interval (82.1% to 90.6%) | 633/732 | $0.043 | Mistral / openrouter:chat-completions-v1 | 0 / 0 | +6 / −1 |
| mistralai/ministral-3b-2512 | 84.3% ; 95% cluster-bootstrap interval (79.8% to 88.5%) | 617/732 | $0.014 | Mistral / openrouter:chat-completions-v1 | 0 / 0 | +5 / −1 |
| mistralai/ministral-8b-2512 | 84.4% ; 95% cluster-bootstrap interval (79.8% to 88.8%) | 618/732 | $0.022 | Mistral / openrouter:chat-completions-v1 | 0 / 0 | +6 / −1 |
| nvidia/nemotron-3.5-lightning | 84.4% ; 95% cluster-bootstrap interval (79.6% to 88.8%) | 618/732 | $0.033 | DeepInfra / openrouter:chat-completions-v1 | 2 / 0 | +5 / −1 |
| qwen/qwen3.5-9b | 82.7% ; 95% cluster-bootstrap interval (77.7% to 87.3%) | 605/732 | $0.039 | DeepInfra / openrouter:chat-completions-v1 | 0 / 0 | +5 / −1 |
| z-ai/glm-5.2 | 90.6% ; 95% cluster-bootstrap interval (87.0% to 93.9%) | 663/732 | $0.119 | DeepInfra / openrouter:chat-completions-v1 | 0 / 0 | +11 / −0 |
These results are not a ranking. The separates column is +ahead / −behind and counts only paired 95% case-cluster bootstrap intervals that exclude zero. Every comparison whose interval includes zero is unresolved; unresolved comparisons are unresolved.
Release 20260903-plicara-v1-expanded covers 244 cases × 3 repetitions (732 case-repetitions; repetition numbers 1, 2, 3), collected from 2026-09-02 through 2026-09-03 (UTC). Benchmark version: 1.0.0. Open the audited release artifact.
Failure policy: any transport failure invalidates a collection run and prevents it from entering this release. Parsing failures are counted separately after the frozen malformed-output retry and remain visible in the table. See the full methodology.
Every number above regenerates offline from committed raw responses. The release pins every manifest and response hash plus dataset f38e4650ecbe… and prompt a5b12544f699…. Audit it with make site-check RELEASE_ID=20260903-plicara-v1-expanded.
| Tag | google-gemma-3-4b-it | ibm-granite-granite-4-0-h-micro | ibm-granite-granite-4-1-8b | ibm-granite-granite-4-2-8b-20260831 | meta-llama-llama-3-1-8b-instruct | meta-llama-llama-3-2-3b-instruct | mistralai-ministral-14b-2512 | mistralai-ministral-3b-2512 | mistralai-ministral-8b-2512 | nvidia-nemotron-3-5-lightning | qwen-qwen3-5-9b | z-ai-glm-5-2 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| abbreviation | 12/30 | 12/30 | 30/30 | 21/30 | 24/30 | 12/30 | 27/30 | 13/30 | 24/30 | 21/30 | 21/30 | 27/30 |
| absent-object | 18/30 | 12/30 | 27/30 | 21/30 | 12/30 | 12/30 | 30/30 | 30/30 | 30/30 | 30/30 | 30/30 | 30/30 |
| adverb | 21/21 | 21/21 | 21/21 | 21/21 | 21/21 | 18/21 | 21/21 | 21/21 | 21/21 | 21/21 | 21/21 | 21/21 |
| ambiguity | 9/30 | 12/30 | 12/30 | 21/30 | 18/30 | 15/30 | 30/30 | 18/30 | 16/30 | 24/30 | 27/30 | 30/30 |
| classic-command | 30/48 | 27/48 | 21/48 | 33/48 | 29/48 | 18/48 | 44/48 | 44/48 | 39/48 | 40/48 | 39/48 | 45/48 |
| compound | 24/27 | 27/27 | 18/27 | 27/27 | 27/27 | 21/27 | 24/27 | 22/27 | 24/27 | 21/27 | 20/27 | 26/27 |
| direction-as-place | 42/45 | 42/45 | 42/45 | 42/45 | 42/45 | 42/45 | 45/45 | 42/45 | 42/45 | 42/45 | 45/45 | 45/45 |
| exact-verb | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 | 36/36 |
| flavor-noun | 6/42 | 9/42 | 3/42 | 6/42 | 0/42 | 3/42 | 17/42 | 18/42 | 11/42 | 18/42 | 15/42 | 14/42 |
| frustration | 9/12 | 12/12 | 12/12 | 12/12 | 9/12 | 9/12 | 12/12 | 12/12 | 12/12 | 9/12 | 12/12 | 10/12 |
| full-sentence | 24/24 | 24/24 | 21/24 | 24/24 | 24/24 | 24/24 | 21/24 | 24/24 | 24/24 | 24/24 | 21/24 | 24/24 |
| go-to-place | 9/12 | 9/12 | 9/12 | 9/12 | 9/12 | 9/12 | 11/12 | 9/12 | 9/12 | 7/12 | 12/12 | 9/12 |
| greeting | 3/9 | 9/9 | 9/9 | 0/9 | 0/9 | 0/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 | 8/9 |
| guess-the-noun | 6/12 | 12/12 | 12/12 | 6/12 | 12/12 | 12/12 | 12/12 | 12/12 | 12/12 | 9/12 | 12/12 | 12/12 |
| help-seeking | 9/15 | 12/15 | 9/15 | 15/15 | 9/15 | 9/15 | 15/15 | 15/15 | 15/15 | 9/15 | 12/15 | 15/15 |
| meta | 15/30 | 18/30 | 27/30 | 30/30 | 27/30 | 9/30 | 27/30 | 27/30 | 30/30 | 27/30 | 30/30 | 27/30 |
| missing-preposition | 18/27 | 27/27 | 27/27 | 27/27 | 27/27 | 27/27 | 24/27 | 24/27 | 27/27 | 27/27 | 27/27 | 27/27 |
| multi-object | 15/15 | 15/15 | 3/15 | 15/15 | 15/15 | 15/15 | 15/15 | 13/15 | 15/15 | 15/15 | 12/15 | 15/15 |
| nonsense | 9/12 | 9/12 | 9/12 | 12/12 | 12/12 | 6/12 | 12/12 | 10/12 | 9/12 | 12/12 | 12/12 | 12/12 |
| out-of-vocab | 60/126 | 42/126 | 45/126 | 63/126 | 50/126 | 48/126 | 95/126 | 98/126 | 80/126 | 97/126 | 71/126 | 97/126 |
| paraphrase | 42/51 | 51/51 | 48/51 | 48/51 | 48/51 | 42/51 | 45/51 | 44/51 | 45/51 | 45/51 | 48/51 | 51/51 |
| politeness | 18/18 | 18/18 | 18/18 | 18/18 | 15/18 | 18/18 | 15/18 | 18/18 | 18/18 | 15/18 | 18/18 | 16/18 |
| pronoun | 12/27 | 15/27 | 15/27 | 21/27 | 18/27 | 15/27 | 21/27 | 21/27 | 19/27 | 24/27 | 24/27 | 25/27 |
| question | 42/45 | 36/45 | 45/45 | 45/45 | 36/45 | 30/45 | 35/45 | 45/45 | 45/45 | 42/45 | 42/45 | 45/45 |
| relative-direction | 3/12 | 3/12 | 3/12 | 6/12 | 3/12 | 3/12 | 9/12 | 3/12 | 6/12 | 3/12 | 3/12 | 12/12 |
| story-mode | 9/9 | 0/9 | 9/9 | 9/9 | 9/9 | 6/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 | 9/9 |
| synonym | 60/63 | 63/63 | 63/63 | 63/63 | 63/63 | 60/63 | 60/63 | 60/63 | 63/63 | 57/63 | 60/63 | 60/63 |
| typo | 36/36 | 36/36 | 33/36 | 36/36 | 33/36 | 36/36 | 30/36 | 31/36 | 36/36 | 34/36 | 33/36 | 36/36 |
| verb-alias | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 |
Per-tag cells are numerators and denominators, not pooled rankings; they show which grounded mapping and calibration patterns each audited run passed.