OenoBench Leaderboard
Top to bottom the leaderboard spans 30.3 percentage points on the same 3,266 multiple-choice wine questions: o3 at 83.6%, Claude Haiku 4.5 at 53.3%. Last place goes to a frontier lab’s small model, 7.2 points behind an open-weights 8B Llama. The filters slice the ranking by domain, by difficulty tier and by closed-book against contextual questions; the other tabs — reasoning lift, self-preference and cost per correct answer — are computed from the same run of 16 configurations.
Step plot. The corpus is ordered by the closed-book split: the first 1,601 questions are answerable from parametric knowledge alone, the remaining 1,665 require contextual reasoning. The pen holds 0.974 in the first regime and 0.703 in the second — a drop of 0.271 for o3 — inside a ±1 binomial standard-error band of 0.0159 and 0.0457 at a reference window of 100. The dashed rules are the unweighted mean over all configurations (16): 0.896 and 0.570, a drop of 0.326. Accuracy runs from 0.30 to 1.00. The trace is a step, not a rolling curve: the run file publishes regime aggregates rather than per-question results.
Dumbbell chart. The configurations (16) are ranked first to last by overall accuracy, one row each. A rule joins the row's accuracy on questions that require contextual reasoning (n 1,665) to its accuracy on closed-book questions (n 1,601), on a scale from 0.30 to 1.00, with the drop between the two printed at the right of the row. The smallest drop is GPT-5 at 0.266; the largest is DeepSeek-V3 at 0.396. The overall leader is o3.
- 1o3OpenAIeffort83.6%
- 2GPT-5OpenAI82.8%
- 3Gemini 2.5 Pro (thinking)Googlethinking82.6%
- 4Gemini 2.5 ProGoogle81.7%
- 5Claude Opus 4.7Anthropic81.0%
- 6Claude Opus 4.7 (thinking)Anthropicthinking81.0%
- 7GPT-5 miniOpenAI78.4%
- 8DeepSeek-R1DeepSeekthinking77.1%
- 9Gemini 2.5 FlashGoogle75.1%
- 10DeepSeek-V3DeepSeek70.3%
- 11Mistral Large 2411Mistral AI69.1%
- 12Qwen 2.5 72BAlibaba67.4%
- 13Llama 3.3 70BMeta67.1%
- 14Llama 3.1 8BMeta60.5%
- 15Qwen 2.5 7BAlibaba57.0%
- 16Claude Haiku 4.5Anthropic53.3%
Every figure above is read from a single versioned run file, published under CC BY-SA 4.0: download the full results (JSON, 23 KB). Run files accumulate and none is overwritten, so this URL keeps resolving after later runs are published — cite the dated file, not the leaderboard page.
Leaderboard updated