Skip to content
VinumExMachina Atlas of AI in Wine
Select language: Русский

Benchmarks

OenoBench Leaderboard

Top to bottom the leaderboard spans 30.3 percentage points on the same 3,266 multiple-choice wine questions: o3 at 83.6%, Claude Haiku 4.5 at 53.3%. Last place goes to a frontier lab’s small model, 7.2 points behind an open-weights 8B Llama. The filters slice the ranking by domain, by difficulty tier and by closed-book against contextual questions; the other tabs — reasoning lift, self-preference and cost per correct answer — are computed from the same run of 16 configurations.

Fig.02
Fig. 02Where the instrument dropsreading 1 of 3 · o3 · two regimes · n 1,601 / 1,665

Step plot. The corpus is ordered by the closed-book split: the first 1,601 questions are answerable from parametric knowledge alone, the remaining 1,665 require contextual reasoning. The pen holds 0.974 in the first regime and 0.703 in the second — a drop of 0.271 for o3 — inside a ±1 binomial standard-error band of 0.0159 and 0.0457 at a reference window of 100. The dashed rules are the unweighted mean over all configurations (16): 0.896 and 0.570, a drop of 0.326. Accuracy runs from 0.30 to 1.00. The trace is a step, not a rolling curve: the run file publishes regime aggregates rather than per-question results.

Δ 0.2710.300.400.500.600.700.800.901.0001,0001,6012,0003,0003,266CLOSED-BOOK (PARAMETRIC) · n 1,601CONTEXTUAL REASONING · n 1,6650.9740.7030.8960.570ordinate: accuracy · abscissa: corpus ordered by closed-book split
±1 SE, binomial, at a reference window of 100mean of all configurations (16)solid pen: o3 · OpenAI · effort
Finding 1. Every configuration collapses on questions that require contextual reasoning rather than recall. o3 loses 0.271; the mean across all configurations (16) loses 0.326. The noise band widens with the drop — the contextual regime is measured less precisely at the same reference window. The trace is a step, not a rolling curve: the run file publishes regime aggregates rather than per-question results.
Fig.03
Fig. 03The gap does not close with rankall 16 configurations · sorted by overall accuracy

Dumbbell chart. The configurations (16) are ranked first to last by overall accuracy, one row each. A rule joins the row's accuracy on questions that require contextual reasoning (n 1,665) to its accuracy on closed-book questions (n 1,601), on a scale from 0.30 to 1.00, with the drop between the two printed at the right of the row. The smallest drop is GPT-5 at 0.266; the largest is DeepSeek-V3 at 0.396. The overall leader is o3.

0.300.400.500.600.700.800.901.00ACCURACYΔ COLLAPSECONFIGURATIONo30.271GPT-50.266Gemini 2.5 Pro (thinking)0.290Gemini 2.5 Pro0.300Claude Opus 4.70.319Claude Opus 4.7 (thinking)0.323GPT-5 mini0.319DeepSeek-R10.337Gemini 2.5 Flash0.343DeepSeek-V30.396Mistral Large 24110.358Qwen 2.5 72B0.381Llama 3.3 70B0.364Llama 3.1 8B0.304Qwen 2.5 7B0.311Claude Haiku 4.50.331
Closed-book (parametric) · n 1,601Contextual reasoning · n 1,665
Ranked first to last, the bar never shortens by more than a few points. The smallest gap belongs to GPT-5 (0.266); the largest to DeepSeek-V3 (0.396).
OenoBench · v2026-05-0416 model configurations · 3,266 questions across six wine domains
Wine domain
Question difficulty
Closed-book vs contextual
ModelAccuracy
  1. 1
    o3OpenAIeffort
    83.6%
  2. 2
    GPT-5OpenAI
    82.8%
  3. 3
    Gemini 2.5 Pro (thinking)Googlethinking
    82.6%
  4. 4
    Gemini 2.5 ProGoogle
    81.7%
  5. 5
    Claude Opus 4.7Anthropic
    81.0%
  6. 6
    Claude Opus 4.7 (thinking)Anthropicthinking
    81.0%
  7. 7
    GPT-5 miniOpenAI
    78.4%
  8. 8
    DeepSeek-R1DeepSeekthinking
    77.1%
  9. 9
    Gemini 2.5 FlashGoogle
    75.1%
  10. 10
    DeepSeek-V3DeepSeek
    70.3%
  11. 11
    Mistral Large 2411Mistral AI
    69.1%
  12. 12
    Qwen 2.5 72BAlibaba
    67.4%
  13. 13
    Llama 3.3 70BMeta
    67.1%
  14. 14
    Llama 3.1 8BMeta
    60.5%
  15. 15
    Qwen 2.5 7BAlibaba
    57.0%
  16. 16
    Claude Haiku 4.5Anthropic
    53.3%

Every figure above is read from a single versioned run file, published under CC BY-SA 4.0: download the full results (JSON, 23 KB). Run files accumulate and none is overwritten, so this URL keeps resolving after later runs are published — cite the dated file, not the leaderboard page.

Leaderboard updated