Spec sheet · run v2026-05-04 · OenoBench release_v1.2
How much does a language model actually know about wine?
A research project on the present and future of artificial intelligence in viticulture, winemaking and the wine business. It holds OenoBench, the wine-knowledge benchmark for large language models, and the Casebook, the most complete audited register of AI deployments in wine.
Instrument record
- Run
- v2026-05-04
- Released
- 2026-05-04
- Paper
- OenoBench release_v1.2
- Questions
- 3,266
- Evaluations
- 52,256
- Configurations
- 16
- Evaluation spend
- $98.33
Domain tags sum to 3,329 and difficulty tags to 3,329 against a corpus of 3,266 — unreconciled, reported as published.
Casebook record
- Audited cases
- 299
- Reached operation
- 147 · 49%
- Of those, on the vendor's word alone
- 40
Measured answer
o3 · OpenAI · best of 16 configurations · 3,266 questions
Chromatogram. 6 peaks, one per wine domain, on a shared baseline with 5 dashed drop lines marking the integration boundaries between them. Each peak's area is proportional to the number of questions carrying that domain tag, and each is labelled with its domain, its question count and its share. The largest is Wine regions at 1,108 questions (33.3%); the smallest is Winemaking at 188 (5.6%). Shares are of the 3,329 domain tags the run publishes, not of its 3,266 questions: 63 questions carry more than one tag. The vertical axis is unlabelled detector signal.
Contents
§1 About the project
What the atlas found, on one page.
§2 Keynote
The long read on what already works in wine and what stays in the papers.
§3 Benchmarks
- 3.1 Leaderboard v2026-05-04
§5 Research
Whether a deployment outlives its pilot, and what a benchmark score predicts. No studies yet.
§6 About the author
Who checked all this: a wine expert who deploys AI for a living.