
Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
Quick Answer
Lukas Petersson and Axel Backlund from Andon Labs discuss their development of VendingBench, a benchmark that evaluates AI models from Haiku to Mythos.
Quick Take
They emphasize the importance of creating robust evaluation frameworks to ensure lasting performance metrics in AI advancements.
Key Points
- VendingBench evaluates AI models, including Haiku and Mythos.
- Andon Labs focuses on building lasting evaluation frameworks.
- Robust evaluations are crucial for measuring AI performance.
- The discussion highlights the evolution of AI model assessments.
Source Excerpt
We talk with the VendingBench authors on evaling Claudes from Haiku to Mythos, and how they build leading, and lasting, frontier evals from scratch.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Latent Space
See more →![[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition](https://substackcdn.com/image/fetch/$s_!8D6O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHMuQw2BXUAAJaQd.png)
[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition
SpaceXAI has launched Grok 4.5, a new coding-focused model that is 3x larger than Grok 4.3, priced at $2 per million input tokens. Positioned as an Opus-class model, it aims for efficiency and speed, outperforming competitors like GPT-5.6 and Opus 4.8 in cost-effectiveness.

