DeepSeek-R1超级外挂!“人类最后的考试”首次突破30分,上海交大等开源方案碾压OpenAI、谷歌
Quick Answer
The Shanghai Jiao Tong University team achieved a groundbreaking score of 32.1 on the Humanity's Last Exam benchmark, surpassing previous records and outperforming OpenAI and Google.
Quick Take
Their open-source tool-enhanced reasoning agent, X-Master, utilizes the DeepSeek-R1-0528 model and introduces a novel workflow to enhance problem-solving capabilities.
Key Points
- X-Master scored 32.1, the first system to exceed 30 on .
- The model uses DeepSeek-R1-0528 and features a novel multi-agent workflow.
- X-Masters outperformed existing models in all tested categories.
- The approach combines tool-enhanced reasoning with structured exploration.
- X-Master achieved 67.4% accuracy in complex biology tasks.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →lila ayu
The 8-week hands-on track focuses on mastering Generative AI, , QLoRA fine-tuning, and AI Agents, enabling participants to build 8 real-world applications. This program emphasizes practical skills with over 20 Frontier and Open models, catering to developers and AI enthusiasts aiming to enhance their expertise in AI technologies.