OpenAI炮轰AI评测"标杆": 731 道题近三成有缺陷 - AI NEWS
Quick Answer
OpenAI challenges the SWE-Bench Pro benchmark, claiming 30% of its 731 tasks have flaws, undermining its credibility.
Quick Take
The rapid increase in pass rates from 23.3% to 80.3% in just eight months raises concerns about the benchmark's ability to accurately assess AI models' programming skills.
Key Points
- OpenAI identifies 200 flawed tasks, representing 27.4% of Pro's total.
- A second review found 249 flawed tasks, increasing the defect rate to 34.1%.
- Issues include overly strict testing and misleading prompts affecting model evaluations.
- OpenAI calls for a new assessment framework designed by experienced software developers.
- The current benchmark's flaws challenge the credibility of the entire AI evaluation system.
Source Excerpt
OpenAI公开质疑 Pro基准,指出其731个测试任务中约30%存在评测缺陷。 该基准由Scale AI推出,是衡量大模型编程能力的行业权威。 但OpenAI警示,前沿模型通过率8个月内从23. 3%飙升至80. 3%,进步速度异常,暗示评测可靠性存疑。
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →It's true that Anthropic is the fastest growing company ...
Anthropic is rapidly emerging as a leader in AI with its Claude model, which outperforms competitors in various benchmarks. Their coding agents are also recognized as the best in the industry, contributing to their unprecedented revenue growth and solidifying their position in the enterprise market.
WSJ: OpenAI is considering deep price reductions as competition ...
OpenAI is contemplating significant price cuts in response to competitive pressure from Anthropic, particularly due to the success of Claude Code in developer and coding workflows. This shift could affect pricing strategies in the AI market as companies vie for dominance in coding solutions.