当前位置:首页 >英文主页 >中英对照 > 报告详情

OpenAI:2025 FrontierScience:评估人工智能执行专家级科学任务的能力研究报告(英文版)(15页).pdf

上传人: 1****1 编号:1000490 2025-12-26 15页 11.82MB

下载:

1、FRONTIERSCIENCE:EVALUATINGAIS ABILITY TOPERFORM EXPERT-LEVEL SCIENTIFIC TASKSMiles WangJoy JiaoNeil ChowdhuryEthan ChangTejal PatwardhanOpenAIABSTRACTWe introduce FrontierScience,a benchmark evaluating AI capabilities for expert-level scientifi c reasoning.FrontierScience consists of two tracks:(1)O

2、lympiad,which contains international olympiad problems(at the level of IPhO,IChO,and IBO),and(2)Research,which contains PhD-level,open-ended problemsrepresentative of sub-problems in scientifi c research.In total,FrontierScience iscomposed of several hundred questions(160 in the open-sourced gold se

3、t)coveringsubfi elds across physics,chemistry,and biology,from quantum electrodynamics tosynthetic organic chemistry.Recent model progress has nearly saturated existingscience benchmarks,which often rely on multiple-choice knowledge questions oralready published information.In contrast,all Olympiad

4、problems are originallyproduced by international olympiad medalists and national team coaches to ensurestandards of diffi culty,originality,and factuality.All Research problems areresearch sub-tasks written and verifi ed by PhD scientists(doctoral candidates,post-doctoral researchers,or professors).

5、For Research,we also introduce a granularrubric-based architecture to evaluate model capabilities throughout the processof solving a research task,as opposed to judging a standalone answer.In initialevaluations of several frontier models,GPT-5.2 is the top performing model onFrontierScience,scoring

6、77%on the Olympiad set and 25%on the Research set.1INTRODUCTIONLanguage models reasoning capabilities have signifi cantly advanced in scientifi c domains.WhenGPQA,a“Google-Proof”multiple-choice science benchmark written by PhD experts,was releasedin November 2023,GPT-4 scored 39%,below the expert ba

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
根据《FrontierScience: Evaluating AI’s Ability to Perform Expert-Level Scientific Tasks》一文,主要内容如下: - FrontierScience是一个评估AI在专家级科学推理能力的新基准,包含两个部分:Olympiad(国际奥林匹克竞赛问题)和Research(博士级开放性问题)。 - 该基准由物理、化学和生物学领域的专家编写和验证,问题难度高、原创性强、可验证。 - GPT-5.2在Olympiad部分得分77%,在Research部分得分25%,表现最佳。 - FrontierScience旨在衡量AI在科学推理、判断和解决实际研究问题方面的能力,为AI加速科学进步提供评估工具。
"AI科学难题挑战赛,你敢来吗?" "揭秘AI如何解决科学难题!" "前沿科技,AI科学推理能力大考验!"
客服
商务合作
小程序
服务号
折叠