1、FRONTIERSCIENCE:EVALUATINGAIS ABILITY TOPERFORM EXPERT-LEVEL SCIENTIFIC TASKSMiles WangJoy JiaoNeil ChowdhuryEthan ChangTejal PatwardhanOpenAIABSTRACTWe introduce FrontierScience,a benchmark evaluating AI capabilities for expert-level scientifi c reasoning.FrontierScience consists of two tracks:(1)O
2、lympiad,which contains international olympiad problems(at the level of IPhO,IChO,and IBO),and(2)Research,which contains PhD-level,open-ended problemsrepresentative of sub-problems in scientifi c research.In total,FrontierScience iscomposed of several hundred questions(160 in the open-sourced gold se
3、t)coveringsubfi elds across physics,chemistry,and biology,from quantum electrodynamics tosynthetic organic chemistry.Recent model progress has nearly saturated existingscience benchmarks,which often rely on multiple-choice knowledge questions oralready published information.In contrast,all Olympiad
4、problems are originallyproduced by international olympiad medalists and national team coaches to ensurestandards of diffi culty,originality,and factuality.All Research problems areresearch sub-tasks written and verifi ed by PhD scientists(doctoral candidates,post-doctoral researchers,or professors).
5、For Research,we also introduce a granularrubric-based architecture to evaluate model capabilities throughout the processof solving a research task,as opposed to judging a standalone answer.In initialevaluations of several frontier models,GPT-5.2 is the top performing model onFrontierScience,scoring
6、77%on the Olympiad set and 25%on the Research set.1INTRODUCTIONLanguage models reasoning capabilities have signifi cantly advanced in scientifi c domains.WhenGPQA,a“Google-Proof”multiple-choice science benchmark written by PhD experts,was releasedin November 2023,GPT-4 scored 39%,below the expert ba