Princeton Gives AI Agents Unpublished Questions: Original Scientists Grade Results
Summary
AI research benchmark evaluation has a structural flaw: any task precise enough to grade is also precise enough to optimize for, which means benchmark scores and genuine AI capability come apart over time. A Princeton-led team’s CRUX project introduces the first evaluation paradigm to escape this