[go: up one dir, main page]

Read the Frontier AI Trends Report
Please enable javascript for this website.

Science of Evaluations

AISI brand artwork

Blog

Optimal stopping: spending evaluation compute where it counts

Science of Evaluations

•

August 27, 2026

We introduce optstop, an open-source tool for LLM evaluations that keeps running where uncertainty is high, and stops where estimates are precise or stable enough.

More compute, more capability: Why AI agent evaluations need to account for test-time compute

Science of Evaluations

•

July 2, 2026

Standard evaluations cap how much compute AI agents can use. We show that raising those caps changes measured capability, the difficulty of tasks agents can solve, and how fast the frontier appears to move.

Evidence for inference scaling in AI cyber tasks: Increased evaluation budgets reveal higher success rates

Science of Evaluations

•

March 5, 2026

Alongside Irregular, we found evidence demonstrating that evaluators need to use large token budgets to understand the cyber capabilities of recent Large Language Models (LLMs).

A pipeline for transcript analysis using Inspect Scout

Science of Evaluations

•

February 25, 2026

We outline a step-by-step pipeline for using our open-source transcript analysis tool, Inspect Scout.

Transcript analysis for AI agent evaluations

Science of Evaluations

•

October 10, 2025

Why we use transcript analysis for our agent evaluations, and results from an early case study.

A structured protocol for elicitation experiments

Science of Evaluations

•

July 16, 2025

Calibrating AI risk assessment through rigorous elicitation practices.

LLM judges on trial: A new statistical framework to assess autograders

Science of Evaluations

•

July 9, 2025

Our new framework can assess the reliability of LLM evaluators, while simultaneously answering a primary research question.

HiBayES: Improving LLM evaluation with hierarchical Bayesian modelling

Science of Evaluations

•

May 12, 2025

HiBayES: a flexible, robust statistical modelling framework that accounts for the nuances and hierarchical structure of advanced evaluations.

Long-Form Tasks

Science of Evaluations

•

December 3, 2024

A Methodology for Evaluating Scientific Assistants

Early Insights from Developing Question-Answer Evaluations for Frontier AI

Science of Evaluations

•

September 23, 2024

A common technique for quickly assessing AI capabilities is prompting models to answer hundreds of questions, then automatically scoring the answers. We share insights from months of using this method.