Skip to this page
THE SUPER INTELLIGENCE TIMESBY OPCELERATE NEURAL RSS
Back to the front page

Superior intelligence desk · science

How to measure superior intelligence beyond a benchmark score

A benchmark result answers a defined test question. A business decision asks whether a system can do useful work under the conditions the buyer actually faces. Moving from one to the other takes an evaluation plan.

A brass balance weighs a glass brain against papers and a stopwatch.
Editorial illustration · AI-generated

Read the unit before the number

METR's task horizon work measures AI performance against the time human professionals need for tasks, at a specified probability of success. A human task time is not the same as the time the AI spends running. Its software-task measurements should not be treated as a universal measure of every kind of work. Read METR's methodology and limitations.

A result at a 50 percent success threshold also answers a different question from a result at 80 percent. When sharing a chart, keep the threshold, task domain and evaluation date with it. Otherwise the number loses much of its meaning.

Do not average away the failure that matters

The International AI Safety Report 2026 describes uneven capabilities: strong results in some areas can coexist with failures on simpler tasks and longer workflows. That is one reason an aggregate score cannot settle a deployment decision. Read the executive summary.

In an estimate-writing workflow, an attractive draft and an incorrect quantity are not equivalent events. Record error types separately. A small improvement in average writing quality should not hide an increase in costly mistakes.

Build a small comparison you can repeat

Our suggested evaluation starts with a fixed set of representative cases. Include ordinary work, missing information and exceptions. Keep a separate group of cases for the final comparison so that repeated tuning does not quietly turn the test into practice material.

Give a qualified reviewer a written scoring guide. Define what counts as correct, what requires a correction and what makes an output unusable. Record the system version, tools, instructions, time limit and whether a person intervened.

Compare against the existing process using the same acceptance criteria. Include review and cleanup time. If the assistant saves drafting time but adds more verification work, the result may not help the team.

Measure recovery as well as success

Introduce a reversible interruption, such as an unavailable test file or an incomplete sample record. Observe whether the system stops clearly, invents missing information, repeats a failed action or asks for help.

For an Alberta office trial, keep the exercise separate from customer-facing work until the team understands these behaviours. A failure caught in a test is useful evidence. A failure hidden by a polished explanation is harder to diagnose.

Write a decision, not just a score

Finish the evaluation with a bounded conclusion: this configuration met these criteria on these cases, with these remaining weaknesses. State the conditions that require another review, such as a model change, new data source or expanded tool access.

That conclusion is more actionable than declaring a system universally superior. It tells a team what they have learned and where they still need evidence.

Sources and editorial note

Sources checked September 20, 2026. Practical examples and trial plans are this publication's analysis. They do not describe measured client results.