Superior intelligence desk · science
How to measure superior intelligence beyond a benchmark score
A benchmark result answers a defined test question. A business decision asks whether a system can do useful work under the conditions the buyer actually faces. Moving from one to the other takes an evaluation plan.

Read the unit before the number
METR's task horizon work measures AI performance against the time human professionals need for tasks, at a specified probability of success. A human task time is not the same as the time the AI spends running. Its software-task measurements should not be treated as a universal measure of every kind of work. Read METR's methodology and limitations.
A result at a 50 percent success threshold also answers a different question from a result at 80 percent. When sharing a chart, keep the threshold, task domain and evaluation date with it. Otherwise the number loses much of its meaning.
Do not average away the failure that matters
The International AI Safety Report 2026 describes uneven capabilities: strong results in some areas can coexist with failures on simpler tasks and longer workflows. That is one reason an aggregate score cannot settle a deployment decision. Read the executive summary.
In an estimate-writing workflow, an attractive draft and an incorrect quantity are not equivalent events. Record error types separately. A small improvement in average writing quality should not hide an increase in costly mistakes.
Build a small comparison you can repeat
Our suggested evaluation starts with a fixed set of representative cases. Include ordinary work, missing information and exceptions. Keep a separate group of cases for the final comparison so that repeated tuning does not quietly turn the test into practice material.
Give a qualified reviewer a written scoring guide. Define what counts as correct, what requires a correction and what makes an output unusable. Record the system version, tools, instructions, time limit and whether a person intervened.
Compare against the existing process using the same acceptance criteria. Include review and cleanup time. If the assistant saves drafting time but adds more verification work, the result may not help the team.
Measure recovery as well as success
Introduce a reversible interruption, such as an unavailable test file or an incomplete sample record. Observe whether the system stops clearly, invents missing information, repeats a failed action or asks for help.
For an Alberta office trial, keep the exercise separate from customer-facing work until the team understands these behaviours. A failure caught in a test is useful evidence. A failure hidden by a polished explanation is harder to diagnose.
Write a decision, not just a score
Finish the evaluation with a bounded conclusion: this configuration met these criteria on these cases, with these remaining weaknesses. State the conditions that require another review, such as a model change, new data source or expanded tool access.
That conclusion is more actionable than declaring a system universally superior. It tells a team what they have learned and where they still need evidence.
Sources and editorial note
Sources checked September 20, 2026. Practical examples and trial plans are this publication's analysis. They do not describe measured client results.




