Hands Off, Eyes Open, Brain Engaged

What should assessment validate when the work already involves AI?

Published: 8/20/2026
Hands Off, Eyes Open, Brain Engaged

Something is shifting in assessment, and the conversation about it keeps landing in the wrong place.

Ask most organisations what "AI in assessment" means and they will more than likely talk about marking, AI-generated items, simulated standard setting. These are things that matter now, but what we've not been focusing on is that whilst AI presents an opportunity to enhance the 20th century test development and marking approaches, it's also completely undermining the foundations of the very system.

We're still testing on the assumption that AI is not present in the day to day, but the reality is that AI is already inside the work. A present reality. Analysts, clinicians, engineers, teachers, lawyers, developers, project managers: across professions, AI is embedded in the daily practice of roles that our qualifications are supposed to certify.

If AI is part of how the work is done, then an assessment that removes it is not measuring the job. It is measuring a historical version of the job. The evidence it produces may be clean and reliable, but it is evidence of the wrong thing.

Assessment with AI does not just mean marking with AI. It means AI being present in the assessment the way it is present in the work. If someone uses AI every day in their role, and we assess them in a locked-down environment where AI does not exist, we are not measuring the job. We are measuring a historical re-enactment of it.

That is not a minor technical point. It is a validity problem. And it leaves assessment organisations with a question that cannot be answered with a blanket rule: for this particular task, what relationship with AI is appropriate, and what should the assessment be validating as a result?

The AI Reliance and Assessment Resilience Grid is a framework for answering that question, task by task.

Where this started

In August 2026, the Association of Test Publishers published "Assessment in 2040: Voices from the Field," a collection of expert perspectives on how assessment will change over the next decade. Contributors included psychometricians, chief assessment officers, test security specialists, CEOs of credentialing bodies and certification organisations, and three AI models invited to comment on the humans' predictions.

My contribution used the levels of vehicle automation as a way into the question. As AI takes over more of the work humans do, assessment moves through stages: hands off, eyes off, brain off. As with all contributions, content was short, so I wanted to expand my thinking and show how it is still very much work in progress.

The ATP piece planted a seed. I am increasingly interested in what an AI collaborator looks like in assessment, both as a partner through the process and as an evaluator. The grid is where that thinking has taken me.

Why driving?

The grid borrows its structure from the levels of vehicle automation. It is a model almost everyone already understands, which is exactly why it was chosen. I first read about the model in the amazing book, Hello World by Professor Hannah Fry, a book I recommend to all my students on my AI and Automation Apprenticeship.

In driving, the progression runs from full human control through several stages of shared responsibility to full automation. At each stage, the human role changes, and so does what that human needs to be capable of. The person supervising a self-driving car needs different competencies from the person driving manually. Not fewer. Different.

The analogy does more than make the levels accessible. Automated driving carries unresolved questions that assessment shares. Who is accountable when the system fails? How are decisions made, and are they explainable? Should everyone have access to the same level of automation? What is lost when the human stops doing the task themselves? Is there value in the human involvement that disappears when we automate?

Some people get great value from driving itself. That brings us to the question: just because we can automate something, should we? But there is a second side to it. Other people value having a human driver. They want to know a person is in control. Sometimes the value is in doing the task. Sometimes the value is that others know a human did it.

The grid does not resolve those questions. It gives you a structured way to work through them, one task at a time.

The five levels

Each level describes a different relationship between the person and AI, and each requires a different kind of evidence.

Level 1: Hands On, Unaided. The person performs the task without AI. They are in continuous control. The assessment should validate independent competence: not just that they performed correctly, but that they can explain the reasoning behind their decisions.

This level exists wherever independent human capability still matters. A pilot managing an emergency. A nurse performing CPR. A trainee teacher planning a lesson. A trainee building foundational knowledge they will need before they can meaningfully supervise AI later. The task sits here because unaided competence is the thing being certified.

Level 2: Hands On, AI Assisted. The person performs the task with AI support but remains in charge throughout. The assessment should validate that they use AI effectively while adding professional value the AI alone could not provide.

This is where most professional work is heading, and it is the hardest level to assess well with traditional approaches.

Two people can use the same AI tool on the same task and produce plausibly similar outputs. One contributed professional expertise throughout. The other accepted what the AI generated with little input and even less editing. Without evidence of the value added, you cannot tell them apart. The difference between them is exactly what competence means at this level.

The bar here is deliberately higher than "retaining control." Simply being present while AI does the work is not enough. The expert needs to be adding something: judgement, context, challenge, professional knowledge that changes the outcome. Expert in the loop, value in the loop, effort in the loop.

Level 3: Hands Off, AI Controlled. AI performs the task while the person actively monitors. The assessment should validate that they can check AI output, identify errors and take corrective action. The evidence needs to show how they recognised the problem, not just that they found it.

Think of the pharmacist reviewing AI-generated dispensing recommendations, or the engineer monitoring an AI-controlled process. They are not doing the task. They are watching it being done and deciding whether the output is safe.

Level 4: Eyes Off, AI Managed. AI performs the task without continuous supervision. The person checks in periodically and intervenes when required. The assessment should validate exception handling: the ability to recognise something has gone wrong, investigate and act under pressure.

The difference between Level 3 and Level 4 is attention. At Level 3, you are watching continuously. At Level 4, you have looked away, focusing on another task, trusting the AI. The assessment has to prove you can step back in when the situation demands it, and that you can do so with the reasoning to back up your decisions.

Level 5: Brain Off, AI Delegated. The task is fully delegated to AI. The person sees the results. The assessment should validate that they understand the output, its consequences and what to do when the process surfaces a problem. At this level, the reasoning is the competence.

Read across the five levels and two progressions run in parallel.

The person's relationship with AI: I do it. AI helps me. AI does it and I watch. AI does it and I check in. AI does it all and I review the results.

The assessment response: do it, use it, check it, step in, understand and use the result.

The levels are not a league table. Level 5 is not worse than Level 1. A senior clinician may operate at Level 4 for a task a trainee must perform at Level 1. The clinician is not less competent. The evidence required is simply different.

Why a task sits where it does

Placing a task on the grid is the first step. The framework then asks why it sits there, because the reason affects the evidence required and how stable the placement is over time.

A task might need human involvement because the consequences of failure are severe, or because doing the task is the very foundation of true understanding. Because the person learns something essential by doing it themselves. Because others value the fact that a human did it. Or simply because AI cannot yet do it reliably.

That last reason behaves differently from the others.

Consequence does not change as AI improves. But capability gaps close. If a task sits at Level 1 only because AI is not good enough yet, that placement has a shelf life. Any assessment designed around a capability gap needs a review point built in.

Most tasks sit where they sit for more than one reason. Identifying all of them strengthens the justification and makes it clearer which parts of the placement are durable and which will need revisiting as AI develops.

The grid is the starting point, not the whole answer

The full framework includes a task analysis loop, a resilience test for checking whether an assessment format can support what it claims to validate, and a set of questions about what humans lose when a task is automated. Those elements are available in the downloadable grid, which is designed to be used in practice rather than read as theory.

This article is the first in a series. The next will deal with the evidence problem: why capturing what someone did and what they produced is not enough, and what a third layer of evidence makes possible. Later pieces will explore what is lost when tasks are handed to AI, and what recent research into AI-enabled cheating means for assessment format and design.

The grid is deliberately unfinished. Different professions, different sectors, different regulatory environments will stress-test the levels in ways I cannot do alone. If it helps assessment organisations ask better questions about what they are actually validating, it has done its job. If others improve it, even better.

The question that started this remains open. Much of the work we do may become hands off. But it must remain eyes open and brain engaged.

Tim Burnett is the founder of the Test Community Network. The AI Reliance and Assessment Resilience Grid is available to download from testcommunity.network.

This is Article 1 of a four-part series. Article 2 will explore why process and product evidence are not enough, and what a third layer of evidence makes possible.

Read the full Assessment in 2040 voices article >