Can AI Write Exam Questions?

What AQA's Research Gets Right, and the One Thing It Misses

Published: 7/16/2026
Can AI Write Exam Questions?

TL;DR: AQA's report Evaluating AI‑Authored Assessment Items for High-Stakes Exams is a genuinely valuable piece of research — strong on command words, item difficulty and the idea of "surface plausibility". But it generated its items with AI more or less out of the box, and that's the crux: with AI item authoring, the process (the prompts, the expert input, the source anchoring) matters more than the model. Treat generative AI as a drafting accelerator, not an autonomous author.

What did AQA's research actually find?

AQA set out to test how AI-generated assessment items stand up against human-authored ones. They built a GCSE-level biology mock, mixed AI items with human items, had students sit them, and asked teachers to judge which was which and to spot flaws.

The findings worth holding onto: one-mark items fared noticeably better than three-mark items, mark schemes emerged as one of the biggest risks in the process, and the report's conclusion landed exactly where it should.

"Generative AI is best treated as a drafting accelerator within a controlled assessment lifecycle, not as an autonomous author."
— AQA, Evaluating AI‑Authored Assessment Items for High-Stakes Exams

So can you trust AI to author exam questions "out of the box"?

This is where the one nagging question sits. The research generated its items with generative AI used straight out of the box, and to my mind, how you create an item is fundamental to whether it's any good.

"It's a bit like asking people to compare holidays without any consideration of how they booked it, who with, how much they spent, or what research they did."
— Tim Burnett, Test Community Network

The prompts, the harness, the expert knowledge you feed in, that is the work. Judge the output without it and you're measuring the wrong thing. To be fair to AQA, they recognise this in the report; it simply wasn't their focus.

Why are mark schemes the biggest risk?

AQA are right that mark schemes are a major risk, but the problem cuts both ways. It isn't only that AI writes weak mark schemes; it's that human mark schemes are often vague, and that vagueness trips the model up too.

"We've tried to use human-authored mark schemes in the AI space to mark items, and the AI gets it wrong, because the mark scheme is so vague in places."
— Tim Burnett, Test Community Network

A mark scheme is supposed to hold all the expertise on its own. When it quietly assumes a human expert will fill the gaps, it becomes too subjective for a model to apply reliably.

What is "surface plausibility" and why should reviewers worry about it?

The term I liked most in the report is surface plausibility. AI items can look highly polished and read confidently, which is exactly the trap.

"The risk is that reviewers get weak at spotting poor quality, because it looks so polished. It can be confidently wrong, and you don't spot that from the polish."
— Tim Burnett, Test Community Network

How do you actually author good items with AI?

Here's what tends to separate a workable process from a bottleneck factory of plausible-but-weak items:

Capture expert insight early. Models are full of correct information, which is precisely why they're poor at distractors. Feed in what candidates plausibly get wrong before you generate, not after.

Use personas for the minimum competent candidate. "Create me a Level 2 question" invites vagueness; a persona anchors the level and context.

Anchor to specific source material. Point the model at the exact passage rather than handing it a 10,000-page workbook and internet access, that's how you avoid hallucinated references and hours lost tracing them back.

Check for test-wiseness. The classic tell is a fifteen-word correct answer sitting among one-word distractors. It should be a basic automated check; too many platforms skip it.

Put effort in the loop, not just an expert. One technique: have the SME type their answer to the generated stem before seeing the options. It surfaces single-best-answer problems and context mismatches early, and saves rework downstream.

Does the choice of AI model matter?

It matters a lot, and it's the part most people overlook. Many teams are defaulting to whatever instant model sits inside their office suite, missing the value of reasoning models, and unaware that some models suit some domains better than others.

That's why I built the MCQ Leaderboard: a community project that pits the available models against each other on item quality, so assessment professionals can see what actually performs — and drill down by topic.

What should assessment teams do next?

Three practical steps to try this week:

1. Read the AQA report for its command-word and difficulty analysis, the psychometric work is strong.
2. Redesign your authoring flow around early expert input, source anchoring and effort-in-the-loop, rather than one-shot generation.
3. Contribute to the MCQ Leaderboard so the model comparisons reflect real practitioner judgement.

Read the AQA report: Evaluating AI‑Authored Assessment Items for High-Stakes Exams

TCN guide: Understanding AI — A Guide for Assessment Professionals and Awarding Organisation Leaders (Edition 1, June 2026)

TCN guide: A Practical Guide to Authoring Test Items with Artificial Intelligence (Edition 1, March 2026)

Try the MCQ Leaderboard: leaderboard.testcommunity.network

Connect with Tim Burnett on LinkedIn: linkedin.com/in/tburnett