AI in assessment design
TLDR
AI in assessment design is about using AI to shape assessment tasks so they still measure the intended construct while better reflecting real-world practice. The central question is not whether AI is present, but whether the resulting task still produces valid, authentic, and defensible evidence of learner performance. Stronger sources point towards explicit design choices, clear boundaries around permitted AI use, and more visible process evidence rather than reliance on detection alone. The main unresolved issue is where AI support ends and the assessed performance begins. New sector signals, including educator-growth recognition and national AI literacy initiatives, reinforce that assessment design now sits inside a broader AI ecosystem rather than a narrow classroom question.
Definition
AI in assessment design refers to using AI to shape, adapt, or support assessment tasks so that the task better matches the intended construct, context, and real-world use of the skill. The key issue is whether the assessment still produces valid, authentic, and defensible evidence of learner performance. Prompt engineering, model choice, and governance now matter because they influence how reliably AI can be used in design workflows.
Why It Matters
Assessment teams are increasingly facing a design question as well as a control question: if AI is part of working life, should some assessments allow learners to use it in a structured way? That can improve realism and future relevance, but it can also weaken comparability if the role of AI is not defined clearly. The deeper assessment issue is deciding when AI should be ignored, prohibited, permitted, or deliberately built into the construct. Recognition and national AI programmes matter here because they suggest the wider education system is normalising AI capability faster than many assessment policies are changing.
The source set also suggests that process evidence, such as drafts or revision history, may help make learner judgement visible in AI-rich tasks. That matters because the more AI is involved, the more important it becomes to show where the learner’s own contribution sits.
Key Concepts
- **Authentic assessment**: a task that reflects how a skill is used in practice.
- **Construct**: the capability the assessment is meant to evidence.
- **Permitted AI use**: AI support that is allowed because it does not undermine the assessment purpose.
- **AI-integrated task**: an assessment designed so that some AI use is part of the expected performance.
- **Process evidence**: drafts, revision history, checkpoints, or other traces that make learner judgement visible.
- **Prompt engineering**: shaping prompts so that AI output is more useful, reliable, or constrained.
- **Humanity in the loop**: keeping human judgement central in AI-supported educational design and decision-making.
What Experts Agree On
The strongest evidence points towards a common view: AI should be handled as a design issue, not only as a misconduct issue. Across academic, practitioner, and policy-facing material, the recurring theme is that assessment validity depends on being explicit about what AI is doing in the task and why that is acceptable.
There is also broad agreement that task design matters more than post-hoc anxiety about AI use. Sources repeatedly point towards clearer boundaries, learner-facing expectations, and evidence of process as more useful than trying to infer authorship after the fact.
A further convergence is that human judgement remains central. The more contemporary material does not argue for removing people from assessment decisions; it argues for defining which parts of the workflow can be AI-assisted and which parts must remain human-led.
What Is Contested
What remains unsettled is how far authenticity can be preserved when AI becomes part of the workflow rather than just an external tool. Some designs may become more realistic, while others may drift away from the intended learning outcome. The open question is which parts of the assessment should remain strictly independent and which may legitimately include AI support. Recognition can pull the conversation towards innovation, but it does not itself answer the design question.
There is also a live tension between pedagogy, employability, and operational convenience. Those aims are not always aligned, and the evidence base does not yet show a single settled model across sectors. National-level AI rollout adds another layer of uncertainty because it changes learner expectations faster than assessment design can adapt.
Risks
- Assessment tasks may drift away from the intended construct.
- Learner support may become hidden assistance.
- Policy may be too vague to guide staff on permitted AI use.
- Comparability may weaken if AI-integrated tasks are not carefully designed.
- Over-reliance on after-the-fact detection can miss the design problem.
- Recognition, awards, or product launches may be mistaken for validation.
Good Practice
A practical decision framework is to start with the construct and work forwards.
1. Define the skill or knowledge the assessment is meant to evidence.
2. Decide whether AI is being used in task design, learner support, or the task itself.
3. Set the permitted role of AI in plain language.
4. Build process evidence where learner judgement matters.
5. Check whether the task still produces comparable evidence across the cohort.
6. Ask how the design will be explained to candidates, staff, and regulators.
7. Review the assessment after pilot use and adjust the design if AI changes the meaning of the evidence.
Where design is procedural, the key is to tell staff what to do. If a task needs a viva, a draft history, or a sign-off step, say so explicitly and explain why. If AI is permitted only for preparation or editing, state that boundary clearly.
Options or Comparison
### Common assessment stances on AI use
| Option | What it means | Main benefit | Main concern |
|---|---|---|---|
| Prohibit AI | Learners must work without AI support | Clear independence signal | Can be hard to enforce and may not fit real practice |
| Permit AI with controls | AI is allowed in defined ways | More realistic and flexible | Needs clear rules, disclosure, and review |
| Integrate AI | AI use is part of the assessed capability | Better fit for AI-rich work | Harder to compare performance across learners |
| Design for process evidence | The task requires drafts, checkpoints, or justification | Makes learner judgement visible | Can increase workload |
### Comparison of design approaches
| Design approach | What it optimises | Main value | Main risk |
|---|---|---|---|
| Final-product only | Convenience and speed | Easy to run | Harder to know who did the thinking |
| Process-rich assessment | Visibility of judgement | Stronger authenticity | More marking and management effort |
| AI-integrated construct | Real-world AI use | Better alignment with practice | Needs very clear evidence criteria |
Example in Practice
A university module team wants students to use AI to support research and drafting, but not to replace independent reasoning. It rewrites the assignment so students must submit a short draft log, a justified source selection note, and a viva-style reflection on one key decision. That preserves the learner’s contribution while still allowing AI use in a controlled way.
Key Sources
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- The AI Assessment Scale Revisited.
- Anthropic education report on how university students use Claude.
- RMIT source note on contextual AI assessment design.
Vendor Landscape
The vendor footprint is strong and varied. Suppliers increasingly present AI as a design, drafting, or optimisation layer rather than only a detection threat. That means assessment teams need to be careful not to adopt product language too quickly; the question is whether the tool changes the evidence claim, not whether it sounds innovative.
FAQs
### Can AI be part of assessment design without hurting validity?
Yes, if the task is designed so AI use fits the intended construct and the learner’s own contribution remains visible.
### What is the difference between AI-assisted and AI-integrated assessment?
AI-assisted assessment uses AI to support learning or production around the task; AI-integrated assessment expects some AI use as part of the performance being judged.
### Should universities redesign coursework because of AI?
Often yes, but the right redesign depends on the construct. Not every task needs more AI; some need clearer process evidence or more direct demonstration of judgement.
### What should design teams ask first?
What is the assessment meant to prove, and what role should AI play in producing that evidence?
Last Reviewed By
Tim Burnett (Admin)
Suggested Citation
Test Community Network. "AI in assessment design." TCN AI & Assessment Wiki. Last reviewed 2026-08-12. https://www.testcommunity.network/wiki/ai-in-assessment-design.html
Sources
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- The AI Assessment Scale Revisited.
- Anthropic education report on how university students use Claude.
- RMIT source note on contextual AI assessment design.
Sources
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Anthropic education report on how university students use Claude.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Cambridge Assessment Network workshop on authentic assessment in the age of AI.
- Microsoft
- Thelearningawards
- The AI Assessment Scale Revisited.
- Anthropic education report on how university students use Claude.
- Anthropic education report on how university students use Claude.
- Deeplearning
- Inside Higher Ed
- Link
- The AI Assessment Scale Revisited.
- The AI Assessment Scale Revisited.
- The AI Assessment Scale Revisited.
- RMIT source note on contextual AI assessment design.
- Jiscpodcast
- Anthropic education report on how university students use Claude.
- Anthropic education report on how university students use Claude.
- The AI Assessment Scale Revisited.
- RMIT source note on contextual AI assessment design.
- President
- Thelearningawards
- Digitaleducation
- Digitaleducation
- The AI Assessment Scale Revisited.
- President
- RMIT source note on contextual AI assessment design.
- President
- Thelearningawards
- Sites
- RMIT source note on contextual AI assessment design.
- Thelearningawards
- Test Community Network
- Inside Higher Ed
- Sites
- President
- Link
- Digitaleducation
- Digitaleducation
- Media