Why Better Assessment Data Isn't Improving Schools, and What Might Actually Close the Gap

Amina Afif on trust, technology, and the human capacity that assessment systems still depend on

Published: 7/28/2026
Why Better Assessment Data Isn't Improving Schools, and What Might Actually Close the Gap

Education systems around the world are generating more assessment data than ever before. Platforms are more sophisticated, item banks are richer, and AI is accelerating content creation at pace. Yet a persistent question hangs over the sector: if we have all this data, why aren't schools getting noticeably better at using it? In a conversation for the Test Community Network podcast, Amina Afif, statistician, education researcher, and Executive Secretary of FLIP+, offered a disarmingly simple answer: the problem was never really the data.

Watch the interview in full >

Afif has spent over twenty years working at the intersection of education policy, school improvement, and assessment in Luxembourg, where she is now based, and also in her native Seychelles. Her career has taken her through the OECD's PISA governing board, AEA-Europe, and the IEA, and she currently leads FLIP+, an international e-assessment association of twenty-two public assessment institutions spanning Europe, Brazil, and beyond. She also serves as the IAEA 2026 Conference Ambassador for Western Europe. What makes her perspective distinctive is that she has always sat between two worlds: the technical side of assessment design and the human side of what happens when that data reaches a school leader's desk.

"We can have very good assessment systems, but if teachers cannot interpret the results, if school leaders cannot use that evidence in a conversation, and if the reporting does not help learners understand where they are, I think we are missing something very big."
Amina Afif, Executive Secretary, FLIP+

What struck me about this observation is that it holds true regardless of a country's wealth. Afif has worked in both Luxembourg, one of the richest countries in Europe, and Seychelles, a small island nation in the Indian Ocean. Both are small populations. Both face remarkably similar challenges when it comes to translating assessment data into meaningful school improvement. Money helps up to a point, she suggested, but beyond a certain threshold, the gap between technically excellent data and practical use remains stubbornly wide. The question isn't whether the data is valid. It's whether the people receiving it have the capacity, confidence, and conditions to do something meaningful with it.

This is a thread that runs through much of the current debate in educational assessment, and it surfaces with particular urgency as AI enters the picture. Afif pointed out that we are now preparing students for jobs that do not yet exist, using tools that are evolving faster than curricula can keep pace with. Prompt engineering, she noted, was considered a highly skilled job not long ago, and already that has shifted. So what are we actually measuring, and does it still mean what we think it means?

"Technical quality and the meaning we give to assessment data cannot be competing priorities. They need both."
Amina Afif

The Technology Question: Ban It or Balance It?

The conversation turned naturally to one of the sharpest debates in education right now: whether technology, and AI in particular, belongs in schools at all. Norway announced in June 2026 that it would ban generative AI tools for primary-aged children from the new school year, following its earlier smartphone ban in 2024 that produced measurable reductions in bullying and improvements in grades. Sweden had already moved to ban mobile phones from compulsory schools, and Denmark has allocated significant funding to reintroduce printed textbooks.

Afif's response was characteristically balanced. She can understand why Scandinavian countries are pulling back, especially for the youngest learners. But she is uneasy with the idea that banning the tool addresses the real problem. The issue, she argued, is not the technology itself but the conditions under which it is used, and the absence of any shared playbook for using it well.

"Sometimes we are very quick to blame the tool, instead of thinking about what conditions are necessary. Do we have a playbook for how to best use that tool? We don't. We just use it. And then when it doesn't work, we say the tool is not good."
Amina Afif

She also offered a quietly powerful challenge to adults who blame children for excessive screen time: how much time do we spend on our own devices? Her three-year-old grandson, she recounted, was asked about a dinosaur and immediately suggested asking his mother to check on her phone. He was three. He had already internalised the behaviour he saw around him. The implication was clear: children are copying us, and if we want to change their relationship with technology, we might need to start with our own.

There are some schools, Afif noted, that are building AI systems designed to teach students rather than simply give them answers, models that guide thinking rather than replace it. This, she suggested, is the more productive conversation: not whether to allow AI in education, but how to create the conditions under which it supports rather than undermines learning.

FLIP+: What Happens When Countries Share Their Failures

Some of the most interesting territory in our conversation was around FLIP+, the international e-assessment association that Afif leads. Founded in 2017 when colleagues from France, Luxembourg, Italy, and Portugal sat around a table and recognised they were all solving the same problems independently, FLIP+ has grown to include twenty-two public institutions from countries including France, Italy, Norway, Lithuania, Brazil, Morocco, Ireland, Spain, the Netherlands, and England, with members such as AQA and the ERC among them.

The organisation's core mission is sharing assessment content, technology solutions and experiences. Not just success stories, but failures, and this, Afif argued, is what makes FLIP+ different from the typical conference circuit where vendors present polished narratives of commercial success. At its recent 9th FLIP+ event in Vilnius, a presenter opened her session by telling the audience she had failed, and then spent the rest of her time explaining why it had failed and what she planned to do next. That kind of candour, Afif suggested, is where the real learning happens.

"I think failure is where we grow. It's a platform. We fail fast and then we grow. So we allow people to come and see how they failed, and then we can all grow together."
Amina Afif

This matters particularly for countries that are earlier in their digital assessment journey. A developing nation looking at the UK or French system might see only the polished version and follow the same path, including into the same traps. FLIP+ tries to short-circuit that by making implementation lessons, procurement experiences, and even failed calls for tender available across its membership. The sharing extends beyond item creation to cover platform implementation, communication strategies, press engagement, and reporting approaches.

At the heart of FLIP+ is the item library project, the development of a shared bank of trusted digital assessment content that member institutions can contribute to, consult, and use. The ambition is significant: building a multilingual, cross-cultural resource that serves public education institutions who would otherwise be solving identical problems in isolation. Afif will be presenting a forty-five-minute session on the library at the IAEA 2026 conference in Toronto this September, where she plans to share the story of what FLIP+ has achieved over nine years, including, true to form, the challenges.

Trust as the Deciding Factor

Throughout the conversation, Afif returned to a single idea that underpins her work: trust. Technology can accelerate assessment, she argued, but trust determines whether people will actually use it, whether teachers believe the data is meaningful, whether school leaders feel confident acting on it, whether the public accepts that digital assessment is credible.

"I believe that technology can accelerate assessment, but trust will determine whether people use it smartly, whether they believe in it, and how students learn."
Amina Afif

This is not a new insight, but Afif's framing of it, grounded in two decades of cross-cultural experience, carries a particular weight. She is not arguing against technology or against sophisticated assessment systems. She is arguing that the human infrastructure around those systems deserves equal investment: capacity building, leadership development, transparency about how data is generated and what it can and cannot tell us. Without that, even the most technically excellent assessment is a missed opportunity.

Over time, her work with school leaders led her into leadership coaching and human-centred leadership development, a natural extension of her realisation that good data and strong assessment systems only create change when the people using them have the capacity, confidence, and trust to do so well.

Looking Ahead: IAEA 2026 in Toronto

Both Afif and I will be at the 51st IAEA Annual Conference in Toronto from 27 September to 2 October 2026, under the theme "Trust, Transparency and Technology in Educational Assessment." With a third of participants expected from low- and middle-income countries, Afif is particularly looking forward to hearing how different systems are connecting technical expertise with strategies for school improvement, and how they are building the human capacity to use assessment data meaningfully.

Her hope is that the human-centred part of the conversation takes just as much weight as the technological side. Given the pace at which AI is reshaping assessment, and the growing pushback from governments asking whether that pace is appropriate, it feels like a conversation whose time has very much arrived.

Connect with Amina Afif on LinkedIn.

Connect with Tim Burnett on LinkedIn.

Listen to the full conversation on the Test Community Network podcast.