As AI Shapes Assessment, What Should We Pay Attention To?
Author Jilliam Joe, Ph.D.
Artificial Intelligence 6 min read

FullScale is learning alongside K-12 networks that are exploring how AI can support teaching and learning that is personalized to students’ identities and strengths, competency-based, and whole-child in meeting individual students’ needs. AI in the context of assessment, therefore, is a particularly relevant topic for us. In this blog, I consider what the growing use of AI means for assessment practice. 

AI has moved from the periphery to the center of assessment discourse. Long before the arrival of Claude or ChatGPT, AI operated quietly in the background of adaptive learning technologies and large-scale assessments, supporting automated scoring. Now, as generative AI makes its way from journal articles and technical reports to tablets and Chromebooks in everyday learning, the question is no longer whether AI will be used in teaching and learning. Instead, we need to ask: What should we pay attention to as AI plays a greater role in what counts and is trusted as evidence of learning?

This question guided the discussions at the 51st Annual International Association for Educational Assessment (IAEA) Conference in Toronto, Ontario, Canada, where the theme—Trust, Transparency, and Technology in Educational Assessment—invited us to examine the complex relationship between human intention and technological innovation. Trust and transparency are not abstract ideals; they are practical necessities that prompt us to scrutinize both the architects of new technologies and the technologies themselves. As the presenters reflected on what constitutes trustworthy and valid assessment in an era shaped by AI, I observed three central issues.

As AI changes how students demonstrate learning, what happens to the claims we make about what learners know and can do?

The opening plenary with Eunice Eunhee Jang (Professor, Ontario Institute for Studies in Education) prompted the idea that a student’ use of AI during a demonstration of learning does not automatically invalidate a performance as representative of a learner’s knowledge or skills. Instead, it complicates the conclusion or claim about what skills and knowledge are actually represented. Consider what that means as AI becomes embedded as a feature of learning and assessment, rather than a novelty. If a student uses generative AI to brainstorm project ideas, receive feedback, translate and process complex information, revise writing, solve part of a problem, or produce a final project in the sciences and arts, what knowledge and skills does the resulting performance represent? There is unlikely to be one answer that applies across every assessment or content domain. The more useful questions are familiar ones that must be addressed before the assessment is built, before scores are interpreted:

What construct(s) are we trying to measure? What evidence do we need? How much of that evidence is sufficient? What role do key conditions like AI use play in the trustworthiness of that evidence? And what claims can we reasonably make as a result?

The answers to these questions matter to both the teacher interpreting their students’ score reports and to the state leader evaluating assessment and accountability policies to support student achievement across complex and diverse learning environments.

As assessment technology advances, who benefits?

Technology can expand how learners access assessments and demonstrate what they know, but those benefits are not necessarily experienced equally. At IAEA, conversations about building trust in AI were held alongside attention to Universal Design for Learning and culturally responsive and sustaining educational approaches to assessment. Confidence Dikgole (Chief Executive Officer, Independent Examinations Board, South Africa) posed a particularly useful question: When we introduce technological advances in assessment, are we also introducing new sources of inequity? Even something as basic as whether a technology depends on reliable high-speed connectivity or devices capable of supporting new technologies can affect who can use it as intended. 

As the field explores ways to advance assessment through AI, accessibility and fairness cannot be downstream checks. We should be asking from the beginning: Who was this assessment designed for? Under what conditions does it work? Whose ways of demonstrating knowledge and in what contexts does it enable or preference? Attention to these considerations heightens the importance of student, educator, and community voices as part of the assessment design process. 

As AI use expands, are we building the literacy needed to use it responsibly in assessment?

Assessment literacy was also a recurring theme at IAEA. AI use, however, adds another layer of demand for literacy. People may need to understand not only what a score or performance means, but also where AI was involved (such as in scoring), what it contributed (such as in feedback to learners), what limitations it introduces, and what conclusions the resulting evidence can reasonably support. That makes assessment and AI literacy part of the infrastructure needed for responsible innovation. The challenge for the field is therefore not only to build better tools, but also to build people’s capacity to understand and use them well. And that capacity must extend beyond assessment professionals to students, families, educators, leaders, policymakers, and employers as competency-based assessments are used to credential the skills and competencies valued in the workforce.

Why These Questions Matter

Assessment, when done well, is an important source of evidence of learning, instructional quality, and how well educational systems are meeting the needs of individual learners. Assessment systems can do something even more consequential: impact access to opportunity, as they are often a powerful gatekeeper to the futures young people and adults imagine for themselves. Therefore, the choices we make now about AI in the context of assessment are more significant than ever. And while the technology will continue to evolve, the questions that should guide us are far more enduring:

  • What are we measuring, not just what did we intend to measure? 
  • Who has meaningful opportunities to demonstrate what they know and can do?
  • Do the people using assessment information understand what it can and cannot tell them?
  • Are we designing assessment systems in ways that earn the trust of the communities they serve?

These aren’t questions for assessment experts alone. They are questions for anyone involved in designing, using, or being affected by assessment. As AI continues to play a critical role in shaping the now and future of assessment, the field has an opportunity to be equally intentional about what should not change: our attention to validity, access, responsible interpretation, and the consequences assessments have for learners.

Explore Related Work 

AI Literacy: AI literacy emerged as an important consideration for district design teams participating in FullScale’s Rural AI Strategy Lab as they evaluated and selected AI tools. Teams worked through an adapted version of this EdTech Systems Guide’s selection module to consider which tools best aligned with their needs and priorities. FullScale’s AI tool selection guide will be released in January as part of the complete Rural AI Toolkit.

Digital Divide: See what guidance FullScale has shared about the persistent divides in who benefits from technological advances, and who does not (see From Digital Access to Digital Equity: Critical Barriers That Leaders and Policymakers Must Address to Move Beyond “Boxes & Wires”).

Related Posts

From Ideas to Action: Rural AI Strategy ...

The Sound We Keep Missing: What Radical ...

Attendees talking in group excitedly

Functions of State K-12 Leadership: What...

About the Author

Jilliam Joe, Ph.D. Dr. Jilliam N. Joe is the Managing Director of Evaluation and Measurement at FullScale, the national nonprofit formed by the […]

Read more from Jilliam Joe, Ph.D.
Scroll to Top