Implementation guide
Practical guidance for implementing the AI Assessment Scale effectively in your institution, faculty, or classroom.
The AIAS is an assessment design tool, not a security tool.
Focus on redesigning tasks to match your chosen level. Labels alone are not enough: our implementation research finds that the scale can become a compliance layer when it is disconnected from learning outcomes, disciplinary context and staff capacity (Perkins et al., 2026). For a guided version of everything on this page, use the AIAS Custom GPT.
Before you start
Conduct a validity audit +
Before implementing the AIAS, audit your existing assessments for validity in the context of generative AI:
Content validity: Does the assessment still measure what it claims to measure if students use AI? If an essay tests writing ability but AI can write it, the content validity is compromised.
Construct validity: Does the assessment actually measure the underlying skill or knowledge you care about? Consider whether AI use would undermine or support this.
Key questions to ask:
- What skills and knowledge does this assessment actually measure?
- Would AI use invalidate these measurements?
- Can the assessment be redesigned to remain valid with appropriate AI use?
- Is there external validity (does this reflect real-world practice)?
Work at the faculty level +
The most effective AIAS implementations happen at the faculty or discipline level, not through individual instructors working alone or institution-wide mandates.
This “leading from the middle” approach works because:
- Discipline-specific needs can be addressed appropriately
- Colleagues can share approaches and learn from each other
- Students experience consistency across their programme
- Assessment mapping becomes more coherent
- Local champions can drive change effectively
Consider your discipline context +
Different disciplines have different relationships with AI. Consider:
- Professional practice: How is AI used in the profession your students are entering?
- Core competencies: What skills must students demonstrably possess themselves?
- Assessment traditions: What forms of assessment are typical in your field?
- Accreditation requirements: Are there external requirements that constrain your choices?
A nursing programme may need more Level 1 assessments for clinical skills, while a digital marketing programme might emphasise Levels 3-4 to reflect industry practice.
Plan for more than one cycle +
It helps to be realistic about what a first year of using the scale delivers. In our own implementation studies, the reliable early win is a shared vocabulary: AI use becomes something staff and students can discuss openly, and students report much greater clarity about what is expected of them.
Two things take longer. Students understand the levels less well than they understand the expectations, so clarity arrives before comprehension. And genuine task redesign, as opposed to attaching a level to an existing task, tends to happen only where staff already have some confidence with the tools. Redesign does not follow automatically from adopting the framework; it needs support to activate.
Plan accordingly. Treat the first cycle as establishing language and surfacing practice, and schedule the redesign work, with staff development behind it, for the cycle after.
Key principles
Structural vs discursive changes +
The distinction between structural and discursive change comes from Corbin, Dawson and Liu (2025), and it is the most important principle for putting the AIAS into practice. We take it up in How (not) to use the AI Assessment Scale.
Discursive changes are statements or labels: “This is a Level 2 assessment” or “You may use AI for planning only.” Research shows students often ignore these if the task structure does not match.
Structural changes redesign the task itself so that the appropriate level of AI use is built into the assessment mechanics:
- Redesigned briefs that specify what students must produce at each stage
- Rubrics that assess the skills appropriate to the level
- Required evidence (drafts, prompt logs, reflections)
- Checkpoints that make the process visible
- Task designs that only work with the intended level of AI use
Replace detection with design +
AI detection tools are unreliable and create adversarial dynamics. Instead of trying to catch inappropriate AI use after submission, design assessments where:
- The appropriate level of AI use is built into the task structure
- Evidence of the student’s process is required as part of the submission
- The assessment only works when completed as intended
- Multiple assessment points build a picture over time
This shifts the focus from policing to pedagogy, and from catching cheaters to supporting learning.
In practice, staff who make this move settle on visible process as the evidence of permitted use: prompt submission, learning logs, document history. What none of them claim is the ability to verify which level a student actually worked at. Design on the assumption of use, and treat process evidence as the integrity mechanism rather than something to check up on.
The key question +
“Is this a bad use of technology, or a bad use of your brain?”
When deciding on the appropriate AIAS level, ask whether using AI would:
- Undermine learning: If the skill you are trying to develop requires practice without AI assistance, then Level 1 may be appropriate (but only in controlled environments)
- Reflect professional practice: If professionals in the field use AI for this type of task, then higher levels may better prepare students for their careers
- Add genuine value: AI should enable students to achieve better outcomes, not just make tasks easier
Secure and non-secure conditions +
The five levels describe the role AI plays in a task. Whether the task runs under secure, controlled conditions is a separate question, and keeping the two apart is what makes the AIAS work as a design system rather than a set of rules.
The padlock in the AIAS artwork sits on Level 1 only. That is deliberate: a No-AI task is only valid where AI-free conditions can genuinely be assured, so security is built into what Level 1 means. The padlock is not shown on the other levels, because a padlock on “Full AI” would read as though AI were restricted there, which is the opposite of the intent.
Any level can be run under secure, supervised conditions if the task calls for it. Level 1 is the only level that must be. A Level 4 task on locked-down lab machines, where students direct approved AI tools under invigilation, is a coherent and potentially useful design, so choosing a secure environment never obliges you to drop to Level 1.
Two decisions, taken separately:
- Which level? What role should AI play, and what is the task assessing?
- Secure or not? Can the environment be controlled, and does this task need it?
Take the two decisions in order. The conditions tell you what is enforceable; the level tells you what the task is built to assess once those conditions are set.
Implementation steps
Step 1: Audit current assessment validity +
Review each assessment in your module or programme and ask:
- What does this assessment actually measure?
- Would AI use compromise this measurement?
- Can students complete this with AI in ways that undermine learning?
- Is the current format still fit for purpose?
Step 2: Decide the appropriate AIAS level +
For each assessment, determine which level best fits your learning outcomes:
- Level 1: Only if you can enforce it in a controlled environment
- Level 2: AI for planning and research, independent execution
- Level 3: AI assists specific tasks, student evaluates and modifies
- Level 4: AI may complete any elements, student directs
- Level 5: Student explores novel AI applications
Consider how assessments across the module or programme create a balanced picture of student capability. The AIAS Custom GPT can help you weigh the options for a specific task.
Step 3: Redesign task structure +
Make structural changes that embed the level into the assessment:
- Rewrite the assessment brief with clear expectations for each stage
- Update rubrics to assess the skills appropriate to your chosen level
- Specify required evidence (process documentation, drafts, reflections)
- Build in checkpoints that make student thinking visible
- Design tasks that only work when completed at the intended level
Step 4: Build evidence over time +
Use the “Swiss cheese” approach: multiple assessment points at various levels build a fuller picture of student capability than any single assessment could provide. The metaphor comes from Reason’s model of layered defences, where no single slice is sound but the holes rarely line up. Phillip Dawson and colleagues have applied it to assessment security, and it is a more useful mental model than trying to make one task airtight.
For example, a module might include:
- Level 1 in-class test (controlled environment)
- Level 2 research proposal (AI for planning)
- Level 3 draft with peer review (AI for refinement)
- Level 1 oral defence of final work
A more structured version of this is the assessment twins approach: two deliberately linked components that address the same learning outcomes through different kinds of evidence, scheduled close together so each can be read against the other. Pairing an open task with a secured one is the most common form, and it gives you cross-verification without making the whole assessment secure.
Step 5: Communicate clearly to students +
Clear communication reduces confusion and supports student success:
- Explain which AIAS level applies and what this means in practice
- Provide examples of appropriate and inappropriate AI use
- Clarify what evidence students must submit
- Explain why these choices have been made (connect to learning outcomes)
- Ensure the task structure reinforces the message (not just labels)
- Give students their own guidance, not only staff. Where understanding of the levels reaches students entirely through their teachers, comprehension lags well behind clarity. Plain-language student materials close that gap directly.
Step 6: Address equity issues +
This matters more than it might appear. Where we have surveyed students, whether they believe the scale makes assessment fairer is the strongest single predictor of whether they support it at all, and fairness rests on provision decisions taken outside the framework: devices, accounts, age-appropriate tools and consistent application across a class. Settle these before the first task goes out:
- If AI tools are required, guarantee free access for all students
- Consider institutional subscriptions or design for freely available tools
- Be transparent about which tools are permitted
- Distinguish AI tools from assistive technologies
- Provide support for students with varying AI literacy levels
If you require AI use but cannot guarantee access, you are creating an inequitable assessment.
Step 7: Build faculty capability +
This is the step that determines whether the scale changes task design or only task labelling. Where staff have some prior familiarity with the tools, they reach the upper levels with intent; where they do not, adoption tends to stop at the lower levels and at relabelling. Budget for it accordingly rather than treating it as optional. Effective implementation requires staff development in:
- Understanding generative AI capabilities and limitations
- Assessment design principles for the AI era
- Practical experience with AI tools students are using
- Sharing approaches across the faculty
- Ongoing learning as technology evolves
Common mistakes to avoid
Unenforceable “No AI” policies +
The mistake: Applying Level 1 to take-home work or unsecured environments.
There is no realistic way to prevent AI use outside controlled environments. Policies that cannot be enforced create validity problems and inequity (students who follow the rules are disadvantaged).
Instead: Reserve Level 1 for genuinely controlled environments (exams, in-class work, observed assessments). For unsecured work, choose Levels 2-5 with appropriate structural design.
AI for AI’s sake +
The mistake: Mandating AI use without clear pedagogical purpose.
Requiring AI use just to appear innovative, without connecting it to learning outcomes or professional relevance, does not serve students well.
Instead: Enable AI where it adds genuine value and reflects professional practice. Always connect AI use to learning outcomes and external validity.
Labelling without redesign +
The mistake: Adding AIAS level labels to existing assessments without structural changes.
Simply labelling an assessment “Level 2” without changing the task structure is a discursive change, in Corbin, Dawson and Liu (2025)’s terms. Students will ignore labels that do not match task requirements.
Instead: Redesign the brief, rubric, required evidence, and checkpoints so the assessment structure embeds the intended level.
Relying on detection tools +
The mistake: Using AI detection tools as the primary enforcement mechanism.
AI detection tools are unreliable (high false positive rates, bias against non-native speakers, easily bypassed) and create adversarial dynamics.
Instead: Replace detection with design. Build appropriate AI use into task structure rather than policing it afterward.
Ignoring equity concerns +
The mistake: Requiring AI use without ensuring equitable access.
Students have varying access to AI tools (paid subscriptions, quality of free tiers) and varying AI literacy levels.
Instead: Plan for equitable access. Guarantee free tool access if required, design for freely available resources, and provide support for developing AI literacy.
Sector-specific guidance
Higher education +
- Identify discipline champions to lead implementation within faculties
- Balance top-down policy guidance with bottom-up assessment design
- Design authentic tasks that reflect professional practice in the discipline
- Use the Swiss cheese approach across modules and programmes
- Consider accreditation requirements and professional body expectations
- Build staff development into implementation plans
K-12 education +
- The Swiss cheese approach is particularly important as students develop foundational skills
- Avoid assessment designs that drift entirely toward examinations
- Consider developmental appropriateness when introducing AI tools
- Ensure age-appropriate AI literacy education accompanies tool use
- Communicate clearly with parents about AI policies
- Do not assume student AI fluency. Where we have looked, students know one or two mainstream tools and little beyond them, so effective use has to be taught rather than presumed.
Primary/elementary contexts will typically use more Level 1, while secondary students may engage with higher levels as appropriate.
TVET and vocational education +
- Practical assessments and demonstrations are naturally less vulnerable to AI misuse
- Level 1 works well for observed practical skills that must be demonstrated in person
- Level 3 may suit written components where AI can assist with documentation
- Consider how AI is actually used in the trades and industries you are preparing students for
- Focus on competency-based assessment that requires demonstration of practical skills
English as a Foreign Language (EFL) and English for Academic Purposes (EAP) +
- Level 1 is often appropriate for core language production skills (writing, speaking) that students must develop themselves
- Levels 2-3 may suit broader academic tasks where AI can assist with research or editing
- Distinguish between assessing language proficiency and assessing content knowledge
- Be aware that AI detection tools have higher false positive rates for non-native English speakers
- Consider how AI translation and writing tools are used professionally
Two published adaptations give detailed guidance for these contexts: the EAP-AIAS for English for Academic Purposes, and the AIAS for EFL writing and translation.
Choosing the right level
Level selection questions +
Ask these questions when selecting an AIAS level:
- What are the learning outcomes? What skills and knowledge must students demonstrate?
- Would the assessment remain valid with AI? Does AI use undermine what you are measuring?
- Can you enforce the restriction? If you cannot genuinely control AI use, do not prohibit it.
- Does it fit the programme? How does this assessment relate to others across the module?
- Is there equitable access? Can all students participate fairly at this level?
- Does it reflect professional practice? How would professionals approach this task?
Example assessment sequence +
A well-designed module might include multiple assessments at different levels:
Level 1 Week 4: In-class test to establish baseline understanding
Level 2 Week 6: Research proposal with AI-assisted literature review
Level 1 Week 8: Draft written independently without AI
Level 3 Week 12: Final submission with AI-assisted refinement and reflection
This sequence builds evidence of student capability across multiple points and levels, so no single point of AI misuse distorts the overall picture.
Implementation checklist +
Cover as many of these as your context allows. Few teams will manage all of them, and the AIAS still works when some are out of reach, but the more you can put in place, the better the implementation tends to hold up.
- ☐ Validity audit completed for existing assessments
- ☐ Faculty or discipline colleagues engaged
- ☐ AIAS levels selected for each assessment
- ☐ Task structures redesigned (not just labelled)
- ☐ Student guidance materials prepared
- ☐ Evidence requirements specified
- ☐ Equity issues addressed (tool access, AI literacy support)
- ☐ Staff development provided
- ☐ Assessment sequence builds dependable evidence over time
Ready to get started?
Work through a task with the Custom GPT, or explore the research behind the framework.
If your institution wants support putting this into practice, Learning Innovation Practice runs assessment reviews, policy development, and staff workshops built around the AIAS. Get in touch.
The AI Assessment Scale (AIAS) was originally developed by Mike Perkins, Jasper Roe, Leon Furze and Jason MacVaugh, and is owned and maintained by Learning Innovation Practice Ltd. AIAS v2.1 © 2026 Learning Innovation Practice Ltd. Based on the original AIAS, developed by Perkins, Furze, Roe and MacVaugh. Licensed under CC BY-NC-SA 4.0.
