Skip to content

Assessment that resists gaming: multi-axis rubrics for adult learners

گیمنگ سے بچنے والی جانچ: بالغ سیکھنے والوں کے لیے کثیر محور روبرک

35 min read

Three ways to see it

  1. Adults game any single-axis assessment within two cohorts. If the only thing that counts is the multiple-choice score, they memorise the question bank. If the only thing that counts is the prompt count, they paste throwaway prompts. If the only thing that counts is hours logged, they stay logged in while doing other work. The defence is multi-axis assessment, where a learner has to perform on three or four uncorrelated dimensions to pass. Gaming one dimension does not help, because the others will catch them. The AIzaadi assessment uses four axes: comprehension, application, metacognition, and pedagogical transfer.

  2. Way one: comprehension is measured by oral explanation, not by writing. The learner sits across from the trainer for five minutes and explains a concept of the trainer's choice. The trainer asks one why-question per minute. A learner who memorised will collapse by the second why. A learner who understood will reach the third level. This is hard to scale, which is exactly why it works as an anti-gaming layer. You can run twenty-five oral check-ins in two hours if you batch them in pairs and let the partner act as scribe.

  3. Way two: application is measured by a real workplace artefact. Not a contrived case study, a real one the learner brings from their own job. A FBR officer brings a draft tax circular. A teacher brings a lesson plan. A bank operations manager brings a customer complaint email. The learner uses AI to improve the artefact in front of you for forty-five minutes. You score on four sub-axes: did the prompt match the goal, did the output need revision, did the learner spot the errors, did the final artefact actually improve. The artefact must be returnable to the workplace immediately. If it is not workplace-grade, the learner has not passed application.

Quick check

Quick check: what makes modern AI different from a rule-based program?

The why-tree

Why-tree level one: why does multi-axis beat single-axis? Because correlation is the gaming signal. If two axes are correlated, gaming one accidentally helps the other. Uncorrelated axes punish gaming because effort spent on one is effort lost on the others.

Try this with Claude

AI-edge prompt: 'I am building an anti-gaming assessment for a 21-hour AI literacy cohort of mid-career civil servants in Pakistan. Design four uncorrelated axes (comprehension, application, metacognition, peer-teach) with rubrics, sample tasks, and scoring criteria. For each axis, describe two specific gaming attempts a clever learner might make and how the design defeats them.'

Sources

Sources and further reading. Wiggins and McTighe, Understanding by Design. Black and Wiliam, formative assessment research papers. NAVTTC competency-based assessment framework. Pakistan Engineering Council outcome-based education guidelines. AKU Examination Board assessment design notes. Lessons from Codex gaming patterns observed in Polymath managed-agents work. UK Quality Assurance Agency authentic assessment guidance.