Part 2 of What is Rigor Clearing Seven Misconceptions
Introduction
A teacher posts a math problem on LinkedIn and asks a simple question: “What DOK level is this?” It’s a good problem — three walkers, each with a distance and a pace, and the task is to rank them by how long they spent walking. It looks demanding. It has that higher-order shine.
Two frontier AI models weigh in. Both return DOK 3.
Jeffrey Ringer, the teacher who posted it, doesn’t buy it. He says what anyone who has interrogated a language model already knows: you can rarely trust the response. He won’t take the 3. Then Erik Francis — who wrote the book on Depth of Knowledge — settles it. It’s a DOK 2. You’re comparing and organizing information, not reasoning across sources toward a non-routine conclusion. And then the line that should stop every one of us cold: it’s higher-order thinking in complexity, Francis says, but it is not Depth of Knowledge in demand.
Two of the most capable AI systems on earth looked at a task and misread its demand. Sit with that. The seven misconceptions in Part 1 were all about human beings misreading rigor. This eighth one is different. It’s about misplacing it — and it’s the misconception the age of AI was built to produce.
Here’s the version I want to take apart: the assumption that when a student uses AI, the task’s rigor drops. It feels intuitive. The machine did the hard part, so the task must have gotten easier. Francis has warned about a nearby problem — that generative AI produces DOK examples that are inconsistent and inaccurate (Francis, 2026a), which is exactly what those two models just demonstrated. But I want to go one level past that. The mislabeling is the small danger. The large one is this: AI doesn’t lower a task’s DOK. It quietly removes the student from the task while the rigor label stays exactly where it was. That’s why the labels confuse. And it’s why the eighth misconception is the most dangerous one yet.
Back to the Through-Line
Part 1 spent seven misconceptions on a single correction: rigor isn’t difficulty, and DOK doesn’t measure how hard a task feels. DOK names the cognitive demand a standard places on a learner (Francis, 2022). I won’t re-argue it here — that’s what Part 1 is for. Carry one idea forward, and we can build the rest: the demand belongs to the task, not to the person doing it.
DOK Is a Property of the Demand, Not the Doer
This is the hinge, so it’s worth being precise. Depth of Knowledge describes a relationship — between the complexity of what a standard asks and the complexity of the assessment meant to measure it (WebbAlign, 2025; Webb, 1997). It’s a statement about alignment, not about authorship.
Read that twice, because everything follows from it. If DOK lives in the alignment between standard and task, then a DOK 3 task is DOK 3 no matter who — or what — completes it. A student reasons through it: DOK 3. A student pastes it into a chatbot and copies the answer: still DOK 3. The label describes the demand the task makes. It cannot certify that anyone actually met it.
That’s the gap. And it’s where the eighth misconception lives.
Characteristic vs. Criterion: The Move That Matters
Here’s the distinction that does all the work. Erik Francis draws a line in Deconstructing Depth of Knowledge that I keep returning to: the difference between a characteristic of a DOK level and its criterion (Francis, 2022).
Multiple steps often show up in DOK 2 work — but counting steps isn’t what makes a task DOK 2. Time often shows up in DOK 4 work — but duration isn’t what makes a task DOK 4. The characteristic is what we notice. The criterion is what actually defines the demand. A student can spend a week on a task and never once leave DOK 1.
Now ask what happens the moment AI enters the room. When we go looking for proof that real learning happened, we reach for characteristics. Did the student put in the time? Move through the steps? Look like they wrestled with it? Did they struggle?
Every one of those is a characteristic. Not one is a criterion.
And that is the crack the eighth misconception slips through. A student can produce every visible characteristic of rigor while AI quietly meets the criterion in their place. The time on task, the multistep process, the furrowed-brow effort — all intact. Meanwhile the reasoning, the evidence, the synthesis — the actual cognitive work the standard demands — came from the machine.
So let me say plainly what I’ve come to believe: productive struggle was never the test. Struggle is a characteristic. It tells you a student is working. It cannot tell you a student is doing the thinking the standard requires — and under DOK, the thinking is the only thing that counts (Francis, 2022).
Picture how this plays out above the student level, where it’s easiest to see. A school leader sets out to answer a real question: where is our practice actually heading? Answering it means reading the faculty’s end-of-year reflections against the Portrait of a Graduate, last year’s data, and the building’s goals — then pulling all four together into a direction worth committing to. That is genuine DOK 4 work. It connects separate bodies of evidence, across time, into something new.
Instead, the leader hands the pieces to a chatbot and asks for the synthesis. Out comes a confident strategic summary — themes named, priorities ranked, next steps proposed. It reads like a plan.
But integrating four sources into one direction is the criterion: holding the reflections, the portrait, the data, and the goals in mind at once, and deciding what they add up to. The chatbot produced the characteristic — an output that looks like synthesis without anyone having synthesized. The criterion was never met, because no person connected the threads the task required.
And notice: this wasn’t a student cutting a corner. It was a capable professional. Which is the quiet point of this whole post. Offloading doesn’t lower a task’s DOK. It removes the person from it.
The Third Machine in the Room
Here’s what the thread doesn’t show at a glance, but should stop you: there were three AI systems in that conversation, not two.
Two were frontier models — Claude and Grok — out of the box. Both said DOK 3. Both missed.
The third was Francis’s own. He built it on the DOK descriptors he developed, and fed it nothing but the photo of the problem. It came back DOK 2 — and it showed its work.
First it named what the student actually has to do: calculate each walker’s time, work the fractions, convert the units, order the results. Then it ran the task against a test. Is the answer predetermined and confirmable? Yes — one correct ordering. Does the student apply a known procedure to a familiar problem? Yes — divide, then order. Must the student build a non-predetermined interpretation, defend a position with evidence, find a path that isn’t routine? No, no, and no. So it classified: DOK 2. Complex, not complicated — many steps, one clear path, a determined answer. To reach DOK 3, it noted, the problem would have to ask what it never asks: is it possible to walk fewer miles but spend more time walking — and can you defend that with evidence? That question demands reasoning. This one demands procedure.
Now hold the results side by side. Two machines said 3. One said 2. What separated them?
Not horsepower. The frontier models outmuscle a tool one educator built on the side — and it didn’t save them. What separated them was the anchor. Francis (2026b) named it in the thread: his tool runs on his own descriptors — a primary source — while the frontier models synthesize whatever they can reach, and much of it is neither primary nor verified. One machine was tied to a criterion. Two were adrift among characteristics.
So notice what reliability actually tracks here. Not the machine’s power — the anchor. That’s the real divide: anchored versus unanchored. And the anchor isn’t the software. It’s the criterion — the answer to the one question Francis put to that thread, the question you can ask of any task, any tool, any student: what is the knowledge source?
Yes, I’m stretching Francis’s word a notch — on purpose. In the book, the criterion defines a task’s demand. Here it defines a tool’s answer: the source it’s held to. Same discipline, one level up.
That question is Teacher Clarity in work clothes. Clarity was never only about telling students where they’re headed. It’s about naming the criterion so plainly that anyone — a colleague, a student, a machine — can be held to it. The frontier models had no one holding them to a source. Francis’s tool did.
And the student comes back in right here. A learner who can see the criterion — who can ask of their own work, whose thinking is this, mine or the machine’s? — is holding their own anchor. That isn’t compliance. It’s agency. The same clarity that keeps a task from drifting is what keeps a student from drifting out of it.
Looks Like, Sounds Like, Feels Like
Francis gives us a way to see exactly where the person gets removed. He describes DOK through three descriptors (Francis, 2026a): the DOK Skill is what learning looks like — the mental processing the task requires. The DOK Response is what learning sounds like — the extent and quality of the evidence-based answer. The DOK Demand is what learning feels like — the learner’s own perception of the work.
One word is about to do two jobs, and I don’t want it to slip past you. Everywhere above, demand has meant one thing: the cognitive requirement the task makes — the thing that belongs to the task, not the doer. Francis’s DOK Demand means the opposite pole — the felt register, what the work feels like to the learner. Same word, two ends. And the gap between those ends is the entire mechanism of the offload: AI meets the demand the task makes while the demand the student feels sits there untouched.
Line those up against AI offloading, and the mechanism jumps out. The offload happens in the “looks like” and the “sounds like.” AI performs the Skill and generates the Response. What it leaves untouched is the “feels like” — the student’s felt sense that they worked hard. So the student feels the effort while the machine meets the task’s demands. The one register that stays intact is the one that was never evidence in the first place.
Four Use Cases, One Question
This is why you can’t judge an AI task by how rigorous it looks. You have to ask what it’s doing to the Skill and the Response. Take four common classroom uses from the Student Achievement Partners literacy guidance (AI for Education & Student Achievement Partners, 2024), and run each through a single question: is AI lowering difficulty, which is legitimate, or absorbing the Skill and Response, which is the loss?
Leveling a text down so a striving reader can reach grade-level ideas lowers difficulty while leaving the thinking to the student — legitimate. Generating rubric-based feedback that the student must interpret and act on is difficult; the responsibility for revision remains theirs. But having AI “chat” as a character so the student skips the inference work, or assembling a text set so the student never does the sourcing and synthesis — those absorb the Skill and the Response. The task still wears its DOK label. The student just isn’t inside it anymore. Francis’s descriptors are already written in reading-comprehension language, so they map onto each case with no translation (Francis, 2026a).
What Offloading Does to a Brain
If this were only a labeling problem, it would be an annoyance. The research says it’s more. Gerlich (2025) found that frequent use of AI tools is associated with cognitive offloading and reduced engagement in critical thinking — the more we hand off, the less we do. And Kosmyna et al. (2025) ran an EEG study and found what they called cognitive debt: students who used an AI assistant from the start of an essay task showed weaker neural connectivity than those who worked unaided first and brought the tool in later.
That last detail is the whole design principle in miniature. It isn’t AI or no AI. It’s sequence. Let students meet the demand unaided first, then bring AI in after the struggle — not in place of it (Kosmyna et al., 2025). The struggle has to happen in a brain before a tool can help that brain do more.
The Teacher’s Move
All of which reduces to one change in the question we ask. For any task a student completes with AI, the diagnostic is not “Is the student struggling?” That’s the characteristic, and we’ve seen it can be faked or felt without any underlying learning. The diagnostic is: “Is the student performing the DOK Skill and producing the DOK Response?” That’s the criterion, and it’s the only thing that tells you the demand was met.
Naming that criterion out loud is a move you already know. Making the expectation visible — telling students exactly what thinking the task requires and how they’ll show it — is Teacher Clarity at its core (Hattie, 2023; Francis, 2026a). Clarity was always the antidote to confusion about rigor. It turns out to be the antidote to offloading, too, because a student who can see the criterion knows when they’ve handed it away.
The Correction
So here is the eighth misconception, named and corrected. AI doesn’t lower a task’s rigor. The DOK never moves. What moves is the student — out of the demand and into the role of editor, approver, passenger — while the rigor sits right where it always was: untouched, and unmet.
Which brings us back to that LinkedIn thread. Three machines read the same walking problem. The two most powerful missed. The one anchored to a verified human criterion did not — but even it only confirmed what Francis, holding that criterion himself, had already named. That’s the shape of the whole thing. On DOK, reliability was never in the machine’s horsepower. It’s in whether a human criterion is holding — and in refusing to mistake the characteristic for it. The task’s rigor will take care of itself. Our job is to keep the student inside it.
References
AI for Education, & Student Achievement Partners. (2024). A guide to integrating generative AI for deeper literacy learning. https://www.aiforeducation.io/ai-resources/guide-to-integrating-generative-ai-for-deeper-literacy-learning
Francis, E. M. (2022). Deconstructing Depth of Knowledge: A method and model for deeper teaching and learning. Solution Tree Press.
Francis, E. M. (2026a, January 14). Why DOK labels confuse and how to fix them. Solution Tree Blog. https://www.solutiontree.com/blog/why-dok-labels-confuse-and-how-to-fix-them/
Francis, E. M. (2026b, August 8). DOK 2 at most because it’s a word problem. You’re using information and basic reasoning to establish and explain [Comment on the post “What DOK level is this?”]. LinkedIn. https://www.linkedin.com/feed/update/urn:li:activity:7491717886055444481/
Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), 6. https://doi.org/10.3390/soc15010006
Hattie, J. (2023). Visible learning: The sequel: A synthesis of over 2,100 meta-analyses relating to achievement. Routledge.
Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. arXiv. https://doi.org/10.48550/arXiv.2506.08872
Webb, N. L. (1997). Criteria for alignment of expectations and assessments in mathematics and science education (Research Monograph No. 6). Council of Chief State School Officers and National Institute for Science Education.
WebbAlign. (2025). Depth of Knowledge (DOK) primer. Wisconsin Center for Education Products and Services. https://www.webbalign.org/dok-primer



