top of page

Our Test Data Tells Us Who Got What Wrong, Not Why. How Are We Supposed to Reteach From That?

A teacher sent this in after a data meeting, and I want to give the direct answer first, because I think it is the most useful thing I can say.


You cannot. Item-level data from a benchmark test genuinely cannot support a reteaching decision, because a right answer arrived at for a wrong reason and a wrong answer arrived at through sound reasoning look identical in that data.


That is not a complaint about your particular benchmark. It is a description of what that kind of data is, and no amount of disaggregating, color coding, or sorting by standard will fix it, because the information you need was never collected.


One Item, Seven Different Students


Freezing and Melting Water is a Phenomenon-Driven Task for fifth grade. Students watch a timelapse of a glass of water sitting outside on a very cold day, slowly turning to ice, and then pick which of four graphs best shows what happened to the weight of the water while it froze. On a benchmark test that item produces one bit of information per student: correct or not.


Here is what was actually in the heads of fifteen fifth graders who answered it. Among the students who chose a graph showing no change, the reasoning split cleanly in two. One group wrote something like “the amount of water does not change when it freezes, so its weight does not change.” That is the target reasoning. Another group wrote something closer to “things look or act different when they freeze but stay the same weight.” Those students have a rule about freezing. They do not have an idea about matter being conserved, and the next phenomenon that is not about freezing will catch them out.


Both groups are scored correct. On your data wall they are the same student.


The students scored incorrect were not making one error. They were making four:


●       Solids weigh more than liquids.

●       Bigger things are heavier than smaller things.

●       Water is lost when it freezes and is replaced by ice.

●       Ice stacks up and gets heavier, so the weight keeps going up.


If the plan is “go back over conservation of matter with the students who missed item 14,” you are planning one lesson for four different problems. The student who thinks bigger means heavier needs a very different conversation from the student who thinks the water is being replaced. And the student who got it right with a rule about freezing is not on your list at all.


Why Item Analysis Cannot Get You There


Percent correct is a measurement of a group, not a description of thinking. Distractor analysis gets closer, because it tells you which wrong option was popular, but it still only tells you what students chose. It is also silent about the students who answered correctly, which is where a surprising amount of shaky reasoning hides.


This is not a sign that anyone built a bad test. It is what happens when one assessment is asked to serve several purposes. A benchmark exists to support comparison across classrooms, satisfy accountability requirements, and evaluate programs, and it may do all of that well. The moment it is also asked to guide reteaching, it is being asked for information it was never designed to produce.


What You Need Is a Second Question


The change is smaller than you would expect. Every part of a Phenomenon-Driven Task has two questions instead of one. The first asks for an answer: choose a statement or image, sort statements or images into categories, or sequence them into a process. The second asks why. It is always short answer, because the point is to let students explain their thinking however they can. Some write, some draw, some record audio, some respond in their home language.


The first question tells you what. The second tells you why. Reteaching decisions require the second.


How to Analyze It Without Drowning


When the purpose is to inform instruction, I use what I call a Thinking Analysis. It is a sorting task, and it takes about twenty minutes the first time.


●       Gather all the handouts for Part 1 into one pile.

●       Sort them into groups based on the answer selected in question one.

●       Sort each of those groups into subgroups based on similarities in the reason given in question two.

●       Name each subgroup for how those students were thinking. “Solids weigh more than liquids.” “Water is lost when it freezes.”

●       Repeat for each part of the task.


You end up with four to eight named ways of thinking per part, and the number of handouts in each pile tells you how common each one is. Notice what you are not doing. You are not marking anything right or wrong, and you are not tracking individual students. You are characterizing the ideas that are in the room.


What You Can Do With It


A Thinking Analysis beats an item analysis not because it is more precise but because it produces something you can build a lesson out of. Once the ways of thinking are named, you can use them:


●       as alternative hypotheses for students to test during an investigation

●       as claims for students to evaluate against evidence

●       as a connection to historical models that scientists themselves revised over time

●       as the starting explanatory model the class revises across a unit

●       as a basis for grouping students, either with peers who share an idea or deliberately across different ones


That is a different posture than reteaching. You are not refilling students who came up short. You are starting from the ideas they already have and moving them.


How ADI Makes This Easier


Writing a two-question prompt that reliably surfaces reasoning is harder than it looks. The phenomenon has to be engaging, explainable both with and without formal science knowledge, and tightly aligned to the ideas and practices you are targeting, and the answer options have to be built around how students genuinely think rather than around plausible-sounding distractors.


Phenomenon-Driven Tasks are designed and pilot tested to do that. Each is anchored in a real phenomenon, includes two or more connected parts, pairs every answer question with a reason question, and comes with guidance for running a Thinking Analysis, a Feedback Analysis, or a Performance Analysis depending on which question you are asking.


I am not suggesting your school drop its benchmark. I am suggesting you stop asking it a question it cannot answer, and add something small that can.


Want to Try One?


Browse the Phenomenon-Driven Tasks books at shop.argumentdriveninquiry.com/collections/phenomenon-driven-tasks, or learn more about ADI curriculum materials, the ADI Learning Hub, and our professional learning options at argumentdriveninquiry.com.

 
 
 

Recent Posts

See All

Comments


bottom of page