Skip to content

When the Measure Becomes the Goal

Goodhart's law is highly relevant for eduactors. It is an old adage that observes how when a metric is used for accountability, behavior adapts in response to the metric so that it no longer measures what it intended. Once one sets out to control the behavior of others through a particular measure, the people within that system have an incentive to improve that measure regardless if that which it represents has improved or not.

A rising score can represent learning. But it can also represent familairity with the measurement tool.

This is my current lens going into Koretz's The Testing Charade. For context, I faciltate the 7th Grade ELA professional learning community (PLC) at the public middle school I work at. We meet for one hour during our planning and are guided by the school's Literacy Coach to review grade-level data from schoolwide assessments. The largest assessments we have access to are the district summatives, where all students in the district take the same assessment. So us teachers can compare our students' scores to the schools around us and to the district as a whole, whereas the schoolwide assessments only let us compare among ourselves.

During these meetings, we identify standards where our students underperformed and plan how we might respond. There is nothing wrong with this. On the contrary, I think students should be continuously looking at their students' data, and I would be suspicious of educators who claim that assessment results had nothing useful to tell them.

But I've become more and more interested (perhaps "bothered" is too antagonistic) in what happens when these assessments turn into the primary way of determining whether instruction was successful. The assessment pressures instruction.

If students perform lowly on a certain standard, the obvious response is to reteach it and reassess. If another teacher's students outperform mine, I want to know what they did differently. If our school performs below average across the board, there is an obvious and justifable desire to identify the gap and close it. Each of these decisions make sense. They approach the assessment as a source of information.

However, all of these decisions taken together help create a system where the assessment dictates what receives our attention. What do we study in class? The standards that appear on the test. Students are familiarized with the language and structure likely to appear on the test. Teachers spend planning time looking for patterns in distractor items and response-frequency reports. After all, the better we understand the assessment, the better positioned we are to improve performance on it.

Enter Goodhart's law.

Fast-forward. After all that analysis, deliberating, planning, teaching, learning... scores improved!

What exactly have we demonstrated?

It is possible that our students have become stronger readers and writers. But it is also possible that they also became more familiar with the processes and typologies chosen by the district. Perhaps we simply became better at preparing students to demonstrate what they know through that particular assessment.

The difficulty is that the score cannot tell us exactly what happened. And once improving it becomes our goal, a higher score is no longer as valuable as it was supposed to be. And the further removed the people looking at the data are from the classroom it was produced in, the less informed they are to make conclusions about the infinite variables to account for.