Measurable and Meaningful Are Not Always the Same Thing: An Early Lesson From A Network of Schools Redesigning Learning With AI
Beth Rabbitt and Virgel Hammonds (FullScale) and Caroline VanderArk (Playlab)
Author Beth Rabbitt and Virgel Hammonds (FullScale) and Caroline VanderArk (Playlab)
Artificial Intelligence 7 min read

When Secretary of Defense Robert McNamara ran the war in Vietnam, he ran it by the numbers. He had come to the Pentagon from Ford, one of the “Whiz Kids” who had rebuilt an ailing automaker on the strength of statistical control. He brought the same faith in metrics to the Department of Defense. The trouble was that some of the things that mattered most in a war, such as the will of an adversary, the loyalty of a village, whether you were actually winning, did not lend themselves to a spreadsheet. So the war was ultimately measured by the things that did. Body counts. Tonnage dropped. The data was precise, abundant, and, as history would judge, beside the point.

This pattern of management has come to be known as the McNamara fallacy: measure whatever can be easily measured, disregard what can’t, or assign it an arbitrary value. Then presume that what can’t be measured easily isn’t very important, or conclude that what can’t be measured doesn’t really exist at all. 

For a generation, and with the best intentions, American education has run a version of McNamara’s war room. We have gotten very good at counting the countable: test scores, attendance rates, course completion, graduation percentages. These numbers are legible, comparable, and (relatively) cheap to collect. They have done real work in holding systems accountable for children they were quietly failing. 

But the things most of us most want for young people, particularly that they become curious, capable, connected, self-directed people who can build a good life, sit almost entirely outside the reach of the instruments we use to judge whether schools are working. Because those things are hard to measure, we have learned to treat them as soft, and outside the purview of what we can actually achieve. 

This is the false choice the field keeps handing itself: rigorous numbers that miss the point or meaningful aspirations that seem uncountable. At FullScale, we think that choice is wrong, and in July we had the opportunity to work with a room full of educators who refused to make it. We are working alongside Playlab to support its AI Lab Schools, a network of schools that are asking  and demonstrating what school could become when intelligent tools are treated as a new design material for K-12 learning. 

As part of kick-off, we worked with schools to name the questions worth asking over the next two years, define the outcomes they were actually chasing, and figure out what evidence would tell them whether they were getting there. To do that, we asked participants to brainstorm every measure they could imagine mattering (wish-lists allowed!) and then to sort each one onto a simple grid. One axis was importance: how much does this evidence actually matter to understanding whether we’re succeeding?

The other was feasibility: how realistically can we collect it right now? That gave us four quadrants:

  • “High-Leverage” – evidence that is high-importance and feasible 
  • “Collect it anyways” –  evidence that is easy to collect (perhaps because we were already collecting it) but of lower importance
  • “Big bets” –  evidence that is high-importance but hard to measure
  • “Maybe laters” – evidence that is low-importance and hard to measure

The big-bets quadrant filled first and filled fastest, with the most human outcomes: agency, belonging, ownership of learning, executive function, trust, ambition, and, more than once, the word flourishing. Across the roughly two-and-a-half dozen design questions the schools generated, participants kept circling back not to curriculum or software or standards, but to the learner’s relationship to learning itself. 

Perhaps most counter-intuitively, our maybe later quadrant stayed empty. Given a category explicitly reserved for the things that don’t much matter and are hard to measure besides, a room of practitioners declined to put a single thing in it. Rather than concluding that what’s hard to measure doesn’t matter, these educators insisted that what matters must somehow be measured, with deep urgency to work on the harder problem of “how” now.

When we stood back and tried to make sense of what we saw, a theory of change emerged that connects the metrics we already collect to the outcomes we’re really after, with the honest acknowledgment that most of what lives between them is proxy and inference.

At the far end sit the major outcomes: agency and capability, belonging, flourishing, ownership of learning, and long-term success in a life beyond school. This is the destination, the horizon these schools are actually moving toward, and almost none of it can be read directly off a dashboard.

In the middle sit the observable signals, the proxies we can see now if we look well: portfolios and exhibitions of real work, growth against competencies, student voice and story, wellbeing, genuine engagement. These are not the outcomes themselves. They are the visible smoke of a fire we care about, and a network that gets good at reading them can learn something long before the longitudinal data comes in.

And at the near end sit the early and operational indicators: attendance, achievement scores, completion rates, participation, whether the model is even being implemented as designed. The numbers we already have.

The reframe this theory of change offers is simple and important: familiar numbers are not the enemy of the outcomes we want. They are leading indicators on the path toward them, but only if we are honest about the theory that connects them, and only if we stop mistaking the near end of the chain for the far end outcomes. Attendance matters not because presence is the goal, but because a child who isn’t there cannot be flourishing. A proficiency score is a floor, not a ceiling. The point was never to throw out the countable. It was to stop letting the countable quietly redefine the goal.

Why does this matter beyond one network of ambitious schools? Because the fallacy has an equity cost that rarely makes it onto the scorecard. When easy metrics stand in for the real goal, the students already least well-served are the ones whose belonging, agency, and flourishing go uncounted. In systems built the way ours are, uncounted too often becomes uninvested-in. A measurement regime that can only see test scores will keep optimizing for test scores, for every child, and will keep missing the ones furthest from opportunity in precisely the dimensions that would change their lives. 

Building the connective chain, so that the human outcomes finally have equal weight to the easily countable concrete measures, is not a soft exercise. It is an act of accountability of the most serious kind. This is also an act of accountability that needs urgent solving. The outcomes we care about most remain the hardest to see, and proxies can mislead (a measure, once it becomes a target, has a way of going crooked). 

We’re offering this as an early lesson, not a conclusion. The AI Lab Schools will follow the first  two years of a much longer arc, the learning agenda will keep changing as the schools do, and the work of building evidence systems worthy of the learning we actually want for young people is, honestly, just getting started. We’d love company in it. If your school or system is wrestling with the same question we are (how do we know we’re succeeding when the things we care about most are the hardest to measure?), tell us what you’re finding. 

Learn More:

AI Lab Schools Cohort

Stay Connected

Related Posts

Infrastructure Is Not Instruction: Befor...

Students at Sutton Middle School listen to songs representing different eras in history during an International Baccalaureate immersion day.

What Can We Learn from Student Experienc...

Starting and Learning at the Margins

Students assemble a model wind turbine they constructed.

About the Author

Beth Rabbitt and Virgel Hammonds (FullScale) and Caroline VanderArk (Playlab) Dr. Beth Rabbitt is Co-CEO of FullScale, a national nonprofit advancing innovation, equity, and systems change in public education. She […]

Read more from Beth Rabbitt and Virgel Hammonds (FullScale) and Caroline VanderArk (Playlab)
Scroll to Top