← FIELD NOTES

EVIDENCE LADDERS · CH. 5

Evidence Ladders

JEFF NICHOLSON · 5 MIN READ · FROM INTENT

Evidence Ladders

Most product conflict isn't about taste. It's about truth, wearing taste as a costume. Two people argue about where to put a button or whether to build a feature, and it looks like a disagreement about preference, but underneath it's a disagreement about what's actually true, about the user, the problem, the likely impact, and neither of them has a shared way to settle it. So the argument gets decided by something else: seniority, confidence, or whoever made the cleaner slide.

The reason teams get stuck isn't that they lack data. It's that they treat all data as equal the moment it enters a slide. A hallway comment, a survey statistic, a behavioral log, and a measured outcome are not the same weight. But once they're each rendered as a bullet point, they look identical, and the team ends up debating with a pile of "evidence" where a vague impression and a hard result carry the same authority.

That's the core problem: "data" is not a category. It's a pile. And a pile doesn't tell you what to do. You can reach into it and pull out support for almost anything, which is exactly why product debates go in circles and get won by whoever is most persuasive rather than whoever is most right.

An evidence ladder is the fix, and it's almost embarrassingly simple. It's just an agreement that some forms of evidence carry more weight than others, especially when you're about to spend real money and time building something. At the bottom is intuition and anecdote, someone has a feeling, someone heard a thing. Above that, statements from multiple users saying the same thing. Above that, observed behavior, what people actually do, not what they say they'll do. At the top, outcomes, evidence that a change actually moved the job. Nothing academic about it. It's a way to stay honest about how much you really know.

The ladder does two things. First, it ends the false equivalence. When someone says "the data shows," you can ask which rung, and a hallway comment stops being able to outvote a measured result just because it was stated with conviction. Second, and more useful, it lets you size your bets honestly. If all you have is intuition, you don't launch a quarter-long project, you run a small test. If you have a pattern in support tickets, you build a small fix and observe. If you have real behavioral data and a clear moment, you can place a bigger bet with your eyes open. The ladder isn't just a way to win arguments. It's a way to match the size of the work to the strength of the evidence.

Here's what it looks like when a team skips this. A customer asks for a custom rule. That's a rung-two signal, a statement, one customer, one request. The team takes it at face value and builds. Eighteen weeks later the custom rule ships. It's beautiful, it's specific, the customer uses it twice and then goes back to their workaround. What happened? The request was real, but the solution wasn't. If anyone had moved one rung up the ladder before building, from "what did the customer ask for" to "what is the customer actually doing", they'd have found a workaround that took forty-five minutes every cycle, and a real problem that was a general-ledger mapping issue, not a missing rule. A different project. Half the time. Solving the thing that was actually broken.

That's the cost of treating all evidence as equal: you build confidently in the wrong direction, because a confident request and a true diagnosis felt the same on the slide.

The ladder also protects you from the opposite failure, dismissing a measured outcome because it's inconvenient. When the data finally comes in and it says the thing you shipped didn't help, the pile lets you wave it away with another bullet point. The ladder doesn't. A real outcome sits at the top for a reason. If you're going to override it, you have to do so explicitly, not by drowning it in lower-quality evidence that happens to agree with what you already wanted to do.

None of this requires a research function or a statistics degree. It requires one shared agreement: that not all evidence is equal, and that the size of a bet should track the strength of the proof behind it. That single agreement changes how a team argues. The question stops being "who's more convincing" and becomes "what do we actually know, and how much does that justify building." That's a slower question to ask and a much faster way to build, because you stop spending quarters on hunches and stop ignoring the results that could have saved you.

This is one of the core tools in my book, Intent: How to Build Products That Last in the AI Era.