English HomeEssaysThe Discipline of Waiting

mrpink essays

timeless PinkLetters, translated from our Spanish Substack for English-speaking readers.

Read in original language

The Discipline of Waiting

Almost half the days my scale said I was gaining. The trend only went down. On noise, sample size, and why the bar for evidence should rise with irreversibility.

The CEO of a B2B startup had a problem with his sales team. The numbers didn't add up, the pipeline looked good in meetings and bad in the bank account, and nobody could explain why.

He did what any sensible person would do: he realized he was deciding blind. His analytics were poor, so he invested in fixing them. He migrated to a decent CRM, cleaned up the funnel stages, forced the team to log their activity. Three months later he had something he'd never had before: win rate by sales rep.

And what he saw unsettled him, because it was the opposite of what he expected.

His hypothesis, the one anyone would have, was that the veterans would come out ahead. Years of craft, relationships built, an instinct for discarding quickly whatever was going nowhere. And that the new hires would improve their win rate as they accumulated experience, the way everyone improves. The numbers said otherwise: the reps he'd been carrying since the early years were closing proportionally less than the three he'd brought on recently.

The difference wasn't enormous, but it was there, and it was consistent with a suspicion he'd already been chewing on. What he didn't know — what he had no way of knowing from that dashboard — was that the new hires' number was already inflated, and that the room for improvement he was projecting onto them was, in fact, air that was going to leak out on its own.

He acted. He let the veterans go, hired more profiles like the new ones, and kept measuring. It felt good: he was finally making decisions with data instead of impressions.

Six months later the business was considerably worse than it had ever been.

Three ways that number lies

The first thing to say is that the CEO didn't do anything unhinged. He was wrong, and the mistake was expensive, but it's the mistake of someone who was trying to do things properly. He measured, found a difference, acted on it. That is literally what we ask of a good executive. The problem isn't in the decision but in a step before it, so invisible we almost never discuss it: how much evidence was needed for that difference to mean anything.

A win rate is not a fact. It's an estimate, and like every estimate it carries a spread. With twenty closed opportunities, a rep who truly closes at 30% can show 20% or 40% without anything at all having changed. The veterans' distributions and the new hires' weren't two separate points: they were two wide bells stacked on top of each other, sharing an enormous amount of surface. The CEO read the width of the bell as signal.

This is what statistics calls a type I error: concluding there's a difference when there isn't. In a staffing decision it has a concrete face, and it's exactly the CEO's in this story — letting go of someone who was doing fine over a difference that was, in reality, the width of the bell. Its twin, the type II error, is the rep who genuinely isn't performing but whose numbers landed in the middle of the pack, so nobody looked twice.

The first leaves an empty chair and an uncomfortable story someone will tell in the hallway. The second stays on payroll for years and generates no conversation at all. That's why almost every company I know is miscalibrated in the same direction: it isn't that they chose the wrong threshold, it's that one error is visible and the other isn't, and we end up protecting ourselves from the one that leaves a trace.

And there's something more uncomfortable in this case. With twenty closed opportunities per rep, that CRM had no capacity to detect a real difference even if one had existed. The number required is worse than intuition suggests: to distinguish a 20% win rate from a 30% one — a gap nobody in a board meeting would hesitate to call significant — you'd need somewhere around three hundred closed opportunities per rep, at the confidence and power levels normally used. Even to separate 20% from 40%, an enormous difference, you need more than eighty.

He had twenty. The same instrument that handed him a false positive was, at the same time, incapable of giving him a true one. It isn't that the answer was wrong: with that data, the question simply had no answer. Buying the instrument and never asking what it can and cannot resolve is the part of the problem that almost never gets discussed.

The second way the number lies has to do with time. If the sales cycle is long, the win rate measures what has already resolved, and what has already resolved is not a representative sample of what's in the pipeline. Deals close fast when they go well and rot slowly when they go badly. A rep with six months of history has a denominator built almost entirely from the cases that resolved quickly. One with six years carries all the corpses.

And the third, the one that interests me most, is no longer statistical.

When the measured controls the measurement

To mark an opportunity as lost you have to do something uncomfortable: accept that it's lost. Nobody forces you. The system will let you leave it sitting there, in "negotiation," indefinitely.

An experienced rep knows that a client who hasn't answered in five weeks already said no. He closes it as a loss, lowers his own win rate, and moves on. A new one leaves it open. Not necessarily out of bad faith: because it stings, because he still believes, because he hasn't learned to tell silence from delay. Sometimes, yes, because he figured out quickly what the company looks at.

The effect is the same in all three cases. The numerator and the denominator of the metric are controlled by the person being evaluated with that metric. And since the one with the least experience is also the one slowest to be honest about losses, the bias points systematically in favor of the new hires.

This is different from noise. Noise is symmetric and averages out. Bias never averages out: no matter how much data you accumulate, it stays there, with the same force and in the same direction.

And what happened next was almost worse.

As the new hires accumulated volume — and as their own parked opportunities aged past the point where they could keep being propped up — their win rates dropped. Not because they'd relaxed, nor because the market had shifted. They dropped because an extreme number calculated over few cases tends to move toward the average when more cases arrive. It is the most predictable thing in statistics and the hardest to accept when it happens to you.

It's the trap Kahneman describes in Thinking, Fast and Slow with the flight instructors of the Israeli air force. They berated a pilot after a bad landing and the next one came out better. They praised him after a good one and the next came out worse. They concluded, with impeccable empirical evidence, that punishment works and praise ruins. What they were watching was a noisy series returning to its mean.

The CEO was left with two possible readings: accept that the difference that justified the firings never existed, or conclude that the new hires had gotten lazy. It's uncomfortable, but the second is the one people pick almost every time, and it usually comes with another round of cuts.

And there's an irony that strikes me as the cruelest part of the whole episode. By letting the veterans go, he destroyed the only baseline against which he could have discovered his own mistake. The action ate the possibility of measuring it. There is no control group when the control group no longer works here.

This isn't exclusive to firing people. Changing pricing every six weeks, moving the ideal customer profile, rotating agencies, killing a campaign before it matures: every intervention generates new data and simultaneously ruins the comparability of the old. You end up with an enormously long series in which no two periods were ever alike.

Same problem, less drama

Six weeks ago I started a body recomposition process. Nothing dramatic: lose fat slowly, hold lean mass, no heroic deadlines.

The standard recommendation, the one you'll find everywhere, is not to weigh yourself every day. That you'll obsess, that the number will run your mood, that once a week is enough. I did exactly the opposite. I weigh myself every day, always at the same hour, fasted, on a scale that syncs to my phone on its own so the record doesn't depend on my willpower.

And I don't weigh myself daily despite knowing that daily weight is noisy. I weigh myself daily precisely because daily weight is noisy.

The sodium from last night's dinner drags water along. Every gram of glycogen carries about three of water with it. There's gut contents, cortisol, what time you went to bed. None of that has anything to do with body fat, and all of it is inside the number on the floor. The only way to remove it is to average it out, and to average it out you need data.

High frequency isn't what feeds the obsession. It's what allows me to be indifferent to any individual measurement. I measure every day so I don't have to look at any single one.

46, 33, 7

Six weeks of daily weigh-ins. The real trend, fit by regression, is a decline of 306 grams per week, with an R² of 0.74 and a p-value small enough that the hypothesis of no decline at all is hard to sustain. One point nine kilos in six weeks. There's no ambiguity: the process worked, steadily, the whole time.

Now, the same measurements read three different ways.

Against the previous day: 46% of days the scale said I had gained. Eighteen out of thirty-nine. Almost half. There was a three-day streak of "gaining weight" inside a stretch where I only lost.

Against the same day the previous week — the discipline almost everyone recommends, and which sounds far more sensible: 33%. A third of the weeks would have looked like stagnation or backsliding.

Against a seven-day moving average, compared to the moving average seven days earlier: 7%. Two cases, both from this past week, both with differences so small they're rounding.

Forty-six, thirty-three, seven. The same data. The same physical reality. Three different reading rules.

I chose a seven-day window for a reason that isn't statistical but biographical: human life has a weekly periodicity. There's a weekend every seven days, you eat differently, the gym routine changes, you sleep differently. A window of seven covers a full cycle of that, whatever its exact shape. An honest caveat is due: when I explicitly looked for a day-of-week effect in my data, I found nothing conclusive. Seven isn't a number I've proven; it's a number that spans the structure I know exists even though I haven't been able to isolate it yet.

And for a commercial process, seven is almost certainly not the number. The window has to be the size of the cycle, and the cycle of a B2B funnel does not last a week.

What the daily reading would have done to me

This is where the two stories touch.

If I had been reading the number against the previous day's, I would have had apparent evidence that something wasn't working on almost half of all days. And I wouldn't have held out. I'd have cut calories, dropped a meal, tightened up. Then I'd have seen three consecutive days of losses and eased off, with the same data-driven justification.

The result wouldn't just have been walking around in a bad mood. I'd have injected more variability into the system, on top of what it already had. Meals shifting week to week, glycogen rising and falling, each correction generating its own oscillation. I'd have converted measurement noise into real process noise, and by that point there would be no way to know what was happening.

That is exactly what happened to the CEO. He read noise, acted, and his action generated new noise — a destabilized team, orphaned accounts, years-long relationships that walked out with the people — which he then read as information about the quality of the people he'd hired.

Measure often, decide slowly

Here's the principle, and it's stranger than it sounds because it separates two things we instinctively fuse.

The frequency at which you measure and the frequency at which you decide don't have to be the same. You measure fast, so you have something to average. You decide slowly, because the signal needs time to emerge from the noise. The CEO's entire problem fits in having equated the two: the CRM gave him, for the first time in his life, the ability to look. And he assumed that looking more often entitled him to decide more often.

It's an understandable instinct. Instrumentation feels like power, and power asks to be exercised.

And there's an entire culture pushing in that direction. Move fast and break things works reasonably well as an antidote to paralysis, which is a real and well-documented problem in startups. The trouble appears when we apply it where it doesn't belong. Moving fast on little information is fine, and sometimes it's the only option available: you take the risk, you use intuition, you explicitly accept that you're betting. That's honest and it's part of the craft.

When that belief loosensAll of this depends on having chosen well before getting on. But from the inside, while you're running, determination and stubbornness feel exactly the same. Both are the same absence of doubt. The difference isn't in the experience, it's in a decision already behind you, one you committed to not revisiting. — not as intellectual pose but as genuine recognition — decisions keep happening, leadership keeps being exercised, things keep moving forward. But there's less internal friction around all of it.

What isn't fine is believing you're deciding with data when you're actually playing a lottery with a dashboard in front of you. That's the trap: not the bet, but the bet dressed up as precision. The CEO in this story would have been on firmer ground saying "my gut says these three are better and I'm going to bet on that." At least he'd have known what he was doing, and he probably would have asked for more evidence before turning a hunch into layoffs. The CRM didn't give him information: it gave him permission.

But there's an asymmetry the CEO inverted completely, and which protected me without my doing anything especially clever.

Where that line sits, I don't know. And I suspect you find out which side you landed on quite a while later.

If I misread my scale, the cost is a week eating slightly less than I needed. It reverses on its own. It's a cheap, reversible error, which is why I can afford a low bar for evidence.

If the CEO misreads his CRM, the cost is people who don't come back, clients who leave with them, and a destroyed basis for comparison. It's an expensive, irreversible error, and it demands an extremely high bar.

He was more casual with noisy evidence when letting five people go than I was with comparable evidence about skipping dessert. Put that way it sounds absurd. And yet it's the default posture in most companies I know: the bigger the decision, the more pressure there is to make it quickly, and the less time is left to ask whether the number justifying it says what it appears to say.

Setting a bar for evidence is, precisely, choosing how you distribute the risk between the two kinds of error. A low bar fills you with type I; a high one, with type II. There is no option of not choosing. The only options are choosing deliberately, or letting the bar be set by whatever urgency you happened to feel that day.

The rule that comes out of this is simple to state and hard to hold: the bar for evidence has to scale with the irreversibility of the decision, not with the urgency you feel.

Coda

Something has been circling me since I started looking at my own data with more attention than it deserved.

Almost everything that matters to us — a diet, a sales team, the health of a relationship, the trajectory of a startup — is a noisy process. Not sometimes: always. It's a property of complex systems, not a defect of measurement. And yet our formal training prepares us to read numbers that don't move, while life hands us numbers that move all the time.

That's why I think statistics should be taught from primary school, and not as a branch of mathematics but as part of logical thinking, alongside learning to tell a cause from a correlation or an argument from an opinion. Not to calculate anything. To have the intuition, from childhood, that an isolated data point almost never means what it appears to mean.

And to develop what is, at bottom, the hardest discipline of all, and the least technical: looking at a number that's screaming at you to do something, understanding that it isn't telling you anything yet, and doing nothing.

How many of the decisions we defend as data-driven are, in fact, reactions to variance?

▲  volver arriba