Designing the Experiment

The cheapest test that moves the decision

A founder with a new product almost always reaches for the same plan: build it, launch it, and see what happens. It feels like progress. It feels brave. And it is the single most expensive way to learn anything.

When you launch to “see what happens,” the market answers slowly, charges you for every word, and never tells you why. Sales are weak — was it the price, the message, the timing, the product, or the fact that you launched in July? You cannot tell. You have spent months and real money to buy one muddy data point, and the fog is exactly as thick as before. This is what “fail fast” actually buys: failure, at full retail, with the lesson redacted.

There is a better way to take a step in the fog. You do not have to walk until you hit a wall to learn the wall is there. You can throw a stone ahead of you and listen. That is what an experiment is.

Experiment — a deliberate, low-cost action designed to produce evidence on one uncertain assumption, built so that the result will change what you do next.

An experiment is a question, not a launch

The difference between a launch and an experiment is the difference between charging ahead and asking a question. A launch commits the whole venture and waits for the world’s verdict. An experiment isolates one thing you are unsure about, tests it cheaply, and comes back with a clean answer. The launch confounds everything; the experiment controls for it.

This is the move that separates entrepreneurs who guess from entrepreneurs who learn. It is not caution — a good experiment is often faster and bolder than a launch, because it refuses to wait months for an answer it could have in a week. It is design: the discipline of asking the world a question sharp enough that its answer means something.

It helps to see where this sits. Gathering what already exists carried you up the low rungs of the evidence ladder, but the strongest evidence, what your customer actually does, waits at the top, and no one has collected it for you. Designing an experiment is how you climb the last rungs and generate it yourself.

Figure 15.1: You are here, at the top of the evidence ladder. The lower rungs gathered what already existed; now you generate the strongest evidence there is, what your customer does in a test you design.

The one rule: test only what could change your mind

Here is the rule that governs every experiment worth running, and it is easy to say and hard to obey:

Run a test only if a possible result would change what you do.

If you would launch the product whether the survey came back warm or cold, the survey is theater. If you would keep building no matter what the prototype test showed, the prototype test is a way of feeling busy while avoiding the decision. An experiment earns its cost only when its outcomes point to different actions. A test whose every result leads to the same next step has taught you nothing — it has only dressed up a decision you had already made.

This is why “designing the experiment” begins with the decision, not the data. Before you choose a method, name the call you are trying to make and ask: what would I have to learn to choose differently? That answer is your experiment’s target. Everything else is motion.

For the Curious — Value of information

Decision scientists make this rule precise with the value of information: a test is worth running only up to the amount it would improve the decision it informs. A piece of evidence has value only if it could change your choice and that changed choice is worth more than the test costs. Evidence that cannot change your action, however interesting, has a value of zero. You will never need to compute this to use it. The instinct is the whole point: would any result move me?

Learn From Your AI

I want to use the value of information on a real decision: [state the call I’m facing]. Walk me through what’s at stake, what I might learn from the test I’m considering ([describe it]), and whether any result would actually change what I do — so I can tell whether the test is worth its cost.

Three properties of an experiment worth running

Once a test could change your mind, three properties decide whether it is well designed. Each is a question you can ask before you spend a dollar.

Is it discriminating? A good experiment can come out more than one way, and its outcomes point to different decisions. “Discriminating” means the result separates the worlds — supported sends you forward, not-supported sends you back or sideways. The classic failure here is the test that can’t produce a “no”: you show ten friends your idea and they smile, you call it validation, and you have learned nothing, because they would have smiled at anything. Design the test so that reality has a real chance to tell you “no.”

Is it affordable? Match the cost of the test to the size of the decision. A reversible choice you can undo next week deserves a cheap, fast test; an irreversible bet that commits the company deserves a more careful one. The error in both directions is real — founders over-spend to confirm small choices and under-spend before betting everything. The discipline is to climb only as high as the decision is worth.

Is it high-information? Of all the things you are unsure about, test the one where being wrong is both most likely and most costly. This is where your work in the earlier chapters pays off: you already made your starting beliefs explicit, so you know which assumptions are thin. Aim the experiment at the thin assumption whose failure would sink the venture — not the comfortable one you already half-believe. The most informative test is the one most likely to kill the idea, because if the idea survives it, you have learned something that matters.

Trap to Avoid — The confirmation test

The most common bad experiment is the one secretly designed to pass. Leading survey questions, friendly audiences, vanity metrics, success criteria set after the results come in — all of it manufactures a “yes.” If you find yourself relieved by a result, ask whether the test could ever have produced a “no.” If it couldn’t, you ran a ceremony, not an experiment.

Explore first, confirm later

Not every experiment is a clean test of a stated hypothesis, and trying to force one too early is its own mistake. When uncertainty is high and you can barely name the frame, when you don’t yet know which customer, which problem, or which version matters, you are not ready to confirm anything. You are ready to explore.

An exploratory experiment is generative. You step into the world to widen what you can see: open-ended conversations, observation, small scrappy probes that don’t prove anything on their own but surface the surprise that becomes your real question. A confirmatory experiment is the opposite. Once exploration has given you a grounded hypothesis, you narrow: you design the sharp, discriminating test that moves your belief up or down.

The mistake is to run a confirmatory test on an unexplored question. You will get a crisp, confident answer to the wrong question — efficient failure. If you cannot yet write a grounded hypothesis (you saw how in the chapter on writing a proper hypothesis, and how to build the grounding through exploration), your next experiment is an exploration, not a test. Confirm only what exploration has earned.

Isolate the one thing

The reason a launch teaches so little is that it changes everything at once. New product, new price, new message, new channel, new season — when the numbers come in, every cause is tangled with every other, and you cannot say which one moved the result. The whole power of a designed experiment is that it isolates: it holds the venture still and varies one thing, so that whatever changes in the result can be traced to that one thing.

This is why the cheap, narrow test often teaches more than the expensive, broad one. A landing page that varies only the price tells you something clean about price. A full launch that varies the price and the packaging and the timing tells you almost nothing about any of them. When you design a test, find the single assumption you are testing and hold everything else constant. One clean answer beats five muddy ones.

Halo Alert — A first test

Halo Alert is a ring. On a walk home, a woman presses it and it silently sends a pre-written text to a trusted contact — without reaching for a phone an assailant could grab or smash. By now the team has settled the more basic calls: who this is for (women whose commute means walking alone) and that the fear is real, sharp, and recurring. The fail-fast plan writes itself: design the ring, build it, launch, and see whether women buy.

But “build it and see” puts months of hardware between the team and a single muddy answer — and it tests the wrong thing. With the target and the pain already settled, the live question is no longer whether women hurt; it is whether any pain remains for the ring to solve. Most women who feel that fear already do something about it — they text a friend “home safe,” they share their location with a partner. The assumption that now decides the venture is whether those free workarounds already close the gap, or whether they leave a residual pain1 sharp enough that a woman would adopt a dedicated device. If the workaround is good enough, the ring is dead no matter how well it is built. So the team names the call (should we build this?) and aims its first test straight at that residual pain.

The test is cheap and it can fail. A dozen structured conversations with women who walk home uneasy, plus a few evening observations near transit stops — not pitching, asking: what do you do now, and where does it fall short? The answers discriminate cleanly. If women describe the workaround as enough (they feel safe, it is one tap away), the residual pain is small and the venture is in trouble, and the team has learned it for the price of a few conversations. If instead they describe the gaps (the phone that has to come out, the hand already full, the moment when there is no time to type), that residual pain is the seam the ring lives in. Only then do they take up the next call: what the ring should say when she presses it, and whether she chooses the words or they do.

Notice the sequence. They did not build the ring to find out whether anyone would leave their phone in their pocket for it. They isolated the cheapest fatal assumption, the residual pain left after the free workaround, and answered it before risking a dollar on hardware.

The experiment exists to inform the call

An experiment is never the point. It is a means to a decision. The reason you isolate one assumption, design it to discriminate, and keep it affordable is so that the result will move your belief — and a moved belief will change, or confirm, the call you are about to make. When you read the result, you are not collecting a fact; you are updating, and then deciding whether you have learned enough to act. A test that produces a number nobody acts on is the most expensive kind of nothing.

Working with your AI — where you step in

Your AI is good at the mechanics here: brainstorming designs, estimating how many responses you need, drafting the landing page, flagging the confounds you missed. Let it. But three judgments stay yours, and they are the ones that decide whether the experiment is worth running:

  • Which assumption to test. The AI can list your assumptions; only you can weigh which failure would actually sink the venture.
  • Whether the test discriminates. Ask it directly: what result would tell me “no”? If no result would, the design is theater — redesign it until some outcome could kill the idea.
  • Whether the spend fits the stakes. Approve the cost against the size of the decision, not the appeal of the test.

Hand the AI the mechanics. Keep the design judgment.

Ask Your AI

I’m about to test this assumption: [state your one assumption]. Here’s the decision it informs: [state the call]. Help me design the cheapest experiment whose result would actually change that decision. Then push back: what result would tell me “no,” and is my design secretly built to pass?

Putting It to Work

Try This — Design one experiment

Take the decision in front of you right now.

  1. Write the call you’re trying to make in one sentence.
  2. List the assumptions it rests on. Circle the one that is both most likely to be wrong and most costly if it is.
  3. Ask: what result would change my decision? If nothing would, you don’t need a test — you need to admit you’ve already decided.
  4. Design the cheapest action that could produce that result — and name, in advance, the outcome that would tell you “no.”

If your test has no possible “no,” redesign it until it does. That is the difference between asking the world a question and asking it for applause.

The move: Before you build, find the one assumption whose failure would end the venture, and buy the answer to it for the price of a question — not the price of a launch.

You don’t clear the fog by walking into it until you hit something. You clear it one honest question at a time. Next, we turn from running the test to reading it: how to weigh whether the evidence is credible, and what it should do to what you believe.


  1. Residual pain is what survives existing solutions — the pain that remains after people use or abandon the alternatives already available to them. It is the only pain a new venture can actually sell into, and isolating it is a central move in Expeditionary Innovation.↩︎