← LOGS

Shipping Was Not Traction Evidence

Why shipping and launch activity prove output but do not, without separate evidence, prove demand, traction, or distribution fit.

A brightly lit convenience store at night, surrounded by a vast, empty dark landscape.
Fully built, fully lit, and now open to the world beyond the glass.

IN BRIEF

Shipping proves that an artifact reached the world. It does not, by itself, prove demand, retention, repeatable acquisition, or distribution fit. The same discipline also prevents the opposite mistake: a weak launch can be too small a sample to justify declaring the product dead. The useful move is to replace vague words like “traction” with the exact claim being made, then ask which observation, denominator, time window, and repetition actually support it. Stronger language should arrive only when stronger evidence does.

A shipping receipt proves that a package left the warehouse.

It does not prove somebody wanted the package. It does not prove they opened it, used it, came back for another one, or that the same route can deliver the next hundred packages at a cost that makes sense.

That comparison has a limit — products are not parcels, and markets are not delivery networks — but it fixes one distinction quickly: shipping is evidence of shipping.

That is already useful. A release exists. A feature went live. A store listing appeared. A launch post was published. An artifact crossed from private work into a state somebody else could reach. I do not think any of that should be dismissed as vanity by default.

The mistake is asking those facts to prove what happened after release.

Ten shipped apps do not, by themselves, establish demand. A launch-day spike does not establish repeatable acquisition. Downloads do not establish retained use. Revenue does not automatically establish a durable distribution channel. The observation can be completely real while the conclusion outruns it.

The reverse mistake is quieter and just as useful to catch. A launch with no sales can feel definitive while still being a very small experiment. If almost nobody reached the product, “nobody bought it” describes the observed outcome. It may not tell you very much about demand.

So I would rather retire the question “Do we have traction?” as early as possible and replace it with a harder one:

What exact claim am I making, and what evidence would have to exist for that claim to be true at the confidence level I am using?

Once the claim is explicit, the missing evidence gets much easier to see.


“Traction” compresses several different claims

The word is convenient because it can mean almost anything good that happened after launch.

People saw the product. People visited. People signed up. People completed a useful action. People came back. Someone paid. Revenue grew. A channel produced customers twice. The same acquisition motion kept working. A product started to look durable.

Those are not synonyms.

They are different observations with different evidentiary weight. I do not think we need one universal startup taxonomy to make that useful. We only need to stop letting one vague noun conceal the proposition underneath it.

Instead of:

We have traction.

try the sentence you would actually need to defend:

We have evidence that some users return after first use.

Or:

We have evidence of willingness to pay from more than one isolated transaction.

Or:

Reddit has produced qualified users repeatedly, but we do not yet know whether those users retain.

Or even:

This launch reached more people. We do not yet know whether that changed product use at all.

The language gets less impressive. It also gets much more useful.


An evidence ladder, without pretending it is universal

I find it helpful to separate five broad evidence classes. This is an editorial framework, not an industry standard, and the boundaries can overlap.

Output activity includes releases, deployments, submissions, launch posts, campaigns, articles, and other completed artifacts. It can establish that production happened and that a release path exists. It cannot establish demand, willingness to pay, retention, or repeatable acquisition on its own.

Exposure and response opportunity includes impressions, visits, store views, downloads, sign-ups, and inbound replies. It can establish that people encountered the product or message. It still does not tell you whether the right people arrived, reached value, returned, or paid.

Use and value behaviour includes activation, meaningful task completion, repeat sessions, recurring feature use, and other behaviour that suggests the product is doing more than attracting curiosity. That can strengthen a value claim. It does not automatically establish willingness to pay, durable retention, or a working acquisition channel.

Commercial response includes purchases, subscriptions, renewals, expansion, and qualified pipeline converting into customers. This can establish willingness to pay in the observed cases. Revenue is stronger evidence than an impression for many commercial questions. It is not proof of retention or repeatable distribution.

Retention and repeated acquisition begin to support stronger claims because repetition is finally visible. Cohorts return or renew. Churn becomes measurable. A channel produces qualified users or customers across several attempts. The effort or spend required becomes clear enough to judge whether the motion is operationally useful.

Even there, the categories do not collapse into one grand “traction” score. A product can retain the users it reaches and still have weak distribution. A channel can acquire users who do not retain. One large customer can create real revenue while leaving acquisition repeatability unknown.

The point is not to rank metrics from fake to real.

The point is to keep each observation attached to the claim it can actually support.


Small denominators can create false despair

Most criticism of vanity metrics focuses on overclaiming. Five launches become “a repeatable product engine.” A traffic spike becomes “distribution fit.” A few purchases become “validated demand.”

The opposite error starts from the same lack of discipline:

We launched and got no sales, therefore the product has no demand.

A recent Indie Hackers community post makes the denominator problem concrete. The author describes a small audience and no sales, then uses a rough funnel to argue that a zero at the bottom may be unsurprising when exposure is tiny.

I would not reuse the post's percentages as benchmarks. They are the author's assumptions, not validated conversion norms. The useful part is simpler:

If the opportunity base is small, a zero can contain less information than it feels like it contains.

That matters for independent builders because small samples are normal. Early evidence is incomplete almost by definition.

A small sample can still support a small claim:

  • this message attracted qualified replies;
  • this onboarding step lost most testers;
  • one user paid after asking for a specific capability;
  • this launch generated visits but almost no activation;
  • this channel produced one promising customer and needs repetition before it is called repeatable.

A hypothesis is allowed to be useful before it becomes a conclusion.


A founder story shows when the evidence class changes

An Indie Hackers profile about Leadverse describes a founder who built many apps before one later accumulated paying customers and recurring revenue.

The business figures in that profile are self-reported. I am not treating them as audited benchmarks, and copying the founder's tactics is not the argument here.

The useful distinction is structural.

“Built many apps” is output evidence.

The later account adds different observations: a problem encountered through a channel, people willing to pay, changes to pricing and business model, and recurring revenue. The story becomes more informative because the evidence class changed. The business claim is no longer being inferred from shipping alone.

That does not retroactively make every earlier launch evidence that the portfolio strategy was validated. Those earlier launches remain what they were: shipped experiments.

The stronger conclusion arrives with stronger observations.


Distribution fit is a repeatability claim

Distribution is especially easy to overstate because one conspicuous event is memorable.

A post can go viral once. A directory can produce a burst of sign-ups. A newsletter mention can send a spike. A Reddit thread can produce the first useful customer.

Those are channel events.

A stronger distribution claim needs repetition. The exact threshold will vary by product and consequence, but the question is stable:

When we run this acquisition motion again, does it keep bringing the right people, and do enough of them progress toward the outcome we care about at an acceptable level of effort or cost?

That does not require a giant company or a pristine experiment. It does require more than a memorable spike.

The wording should move with the evidence:

  • “Reddit produced the first qualified user.”
  • “Several separate Reddit conversations produced users with the same problem.”
  • “Repeated campaigns using the same problem framing produced activated users.”
  • “The channel has now produced paying customers repeatedly.”
  • “Acquisition has repeated often enough, with visible effort or cost, that we are treating it as a working channel.”

Each sentence asks more of the evidence than the one before it.

That is the discipline. Not pessimism. Not refusing to celebrate. Just making the language wait for the observation.


Build evidence and business evidence need different contracts

Build evidence is usually easy to collect. Commits. Pull requests. Releases. Deployment records. Store listings. Screenshots. Published URLs.

Business evidence is messier. Attribution. Qualified traffic. Activation. Retention. Conversion. Revenue quality. Cohorts. Acquisition effort. Repeatability across time.

The convenience difference creates a temptation: use the abundant build record as a proxy for the harder business record.

Dream Atlas's Product Build History process already treats historical claims this way — different records prove different things, and collapsing them into one story manufactures certainty. Business claims need the same restraint, but they need a different evidence contract.

A build claim can ask:

What proves that this artifact shipped, and what exactly shipped?

A business claim can ask:

What market behaviour are we claiming, what observation supports it, over what period or cohort, and what remains unknown?

Those contracts can sit beside each other without one diminishing the other.

This matters even more when AI-assisted building makes output faster. If another app, landing page, article, or launch becomes cheaper to produce, the number of shipped artifacts becomes even less informative about whether the distribution problem has been solved.

High output can be a strength.

It can also widen the distance between “we made a lot” and “the market repeatedly chose it.”


Use a claim-to-evidence worksheet before using a loaded word

The worksheet is intentionally boring. That is a feature.

Claim — What exactly are we asserting?

Observation — What did we actually see?

Denominator — What was the opportunity base: impressions, qualified visits, sign-ups, activated users, trials, campaigns?

Time and repetition — Over what period, and across how many independent attempts or cohorts?

Qualification — What does the observation still not prove?

A launch spike and a recurring pattern are different evidence. A numerator without a denominator invites stories. One purchase can establish that one observed person was willing to pay under those conditions; it does not establish broad willingness to pay. Retention can look promising while the sample is still small. A channel can produce customers while its economics remain unknown.

The worksheet changes the default from “find a metric that sounds like traction” to “state a proposition the evidence can carry.”


Enough evidence depends on what the claim is for

There is no dashboard moment where a system flips from “not traction” to “traction.”

Evidence sufficiency depends partly on the consequence of the claim.

A private working hypothesis can tolerate more uncertainty:

This channel looks promising enough to run again.

A public portfolio claim should be narrower:

This launch produced a measurable increase in qualified visits.

A spending or hiring decision may need stronger support:

Repeated acquisition from this channel has been stable enough to justify increasing spend.

This is not an argument for paralysis. Builders have to act before certainty.

Action thresholds and truth claims can move at different speeds.

You can run the next experiment on weak evidence while still calling the evidence weak. You can keep building because one user cared without announcing product-market fit. You can invest more attention in a channel because the early pattern is encouraging without claiming distribution fit has been established.

That distinction is small, but it prevents a decision from quietly upgrading the evidence that motivated it.


Qualitative evidence still counts

None of this means “only dashboards are real.”

A customer describing an urgent problem can be meaningful evidence. An unsolicited request can be a demand signal. A buyer accepting a price can be stronger evidence than a thousand impressions for a willingness-to-pay question.

The same rule applies.

One interview establishes that one person had a problem and expressed it in a particular way. It may justify a broader hypothesis. It does not establish market prevalence.

Several unsolicited requests can justify investigating a pattern. They do not automatically establish a repeatable acquisition channel.

Qualitative evidence becomes dangerous when its scope disappears during retelling, not because it is qualitative.


What each milestone proves

The cleanest version I know is to record two statements beside a milestone:

  1. What this proves.
  2. What this does not prove.

A release can prove that the product reached a public release state. It does not prove demand, retained use, or repeatable acquisition.

A launch can prove that an acquisition attempt occurred and produced the recorded exposure or response. It does not prove the channel will repeat or that reached users will retain.

A first payment can prove that at least one observed user was willing to pay under those conditions. It does not prove broad willingness to pay, retention, or scalable acquisition.

A retention cohort can prove that the observed cohort returned at the recorded rate over the measured window. It does not prove acquisition is repeatable or economical.

Repeated channel performance can prove that the same acquisition motion reproduced a defined result across the observed attempts. It does not prove permanence or every downstream product outcome.

This is not humility as decoration. It is useful information.

And it also prevents the opposite problem: underselling real evidence. When the observation genuinely supports a stronger claim, the archive should say so.


Keep the conclusion inside the observation

The rule is not “traction metrics are fake.” It is not “revenue is the only metric that matters.” It is not “never celebrate a launch.” It is not “wait for statistical certainty before deciding anything.”

It is much less dramatic:

Keep the conclusion inside the boundary of the observation.

Shipping is evidence of output. Exposure is evidence of exposure. Use is evidence of use. Payment is evidence of willingness to pay in the observed cases. Retention is evidence about repeated value in the observed cohort. Repeated acquisition is evidence about a channel's repeatability under the observed conditions.

A stronger business conclusion can be assembled from those observations as they accumulate. It should not be smuggled into the first one.

The builder still gets to move quickly.

The evidence just gets to keep its name.


// End of transmission. State the smaller claim. — ZYANE