PRODUCT · notebook spread

Build to learn vs build to earn.

a 60% prototype, scoped to ten campaign managers

Marty Cagan's distinction between build-to-learn and build-to-earn has been getting passed around in product Slacks for a year now, mostly read as advice for early-stage teams. The argument I want to add, after a quarter of running the experiment myself on an established B2B marketing-workflow platform with thousands of paying enterprise marketing teams attached to it, is that the biggest scope for build-to-learn in 2026 is sitting inside a working build-to-earn product, not in the lean-startup garage everyone keeps drawing it in.

read the spread

two pages, same desk, same Sunday-night notebook.

The trigger for this essay was a 40 minute call last Sunday night with two PMs on my own team about whether a feature we'd been planning belonged in the roadmap deck or in the experiments backlog. Somewhere around minute 26 it became clear that the framing we'd inherited from senior leadership did not have a category for what we actually wanted to do. The plan was to ship an inline AI for the long, multi tab campaign setup forms that marketing operators spend half a working day filling in. It would be scoped to a single tier of enterprise customers, run for ninety days behind a flag, and then either hardened into a permanent slot in the workflow or quietly removed altogether. The engineering lead on the call kept calling it a feature, with the launch plan and the design system inheritance and the marketing brief that vocabulary drags behind it.
Nobody on the call had the better word for what the thing actually was.
What it actually was is a build to learn experiment running inside a build to earn product, with all the political headwind that combination produces inside an established platform.
Marty Cagan's piece on SVPG is the cleanest articulation I have read of the distinction the team was groping for on that Sunday call. His argument is that product discovery and product delivery are different activities, with different success criteria and an almost opposite relationship to quality. Confusing the two is, in practice, the most common way a team will burn a quarter while shipping nothing that learns and nothing that earns. The shape of discovery is a 60% prototype in front of ten real customers in two weeks. The shape of delivery is a 99.9% production system in front of every paying customer indefinitely, and a team running both as if they were the same activity will be bad at both.
The cost of delivery is collapsing while the cost of doing discovery badly stays exactly where it was, which means the competitive advantage of a product team in 2026 sits almost entirely in how quickly and how cheaply it can learn whether the thing is worth building before the thing gets built.
Marty Cagan, SVPG
The version of Cagan's argument that gets passed around in product Slacks tends to land most heavily for founders before product market fit, because the rhetoric of build to learn maps cleanly onto what Eric Ries was already saying in The Lean Startup about validated learning, and onto what Steve Blank was saying a decade earlier about customer discovery. The mental picture is a small team in a garage with no customers yet, learning by shoving prototypes in front of strangers in coffee shops until somebody pulls out a credit card. It is one valid shape of the work, and it happens to be the shape with the worst possible discovery substrate. The team at that early stage has to spend most of its energy finding someone willing to stand still long enough to look at the prototype.
The inverse case, which Cagan's piece does not spend much time on, is that the team with the best possible discovery substrate in 2026 is the established B2B SaaS company with thousands of customers running real workflows every day. Any 60% prototype dropped behind a feature flag on a Monday morning will have ten cohorts of usage data attached to it by the second Wednesday. The audience a founder before product market fit spends most of the week trying to manufacture is already on the rails for the established platform. The cost of getting signal on a discovery experiment is the cost of writing the flag plus a prompt inside the app asking ten campaign managers across three enterprise marketing customers whether they want to try the new thing.

Build to learn

Pencil sketch

A 60% prototype dropped behind a flag for ninety days, scoped to a small cohort of campaign managers who know the thing might be removed on day 91, with the success criterion expressed as one number the team is willing to lose against and walk away from.

Build to earn

Inked and stamped

A permanent slot in the platform with an SLA behind it, a deprecation cost that runs for years, every paying customer enrolled by default, and a migration plan for the day the team eventually has to move the surface area around.
Most established platforms do not use this substrate, which is the part that frustrates me into writing the essay. The product culture inside a Series C company or one that has already gone public hardens, somewhere between two hundred and four hundred engineers, around build to earn as the default mode. Every release runs through a launch plan, a security review, a marketing brief, a localisation pass, and three layers of stakeholder signoff, regardless of whether the thing being released is meant to live for a decade or for ninety days. The overhead of build to earn becomes the overhead of building anything at all, and the cheap experimental loop the platform should be running every quarter just stops happening, because the cost of any new thing has become the cost of a permanent thing.
The argument I want to make is that the established platform should be shipping five to ten build to learn experiments inside the product every quarter. Each should sit behind a feature flag, scoped to a small customer cohort, explicitly framed with a ninety day window and a "this might be removed" notice in the UI itself. That cadence is achievable and also unusual. The enterprise marketing software company I currently work at is somewhere in the middle nine figures of annual revenue, and roughly three quarters of the platform is firmly build to earn. It has been running an agentic AI initiative built in house for marketing workflows for the better part of two years, work I helped shape from the product side. The parts of that initiative which worked, in retrospect, were the parts that ran as build to learn experiments dropped inside the existing build to earn surface, with small cohorts of marketing customers opting in for ninety days. The parts which struggled were the parts where the launch machine demanded a full release for what should have been a discovery prototype. The embedded experiment is structurally faster to learn from than the standalone one, because the marketing team is already on the platform running real campaigns and the signal arrives without anyone having to recruit anything.
The Bangladesh examples that taught me what this looks like in the wild are mostly the consumer ones, because the consumer companies here have been forced to learn discovery faster than the B2B ones. bKash runs as a build to earn rail for mobile money at the scale of three hundred thousand agents. The merchant and QR features bolted onto that rail over the last four years all shipped first as quiet build to learn flags to small cohorts of retailers in specific neighbourhoods, before any rollout at the scale of the platform. At Pathao the same play shows up around the food and pharmacy verticals, both of which started as build to learn experiments inside the ride share rails. Pathao Pay was the verticals experiment that did not stick, and the platform had the institutional posture to kill it when the numbers did not show up, rather than spending another two years dragging it across the finish line. The same pattern repeats at Daraz with Live, Mart, and Pharmacy, each of which lived as a small cohort experiment inside the marketplace for months before being given permanent dashboard real estate.
What those three companies share, and what an established B2B SaaS platform usually does not, is an honest culture around the kill. A build to learn experiment that is never killed is a build to earn launch that was lying about its category. The discipline of running ten experiments a quarter only works if six of them are removed on day 91 with no apology owed to anyone, because the contract the customer signed up to when they enrolled was an experimental one. The agentic marketing workflow work was, for its first eighteen months, a real example of this discipline inside a B2B context. Features that did not earn their adoption curve with the enrolled marketing teams in the window were quietly removed and the team moved on; the ones that did earn their curve were promoted into the build to earn surface with a proper launch.
The cornered resource argument from Hamilton Helmer's 7 Powers is the cleanest way I have found to explain to a CFO why this matters. Helmer treats the customer base of a B2B SaaS company not as an acquisition cost that gets paid once but as a renewable asset that produces compounding returns over time. The version of the argument Cagan would make on top of Helmer is that the customer base is also the rarest possible discovery substrate. The platform which does not run its experiments inside that base is simply paying the cost of the moat without using its primary advantage. Patrick McKenzie has been making a neighbouring argument on Bits about Money about how the company already sitting inside the customer's accounting close has the cheapest path to test the next adjacent product. The platform already sitting inside the marketing team's workflow from brief to publish can test the next adjacent feature for the price of a flag, while the startup testing the same feature has to acquire the enterprise marketing customer first.
The continuation of last week's argument lives here too. The tactical version of that argument was about how to respond to a single customer's feature request without committing to ten years of maintenance. The strategic version is the same shape at a different scale, where the platform itself runs build to learn experiments instead of build to earn launches as the default response to any new product direction. The marketing ops lead at a global insurance customer who once asked us for a brand compliance review dashboard would have been better served by a build to learn flag scoped to twenty enterprise customers for ninety days than by the build to earn project we actually ran, which took fourteen months. The platform would still have the engineering capacity the unused dashboard has been quietly consuming since.

The lifecycle of a build to learn flag inside a build to earn product

  1. proposed
    one-sentence hypothesis, a real number attached, a 90-day window written into the deck
  2. prototype
    60% build behind a flag, no SLA, no localisation, no security review beyond table-stakes
  3. cohort
    10-50 enrolled customers see the experimental UI, with a banner that says this might be removed
  4. signal
    the success metric either shows up by day 60 or it does not
  5. kill or harden
    day 91, the flag is either turned off and the code is removed, or it is promoted to a real build-to-earn launch
Six flags out of ten should be killed on day 91, and that is a healthy ratio.
The senior leadership at most platforms has to decide, before any of this works, that the kill ratio is the metric. A team promoting nine out of ten experiments has stopped running build to learn and started running build to earn with extra steps. A team killing nine out of ten has, more often than not, the substrate of a real discovery practice, and is using its customer base the way Helmer's framework would predict. The discomfort senior leaders feel about authorising a kill rate of sixty percent has to be unwound before any of the cadence I am describing actually starts to happen.
My own team has been running roughly this practice for two quarters, partly as an experiment on the people running the experiments. We have shipped eleven flags inside the build to earn marketing workflow platform in that window. Four were killed on day 91 with no replacement, three were killed with a redesigned successor proposed for the next quarter, two were promoted to permanent slots in the workflow with a real launch behind them, and two are still inside their ninety day windows. The eleven flags between them have collected real usage signal from north of two hundred and fifty enterprise campaign managers sitting across customer marketing orgs in Berlin, Boston, London, and Singapore. That is a substrate no startup before product market fit in the same space could realistically manufacture in a year of focused work on customer development. The cost of the kills, when they come, is the politeness of a small banner inside the app thanking the cohort and explaining the flag is going away on the 21st of next month.