New to Marqeable? See how it generates leads and wins customers. See the platform

How Many Growth Experiments Can You Actually Afford?

Most growth plans treat experiments as free. They sit at the bottom of the budget as a line called “testing” with a small number next to it, or with no number at all, on the assumption that someone will run them in the gaps.

They are not free, and they are not cheap. An experiment costs media spend, production hours, analysis time, and a slot in a queue that has very few slots. At a company with one or two marketers, the number of experiments you can genuinely run in a quarter is usually a single digit, and treating it as unlimited is how a team ends up with a dozen inconclusive tests and no decisions.

Here is how to size the line, what success rate to plan for, and why the answer is fewer and bigger rather than more and smaller.

The success rate you should plan for

Start with the number that reframes the whole exercise, because it comes from an organization with vastly more rigor than a Series A marketing team.

In Online Experimentation at Microsoft, Ronny Kohavi and colleagues report: “only about 1/3 of ideas improve the metrics they were designed to improve.” The paper is blunt about the distribution: roughly one third of experiments are positive, about one third are flat, and about one third actively hurt the metric they were designed to improve. It also notes that “the literature is filled with reports that success rates of ideas in the software industry, when scientifically evaluated through controlled experiments, are below 50%.”

The paper’s own framing of this finding is worth quoting for the meeting where you present it: “Evaluating well-designed and executed experiments that were designed to improve a key metric, only about one-third were successful at improving the key metric.”

Two implications for your budget.

Budget for a portfolio, not for a plan. If two thirds of well-designed ideas fail to move their target metric, an experiment line funded to produce one specific outcome is not an experiment line. It is a bet with a story attached.

The value is in the ones you stop. Kohavi’s paper makes this argument directly: a team that launches ten ideas without measuring gets roughly a third good, a third flat, and a third negative, and the net value is small. The same team that tests ten and ships only the three or four that worked captures most of the upside and avoids the damage. The return on experimentation comes from the kills, not the wins, which means the budget item you are really funding is the ability to stop things.

Microsoft’s one-in-three came after heavy upfront pruning of ideas before they ever entered a test. A startup with a looser filter should expect a lower hit rate, not a higher one. Plan on one in three at best, and treat a quarter where two of five experiments won as a good quarter rather than a disappointing one.

What an experiment actually costs

The reason experiment budgets are wrong is that they count media and ignore everything else. Cost one properly, and the number of them you can run drops sharply.

Cost componentWhat it includesUsually the bigger line at Series A
MediaAd spend, list rental, sponsorship feeOnly for paid channel tests
ProductionCreative, landing page, sequence build, tracking setup, QAAlmost always yes
AnalysisWaiting, reading, writing up, decidingUnderestimated, rarely zero
Opportunity costThe thing the same person did not do that monthThe real constraint, never in the spreadsheet

For a team of one or two, production plus opportunity cost dominates, which means the honest per-experiment price is measured in weeks of a marketer’s time rather than in dollars of media. Divide available capacity by that, and the quarterly experiment count for most Series A teams comes out at two or three, not twelve.

That is not a failure of ambition. It is the same arithmetic that makes running four channels at once unreadable, and it points at the same fix. Where the experiment line sits relative to the rest of the budget is covered in our payback horizon framework: most experiments belong in the short horizon, because an experiment you cannot read inside your review cycle is not an experiment.

Why small tests on small traffic are worse than none

There is a hard statistical floor underneath all of this, and startups run into it constantly.

The sample you need to detect an effect grows as the effect you want to detect gets smaller. Evan Miller’s How Not To Run An A/B Test gives the standard sizing rule, n = 16σ²/δ², where δ is the minimum detectable effect. The practical consequence: halving the effect size you want to detect roughly quadruples the sample you need.

A site with a few thousand monthly visitors and a low single-digit conversion rate cannot detect a 10% relative improvement in any reasonable timeframe. So teams do the natural thing and stop the test when it looks significant, which is exactly the behavior that breaks the statistics. Miller’s analysis shows that continuously monitoring a test and stopping on significance can turn a nominal 5% significance level into a 26.1% false positive rate, more than five times what the dashboard reports, and that peeking ten times makes a reported 1% level closer to 5%.

The result is not neutral. It is a confident wrong answer that gets implemented, becomes received wisdom, and quietly shapes the next four decisions.

Three ways out, in order of usefulness at startup scale:

  1. Test bigger swings. If you cannot detect a 5% change, do not test things that produce 5% changes. Test a different offer, a different audience, a different mechanism. Large effects need small samples, and a startup’s real uncertainty is usually about big things anyway.
  2. Move the measurement upstream. Cost per qualified conversation reaches usable volume far sooner than closed-won revenue does. Pick the earliest metric that is genuinely causal, per the ladder in our pipeline lag post.
  3. Accept qualitative evidence where quantitative is impossible. Ten recorded sales calls where buyers say the same sentence is real evidence. A conversion test with n=40 is not, no matter what the dashboard says about it.

Kill criteria, written before launch

The single highest-return habit in experimentation costs nothing: decide how it ends before it starts. Write these five lines into a shared document and get them agreed by whoever will be in the review.

LineExampleWhy it matters
Hypothesis”Mid-market buyers will book a meeting from a comparison page if pricing is visible”Forces a mechanism, not a wish
Primary metric, singleQualified conversations per week from that pageOne metric. Two metrics means you will pick the flattering one afterwards
Minimum interesting result8 per week, up from 3Below this, you would not act, so it is not worth running
Decision date and sample floor6 weeks or 2,000 sessions, whichever is laterRemoves peeking, which is where the false positives come from
What happens either wayWin: build 4 more. Loss: stop, do not iteratePrevents the zombie experiment that never quite ends

The last line is the one teams skip and the one that saves the most money. Without a pre-agreed loss branch, a failed experiment becomes an iteration, then another iteration, and six weeks of budget turns into a quarter of attention with no decision at the end of it.

Where AI execution actually changes the math

Worth being precise about, because the category is full of overclaiming.

AI does not change the success rate. Kohavi’s one-in-three is about how often ideas are right, and no tool fixes that. What it changes is the cost per experiment, and specifically the production component, which is the line that dominates at a small company.

If drafting the campaign, the emails, the landing page copy, the social posts, and the images is the reason an experiment takes three weeks instead of three days, then compressing that work raises the number of experiments a fixed team can run per quarter. More attempts against an unchanged one-in-three hit rate is a real improvement, and it is a defensible claim because it is arithmetic rather than magic.

Two honest caveats. First, more experiments only helps if each one is still powered enough to read, and volume constraints do not go away. Second, faster production increases the temptation to run many small tests, which is the failure mode this whole post argues against.

That is the lane Marqeable is built for. Campaigns draft the emails, social posts, blog pieces, and images for a full plan with a human approving every piece, so an experiment costs days of production rather than weeks. Automations run the follow-up so a test does not decay when attention moves to the next one. AI website chat answers the buyers each test brings in from your own business information rather than queuing them behind a form, and attribution ties revenue back to the exact message, which is what makes an experiment readable at all. We are in private beta with a small early cohort, so weigh that accordingly.

When not to experiment

Frequently asked questions

What percentage of growth experiments succeed?

Plan for roughly one in three, and treat that as optimistic. Kohavi’s Online Experimentation at Microsoft reports that only about one third of ideas improve the metrics they were designed to improve, with about one third flat and one third negative, and notes that success rates for software ideas evaluated through controlled experiments are generally below 50%. Those figures come from teams that pruned ideas heavily before testing, so a looser filter should expect worse.

How do you budget for marketing experiments?

Cost each one fully loaded: media, plus production hours to build it, plus analysis time to read it, plus the opportunity cost of what that person did not do. At a small company production and opportunity cost dominate, so the honest unit is weeks of a marketer’s time rather than dollars of media. Divide capacity by that number and you usually get two or three experiments per quarter, not twelve.

Why are underpowered A/B tests worse than no test?

Because they generate confident conclusions from noise and those conclusions get implemented. Sample requirements scale with the inverse square of the effect you want to detect, so halving the detectable effect roughly quadruples the sample needed. Stopping early compounds it: Evan Miller shows continuous monitoring can turn a nominal 5% significance level into a 26.1% false positive rate.

How many experiments should a startup run per quarter?

Fewer and larger than the plan usually says. For a team of one or two, two or three properly sized experiments per quarter is a realistic number. The right response to that limit is to increase the size of the swing rather than the count, because low traffic cannot power tests of small effects, and testing bigger changes is the only route to a readable result.

The bottom line

Experiments are a budget line whose real currency is your marketer’s weeks, not your media spend. Cost them that way and the quarterly count drops to something honest.

Plan for one in three to work, because that is what a heavily pruned pipeline of ideas produced at Microsoft, and understand that the return comes from the two you stop rather than the one you scale. Size the swings big enough that your traffic can actually detect them, because an underpowered test that gets stopped on a good-looking day is a confident wrong answer with a 26.1% chance of being noise. And write the kill criteria, including the loss branch, before anything launches.

Fewer, bigger, and pre-committed beats a long backlog of small inconclusive tests, every quarter.

See it live: Marqeable’s campaigns cut the production cost that limits how many experiments you can run, automations keep each one running without babysitting, and attribution tells you which one actually produced revenue.


Marqeable runs your campaigns, answers every visitor, text, and email in seconds, and turns them into booked jobs and meetings - even at 9pm on a Saturday. We’re in private beta with a small early cohort. Get early access

Marqeable
© 2026 Marqeable. All rights reserved.