How to test 100 ideas, shortlist 3 and launch 1

How to test 100 ideas, shortlist 3 and launch 1

By Ontevo · Published October 10, 2026

Test 100, shortlist 3, launch 1 means screening a broad set of concepts cheaply, taking a few distinct candidates into real customer and prototype testing, then launching the strongest supported offer in stages. The numbers are a memorable planning framework. Each stage requires different evidence, and stopping without a launch is a valid result.

A beautifully presented idea can reach a budget meeting before anyone asks why a buyer would choose it over the current alternative. Simulation makes it easier to explore options before that commitment. Its commercial value depends on what happens after the screen.

Our guide to testing your offer before increasing ad spend explains how to identify the buying obstacle. This playbook starts with that obstacle and turns competing fixes into a funding decision.

What do you need before screening concepts?

Start with a specific buying situation, the offer available today and evidence of an unresolved need. Add realistic delivery constraints so the shortlist can survive contact with the business.

Prepare recent buyer questions, lost-sale explanations, relevant reviews, current alternatives and basic delivery costs. Keep the source beside each observation. Separate what buyers said from what the team inferred.

For an illustrative example, suppose a lunch-container brand is investigating complaints about carrying meals on a crowded commute. Possible explanations include leakage, bulky shape, awkward cleaning and replacement-part availability. This is a fictional planning example; no customer results or market-size estimates are implied.

The proposed 100/3/1 Evidence Ladder connects a broad concept screen to a small customer-validation shortlist and a bounded launch. Each transition requires stronger evidence and authorizes a larger commitment. The counts describe a way to organize work; the decision rules determine whether any candidate moves forward.

StageActionEvidence needed to advanceStop or revise when
Test 100Explore materially different concepts using desk research and simulationA plausible buying reason, traceable evidence and a feasible testThe idea duplicates another, depends on invented facts or violates delivery constraints
Shortlist 3Test distinct candidates with relevant people and working prototypesEvidence matched to the claim: use tests for function; real choices at stated terms for demandInterest disappears at the real price, the prototype fails or uncertainty remains too large
Launch 1Release the best-supported candidate within explicit limitsAcceptable contribution and delivery performance through the buying cycleReturns, costs or operating strain breach the agreed limits

1. Make the broad screen answer a real decision

Use the buying obstacle to create alternatives with meaningful tradeoffs. A hundred cosmetic variations teach little when the unresolved question is whether the product fits the buyer's daily routine.

Input: the evidence packet and current offer. Output: comparable concept cards. Failure condition: a card needs an unsupported claim to make it attractive.

Give every card the same fields: buying situation, proposed change, total price, supporting evidence, delivery constraint and reason to keep the current alternative. Vary the feature, bundle or service condition deliberately. Remove duplicates before counting them as separate ideas.

For the lunch-container example, a narrower body, a simpler lid and an available replacement seal address different obstacles. Give the existing product a card too. Include buying a competing product or keeping the container already owned. Use comparable competitor offers with their actual price and availability.

“100” encourages breadth. Stop earlier if further ideas repeat the same assumptions. There is no prize for filling the grid.

2. Use simulation to expose objections and contradictions

Ask simulation to challenge each concept's reasoning, compare tradeoffs and suggest what a human test needs to resolve. Keep the source material, instructions and evaluation criteria consistent enough to inspect why a ranking changes.

Input: standardized concept cards. Output: an objection log and provisional shortlist. Failure condition: enthusiasm rises only after adding flattering wording or unsupported benefits.

Useful prompts include: “What would make the existing alternative preferable?” and “Which claim needs direct evidence?” For each objection, label it observed in source material, generated by the model, or unresolved. Ask a reviewer to inspect the strongest rejection as carefully as the highest score.

Research gives this use a basis. Maier and colleagues' consumer-research preprint found that a particular method for translating generated text into survey ratings could reproduce aspects of human concept rankings in personal care. Its reference statements were optimized on the studied surveys, and the authors identify limits to transfer across categories and population groups. Similarity to stated purchase intent doesn't establish actual sales.

Agreement among synthetic personas adds no independent consumer-panel evidence. The answers can share model assumptions even when the personas have different names. Bisbee and colleagues' study of political survey responses found that plausible averages could coexist with distorted variation and subgroup relationships. Its domain differs from shopping, but the measurement warning matters.

A separate preprint comparing generated preference distributions found stable answers within models alongside disagreement across models. That study didn't benchmark against human choices. Repeatability alone leaves the customer question open.

In practice, use Ontevo Simulation Engine (beta) to support early offer scrutiny. A favorable synthetic verdict earns a place in the investigation; the next gate requires fresh evidence from people.

3. Shortlist different bets and write the experiment brief

Choose candidates that test different explanations of the obstacle. Taking near-identical winners into validation can leave the team's central assumption untouched.

Input: the objection log, feasibility review and source evidence. Output: a customer-test brief for each shortlisted candidate. Failure condition: the team can't name an observation that would change its mind.

In the illustrative example, shortlist a slimmer container, a lid redesign and the existing container with a clearer closure demonstration. The last candidate challenges the assumption that manufacturing needs to change. Include a promising concept with uncertain simulation feedback when customer evidence supports investigating it.

Write the brief before recruiting:

Brief fieldWhat to record
Buyer and occasionPeople who actually carry meals on the relevant commute
HypothesisThe precise obstacle the candidate is intended to remove
ComparisonCurrent product or credible alternative at the stated price
Evidence to collectObserved handling, use, choice and reasons for rejecting the offer
Decision thresholdMinimum worthwhile behavioral improvement and acceptable uncertainty
Operating limitsDelivery cost, returns, support needs and conditions that stop the test

The causal comparison needs deliberate design. Microsoft's experimentation guidance recommends a falsifiable hypothesis, defined success measures, attention to statistical power and a suitable randomization unit. Agree those details before viewing the result.

4. Validate the shortlist with people and working prototypes

Match the test to the uncertainty. A description can test comprehension; a functioning product can test use; a real offer at real terms can test purchasing behavior.

Input: the experiment briefs and appropriate prototypes. Output: a documented advance, revise or stop decision. Failure condition: the evidence answers an easier question than the commitment requires.

Let relevant participants pack, carry, open and clean the lunch-container prototypes. Observe where they struggle before explaining the design. Compare against their current solution. Ask what they would give up to switch, and investigate people who prefer the old option.

Then test the viable offer with qualified buyers at its intended price and honestly stated availability. Where feasible, randomly assign comparable buyers to the current and changed offers while holding other material conditions stable. Plan the sample and buying-cycle coverage around the decision; repeated model outputs cannot supply missing human participants.

Treat an underpowered pilot as learning about use, objections and feasibility. A waitlist helps recruit future testers, but it leaves payment and delivery untested. If none of the candidates clears the gate, keep the current offer or investigate another obstacle.

5. Launch within delivery limits and earn expansion

Release the supported candidate at a scope the team can deliver and reverse. Write the expansion rule before orders arrive: the observation window, minimum useful result, operating limits and person authorized to hold the next commitment.

Input: customer-test results and delivery economics. Output: a bounded launch followed by an expand, hold or stop decision. Failure condition: early conversion hides delivery costs, returns or an unusually favorable audience.

For the illustrative lunch container, a simulation objection that “the narrower body may be harder to clean” should become an observed-use task with a working prototype. Ask relevant users to pack, empty and clean it under realistic conditions. Compare with the current product. Agree acceptable cleaning effort and leak performance before observing the result; changing the shape must preserve both.

That test answers a usability question. A separate offer test at real price and delivery terms answers the purchasing question. During a bounded release, track kept purchases and contribution after fulfillment, support and returns, including buyers switching from existing offers. Keep the current offer available as a comparison where practical.

Be careful with the best result from many comparisons. It may partly reflect noise. Give the selected candidate a fresh, prespecified confirmation before a large commitment. Use acquisition-cost pressure as context too: success with unusually cheap traffic may not survive continued acquisition.

Where does this leave the product team?

The method concentrates expensive learning on the decisions that survive cheaper scrutiny. Simulation can sharpen the shortlist; technical testing, purchasing behavior and delivery performance determine what deserves a larger commitment.

Ontevo Concept Studio (beta) fits the work of developing offer concepts, while Revenue Leak Intelligence connects the investigation to the gap between customer demand and the business's offering. Neither a concept presentation nor a simulated preference establishes demand, product safety or manufacturing readiness.

Questions before the next test

Must the team screen exactly 100 ideas?

No. Use enough distinct alternatives to challenge the first proposed solution. Stop when additional concepts repeat the same assumptions or the next useful evidence must come from customers.

Does “shortlist 3” mean three research participants?

No. It describes candidate concepts. Research design determines recruitment and sample requirements; complex segments or small expected differences can require more evidence.

Can a lower-scoring simulated concept still advance?

Yes. Traceable customer evidence or a valuable unresolved question can justify testing it. Preserve the disagreement and define the real-world observation that would settle it.

What if the launch result is inconclusive?

Keep the commitment bounded. Complete the prespecified collection, improve measurement for a separately planned follow-up, or stop. Don't keep adding observations until a desired result appears. A shortlist creates options; it doesn't create an obligation to launch.

Ontevo Research. Where this post carries figures, they come from Ontevo's own scan corpus or are modeled from scan patterns across the category. No figure is measured from a named customer.

Find Your Revenue Leakage
RELATED PAIN POINTS
AGENTS THAT FIX THIS
Free Rapid Scan. Your website and your email. Nothing to connect.
Scan running. Your 3 Truth Cards are being built. Keep this tab open. A few extras go to your inbox.
We could not start that scan. Check your website address and try again, or email us at hello@ontevo.ai and we will run it for you.

See what your competitors are taking. Priced in dollars.

Your top 3 revenue leaks, on screen. Your website and your email are the whole input.