MVP & Experimentation

The Five-Month MVP

Engineering wants five months to build the recommendation engine. You can get the same answer in a week — with a human pretending to be one.

You're advising Pantryline, a meal-planning startup. The pitch: tell us your diet, budget, and what's in your fridge, and our AI will build your weekly meal plan and auto-order the missing groceries from a partner grocer. The founder, Teo, is convinced this is a $50/month habit for busy parents. Engineering has scoped the real thing: a recommendation engine trained on recipe and pricing data, plus live integrations with two grocery delivery APIs. Estimate: five months, two engineers, roughly $180,000 in fully-loaded cost, before a single real customer has used it.

Teo is already deep in the mechanics — model architecture, ingredient substitution logic, inventory sync edge cases. When you ask "how do you know people will actually let an algorithm choose their groceries and pay $50/month for it," Teo says: "That's what the MVP will tell us. We just have to build it first."

Here's what you know from ten minutes of digging:

FactDetail
Core untested assumptionPeople will trust an automated system to pick their groceries and pay $50/mo for it
Time to test that assumption via the planned build~5 months
Cost to test it that way~$180,000
Landing page signups (soft launch, no product)640 in 3 weeks
Signups who clicked "see my plan" (no plan exists yet)640 (100% — the button just says "coming soon")
Signups who replied to a manual follow-up email4%
Team's current confidence the AI model will even beat a competent human at week oneLow — recipe/pricing data is thin
Runway remaining7 months

The team has real distribution (640 signups) but zero evidence that the automation is what people want, versus just wanting the outcome — a done-for-you weekly meal plan and grocery order, however it's produced.

Data snapshot

Planned build time
5 months
Recommendation engine + 2 grocery API integrations
Planned build cost
~$180,000
2 engineers, fully loaded, before any real user touches it
Runway left
7 months
The 5-month build would burn most of it before learning anything
Landing page signups
640 in 3 weeks
Real interest — but interest in the outcome, not proof of the AI mechanism
Untested core assumption
Will people trust + pay for an automated pick?
Nothing built yet has tested this specific claim
Manual follow-up reply rate
4%
Weak signal on urgency — worth investigating before or instead of building
Your moveGraded against a 5-point rubric · pass at 7/10

Teo asks you: "We've validated interest with the landing page — can we start the five-month build?" Write your answer. Take a clear position on whether Pantryline should start the five-month build now, name the specific untested assumption that build doesn't actually de-risk, propose a concrete pretotype or concierge test (who does what manually, for how long, with what pass/fail threshold) that could answer the real question in about a week, and say what result would actually earn the five-month investment.

0 / 100 words minimum

Rubric

  • Clear verdict. Takes an unambiguous position against starting the five-month build now, stated early, without retreating into 'do both in parallel' as a way to avoid the call.
  • Names the real risk. Identifies that the landing page validates interest in the outcome (a done-for-you meal plan) but not the specific mechanism (automated AI picks) or willingness to pay $50/mo for it — the actual thing the 5-month build is meant to prove.
  • Concrete pretotype design. Proposes a specific mechanical-turk/concierge/Wizard-of-Oz test: a human (founder or hire) manually builds meal plans and places real grocery orders behind the same UI/promise, for a defined small cohort, within about a week.
  • Testable pass/fail bar. States a specific, falsifiable threshold (e.g. an XYZ-style hypothesis: 'at least X% of Y signups will pay $Z for a hand-built plan and reorder next week') rather than a vague 'see how it goes.'
  • Ties test result to the build decision. Explains what result would justify greenlighting the five-month engineering investment and what result should kill or redirect it, protecting the 7 months of runway.