The pitch as first stated: a native-only app (no web), a personal AI concierge, privacy-first, covering travel, dining and outdoor.
01Is the problem real?
Real but shallow, and three unrelated pains are bundled into one product.
- Travel planning is genuinely painful but episodic. Most people plan one to three trips a year. A subscription with two annual usage events has a retention problem you cannot out-design.
- Dining decisions are high-frequency but low-stakes. "Where should we eat" gets solved in forty seconds on Google Maps or in a group chat. Low pain intensity means low willingness to pay.
- Outdoor is the one with real substance: permits, conditions, weather windows, gear, access rules. AllTrails proved people pay. That also means the pain is already monetised by someone with a decade of user-generated content.
A concierge is only useful if it knows your calendar, location, contacts, budget and history. "Privacy-first personal concierge" is close to a contradiction unless you go on-device, which caps quality.
Revealed preference says consumers don't pay for privacy. They pay for a better product that happens to be private. Proton and DuckDuckGo work because they substitute for something people already do daily. A concierge is net-new behaviour, so privacy can't carry it.
What they do today
ChatGPT or Gemini for ideas, Google Maps for the decision, Instagram and TikTok saves for inspiration, a group chat or Notes for the itinerary, Resy or OpenTable to book, AllTrails for trails. Fragmented, but every piece is free and good enough.
02Who is it for, and why now?
Three verticals imply three people who share almost nothing.
- The six-trips-a-year leisure traveller, 28–45, household income $150k+, already carrying an Amex with a concierge attached.
- The city diner who uses Beli and Resy notify and has opinions about pasta.
- The serious outdoors person who pays for AllTrails Plus and Gaia and loses permit lotteries.
That overlap is you, not a market.
On urgency this is the weakest part of the pitch. Nobody has a concierge emergency. The only genuine urgency triggers in your space are a trip booked three weeks out, a permit lottery deadline, and a hard-to-get reservation drop. Each already has an owner: Wanderlog, Recreation.gov, Resy.
03Competition, including the indirect kind
Filter the landscape by type. The category incumbents are the obvious threat; the platform layer is the real one.
Free, already installed, review network effects, booking integration, and the cost of re-teaching a new assistant everything about you.
04First ten paying customers
Don't build an app for this. Be the concierge.
- Pick one vertical and one hard trip type. Dolomites hut-to-hut, JMT permits, Japan in cherry blossom season, a first Patagonia trip. Hard problems justify payment.
- Set a price today. $75–150 for one planned trip. Money is the only signal that counts; free users teach you nothing.
- Go where they already are. Subreddits, Facebook groups for specific trails and regions, Strava clubs, climbing gyms, running clubs, alumni Slacks. Participate genuinely for a week before pitching.
- Answer questions publicly, convert privately. Post a thorough, free, obviously useful plan for someone's trip. Do it five times. People will message you. That's your funnel.
- The pitch: "I'll plan your trip in 48 hours — permits, conditions, reservations, day by day. $100, full refund if it isn't useful."
- Use AI heavily behind the curtain. You're testing demand for the outcome, not the technology.
- Ask every buyer three questions: what did you try before, what did you almost pay for instead, would you pay again next trip.
05MVP scope and speed
Weeks 1–2: no app at all
A landing page, a Stripe link, and a Telegram or iMessage thread where you do the work by hand. Target: ten paid deliveries. If ten people won't pay $100 for something you do manually, an app doesn't fix that.
Weeks 3–6, only if that works
A TestFlight-only native build. One vertical, one flow, no accounts, local storage, manual backend.
Exclude aggressively
Android, web, booking, payments, multi-vertical, social and sharing, on-device inference, end-to-end-encryption architecture, a memory system. Privacy engineering is expensive and no pre-product-market-fit user is paying for it.
Validation bar before building more
At least 30% of buyers repeat within 90 days, and at least one in five refers someone unprompted.
Two things I'd change about the pitch
Travel planning happens on laptops. Shared itinerary links were Wanderlog's entire growth loop. No web means no search traffic, no shareable artefact, no viral surface, and total dependence on App Store discovery. The privacy rationale doesn't hold — you can build a private web product.
Google, Yelp, Booking and TripAdvisor all monetise through ads and paid placement. A concierge that takes no commissions and no placement money works for the user, not the restaurant. Sharper, more sellable, and it survives the question "but Apple says they're private too."
The clarification: critical personal data stays on the device, only metadata goes to the cloud; an open-weight LLM runs locally on the phone; the agent learns your behaviour and improves over time.
What the clarification actually changed
The architecture. Not the problem. And it created a new one.
Both pillars became operating-system features on your only platform, and the timing is not in your favour.
On-device inference
Apple's Foundation Models framework already gives third-party apps direct access to a ~3B on-device model with no API key, no network and no per-token cost. The iOS 27 version rebuilt that model to be better at instruction-following, added image input, opened developer access to Private Cloud Compute with a 32K context window, and added a Core AI framework for running custom models locally. Shipping your own open-weight model is now a supported integration path rather than an edge.
A private assistant that learns you
Siri AI routes queries across on-device models, Private Cloud Compute and a custom Gemini model. It draws on personal context across messages, emails and photos, acts across apps, pulls from the web, and syncs history privately through iCloud. One walkthrough of the new Siri uses this exact example: finding well-reviewed Japanese restaurants under a price cap and checking Saturday availability, collapsing a Maps–TripAdvisor–OpenTable workflow into one query.
That is your pitch, your architecture and one of your three verticals — free, preinstalled, from the most trusted privacy brand in consumer technology. Verify the current shipping state yourself before acting on it, but plan as if it's true.
01Problem, re-examined
You picked the three categories where on-device helps least
Value in travel, dining and outdoor is overwhelmingly fresh external world data: hours, availability, prices, closures, trail conditions, whether the place is still good. A 3B local model has close to zero reliable world knowledge here and will confidently invent restaurants. So you call the cloud for almost every useful answer, and the query itself leaves the device.
Metadata is the sensitive part
"Gluten-free, stroller-friendly, $$, within ten minutes of these coordinates, Tuesday 7pm" is more identifying than a name attached to it. If a researcher or journalist pulls your traffic and finds that, the privacy claim doesn't survive the write-up. Privacy-positioned companies are held to a far higher standard than everyone else.
Privacy premiums appear where the payload is sensitive
Signal, Proton and Standard Notes work because the content itself is the secret. In dining and travel the payload is public information about businesses. Users don't feel exposed, so they won't pay to be un-exposed.
02The data-density problem
Personalisation only compounds where the signal is dense enough to compound before churn.
| Category | Signal events per year | Time to know you | Verdict |
|---|---|---|---|
| Travel | 2–4 | ~3 years | Too sparse |
| Outdoor | 20–40 | ~1 season | Workable |
| Dining | 200–500 | Weeks | Dense enough |
- Cold start lands on the churn cliff. Consumer subscriptions live or die on day-1 to day-7 retention. Your product is at its worst precisely then, and "trust me, it improves" asks users to fund a promise before seeing evidence.
- Local-only forfeits collaborative filtering. The strongest signal in any recommender is other people like you. You'd learn from one device while competitors learn from hundreds of millions. You are buying privacy with product quality, and the buyer didn't ask for that trade.
- Device gating is worse than it looks. Ship your own quantised model and you're on recent flagships with a 1–2GB post-install download before first use. Large drop-off at the worst moment, most of the installed base excluded, and the survivors skew toward people already paying for ChatGPT or Gemini.
Honestly, the privacy-native segment: Proton, Mullvad, GrapheneOS, the de-Google crowd. Real, reachable, and some of them genuinely pay. Also small, churn-prone, allergic to subscriptions, and Apple just told them they're covered. Build a model where that segment is enough, or find a different one.
03Competition, re-examined
- The platform layer, existential. Foundation Models for the infrastructure, Siri AI for the product. Free, preinstalled, no download, no battery cost, works on older phones, and bundleable with iCloud+ tomorrow.
- A wrinkle worth thinking about. The next generation of Apple Foundation Models is being built on Google's Gemini models and cloud technology under a multi-year deal. That leaves a narrow opening for people who read "privacy" as "never leaves my device, not even to a trusted enclave." Real, but small and ideological — Private Cloud Compute's attestation story is strong enough that most users won't split the hair.
- Cloud assistants. ChatGPT with memory, Gemini with personal context, Perplexity. Better world knowledge, better reasoning, cross-device sync, free tiers.
- Category incumbents. Google Maps, Beli, Resy, Wanderlog, AllTrails, Gaia, onX, Recreation.gov. They own the data and the booking rails.
04First ten customers, re-aimed
You're no longer testing whether planning help is valuable. You're testing whether accumulated memory is what people pay for.
- Run a longitudinal concierge, not a one-off. Sell a six-to-eight-week subscription at $30–50 a month, dining-led, to ten people. Do it by hand, keep a real profile, and reference it out loud: "skipping this one, you didn't love the last natural-wine place."
- Measure the slope, not the level. Score satisfaction and willingness-to-pay at engagement one versus engagement five. If the curve is flat, the learning thesis is dead and no amount of on-device engineering fixes it. That single number is worth more than your next six months of code.
- Price-test privacy directly. Two landing pages, same offer, one privacy-framed. Measure the conversion delta. Two hundred dollars and a weekend to falsify or validate your core positioning.
- Recruit where the thesis should be strongest. Privacy-native communities intersected with the category. If those users won't pay, nobody will — and you learned it free.
- Ask one question before taking their money: what would make you cancel. That answer is your roadmap.
05Two parallel tracks, not one MVP
You cannot ship a native app with a bundled on-device LLM in two weeks. Run demand and feasibility at the same time.
Track A — demand, weeks 1–2
Landing page, Stripe, manual delivery over iMessage or Telegram. Target ten paying subscribers. No code.
Track B — technical spike, weeks 1–2
A throwaway build that answers whether this is feasible at all. Set the pass marks before you start:
| Measure | Pass mark |
|---|---|
| Time to first token | Under 1.5s on target device |
| Peak memory | Comfortably under the iOS jetsam budget, with headroom |
| Thermal behaviour | No throttling across a 10-minute session |
| Battery drain | Under ~3% per session |
| Install-to-usable download | Measured, with the drop-off modelled |
| Device coverage | % of installed base clearing the hardware bar |
Start on Apple's framework, not your own model. Free, optimised, ships with the operating system, costs you no download. Swap in your own weights later, and only if the spike proves the system model is insufficient for your specific task. Shipping 2GB to make an architectural point is a decision you make after you have users.
These aren't three versions of one company. They're three companies with different customers, moats and failure modes. Arguing about which is best in the abstract will burn a month.
The criteria that predict survival
In rough order of how much each one determines the outcome.
- Task–model fit. Does a ~3B local model actually do this job, or are you fighting your own architecture? Small models are weak at world knowledge and decent at extraction, classification and retrieval over text you hand them.
- Is local functionally required, or ideological? If the customer would still buy it with a cloud backend, local is a cost decision, not a product decision.
- Apple-proof? Structurally blocked, or merely not there yet?
- Frequency. Determines whether personalisation compounds before churn.
- Who holds the budget, and how big is it?
- What compounds? Data, workflow lock-in, or nothing.
- Time to first dollar.
Path A — Informed, not private
The argument for on-device isn't secrecy, it's access.
The strongest version is not capture. It's resurfacing. Capture is already commoditised. A crowded category does social-media-to-map extraction: Plotline, Stasht, Rodeo, Dream Trip, Wanderlist, DocentPro, Mapstr and more. Plotline reports around 30,000 travellers mapping two million places from half a million posts. Mapstr is described as the veteran with years of track record. Google Maps has added screenshot matching for text already captured in a frame.
A dozen entrants, several years, no breakout. Plotline's scale is small. Mapstr has around 2,500 ratings. That usually means retention is bad or distribution is the wall. My suspicion is retention: saving is a dopamine act, retrieving is a chore nobody does. Notably, one of these companies says the differences between them show up in resurfacing — nearby alerts, calendars, reminders, daily picks, or nothing. That's the competitors telling you where the unsolved problem is.
Why local wins here, if it wins
Every good resurfacing trigger lives on the device: where you are, what's on your calendar, the time, who you're with, what you screenshotted eleven minutes ago, whether it's raining. A cloud app gets that only if you hand it over continuously. And the task shape — extract, classify, retrieve over a personal corpus — is exactly what a small model is good at. Best task–model fit of the three.
Where it breaks
- It's a feature, not a company. Google and Apple are both walking toward it.
- The crowded-but-flat category is evidence the underlying behaviour doesn't retain.
- Your iOS-native-only instinct is backwards here. Android is the permissive platform — notification access, share targets, background services, accessibility APIs. The agent that sees everything is far more buildable on Android, and Apple has already claimed the iOS version.
Find twenty heavy savers. Ask them to open their Instagram saves or Maps lists in front of you and name the last time they retrieved something and acted on it. If most can't name one, the resurfacing thesis is dead and you saved a year.
Path B — Go where Apple won't
The only path where your architecture has a non-ideological justification.
In the backcountry there's no signal. Local inference isn't a privacy stance there, it's a functional requirement the user feels immediately. You stop having to convince anyone that privacy is worth paying for, because what you're selling is "it works when nothing else does."
The wedge isn't trails
AllTrails owns trails and has a decade of user-generated content you can't replicate. The wedge is conditions and permits, which is genuinely fragmented and genuinely painful:
- Conditions synthesis across weather models, snowpack, streamflow, fire and smoke, road and area closures, agency alerts, and recent trip reports. Every one lives somewhere different and nobody assembles them.
- Permit tracking across Recreation.gov, state agencies, individual parks, and international systems like Dolomite rifugi, New Zealand Great Walks, Japanese huts. No international aggregator exists.
- Multi-day logistics: shuttles, resupply, bear canister rules, group size limits, campsite reservations.
Frequency and willingness to pay are both adequate
The serious user does ten to thirty outings a year plus obsessive pre-trip research — enough for personalisation to compound. And they already pay: AllTrails Plus, Gaia, onX and CalTopo all have real paying bases.
What compounds
The data pipeline. Normalising dozens of agency feeds is unglamorous, tedious work that neither Apple nor an AI lab will ever do. A moat made of effort rather than technology, which is the most durable kind available to you.
Where it breaks
- Consumer price ceiling is low, roughly $30–60 a year. You need volume or an adjacent business.
- AllTrails and Outside Inc. can build conditions synthesis. Your defence is speed and depth, not exclusivity.
- Liability is real and you should design for it now. If your agent summarises conditions as favourable and someone gets hurt, a disclaimer may not save you. Present the inputs and what they say rather than "you should go." That's also a better product.
Sell a $25 go/no-go conditions and permit brief for specific named objectives, delivered by hand in 24 hours. Post in the relevant subreddits and Facebook groups. Measure repeat purchase within 60 days. Repeat rate is the whole signal.
Path C — Privacy as a requirement
Be honest with yourself: this is a different company. The consumer concierge doesn't survive the transition.
C1 — Executive and ultra-high-net-worth discretion
Family offices, chiefs of staff, executive assistants, corporate security. Here privacy genuinely is a purchase criterion: a principal's movement patterns are a security asset, and a deal team's travel patterns leak M&A activity — people track corporate jets for exactly this reason. Buyers have real budget and already pay for human concierge services. The problem: tiny market, relationship-driven sales, and the buyer's honest preference is a human who answers the phone at 2am.
C2 — On-device AI for field and regulated work
Insurance adjusters, home health aides, utility inspectors, maritime, law enforcement. Offline plus sensitive data plus structured extraction. A real and growing category that fits your architecture well — and has nothing to do with travel, dining or outdoor.
You become an enterprise software company with six-to-twelve-month sales cycles, and none of your consumer instincts transfer. Private Cloud Compute and its attestation guarantees also satisfy many enterprise buyers who would otherwise need on-device, which narrows the set of customers who structurally require what you're building.
Ten discovery calls with one named buyer type. Not "would this be useful" — ask what they use today, what it costs, and who signs. If three ask for a quote, it's real. If they all say "interesting, keep me posted," it isn't.
How they score
| A — Informed | B — Outdoor | C — Regulated | |
|---|---|---|---|
| Task–model fit | Strong | Good | Good |
| Local required? | Preferred only | Functionally required | Required for some buyers |
| Apple-proof | Weak on iOS | Strong | Strong |
| Frequency | High | Medium–high | Not applicable |
| Budget size | Low, consumer | Low–medium, consumer | High, enterprise |
| What compounds | Personal data, weakly | Data pipeline | Contracts, switching cost |
| Time to first dollar | Weeks | Weeks | Months |
A personal memory layer inside an outdoor product is coherent: your saved objectives, past trips, pace and tolerances, surfaced against live conditions. That's a real product using the best parts of both paths. C composes with neither — choosing it means restarting.
Answering the question left at the end of round two
When Siri AI ships with personal context and free on-device inference, what can you do that it can't?
- Path A: on iOS, honestly not much. On Android, a lot. This path only holds up if you invert your platform strategy.
- Path B: Apple will never scrape snowpack sensors, fire incident feeds, streamflow gauges and agency closure notices, or model permit lottery odds. It won't work offline in a canyon with your gear list and your pace. A real answer with a specific capability attached.
- Path C: Apple doesn't sell to compliance officers or sign a business associate agreement. Also a real answer — for a different company.
Only B answers the question without requiring you to change platform or industry. That's why I'd weight it highest despite the smaller market.
Score them against what you actually care about
Weights start where I'd set them. Move them and watch the ranking change.
If your weights produce a different winner than mine, the disagreement is about what matters, not about the facts. That's a much more productive argument to have with a co-founder.
Don't pick by argument. Run all three tests concurrently, cheaply, with pass marks set before you start. Total cost: thirty days and a few hundred dollars, against a year spent building the wrong thing.
Week one
- Twenty saved-folder interviews (A).
- Ten discovery calls with one enterprise buyer type (C).
- Post three free, genuinely excellent conditions briefs in outdoor communities (B).
Week two
Put up three one-page offers with Stripe links. Charge for all of them. Free signups are noise.
Weeks three and four
Deliver by hand. Track one number per path.
Pass marks
Tick them off as they clear. Nothing is saved when you reload — this is a worksheet, not a database.
Nothing cleared yet. Set the pass marks before you start the tests, not after you see the results.
Deciding
If two paths pass, pick the one where you'd enjoy the next five years, because that genuinely becomes the deciding variable once the evidence ties.
If none pass, you've spent thirty days and a few hundred dollars to avoid spending a year — which is the best possible outcome of a reality check, and the one people are least willing to accept.
Local inference means no per-user token cost and near-100% gross margin. That's a genuinely good economics story whichever path you take. It belongs in the model, not the marketing.