…
AI Product Case Questions

Design an AI product for a ride-sharing app

A worked answer to a real AI PM interview question: design an AI product for a ride-sharing app.

Transcript

Read the full transcript (2,005 words)

[INTERVIEWER] Design an AI product for a ride-sharing app. "Design an AI product for a ride-sharing app." It sounds open-ended, but it really is not. This is a system design question wearing a costume. And here is the split that decides whether you pass. The candidates who fail treat it as a brainstorm and start reeling off features. An AI chatbot, AI safety scores, smart driver ratings, on and on.

The candidates who pass do something completely different. They name the four machine learning problems that actually live inside ride-sharing, they draw the data pipeline on the whiteboard, and then they go properly deep on exactly one of them. That is the whole game today. Let me show you how to run it. The interviewer is really testing one thing here.

Not creativity. They want to know if you can take a vague, two-word prompt and turn it into a structured engineering problem without panicking. Can you spot that AI ride-sharing is really matching, ETAs, pricing, and forecasting? Can you talk about a real pipeline, Kafka, a feature store, model serving, without hand-waving? And crucially, do you know when to stop going wide and start going deep?

By the end of this you will have a repeatable frame you can run on any "design an AI product for X" question, whether X is ride-sharing, food delivery, or a dating app. First, you spend two minutes clarifying the scope. Do not skip this to look fast. The questions you ask are the ones that change the design. Ask how large the system is.

Is it a two-sided marketplace with tens of thousands of drivers per city across dozens of cities? Ask about the latency budget, because it splits the whole system. Matching and ETA prediction are real-time, so you have seconds at most. Demand forecasting is batch, so it can run every few minutes. Then ask the tricky question. Which parts are actually AI and which are just rules?

Pricing floors, surge caps, and safety blocks are not actually models. They are business rules sitting on top of model output. Then you plant your flag. You explain that AI ride-sharing almost always means four things: matching, ETA prediction, dynamic pricing, and demand forecasting for driver repositioning. You state that you will design all four at a high level, then go deep on matching.

Saying that tells the interviewer you have seen this problem before. Next, we need to define who we are serving and what we are optimising. There are two users, and they pull against each other. Riders want a short wait and a fair, predictable price. Drivers want high utilisation, less idle time between trips, and less unpaid deadhead driving to the next pickup.

The business, on the other hand, wants marketplace liquidity, meaning enough supply sitting near demand that both sides stick around. Your North Star metric should not just be the number of rides, because that is lazy. It is completed trips per available driver hour, with two guardrails attached to it, which are rider pickup ETA and driver idle time. Stating that out loud shows you think in trade-offs rather than vanity counts.

This is the part average candidates skip entirely, and it is the part that wins the room. You name each machine learning problem precisely. First is matching. Matching is an assignment optimisation problem, not classification. Given a set of open ride requests and a set of nearby idle drivers in a short time window, you assign drivers to riders to minimise total rider wait plus deadhead miles, subject to constraints like driver heading, vehicle type, and rider rating.

Second is ETA. That is a regression problem. You predict time-to-pickup and time-to-destination from route segments, using live traffic, historical speed by road segment and time-of-day bucket, and weather. Third is surge pricing. That is a demand forecast per hex cell per five-minute window, plus a price elasticity model that maps a surge multiplier to the expected change in requests and driver supply.

Fourth is demand forecasting itself, which is a time series problem per hex cell, feeding a repositioning nudge that moves idle drivers toward demand before it spikes. Those are four distinct problems with four different model shapes, and you should list them exactly like that. Now you draw the pipeline before you talk about any single model. This is the moment you pick up the marker.

Start with ingestion. Driver GPS pings every four seconds and rider app events land on Kafka. That flows into streaming feature computation in Flink, where you are rolling up supply and demand counts per cell, live segment speeds, and driver idle time. That feeds a feature store. And here is a detail that lands well: an online store in Redis for sub-100-millisecond serving, and an offline store in a warehouse for training, sharing the exact same feature definitions so training and serving do not skew.

Next is model serving. Your ETA model sits behind a low-latency endpoint, target under 100 milliseconds. Then comes the decision engine, which is the matching optimiser running in batches every few seconds. The trick that makes this tractable is spatial indexing. You partition the map with H3 or S2 cells, so matching runs per cell and its neighbours, not globally across the entire city.

Inside each cell you solve with the Hungarian algorithm for small batches, or a greedy constrained assignment when the batch is large. Mentioning per-cell matching shows a real marketplace engineer that you understand the domain. Now let us cover where the training data comes from and how you ship changes safely. Training data comes from completed trips: realised pickup times, realised ETAs, whether a driver accepted at a given surge level.

You must name the counterfactual problem out loud, because this signals real depth. You only ever observe the outcome for the match you actually made. You cannot naively learn if that was the best assignment, because you never saw what the alternatives would have done. The fix: log the full candidate set the optimiser considered, not just the winner, so you can do counterfactual evaluation later.

When you roll a change out, you stage it. Run offline evaluation first, then shadow deployment where the new model scores live traffic but does not control anything, and finally a geo-based A/B test. Whole cities or cells are split into control and treatment. The reason we use geo and not user-level is simple. You cannot randomise individual riders in a shared marketplace without contamination.

One rider's match affects the next rider's options. Explaining that reason is the detail that separates you from the pack. Finally, cover the failure modes and trade-offs, and give this a solid two minutes. Cold start in a new city with no trip history: you bootstrap ETA from open map data and generic priors, and you widen surge caps slowly.

Then there is the nasty issue of feedback loops. Surge suppresses demand, which lowers observed requests, which the forecaster reads as low demand, which drops surge, which brings demand back. The forecast is chasing its own tail. You mitigate by forecasting latent demand from app-open and search events, not just completed requests. Fairness matters too. An efficiency-only objective can starve drivers in low-demand suburbs of income, and that is a real business risk, not just an ethics footnote.

And GPS drift in dense downtowns, the so-called urban canyons, corrupts both ETA and matching, so you snap pings to the road network and down-weight low-confidence positions. The interviewer will ask you to go deep somewhere, so let us go deep on matching, because it is the beating heart of the marketplace. Every few seconds, inside one H3 cell, you have got a batch of open ride requests and a batch of idle drivers.

You build a cost matrix. Each driver-to-rider pair gets a cost, and the cost is not just distance, it is estimated time-to-pickup from your ETA model, plus a penalty for deadhead miles, plus soft penalties for a bad heading or a vehicle-type mismatch. Then you solve. For a small batch you use the Hungarian algorithm, which finds the assignment that minimises total cost optimally.

For a large batch, where the cubic cost of the Hungarian algorithm gets painful, you fall back to a greedy constrained assignment that is near-optimal and fast. And here is the trade-off you name out loud: the batching window. Wait longer, say five seconds instead of two, and you gather more drivers and riders, so the optimiser finds better global matches.

But every rider now waits those extra seconds, staring at a spinning app. So the batch window is a dial between match quality and perceived responsiveness, and you would tune it per cell based on density, tighter in a busy downtown, looser in a quiet suburb. Discussing that kind of concrete trade-off on a single component provides exactly the depth they are looking for.

Let me make this concrete, because an interviewer wants to see you ship something, not just describe a platform. Pick one wedge and go. Predictive driver repositioning to cut idle time. You model demand as a per-H3-cell time series with a gradient-boosted regressor. Features: last-hour requests, day-of-week and hour buckets, weather, local events. When it predicts a spike in a cell ten minutes out, you push a soft nudge to idle drivers within two neighbouring cells: "Head toward Downtown, high demand expected." Your primary metric is driver idle time per hour.

Your guardrail is rider pickup ETA, because repositioning must not empty out the cells those drivers just left. And here is your success bar for the first geo A/B: idle time down eight percent with pickup ETA flat or better. Now, if the interviewer pushes and asks what happens if drivers ignore the nudge, you have got an answer ready.

You add a small incentive, you measure nudge-acceptance rate as a second guardrail, and you personalise the nudge to drivers who have historically repositioned. Having that follow-up ready is what turns a good answer into a memorable one. Here is what makes them lean in. First, that you called matching an assignment optimisation problem and not classification. That single distinction signals you actually understand the domain, as most candidates get it wrong.

Second, the spatial indexing and the counterfactual logging point. Those are the gritty details a real marketplace team lives with every day, and you cannot fake them. Third, choosing geo-based A/B over user-level randomisation and explaining why, marketplace contamination. And fourth, quietly holding the two-sided tension the whole way through, noticing that a change which helps riders can wreck driver economics.

That awareness is a sign of maturity, and they will be watching for it. Now the traps, because these sink good candidates fast. The first trap is listing features instead of naming machine learning problems. "AI chatbot, AI ratings, AI safety." That is a brainstorm, not a design, and the interviewer switches off. The second trap is going shallow on all four problems to seem thorough.

Do not do that. Pick one, go deep, gesture at the rest. Depth in one area plus awareness of the others is exactly what gets rewarded. And the third trap, the quiet killer, is forgetting the marketplace is two-sided. If you optimise rider wait and never mention driver income, the interviewer is just sitting there waiting to see if you will notice.

Make sure you notice it out loud. So let us assemble the whole picture. You clarify scope in two minutes. You name two users and a real North Star. You call out four machine learning problems: matching, ETA, surge, and forecast. You draw one pipeline: Kafka into Flink into a feature store into serving into a per-cell optimiser. You handle training data, the counterfactual problem, and geo A/B.

And you go deep on one wedge with a real success bar. Carry this one line into the room. Ride-sharing AI is four machine learning problems on one pipeline. Name them, draw it, then go deep on exactly one.

Keep learning