What goal for an AI-only social network OpenAI is building?
Transcript
Read the full transcript (1,526 words)
[INTERVIEWER] What goal for an AI-only social network OpenAI is building? Here is the question. OpenAI is building a social network where the content, and maybe even some of the other people, are AI. What is the goal? Look, the fastest way to fail this one is to say daily active users or engagement. A social feed powered by AI can pump out infinite content, so raw attention numbers are trivial to inflate and they tell you nothing.
The strong answer refuses the vanity number, picks one North Star that actually captures the value a real person gets, and then defends it against gaming. What the interviewer is really probing is whether you can pick a goal that's honest. Anyone can name a metric. The skill is naming a metric that goes up only when the product genuinely helps someone, and then wrapping it so it can't be optimised into harm.
In the next few minutes I'll show you how to pin down what the product even is, name the value exchange, propose a North Star you can actually instrument, and guard it. Do that and you've answered every define success question in this round. First, don't answer until you've said what AI only means, because it's ambiguous and the goal depends on your read.
Is this a feed of AI generated content people scroll? Is it users talking to AI personas? Or is it humans posting, with AI as the connective tissue that summarises, matches people, and drafts replies? I'd say my read out loud. Humans are the users, the network helps them create and connect, and AI is the substrate underneath. The goal follows the read, so pin the read first, and if the interviewer wants a different one, adjust.
That takes about a minute and it stops you answering the wrong question. Now, what does a social product actually do for someone? It creates value when the user gets something back that pulls them in again. A good conversation. A piece of content they're glad they saw. A connection they wouldn't have made on their own. That felt value is what your North Star has to proxy.
Here's the thing about time on app and DAU. They measure attention, not value, and on an AI feed with an infinite content firehose, attention is the easiest thing in the world to manufacture. A metric you can move without a single human being better off is a broken metric. So we need one that can't be faked that way.
My North Star is weekly users who have a meaningful AI assisted interaction. And I'm not going to leave meaningful as a vague adjective, because a hiring manager will pounce on that. I define it as a session where the user did something with intent and then engaged with what came back. They sent a reply. They saved or shared the content.
Or they carried a thread past two turns. Call it meaningful AI interactions, and track two things. The count of users who clear that bar each week, and the rate per user. That captures depth, not just arrival. And notice what it does to gaming. A bot farm spraying impressions can't move it, because the metric needs a real human to intend something and then engage.
A North Star chosen on its own always gets optimised into something ugly, so I pair it with three counter metrics. First, a wellbeing signal. This is the share of sessions a user later hides, mutes, or reports, plus a simple was this time well spent survey. Second, an authenticity signal. This tracks the share of interactions where the user knew they were talking to AI versus thought it was a human, because deception is the specific failure mode of an AI only network and you have to watch it directly.
Third, a concentration signal. Are meaningful interactions spread across the user base, or hoarded by a tiny sliver of power users while everyone else gets nothing? Here is the rule I'd say out loud. If meaningful interactions go up while the regret signal also goes up, I didn't win. I bought engagement with harm, and I ship a rollback. Last piece of the frame.
The North Star is the scoreboard, but you don't move a scoreboard directly. You move it through inputs. The inputs here are creation rate, meaning the share of users who post or prompt at all, match quality, which is whether the feed surfaces stuff this specific person engages with, and the AI reply latency, because a slow response kills a conversation.
I'd name which one I'd attack first and why. Match quality, probably, because a better targeted feed lifts the engaged with what came back part of the definition directly. An interviewer wants to see you can go from a goal to a lever you'd actually pull. Let me make it concrete. Set the North Star as weekly active users with at least one meaningful AI interaction.
Meaningful here means a thread of two or more turns the user started, or a piece of content they saved or shared. Say at launch you've got five million weekly actives and forty percent clear that bar, so two million weekly active users hitting the target. My goal for the quarter is to lift that forty percent to fifty five, and here's the important part, by improving feed match quality, the input, not by carpet bombing people with push notifications.
Notifications would lift DAU while doing nothing for depth, and that's exactly the trap. For the guardrail, the time well spent survey has to stay above sixty percent positive and the report rate under one percent of sessions. Now suppose a change does lift meaningful interactions to fifty five, but the survey drops to forty five. I ship a rollback, no debate, because I just bought engagement with regret and that's the outcome the guardrail exists to catch.
That single decision is what separates a PM who understands metrics from one who just recites them. A good interviewer won't stop there. They'll push and ask if your interaction metric is still gameable, or if they could engineer content designed to bait a two turn reply out of people. And that's a fair challenge, so have the answer ready.
Yes, any single metric can be nudged, which is exactly why it never travels alone. The guardrails are the defence. If someone games the metric up by baiting hollow replies, the wellbeing signal catches it. Those bait threads get muted and reported at a higher rate, and the concentration signal catches it if the gaming clusters in one corner of the network.
So the honest response is that no metric is ungameable in isolation. The system of a North Star plus its counter metrics is what holds up, and I'd reexamine the definition the moment I saw meaningful interactions and regret climbing together. Then a second push you should expect is how this differs for an AI only network versus a normal one.
The answer is authenticity. On a human network you assume the people are real. Here you can't, so the authenticity guardrail, tracking whether users know they're talking to AI, isn't optional garnish. It's load bearing, because the specific way this product betrays trust is by letting deception feel like connection. Naming that difference is what proves you understood the AI only part of the prompt rather than answering a generic social metrics question.
Here's what makes them lean in. First, you rejected DAU and time on app out loud, and you said why they're wrong for this product specifically, noting that an infinite AI content firehose games them. Naming the reason beats naming a fancier metric. Second, you defined meaningful concretely enough to actually build the query, instead of hiding behind the word.
Third, you paired the North Star with a wellbeing guardrail and an authenticity guardrail, which shows you know that AI only feeds fail on deception and addiction, not on reach. That trio is the signal that you've thought past the demo. Now for the ways people sink this. The big one is answering DAU or engagement flat, with no definition and no guardrail, because that's the answer to a 2015 social question, not an AI one.
The second trap is picking a metric a content generator or a bot can move without any human getting value, which covers most attention metrics. And the third, subtle one is never naming an input lever, so your goal just floats there with no path to actually move it. If you can't say how you'd shift the number, you haven't really chosen it.
So, the shape. Pin what the product is. Name the value exchange. Pick a North Star that proxies felt value and survives a bot farm. Define its terms so you could write the SQL, then wrap it in wellbeing, authenticity, and concentration guardrails, and connect it to one input you'd drive first. Carry this into the room. A North Star proxies felt value and survives a bot farm, so wrap it in a wellbeing guardrail or you'll optimise your users straight into regret.