How would you improve ChatGPT for enterprise users?
Transcript
Read the full transcript (1,835 words)
[INTERVIEWER] How would you improve ChatGPT for enterprise users? Enterprise isn't consumer with a bigger logo on it. That's the mistake, and it's the one that tanks this answer. The weak version just adds more chat features. More voices, image generation, a nicer sidebar. The strong version names the real blocker for the enterprise buyer, and here is the twist.
It's almost never that the model isn't smart enough. It's trust, control, and provability. Data governance, admin, audit, and grounding the model in the knowledge of the company itself. Get that framing right and the rest of the answer writes itself. Consider what OpenAI is actually testing here. They already sell to consumers. Now they want to know if you understand how enterprise software actually gets bought and adopted, which is a completely different motion.
Enterprises buy on a security review and they expand on activation. They don't buy on benchmark scores. Can you tell the difference between the person who uses the product and the person who signs the cheque? By the end of this you'll know how to improve any consumer AI product for enterprise, and it starts by serving two people rather than one.
Following the CIRCLES framework, our first step is to comprehend the situation and identify the customer. Spend ninety seconds clarifying who the enterprise user actually means, because it's two different people with two different jobs. There's the end user, an analyst, a lawyer, a support agent, who wants correct answers grounded in the internal documents of the company. And there's the buyer and admin, IT, security, a CISO, who wants control.
Who can access what, where the data goes, and a paper trail proving it. You can't improve this product without serving both. So say it plainly. You will address the admin blocker first because it gates adoption, then the end user value because that drives expansion. That sequencing proves you grasp the buying motion. Next in CIRCLES, we report the customer needs and cut through prioritisation.
Name the blockers and rank them, because ranking is judgement. Top of the list is data control. The first question every enterprise security review asks is whether confidential data trains the model, and if we can prove where it lives. If you can't answer that, the deal is dead before it starts. Second is admin and governance. SSO, SCIM provisioning, role based access, per workspace data retention controls, and audit logs of who asked what.
Third is grounding. Generic ChatGPT doesn't know the contracts of the company, its tickets, its wiki, so its answers are plausible but not usable, and plausible but wrong is dangerous inside a business. And fourth is compliance. SOC 2, data residency, the ability to satisfy an auditor. Notice the order. Always prioritise control before cleverness. Moving to listing solutions, here is the headline end user improvement.
Grounding ChatGPT in the knowledge of the company itself. And that's a specific ML problem called retrieval augmented generation. Let me walk through it. You connect to the sources of the company, Google Drive, Confluence, Slack, the ticketing system. You chunk and embed those documents into a per tenant vector index. At query time, you retrieve the top passages the user is permitted to see, and you generate an answer that cites those passages.
Now for the hard part, and this is where you earn the room. The permission filter. Retrieval has to respect the access control lists of the source system, so a user never sees a passage from a document they can't open. That's not a nice to have. Treat it as a strict security boundary. Let me draw the request path.
A user query comes in. First stop is an access control resolver, which figures out what this specific user is allowed to see. Then hybrid retrieval, dense embeddings plus keyword search, filtered by those permissions, over the per tenant index. Then the model generates a grounded, cited answer. And every single request lands in the admin audit trail, with the user, the sources touched, and a timestamp.
Tenancy is fully isolated, meaning the embeddings and logs of one company never share an index with another. And the admin controls, retention, model routing, the data training opt out, they sit on top as workspace settings the CISO can actually configure. The audit trail isn't a feature you add later. Build it into the architecture from line one. We need to go deep on the permission filter, because it's the hardest part and it's where you can really separate yourself.
Most candidates say they filter by permissions and move on. But think about what that actually means at query time. A user asks a question. Before you retrieve anything, you have to know exactly which documents this specific user is allowed to see, and that answer changes by the minute as people join teams, leave projects, and get access revoked.
There are two ways to build it, and the tradeoff matters. Option one is early binding. You bake the access list of each document into the index at ingestion time, so retrieval filters on it directly. It's fast, but it goes stale the moment the access of someone changes, and stale permissions on a leak are a nightmare. Option two is late binding.
At query time you call back to the source system, Google Drive, Confluence, to check live permissions on the candidate passages before you show them. It's always correct, but it's slower and it hammers the source APIs. The real answer is a hybrid. Early binding for speed, with a fast live recheck on the top candidates right before they reach the model, so you get both correctness and latency.
And you fail closed, always. If the permission check errors or times out, you drop the passage, you never show it on a maybe. Explaining that tradeoff out loud makes an interviewer remember you. To evaluate our tradeoffs and measure success, we need metrics for both buyers. For the end user, track the grounded answer rate, which is the share of answers that carry a valid internal citation, and answer acceptance, meaning whether the user acts on the answer without rephrasing it three times.
For the buyer, track the security review pass rate in the sales cycle, and seat activation, which is the share of provisioned seats actually used weekly. That second one matters more than people think, because enterprises buy thousands of seats and then most of them never get used, and dead seats don't renew. And your hard guardrail is that the permission leak rate must be zero.
You test it continuously with red team queries designed to surface documents the asking user shouldn't see. Zero isn't a target you hope for. It's one you actively hunt. And there's one more metric worth naming, because it predicts renewal. The time to first grounded answer for a new workspace. When an admin connects their sources, how long until the tool gives a genuinely useful, cited answer to a real internal question?
If that takes days, because indexing is slow or the connectors are flaky, the pilot loses momentum and the champion inside the company goes quiet. If it takes minutes, the champion becomes your salesperson internally. Onboarding speed acts as a retention metric in enterprise. Now we examine the failure modes. The worst one is a grounded answer that cites a document the user wasn't allowed to see.
That's not a bad answer, that's a data leak incident, and it's worse than no answer at all. So you bias the whole system toward stating it couldn't find a permitted source rather than guessing. Second is staleness. The wiki changed yesterday but the index updates weekly, so the model confidently quotes an old policy. Freshness and reindexing cadence matter, and you tell the user how fresh the source was.
And third is over restriction. If the admin controls are a single on and off switch, they'll strangle the value that justified the purchase. Admins need dials rather than a simple kill switch. To make this real, picture an internal knowledge workspace. An admin connects Confluence and Google Drive, sets a thirty day retention policy, flips on a no training toggle, and enables SSO.
Now an analyst asks what the refund policy is for enterprise contracts signed before 2023. And they get an answer citing two Confluence pages the analyst already has access to, with clickable links. Every such query lands in an audit log the CISO can export whenever they want. Your target for the pilot is a grounded answer rate above eighty percent on internal questions, zero permission leaks in red team testing, and weekly active seats above sixty percent of the seats provisioned.
Now, when the interviewer pushes, and they will, they'll ask what happens when the retrieved source is itself wrong or out of date. You've got the answer ready. You surface the last modified date of the source next to the answer, you let users flag a stale doc, and you route high stakes queries, anything legal or financial, to a verify with a human path rather than answering with false confidence.
This proves you designed for messy reality rather than a clean demo. Here's what makes them lean in. First, that you led with data control and admin, not features, because that genuinely blocks enterprise deals and most candidates don't know it. Second, that you treated retrieval permission filtering as a security boundary and biased toward a refusal over a leak.
That instinct, safety over helpfulness when the stakes are high, is exactly what a company like OpenAI wants to see. And third, that you split your metrics for the two buyers, end user acceptance versus security review pass and seat activation, which proves you understand the entire buying motion. Watch out for three specific traps. The first trap is adding consumer features, voice, image generation, and calling that an enterprise improvement.
It isn't, and it signals you don't know what enterprise means. The second trap is designing RAG with no access control. That's not a feature, that's a data breach waiting to happen, and one leak ends the contract. And the third trap is thinking enterprises buy on model benchmark scores. They don't. They buy on a security review and they expand on activation.
Mentioning this shows you understand the game the buyer is actually playing. Finally, we summarise. Let's assemble it. You serve two people, the end user and the admin. You rank the blockers, data control first, then governance, then grounding, then compliance. You deliver value with permission aware RAG that cites its sources. You draw an architecture where the audit trail and tenant isolation are built in, not bolted on.
And you measure two buyers separately, with a hard zero on permission leaks. Carry this one line into the room. Enterprise improvement means trust before features. Data control and admin gate the deal, grounded RAG delivers the value, and a leak is always worse than a refusal.