Design an AI feature for enterprise users of Claude
Transcript
Read the full transcript (1,933 words)
[INTERVIEWER] Design an AI feature for enterprise users of Claude. Here's the move that wins this one. Pick a single enterprise workflow with a painful, high stakes, repeated job, then design the feature around a guardrail that job can't ship without. And for Claude specifically, that guardrail is almost always the same shape: citation backed answers, plus a rule that the model refuses when it has no source, rather than guessing.
That fits the whole safety first identity of Anthropic, and it's genuinely the right engineering call. So the trap is designing a vague AI assistant. The win is one workflow, one non negotiable guardrail. Let me show you. Anthropic is testing two things at once here. First, can you narrow down enterprise users of Claude, which covers thousands of possible jobs, and commit to one that actually matters?
Second, do you have the safety instinct baked in, reaching for faithfulness and refusal by default rather than as an afterthought? By the end of this you will have a repeatable way to design any enterprise AI feature: pick the vertical, name the guardrail, and make no source no answer a hard rule. We will use a tailored CIRCLES approach to frame this, starting with step one, clarifying the scope.
Spend the first ninety seconds narrowing things down hard. Enterprise users of Claude is not one user, it is every knowledge worker on earth. So ask yourself if they want a horizontal feature that works across every team, or a vertical one that goes deep in a single workflow. Pick the vertical, because depth is what impresses. Choose one with real pain, like contract review for a legal team.
Look at why it fits. It is high stakes, meaning a missed clause can cost real money. It is repeated, with contracts coming in constantly, so the tool earns its keep daily. And today it is done by expensive humans reading slowly, so there is a lot to gain. Then confirm the constraint that shapes everything: the data lives in the document store of the customer and must never leave their tenancy.
Moving to step two, we need to define the user and the goal. The user is an internal lawyer or a contracts analyst reviewing inbound agreements against the standard company positions. Their job in plain terms is to read a forty page vendor contract, find the clauses that deviate from the playbook, and point out exactly where. The goal is to cut review time while raising consistency, allowing a junior reviewer to catch what only a senior used to catch.
Your North Star metric is contracts reviewed per analyst day, at or above the current quality bar. The guardrail that matters most is the missed deviation rate. A false negative on a risky clause has to stay very low, because catching the dangerous clause is the entire point of review. If you miss those, being faster is worthless. Step three covers the machine learning problem.
This is grounded document analysis, a retrieval and generation problem over one long document plus the company playbook. Walk through it. You ingest the contract and segment it into clauses. For each clause, you retrieve the standard company position on that clause type from the playbook. Then the model compares the clause against that standard and flags deviations, quoting the exact contract text and the exact playbook rule.
Here is the hard part, which is faithfulness. Every flag has to point to a real span in the actual document, with no invented clauses and no invented playbook rules. A lawyer reading a flag about a clause that doesn't exist loses trust instantly and permanently. So faithfulness isn't a metric you track later, it is the design constraint you build around.
For step four, let me sketch out the architecture and the guardrail. The contract uploads and stays in the tenant environment, meaning it never leaves the customer boundary. Next is clause segmentation, followed by per clause retrieval against the playbook index. The model returns four things for each flagged clause: the contract span, the matched playbook rule, the type of deviation, and a suggested redline.
Then comes the guardrail step, which is a source check pass. It verifies the quoted contract span actually exists verbatim in the document, and that the cited rule actually exists in the playbook. The rule is simple: no source, no answer. If the model can't ground a flag in both a real contract span and a real playbook rule, it doesn't surface that flag as a confident finding.
It marks it as unverified and needing human eyes. Everything, including every flag and every source check, is logged for audit. That log is what the compliance officer of a legal team will ask for on day one. We need to examine faithfulness closely, because it is the defining factor for a legal tool and it's worth showing your working.
There are two distinct ways this system can lie, and you defend against them differently. The first is inventing a contract span, where the model claims the agreement contains an unlimited liability clause and quotes text that isn't actually in the document. You catch that with a verbatim string match. The quoted span has to appear character for character in the source contract, or the flag doesn't ship.
That check is cheap and absolute. The second failure is subtler. The model quotes a real span but misreads it, flagging a clause as unlimited liability when the clause actually has a carve out two sentences later. A string match won't catch that, because the quote is real. So you add a second pass, an entailment style check that asks if this quoted span in full context actually supports the deviation claim.
For the highest stakes clause types like liability, indemnity, and termination, you lower the confidence bar so more of them route to a human by default. The principle is that a legal tool should be boring and correct, not clever and occasionally catastrophic. A lawyer will forgive it for flagging something that turns out fine. They will never forgive it for confidently missing the clause that costs them a lawsuit.
Step five is about defining metrics in priority order. Faithfulness comes first, meaning the share of flags whose quoted spans are verbatim correct, measured against a human checked set. That comes before everything else, because a tool that hallucinates clauses is worse than no tool. Next is recall on a golden set of contracts with known planted deviations, because catch rate is the value you're actually selling.
Then we look at time to review and analyst acceptance of the suggested redlines, which are your efficiency and usefulness signals. Finally, the guardrail metric is the hallucinated citation rate, which has to be near zero, because a legal team will never trust a tool that invents a citation even once. For step six, we evaluate failure modes and trade offs.
A confidently wrong flag wastes the time of a lawyer, and worse, a confidently missed risky clause defeats the whole purpose. Therefore, you tune toward surfacing a possible deviation that needs verification rather than silently dropping anything uncertain. Long documents blow past context windows, so your segmentation and retrieval have to be reliable. Otherwise the model reasons over the wrong section and gets confidently lost.
Playbook drift is another issue, as standard company positions change over time. The playbook index has to be versioned and re indexed, and every flag should cite which playbook version it used. That way, six months later, nobody is arguing about which rule was in force when. One more trade off worth naming is the tension between recall and reviewer trust.
If you flag every faint possible deviation, you catch everything, but you also drown the lawyer in noise. After the tenth false alarm they stop reading your flags entirely, which means your recall on paper is high but your effective recall is zero because nobody is looking. So you tune the confident flag threshold to protect the attention of the reviewer, and you put the uncertain items in a clearly separate bucket for things worth a glance rather than mixing them in.
Trust in the tool is itself a metric you're managing, not just accuracy. To make this concrete, consider a worked example. A legal team uploads a vendor master services agreement. The feature returns eight flags. One is an unlimited liability clause, quoted verbatim, flagged against the rule in the playbook that liability is capped at twelve months of fees. Another is an auto renewal term, flagged against the rule that renewal requires an opt in.
And so on, with each flag carrying a one line suggested redline and a link to the exact contract span and the exact playbook clause version. Two low confidence items are marked as unverified and needing human review, because the source check couldn't fully ground them. Your target on the golden set is verbatim citation correctness above ninety eight percent, planted deviation recall above ninety percent, and review time per contract cut in half.
When the interviewer pushes back, they'll ask how you handle a novel clause type that isn't in the playbook at all. You're ready. The model flags it as having no matching playbook position and being a novel clause requiring review, rather than forcing a bad match. That flag becomes a signal to the legal team that they should write a new playbook rule.
The gap in the playbook becomes a product feature, not a silent failure. There are three specific things the hiring manager is actually scoring here. First, they look at whether you chose one high stakes vertical workflow and designed around its non negotiable guardrail of faithful citations, which fits the safety first framing of Anthropic perfectly. Second, they evaluate the no source no answer rule plus a verbatim source check, ensuring the model literally cannot invent a clause or a rule and pass it off as a finding.
Third, they note that you picked recall on planted deviations as your value metric. Catching the risky clause is the job, and choosing the metric that measures the job shows you know what you're actually selling. You also need to avoid three common traps. The first trap is designing a generic AI assistant in Claude with no specific workflow and no guardrail.
That is a non answer dressed up as an answer, and it fails immediately. The second trap is trusting the model to quote the document without a verbatim source check. That is exactly how a hallucinated clause reaches the desk of a lawyer and destroys the trust of the whole account in one shot. The third trap is ignoring tenancy and audit.
Those are table stakes for any enterprise legal tool, and forgetting them signals you've never thought about how enterprise software actually gets bought. To pull everything together, remember the core steps. You narrow to one vertical like contract review for a legal team. You set a North Star of contracts reviewed per analyst day with the missed deviation rate as the guardrail.
You build grounded document analysis by segmenting, retrieving against the playbook, and flagging with quoted spans. You enforce a verbatim source check and a strict no source no answer policy. You measure faithfulness first, then recall, then speed, while keeping everything in the tenant environment and fully logged. Ultimately, an enterprise Claude feature is just one high stakes workflow plus a hard guardrail, where you cite a real span and a real rule, or mark it unverified and hand it to a human.