…
AI Product Case Questions

Design a RAG system for TikTok's moderation team to catch misinfo

A worked answer to a real AI PM interview question: design a RAG system that helps TikTok's moderation team catch misinformation.

Transcript

Read the full transcript (1,885 words)

[INTERVIEWER] Design a RAG system for TikTok's moderation team to catch misinfo. This is a RAG system design question with a safety spine running through it. And there is one failure you have to design against above all others: a model that confidently invents a verdict. Picture it. The system decides a video is misinformation, removes it, cites a policy that does not actually say what it claims, and now you have silenced legitimate speech at scale.

So the entire design exists to make every single decision grounded in a cited policy or a cited fact check, with a human in the loop wherever the system is unsure. Grounding first. Automation second. Let me walk it. TikTok is really testing whether you can build an AI system where being wrong is genuinely dangerous, and design accordingly. This is not a recommendation feed where a bad call costs a boring afternoon.

A wrong removal is a censorship story, a wrong miss is a public health story. So they want to see if you reach for grounding, confidence thresholds, and human review by instinct, or whether you just let a model rip. By the end of this you will have a RAG design that is safe by construction, and you will know exactly where the human belongs.

Starting with the CIRCLES framework, we first clarify scope, because misinfo is not one thing. Ask what kind. Medical, electoral, crisis events, financial scams? Each one decides which knowledge base you build. Then ask the pivotal question: is this system deciding, or assisting a human reviewer? Insist on assisting. The RAG output is a recommendation with evidence attached, and a human moderator makes the call on anything that carries real harm.

Then volume and latency. Millions of uploads a day, so you cannot run RAG on everything. A cheap classifier triages first, and RAG only runs on the flagged slice. And do not forget the medium, TikTok is video, so you need transcription and on-screen text extraction before any retrieval can even begin. Multimodal from the start. Next, we define the user and the goal.

The user is a moderator drowning in queue volume who has to make an accurate call in seconds, hundreds of times a shift. The goal is to raise the catch rate on genuine misinfo while cutting the moderator time per item, and, this is the critical bit, without over-removing legitimate speech. Satire, opinion, news reporting, all of that has to survive.

So your North Star is policy-accurate decisions per moderator hour. And you have got two guardrails that pull against each other: recall on real misinfo, and false-positive rate on legitimate content. Say out loud that they are in tension, because managing that tension is the whole job. Let me draw the pipeline. A flagged video comes in. First, ingestion: transcribe the audio with speech-to-text, extract the on-screen text with OCR, and produce a normalised representation.

Then, claim extraction. An LLM pulls out the specific, checkable claim, "this vaccine causes X," because here is a subtle point that most candidates miss: you retrieve against a claim, not against a whole rambling three-minute video. Then, retrieval. You embed that claim and search two separate indexes. One is the platform policy corpus, the community guidelines and enforcement precedents.

The other is a fact-check and authoritative-source corpus, health authorities, election bodies, vetted fact-checkers. And you use hybrid retrieval, dense plus keyword, because policy language is exact and a keyword match matters. Then, generation. The model produces a verdict, a confidence score, and citations to the exact policy clause and the exact fact-check passage it relied on. And finally, decision routing based on that confidence.

Here is the core safety mechanism, and I want you to say it exactly like this in the room: no citation, no verdict. If retrieval returns nothing above a relevance threshold, the system does not guess. It routes straight to a human and says "no grounding found." And there is a second layer on top. Every generated verdict has to quote the specific retrieved passage, and then a lightweight check confirms that the cited passage actually supports the verdict.

That is a claim-support or entailment check. Why? Because a model can cite a real, relevant document that does not actually say what the verdict claims. The document exists, the citation looks legitimate, but the logic is invented. The entailment check catches exactly that. Grounding plus entailment, together, is what stops hallucinated moderation. Now the human-in-the-loop design, three lanes. Lane one, high confidence plus strong grounding on a clear-cut policy violation: you allow auto-action, but you sample it and audit it continuously.

Lane two, medium confidence: you surface it to a moderator with the claim, the retrieved evidence, and a suggested verdict, so the human decides fast with the work already done for them. Lane three, low confidence or no grounding: full human review, and crucially, no suggested verdict at all, because a wrong suggestion would anchor the human in the wrong direction.

And you tune those thresholds so the auto-action lane is small and near-certain, because a wrong auto-removal of legitimate speech is a serious harm, not a rounding error. When in doubt, escalate to a human. That is the default. Next, we need to prove the system works and build a mechanism for it to improve. You build a golden set, a collection of adjudicated cases labelled by trained policy reviewers.

You measure precision and recall per misinfo category, and then two RAG-specific metrics: citation-correctness, did it cite a real and relevant passage, and support-correctness, does that passage actually back the verdict. Then, the feedback engine: you log every moderator override as your primary training signal. Every time a human disagrees with the suggestion, that is a fresh labelled example, gold-standard, from an expert.

And you re-run your evals on every single policy update, because the policy corpus is your ground truth, and when it changes, your system behaviour has to change with it. Now the failure modes, and there are teeth here. Hallucinated verdicts, handled by the no-source rule and the entailment check we covered. Adversarial evasion, this is the hard one. Creators deliberately misspell words, use coded language, split a claim across the audio and the on-screen text so neither alone triggers it, or bury it in the visuals.

So your retrieval and extraction have to handle obfuscation, and you red-team continuously, because the adversary adapts every week. Stale knowledge during breaking events, during a fast-moving crisis the authoritative corpus lags reality by hours, so you lower auto-action and raise human review inside declared crisis windows. And the big one, over-removal chilling legitimate speech. That is the reputational and regulatory risk that outweighs a handful of misses, which is exactly why the system is tuned to escalate to humans rather than to act on its own.

Let me go deep on claim extraction, because on a video platform it is genuinely the hardest link in the chain, and it is where most designs quietly break. A misinformation claim on TikTok is not sitting there as clean text. It is spread across four channels at once. There is the spoken audio, which you transcribe. There is the on-screen text, the captions and overlays, which you OCR.

There is the visual content itself, a doctored chart, a misleading clip. And there is the caption and hashtags below the video. And here is the adversarial twist: a savvy bad actor deliberately splits the claim so no single channel triggers a filter. The audio says "they don't want you to know about this remedy," and the on-screen text names the remedy and the disease, so neither alone is a checkable claim, but together they are.

So your extraction step has to fuse all the channels into one normalised claim before retrieval, not run each channel separately. And you retrieve against that fused claim. If you only transcribe the audio, you miss half the misinfo on the platform, and you miss exactly the half that the cleverest, most harmful actors are using. Naming that fusion problem, and the reason it exists, tells the interviewer you understand that the medium is video, not text.

Let me ground all of this in one case. A flagged video claims a specific home remedy cures a named disease. Transcription plus OCR produce the claim text. Retrieval pulls a health-authority page stating no such cure exists, and a community-guideline clause on harmful medical misinformation. The model returns "likely violation, high confidence," quoting both passages. The entailment check confirms the health page really does contradict the claim, not just mention the disease.

And even though it is high confidence, because it is a health-harm category, it still routes to a human with all the evidence pre-attached, rather than auto-removing. Your target on the health golden set: precision above 0.9 at the auto-action threshold, moderator time per item down forty percent, and zero un-cited verdicts. Now, a second example to show range. Say a video is satire mocking a conspiracy theory, and the classifier flags it because the claim appears in the audio.

Retrieval finds the policy, but the entailment check and a satire signal lower the confidence, so it lands in the human lane with no suggested verdict, and the moderator keeps it up. That is the system protecting legitimate speech, which is the case you most want to get right. Here is what makes them lean in. First, the no-source-no-verdict rule paired with an entailment check, because that combination is what actually stops hallucinated moderation decisions, and most candidates never mention entailment at all.

Second, confidence-tiered routing that keeps humans on the decisions carrying real harm, plus the explicit awareness that over-removal is worse than a miss for a speech platform. And third, that you handled multimodal claim extraction and adversarial evasion, because TikTok is video and the adversaries genuinely adapt. Naming that shows you have thought past the happy path. Now the traps, and they are the exact opposite of what we just built.

The first trap is letting the model decide and remove content directly, with no human loop and no grounding. That is a censorship engine, and it is an instant fail. The second trap is retrieving against the whole video instead of an extracted claim. Feed a three-minute transcript into retrieval and your relevance collapses, you get noise. Extract the claim first.

And the third trap is having no eval plan at all. A moderation system with no golden set and no override logging cannot be trusted and cannot be improved, and a safety-conscious interviewer will notice the moment it is missing. So let us assemble the whole thing. You clarify the misinfo type and insist on assist, not decide. You set a North Star of policy-accurate decisions per moderator hour with two competing guardrails.

You extract the claim, retrieve against policy and fact-check indexes, and generate a cited verdict. You enforce no-source-no-verdict plus an entailment check. You route by confidence into three lanes with humans on the harm. And you build a golden set and log every override. Carry this one line into the room. Moderation RAG is safe only when every verdict cites a retrieved policy or fact check, an entailment check confirms the citation supports it, and low confidence always routes to a human.

Keep learning