InterviewAdvancedResponsible AI & Advanced Practice / AI product case study teardowns / #19

Pick an AI product and argue it should not exist.

GUARD the product is BrackenVale Proctor, an AI exam-surveillance tool that flags "suspicious" test-taker behavior

BrackenVale watches a student's webcam during a remote exam and flags sessions with eye movement, background noise, or extra faces it reads as suspicious. Yaw Brennecke, dean of students at a mid-size university, champions it as a fair, consistent alternative to overworked human proctors. Solveig Achterberg runs disability services at the same school, and reviews every appeal a flagged student brings to her.

The direct answer
BrackenVale Proctor should not exist in its current form. It flags students with documented disabilities at roughly nine times the rate of everyone else, and the flagged student has no real way to contest a flag before it becomes an academic integrity case on their record. Replace it with exam designs that don't need behavioral surveillance at all, not a better version of the same camera.
Do this, in order
  1. Stop using automatic behavioral flagging as the trigger for an academic integrity case.Why: the flagged student has no lever to contest it before real consequences start.
  2. Redesign the exams themselves, open-book, randomized questions, applied problems, instead of surveilling the test-taker.Why: it solves the actual integrity problem without needing to watch anyone's face at all.
  3. If any flagging tool remains, require a human review before a case is even opened, not after.Why: right now the flag itself starts the case, and the human only shows up once the student is already presumed suspect.
  4. Audit flag rates by disability status, language background, and housing situation every term.Why: this is the one number that would have caught the problem before students had to file the complaints themselves.
  5. Keep a real, penalty-free opt-out for any student who doesn't want to be recorded.Why: an opt-out that quietly costs a student their exam slot isn't really an opt-out.
  6. Don't assume a "less biased" future version fixes this.Why: the deeper problem is a design that puts the burden of proof on the student, not the specific accuracy of this one release.

How to answer this, stage by stage

Eight moves for the broadest question in this set. Handle it plainly. Don't moralize, and don't soften a real harm with process language.

Stage 1
Pick one real product, not a category
Say it like this
"I'll argue against one specific product: BrackenVale Proctor, an AI exam-surveillance tool, not remote proctoring in general."
Why this works
A named product with real behavior is arguable. A category is just an opinion about the news.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD: groups, who's affected. Unequal, where the harm actually lands. Ability to contest, who has no lever. Reduce, the specific design change. Detect, how you'd know it's happening."
Why this works
Shows the interviewer a real method, not just an opinion dressed as an argument.
Stage 3
Name both people, plainly
Say it like this
"Yaw holds the switch, he decides the tool stays in place. The student being watched holds nothing. They can't inspect the flag, appeal it before it's filed, or opt out without losing their exam slot."
Why this works
Names the actual power imbalance instead of talking about "bias" as an abstraction.
Stage 4
State where the harm lands unevenly
Say it like this
"Students registered with disability services get flagged at 19%. Everyone else, 2%. It's not an equally distributed inconvenience, it's a tool that singles out exactly the students least able to absorb the cost of a false accusation."
Why this works
A real number, stated plainly, does more work than any adjective could.
Stage 5
Say who never gets to push back
Say it like this
"The flag itself opens the academic integrity case. The student's first real chance to explain anything comes after the case is already filed, on their record, whether or not it's later cleared."
Why this works
This is GUARD's strongest move: the subject who can't inspect, appeal, or opt out before the damage starts.
Stage 6
Give the direct answer
Say it like this
"BrackenVale Proctor shouldn't exist in this form. Redesign the exams so they don't need behavioral surveillance at all, that's a bigger project than patching the model, and it's the honest fix."
Why this works
Matches deliverable 0, a real position, not a call for "more oversight" that leaves the tool running.
Stage 7
Name the detection signal
Say it like this
"An honest audit of flag rates by disability status and language background would have caught this before a single student had to file a complaint to be heard."
Why this works
Shows you'd catch this in production, not just react to it after a lawsuit.
Stage 8
Close on what wouldn't fix it, and stop
Say it like this
"A less biased version of the same camera doesn't fix this. The design puts the burden of proof on the student by default, and that's the part that has to go."
Why this works
Refuses the easy answer of "improve the model" and names the actual structural problem.

Let's learn

BrackenVale watches a student through their webcam during a timed remote exam and flags the session if it reads eye movement, background sound, or an extra face as suspicious.

Before BrackenVale, the university used a mix of live human proctors and an honor-code pledge, catching a small number of genuine incidents a term, roughly a dozen, most resolved with a short conversation.

With BrackenVale, flagged sessions jumped into the hundreds each term. Most flags aren't cheating at all.

Knowledge spark: what's a false positive here? A flag on a student who wasn't actually cheating. In BrackenVale's case, a tic, a glance at a clock, a sibling walking through the room, or reading a question aloud in another language can all trigger one, with no way for the student to know which behavior caused it.
Flag rate by student group, one term
20% 10% 0 Baseline 19% Disability-registered 13% Non-English home 9% Shared housing
Every one of these groups already carries more risk in their daily life. BrackenVale adds an academic integrity case on top of it.

The turn: the flagged sessions aren't the real harm. The real harm is that a flag alone opens a case, and the student flagged has no way to explain themselves before that happens.

At its worst: a student with a documented anxiety disorder gets flagged twice in one semester for the exact tic their accommodation paperwork already describes, and spends more hours defending an exam they passed than the exam itself took to complete.

The decision that put the burden on the student We designed the flag to automatically open a formal academic integrity case, since it seemed efficient to skip a manual triage step for every one of hundreds of flags a term. That made sense when the tool was expected to flag only genuinely rare, serious cases. It stopped making sense once ordinary behavior, tics, background noise, another language, started generating hundreds of flags nobody had reviewed by a human first.

What I would leave alone: the university's underlying honor-code system, and human proctors for the exams that genuinely need live, in-person oversight. Neither of those carries this same automatic, unreviewed presumption of guilt.

The problem was never that some students got flagged. The problem is that a flag alone was enough to open a case, with no person and no chance to respond standing in between.

The lesson: a fairness problem in an AI product isn't fixed by a more accurate model. It's fixed by asking who the design left with no lever at all, and building one before the harm starts, not after.

Now here is the same thing as a story

Use the short version above when arguing this live. Read this one for what the appeal process actually looked like from Solveig's side of the desk.

Solveig has run disability services at this university for nine years, and can usually tell within the first sentence of a student's story whether an accommodation was properly filed and ignored.

Hand sketched comparison diagram titled Two people one lever. Left panel a person icon labeled Dean, caption holds the flag button. Right panel a person icon labeled Student, caption no way to contest it.
One of these two people can turn the tool off. The other can only wait for a case to close.

The rollout looked reasonable on paper. BrackenVale promised consistency, no tired human proctor missing a real incident at 2am, and a lower cost than staffing live exams around the clock.

Hand sketched flow diagram titled Where the appeal should sit and doesn't. Five boxes: exam taken, AI flags session, appeal step question mark highlighted, integrity case filed, grade held.
The appeal step, third in line, is the one that doesn't actually exist before a case gets filed.

The pattern built quietly across a semester. A student with a documented eye-tracking condition got flagged twice. A student subvocalizing a translation of an exam question in her first language got flagged once and nearly withdrew from the course out of embarrassment before Solveig's office intervened.

Hand sketched decision tree titled What happens after a flag. Root: BrackenVale flags a session. Four branches: student has an advisor leads to slow costly sometimes cleared, student has no advocate leads to case proceeds grade held, student has a disability tic leads to flagged again next exam, student opts out leads to no real alternative offered.
Every branch on this tree runs through the student. None of them run through a human review before the case opens.

The trigger wasn't one dramatic case. It was a routine end-of-term meeting where Solveig laid out, side by side, the flag rate for her registered students against the university-wide average, a comparison nobody had run before because nobody had asked.

Hand sketched quadrant titled Sorting students by risk of a false flag. Axes control over how they look on camera from little to a lot, and chance of a false flag from low to high. Quiet private room sits lower right, low risk high control. Registered disability sits upper left, high risk low control. Shared housing sits middle left. Non English home language sits middle.
The students with the least control over how they appear on camera are the ones carrying the highest risk of a false flag.

She pulled a term's worth of appeal outcomes. Nearly every flagged case involving a registered disability was eventually cleared, but only after weeks of paperwork, a formal hearing, and real academic stress for a student who had done nothing wrong.

Hand sketched labeled parts diagram titled What a fair integrity process needs. Center scale icon labeled Process, four callouts: a named reviewer, a chance to respond first, no grade hold by default, a real opt out path.
BrackenVale's rollout shipped with none of these four. All four existed in the old, slower, human-only process it replaced.

The old process asked the tool to flag first and let a case run on autopilot from there. Any redesign worth building asks a human to look before a student's academic record is touched at all.

The rollout team approved automatic case-opening because it looked like an efficiency win on a slide, one less manual step for hundreds of flags a term. It took watching students with documented, on-file conditions get treated as suspects by their own paperwork to see that the step being skipped was the only one that mattered.

GUARD, in one screenNot a review board. GUARD is what forces you to name who never gets a lever at all.

Hand sketched icon list titled GUARD in one screen. Five items: a person icon labeled Groups name operator and subject, a gauge icon labeled Unequal where the harm lands, a question mark box icon labeled Ability to contest who has no lever, a box icon labeled Reduce the actual design change, a document icon labeled Detect how you'd know in production.
Five letters, and ability to contest is the one a policy-document answer always skips.
G
Groups.
Yaw Brennecke, the dean who decides whether the tool stays. Every remote-exam student, who has no say in whether they're watched.
Names the operator and the subject by role, not as an abstract "users."
U
Unequal.
Disability-registered students flagged at 19%, non-English-home students at 13%, versus a 2% baseline.
A real, checkable disparity, not a vague concern about "bias."
A
Ability to contest.
A flag automatically opens a formal case before any human reviews it, so the student's first real chance to respond comes after the harm has already started.
GUARD's strongest, hardest move: naming who has no lever at all before the damage begins.
R
Reduce.
Redesign exams to not need behavioral surveillance at all, rather than trying to patch the flagging model's accuracy.
A specific product decision, not a training session or a review committee.
D
Detect.
A per-term audit of flag rates by disability status, language, and housing would have surfaced this before students had to file complaints to be heard.
Catches the harm in production, before an external complaint or lawsuit forces the issue.
Withdrawal rate among falsely flagged students, weeks into the term
24% 12% 0 Wk2 Wk7 Wk12
These are students who did nothing wrong and were eventually cleared. A fifth of them still left the course before that clearance came.

The recap, one line per letter: groups is the dean who holds the switch versus the student who doesn't, unequal is the nine-times flag-rate gap by disability status, ability to contest is a case opening automatically with no human review first, reduce is redesigning the exam instead of the camera, and detect is a per-term audit by group nobody had run.

And if you want to be sure it really works, try it somewhere elseSame five letters, gig delivery driving instead of a university.

FerrousLane Driver Score is an AI system that scores delivery drivers on "safe driving behavior" from an in-cab camera, used to allocate the best-paying routes. Tobias Umberfield drives for a regional delivery company using it.

Mapped onto GUARD: groups is the dispatch manager who sets the scoring thresholds, and the driver, who can't see or contest an individual score before losing access to better routes. Unequal: drivers on older vehicles with worse camera angles and drivers who talk to themselves or sing while driving, a normal habit for many, get scored as "distracted" far more often than drivers in newer vehicles. Ability to contest: a low score silently reduces route access with no explanation given and no appeal offered before the next week's assignments go out. Reduce: base route allocation on completed-delivery safety records instead of an unreviewed in-cab behavior score. Detect: a quarterly audit comparing score distributions by vehicle age and route type would show the same kind of gap found in BrackenVale's data.

Hand sketched labeled parts diagram titled What a fair integrity process needs, reused for the delivery driver example. Center scale icon labeled Process, four callouts: a named reviewer, a chance to respond first, no grade hold by default, a real opt out path.
Swap "grade hold" for "route access" and the same four missing pieces show up in a driver-scoring system too.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "it flags disabled students at nine times the baseline rate with no review before a case opens, that alone is enough to pull it," and stop.
Cost: there's no budget this term to redesign every exam format. Say so honestly, and start with the cheapest real fix: require a human review before any flag opens a formal case, even without touching the exam formats yet.
The model gets better, for real: if BrackenVale's next version genuinely reduces the disparity to 2x instead of 9x, that's still not a reason to declare it fixed, since the design still puts the burden of proof on the student by default.

Where people run it wrong.
They treat this as a bias problem to fine-tune away, instead of a structural problem in who gets to contest a decision.
They propose a review board or a training session, which sounds responsible but changes nothing about the actual flag-to-case pipeline.
They wait for an external complaint or a lawsuit to reveal the disparity, instead of auditing for it proactively.

How to use it live. When someone asks you to argue an AI product shouldn't exist, ask yourself: who does this product act on that has no way to inspect, appeal, or opt out of that action. If you can name that person and the number is real, you have your argument.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "pick an AI product and argue it should not exist"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It's a risk and fairness question about who bears the harm.
2 · THE PEOPLE
Who are the two people named in this answer?
Tap to flip
ANSWER
Yaw Brennecke, the dean who holds the decision to keep BrackenVale running, and Solveig Achterberg, who runs disability services and handles the appeals.
3 · THE UNEQUAL HARM
Which groups get flagged at higher rates, and by how much?
Tap to flip
ANSWER
Disability-registered students at 19%, non-English-home students at 13%, versus a 2% baseline for everyone else.
4 · ABILITY TO CONTEST
What's the actual structural problem, beyond the flag rate itself?
Tap to flip
ANSWER
A flag alone automatically opens a formal academic integrity case before any human reviews it, so the student can't respond before the harm starts.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting a flag automatically open a case with no human triage step, which seemed efficient when flags were expected to be rare.
6 · THE NUMBER
Fill in the blank: falsely flagged students' course withdrawal rate climbed to ___% by week 12.
Tap to flip
ANSWER
21%. Among students who did nothing wrong and were eventually cleared.
7 · THE REDUCE STEP
What's the actual recommended fix, not just a patch?
Tap to flip
ANSWER
Redesign the exams so they don't need behavioral surveillance at all, rather than trying to make the flagging model less biased.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and who's the equivalent powerless subject?
Tap to flip
ANSWER
FerrousLane Driver Score, a gig-delivery driver scoring system. The delivery driver, who can't see or contest a low score before losing route access.

Check yourself Score: 0 / 0

Short answer, name the harm
1. Why does this answer argue BrackenVale should not exist, rather than just be improved?
Show hint
Look at the direct answer and the "ability to contest" step.
Show answer
Model answer: Because the structural problem, a flag alone opening a case with no human review first, isn't fixed by a more accurate model. The design itself puts the burden of proof on the student.
Multiple choice
2. Why does this answer treat the 19% vs 2% flag-rate gap as the strongest piece of evidence?
  • A. It proves the camera hardware is defective.
  • B. It's a real, checkable number showing the harm falls disproportionately on students with the least power to contest it.
  • C. It shows disability-registered students cheat more often.
  • D. It proves the university broke a specific law.
Show hint
Look at the "Unequal" step.
Show answer
B. GUARD's move is naming the disparity in checkable numbers, not moralizing about "bias" in the abstract.
True or false
3. True or false: this answer recommends keeping BrackenVale but adding a review board to oversee it.
  • True
  • False
Show hint
Look at "where people run it wrong."
Show answer
False. A review board doesn't change the flag-to-case pipeline. The answer calls for redesigning the exams instead.
Fill in the blank
4. Fill in the blank: students with a non-English home language were flagged at ___%, versus a 2% baseline.
Show hint
Look at the bar chart of flag rates by group.
Show answer
13%. Roughly six and a half times the baseline rate.
Short answer, apply it yourself
5. Pick an AI product you know of that makes a decision about a person. Who is that decision made about, and do they have any real way to contest it before it takes effect?
Show hint
Think about a hiring screen, a credit decision, or a content moderation flag.
Show answer
Model answer: Many people point to automated resume screening, where a rejected applicant often has no way to know which part of their resume the system flagged, let alone contest it.
Short answer, where it wouldn't matter
6. Name a use of behavioral monitoring where this same power imbalance genuinely wouldn't apply.
Show hint
Think about a setting where the person being monitored requested it and can turn it off freely.
Show answer
Model answer: A student voluntarily using a focus-tracking app on their own device, purely for their own feedback, with no case or consequence attached to the data at all.
Before you close the answer
Why this works
Tests whether you can build a real, evidence-backed argument against a product instead of a vague ethical objection, and whether you can name a structural fix instead of settling for "add more oversight."
Follow-up traps
"Isn't some proctoring necessary to stop real cheating?" Response: yes, and the answer doesn't argue against academic integrity, it argues that surveillance-based exams are the wrong tool, since exam redesign solves the integrity problem without this disparity.

"Couldn't you just retrain the model to reduce the disability flag rate specifically?" Response: that treats the symptom, not the structure. Even a perfectly calibrated model still opens a case automatically with no human review, which is the actual harm.
If pressed
BrackenVale's flagging model was trained primarily on footage from a single pilot campus with high broadband access and mostly private testing rooms, meaning the disparities seen elsewhere were partly baked in at the training-data stage, before the model ever reached a more varied student population.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more