How do you handle a vendor who will not disclose which model they use?
Rentwise AI sells a tenant-screening score to property managers, built on a model it calls proprietary. Wren Okafor is VP of Product at Fernhollow Residential, a mid-size apartment operator. Marisol Prieto is a rideshare driver and freelance bookkeeper who applied for a unit and never learned why she was denied.
- Require adverse-action reasons with every score, not just a number.Why: without reasons, nobody, not your staff, not the applicant, can tell a fair denial from an unfair one.
- Build a real appeal path with a person who can override the score.Why: the people harmed most by a bad score are the ones the vendor never has to answer to directly.
- Track denial rate by applicant income type, not just an overall approval rate.Why: unequal harm hides completely inside a number that looks fine on average.
- Get the vendor to confirm in writing that no proxy for a protected class feeds the score.Why: income type and credit-file thinness can quietly stand in for exactly the categories the law protects.
- Keep a written record of every denial reason.Why: without a record, nobody can find the pattern until it has already hurt someone.
How to answer this, stage by stage
Nobody is scoring you on whether you can say "AI transparency." They're scoring whether you can name who actually pays when a vendor keeps its model a secret.
Let's learn
Every month, Fernhollow's leasing staff reviewed about 340 rental applications by hand: a credit check, income verification, rental history, roughly 25 minutes of work per file, with decisions that varied a little from one regional office to the next.
Then Fernhollow signed Rentwise AI. One score, generated in seconds, replaced the 25-minute review. Applicants below a set score were denied automatically, no further look. Review time dropped to about 4 minutes per file, mostly spent confirming the number matched what the tenant had reported. Wren's team liked the consistency. The old process had a real problem: two regional managers reading the same file could land on different answers. Rentwise's score never did that.
Rentwise would not say which model produced the score, or which factors mattered most, calling both proprietary. Wren's team accepted that. Every applicant got a number. Nobody at Fernhollow could tell you why a specific applicant scored low, and neither could the applicant.
Here's the turn: the denials themselves were never the real problem. Denying an applicant who genuinely can't afford the rent is the system working. The real problem showed up the day a denied applicant asked why, and nobody, not Wren's staff, not the applicant, had an answer, because the vendor's number came with no reasons attached.
At its worst: an applicant like Marisol Prieto, with a genuinely strong payment history but a nontraditional income file, gets denied on a number nobody can explain, with no path to correct the record and no way to know if the denial was even right.
What I would leave alone: the underlying speed of the score itself doesn't need to change. Four minutes instead of twenty-five is a real, honest gain, and slowing every single application down to build a case file would waste everyone's time, including the applicants who are clearly, uncontroversially, a good fit.
The lesson: a fast decision and a fair decision are not the same test, and a vendor who will only pass one of them is only halfway trustworthy, no matter how good the speed feels on day one.
Now here is the same thing as a story
The short version above is what you'd say defending this to Fernhollow's leadership. Read this one for how the harm actually reached someone.
For six years, Fernhollow's leasing office decided who got an apartment by a person reading a file: pay stubs, a credit report, a call to the last landlord. Then it was a score on a shared tablet, glowing green or red at the front desk of every regional office.
Marisol Prieto drove rideshare four nights a week and did freelance bookkeeping for two small businesses on the side. Her rent had never once been late in five years. Her income, on paper, looked like four different part-time jobs instead of one steady one, because that's what gig work looks like on paper.
Rentwise's score came back low. The denial letter said only that she "did not meet the property's screening criteria." Wren's regional staff couldn't tell her more, because Rentwise never told them more either. Marisol called twice. Both calls ended the same way: a form letter, restated.
Nine months later, Fernhollow ran a routine fair-housing self-audit, the kind most operators run every couple of years to stay ahead of a real complaint. The audit pulled every denial from the past year and sorted it by income type against credit tier.
At the same credit tier, gig workers and the self-employed were denied 34 percent of the time. Salaried W2 applicants were denied 11 percent. Not one of those 340 monthly files had a written reason attached, because Rentwise had never been asked to produce one, and nobody at Fernhollow had ever asked what would happen if they were wrong.
Wren reopened the Rentwise contract that week. The new terms required adverse-action reasons with every score and a named person at Fernhollow with authority to override a denial on appeal, with no change to Rentwise's underlying model at all.
Here's what I'd take back. Fernhollow's leasing office accepted a vendor's proprietary score with no reasons and no appeal, because the speed gain was real and the aggregate numbers looked fine. That was a reasonable trade to make with limited information. It stopped being reasonable the moment one group of applicants started absorbing three times the denial rate of another, with no record anywhere explaining a single one of those decisions.
I would go back and put the reasons and the appeal path in the contract from day one, not nine months and one audit later. And the part I'd tell myself: we didn't need Rentwise's model. We needed Rentwise to be answerable, and we never once asked for that.
GUARD, in one screenNot "is the vendor hiding something." GUARD is what tells you who actually carries the cost of that secrecy, and what to demand instead.
The recap, one line per letter: groups is naming the applicant, not just the leasing staff, unequal is the three-times denial gap by income type, ability to contest is the form letter with no reason attached, and reduce is the reasons-plus-appeal contract change that fixed it without touching the model.
Detect. Track denial rate by income type every month, not just the overall approval rate, and flag any group whose gap against a comparable credit tier crosses ten points. That's how the next version of this gets caught before an audit has to find it.
And if you want to be sure it really works, try it somewhere elseSame five letters, a résumé-screening tool for an hourly-staffing agency instead of a rental application. A different history the model reads as risk.
Fieldstone Staffing places workers into hourly warehouse and retail roles using a vendor's résumé-screening tool that auto-rejects candidates below a fit score, with no visible reasons. Isabel Duarte runs vendor evaluation there. Mapped onto GUARD: groups is staffing coordinators, who read a score and move a candidate along, against rejected job applicants, who get an automated "not selected" email. Unequal is candidates with employment gaps, caregivers and people with a prior conviction, whose gaps read as risk to the model even when the underlying reason has nothing to do with job readiness. Ability to contest is the same shape as Fernhollow's: an applicant has no way to see the score or push back on it. Reduce is the identical fix: require the vendor's reasons and add a person who can pull a borderline file for a real human look.
The old decision here isn't an unwritten record, it's a different reversal: Fieldstone's original contract let the vendor keep its scoring reasoning fully proprietary, "for competitive reasons," and Isabel's team accepted that because the agency's overall placement numbers looked healthy in aggregate. That made sense when nobody had broken the number apart. It stopped making sense once employment-gap candidates turned out to be rejected at a much higher rate than candidates with continuous work histories, at the same experience level.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "require reasons with every denial, and give a person the power to override it," and stop.
Cost: there's no budget to build a full appeal system on day one. Say so honestly, and start with a single named reviewer for the highest-volume denial reason, rather than skipping the fix entirely.
The model gets better, for real: if a vendor's update genuinely narrows the group gap, that's the monthly detection number doing its job, telling you the appeal volume can shrink, not that the monitoring can stop.
Where people run it wrong.
They treat "the vendor won't disclose its model" as the whole problem, when the real fix never required the model at all.
They check an overall approval rate and miss a group-level gap sitting quietly underneath it.
They build a review process for the buyer's convenience and forget the subject never gets to use it.
How to use it live. The moment someone says "the vendor won't tell us how it works," ask back: whose life gets harder when we can't answer for it? That's the group this whole framework exists to find.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't a human appeal path just adding the inconsistency back that the score was supposed to fix?" Response: no, because it only applies to a rejected applicant asking for review, not to the initial decision, so most of the consistency gain stays intact.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Evaluating AI vendors as a buyer
- #1 List the ten questions you would ask every AI vendor before a pilot.
- #2 How do you evaluate a vendor's quality claims without running your own eval?
- #3 Design the pilot you would run to evaluate two competing AI vendors.
- #4 What contractual terms matter specifically for AI vendors and not for other software?
- #5 How do you assess a vendor's model dependency and what happens if their provider changes terms?
- #6 Describe the data handling questions you would put to a vendor on behalf of your security team.