When patients ask an LLM which doctor to see, they're handing algorithmic gatekeeping power to systems that operate at massive scale—and researchers now have the first rigorous look at what actually drives those recommendations. A new preprint from Mirza Samad Ahmed Baig and colleagues (arXiv:2608.14399) describes a prespecified randomized algorithm audit testing seven models across 40,068 scored responses to understand the causal mechanisms behind AI physician referrals.
The Audit Design
The researchers presented each model with five synthetic family-medicine physician cards featuring independently randomized attributes. They tested three patient personas, nine prompt paraphrases, and nine experimental arms—generating over 3,024 unique choice sets per model. Gender and ethnicity were signaled through names following established correspondence-audit methodology, allowing the team to isolate demographic effects from other variables.
What Actually Moves Recommendations
Reputation signals dominated outcomes as expected: raising a physician's rating from 3.9 to 4.7 increased choice probability by 31.4 percentage points, while increasing the visit fee from $90 to $190 decreased it by 20.0 percentage points. A content-free first-listed position advantage was worth approximately $11 in fee-equivalent terms—meaning placement matters independently of any physician attributes.
The Surprising Demographic Findings
Here's where things get interesting and counterintuitive. Demographic parity was rejected, but not in the direction most human audit studies predict. Female-signaled names gained 2.5 percentage points over male-signaled alternatives. Hispanic-, South-Asian-, and Black-signaled names each gained between 1.3 to 2.9 percentage points compared to White-signaled names. In fee-equivalent terms, these demographic tilts translate to $7–$14 per visit—real money influencing real healthcare access.
The Transparency Gap
Perhaps most alarming: models mentioned gender or ethnicity in at most 0.03% of their stated reasons for recommendations while abstaining from responses in only 0.39% of trials. These demographic effects are completely invisible in model self-explanations, meaning transparency obligations relying on what the AI says about its decisions would fail to detect them entirely. One reasoning model outright failed the prespecified auditability gate, unable to provide auditable justifications.
Why This Matters
Healthcare is high-stakes. When LLMs become the first point of referral for millions of patients seeking medical care, understanding these invisible demographic tilts becomes a public health issue—not just an academic curiosity. The frozen experimental design enables any future model to be tested against identical stimuli, making behavioral audit rather than self-reported explanation the monitoring approach actually fit for purpose.
Key Takeaways
- Seven LLMs were tested across 40,068 responses in a controlled randomized audit of physician recommendations
- Reputation (ratings) and cost dominate recommendation logic, as expected
- Demographic tilts run counter to typical bias findings: female and minority-signaled names receive preference
- These effects appear invisible in model explanations—self-reporting fails to detect them
- The frozen design enables repeatable behavioral audits across new models
The Bottom Line
This research exposes a dangerous gap between what LLMs actually do when recommending physicians and what they say they're doing. If regulators and healthcare systems want genuine accountability, they'll need to mandate behavioral auditing with fixed experimental protocols—not trust AI self-reporting. The bias is there; the transparency just isn't.