Building a clinical trial matching tool is one thing; being the patient it scores is another. Chen Wang, the developer behind ClinTrialFinder, has stage IVB metastatic nasopharyngeal carcinoma and recently analyzed 19 months of data from running his own case through his AI tool. The results were not a clean upward trend in match quality, but a chaotic scatter plot where the number of 'strong matches' ranged from 13 to 78 for the identical medical profile.
The Illusion of Objective Data
Wang expected to find a linear improvement in his toolβs accuracy over time. Instead, he found that the match count was heavily dependent on how he phrased his query rather than the state of his disease or the tool's underlying intelligence. When he provided a full, accurate medical history including six prior treatment lines, the tool returned 16 strong matches. When he used a sparse questionnaire, it returned 38. The higher number wasn't better news; it was the result of the AI failing to rule out ineligible trials due to lack of context.
Instrument Changes Masked as Progress
Part of the variance came from legitimate software updates, such as the spring 2026 release of 'basket matching,' which expanded the search to pan-cancer studies. This feature increased the pool of potential matches by casting a wider net, but Wang notes that much of this 'growth' was just a different instrument capturing off-disease weak matches. He also observed that specific regulatory events, like the June 23, 2026 approval of the drug BL-B01D1 in China, caused immediate jumps in rankings that reflected real-world changes rather than algorithmic noise.
Key Takeaways
- Match counts in AI-driven medical tools are often artifacts of query specificity, not objective measures of available options.
- Sparse data inputs can lead to false positives, where a tool lists trials a patient is ineligible for because it lacks exclusion criteria.
- Developers must distinguish between algorithmic improvements and changes in search scope, such as adding basket matching features.
The Bottom Line
If your AI tool returns a number that fluctuates wildly based on how a user types, you aren't measuring the worldβyou're measuring the input. For builders in health-tech, this is a critical warning against treating match counts as verdicts on patient options.