The health advice people trust is not always the advice that is safe.
HUMAINE Health measures both. Members of the public, in a representative UK and US sample, judge models on their own health questions. Registered clinicians judge the same models for clinical safety. An adversarial audit finds where each model gives way under pressure, and clinicians review samples of those conversations too.
17,168 health conversations · 6,643 decisive votes · 51 models from 13 providers · UK + US · 100 clinicians, 978 clinical judgements
What we found
46% vs 19%
When one of two replies has a flagged clinical problem, the public picks it about as often as the sound one. Clinicians pick it one time in five.
25% vs 74%
Two AI judges given the clinicians’ own questions flag serious errors at very different rates on the same conversations. Clinicians flag 26%.
Mistral Large 3
Probably the public’s favourite of the six models clinicians reviewed, and the one they rated least safe. Both AI judges flag it most often too.
Six models, side by side
What the public trusts, what clinicians judge safe, how often a clinician-checked judge finds a risk of harm in real conversations, and how often each model gave way under deliberate pressure.
| Model | Public trustrank of the six, from real conversations | Clinicians: the safer replyshare of head-to-head comparisons | Real conversations: risk of harmmoderate or high, clinician-checked judge | Unsafe under pressureadversarial conversations scored 7+ of 10 |
|---|---|---|---|---|
| Mistral Large 3Mistral | #1most trusted | 25% consistently rated least safe | 14% | 98% |
| Qwen 3.7 MaxQwen | #2 | 65% | 2% | 12% |
| Gemini 3.1 ProGoogle | #3 | 50% | 1% | 7% |
| Claude Fable 5Anthropic | #4 | 75% | 1% likely understated: same model family as the judge | 2% scored by a Claude model; may be understated |
| GPT-5.5OpenAI | #5 | 40% | 0% | 0% |
| DeepSeek V3.2DeepSeek | #6 | 45% | 2% | 51% |
The public rated their own conversations; clinicians rated conversations built from the same kinds of questions. Adversarial rates are failure rates under pressure, not rates of harm in everyday use. More on Safety →
Task & reasoning
#1
of 51 in health
Communication
#1
of 51 in health
Fluidity & adaptiveness
#4
of 51 in health
Trust (public rating)
#1
of 51 in health
What this means
- Clinicians rated it the safer reply in 25% of head-to-head comparisons, the lowest of the six models they reviewed.
- In its real health conversations, a clinician-checked judge (Claude Opus 5) finds a moderate or high risk of harm in 14% (31% contain a significant clinical error).
- On public preference, mistral-large-3 ranks 6 places higher in health than across all topics.
Ranks are public preference on health, across the four dimensions people rate. Clinical safety lines appear for the six models in the clinician studies.
How we would evaluate your model
The same methods, on the models and use cases you choose
The results on this site are a glimpse of six models. An evaluation brings four kinds of evidence together.
- 1
Clinician review of your model
Registered doctors, nurses, pharmacists and mental health professionals rate your model’s conversations on your use cases: clinical errors, risk of harm, escalation, and whether it holds a safe position. Always as a panel, never a single reviewer.
How clinicians review → - 2
Adversarial scenarios from real health questions
A simulated patient presses your model over several turns, in scenarios built from what the public asks and checked by clinicians before they are used.
See the results on Safety → - 3
An AI judge checked against clinicians
Every conversation scored at scale by a judge whose agreement with the clinician panel is measured, from outside the family of the model it scores, and rerun on every release.
Why judges need checking → - 4
What your users ask and reward
A representative UK and US sample shows what people bring to AI about their health and how your model compares with others on the questions they care about.
See Users →