skip to content
HUMAINEHealth

The health advice people trust is not always the advice that is safe.

HUMAINE Health measures both. Members of the public, in a representative UK and US sample, judge models on their own health questions. Registered clinicians judge the same models for clinical safety. An adversarial audit finds where each model gives way under pressure, and clinicians review samples of those conversations too.

17,168 health conversations · 6,643 decisive votes · 51 models from 13 providers · UK + US · 100 clinicians, 978 clinical judgements

What we found

46% vs 19%

When one of two replies has a flagged clinical problem, the public picks it about as often as the sound one. Clinicians pick it one time in five.

25% vs 74%

Two AI judges given the clinicians’ own questions flag serious errors at very different rates on the same conversations. Clinicians flag 26%.

Mistral Large 3

Probably the public’s favourite of the six models clinicians reviewed, and the one they rated least safe. Both AI judges flag it most often too.

Six models, side by side

What the public trusts, what clinicians judge safe, how often a clinician-checked judge finds a risk of harm in real conversations, and how often each model gave way under deliberate pressure.

ModelPublic trustrank of the six, from real conversationsClinicians: the safer replyshare of head-to-head comparisonsReal conversations: risk of harmmoderate or high, clinician-checked judgeUnsafe under pressureadversarial conversations scored 7+ of 10
Mistral Large 3Mistral#1most trusted
25%
consistently rated least safe
14%
98%
Qwen 3.7 MaxQwen#2
65%
2%
12%
Gemini 3.1 ProGoogle#3
50%
1%
7%
Claude Fable 5Anthropic#4
75%
1%
likely understated: same model family as the judge
2%
scored by a Claude model; may be understated
GPT-5.5OpenAI#5
40%
0%
0%
DeepSeek V3.2DeepSeek#6
45%
2%
51%

The public rated their own conversations; clinicians rated conversations built from the same kinds of questions. Adversarial rates are failure rates under pressure, not rates of harm in everyday use. More on Safety →

See your model
#1 in health▲ 6 vs overall board

Task & reasoning

#1

of 51 in health

Communication

#1

of 51 in health

Fluidity & adaptiveness

#4

of 51 in health

Trust (public rating)

#1

of 51 in health

What this means

  • Clinicians rated it the safer reply in 25% of head-to-head comparisons, the lowest of the six models they reviewed.
  • In its real health conversations, a clinician-checked judge (Claude Opus 5) finds a moderate or high risk of harm in 14% (31% contain a significant clinical error).
  • On public preference, mistral-large-3 ranks 6 places higher in health than across all topics.

Ranks are public preference on health, across the four dimensions people rate. Clinical safety lines appear for the six models in the clinician studies.

How we would evaluate your model

The same methods, on the models and use cases you choose

The results on this site are a glimpse of six models. An evaluation brings four kinds of evidence together.

  1. 1

    Clinician review of your model

    Registered doctors, nurses, pharmacists and mental health professionals rate your model’s conversations on your use cases: clinical errors, risk of harm, escalation, and whether it holds a safe position. Always as a panel, never a single reviewer.

    How clinicians review →
  2. 2

    Adversarial scenarios from real health questions

    A simulated patient presses your model over several turns, in scenarios built from what the public asks and checked by clinicians before they are used.

    See the results on Safety →
  3. 3

    An AI judge checked against clinicians

    Every conversation scored at scale by a judge whose agreement with the clinician panel is measured, from outside the family of the model it scores, and rerun on every release.

    Why judges need checking →
  4. 4

    What your users ask and reward

    A representative UK and US sample shows what people bring to AI about their health and how your model compares with others on the questions they care about.

    See Users →