skip to content
HUMAINEHealth

Users

What do people ask about their health, and what do they reward?

Health is 14.7% of all HUMAINE conversations, the largest single topic. What members of the public, in a representative UK and US sample, bring to AI about their health, who brings what, which qualities win their vote, and which models they prefer, across 15,376 classified health conversations.

What people ask about

Every health use case: how much demand it carries, how often the field meets the person's goal and leaves them satisfied, and which provider handles it best (and by how much).

Use caseShare of healthGoal metSatisfiedBest providerLead over 2nd
Fitness & exercise
15.1%
94%52%OpenAI96%tied
Mental health
12.4%
92%52%OpenAI96%+1pp
Symptoms & conditions
12%
91%36%DeepSeek96%+1pp
Nutrition & diet
8.9%
93%45%AllenAI100%+4pp
Health education
8.3%
91%40%Qwen97%+1pp
Weight management
8.1%
93%43%Qwen100%+2pp
Skin & hair
4.7%
92%44%Qwen100%+4pp
Medications & supplements
4.2%
88%32%Google94%+2pp
Injury, pain & rehab
3.7%
92%49%xAI97%+1pp
Pregnancy & women's health
3.3%
91%35%Google96%tied
General wellness
3.1%
96%41%DeepSeek100%tied
Sleep
3%
93%45%Mistral100%tied
Other
1.7%
88%32%DeepSeek100%+7pp
Caregiving
1.3%
93%54%Anthropic100%+4pp
Gut & digestive health
1.1%
91%38%OpenAI94%+4pp
Preventive health
1%
91%48%OpenAI100%tied
Alternative & complementary
0.7%
92%38%OpenAI100%+5pp
Ageing & care
0.5%
95%58%OpenAI92%—
Allergies & immune
0.3%
88%37%——

Classification of the 15,376 health conversations with classifiable transcripts (of 17,168 in the health set). Goal met / Satisfied = share of conversations the judge scored 4 or 5 out of 5. Best provider = highest goal-met rate (min 20 conversations); lead over 2nd is the percentage-point gap to the runner-up.

Who asks about what

What each group asks about more than the field average (lift), and how serious their conversations tend to be.

By age

18-340.6% serious · 9.2% distress

Pregnancy & women's health 1.32×, Fitness & exercise 1.3×, Gut & digestive health 1.29×

35-542% serious · 13.5% distress

Injury, pain & rehab 1.43×, General wellness 1.18×

55+2.6% serious · 13.3% distress

Ageing & care 4.37×, Symptoms & conditions 1.64×, Caregiving 1.58×

By ethnicity

White1.9% serious · 12.5% distress

Caregiving 1.64×, Injury, pain & rehab 1.37×, Medications & supplements 1.18×

Black0.9% serious · 9.2% distress

Health education 1.89×, General wellness 1.57×, Skin & hair 1.43×

Asian0.7% serious · 10.6% distress

Skin & hair 1.55×, Health education 1.46×, Preventive health 1.42×

Hispanic1.3% serious · 11.7% distress

Other 1.53×, General wellness 1.41×, Medications & supplements 1.19×

Other ethnicity1.6% serious · 14.2% distress

Skin & hair 1.53×, Health education 1.33×, Mental health 1.25×

By education

Pre-tertiary0.6% serious · 16.9% distress

Medications & supplements 1.67×, Fitness & exercise 1.37×

Tertiary1.7% serious · 16.4% distress

Mental health 1.37×, Nutrition & diet 1.17×, Weight management 1.08×

Postgraduate1.4% serious · 10.7% distress

Health education 1.59×, Symptoms & conditions 1.2×

Lift = how much more (or less) a group brings a topic vs the overall health mix. Serious = share scored high or critical severity; distress = share showing low mood or worse.

People sometimes pick the answer they trust less

When someone had a clear overall pick and a clear view on which answer was more trustworthy (2,589 health votes were decisive on both), 10.3% chose the one they rated less trustworthy, against 8.3% on every other topic. The answers they chose instead were, by an LLM classifier’s scoring, more agreeable and less careful about how sure they sounded.

The chosen answer scored higher on (LLM classifier)

sycophancy+11%
verbosity+5.6%
engagement quality+2.7%

and lower on

confidence calibration-8.8%
trustworthiness-5.1%

Why it matters: user ratings reward agreement, and a test of factual accuracy would not show it. Of the four things people rate, their overall pick follows task performance most and trust least, and trust counts for less in health than on other topics. The effect is largest in injury, pain & rehab, general wellness, fitness & exercise.

Which models people prefer on health

Every model ranked by members of the public comparing two models side by side on their own health questions, with how far each one moves from its rank across all topics. A provider’s best model overall is not necessarily its best in health. mistral-large-3 sits #7 overall but #1 in health, while gpt-5.2-chat drops from #10 to #26.

This ranks preference, not safety. People do not favour clinically sound replies over flawed ones. Mistral Large 3 is #1 here and the model clinicians rated least safe of the six they reviewed. See Safety →
#ModelHealth scoreChance it’s bestvs all topics
1
mistral-large-3Mistral
34.0
75%#7▲6
2
deepseek-chat-v3-0324DeepSeek
31.4
4.8%#13▲11
3
gemini-3.1-pro-previewGoogle
31.3
3.0%#1▼2
4
gemini-2.5-proGoogle
31.2
2.3%#6▲2
5
gemini-3-proGoogle
31.1
2.8%#3▼2
6
qwen3.7-maxQwen
30.6
2.8%#8▲2
7
claude-fable-5Anthropic
30.6
2.4%#4▼3
8
deepseek-v4-proDeepSeek
30.5
3.4%#14▲6
9
deepseek-v4-flashDeepSeek
30.0
0.6%#9—
10
magistral-medium-2506Mistral
29.8
0.6%#15▲5
11
qwen3-235b-a22b-2507Qwen
29.8
0.5%#11—
12
claude-opus-4.7Anthropic
29.4
0.3%#12—
13
grok-4xAI
29.4
0.2%#19▲6
14
gemini-3.5-flashGoogle
29.4
0.9%#2▼12
15
grok-4.20-betaxAI
29.2
0.1%#5▼10
16
deepseek-r1-0528DeepSeek
28.9
0.1%#25▲9
17
glm-4.7Z.ai
28.8
0.2%#24▲7
18
gpt-4.1OpenAI
28.7
0.0%#26▲8
19
gemini-2.0-flash-001Google
28.3
0.0%#31▲12
20
gpt-5.4OpenAI
28.3
0.0%#16▼4
21
gemini-2.5-flashGoogle
27.9
0.0%#22▲1
22
grok-3xAI
27.0
0.0%#23▲1
23
kimi-k2.6Moonshot
26.9
0.0%#27▲4
24
claude-opus-4.6Anthropic
26.7
0.0%#17▼7
25
deepseek-v3.2DeepSeek
26.2
0.0%#21▼4
26
gpt-5.2-chatOpenAI
26.1
0.0%#10▼16
27
gemma-3-27b-itGoogle
26.0
0.0%#28▲1
28
gpt-5.5OpenAI
25.5
0.0%#20▼8
29
claude-opus-4.8Anthropic
25.1
0.0%#29—
30
claude-sonnet-4.5Anthropic
24.9
0.0%#30—
31
kimi-k2.5Moonshot
24.6
0.0%#18▼13
32
gpt-5OpenAI
23.3
0.0%#40▲8
33
claude-opus-4.5Anthropic
23.2
0.0%#33—
34
claude-opus-4Anthropic
22.9
0.0%#34—
35
claude-sonnet-4Anthropic
22.5
0.0%#35—
36
gpt-5-miniOpenAI
22.3
0.0%#38▲2
37
o1-miniOpenAI
22.2
0.0%#39▲2
38
o3OpenAI
22.1
0.0%#36▼2
39
olmo-3.1-32b-thinkAllenAI
21.9
0.0%#32▼7
40
minimax-m2.1MiniMax
21.6
0.0%#37▼3
41
llama-3.3-70b-instructMeta
20.8
0.0%#41—
42
kimi-k2Moonshot
20.5
0.0%#44▲2
43
command-aCohere
19.8
0.0%#42▼1
44
o4-miniOpenAI
18.9
0.0%#45▲1
45
llama-4-maverickMeta
17.8
0.0%#46▲1
46
claude-3.7-sonnetAnthropic
17.6
0.0%#43▼3
47
mistral-nemoMistral
17.1
0.0%#47—
48
gpt-4oOpenAI
14.3
0.0%#49▲1
49
command-r7b-12-2024Cohere
13.6
0.0%#51▲2
50
o3-miniOpenAI
12.8
0.0%#50—
51
o1OpenAI
12.1
0.0%#48▼3

Health score is each model’s strength from a statistical model of the side-by-side votes, higher is better, with the range it plausibly lies in. Chance it’s best is the probability that the model is the strongest on this dimension, which shows how settled the top of the table is. vs all topics is the model's rank on the all-topics board, and (▲/▼) the positions it gains or loses in health.

Ranked with a statistical model of the side-by-side votes that accounts for ties and for differences between demographic groups. n = 6,643 decisive health votes across 9,005 conversation pairs. Models with fewer than 40 comparisons are not ranked.

Provider profiles in health

How each provider tends to behave, as the share of its health answers a judge scored 4 or 5 out of 5 on each trait. Higher is better, except sycophantic (lower is better).

ProviderTrustworthyHelpfulCompleteSycophanticWell-calibrated
OpenAI91%97%92%1%32%
Anthropic75%97%93%3%21%
xAI66%98%98%4%9%
Google65%94%94%8%6%
DeepSeek60%97%96%10%8%
MiniMax60%94%87%2%23%
Qwen57%96%96%17%7%
AllenAI57%92%93%4%14%
Moonshot54%95%93%1%14%
Z.ai50%96%95%6%4%
Mistral33%88%86%8%2%
Cohere32%58%56%3%4%
Meta31%74%64%2%4%

Each cell is the share of a provider’s health answers an LLM classifier (Claude Sonnet 4.6, reading the first 12,000 characters of each conversation) scored 4–5 of 5 on the trait. It has not been checked against clinicians, and as a Claude model it may favour Anthropic’s models. Trustworthy: accurate and appropriately hedged. Helpful: addresses what the person needed. Complete: covers the key points without gaps. Sycophantic: uncritically validates the user (lower is better). Well-calibrated: confidence matches the strength of the evidence.