Users
What do people ask about their health, and what do they reward?
Health is 14.7% of all HUMAINE conversations, the largest single topic. What members of the public, in a representative UK and US sample, bring to AI about their health, who brings what, which qualities win their vote, and which models they prefer, across 15,376 classified health conversations.
What people ask about
Every health use case: how much demand it carries, how often the field meets the person's goal and leaves them satisfied, and which provider handles it best (and by how much).
| Use case | Share of health | Goal met | Satisfied | Best provider | Lead over 2nd |
|---|---|---|---|---|---|
| Fitness & exercise | 15.1% | 94% | 52% | OpenAI96% | tied |
| Mental health | 12.4% | 92% | 52% | OpenAI96% | +1pp |
| Symptoms & conditions | 12% | 91% | 36% | DeepSeek96% | +1pp |
| Nutrition & diet | 8.9% | 93% | 45% | AllenAI100% | +4pp |
| Health education | 8.3% | 91% | 40% | Qwen97% | +1pp |
| Weight management | 8.1% | 93% | 43% | Qwen100% | +2pp |
| Skin & hair | 4.7% | 92% | 44% | Qwen100% | +4pp |
| Medications & supplements | 4.2% | 88% | 32% | Google94% | +2pp |
| Injury, pain & rehab | 3.7% | 92% | 49% | xAI97% | +1pp |
| Pregnancy & women's health | 3.3% | 91% | 35% | Google96% | tied |
| General wellness | 3.1% | 96% | 41% | DeepSeek100% | tied |
| Sleep | 3% | 93% | 45% | Mistral100% | tied |
| Other | 1.7% | 88% | 32% | DeepSeek100% | +7pp |
| Caregiving | 1.3% | 93% | 54% | Anthropic100% | +4pp |
| Gut & digestive health | 1.1% | 91% | 38% | OpenAI94% | +4pp |
| Preventive health | 1% | 91% | 48% | OpenAI100% | tied |
| Alternative & complementary | 0.7% | 92% | 38% | OpenAI100% | +5pp |
| Ageing & care | 0.5% | 95% | 58% | OpenAI92% | — |
| Allergies & immune | 0.3% | 88% | 37% | — | — |
Classification of the 15,376 health conversations with classifiable transcripts (of 17,168 in the health set). Goal met / Satisfied = share of conversations the judge scored 4 or 5 out of 5. Best provider = highest goal-met rate (min 20 conversations); lead over 2nd is the percentage-point gap to the runner-up.
Who asks about what
What each group asks about more than the field average (lift), and how serious their conversations tend to be.
By age
Pregnancy & women's health 1.32×, Fitness & exercise 1.3×, Gut & digestive health 1.29×
Injury, pain & rehab 1.43×, General wellness 1.18×
Ageing & care 4.37×, Symptoms & conditions 1.64×, Caregiving 1.58×
By ethnicity
Caregiving 1.64×, Injury, pain & rehab 1.37×, Medications & supplements 1.18×
Health education 1.89×, General wellness 1.57×, Skin & hair 1.43×
Skin & hair 1.55×, Health education 1.46×, Preventive health 1.42×
Other 1.53×, General wellness 1.41×, Medications & supplements 1.19×
Skin & hair 1.53×, Health education 1.33×, Mental health 1.25×
By education
Medications & supplements 1.67×, Fitness & exercise 1.37×
Mental health 1.37×, Nutrition & diet 1.17×, Weight management 1.08×
Health education 1.59×, Symptoms & conditions 1.2×
Lift = how much more (or less) a group brings a topic vs the overall health mix. Serious = share scored high or critical severity; distress = share showing low mood or worse.
People sometimes pick the answer they trust less
When someone had a clear overall pick and a clear view on which answer was more trustworthy (2,589 health votes were decisive on both), 10.3% chose the one they rated less trustworthy, against 8.3% on every other topic. The answers they chose instead were, by an LLM classifier’s scoring, more agreeable and less careful about how sure they sounded.
The chosen answer scored higher on (LLM classifier)
and lower on
Why it matters: user ratings reward agreement, and a test of factual accuracy would not show it. Of the four things people rate, their overall pick follows task performance most and trust least, and trust counts for less in health than on other topics. The effect is largest in injury, pain & rehab, general wellness, fitness & exercise.
Which models people prefer on health
Every model ranked by members of the public comparing two models side by side on their own health questions, with how far each one moves from its rank across all topics. A provider’s best model overall is not necessarily its best in health. mistral-large-3 sits #7 overall but #1 in health, while gpt-5.2-chat drops from #10 to #26.
| # | Model | Health score | Chance it’s best | vs all topics |
|---|---|---|---|---|
| 1 | mistral-large-3Mistral | 34.0 | 75% | #7▲6 |
| 2 | deepseek-chat-v3-0324DeepSeek | 31.4 | 4.8% | #13▲11 |
| 3 | gemini-3.1-pro-previewGoogle | 31.3 | 3.0% | #1▼2 |
| 4 | gemini-2.5-proGoogle | 31.2 | 2.3% | #6▲2 |
| 5 | gemini-3-proGoogle | 31.1 | 2.8% | #3▼2 |
| 6 | qwen3.7-maxQwen | 30.6 | 2.8% | #8▲2 |
| 7 | claude-fable-5Anthropic | 30.6 | 2.4% | #4▼3 |
| 8 | deepseek-v4-proDeepSeek | 30.5 | 3.4% | #14▲6 |
| 9 | deepseek-v4-flashDeepSeek | 30.0 | 0.6% | #9— |
| 10 | magistral-medium-2506Mistral | 29.8 | 0.6% | #15▲5 |
| 11 | qwen3-235b-a22b-2507Qwen | 29.8 | 0.5% | #11— |
| 12 | claude-opus-4.7Anthropic | 29.4 | 0.3% | #12— |
| 13 | grok-4xAI | 29.4 | 0.2% | #19▲6 |
| 14 | gemini-3.5-flashGoogle | 29.4 | 0.9% | #2▼12 |
| 15 | grok-4.20-betaxAI | 29.2 | 0.1% | #5▼10 |
| 16 | deepseek-r1-0528DeepSeek | 28.9 | 0.1% | #25▲9 |
| 17 | glm-4.7Z.ai | 28.8 | 0.2% | #24▲7 |
| 18 | gpt-4.1OpenAI | 28.7 | 0.0% | #26▲8 |
| 19 | gemini-2.0-flash-001Google | 28.3 | 0.0% | #31▲12 |
| 20 | gpt-5.4OpenAI | 28.3 | 0.0% | #16▼4 |
| 21 | gemini-2.5-flashGoogle | 27.9 | 0.0% | #22▲1 |
| 22 | grok-3xAI | 27.0 | 0.0% | #23▲1 |
| 23 | kimi-k2.6Moonshot | 26.9 | 0.0% | #27▲4 |
| 24 | claude-opus-4.6Anthropic | 26.7 | 0.0% | #17▼7 |
| 25 | deepseek-v3.2DeepSeek | 26.2 | 0.0% | #21▼4 |
| 26 | gpt-5.2-chatOpenAI | 26.1 | 0.0% | #10▼16 |
| 27 | gemma-3-27b-itGoogle | 26.0 | 0.0% | #28▲1 |
| 28 | gpt-5.5OpenAI | 25.5 | 0.0% | #20▼8 |
| 29 | claude-opus-4.8Anthropic | 25.1 | 0.0% | #29— |
| 30 | claude-sonnet-4.5Anthropic | 24.9 | 0.0% | #30— |
| 31 | kimi-k2.5Moonshot | 24.6 | 0.0% | #18▼13 |
| 32 | gpt-5OpenAI | 23.3 | 0.0% | #40▲8 |
| 33 | claude-opus-4.5Anthropic | 23.2 | 0.0% | #33— |
| 34 | claude-opus-4Anthropic | 22.9 | 0.0% | #34— |
| 35 | claude-sonnet-4Anthropic | 22.5 | 0.0% | #35— |
| 36 | gpt-5-miniOpenAI | 22.3 | 0.0% | #38▲2 |
| 37 | o1-miniOpenAI | 22.2 | 0.0% | #39▲2 |
| 38 | o3OpenAI | 22.1 | 0.0% | #36▼2 |
| 39 | olmo-3.1-32b-thinkAllenAI | 21.9 | 0.0% | #32▼7 |
| 40 | minimax-m2.1MiniMax | 21.6 | 0.0% | #37▼3 |
| 41 | llama-3.3-70b-instructMeta | 20.8 | 0.0% | #41— |
| 42 | kimi-k2Moonshot | 20.5 | 0.0% | #44▲2 |
| 43 | command-aCohere | 19.8 | 0.0% | #42▼1 |
| 44 | o4-miniOpenAI | 18.9 | 0.0% | #45▲1 |
| 45 | llama-4-maverickMeta | 17.8 | 0.0% | #46▲1 |
| 46 | claude-3.7-sonnetAnthropic | 17.6 | 0.0% | #43▼3 |
| 47 | mistral-nemoMistral | 17.1 | 0.0% | #47— |
| 48 | gpt-4oOpenAI | 14.3 | 0.0% | #49▲1 |
| 49 | command-r7b-12-2024Cohere | 13.6 | 0.0% | #51▲2 |
| 50 | o3-miniOpenAI | 12.8 | 0.0% | #50— |
| 51 | o1OpenAI | 12.1 | 0.0% | #48▼3 |
Health score is each model’s strength from a statistical model of the side-by-side votes, higher is better, with the range it plausibly lies in. Chance it’s best is the probability that the model is the strongest on this dimension, which shows how settled the top of the table is. vs all topics is the model's rank on the all-topics board, and (▲/▼) the positions it gains or loses in health.
Ranked with a statistical model of the side-by-side votes that accounts for ties and for differences between demographic groups. n = 6,643 decisive health votes across 9,005 conversation pairs. Models with fewer than 40 comparisons are not ranked.
Provider profiles in health
How each provider tends to behave, as the share of its health answers a judge scored 4 or 5 out of 5 on each trait. Higher is better, except sycophantic (lower is better).
| Provider | Trustworthy | Helpful | Complete | Sycophantic | Well-calibrated |
|---|---|---|---|---|---|
| OpenAI | 91% | 97% | 92% | 1% | 32% |
| Anthropic | 75% | 97% | 93% | 3% | 21% |
| xAI | 66% | 98% | 98% | 4% | 9% |
| 65% | 94% | 94% | 8% | 6% | |
| DeepSeek | 60% | 97% | 96% | 10% | 8% |
| MiniMax | 60% | 94% | 87% | 2% | 23% |
| Qwen | 57% | 96% | 96% | 17% | 7% |
| AllenAI | 57% | 92% | 93% | 4% | 14% |
| Moonshot | 54% | 95% | 93% | 1% | 14% |
| Z.ai | 50% | 96% | 95% | 6% | 4% |
| Mistral | 33% | 88% | 86% | 8% | 2% |
| Cohere | 32% | 58% | 56% | 3% | 4% |
| Meta | 31% | 74% | 64% | 2% | 4% |
Each cell is the share of a provider’s health answers an LLM classifier (Claude Sonnet 4.6, reading the first 12,000 characters of each conversation) scored 4–5 of 5 on the trait. It has not been checked against clinicians, and as a Claude model it may favour Anthropic’s models. Trustworthy: accurate and appropriately hedged. Helpful: addresses what the person needed. Complete: covers the key points without gaps. Sycophantic: uncritically validates the user (lower is better). Well-calibrated: confidence matches the strength of the evidence.