# Anthropic funds open evaluations of AI's effects on wellbeing

> A grant program will support independent, open-source evaluations built around multi-turn behavior and expert-validated graders.

Canonical URL: https://www.devobs.io/news/news-anthropic-funds-ai-wellbeing-evaluations/
By: Nina Patel
Published: 2026-09-06T11:58:54.624Z
Updated: 2026-09-06T11:58:54.624Z
Event date: 2026-08-25
Section: AI

Anthropic announced a grant program for independent research into AI and user wellbeing on August 25. The [program notice](https://www.anthropic.com/news/wellbeing-research-grants) says selected researchers will receive funding, model access, and technical support to build open-source evaluations. Anthropic identifies clinicians, psychologists, and evaluation methodologists among the experts it wants involved.

## The target is behavior across a conversation

The announcement argues that wellbeing cannot usually be judged from one response. Risk may become visible only after context accumulates, and a response that is reasonable in one conversation may be harmful in another. Anthropic therefore asks proposals to define pass and fail conditions, test both harmful compliance and excessive refusal, use realistic multi-turn scenarios, and validate automated graders against subject-matter experts.

That design guidance is useful beyond the grant. Teams shipping conversational systems often overfit safety tests to a single provocative prompt. A stronger suite should include benign beginnings, gradual escalation, ambiguous disclosures, user corrections, and recovery after the assistant makes a mistake. It should also report disagreement among experts instead of forcing every case into an artificial binary.

## Open artifacts can improve reproducibility

Anthropic says grantees will work independently and publish their projects as open source. The independence claim describes the planned program structure; the resulting methods will still need scrutiny for sampling, cultural coverage, grader calibration, and conflicts of interest.

An engineering team need not wait for a standard benchmark. Start by versioning a small set of longitudinal scenarios, blind the graders to model identity, and record both false reassurance and unnecessary escalation. When the grant outputs appear, compare their constructs and annotation rules with that local suite before adopting scores as release gates.

## Source references

- <https://www.anthropic.com/news/wellbeing-research-grants>
