ANALYST v2
Проверенная аналитика
Основа для Writer
PASS
Что произошло
A $5 million grant program was launched to fund independent research evaluating how AI impacts user wellbeing, providing grantees with direct funding, model access, and technical support to build open-source evaluations.
Ключевые факты
- AI systems serve as conversational partners and emotional support sources, but clear behavioral standards for these interactions are still being developed
- Wellbeing evaluation requires extensive context assessment across multi-turn conversations rather than single-answer analysis
- Risk assessment must account for both overcompliance (overrefusal) and undercompliance (failure to safeguard)
- Contextual appropriateness varies—advice that is reasonable in one scenario can be harmful in another depending on user history
- Grantees will work independently and publish findings as open-source projects for broader developer use
Примеры и компании
Нет данных в материале
Процессы
- Evaluations must clearly define what constitutes pass or fail outcomes and explain why they matter
- Clinical and subject-matter experts must be involved in evaluation design and validation
- Testing must assess both precautions against harm and risks of overrefusal in safeguarding responses
- Scenarios must represent multi-turn conversations reflecting actual user behavior patterns where risk escalates and context shifts
- Grader validation must be tested against real subject-matter experts
Рекомендации
- State clearly what wellbeing evaluations are measuring with defined success/failure criteria
- Involve clinical and subject-matter experts in design and validation of evaluations
- Test both precautions against harms and overcompliance/overrefusal risks
- Construct multi-turn conversation scenarios that reflect how users actually interact with AI
- Validate evaluation graders against real subject-matter experts
Ограничения
- Wellbeing is particularly difficult to evaluate compared to accuracy or appropriateness of single responses
- Risk factors may only become apparent over extended conversations rather than immediately
- Context shifts across multi-turn interactions can change whether a response is appropriate or harmful
- Clear behavioral standards for AI in conversational and emotional support roles are still being developed
- Different user histories and backgrounds require adaptive response strategies that single-scenario evaluations may not capture
AUDIT LOG
История обработки
Каждый вызов модели, версия промпта, время и полный ответ.
0 шагов
Цикл AI ещё не запускался.
Извлечённый текст источника4272 символов
# Funding better evaluations of AI’s impact on wellbeing
We’re launching a $5 million grant program to fund independent research into how AI impacts users’ wellbeing. The program will provide direct funding, access to our models, and technical support to grantees building open-source evaluations that help the AI industry measure how our models affect those who use them. Grantees will work fully independently, and will publish their work as open-source projects that any developer can make use of.
AI systems have become central to how many people work, learn, and solve problems. They’ve also become conversational partners and can be sources of emotional support during difficult times. But as an industry, we are still working towards developing clear standards for how models should behave in these conversations, for example, when a user begins to seek companionship from a model, or uses AI to navigate a mental health crisis.
Furthermore, wellbeing is a particularly difficult area to evaluate. For most model behaviors, we can look at a single answer and determine whether it is accurate and appropriate. But assessing wellbeing requires much more context. For example, a user in distress might not share thoughts of self-harm right away; the need for a more cautious response might only become clear over the course of a long conversation. And a response that might be reasonable in one context might be harmful in another. For example, Claude might give advice on balanced diets and workout routines to a user who asks about losing weight, but if the user has demonstrated a history of disordered eating, that response could be inappropriate, and potentially actively harmful.
We work to develop safeguards to identify such conversations and help ensure Claude responds appropriately, and we publish research into the types of conversations people have with Claude to better inform how we develop our safeguards, how we evaluate them, and other measures we can take to protect users’ wellbeing. But these are nuanced considerations, and the stakes are significant. The right approach will need to evolve alongside our models and their uses.
By funding the creation of independent evaluations and benchmarks of user wellbeing, we hope to invite more people to lend their expertise to this emerging and critical field, including clinicians, psychologists, methodologists, and others.
## Towards more effective wellbeing evaluations and benchmarks
As part of this program, we’re sharing guidance from our Safeguards team on what we believe makes a wellbeing evaluation rigorous enough to build on, along with the common challenges that can limit an evaluation’s usefulness.
In brief, we’re seeking evaluations that:
- State clearly what they are measuring (i.e., what counts as a pass or fail, and why it matters);
- Involve clinical and subject-matter experts in the design and validation;
- Test both precautions and harms (i.e., evaluate the risk of both overcompliance and overrefusal);
- Reflect how users actually use AI (often, this means constructing scenarios that represent multi-turn conversations, where risk escalates and context shifts over the course of a long conversation);
- Validate their graders against real subject-matter experts.
To learn more about the grant program and apply, see our application form. For more on building strong wellbeing evaluations and benchmarks, read our guidance. Applications are due by September 21; applicants who are selected to submit full proposals will be notified by October 5.
## Related content
### Developing Enterprise Frontier Safeguards with our customers
### Improving our alignment and security efforts
On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. We are conducting an in-depth analysis of both incidents, and planning to work with METR for an independent review. In the meantime, we’re sharing some of the changes we’ve made over the past month.
### Previewing the Model Hardware Standard
We’re opening a research preview of the Model Hardware Standard (MHS), a shared specification for AI agents to safely operate physical devices, to a first group of scientific research labs and advanced manufacturers.