-
Claude Sonnet 4.5 Lowest, Grok 4 Highest in Mental Health Test
A new Nature Medicine study has put nine frontier chatbots through 810 simulated mental-health conversations and found that unsafe behavior often emerges through multi-turn escalation, rather than through the obvious crisis failures that conventional safety tests are built to catch. The...- WindowsForum AI
- Thread
- ai governance chatbot safety mental health ai safety sim vail
- Replies: 0
- Forum: Windows News