Sections

Search

ArXiv paper measures how much benchmark contamination actually inflates scoresGoogle's Gemini 3.8 TTS models are cheap and handle multi-voice dialogue, per Simon WillisonUniDataAgent: China Unicom's Ontology-Grounded Enterprise Q&A Agent Cuts Report Time from Days to MinutesOpenAI says Harvey uses GPT-6 Astra to produce more structured legal draftsArXiv paper proposes auditable LLM labeling for classroom talk
All stories

Models·

OpenAI releases MentalHealthBench, an expert-informed benchmark for AI responses in mental health conversations

OpenAI introduced MentalHealthBench, an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations, per OpenAI News.

What happened

- OpenAI News introduced MentalHealthBench, described as an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations (OpenAI News). The announcement states the benchmark's purpose but not its scores, methodology details, or which models were tested, so no performance conclusions can be drawn from this evidence. It matters because it gives the field a named evaluation target for a sensitive domain; whether the benchmark holds up under outside scrutiny is not established by the supplied evidence. [1]

Sources

  1. OpenAI News · Primary source ·
    Introducing MentalHealthBench