What is OpenAI’s MentalHealthBench, why was it created, how were its conversations and expert scoring criteria developed and evaluated, whatEditorial illustration; MentalHealthBench evaluates AI responses using expert-written criteria, not patient outcomes.
AI 提示詞
Create a landscape editorial hero image for this Studio Global article: What is OpenAI’s MentalHealthBench, why was it created, how were its conversations and expert scoring criteria developed and evaluated, what. Article summary: MentalHealthBench is OpenAI’s open benchmark for testing whether AI gives helpful, safe responses across 1,215 realistic mental-health conversations. OpenAI created it because earlier evaluations focused heavily on crise. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
openai.com
當有人向 AI 談起心理困擾,一句「安全」的回答未必就足夠。它有沒有留意對話背景、問出必要的問題、尊重當事人的選擇,並按風險程度作出適切回應?OpenAI 推出的開放基準 MentalHealthBench,正是要測試 AI 在具體情境中的表現,而不只看它有沒有避開明顯不當的說法。OpenAI 指出,過往相關評估較集中於危機情況及籠統的安全規則,較難呈現人們尋求心理健康支援時可能遇到的各種情況。13