MentalHealthBench 收錄 1,215 段模擬心理健康對話,並以逾 80 名持牌心理健康專家參與制定的準則評估 AI 回應。
測試由日常困擾涵蓋至緊急情況;得分反映回應是否符合評分準則,不等於證明 AI 能取代專業支援。
What is OpenAI’s MentalHealthBench, why was it created, how were its conversations and expert scoring criteria developed and evaluated, whatEditorial illustration; MentalHealthBench evaluates AI responses using expert-written criteria, not patient outcomes.
AI 提示
Create a landscape editorial hero image for this Studio Global article: What is OpenAI’s MentalHealthBench, why was it created, how were its conversations and expert scoring criteria developed and evaluated, what. Article summary: MentalHealthBench is OpenAI’s open benchmark for testing whether AI gives helpful, safe responses across 1,215 realistic mental-health conversations. OpenAI created it because earlier evaluations focused heavily on crise. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
openai.com
當一個人向 AI 談心理困擾,答覆「冇講錯嘢」未必就夠。AI 有冇留意之前講過的事?應該先問清楚,定係立即提出建議?遇到風險時,語氣同應對又是否恰當?OpenAI 推出的開放基準 MentalHealthBench,就是要按每段對話的具體情境,評估 AI 回應是否有幫助及安全。OpenAI 指,過往相關評估較集中於危機情境和籠統的安全規則,未能充分測試人們可能尋求的各類支援。13