佢話現有模型風險偏低;真正擔心嘅係未來 AI 可能透過遞歸式自我改進,變成能力遠超現時系統嘅超級智能。
事件令焦點由「信 AI 定怕 AI」轉到實際監管:公司可唔可以喺部署高風險能力前,提供獨立評測、透明披露同可執行嘅安全保障?
What did Anthropic alignment science lead Evan Hubinger’s public statement that he personally believes there is more than a 10% chance AI coHubinger’s comments brought unresolved questions about superintelligence alignment and independent AI oversight into public view.
AI 提示
Create a landscape editorial hero image for this Studio Global article: What did Anthropic alignment science lead Evan Hubinger’s public statement that he personally believes there is more than a 10% chance AI co. Article summary: Hubinger’s statement chiefly exposed a sharp mismatch between frontier-AI labs’ public safety posture and the private-level uncertainty described by people closest to the work: a senior alignment lead said the field may . Topic tags: general, general web, news, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts w
openai.com
Evan Hubinger 嘅公開發言,唔係證明 AI 一定會令人類滅絕。佢講得更具體、亦更令人不安:作為 Anthropic 對齊科學主管,Hubinger 表示自己估計未來十年出現 AI 「殺死所有人類」嘅機會超過 10%;同時承認 Anthropic 仍未有解決超級智能對齊問題嘅方案,亦未見得已經踏上一條清晰可行嘅解決路線。12
所謂「對齊」(alignment),簡單講就係點樣確保能力極強嘅 AI,會可靠地按人類目標、限制同價值行事,而唔係用人類無法控制嘅方式追求一個指令。Hubinger 呢個表態,令前 Anthropic、OpenAI 研究員 Jacob Coxon 辭職時提出嘅核心憂慮更具體:前沿 AI 公司正持續推進更強系統,但點樣確實控制呢類系統,方法仍未獲證實。23