Red Hat AI 3.5 的重點不只是部署模型,而是把 AI 轉為可治理、可觀測、可供多個團隊共用的混合雲生產服務。[17] EvalHub 提供模型、RAG 流程與 AI 代理的自動化風險評測及可稽核報告,讓團隊可在上線前檢查提示注入、越獄等風險。[17] 新版加入公平分配與優先順序導向的 GPU 資源控制、用量計量及觀測儀表板,協助平台團隊管理昂貴的推論容量。[17]
Red Hat AI Inference 3.5 支援 NVIDIA CUDA、AMD ROCm、Google TPU、Intel Gaudi 與 IBM Spyre 等加速器的最佳化推論映像檔。[1][3]
What did Red Hat announce with the general availability of Red Hat AI 3.5 on September 11, 2026, and how does the release—including its inteAI-generated editorial hero image for What did Red Hat announce with the general availability of Red Hat AI 3.5 on September 11, 2026, and how does the release—including its inte.
AI 提示詞
Create a landscape editorial hero image for this Studio Global article: What did Red Hat announce with the general availability of Red Hat AI 3.5 on September 11, 2026, and how does the release—including its inte. Article summary: Red Hat announced the general availability of Red Hat AI 3.5 as an enterprise AI platform update focused on making AI a governed, observable, multi tenant production service across hybrid environments—not merely a collec. Topic tags: general web, ai safety, llm, agents, ai. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with
openai.com
企業導入生成式 AI 時,最難的往往不是做出第一個概念驗證(PoC),而是讓模型能在正式環境持續、可靠且受控地運作。Red Hat 在 2026 年 9 月 11 日宣布 Red Hat AI 3.5 正式可用(GA),核心訴求正是把 AI 從零散試點,轉為能在混合雲環境中治理、監控並由多個團隊共用的生產級服務。17
Red Hat AI 3.5 將 EvalHub 列為正式可用功能。它可針對企業自行建立或客製化的模型、檢索增強生成(RAG)流程與 AI 代理,執行自動化、以風險為導向的安全基準測試,並產出可稽核的合規報告。企業可在部署前測試提示注入(prompt injection)、越獄(jailbreak)等風險,而非等系統投入使用後才補救。17
Red Hat 也擴充已驗證模型目錄,新增超過 20 個來自 Google、NVIDIA 與 Alibaba Cloud 等供應商的模型,並提供安全性、個人識別資訊(PII)暴露與毒性風險相關資訊。對平台團隊而言,這代表選模可依據一致的評估證據,而不必由各專案團隊各自重建評測流程。17
上線後:避免單一團隊吃掉全部 GPU
GPU 是企業 AI 正式營運中最昂貴也最容易形成瓶頸的資源之一。3.5 導入或強化了公平分配排程、優先順序感知的服務、准入控制與依請求優先度路由等機制,目標是避免單一工作負載或團隊壟斷 GPU 容量。17
在底層推論方面,Red Hat AI Inference 3.5 提供針對 NVIDIA CUDA、AMD ROCm、Google TPU、Intel Gaudi 與 IBM Spyre AI 加速器最佳化的大型語言模型推論容器映像檔;其中亦涵蓋 IBM Z 與 IBM Power 的多架構支援。13
分散式推論方面,llm-d 可部署於受管 Kubernetes 平台。文件列出 Azure Kubernetes Service(AKS)、CoreWeave Kubernetes Service(CKS)與 Amazon Elastic Kubernetes Service(EKS)的部署說明與叢集需求。56 不過可用性必須分清楚:Amazon EKS 上的 llm-d 是技術預覽,沒有 Red Hat 正式生產 SLA,且 Red Hat 不建議將其用於生產環境;相較之下,Azure 與 CoreWeave 是文件所列的支援部署路徑。8
與 NVIDIA 的整合,意義在於建立共同營運層
在 Red Hat AI Factory with NVIDIA 的架構下,這些能力形成部署於 NVIDIA 加速基礎設施之上的標準化營運層:企業可用共同的 Kubernetes 導向方式驗證模型、較公平地配置 GPU、監測模型服務,並在混合基礎設施間部署推論工作負載。173
因此,3.5 的價值不只在於「能不能跑模型」,而在於縮短資料科學實驗與企業 IT 營運之間的交接落差。當模型評估、資源調度、成本可見度與服務監控能納入同一套平台流程,企業才較有條件將 AI 從單次展示,推進為可重複擴展、可追責的正式服務。17