Red Hat AI 3.5 is designed to turn AI from isolated pilots into a governed shared service: EvalHub is generally available for pre deployment safety evaluation, while GPU controls and observability target production op... The release also expands model validation information, multi tenant GPU controls, token metering...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Red Hat announce with the general availability of Red Hat AI 3.5 on September 11, 2026, and how does the release—including its inte. Article summary: Red Hat announced the general availability of Red Hat AI 3.5 as an enterprise AI platform update focused on making AI a governed, observable, multi-tenant production service across hybrid environments—not merely a collec. Topic tags: general, documentation, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fak
Red Hat’s AI 3.5 release is a production-operations update rather than a single new model or chatbot feature. Its focus is giving enterprise platform teams a way to evaluate AI systems before deployment, share costly GPU capacity across users, and monitor inference after it goes live. Red Hat announced the release on September 9, 2026; its September 11 roundup reiterated the emphasis on safety, operational control and performance transparency. 43
15
The clearest production-readiness addition is EvalHub, Red Hat’s model and agent evaluation toolkit. It is generally available for customer-provided or customized models, retrieval-augmented generation (RAG) configurations and AI agents. Teams can run safety-focused tests for risks including prompt injection and jailbreaks, then generate compliance-oriented certifications from the results. 13
That matters because a model that works in a prototype is not necessarily ready for an enterprise workflow. EvalHub is intended to put a repeatable evaluation step ahead of deployment, rather than leaving each application team to build and document its own testing process.
Red Hat also said its validated model catalog gained more than 20 models, including models from Google, NVIDIA and Alibaba Cloud. Catalog evaluations include Garak benchmark results and indicators for toxicity and potential personally identifiable information exposure, giving teams additional evidence when selecting a model. 3
4
AI pilots often run with dedicated infrastructure and a limited audience. Production deployments must handle competing workloads, service priorities and cost accountability. Red Hat AI 3.5 adds controls aimed at that operational problem, including fair-share scheduling, priority-aware serving, admission control and priority-based request routing for shared GPU environments. 4
8
The release also adds per-user token metering and dashboards for inference health, GPU utilization and model performance. Together, those capabilities are intended to help platform operators see who is consuming capacity, identify serving problems and support team-level cost attribution. 4
14
For organizations using NVIDIA-accelerated estates, the practical connection is a Kubernetes-oriented operational layer around model serving: evaluate workloads, manage access to shared GPU capacity and monitor the running service. Red Hat AI Inference 3.5 also provides generally available optimized inference images for NVIDIA CUDA, alongside AMD ROCm, Google TPU, Intel Gaudi and IBM Spyre accelerators. 35
Red Hat AI 3.5 also extends its tooling for retrieval and agent-based applications. AutoRAG adds multilingual capabilities for non-English and mixed-language document collections, including language detection, language-aware chunking and multilingual embeddings in its optimization search space. 1
Red Hat describes AutoRAG as a way to evaluate and tune RAG pipelines with capabilities such as contextual retrieval, conversational testing and visual debugging. However, AutoRAG itself is a Technology Preview, not a generally available feature. Teams should account for that status when deciding which components can underpin a production service. 13
The release additionally includes AI Hub templates intended as starting points for workflows such as code review, document processing and research. 43
The release broadens deployment choices for distributed inference with llm-d on managed Kubernetes. Red Hat documentation identifies Azure Kubernetes Service, CoreWeave Kubernetes Service and Amazon EKS as target platforms, with Kubernetes 1.33 or later and provisioned GPU nodes among the prerequisites. 37
The support status is not identical across those choices. Documentation specifically describes distributed inference with llm-d on Amazon EKS as a Technology Preview, without production SLA support and not recommended for production. The Red Hat AI Inference documentation highlights deployment guidance for Azure and CoreWeave Kubernetes Service, making those the relevant supported paths in this release. 39
40
The release brings four workstreams that are often separate in an AI pilot into one platform direction:
Red Hat positions these functions as the operational foundation for AI that spans hybrid environments and behaves more like shared enterprise infrastructure than a collection of standalone experiments. 43
The wider IBM and Red Hat strategy also includes Lightwell, a separate joint initiative focused on vulnerable third-party open-source dependencies. In September, LTM announced a collaboration with IBM and Red Hat around Lightwell for AI-driven vulnerability remediation. That initiative concerns software supply-chain security, not a Red Hat AI 3.5 feature, but it reflects the companies’ broader emphasis on validation and operational risk reduction. 24
25
Red Hat AI 3.5’s primary promise is operational: give platform teams tools to test AI before deployment, govern shared acceleration infrastructure and observe model serving across hybrid environments. EvalHub’s general availability is the centerpiece, while GPU scheduling, metering and observability address the day-two work that pilot projects typically avoid. The important caveat is that availability varies by feature—AutoRAG and EKS distributed inference are still Technology Previews. 13
39
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Red Hat AI 3.5 is designed to turn AI from isolated pilots into a governed shared service: EvalHub is generally available for pre deployment safety evaluation, while GPU controls and observability target production op...
Red Hat AI 3.5 is designed to turn AI from isolated pilots into a governed shared service: EvalHub is generally available for pre deployment safety evaluation, while GPU controls and observability target production op... The release also expands model validation information, multi tenant GPU controls, token metering, agent and RAG tooling, and inference options for several accelerator and Kubernetes environments.