Red Hat logo.

Red Hat said on Wednesday it had unveiled a new version of its AI platform, Red Hat AI 3.5.

The company said Red Hat AI 3.5 focuses on pre-deployment security validation, multitenancy for shared GPU infrastructure and observability. Agent development features were also strengthened. It added AutoRAG and ready-to-use agent templates.

Red Hat plans to use Red Hat AI 3.5 to help companies run AI workloads at the same level as other core infrastructure, with verifiable security, cost allocation and granular resource management.

To that end, Red Hat formally launched EvalHub. EvalHub checks models, RAG configurations and agents in advance to identify risks and link them to auditable compliance reports. Users can validate models before deployment and generate risk-focused safety benchmarks and regulatory compliance certificates.

The model catalog added more than 20 validated models, including Google Gemma 4, Nvidia Nemotron 3 and Alibaba Cloud Qwen. In addition to performance benchmarks, the models are given Garak security scores and scores for privacy exposure and harmfulness risks.

Red Hat also highlighted shared GPU infrastructure capabilities. Fair-share scheduling divides capacity by tenant, while priority-based serving protects real-time inference and assigns remaining capacity to background tasks. If stricter separation is needed, users can use a hosted control plane based on OpenShift Virtualization.

The company said the distributed inference scheduler added a new feature that selects servers to reuse prior computation results, improving processing speed. LLM-D, an open-source project that supports distributed inference by splitting large language models across multiple servers, now formally supports CoreWeave CKS and Microsoft Azure beyond OpenShift, and is available in preview on Amazon EKS.

Red Hat also introduced AutoRAG for agent applications. AutoRAG connects enterprise data sources directly to agents and includes multilingual document support and context-based search. It also released the embedded Responses API in RAG as a formal version.

Responses API also includes NeMo Guardrails, which intercepts malicious tool calls. On AI Hub, agent templates for code review, document processing and research workflows can be used immediately. In observability, dashboards for inference status, GPU utilization and model performance were added. It measures per-user token usage for use as a basis for billing, and MLflow visually tracks agent behavior.

Keyword

#Red Hat AI 3.5 #EvalHub #AutoRAG #OpenShift #Responses API
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.