VAST Data has extended its partnership with AMD to advance AI infrastructure for cloud and enterprise platforms running model training, inference, retrieval-augmented generation (RAG), and agentic AI workloads. The integrated stack unites the VAST AI Operating System with 6th Gen AMD EPYC processors, AMD Instinct GPUs, AMD networking hardware, and ROCm software.
This collaboration aligns with a major industry shift: AI infrastructure is evolving beyond training-only clusters to support persistent context, high-concurrency inference, and multi-turn agent workflows. Modern AI deployments prioritize efficient data movement, optimized KV cache management, higher GPU utilization, and low-latency access to large models and contextual datasets.
VAST’s Disaggregated Shared Everything (DASE) architecture serves as the unified shared data layer for these modern AI environments. It consolidates file and object storage, databases, event streaming, and core data services under a single global namespace, while delivering multi-tenancy and workload isolation for AI clouds running diverse concurrent customer and application tasks.
AMD EPYC 9006 Powers Next-Gen VAST Hardware
VAST is adopting 6th Gen AMD EPYC 9006-series (Venice) processors for its upcoming 6th-gen CBox and 3rd-gen EBox systems, which form the hardware foundation of the VAST AI Operating System. The EPYC 9006 platform introduces PCIe Gen6 connectivity to VAST’s hardware lineup, doubling per-generation I/O bandwidth.
The upgrade boosts file and object storage throughput and lowers latency for database, data warehouse, and event streaming workloads supported by VAST DataBase and DataEngine. For AI infrastructure, enhanced I/O bandwidth eases bottlenecks across compute, networking and NVMe storage, benefiting model loading, retrieval pipelines, checkpoint access, and external KV cache workflows that exceed native GPU memory limits.
Tri-Player Reference Architecture with AMD and DriveNets
VAST, AMD and DriveNets are jointly building a unified AI infrastructure reference architecture, combining AMD’s rack-scale Helios AI infrastructure, VAST AI OS, and DriveNets AI Fabric networking. The framework delivers standardized deployment guidance for AI training, inference, reinforcement learning and KV cache workloads, providing sizing and availability best practices for enterprises building AI factories with shared data infrastructure and high-performance networking.
VAST has also broadened its ecosystem ties with inference vendors TensorMesh and EmbeddedLLM, focusing on production-grade inference architectures for agentic AI use cases, with specific integration and launch details yet to be disclosed.
Optimized KV Cache Offload for High-Concurrency Inference
A centerpiece of the updated collaboration is enhanced KV cache offloading, leveraging AMD Instinct GPUs, AMD Infinity Context, ROCm software and the VAST AI OS. KV cache stores inference-generated attention state data to accelerate multi-turn interactions and long-context AI workloads, yet large cache sizes rapidly consume limited GPU memory.
Offloading or tiering KV cache to high-performance shared storage frees GPU memory for active tasks while preserving contextual data for subsequent inference requests. Early testing on AMD Instinct MI355X GPUs delivered a 9x faster time-to-first-token and 9.7x higher token throughput for high-concurrency agentic AI workloads via VAST-powered KV cache offload, with performance varying by hardware baseline, workload and storage configuration.
Additionally, VAST’s automated lifecycle policies apply to KV cache data, enabling scheduled expiration and deletion of cached content — a critical feature for handling sensitive, personal or regulated inference data.
High-Speed GPU-Storage Interconnect via Pensando Pollara 400
The architecture deploys AMD Pensando Pollara 400 AI NICs to connect AMD Instinct GPUs to VAST storage clusters. Supporting NFS over TCP and RDMA, the NIC enables efficient GPU-to-storage data transfers, granting low-latency access to NVMe-based VAST storage for KV cache and general inference workloads.
This design decouples context management scaling from physical GPU memory limits. Rather than treating storage as passive persistence, VAST positions its platform as a unified data and execution layer for distributed model management, databases, streaming, file assets and AI context. For AI cloud providers, the architecture elevates GPU utilization and operational efficiency, supporting the industry transition from batch training and GPU rental to persistent inference and agentic AI services.
Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!



