logo
Home Cases

VAST Data and AMD Claim 9x Faster Time-to-First-Token With KV Cache Offload on Instinct

Certification
China Beijing Qianxing Jietong Technology Co., Ltd. certification
China Beijing Qianxing Jietong Technology Co., Ltd. certification
Customer Reviews
The sales staff of Beijing Qianxing Jietong Technology Co.,Ltd are very professional and patient. They can provide quotations quickly. The quality and packaging of the products are also very good. Our cooperation is very smooth.

—— 《Festfing DV》LLC

When I was looking for intel CPU and Toshiba SSD urgently, Sandy from Beijing Qianxing Jietong Technology Co., Ltd gave me a lot of help and got me the products I needed quickly. I really appreciate her.

—— Kitty Yen

Sandy of Beijing Qianxing Jietong Technology Co.,Ltd is a very careful salesman, who can remind me of configuration errors in time when I buy a server. The engineers are also very professional and can quickly complete the testing process.

—— Strelkin Mikhail Vladimirovich

We are very happy with our experience working with Beijing Qianxing Jietong. The product quality is excellent, and delivery is always on time. Their sales team is professional, patient, and very helpful with all our questions. We truly appreciate their support and look forward to a long-term partnership. Highly recommended!

—— Ahmad Navid

Quality: “Great experience with my supplier. The MikroTik RB3011 was already used, but it was in very good condition and everything works perfectly. Communication was fast and smooth, and all my concerns were addressed quickly. Very reliable supplier—highly recommended.”

—— Geran Colesio

I'm Online Chat Now

VAST Data and AMD Claim 9x Faster Time-to-First-Token With KV Cache Offload on Instinct

July 28, 2026
VAST Data has extended its partnership with AMD to advance AI infrastructure for cloud and enterprise platforms running model training, inference, retrieval-augmented generation (RAG), and agentic AI workloads. The integrated stack unites the VAST AI Operating System with 6th Gen AMD EPYC processors, AMD Instinct GPUs, AMD networking hardware, and ROCm software.

This collaboration aligns with a major industry shift: AI infrastructure is evolving beyond training-only clusters to support persistent context, high-concurrency inference, and multi-turn agent workflows. Modern AI deployments prioritize efficient data movement, optimized KV cache management, higher GPU utilization, and low-latency access to large models and contextual datasets.

VAST’s Disaggregated Shared Everything (DASE) architecture serves as the unified shared data layer for these modern AI environments. It consolidates file and object storage, databases, event streaming, and core data services under a single global namespace, while delivering multi-tenancy and workload isolation for AI clouds running diverse concurrent customer and application tasks.

AMD EPYC 9006 Powers Next-Gen VAST Hardware

VAST is adopting 6th Gen AMD EPYC 9006-series (Venice) processors for its upcoming 6th-gen CBox and 3rd-gen EBox systems, which form the hardware foundation of the VAST AI Operating System. The EPYC 9006 platform introduces PCIe Gen6 connectivity to VAST’s hardware lineup, doubling per-generation I/O bandwidth.

latest company case about VAST Data and AMD Claim 9x Faster Time-to-First-Token With KV Cache Offload on Instinct  0

The upgrade boosts file and object storage throughput and lowers latency for database, data warehouse, and event streaming workloads supported by VAST DataBase and DataEngine. For AI infrastructure, enhanced I/O bandwidth eases bottlenecks across compute, networking and NVMe storage, benefiting model loading, retrieval pipelines, checkpoint access, and external KV cache workflows that exceed native GPU memory limits.

Tri-Player Reference Architecture with AMD and DriveNets

VAST, AMD and DriveNets are jointly building a unified AI infrastructure reference architecture, combining AMD’s rack-scale Helios AI infrastructure, VAST AI OS, and DriveNets AI Fabric networking. The framework delivers standardized deployment guidance for AI training, inference, reinforcement learning and KV cache workloads, providing sizing and availability best practices for enterprises building AI factories with shared data infrastructure and high-performance networking.

VAST has also broadened its ecosystem ties with inference vendors TensorMesh and EmbeddedLLM, focusing on production-grade inference architectures for agentic AI use cases, with specific integration and launch details yet to be disclosed.

Optimized KV Cache Offload for High-Concurrency Inference

A centerpiece of the updated collaboration is enhanced KV cache offloading, leveraging AMD Instinct GPUs, AMD Infinity Context, ROCm software and the VAST AI OS. KV cache stores inference-generated attention state data to accelerate multi-turn interactions and long-context AI workloads, yet large cache sizes rapidly consume limited GPU memory.

latest company case about VAST Data and AMD Claim 9x Faster Time-to-First-Token With KV Cache Offload on Instinct  1

Offloading or tiering KV cache to high-performance shared storage frees GPU memory for active tasks while preserving contextual data for subsequent inference requests. Early testing on AMD Instinct MI355X GPUs delivered a 9x faster time-to-first-token and 9.7x higher token throughput for high-concurrency agentic AI workloads via VAST-powered KV cache offload, with performance varying by hardware baseline, workload and storage configuration.

Additionally, VAST’s automated lifecycle policies apply to KV cache data, enabling scheduled expiration and deletion of cached content — a critical feature for handling sensitive, personal or regulated inference data.

High-Speed GPU-Storage Interconnect via Pensando Pollara 400

The architecture deploys AMD Pensando Pollara 400 AI NICs to connect AMD Instinct GPUs to VAST storage clusters. Supporting NFS over TCP and RDMA, the NIC enables efficient GPU-to-storage data transfers, granting low-latency access to NVMe-based VAST storage for KV cache and general inference workloads.

This design decouples context management scaling from physical GPU memory limits. Rather than treating storage as passive persistence, VAST positions its platform as a unified data and execution layer for distributed model management, databases, streaming, file assets and AI context. For AI cloud providers, the architecture elevates GPU utilization and operational efficiency, supporting the industry transition from batch training and GPU rental to persistent inference and agentic AI services.

Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!
Contact Details
Beijing Qianxing Jietong Technology Co., Ltd.

Contact Person: Ms. Sandy Yang

Tel: 13426366826

Send your inquiry directly to us (0 / 3000)