logo
Home Cases

DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories

Certification
China Beijing Qianxing Jietong Technology Co., Ltd. certification
China Beijing Qianxing Jietong Technology Co., Ltd. certification
Customer Reviews
The sales staff of Beijing Qianxing Jietong Technology Co.,Ltd are very professional and patient. They can provide quotations quickly. The quality and packaging of the products are also very good. Our cooperation is very smooth.

—— 《Festfing DV》LLC

When I was looking for intel CPU and Toshiba SSD urgently, Sandy from Beijing Qianxing Jietong Technology Co., Ltd gave me a lot of help and got me the products I needed quickly. I really appreciate her.

—— Kitty Yen

Sandy of Beijing Qianxing Jietong Technology Co.,Ltd is a very careful salesman, who can remind me of configuration errors in time when I buy a server. The engineers are also very professional and can quickly complete the testing process.

—— Strelkin Mikhail Vladimirovich

We are very happy with our experience working with Beijing Qianxing Jietong. The product quality is excellent, and delivery is always on time. Their sales team is professional, patient, and very helpful with all our questions. We truly appreciate their support and look forward to a long-term partnership. Highly recommended!

—— Ahmad Navid

Quality: “Great experience with my supplier. The MikroTik RB3011 was already used, but it was in very good condition and everything works perfectly. Communication was fast and smooth, and all my concerns were addressed quickly. Very reliable supplier—highly recommended.”

—— Geran Colesio

I'm Online Chat Now

DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories

July 17, 2026
At Paris’s RAISE Summit, DDN showcased its ongoing collaboration with Nebul, a European sovereign hybrid cloud provider, focused on boosting efficiency for large-scale AI inference deployments. Unveiled last week, the joint initiative unites Nebul’s inference platform, DDN’s Infinia data intelligence architecture, and NVIDIA accelerated computing to resolve a key production AI bottleneck: data movement costs and performance limitations during inference workloads.

latest company case about DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories  0

DDN frames the collaboration around critical production metrics: GPU utilization, token throughput, cost per token, and latency. While model training builds AI asset value, inference defines its operational and commercial returns. The rising adoption of agentic AI, retrieval-augmented generation (RAG), and high-concurrency inference means storage and data infrastructure directly impact accelerator efficiency and AI response speeds.

This active proof-of-concept project has yielded promising early results. The partners have recorded measurable improvements in time-to-first-token with KV cache enabled and completed validation for RoCE-based infrastructure. Ongoing benchmarking covers longer inference sequence lengths, unlocking further optimization potential for the Infinia platform. The collaboration also expands to joint NVIDIA efforts on benchmarking frameworks, scalability verification, and upcoming technical publications.

The integrated platform leverages distributed KV cache services, GPU-native data movement, intelligent data orchestration, and high-performance storage architecture. KV cache acceleration delivers notable inference gains by preserving and rapidly retrieving pre-computed attention states, cutting redundant calculations and eliminating data delivery delays that cause GPU idling.

latest company case about DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories  1

Leaders from DDN, Nebul, and NVIDIA highlighted a major industry shift: AI infrastructure priorities are moving from raw GPU deployment to operational efficiency, maximizing returns from existing accelerator hardware. DDN CEO Alex Bouzari and Nebul CEO Arnold Juffer noted that past focus on larger model scales has given way to optimizing inference economics to make production AI commercially viable via lower per-token costs. NVIDIA Cloud Infrastructure VP Rod Evans added that large-scale agentic workloads now measure infrastructure success by GPU utilization and latency, rather than sheer compute power.

DDN emphasizes that AI infrastructure must evolve beyond basic storage functions to actively support AI execution workflows. The firm’s infrastructure platforms currently power over one million GPUs worldwide, serving hyperscalers, cloud providers, enterprises, governments, and research institutions.

Modern AI infrastructure teams now prioritize these core production metrics:
GPU utilization: Measures effective accelerator activity during inference, maximizing value of high-end GPU hardware.
Cost per token: Links infrastructure performance directly to AI model output operational costs.
Tokens per watt: Evaluates energy efficiency of AI inference output.
Time to first token: Determines interactive AI application responsiveness and user experience.
Time to production: Quantifies operational effort to migrate AI services from testing to scalable commercial deployment.

As inference becomes the dominant AI workload, delivering cached context and enterprise data to GPUs with low, stable latency will be critical to sustaining high GPU utilization and controlling long-term operational costs.

Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!

Contact Details
Beijing Qianxing Jietong Technology Co., Ltd.

Contact Person: Ms. Sandy Yang

Tel: 13426366826

Send your inquiry directly to us (0 / 3000)