WEKA has launched the exabyte-scale WEKAPod all-flash array, optimized for its sixth-generation NeuralMesh software. The new platform delivers massive data throughput to power production-grade AI training and inference workloads. It positions WEKA alongside premium AI storage vendors including DDN, Everpure, NetApp and VAST Data, featuring industry-leading density of 1.1 exabytes of effective capacity within a single 56 RU rack.

NeuralMesh v6 integrates native hyperscale multi-tenancy, an NVMe-based unified file and S3 object stack, intelligent data mobility with replication and remote caching, performance-guaranteed always-on data reduction, Kubernetes-native workflows, and unified observability. On Oracle Cloud Infrastructure (OCI) H100 hardware, its Augmented Memory Grid delivers 10x higher token throughput, 10x more concurrent users, and 7x higher per-GPU token output than NVIDIA CMX in production. For CoreWeave deployments, it boosts per-GPU token throughput by 4.2x and cuts TTFT latency by up to 6x versus CMX.
WEKA co-founder and CEO Liran Zvibel commented in July 2026: “Legacy AI infrastructure built for traditional workloads cannot support modern inference-era operations. Today’s AI profitability hinges on rack space efficiency, energy consumption, supply chain stability and operational density.
“WEKAPod is purpose-built to fit these constraints. Engineered exclusively for inference-era computing with a self-controlled supply chain, it delivers upfront cost advantages and superior TCO across full-stack AI infrastructure needs, setting a new benchmark for large-scale production AI infrastructure evaluation.”
The WEKAPod is a dedicated storage appliance that feeds data to GPU clusters such as NVIDIA SuperPODs, integrating pre-configured hardware and WEKA’s data platform software. First launched in late 2025 with PCIe Gen4 Prime and Gen5 Nitro models featuring AlloyFlash tiered SSDs, the updated third-generation lineup adds 2RU-based Nitro, Prime and new Prime Max configurations tailored for production AI inference.
Prime Max is WEKA’s highest-capacity model, exceeding 1 EB per rack. Prime delivers balanced price-performance via four 2U rack servers with 70 mixed QLC/TLC NVMe drives managed by AlloyFlash. Nitro targets high-performance AI factories and GPU-as-a-service environments with hundreds to thousands of active GPUs.
The third-gen WEKAPod is the industry’s first single-rack system to surpass exabyte-scale effective capacity. It packs 441.5 PB of raw NVMe storage to deliver 1.1 EB effective capacity via NeuralMesh data reduction, alongside 10.2 TB/s per-rack throughput and 210 million rack-level IOPS.
Its upgraded hardware stack includes PCIe Gen6 internal fabric, backplane-free cable-based drive interconnects, NVIDIA ConnectX SuperNICs with Spectrum-X Ethernet connectivity, and software-defined thermal management. The platform minimizes data center footprint, optimizes rack space utilization and improves tokens-per-watt efficiency without compromising throughput or workload performance. WEKA states it outperforms competing solutions by 267% in effective capacity density and 114% in throughput density per rack unit.
By directly managing component sourcing without OEM channel reliance, WEKA provides AI cloud providers and enterprises with stable pricing and predictable lead times for large-scale infrastructure deployments. All WEKAPod units ship factory-configured and pre-validated with full NeuralMesh v6 capabilities including Augmented Memory Grid, eliminating the need for dedicated on-site storage teams.
NeuralMesh v6, WEKA’s parallel file system, leverages Augmented Memory Grid to extend GPU memory for inference workloads. It externalizes storage as a low-latency (microsecond-level) KV cache with multi-TB/s bandwidth, delivering petabyte-scale extended address space for AI models.
Key NeuralMesh 6 capabilities include:
Composable Clusters: Delivers full hardware isolation with dedicated CPU, memory and storage for anchor tenants, ensuring guaranteed resource allocation, strict workload separation and consistent performance under heavy loads.
Virtual Multi-Tenancy: Offers VPC-grade isolation via Virtualized RDMA Data Fabric (VRDF), supporting private VLANs, overlapping IP spaces, per-tenant QoS, independent KMS encryption and LDAP/AD authentication. It scales to over 1,000 isolated logical tenants per cluster with new tenant provisioning under 30 minutes.
Shared physical NVMe drives across multiple NeuralMesh clusters maximize device parallelism for multi-tenant inference. A single hardware cluster supporting 50 Composable Clusters can host up to 50,000 isolated logical tenants. It unifies S3 and POSIX file access under one namespace, enabling zero-copy S3-over-RDMA data transfers directly to GPU memory with 2,000–5,000 concurrent node connections — roughly five times conventional S3 concurrency.
Its AlloyFlash technology transparently unifies cost-effective QLC and high-performance TLC NVMe flash in one cluster, routing latency-sensitive workloads to TLC and bulk capacity tasks to QLC at 30–40% lower per-TB cost.
Always-on data reduction combines fingerprinting, similarity hashing, deduplication and compression with under 5% write overhead, delivering up to 6x capacity savings for AI training data. WEKA provides contractual guarantees for both reduction ratios and performance impact, with pre-sales workload analysis to forecast real-world efficiency.
All data reduction processes run off the critical write path via background similarity compression and cross-filesystem deduplication, enabling native NVMe-speed writes regardless of capacity usage or background load. This avoids the write performance degradation common to inline reduction architectures during checkpoint-heavy AI training, mirroring ExaGrid’s deduplication-free landing zone design for consistent performance.
NeuralMesh v6 supports full federation and global namespace capabilities. Its asynchronous replication enables AI data mobility, cloudbursting and multi-site collaboration, with priority metadata replication for instant destination environment browsing and on-demand data hydration. This reduces WAN bandwidth consumption via replicated compressed data.
A dedicated Kubernetes Operator automates cluster deployment and lifecycle management for Kubernetes-based AI clouds and research labs. The SaaS-based unified Observe platform delivers multi-cluster dashboards, client-level diagnostics, configurable intelligent alerting, and integrated routing to Slack, PagerDuty or email with direct dashboard links.
WEKA Chief Product Officer Ajay Singh stated: “Today’s production AI operators rely on disjointed, retrofitted third-party stacks with siloed file/object layers, manual data movement and bolted-on multi-tenancy. NeuralMesh 6 solves these pain points with a unified platform that converges high-performance file and high-capacity object storage on shared hardware, with native multi-tenancy, intelligent data mobility and built-in data efficiency. It is purpose-built for production inference workloads, not retrofitted for them.”
NeuralMesh 6 will reach general availability in H2 2026, with free seamless upgrades available to all existing WEKA customers via standard update channels.
Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!