WEKA, the AI data and memory infrastructure company, announced WEKA NeuralMesh 6, the most significant software release in the company’s history, delivering a purpose-built platform designed to help customers run production AI training, inference, and accelerated compute workloads at production scale on a single unified software stack. The release offers a robust set of capabilities that the AI infrastructure market has historically forced operators to assemble from separate vendors: native multi-tenancy at hyperscale, a full S3 protocol stack on NVMe, intelligent metadata-first data mobility that delivers async replication and remote caching, always-on data reduction with contractual guarantees, Kubernetes-native operations, and unified observability across every deployment.
Inference Costs Now Decide Which AI Companies Scale
The release arrives as the AI industry increasingly shifts decisively from training to production inference. Long-context reasoning, agentic workflows, and retrieval-driven AI workloads place sustained pressure on memory, metadata, and storage in ways traditional infrastructure architectures were not designed to handle. NeuralMesh 6 was built to deliver breakthrough inference economics at production scale for AI clouds, frontier model providers, government agencies, and large enterprises at the forefront of AI innovation.
Inference Economics Proven in Production on Oracle Cloud Infrastructure
NeuralMesh powers hyperscale AI inference in production today on Oracle Cloud Infrastructure (OCI). Leveraging WEKA‘s Augmented Memory Grid capability, which extends GPU memory by accelerating persistent KV cache access to NeuralMesh-managed NVMe storage, benchmarks on OCI H100 infrastructure have demonstrated 10x higher token throughput, 10x more concurrent users served, and 7x more tokens per GPU in production deployments.
Also Read: GitLab Appoints Brittany Lutz as Product Marketing Manager for AI
“As agentic AI workloads push context windows and GPU utilization to new limits, WEKA and Oracle Cloud Infrastructure are helping customers scale inference more efficiently without simply adding more GPUs,” Pablo Selem, senior director, software development, Oracle Cloud Infrastructure. “WEKA’s NeuralMesh platform with Augmented Memory Grid on OCI helps remove memory bottlenecks, delivering substantially more throughput and concurrent users from the same GPU footprint. For customers, that means higher ROI on infrastructure investments and a clearer path to cost-efficient AI at scale.”
“WhiteFiber is building a distributed GPU platform across multiple data centers connected by high-speed dark fiber. At our scale, data mobility isn’t a nice-to-have; it’s foundational,” said Sam Tabar, CEO at WhiteFiber. “NeuralMesh’s intelligent replication makes data mobility real at scale: we can make datasets visible across sites and pull exactly the data each job needs to the next GPU allocation as it becomes available. That shifts replication from a back-end protection function to a core part of how our distributed AI infrastructure needs to operate, improving workload mobility, capacity efficiency, and the resiliency our customers depend on. WEKA’s data and memory infrastructure provides the foundation to scale our footprint without compromise. We’re excited to keep building on that together.”
“The infrastructure operators running production AI today have been forced to assemble platforms from vendors that were never designed to work together. Separate stacks for file and object, manual data movement between them, multi-tenancy bolted on after the fact. NeuralMesh 6 delivers what they’ve actually needed all along: a single platform that handles the high-performance file layer and the high-capacity object layer on the same blocks, with native multi-tenancy, intelligent data mobility, and always-on data efficiency built in from the start. This is what production inference infrastructure looks like when it’s designed for the workload, not retrofitted for it,” said Ajay Singh, Chief Product Officer at WEKA.
“Agentic AI and reasoning workloads require infrastructure that thinks as fast as the models running on it. By combining Spectro Cloud PaletteAI and WEKA NeuralMesh, we give enterprises a single, jointly validated path to deploy, scale, and operate AI factories. Our joint solution delivers full-stack orchestration alongside data performance at the speed of inference, with the governance and repeatability that production demands,” said Saad Malik, CTO and co-founder, Spectro Cloud
“Our customers are scaling AI production workloads and asking hard questions about infrastructure efficiency and cost. WEKA’s new NeuralMesh platform capabilities – including Augmented Memory Grid and data reduction with its performance guarantee – give us compelling, concrete answers. We’re pleased to work alongside WEKA to bring these innovations to organizations across our markets,” said Joseph Giam, Managing Director, Glocomp Systems (M) Sdn Bhd.
“The economics of AI infrastructure are shifting fast, and our reseller partners are looking for solutions that help them win deals. WEKA’s NeuralMesh software innovations are designed to help organizations address the performance and efficiency demands of modern AI workloads. We’re excited to support the channel in bringing these capabilities to customers,” said Eunice Lau, Executive Managing Director, Singapore, Ingram Micro.
SOURCE: WEKA























