As generative AI models continue to outgrow conventional computing infrastructure, memory and storage have emerged as critical bottlenecks for enterprise AI deployments. At Future of Memory and Storage (FMS) 2026, ScaleFlux will showcase its vision for next-generation AI infrastructure through a keynote co-presented with NVIDIA, alongside a series of technical sessions exploring flash storage, CXL memory, SSD architecture, and storage security.
The rapid expansion of artificial intelligence workloads is reshaping data center architecture, forcing infrastructure providers to rethink how memory and storage systems support increasingly large AI models. Against this backdrop, ScaleFlux announced a major presence at Future of Memory and Storage (FMS) 2026, where CEO and co-founder Hao Zhong will deliver a keynote with Jason Hardy, Vice President of Storage Technology at NVIDIA, examining how emerging memory technologies can improve AI scalability and reduce infrastructure costs.
Beyond the keynote, ScaleFlux will participate in seven technical sessions across the three-day conference, highlighting developments in AI infrastructure, flash storage, Computational Express Link (CXL) memory, solid-state drive (SSD) architecture, and storage security. The company said the presentations reflect growing industry demand for more efficient approaches to managing the data-intensive requirements of large language models (LLMs), generative AI, and inference workloads.
AI Infrastructure Is Driving a New Memory Hierarchy
The headline keynote, “Memory Solutions to Scale AI Data Pipeline,” will focus on one of the most pressing challenges facing AI infrastructure: memory capacity.
While GPUs continue to deliver increasing computational performance, memory availability has become a limiting factor for training and deploying large AI models. ScaleFlux and NVIDIA plan to discuss how flash storage can evolve beyond its traditional role as persistent storage and function as an additional memory tier for AI inference. Such architectures could enable larger models, improve GPU utilization, and lower infrastructure costs by reducing dependence on expensive high-bandwidth memory (HBM).
The discussion reflects a broader industry trend in which hyperscale cloud providers and AI infrastructure vendors are exploring memory tiering strategies to balance performance, capacity, and cost as enterprise AI deployments expand.
Addressing Storage Challenges for AI Workloads
Several ScaleFlux sessions will examine how storage architectures must evolve to accommodate increasingly heterogeneous AI workloads.
One presentation explores redesigned SSD controllers optimized for AI inference, where storage devices are expected to manage diverse access patterns generated by retrieval-augmented generation (RAG), vector databases, and multimodal AI applications. Another technical session addresses write amplification, a longstanding challenge affecting SSD endurance and performance, by introducing architectural approaches that minimize unnecessary write operations without relying on conventional data placement techniques.
These developments highlight how storage hardware is becoming increasingly intelligent rather than simply serving as passive data repositories.
CXL Memory and Cost-Efficient AI Scaling
Another focus area is CXL (Compute Express Link) memory, an emerging interconnect standard designed to enable memory expansion across servers and accelerators.
ScaleFlux Chief Scientist Prof. Tong Zhang will present research examining how CXL memory can improve reliability, availability, and serviceability (RAS) while reducing total cost of ownership for enterprise and AI infrastructure. As organizations deploy larger AI clusters, CXL is increasingly viewed as a key technology for overcoming memory limitations without proportionally increasing hardware costs.
Prof. Zhang will also lead a session addressing the growing cost of high-bandwidth memory (HBM), which has become one of the most expensive components in modern AI servers. The presentation explores architectural alternatives that could reduce reliance on HBM while maintaining inference performance for enterprise AI applications.
Flash Storage Moves Closer to Memory
One of the conference’s collaborative sessions between ScaleFlux and NVIDIA revisits the classic Five-Minute Rule, a long-standing principle in computer architecture used to determine whether data should reside in memory or storage.
Advances in flash technology, lower latency storage media, and AI-driven workloads are prompting researchers to reconsider those traditional assumptions. The joint presentation will examine how flash increasingly functions as an extension of system memory rather than merely long-term storage, reflecting broader changes in AI infrastructure design.
As memory and storage technologies converge, organizations may gain greater flexibility in balancing performance, capacity, and operational cost across AI deployments.
Security Becomes a Core Infrastructure Requirement
ScaleFlux’s conference agenda also includes a session focused on SSD security, addressing the growing importance of trusted storage in AI environments.
Modern storage devices increasingly incorporate onboard processors, firmware intelligence, and computational capabilities, expanding both their functionality and potential attack surface. The presentation examines architectural requirements for ensuring transparency, integrity, and security as storage becomes an active participant in AI data pipelines.
Market Landscape
Enterprise investment in AI infrastructure continues to accelerate as organizations expand generative AI deployments. According to IDC, worldwide spending on AI infrastructure—including compute, networking, storage, and memory—continues to grow rapidly as enterprises modernize data centers to support increasingly demanding AI workloads. Meanwhile, Gartner projects that AI infrastructure spending will increasingly shift toward memory optimization, storage acceleration, and intelligent data management as organizations seek to improve GPU utilization and reduce operational costs.
This evolution has elevated technologies such as CXL memory, computational storage, intelligent SSD controllers, and flash-based memory tiers from niche innovations to strategic infrastructure components. Alongside companies such as NVIDIA, AMD, Intel, Samsung, Micron, and Kioxia, ScaleFlux is contributing to an industry-wide effort to redesign data infrastructure around the requirements of AI-native computing rather than traditional enterprise applications.
As enterprises evaluate next-generation AI platforms, innovations in memory hierarchy, storage efficiency, and infrastructure economics are expected to play an increasingly important role in determining scalability and total cost of ownership.
Top Insights
- ScaleFlux will deliver seven technical presentations at FMS 2026, highlighting innovations in AI infrastructure, CXL memory, flash storage, SSD architecture, and enterprise storage security.
- CEO Hao Zhong will join NVIDIA’s Jason Hardy to discuss how flash memory can function as an additional memory tier for AI inference, improving GPU utilization and reducing infrastructure costs.
- Technical sessions explore emerging approaches to CXL memory expansion, intelligent SSD controller design, write amplification reduction, and computational storage for AI workloads.
- ScaleFlux is addressing one of AI infrastructure’s biggest challenges—the rising cost of high-bandwidth memory (HBM)—through alternative memory architectures and scalable storage technologies.
- The company’s research reflects broader industry efforts to redesign enterprise infrastructure for large language models, generative AI, and next-generation inference workloads.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI












