Technical Program
13:30-13:35 Opening Remarks
13:35-14:25 Keynote Speech
Rethinking the Data Path: Distributed In-Memory Key-Value Stores with DPUs
Speaker: André Brinkmann (Saarland University)
Abstract: Remote in-memory key-value stores are a cornerstone of modern distributed systems, powering caching layers, real-time analytics, and coordination services at scale. As application demands grow, so do the requirements on these systems, spanning throughput, range query support, fault tolerance, and operational simplicity. Achieving all of these simultaneously remains an open challenge, and the diversity of existing approaches reflects how differently these tradeoffs can be resolved. This talk surveys the landscape of distributed, high-performance KV store designs, from classical host-based architectures limited by the kernel network stack, to kernel-bypass techniques and RDMA-based distributed designs. We then focus on Data Processing Units (DPUs) and SmartNICs, hardware that sits directly on the network data path and processes requests without involving the host CPU or OS. We base this discussion on DPA-Store, our DPU-based KV store, illustrating how on-path processing enables low-latency request handling, stateless clients, and range query support, while avoiding the complications of RDMA-based approaches. We close with open challenges, including consistency under failure, storage hierarchy integration, and learned data structures in network-bound environments.
Biography: Prof. Dr.-Ing. André Brinkmann has been a Full Professor in the Department of Computer Science at Saarland University in Saarbrücken since 2026. Prior to this, he held the position of Full Professor of Computer Science at Johannes Gutenberg University (JGU) Mainz and served as director of the JGU Data Center. He earned his Ph.D. in Electrical Engineering from Paderborn University in 2004 and subsequently served as Assistant Professor in the Department of Computer Science at Paderborn University from 2008 to 2011. His research focuses on applying algorithm engineering to storage systems, data center management, and high-performance computing. He has authored over 150 publications in leading conferences and journals and serves as Senior Associate Editor of ACM Transactions on Storage.
14:25-15:00 Invited Talk
LLM Inference at Wafer Scale: When the Model Lives in On-Chip Memory
Speaker: Luo Mai (University of Edinburgh)
Abstract: LLM inference systems are built on GPU-centric assumptions: large off-chip memory, coarse-grained kernels, and communication among a few devices. Wafer-scale accelerators break all of them, offering hundreds of thousands of cores and tens of gigabytes of distributed on-chip memory with far more bandwidth than HBM. This talk presents WaferLLM, an LLM inference system that treats inference as a distributed on-chip memory management problem: it partitions weights and the KV-cache across the wafer, parallelizes prefill and decode at wafer scale, and introduces two near-memory kernels, MeshGEMM and MeshGEMV. On a Cerebras WSE-2, WaferLLM serves a single user at 2,700 tokens/s — sub-millisecond per-token latency. I will close with ongoing work spanning on-chip and external memory tiers for interactive reasoning and test-time compute.
Biography: PLuo Mai is an Associate Professor at the University of Edinburgh, where he leads the Large-Scale Machine Learning Systems Group. He builds systems for LLM training and inference, including WaferLLM (OSDI '25), ServerlessLLM (OSDI '24), and Tenplex (SOSP '24), with wafer-scale compilation and runtime work appearing at SOSP '26. He co-directs the EPSRC Centre for Doctoral Training in Machine Learning Systems and leads projects in the UK ARIA Scaling AI Compute programme. He received his PhD from Imperial College London and previously worked at Microsoft Research.
15:00-15:30 Coffee and Tea Break
15:30-16:20 Keynote Speech
Two Peas in a Pod: Storage and Memory Heterogeneity in the Age of AI
Speaker: Sudarsun Kannan (Rutgers University)
Abstract: Modern servers no longer have a simple split between memory and storage. They have a continuum of unequal tiers, from HBM and DRAM to CXL-attached memory and flash, differing 10 to 100 times in latency and bandwidth, with compute increasingly embedded network, memory, and storage devices. Software abstractions for memory and storage, however, have barely changed, and today's AI challenges echo classic ones: KV-cache eviction recalls buffer-cache management; model paging, virtual memory; checkpointing, logging; pipeline stalls, prefetching. The solutions transfer too, because the principles were never device specific. In this talk, using cross-layered design as a running example, I will explore what carries over from the past decade of storage and memory systems research and where it falls short, from datacenter and HPC AI to edge cyberinfrastructure for natural-hazard monitoring. I will conclude the talk with open challenges and ongoing work toward treating storage and memory heterogeneity as two peas in a pod.
Biography: Sudarsun Kannan is an Associate Professor in the Computer Science Department at Rutgers University. His research focuses on operating system design and its intersection with computer architecture, distributed systems, and high-performance computing. His group also builds edge cyberinfrastructure for hazard-monitoring systems, including wildfire detection and mitigation. He is a recipient of the NSF CAREER Award, the Google Research Scholar Award, and the 2025 Rutgers Board of Trustees Fellowship for outstanding early-career faculty, and his group's research has received best paper and distinguished paper awards at SOSP, ASPLOS, and SPAA. He co-chaired the HotStorage '22 workshop and serves as an Associate Editor for ACM Transactions on Storage. Before joining Rutgers, he was a postdoctoral research associate at the University of Wisconsin-Madison, and he received his M.S. and Ph.D. from Georgia Tech.
16:55-17:30 BigMem Infrastructure
CXL-SDK: Building a Software Development Kit for CXL Shared Memory
Authors: Fangnuo Wu, Mingkai Dong, Haibo Chen (Shanghai Jiao Tong University)
Fresco: Enabling ACK-Time Searchable Durability for Fresh Vector Index Updates with Tiered Persistent Memory
Authors: Sen Jiang, Yinjin Fu (Sun Yat-sen University)
The rapid evolution of memory-centric computing technologies—including Compute Express Link (CXL), persistent memory, and disaggregated architectures—is fundamentally reshaping system software design paradigms. The International Workshop on Big Memory (BigMem 2026) establishes a premier forum for researchers and practitioners to explore the challenges and opportunities in managing terabyte-to-petabyte scale memory hierarchies in modern computing systems.
Our focuses include:
- OS & Memory Hierarchy Redesign: Traditional OS abstractions, such as virtual memory and caching, are challenged by heterogeneous memory technologies and disaggregated architectures.
- Distributed Memory & Networking: Memory-centric networking technologies, such as RDMA and CXL.mem, blur the line between local and remote memory, necessitating new consistency models and failure handling mechanisms.
- Applications & Workloads: Big-memory applications, such as machine learning, graph processing, and in-memory databases, reveal performance bottlenecks in current OS and architecture designs.
BigMem 2026 aims to foster cross-disciplinary collaboration between operating system researchers, computer architects, and industry practitioners. By bringing together diverse perspectives, we seek to accelerate innovation in big memory systems and establish this workshop as the leading venue for memory-centric computing research.
BigMem 2026 will be co-located with SOSP 2026 and registrations will be handled by SOSP 2026.
Topics of Interest
This workshop will focus on important research directions, including operating system support for heterogeneous memory systems integrating DRAM, CXL and/or persistent memory, memory-centric networking technologies enabling efficient remote memory access, and innovative applications leveraging massive memory capacities. Moreover, we will examine emerging challenges in memory virtualization, coherence protocols, and fault tolerance for next-generation memory architectures.
The topics include, but are not limited to:
Memory Architecture & Systems
- Memory-system design (hardware/software co-optimization)
- Disaggregated memory architectures (e.g., CXL, NVMe-over-Fabrics)
- Operating systems for hybrid/heterogeneous memory (NVM, DRAM, HBM)
- Design and operation of large-scale memory systems
Memory Technologies & Reliability
- Emerging memory technologies (3D XPoint, STT-MRAM, FeRAM, optical memory)
- Memory failure modes, reliability, and fault mitigation
- Energy-efficient memory subsystems
Security & Safety
- Memory security (side-channel attacks, encryption, secure allocators)
- Memory safety for critical systems
Programming & Performance
- Memory-centric programming models and languages
- In-memory and near-memory computing architectures
- Algorithmic memory optimizations (caching, prefetching, compression)
- Non-volatile memory programming paradigms
Applications & Storage
- In-memory databases, NoSQL stores, and analytics
- Memory-driven AI/ML workloads
- Embedded and autonomous systems
Emerging Directions
- Memory disaggregation for cloud/edge environments
- Bio-inspired memory architectures
By bridging architecture, systems, and applications, BigMem 2026 seeks to shape future research directions in operating systems.
Submission Guidelines
We invite original research contributions that have not been published previously or submitted concurrently to other venues, including any other conference or journal. Authors should prepare their work as a two-page extended abstract (references excluded) in English, formatted as a PDF document. The ACM submission template is recommended.
All submissions must be made through the workshop's online submission system and will undergo a double-blind review process by our program committee. Please ensure your submission is properly anonymized to maintain the integrity of the review. Submissions will be evaluated on their technical merit, novelty, relevance to the workshop themes, and clarity of presentation. This workshop focuses on discussion and feedback rather than archival publication, and therefore does not produce formal proceedings.
Submission site: https://bigmem26.hotcrp.com
Important Dates
- Submissions due: June 15, 2026 (AOE)
- Notification of acceptance: July 10, 2026
- Workshop presentation: September 29, 2026
Program Co-Chairs:
Yu Hua, Huazhong University of Science and Technology
Xue Liu, McGill University and MBZUAI
Web/Submission Chair:
Yujie Hu, Huazhong University of Science and Technology
Program Committee Member:
Ashvin Goel, University of Toronto
Bingzhe Li, The University of Texas at Dallas
Daniel S. Berger, Microsoft
Hui Lu, The University of Texas at Arlington
Jiacheng Shen, Duke Kunshan University
Juncheng Yang, Harvard University
Luo Mai, University of Edinburgh
Mai Zheng, Iowa State University
Michael Swift, University of Wisconsin—Madison
Mingxing Zhang, Tsinghua University
Oana Balmau, McGill University
Sanidhya Kashyap, EPFL
Vijay Chidambaram, The University of Texas at Austin
Yao Fu, NVIDIA Research
Youhui Bai, University of Science and Technology of China
Yuxin Ren, Huawei
Zhichao Cao, Arizona State University