Next-Gen Hadoop Storage: BlueField-3 and CSAL for Performance and Resilience

Keynote
Shanghai
3:50 p.m. - 4:30 p.m.
Venue A(Integrated Building, Room 506 Lecture Hall)
  • Wayne Gao Principal Engineer and Storage Solution Architect at Solidigm

    Wayne Gao is a Principal Engineer and Storage Solution Architect at Solidigm, where he led the CSAL (Compute Storage Acceleration Library) project from Intel proof-of-function through its commercial release at Alibaba, and served as the primary developer for its PMem, DSA, and CXL.mem implementations during the Intel-to-Solidigm transition.

    With over 20 years of storage development experience, he previously worked as a Sr. Principal Software Engineer in Dell EMC’s ObjectScale all-flash object storage team and as a P8-level engineer at Alibaba, building deep connections within China’s enterprise and cloud storage communities.

    Wayne holds 4 U.S. patents (filed/granted), co-authored the EuroSys 2024 paper on CSAL, received the USENIX FAST ’26 Best Paper Award, and has first-author acceptances at GTC 2025/2026 (posters) and KubeCon 2025 (presentation).

    He currently serves as an ODCC Expert Committee member representing Solidigm in storage software for the PRC.

    In his spare time, he enjoys badminton, swimming, movies, and gaming.

    waynegao

Abstract

CSAL: open-source user-mode FTL, cache, and I/O trace component accelerating Big Data and AI storage within SPDK, commercially powering Alibaba Cloud.

Details

CSAL: Open-source user-mode FTL, cache, and I/O trace component accelerating Big Data and AI storage within SPDK (upstreamed); commercially powers Alibaba Cloud. See: solidigm.com/csal | Alibaba-Solidigm Eurosys'24 paper | BlueField-3/Solidigm GTC'25/26 session.

Traditional Hadoop storage systems face severe bottlenecks in high-concurrency, throughput, and low-latency demands for large-scale analytics, AI training/inference, and real-time processing due to "share-everything" design. BlueField-3 + CSAL solution: Optimizes data paths for breakthrough throughput, offloads storage tasks to DPU (reduces CPU overhead by 70-80%), and redesigns HDFS with computation-storage separation for superior resource efficiency.

Triple replication, erasure coding (EC), and RAID deliver multi-layered redundancy, ensuring high reliability, data integrity, and enterprise security across failure scenarios. Supports diverse use cases: large-scale analytics, AI model training/inference, real-time streaming etc.