Dr. Tony Skjellum is a Professor of Computer Science with over 40 years of experience designing and developing HPC and scientific computing middleware. A native of California, he graduated from Caltech in 1990, worked at LLNL for a few years, and has since been a professor at a total of five Universities in the Southeast. He works at Tennessee Technological University (since 2023), where he leads the ASCEND Center – Advanced Scalable Computing Extreme Networks and Data. His experience includes original work on the MPI Forum dating from 1992, the initial years of work together with Drs. Rusty Lusk and Bill Gropp on MPICH (an Argonne/Mississippi State collaboration), and a continuing set of contributions to MPI and HPC since then, including the MPI Software Technology, Inc years that produced MPI/Pro, a commercial MPI product used widely and part of a DOE Pathforward program to enhance MPI-2 capabilities on ASCI Leadership class systems of the Terascale era. He and his team are part of the NNSA PSAAP IV Center COMPASS led by the University of New Mexico, focusing on full-stack optimization of HPC applications enabled by AI. He and his colleagues are funded by the National Science Foundation on several projects, notably the MPI4AI and MPI Advance Efforts. Tennessee Tech is an up-and-coming R2 university in middle Tennessee, located between Knoxville and Nashville along I40 (about 100 miles north of Chattanooga).
Presentation Title:
Punctuated Equilibrium: Message Passing Middleware Design, Specification, and Implementation in the AI-Era
Presentation Abstract:
MPI and key implementations offered a stable, performance-portable interface for 30+ years, key HPC enabling technology. The AI era has punctuated it. GPU collective libraries hardened into de-facto interfaces within a few years, with no standard, reference implementation, or vendor neutrality. At the same moment, AI-assisted engineering changed what middleware costs to specify, build, and verify.
We report two responses: OpenCCL is a proposed vendor-neutral collective standard and reference implementation, API-upward-compatible with NCCL, now running multi-node on AMD MI300A over Slingshot-11. zettaMPI combines ExaMPI, MPI/Pro’s commercial-MPI design experience, and machine-readable MPI specifications toward a complete BSD-3 MPI-5.x implementation, the first new one offered the community in twenty years.
And, fabric portability layers have lost pertinence (e.g., libfabric); rather, fast low-level to high-level co-design replaces them, and that NCCL’s thin semantics should borrow MPI’s rigor rather than rediscover it over time. The talk covers the AI-enabled engineering methodology behind both projects.