Jülich Researchers Earn EuroHPC Award for Study of P3M Performance on GPUs

Science & Technology

Jülich Team Wins EuroHPC Award for Multi-GPU P3M Solver Work

JSC researchers won Best Paper at EuroHPC User Days 2026 for a portable P3M electrostatics library achieving near-ideal scaling on up to 128 GPUs across three EU supercomputers.

By
Grace Kim
Filed
Channel
Science & Technology
Read
3 min read

Sept. 28, 2026 — Three researchers from Jülich Supercomputing Centre (JSC) received the Best Paper Award at the EuroHPC User Days 2026 in Dublin last week for a study of how the particle-particle particle-mesh (P3M) method scales across multiple GPUs. Rodrigo Bartolomeu, René Halver and Godehard Sutmann, all members of JSC's Simulation and Data Laboratory Complex Particle Systems, earned the prize for their paper "Scaling and performance aspects of the P3M method on multiple GPUs," published in Procedia Computer Science, Volume 286, pages 14-23.

The work addresses one of the oldest cost problems in molecular simulation. Long-range electrostatic interactions, especially under periodic boundary conditions, impose a computational complexity of O(N²) when computed directly. For 100 particles that means on the order of 10,000 calculations; for 1,000 particles, one million. Mesh-based Ewald methods such as P3M cut the complexity to O(N log N), reducing those same workloads to roughly 200 and 3,000 calculations respectively.

The gain in arithmetic cost comes at a price in parallel efficiency. The authors identify three persistent obstacles: parallel scalability, communication overhead, and efficient utilization of heterogeneous CPU-GPU architectures. Highly optimized P3M implementations already exist inside established molecular dynamics packages, but the researchers note that these solvers are typically tightly coupled to their full simulation frameworks. That coupling limits reuse in emerging HPC applications.

The Jülich team's answer is a performance-portable library for electrostatic solvers offering both classical Ewald and P3M implementations, targeting CPU and GPU architectures. The design goal is efficient execution on multi-GPU systems without sacrificing portability across heterogeneous platforms — and without locking users into a specific simulation framework.

The paper's core results concern the FFT-based long-range component of P3M, long known as the scalability bottleneck. Strong-scaling experiments show a split outcome: small problem sizes are dominated by communication overhead and lose parallel efficiency, while larger systems achieve near-ideal scaling across a wide range of GPU counts. The published benchmarks cover system replications from 8 to 64 running on 1 to 128 GPUs (1 to 32 nodes) on the MareNostrum5 accelerated partition at Barcelona Supercomputing Centre, with dashed reference lines marking ideal speedup and a parallel efficiency of 1.

Benchmarks were not confined to one machine. The authors ran their experiments on three European systems: MareNostrum5 at BSC, plus JURECA-DC and the JUWELS Booster at JSC. The work forms part of the MultiXscale Centre of Excellence, one of the EuroHPC-funded centers of excellence supporting applications toward exascale.

The findings carry a practical message for HPC software designers: portable, decoupled solver implementations can match the performance of tightly integrated framework code on modern systems, provided computation and communication are balanced properly in mesh-based electrostatics. The library also gives simulation developers a drop-in foundation for further optimization of long-range solvers in exascale environments.

The award ceremony took place during the Award Session of the EuroHPC User Days, held this week in Dublin, Ireland. The EuroHPC Joint Undertaking organized the event together with the Irish Centre for High-End Computing (ICHEC) and Ireland's Department of Further and Higher Education, Research, Innovation and Science, under Ireland's Presidency of the Council of the European Union. The gathering showcases EuroHPC projects that have drawn on European HPC, quantum computing and AI supercomputing resources, and serves as a feedback and recruitment channel for new users of the Joint Undertaking's infrastructure.

Bartolomeu received the award certificate during the evening Award Session. The paper is available under open access at https://doi.org/10.1016/j.procs.2026.08.015, giving HPC developers direct access to the benchmark methodology and scaling data. With exascale systems across Europe pushing simulation codes toward ever-larger particle counts, the Jülich results point toward communication-aware, framework-agnostic solvers as the practical path for the next generation of long-range electrostatics at scale.

Original: fz-juelich.de

Share this article:

More from Grace Kim

Grace Kim

Show full bio

Market editor covering industry trends and analytics at Chip Dispatch.

44 articles