LRZ Hackathon Benchmarks GPU Vendors for Agentic AI Efficiency
More than 30 researchers benchmarked GPU vendors and llama.cpp backends at LRZ's Sept. 9-11 TPC hackathon, targeting energy efficiency for agentic AI inference and shaping the EU's AI-for-science roadmap.
- By
- Rebecca Stone
- Filed
- Channel
- AI & Compute
- Read
- 3 min read
More than 30 researchers from eight U.S. and European supercomputing centers convened at the Leibniz Supercomputing Centre (LRZ) from Sept. 9-11, 2026, where they benchmarked GPU vendors and llama.cpp backends for performance and energy efficiency across chat, reasoning, and vision-language model workloads.
The three-day event, organized by EuroTPC under the Trillion Parameter Consortium (TPC), drew specialists from Argonne National Laboratory, the Barcelona Supercomputing Centre (BSC), CSC – IT Center for Science in Finland, CINECA in Italy, Ludwig-Maximilians-Universität München (LMU Munich), the Munich Centre for Machine Learning (MCML), Leipzig University, and the High-Performance Computing Center Stuttgart (HLRS).
What hardware did the teams target?
Four working groups spread across the GPU-accelerated stack. Working group 2 (WG2) ran llama.cpp against multiple GPU vendors, testing how each backend handled chat, reasoning, and vision-language model inference to find the most power-efficient hardware-software-model combination.
WG3 designed cost-aware LLM evolutionary algorithms using multi-fidelity search, Bayesian optimization, and tiered model routing. Researchers then benchmarked the results on GPU kernel-generation tasks running against LRZ's HPC infrastructure. WG4 built an LLM-agent system that profiles HPC hardware, runs targeted micro-benchmarks, and recommends code optimizations grounded in real measurements.
Why does energy efficiency dominate the agenda?
Agentic AI, in which large language models coordinate with external tools to take actions, drives sustained inference calls across hardware. LRZ researcher Ajay Navilarekal Rajgopal stated the problem plainly: "The more Agentic AI is used, the more energy consumption increases due to inference calls. We need to limit that."
WG2's vendor benchmarking measured the same trade-off from the silicon side, comparing watts-per-token across GPU backends. A separate Gauss Supercomputing Centre (GCS) Agentic AI Focus group built secure authentication and authorization workflows for AI agents, citing risky system attacks and early incidents as motivation.
What did the science teams produce?
WG1 developed agentic workflows and implementations for the Pandemic Preparedness Engine (PPX), a project funded through the Coalition for Epidemic Preparedness Innovations (CEPI), and built hypothesis-generation pipelines for protein research.
A second GCS group created code-adaptation workflows that let an AI agent port software to various supercomputers, lowering the barrier to entry across heterogeneous HPC systems. An MCML team studied how researchers can use language models to draft and iterate on algorithms more quickly, while a separate group focused on risk prevention and pandemic preparedness.
Where does this lead next?
LRZ now leads drafting of an AI-for-science strategy roadmap for the European Union under the EuroTPC project. "We are currently gathering suggestions and topics from the EuroTPC's advisory board and from researchers, and are using these to draft a strategy roadmap for the EU," said Dr. Nicolay Hammer, head of LRZ's Big Data & AI (BDAI) team.
Prof. Charles Catlett, a senior computer scientist at Argonne and a TPC Planning and Strategy Team member, urged attendees to consider AI agents that analyze environmental sensor and camera data to safeguard urban environments. As more GPU vendors and inference frameworks enter the test matrix, future TPC hackathons will fold WG2's power-efficiency findings into the EU roadmap and sharpen the case for hardware-software co-design on next-generation accelerators.
Original: tpc.dev
More from Rebecca Stone
Show full bio
Correspondent covering media and advertising at Chip Dispatch.
250 articles
Related articles
carla-2026-latin-america-squeezes-more-from-the-machines-it-has-da7be648
CARLA 2026: Latin America Squeezes More From the Machines It Has
asus-ai-superbuild-puts-local-llms-on-a-172-tops-mini-pc-ef10affc
ASUS AI SuperBuild Puts Local LLMs on a 172-TOPS Mini-PC
cern-openlab-to-test-signaloid-s-uxhw-probability-compute-hardware-cfdce977
CERN openlab to Test Signaloid's UxHw Probability-Compute Hardware
coreweave-puts-nvidia-vera-rubin-nvl72-into-production-with-cognition-first-d77d4696
CoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production, With Cognition First


