Speaker
Description
QUDA is a widely used library of GPU-optimised kernels for lattice QCD whose highly tuned Dirac-operator solvers can accelerate propagator inversions, which usually dominate the overall calculation time. However, not every fermion action maps cleanly onto its interface: the Stout-Link Non-perturbative Clover (SLiNC) fermion action uses distinct gauge fields for the Wilson hopping term (stout-smeared $fat$ links) and the clover term (unsmeared $thin$ links), preventing its direct mapping onto QUDA's single-gauge-field interface. Feynman–Hellmann propagator calculations compound this by introducing a site-dependent, momentum-projected current perturbation into the Dirac operator that lies entirely outside QUDA's solver kernels. We present a hybrid QDP–QUDA strategy implemented in Chroma that addresses both obstacles without modifying the highly optimised QUDA GPU routines. We present computational benchmarks using an NVIDIA A100-40GB cluster that show significant speed-ups of this approach in comparison to non-QUDA based GPU routines.