Contribution to a conference proceedings/Contribution to a book PUBDB-2026-02256

http://join2-wiki.gsi.de/foswiki/pub/Main/Artwork/join2_logo100x88.png
SIMD Vectorization of the Three-Body Axilrod-Teller-Muto Potential

 ;  ;

2026
IEEE
ISBN: 979-8-3195-3432-3

International Symposium on Parallel and Distributed Computing: proceedings, ISPDC, 25th conference, Hamburg, Germany, July 1-2, 2026
2026 25th International Symposium on Parallel and Distributed Computing, ISPDC 2026, HamburgHamburg, Germany, 1 Jul 2026 - 2 Jul 20262026-07-012026-07-02
IEEE : 1st, 39 - 47 () [10.1109/ISPDC69862.2026.00014]  GO

This record in other databases:  

Please use a persistent id in citations: doi:

Abstract: This work investigates the use of SIMD for the computation of three-body interactions in molecular dynamics. Our main focus are non-additive potentials in the form of the Axilrod-Teller-Muto (ATM). Given the high computational load required to calculate these interactions, the use of vectorization is of high relevance. Literature on the vectorization of the ATM potential is limited. In this paper, we propose two different techniques for the SIMD parallelization of these calculations, one based on register broadcast and another on permutation. These techniques are implemented using AVX2 intrinsic functions. Since in SIMD memory read/write operation management is critical to achieve a high performance, our work also explores the use of the two commonly found memory layouts in HPC applications, i.e. array-of-structures (AoS) and structure-of-arrays (SoA). These proposed vectorization techniques were tested at the node level, for which our simulations were implemented with shared-memory parallelization based on OpenMP. Our approach allows us to test the proposed parallelization schemes without the complications of more intricate molecule container implementations. With this purpose, the effect on the SIMD performance of other optimization aspects often used in molecular dynamics are described, such as the implementation of Newton's third law of motion. We present results that describe the achieved speedup and runtime gains for each technique. These were carried out for up to 64 cores in a single node and they illustrate the interplay between our proposed approaches, the use of shared-memory parallelization, and the described memory layouts. We were able to show a runtime speedup of up to 2.6× and up to 12× increase in performance with respect to the scalar calculations, without loss of performance at high core counts. A conclusion wraps up our studies.


Note: BMBF project 3xa, 16ME0653

Contributing Institute(s):
  1. Informationstechnologie (IT)
Research Program(s):
  1. 6G9 - IDAF (DESY) (POF4-6G9) (POF4-6G9)
Experiment(s):
  1. No specific instrument

Appears in the scientific report 2026
Click to display QR Code for this record

The record appears in these collections:
Private Collections > >DESY > >FH > >IT > IT
Document types > Events > Contributions to a conference proceedings
Document types > Books > Contribution to a book
Public records
Publications database


Linked articles:

http://join2-wiki.gsi.de/foswiki/pub/Main/Artwork/join2_logo100x88.png Book/Proceedings  ;  ;  ;  ;  ;  ;
International Symposium on Parallel and Distributed Computing: proceedings, ISPDC, 25th conference, Hamburg, Germany, July 1-2, 2026
2026 25th International Symposium on Parallel and Distributed Computing (ISPDC), ISPDC, HamburgHamburg, Germany, 1 Jul 2026 - 2 Jul 20262026-07-012026-07-02 IEEE ()  GO  Download fulltext Files BibTeX | EndNote: XML, Text | RIS


 Record created 2026-07-28, last modified 2026-07-30


Restricted:
Download fulltext PDF Download fulltext PDF (PDFA)
Rate this document:

Rate this document:
1
2
3
 
(Not yet reviewed)