Batched transpose-free ADI-type preconditioners for a Poisson solver on GPGPUs
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989100%3A27740%2F20%3A10243520" target="_blank" >RIV/61989100:27740/20:10243520 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.sciencedirect.com/science/article/pii/S0743731519307609" target="_blank" >https://www.sciencedirect.com/science/article/pii/S0743731519307609</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.jpdc.2019.11.004" target="_blank" >10.1016/j.jpdc.2019.11.004</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Batched transpose-free ADI-type preconditioners for a Poisson solver on GPGPUs
Popis výsledku v původním jazyce
We investigate the iterative solution of a symmetric positive definite linear system involving the shifted Laplacian as the system matrix on General Purpose Graphics Processing Units (GPGPUs). We consider in particular the Chebyshev iteration for its reduced global communication. The ADI-type preconditioner involves solving multiple (batched) symmetric positive tridiagonal Toeplitz systems along each coordinate direction. We investigate several variants how to solve these tridiagonal systems, the Thomas algorithm, the Thomas combined with the SPIKE algorithm, and a polynomial approximation of the inverse. We test the various implementations numerically by means of two- and three-dimensional examples. It turns out that a combination of the Thomas algorithm and the approximate inverse leads to a solution that does not need either tiling or transpositions. As such none of the kernels uses an extensive amount of shared memory which yields a very high GPU utilization and more importantly optimal coalesced global memory access patterns. (C) 2019 Elsevier Inc.
Název v anglickém jazyce
Batched transpose-free ADI-type preconditioners for a Poisson solver on GPGPUs
Popis výsledku anglicky
We investigate the iterative solution of a symmetric positive definite linear system involving the shifted Laplacian as the system matrix on General Purpose Graphics Processing Units (GPGPUs). We consider in particular the Chebyshev iteration for its reduced global communication. The ADI-type preconditioner involves solving multiple (batched) symmetric positive tridiagonal Toeplitz systems along each coordinate direction. We investigate several variants how to solve these tridiagonal systems, the Thomas algorithm, the Thomas combined with the SPIKE algorithm, and a polynomial approximation of the inverse. We test the various implementations numerically by means of two- and three-dimensional examples. It turns out that a combination of the Thomas algorithm and the approximate inverse leads to a solution that does not need either tiling or transpositions. As such none of the kernels uses an extensive amount of shared memory which yields a very high GPU utilization and more importantly optimal coalesced global memory access patterns. (C) 2019 Elsevier Inc.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
Výsledek vznikl pri realizaci vícero projektů. Více informací v záložce Projekty.
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2020
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Journal of Parallel and Distributed Computing
ISSN
0743-7315
e-ISSN
—
Svazek periodika
137
Číslo periodika v rámci svazku
March
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
12
Strana od-do
148-159
Kód UT WoS článku
000510315300012
EID výsledku v databázi Scopus
2-s2.0-85075774593