All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Assembly of FETI dual operator using CUDA

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989100%3A27740%2F25%3A10258738" target="_blank" >RIV/61989100:27740/25:10258738 - isvavai.cz</a>

  • Result on the web

    <a href="https://ieeexplore.ieee.org/document/11106042" target="_blank" >https://ieeexplore.ieee.org/document/11106042</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1109/IPDPSW66978.2025.00062" target="_blank" >10.1109/IPDPSW66978.2025.00062</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Assembly of FETI dual operator using CUDA

  • Original language description

    FETI is a numerical method used to solve engineering problems. It builds on the ideas of domain decomposition, which makes it highly scalable and capable of efficiently utilizing whole supercomputers. One of the most time-consuming parts of the FETI solver is the application of the dual operator F in every iteration of the solver.It is traditionally performed on the CPU using an implicit approach of applying the individual sparse matrices that form F right-to-left. Another approach is to apply the dual operator explicitly, which involves a simple dense matrix-vector multiplication and can be efficiently performed on the GPU. However, this requires additional preprocessing on the CPU where the dense matrix is assembled, which makes the explicit approach beneficial only after hundreds of iterations are performed.In this paper, we use the GPU to accelerate the assembly process as well. This significantly shortens the preprocessing time, thus decreasing the number of solver iterations needed to make the explicit approach beneficial.With a proper configuration, we only need a few tens of iterations to achieve speedup relative to the implicit CPU approach. Compared to the CPU-only explicit approach, we achieved up to 10× speedup for the preprocessing and 25× for the application.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    2025 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2025 : proceedings : 3-7 June 2025 Milan, Italy

  • ISBN

    979-8-3315-2644-3

  • ISSN

    2639-3867

  • e-ISSN

    2995-066X

  • Number of pages

    10

  • Pages from-to

    365-374

  • Publisher name

    IEEE

  • Place of publication

    Piscataway

  • Event location

    Milán

  • Event date

    Jun 3, 2025

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article