Weight-Rounding Error in Deep Neural Networks
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F67985807%3A_____%2F25%3A00636916" target="_blank" >RIV/67985807:_____/25:00636916 - isvavai.cz</a>
Výsledek na webu
<a href="https://doi.org/10.1007/978-3-032-06078-5_23" target="_blank" >https://doi.org/10.1007/978-3-032-06078-5_23</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-032-06078-5_23" target="_blank" >10.1007/978-3-032-06078-5_23</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Weight-Rounding Error in Deep Neural Networks
Popis výsledku v původním jazyce
Current AI technologies based on deep neural networks (DNNs) are computationally extremely demanding, which limits their widespread deployment in embedded devices with constrained energy resources (e.g. battery-powered smartphones). One possible approach to solving this problem is to reduce the precision of weight parameters, which can save an enormous amount of energy for computation and data transfer at the cost of only a small loss in inference accuracy. In this paper, we provide a theoretical analysis of the effect of any weight rounding (e.g. reduced bitwidth) in a trained DNN on its output. We first derive a global upper bound on the output error of DNN (under the L1 norm) caused by the weight rounding for all inputs from a bounded domain in the worst case, which turns out to be overestimated for practical use. We prove that computing this maximum error is NP-hard for a given weight rounding even for two layers, which follows from the NP-hardness of neuron state domains. Based on the concept of so-called shortcut weights, we propose a method called AppMax that estimates this error using linear programming on convex polytopes around test/training data points, which works for any approximation of DNN (e.g. including pruning). The AppMax method was extensively tested on fully connected and convolutional neural networks (trained on the MNIST database) for decreasing bitwidth of weights. The experiments demonstrate a clear improvement in the error guarantees provided by this method, which can be used to evaluate different approximation strategies and identify those that best balance accuracy and energy efficiency
Název v anglickém jazyce
Weight-Rounding Error in Deep Neural Networks
Popis výsledku anglicky
Current AI technologies based on deep neural networks (DNNs) are computationally extremely demanding, which limits their widespread deployment in embedded devices with constrained energy resources (e.g. battery-powered smartphones). One possible approach to solving this problem is to reduce the precision of weight parameters, which can save an enormous amount of energy for computation and data transfer at the cost of only a small loss in inference accuracy. In this paper, we provide a theoretical analysis of the effect of any weight rounding (e.g. reduced bitwidth) in a trained DNN on its output. We first derive a global upper bound on the output error of DNN (under the L1 norm) caused by the weight rounding for all inputs from a bounded domain in the worst case, which turns out to be overestimated for practical use. We prove that computing this maximum error is NP-hard for a given weight rounding even for two layers, which follows from the NP-hardness of neuron state domains. Based on the concept of so-called shortcut weights, we propose a method called AppMax that estimates this error using linear programming on convex polytopes around test/training data points, which works for any approximation of DNN (e.g. including pruning). The AppMax method was extensively tested on fully connected and convolutional neural networks (trained on the MNIST database) for decreasing bitwidth of weights. The experiments demonstrate a clear improvement in the error guarantees provided by this method, which can be used to evaluate different approximation strategies and identify those that best balance accuracy and energy efficiency
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/GA25-15490S" target="_blank" >GA25-15490S: LEDNeCo: Nízkoenergetické hluboké neurovýpočty</a><br>
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Machine Learning and Knowledge Discovery in Databases. Research Track. ECML PKDD 2025 Proceedings, Part IV
ISBN
978-3-032-06077-8
ISSN
0302-9743
e-ISSN
—
Počet stran výsledku
19
Strana od-do
398-416
Název nakladatele
Springer
Místo vydání
Cham
Místo konání akce
Porto
Datum konání akce
15. 9. 2025
Typ akce podle státní příslušnosti
EUR - Evropská akce
Kód UT WoS článku
—