Multiple Mean-Payoff Optimization Under Local Stability Constraints
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00141698" target="_blank" >RIV/00216224:14330/25:00141698 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1609/aaai.v39i25.34856" target="_blank" >http://dx.doi.org/10.1609/aaai.v39i25.34856</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1609/aaai.v39i25.34856" target="_blank" >10.1609/aaai.v39i25.34856</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Multiple Mean-Payoff Optimization Under Local Stability Constraints
Popis výsledku v původním jazyce
The long-run average payoff per transition (mean payoff) is the main tool for specifying the performance and dependability properties of discrete systems. The problem of constructing a controller (strategy) simultaneously optimizing several mean payoffs has been deeply studied for stochastic and game-theoretic models. One common issue of the constructed controllers is the instability of the mean payoffs, measured by the deviations of the average rewards per transition computed in a finite "window" sliding along a run. Unfortunately, the problem of simultaneously optimizing the mean payoffs under local stability constraints is computationally hard, and the existing works do not provide a practically usable algorithm even for non-stochastic models such as two-player games. In this paper, we design and evaluate the first efficient and scalable solution to this problem applicable to Markov decision processes.
Název v anglickém jazyce
Multiple Mean-Payoff Optimization Under Local Stability Constraints
Popis výsledku anglicky
The long-run average payoff per transition (mean payoff) is the main tool for specifying the performance and dependability properties of discrete systems. The problem of constructing a controller (strategy) simultaneously optimizing several mean payoffs has been deeply studied for stochastic and game-theoretic models. One common issue of the constructed controllers is the instability of the mean payoffs, measured by the deviations of the average rewards per transition computed in a finite "window" sliding along a run. Unfortunately, the problem of simultaneously optimizing the mean payoffs under local stability constraints is computationally hard, and the existing works do not provide a practically usable algorithm even for non-stochastic models such as two-player games. In this paper, we design and evaluate the first efficient and scalable solution to this problem applicable to Markov decision processes.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39 No. 25: AAAI-25 Technical Tracks 25
ISBN
9781577358978
ISSN
2159-5399
e-ISSN
2374-3468
Počet stran výsledku
8
Strana od-do
26551-26558
Název nakladatele
AAAI
Místo vydání
Palo Alto
Místo konání akce
Philadelphia, PA
Datum konání akce
1. 8. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001477487000040