Beneficial AGI: Care and Collaboration Are All You Need
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21730%2F24%3A00380938" target="_blank" >RIV/68407700:21730/24:00380938 - isvavai.cz</a>
Výsledek na webu
<a href="https://doi.org/10.1007/978-3-031-65572-2_9" target="_blank" >https://doi.org/10.1007/978-3-031-65572-2_9</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-031-65572-2_9" target="_blank" >10.1007/978-3-031-65572-2_9</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Beneficial AGI: Care and Collaboration Are All You Need
Popis výsledku v původním jazyce
This position paper conjectures that for beneficial AGI, it is necessary and sufficient for AGI systems to care about people and to employ goals whose success is collaboratively determined by the others involved in the situation. Moreover, I posit that any goal whose success can be determined without the consensual feedback of those concerned is likely to lead to the manifestation of dark factor traits. Integrating care reduces the risk that an AGI will be incentivized to seek harmful shortcuts to obtaining satisfactory feedback. Employing collaborative goals reduces the risk that an AGI will optimize for superficial features of success and proxy goals. Together, these ideas propose a fundamental shift away from the traditional control-centric “AI Safety” strategies. This paradigm not only promotes more beneficial outcomes but also enables AGIs to learn from and adapt to complex moral landscapes, thus continuously improving their capacity to contribute positively to the wellbeing of humans and other sentient beings.
Název v anglickém jazyce
Beneficial AGI: Care and Collaboration Are All You Need
Popis výsledku anglicky
This position paper conjectures that for beneficial AGI, it is necessary and sufficient for AGI systems to care about people and to employ goals whose success is collaboratively determined by the others involved in the situation. Moreover, I posit that any goal whose success can be determined without the consensual feedback of those concerned is likely to lead to the manifestation of dark factor traits. Integrating care reduces the risk that an AGI will be incentivized to seek harmful shortcuts to obtaining satisfactory feedback. Employing collaborative goals reduces the risk that an AGI will optimize for superficial features of success and proxy goals. Together, these ideas propose a fundamental shift away from the traditional control-centric “AI Safety” strategies. This paradigm not only promotes more beneficial outcomes but also enables AGIs to learn from and adapt to complex moral landscapes, thus continuously improving their capacity to contribute positively to the wellbeing of humans and other sentient beings.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
V - Vyzkumna aktivita podporovana z jinych verejnych zdroju
Ostatní
Rok uplatnění
2024
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Artificial General Intelligence, 17th International Conference, AGI 2024
ISBN
978-3-031-65571-5
ISSN
0302-9743
e-ISSN
1611-3349
Počet stran výsledku
5
Strana od-do
84-88
Název nakladatele
Springer
Místo vydání
Cham
Místo konání akce
SEATTLE
Datum konání akce
13. 8. 2024
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—