Assessing Explainability Methods for AI Safety Governance
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00383529" target="_blank" >RIV/68407700:21230/25:00383529 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.kuleuven.be/ethics-kuleuven/chair-ai/conference-ai-risks/xrisk-book_of_abstracts-1.pdf" target="_blank" >https://www.kuleuven.be/ethics-kuleuven/chair-ai/conference-ai-risks/xrisk-book_of_abstracts-1.pdf</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Assessing Explainability Methods for AI Safety Governance
Popis výsledku v původním jazyce
AI safety has not only been a matter of academic but also of public interest, culminating in a call for an AI moratorium in March 2023 (Future of Life Institute, 2023). Governments around the world have since taken action to promote the safe development of AI. For example, the UK, the US, and Japan have founded National AI Safety Institutes (AISIs), and other countries have followed since (Allen & Adamson, 2024). AISIs are tasked with the safety evaluation of advanced AI systems, contributing to standards with technical expertise, and strengthening international cooperation (Araujo et al., 2024). The EU AI Office has largely similar roles to AISIs, with the additional mandate to support the implementation and enforcement of the AI Act. The success of governance initiatives for AI safety depends on three components: i) specification of technical and legal means by which governance initiatives, such as the AI Act, shall be implemented; ii) stringent legal enforcement of regulation; iii) meaningful human oversight. However, safety criteria are impossible to fully specify in technical terms and to enforce at scale for current AI models that are deployed in dynamically changing, complex environments. We thus propose approaching i) standard specification, ii) auditing, and iii) continuous human oversight with explainable AI (XAI) methods (Doˇsilovi´c et al., 2018; Schwalbe & Finzel, 2024) that allow stakeholders to react flexibly as new safety concerns arise. Arguing that not all existing XAI methods are equally beneficial for AI safety, we review their usage in safety-related case studies based on AI Act classification and propose a 5-dimensional framework for their assessment. The first two dimensions are understandability to subjects and auditors. While (non-expert) subjects require simple and succinct justifications, auditors can be expected to process more involved explanations utilizing, e.g., their understanding of statistics. The third dimension is veracity, which measures how accurately an XAI method represents the model’s behavior. The penultimate dimension is actionability, which evaluates the helpfulness of an explanation in resolving possible undesired behaviors of an AI model. Finally, there is scalability, ensuring that even the largest state-of-the-art models can be explained with an XAI method. We suggest that a joint evaluation of the presented (possibly interdependent) dimensions is an essential part of a holistic approach to AI governance, bridging the gap between technical development and regulation.
Název v anglickém jazyce
Assessing Explainability Methods for AI Safety Governance
Popis výsledku anglicky
AI safety has not only been a matter of academic but also of public interest, culminating in a call for an AI moratorium in March 2023 (Future of Life Institute, 2023). Governments around the world have since taken action to promote the safe development of AI. For example, the UK, the US, and Japan have founded National AI Safety Institutes (AISIs), and other countries have followed since (Allen & Adamson, 2024). AISIs are tasked with the safety evaluation of advanced AI systems, contributing to standards with technical expertise, and strengthening international cooperation (Araujo et al., 2024). The EU AI Office has largely similar roles to AISIs, with the additional mandate to support the implementation and enforcement of the AI Act. The success of governance initiatives for AI safety depends on three components: i) specification of technical and legal means by which governance initiatives, such as the AI Act, shall be implemented; ii) stringent legal enforcement of regulation; iii) meaningful human oversight. However, safety criteria are impossible to fully specify in technical terms and to enforce at scale for current AI models that are deployed in dynamically changing, complex environments. We thus propose approaching i) standard specification, ii) auditing, and iii) continuous human oversight with explainable AI (XAI) methods (Doˇsilovi´c et al., 2018; Schwalbe & Finzel, 2024) that allow stakeholders to react flexibly as new safety concerns arise. Arguing that not all existing XAI methods are equally beneficial for AI safety, we review their usage in safety-related case studies based on AI Act classification and propose a 5-dimensional framework for their assessment. The first two dimensions are understandability to subjects and auditors. While (non-expert) subjects require simple and succinct justifications, auditors can be expected to process more involved explanations utilizing, e.g., their understanding of statistics. The third dimension is veracity, which measures how accurately an XAI method represents the model’s behavior. The penultimate dimension is actionability, which evaluates the helpfulness of an explanation in resolving possible undesired behaviors of an AI model. Finally, there is scalability, ensuring that even the largest state-of-the-art models can be explained with an XAI method. We suggest that a joint evaluation of the presented (possibly interdependent) dimensions is an essential part of a holistic approach to AI governance, bridging the gap between technical development and regulation.
Klasifikace
Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
S - Specificky vyzkum na vysokych skolach
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů