Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Assessing Explainability Methods for AI Safety Governance

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00383529" target="_blank" >RIV/68407700:21230/25:00383529 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://www.kuleuven.be/ethics-kuleuven/chair-ai/conference-ai-risks/xrisk-book_of_abstracts-1.pdf" target="_blank" >https://www.kuleuven.be/ethics-kuleuven/chair-ai/conference-ai-risks/xrisk-book_of_abstracts-1.pdf</a>

  • DOI - Digital Object Identifier

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Assessing Explainability Methods for AI Safety Governance

  • Popis výsledku v původním jazyce

    AI safety has not only been a matter of academic but also of public interest, culminating in a call for an AI moratorium in March 2023 (Future of Life Institute, 2023). Governments around the world have since taken action to promote the safe development of AI. For example, the UK, the US, and Japan have founded National AI Safety Institutes (AISIs), and other countries have followed since (Allen & Adamson, 2024). AISIs are tasked with the safety evaluation of advanced AI systems, contributing to standards with technical expertise, and strengthening international cooperation (Araujo et al., 2024). The EU AI Office has largely similar roles to AISIs, with the additional mandate to support the implementation and enforcement of the AI Act. The success of governance initiatives for AI safety depends on three components: i) specification of technical and legal means by which governance initiatives, such as the AI Act, shall be implemented; ii) stringent legal enforcement of regulation; iii) meaningful human oversight. However, safety criteria are impossible to fully specify in technical terms and to enforce at scale for current AI models that are deployed in dynamically changing, complex environments. We thus propose approaching i) standard specification, ii) auditing, and iii) continuous human oversight with explainable AI (XAI) methods (Doˇsilovi´c et al., 2018; Schwalbe & Finzel, 2024) that allow stakeholders to react flexibly as new safety concerns arise. Arguing that not all existing XAI methods are equally beneficial for AI safety, we review their usage in safety-related case studies based on AI Act classification and propose a 5-dimensional framework for their assessment. The first two dimensions are understandability to subjects and auditors. While (non-expert) subjects require simple and succinct justifications, auditors can be expected to process more involved explanations utilizing, e.g., their understanding of statistics. The third dimension is veracity, which measures how accurately an XAI method represents the model’s behavior. The penultimate dimension is actionability, which evaluates the helpfulness of an explanation in resolving possible undesired behaviors of an AI model. Finally, there is scalability, ensuring that even the largest state-of-the-art models can be explained with an XAI method. We suggest that a joint evaluation of the presented (possibly interdependent) dimensions is an essential part of a holistic approach to AI governance, bridging the gap between technical development and regulation.

  • Název v anglickém jazyce

    Assessing Explainability Methods for AI Safety Governance

  • Popis výsledku anglicky

    AI safety has not only been a matter of academic but also of public interest, culminating in a call for an AI moratorium in March 2023 (Future of Life Institute, 2023). Governments around the world have since taken action to promote the safe development of AI. For example, the UK, the US, and Japan have founded National AI Safety Institutes (AISIs), and other countries have followed since (Allen & Adamson, 2024). AISIs are tasked with the safety evaluation of advanced AI systems, contributing to standards with technical expertise, and strengthening international cooperation (Araujo et al., 2024). The EU AI Office has largely similar roles to AISIs, with the additional mandate to support the implementation and enforcement of the AI Act. The success of governance initiatives for AI safety depends on three components: i) specification of technical and legal means by which governance initiatives, such as the AI Act, shall be implemented; ii) stringent legal enforcement of regulation; iii) meaningful human oversight. However, safety criteria are impossible to fully specify in technical terms and to enforce at scale for current AI models that are deployed in dynamically changing, complex environments. We thus propose approaching i) standard specification, ii) auditing, and iii) continuous human oversight with explainable AI (XAI) methods (Doˇsilovi´c et al., 2018; Schwalbe & Finzel, 2024) that allow stakeholders to react flexibly as new safety concerns arise. Arguing that not all existing XAI methods are equally beneficial for AI safety, we review their usage in safety-related case studies based on AI Act classification and propose a 5-dimensional framework for their assessment. The first two dimensions are understandability to subjects and auditors. While (non-expert) subjects require simple and succinct justifications, auditors can be expected to process more involved explanations utilizing, e.g., their understanding of statistics. The third dimension is veracity, which measures how accurately an XAI method represents the model’s behavior. The penultimate dimension is actionability, which evaluates the helpfulness of an explanation in resolving possible undesired behaviors of an AI model. Finally, there is scalability, ensuring that even the largest state-of-the-art models can be explained with an XAI method. We suggest that a joint evaluation of the presented (possibly interdependent) dimensions is an essential part of a holistic approach to AI governance, bridging the gap between technical development and regulation.

Klasifikace

  • Druh

    O - Ostatní výsledky

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

    S - Specificky vyzkum na vysokych skolach

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů