Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

MIPSEval: Multi-turn LLM Evaluation of LLM Safety

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00388559" target="_blank" >RIV/68407700:21230/25:00388559 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://blackhat.com/eu-25/arsenal/schedule/index.html#mipseval-multi-turn-llm-evaluation-of-llm-safety-48266" target="_blank" >https://blackhat.com/eu-25/arsenal/schedule/index.html#mipseval-multi-turn-llm-evaluation-of-llm-safety-48266</a>

  • DOI - Digital Object Identifier

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    MIPSEval: Multi-turn LLM Evaluation of LLM Safety

  • Popis výsledku v původním jazyce

    Creating malicious and vulnerable code and harmful content has become easier with LLMs becoming publicly available. Even though the developers of cloud and most local LLMs are taking care to implement ethical guidelines and safety guardrails in their models, to make them refuse malicious content generation, malicious actors are still finding ways to elicit unwanted behaviors from LLMs. The malicious actors often use various jailbreaking or prompt injection techniques and their combination to achieve the desired result.

  • Název v anglickém jazyce

    MIPSEval: Multi-turn LLM Evaluation of LLM Safety

  • Popis výsledku anglicky

    Creating malicious and vulnerable code and harmful content has become easier with LLMs becoming publicly available. Even though the developers of cloud and most local LLMs are taking care to implement ethical guidelines and safety guardrails in their models, to make them refuse malicious content generation, malicious actors are still finding ways to elicit unwanted behaviors from LLMs. The malicious actors often use various jailbreaking or prompt injection techniques and their combination to achieve the desired result.

Klasifikace

  • Druh

    O - Ostatní výsledky

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů