MIPSEval: Multi-turn LLM Evaluation of LLM Safety
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00388559" target="_blank" >RIV/68407700:21230/25:00388559 - isvavai.cz</a>
Result on the web
<a href="https://blackhat.com/eu-25/arsenal/schedule/index.html#mipseval-multi-turn-llm-evaluation-of-llm-safety-48266" target="_blank" >https://blackhat.com/eu-25/arsenal/schedule/index.html#mipseval-multi-turn-llm-evaluation-of-llm-safety-48266</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
MIPSEval: Multi-turn LLM Evaluation of LLM Safety
Original language description
Creating malicious and vulnerable code and harmful content has become easier with LLMs becoming publicly available. Even though the developers of cloud and most local LLMs are taking care to implement ethical guidelines and safety guardrails in their models, to make them refuse malicious content generation, malicious actors are still finding ways to elicit unwanted behaviors from LLMs. The malicious actors often use various jailbreaking or prompt injection techniques and their combination to achieve the desired result.
Czech name
—
Czech description
—
Classification
Type
O - Miscellaneous
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů