AlquistCoder: A Constitution-Guided Approach to Safe, Trustworthy Code Generation
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00384600" target="_blank" >RIV/68407700:21230/25:00384600 - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/68407700:21730/25:00384600
Výsledek na webu
<a href="https://assets.amazon.science/ea/07/ceb5bfb040a7a3cd5e9bc89466f7/alquistcoder-constitution-guided-approach-trustworthy-code-generation.pdf" target="_blank" >https://assets.amazon.science/ea/07/ceb5bfb040a7a3cd5e9bc89466f7/alquistcoder-constitution-guided-approach-trustworthy-code-generation.pdf</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
AlquistCoder: A Constitution-Guided Approach to Safe, Trustworthy Code Generation
Popis výsledku v původním jazyce
We introduce AlquistCoder, a code-generating system that effectively minimizes the risk of producing malicious content or vulnerable code while maintaining excellent Python coding and question answering standards across a wide range of tasks. The architecture of AlquistCoder employs a sophisticated input guardrail classifier that analyzes whether the user’s intention is benign, potentially harmful, or falls into a security-sensitive domain requiring special handling. Based on this classification, the system’s coding LLM receives an appropriately tailored system prompt and produces a contextually relevant response. This response is then evaluated by an output guardrail classifier to detect any security vulnerabilities that might have been introduced inadvertently. If problems are identified during this evaluation phase, the system automatically regenerates the answer until it meets our safety standards. Although several public datasets were used for training, we primarily utilized synthetically generated data. Our training methodology followed a multi-stage approach: we first aligned the model through supervised fine-tuning on high-quality examples and then further refined its capabilities using Direct Preference Optimization to enhance both code quality and safety aspects. Beyond architectural innovations, we introduce a novel data generation pipeline inspired by Constitutional AI and Constitutional Classifiers principles, resulting in a constitution-focused approach designed specifically for each stage of the training process.
Název v anglickém jazyce
AlquistCoder: A Constitution-Guided Approach to Safe, Trustworthy Code Generation
Popis výsledku anglicky
We introduce AlquistCoder, a code-generating system that effectively minimizes the risk of producing malicious content or vulnerable code while maintaining excellent Python coding and question answering standards across a wide range of tasks. The architecture of AlquistCoder employs a sophisticated input guardrail classifier that analyzes whether the user’s intention is benign, potentially harmful, or falls into a security-sensitive domain requiring special handling. Based on this classification, the system’s coding LLM receives an appropriately tailored system prompt and produces a contextually relevant response. This response is then evaluated by an output guardrail classifier to detect any security vulnerabilities that might have been introduced inadvertently. If problems are identified during this evaluation phase, the system automatically regenerates the answer until it meets our safety standards. Although several public datasets were used for training, we primarily utilized synthetically generated data. Our training methodology followed a multi-stage approach: we first aligned the model through supervised fine-tuning on high-quality examples and then further refined its capabilities using Direct Preference Optimization to enhance both code quality and safety aspects. Beyond architectural innovations, we introduce a novel data generation pipeline inspired by Constitutional AI and Constitutional Classifiers principles, resulting in a constitution-focused approach designed specifically for each stage of the training process.
Klasifikace
Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů