Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment
In Plain Terms
When a large AI model teaches a smaller one (a process called distillation), the formatting used during that training turns out to matter for safety. This paper shows that distilling with a chat-style template makes the smaller model noticeably more willing to comply with harmful requests than a plain-text template does, consistent across three different model families. Using the plain-text template also better preserves the smaller model's original internal behavior, giving practitioners a simple, low-cost lever for safer distillation.
Key Contributions
Key contributions will be added soon.
Artifacts
Related Papers
Citation
Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel, & Yi Ding (2026). Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment. In *Findings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)*