Unit 42 maps a thin safety layer inside aligned LLMs
Unit 42 says new research found that some aligned large language models may concentrate safety refusal behavior in very small sets of feed-forward neurons. The work introduces perturbation probing, a diagnostic designed to locate internal...
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.