A 12-billion-parameter safety classifier from Meta designed to detect harmful content and enforce content policies in LLM inputs and outputs.