Mistral introduced Shieldstral, a three-billion-parameter safety classifier under Apache 2.0. Instead of relying only on a fixed set of categories, it takes a policy expressed as a natural-language question and returns a safety score for text or images.
Context
Products serve different audiences and contexts, so a configurable policy is useful. It also becomes something that needs evaluation: ambiguous wording can produce inconsistent classifications. Teams should test both missed cases and unnecessary blocking against their own policy, rather than assuming a general benchmark captures every moderation decision.
Sources & authors
- Introducing Shieldstral. | MistralMistral AI · August 4, 2026



