Mistral introduced Shieldstral, a three-billion-parameter safety classifier under Apache 2.0. Instead of relying only on a fixed set of categories, it takes a policy expressed as a natural-language question and returns a safety score for text or images.

Context

Products serve different audiences and contexts, so a configurable policy is useful. It also becomes something that needs evaluation: ambiguous wording can produce inconsistent classifications. Teams should test both missed cases and unnecessary blocking against their own policy, rather than assuming a general benchmark captures every moderation decision.

Sources & authors

  1. Introducing Shieldstral. | Mistral
    Mistral AI · August 4, 2026