Mistral AI has unveiled ShieldStral, a lightweight multimodal safety model that classifies AI model inputs and outputs. SiliconANGLE reported on Aug. 5 that the open-weights-based model returns a yes-or-no decision along with a safety score when developers enter policies in natural language at runtime.
ShieldStral processes text and images together without additional retraining. Developers first provide high-level guidelines that set the purpose and criteria for the safety assessment, then enter the user query and the content to be checked.
Mistral AI said ShieldStral was trained to distinguish between similar but different criteria rather than memorising policy wording. For example, it can distinguish between whether something is related to malware and whether it is a cybersecurity discussion, the company said.
The company also stressed performance was high for a lightweight model. ShieldStral posted an average text safety score of 84.9 percent across various safety benchmarks, matching or exceeding models more than 7 times larger, it said. The average score on multimodal image safety benchmarks was 83.8 percent, outperforming all baseline models evaluated, it added.