Not established from the available sources.
MISTRAL AI · MODEL RELEASE TRACKER
Shieldstral
Shieldstral is a 3B-parameter open-weights multimodal safety classifier from Mistral AI. It frames content moderation as a policy-adaptive yes/no question-answering task, accepts plain-language policies at inference time, returns calibrated safety scores, and is released under Apache 2.0.CURRENT SNAPSHOT3/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
Modalities supportedText and images (multimodal evaluation of safety).
Performance vs larger modelsMatches models up to 7× its size on text safety and sets a new state of the art on multimodal moderation benchmarks (per the announcement).
Hardware efficiencyDesigned to run efficiently on a single 16GB NVIDIA GPU.
VERIFIABLE FACTS
Every value stays attached to a source and date.
RELEASE · Release and licenseDEVELOPER CLAIM
Released as open weights under the Apache 2.0 license and available for download.
CAPABILITIES · Model size and typeDEVELOPER CLAIM
3-billion-parameter multimodal safety classifier that frames content moderation as a policy-adaptive question-answering task and returns calibrated safety scores.
CAPABILITIES · Policy-adaptive inferenceDEVELOPER CLAIM
Accepts plain-language policies at inference time (policy supplied as a question/context) and does not require retraining to retarget to a new deployment context.
MODALITIES · Modalities supportedDEVELOPER CLAIM
Text and images (multimodal evaluation of safety).
BENCHMARKS · Performance vs larger modelsDEVELOPER CLAIM
Matches models up to 7× its size on text safety and sets a new state of the art on multimodal moderation benchmarks (per the announcement).
AVAILABILITY · Hardware efficiencyDEVELOPER CLAIM
Designed to run efficiently on a single 16GB NVIDIA GPU.
SAFETY · Safety interface and outputDEVELOPER CLAIM
At inference the model accepts an evaluation context and a yes/no question about safety, reads only the yes/no logits, and softmax-normalizes them into a continuous calibrated safety score (single-token verdict).
WHAT CHANGED