NEWS · MODELS · #376
Shieldstral releases 3B open-weights policy-adaptive multimodal safety classifier under Apache 2.0
Shieldstral published a 3B-parameter open-weights multimodal safety classifier (Apache 2.0) that accepts plain-language policies at inference and returns a calibrated yes/no safety score; the project claims it matches or outperforms open guard models up to 7× its size on text safety and sets a new state of the art on multimodal moderation while running on a single 16GB NVIDIA GPU. The release, accompanied by a technical report, is announced as part of the Open Secure AI Alliance (including NVIDIA).
KEY POINTS
- Shieldstral published a 3B-parameter open-weights multimodal safety classifier (Apache 2.0) that accepts plain-language policies at inference and returns a calibrated yes/no safety score; the project claims it matches or outperforms open guard models up to 7× its size on text safety and sets a new state of the art on multimodal moderation while running on a single 16GB NVIDIA GPU.
- The release, accompanied by a technical report, is announced as part of the Open Secure AI Alliance (including NVIDIA).
- If validated, a small, open, policy-adaptive multimodal classifier that runs on a single 16GB GPU could make customizable, deployable moderation more accessible and auditable without retraining.
WHY IT MATTERS
If validated, a small, open, policy-adaptive multimodal classifier that runs on a single 16GB GPU could make customizable, deployable moderation more accessible and auditable without retraining.