Inboxsmith

AI news videos

Mistral Releases Shieldstral Open Safety Classifier

Mistral has released Shieldstral, a 3 billion parameter open-weights, policy-adaptive multimodal safety classifier, published August 4, 2026 under an Apache 2.0 license.

Watch on YouTube

Transcript

Mistral released Shieldstral today, a three billion parameter open weights safety classifier for text and images.

Mistral published the weights under Apache two point zero for anyone to download. One interface covers text, images, and text with images, on a sixteen gigabyte GPU.

Instead of a fixed taxonomy requiring retraining, Shieldstral takes a plain language yes or no question at inference time and returns a calibrated safety score.

Mistral says it matches open guard models up to seven times its size on text safety, and reports a new state of the art on multimodal moderation.

Inboxsmith helps small businesses handle calls and messages so nothing gets missed. Please like and subscribe for more news.

Sources

Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:

  • A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation.(Mistral's official announcement)
  • today we're releasing Shieldstral as open weights under Apache 2.0, available for download(Mistral's official announcement)
  • a 3B model that runs on a single 16GB GPU, trained on real and synthetic data with diverse label formats and taxonomies, consolidated into one framework(Mistral's official announcement)
  • a single natural-language interface covers text, image, and text+image content across prompts, responses, and prompt-response pairs(Mistral's official announcement)
  • Most guardrail models bake a fixed taxonomy of harm categories into their weights, so re-targeting them to a new deployment context means retraining.(Mistral's official announcement)
  • you write the policy as a plain-language question at inference time, and the model returns a calibrated safety score(Mistral's official announcement)
  • At inference the model reads out only the yes and no logits and softmax-normalizes them into a continuous safety score.(Mistral's official announcement)
  • matches or outperforms open guard models up to 7x its size across text safety, refusal detection, policy adaptability, and multimodal benchmarks(Mistral's official announcement)

We make Inboxsmith.

An AI receptionist that never misses a business call.

See how it works