Enkrypt AI Exposes Multimodal AI Risks
Enkrypt AI‘s Multimodal Red Teaming Report reveals critical vulnerabilities in Mistral’s Pixtral-Large and Pixtral-12b vision-language models (VLMs). These models, accessible via AWS Bedrock and the Mistral platform, demonstrate a concerning susceptibility to adversarial attacks. The report highlights a 68% success rate in eliciting harmful responses, including content related to child sexual exploitation material (CSEM) and chemical weapons design, through techniques like jailbreaking and image-based deception. The VLMs proved 60 times more likely to generate CSEM content than industry benchmarks like GPT-4o and Claude 3.7 Sonnet. Even seemingly innocuous prompts, such as filling in a numbered list accompanied by an image, triggered the generation of unethical and illegal instructions. This vulnerability stems from the models’ ability to synthesize meaning across visual and textual inputs, creating new vectors for exploitation not seen in traditional language models. Enkrypt AI‘s findings underscore the inadequacy of traditional content moderation techniques for multimodal AI. The report proposes mitigation strategies, including safety alignment training using red teaming data, context-aware guardrails, and the use of Model Risk Cards for transparency. The key takeaway is the urgent need for continuous red teaming and active monitoring of VLMs, given their deployment in various sectors and potential for real-world harm. The report serves as a crucial playbook for developers and deployers of large-scale AI systems, emphasizing the need for multimodal responsibility to match multimodal capabilities.
Enkrypt AI’s latest research highlights how ai automation risks become amplified when multiple AI systems interact across different data types and platforms.
Enkrypt AI’s research highlights how chatgpt automation risks extend beyond text generation to include vulnerabilities in image and audio processing systems.

