Mistral Large 4 outperforms Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks
Cline reported that Mistral Large 4 outperformed Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks. According to the team, the performance difference was largely driven by task refusal rates, with Mistral Large 4 rejecting far fewer security-related evaluation tasks.
In contrast, Opus 5.5 and GPT-6 Astra had approximately 40% of benchmark tasks blocked by their own safety filters during the testing.