Codex Reset
AIニュース
GitHub

QQ·微信群

CODEX / SIGNAL STUDIO

Codex の動きを、ひと目で。

Mistral Large 4 outperforms Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks

Cline

Cline reported that Mistral Large 4 outperformed Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks. According to the team, the performance difference was largely driven by task refusal rates, with Mistral Large 4 rejecting far fewer security-related evaluation tasks.

In contrast, Opus 5.5 and GPT-6 Astra had approximately 40% of benchmark tasks blocked by their own safety filters during the testing.