When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.


From: ๐Ÿ”—https://huggingface.co/blog/security-incident-july-2026

Also: ๐Ÿ”—https://openai.com/index/hugging-face-model-evaluation-security-incident/


๐Ÿ˜Ž