Proteus reports over 99 percent phishing recall on a one-week AegisAI sample
AegisAI's Proteus identified more than 99% of 3,020 held-out phishing attacks, while a closed-source model missed 565 and an open-weight model missed 604.
| Proteus | |
|---|---|
| phishing detection | more than 99% |
| sample size | 3,020 |
| median speed vs closed-source | 7.2 times faster |
| slowest 5% speed | 26 times faster |
| expected calibration error | 0.037 |
| confidence band | 82% |
| ablation gain | 8.4% more phishing |
Proteus identified more than 99% of 3,020 human-reviewed phishing attacks in AegisAI's benchmark. No figure appears on the page, only a text table of misses and speed ratios. Closed-source model missed 565, open-weight model missed 604, and Proteus missed far fewer. One week of AegisAI production traffic supplied the held-out sample.
It treats an email as a structured object: message text, headers, sender and recipient history, sanitized link and attachment data, and threat intelligence. A custom tokenizer and a single inference pass without sampling support that design. Ablation results say the same model found 8.4% more phishing when supplied with evidence at inference. Evidence outside the attacker's control, not pattern matching, is the core mechanism.
Die Brief reads this as a recall claim, not a full deployment proof. Built for a three-way inline verdict, the model makes false-positive cost and tail latency the buyer's real constraint, not just missed phishing. AegisAI says false-positive data will come later, which leaves the most expensive failure mode unpriced.
Missing are false-positive rates, absolute latency, price, hardware footprint, and independent testing. The page also does not show how Proteus behaves on benign mail, spoofed but legitimate accounts, or multi-step attacks that require human follow-up. Those gaps matter because email security runs inline and every extra second or false alarm is paid by the user.
The benchmark proves recall on a one-week AegisAI sample, not enterprise-wide false-positive behavior or worst-case latency. The buyer's constraint is that inline email security must hold the message, so speed and false positives matter more than recall alone.
The document does not state false-positive rates, absolute latency, price, deployment footprint, or independent validation.
The numbers show high recall on 3,020 held-out phishing messages and faster median verdicts, but they do not establish precision, cost, or production tail latency. The architecture trades general-purpose flexibility for a narrow three-way inline decision.
After prnewswire.com. We did not report this. The pictures, if any, are theirs.
