Session

300M Parameters vs the Agent: Why Security AI Should Run Locally

As AI agents gain access to tools, files, business systems and infrastructure, a common security pattern is emerging: use another large model to decide whether the first model's behaviour is safe.

We tried a different approach, placing a small, specialised security model close to the runtime boundary. That led to several uncomfortable engineering discoveries.

A static INT8 transformation left the model running while pushing benign false positives towards 90%. Keeping the embedding model's full 768-dimensional representation increased median inference latency by roughly 34% without improving primary threat detection. And a semantic-similarity signal that appeared highly confident turned out to flag 100% of both benign and malicious populations at its configured threshold.

This session explains the architecture that survived those failures: a narrow binary semantic decision, specialised classifier heads for context, deterministic rules where certainty exists, novelty signals for uncertainty, and larger-model adjudication only where a local model has demonstrably reached its limit.

The question is not whether small models can replace frontier models. It is which security decisions ever needed a frontier model in the first place.

Mukund Hirani

RAXE - AI Runtime Security

Dubai, United Arab Emirates

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top