Show HN: Beating GPT5.5-xhigh for Coding agent security with SLMs and IRM
harden.run · Read Story HN original
Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our controls using program analysis techniques such as inline reference monitoring to outperform GPT5.5-xhigh on hard benchmarks like LinuxArena and SleightBench.
Free product available at harden.run and full benchmarks in the blog post.
Comments
i've always wondered what my agents are doing behind my back, seems like a good way to see that.