model roundup
GPT 5.5
-
Veracode ran 100+ models this year. Two numbers from that: - it compiles basically 100% of the time - it passes security 56% of the time, and that number has not moved since last year So if you don’t specifically ask for secure code, you g…
-
Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our controls using program analysis techniques such as inline ref…
-
TL;DR: Web search open router that routes agent web search across providers (Exa, Parallel, Tavily, Brave, SerpAPI etc.), and traces web search results with quality metrics, so you can visualize whether a bad agent run is a web search prob…