model roundup

GPT 5.5

3 items · started 2026-08-27 · closed 2026-08-30

  1. Veracode ran 100+ models this year. Two numbers from that: - it compiles basically 100% of the time - it passes security 56% of the time, and that number has not moved since last year So if you don’t specifically ask for secure code, you g…

  2. Coding agents craft arbitrary code so securing them is more complicated than red-teaming. We post trained a cyber-security small llm, changed how it reasons and supplemented our controls using program analysis techniques such as inline ref…

  3. TL;DR: Web search open router that routes agent web search across providers (Exa, Parallel, Tavily, Brave, SerpAPI etc.), and traces web search results with quality metrics, so you can visualize whether a bad agent run is a web search prob…

← all threads