System card: https://deploymentsafety.openai.com/gpt-5-6-preview
#gpt-5
634 items
Previewing GPT‑5.6 Sol: a next-generation model (openai.com via hn) GPT5.5 slightly outperformed Mythos on a multi-step cyber-attack simulation. One challenge that took a human expert 12 hrs took GPT-5.5 only 11 min at a $1.73 cost (www.reddit.com) Link to tweets: https://x.com/deredleritt3r/status/2049890601236390098?s=20 https://x.com/AISecurityInst/status/2049868227740565890?s=20 Link to associated blogs: https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabil…
GPT-5.6 used a prompt to close a 30-year gap in convex optimization (old.reddit.com via hn) could not extract summary
Caught the massive OpenAI Codex model leak on video before it was patched! (GPT-5.5, Arcanine, Glacier-alpha) (www.reddit.com) Hey everyone, I opened up Codex today and was greeted by this massive list of unreleased and internal models. I managed to get a screen recording of the dropdown right before OpenAI seemingly realized the mistake and patched it out.
GPT-5.5's Unicorn (www.reddit.com) could not extract summary
OpenAI has truly stepped up their game and released some great models in the last three months. (www.reddit.com) I know it’s rare to see a positive post about OpenAI on Reddit, so this often goes overlooked. After the GPT-5 fiasco and all the drama, it feels like they’re finally on the right track.
Gemini 3.2 Flash is capable of solving IMO 2025 P6. Only GPT-5.5-Pro can solve it currently without any scaffolding / harness engineering. (www.reddit.com) could not extract summary
GPT-5.5's SimpeBench scores are out (www.reddit.com) Source: https://simple-bench.com/
🔥BREAKING: OpenAI rolls out GPT-5.4-Cyber to limited group for testing, seeks to rival Claude Mythos (www.reddit.com) OpenAI has officially announced GPT-5.4-Cyber today as part of an expanded Trusted Access for Cyber Defense program. OpenAI describes it as a version of GPT-5.4 that is tuned for legitimate cybersecurity work, with a lower refusal boundary…
DeepSeek V4 Pro beats GPT-5.5 Pro on precision (runtimewire.com via hn) DeepSeek V4 Pro takes this matchup 38.0 to 33.0, and the margin feels earned. Across the scored tasks, the pattern is simple: Model A was tighter, more literal, and more reliable under constraints, while Model B was good but a little too w…
Accelerating GPT-5.6 Sol Ultrafast (www.cerebras.ai via hn) Today, Cerebras and OpenAI are sharing an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast is available initially to a select group of customers, with access expanding ov…
GPT-5.5 autonomously spent 150+ hours improving protein folding models. (www.reddit.com) https://x.com/chrishayduk/status/2055757345506877759?s=46
Kimi K2.6 vs. GPT-5.4 (xhigh) - When will the new OpenAI model be released? This Thursday? (www.reddit.com) Ever since the new $100 Pro plan, they now claim there's a "dynamic usage limits" that can become restricted at anytime, and not reset for indefinitely as long as they deem it "appropriate" (www.reddit.com) U.S. government will decide who gets to use GPT-5.6 (www.washingtonpost.com via hn) https://archive.ph/PCQQl
Qwen 3.8 follows GPT-5.5 Pro reasoning prefills (gist.github.com via hn) A follow-up to Reasoning prefills on a few open models and Stolen Thoughts This v1.1 reruns the reasoning-prefill experiment with GPT-5.5 Pro as the teacher. For each problem, I generated two responses from each target model: - an ordinary…
GPT-5.5 was used to flag fatal errors in FrontierMath problems (www.reddit.com) FrontierMath is supposed to be one of the hard benchmarks for frontier models, and now Epoch is saying an AI-assisted review found fatal errors in about a third of Tiers 1-4. Noam Brown says the initial flags came from GPT-5.5.
On a difficult new SWE benchmark, ProgramBench, GPT5.5 high/xhigh solves a task for first time, significantly outperforms Opus 4.7 (www.reddit.com) Link to tweets: https://x.com/KLieret/status/2054215545663144217?s=20 Link to GitHub: https://github.com/facebookresearch/ProgramBench/ Link to ProgramBench website: https://programbench.com/blog/gpt-5-5-first-solve/
GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review? (entelligence.ai via hn) GPT-5.6 Luna vs GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review? GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens.
GPT-5.5 improves over GPT-5.4 and overtakes Opus 4.6 to take the 2nd place behind Gemini 3.1 Pro on the Extended NYT Connections Benchmark (www.reddit.com) GPT-5.5: xhigh: 94.0→97.5 high: 93.6→96.9 medium: 92.0→95.0 no reasoning: 32.8→37.5 Kimi K2.6 improves over Kimi K2.5 (78.3→91.4) and becomes the #1 open weights model. DeepSeek V4 Pro improves over DeepSeek V3.2 (50.2→75.7).
I’ve used enough AI models to realize they all have wildly different personalities At this point I’m convinced AI models are just coworkers with different levels of talent, ego, and criminal energy. (www.reddit.com) - Claude Opus 4.6 - absolute rogue AI. Does what I want like it’s breaking at least 3 internal policies to make it happen.
First time ever hitting a limit on the new $100 Pro plan for the Pro model (www.reddit.com) It's clearly meant to be unlimited. And I'm definitely not abusing it, just using it extensively.
Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge (thinkpol.ca via hn) By Rohana Rezel I’m running the ongoing AI Coding Contest where I pit major language models against each other in real-time programming tasks with objective scoring. Day 12 was the Word Gem Puzzle.
GPT-5.5 Instant is starting to roll out in ChatGPT. (www.reddit.com) could not extract summary
"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok (www.tryai.dev via hn) "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok Four frontier models, a blank canvas, and colored pencils. We tracked every stroke, dollar, and output as they tried to draw the Mona Lisa.
Decreased Intelligence Density in DeepSeek V4 Pro (www.reddit.com) In the V3.2 paper, they mentioned: Second, token efficiency remains a challenge; DeepSeek-V3.2 typically requires longer generation trajectories (i.e., more tokens) to match the output quality of models like Gemini 3.0-Pro. Future work wil…
LLMs do fine on ARC-AGI-3 if they are allowed to search over game logs (www.reddit.com) I was reading the comments to this post and the overall opinion seemed to be that harness makes little/no difference for ARC-AGI-3. Turns out, it makes a huge difference: Hill-climbing ARC-AGI-3 TLDR: if you save game logs - taken actions,…
Page 15 of the GPT-5.5 System Card: " Our analysis estimates that GPT-5.5 is slightly more misaligned than GPT-5.4 Thinking across several categories, though nearly all of this is low-severity misalignment. " (www.reddit.com) https://deploymentsafety.openai.com/gpt-5-5/gpt-5-5.pdf
FrontierMath: Opus 4.7 improves over Opus 4.6 and Gemini 3.1 but still trails GPT-5.4-xHigh and GPT-5.4-Pro (www.reddit.com) could not extract summary
12M Context Window and some some sprinkle of lies? (www.reddit.com) Spent some time on the SubQ launch today. Some things don't line up.
GPT 5.5 "secret sauce" is just having the thinking be some stupid caveman mode? (www.reddit.com) I think I had GPT-5.5 leak its trace during a normal conversation, and it really reads like the caveman mode fad from a few months back. Maybe we can achieve better token efficiency by taking some high-quality thinking trace from an open m…
We benchmarked TranslateGemma-12b against 5 frontier LLMs on subtitle translation - it won across the board, with one significant catch (www.reddit.com) As part of our ongoing translation quality research at Alconost, we put six models through subtitle translation into six language pairs. At first glance the numbers told a clean story.
Is the AI subscription bubble starting to crack? GPT-5.5 just dropped, prices keep rising, and the “all-you-can-eat” era looks more fake by the month (www.reddit.com) GPT-5.5 just launched, and the pricing is hard to defend. OpenAI’s API pricing now puts GPT-5.5 at $5 / 1M input tokens and $30 / 1M output tokens, while GPT-5.4 is $2.50 / $15.
Just got an email announcing GPT-5.3-Codex-Spark (www.reddit.com) Just got this e-mail from OpenAI, two months too late. I hope they mean March, 20th 2027.
All major LLMs are lib-left. Even Grok, half the time (unslop.run via hn) I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant.
ARC-AGI-3 Update (GPT-5.5 High and Opus4.7) (www.reddit.com) - GPT-5.5: 0.43% - Opus 4.7: 0.18% ARC-AGI-3 is no joke. I can’t wait to see which models finally crack.
DeepSeek V4 isn't beating Opus, but it doesn't need to (www.reddit.com) DeepSeek V4 is not in the same league as GPT-5.5 or Opus 4.7. Benchmarks put it slightly below both of those, roughly on par with Opus 4.6.
Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help? (charlesazam.com via hn) I gave Claude Fable 5 and GPT-5.6 Sol the same unpublished NP-hard optimization problem, with and without their native /goal mode. Fable 5 is a beast; /goal is not a game changer.
Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, largely because it is one of the most cautious models. (www.reddit.com) Does a model maintain the same judgment or does it side with whoever is speaking? This benchmark measures that inconsistency directly.
Running gpt and glm-5.1 side by side. Honestly can’t tell the difference (www.reddit.com) So I have been running gpt and glm-5.1 side by side lately and tbh the gap is way smaller than what im paying for On SWE-Bench Pro glm-5.1 actually took the top spot globally, beat gpt-5.4 and opus 4.6. overall coding score is like 55 vs g…
GPT-5.5 is lowkey blowing my mind (www.reddit.com) Just spent the whole morning testing GPT-5.5 in ChatGPT and the jump in agentic reasoning and complex task handling is ridiculous.It plans multi-step workflows, uses tools properly, checks its own work, and actually gets stuff done instead…
UPDATE: The method from the proof generated by GPT-5.4 Pro for Erdos Problem #1196 was successfully applied to other problems including another 60 year old Erdos conjecture. (www.reddit.com) Link to tweet: https://x.com/jdlichtman/status/2050460077904285789 Links for the talks: https://m.youtube.com/@FoMathematics?ra=m https://events.stanford.edu/event/future-of-mathematics-symposium Link to original post about problem #1196:…
Why did OpenAI stop releasing “chat” api models? (www.reddit.com) I have built an AI Assistant and since last year I have been upgrading the internal LLM from through gpt-5.3-chat but since 5.4 they stopped rolling the chat api. This is my app Sweezy she uses gpt-5.3-chat and in the conversation, you can…
Top open weight models like ds v4 pro max are still like 6-7 months if not more behind closed lab models (www.reddit.com) The best open weight and/or non -American models like Deepseek v4 pro max and kimi k2.6 are still like 3-7 months if not more behind closed lab models .. From ds's technical report- P5-"Nevertheless, its performance falls marginally short…
Construction Spending on Data Centers Again Outpaces Office Construction (www.reddit.com) The Federal Construction Spending Report for Feb and March 2026 was released today by the Census Bureau. It shows that data center construction spending is again higher than office spending, and the gap is still widening.
New LLM Position Bias Benchmark: does an LLM keep the same judgment when you swap the answer order? Judge models compare two lightly edited versions of the same story twice, with the order swapped. The median model flips in 45% of decisive case pairs. GPT-5.4 is worst at 66%. (www.reddit.com) More info, including charts, per-case metrics, raw judge outputs, and the parsed answer dump: https://github.com/lechmazur/position_bias This benchmark isolates one basic and frustrating failure mode. The model-average first-shown pick rat…
OpenAI cuts developer pricing for frontier GPT-5.6 Sol model by more than 20% (www.reuters.com via hn) could not extract summary
GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday (twitter.com via hn) GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.
I stumbled on a Gemma 4 chat template bug for tools and fixed it (www.reddit.com) TLDR: tool parameters using the common JSON Schema pattern `anyOf: [$ref, null]` are rendered into the prompt as empty `type` fields. This strips the useful schema information before the model sees it.
GPT 5.5 outperforming Opus 4.7 on ProgramBench (www.reddit.com) When we released ProgramBench last week, we hadn't included GPT 5.5 yet because it came out after we frozen model selections for our NeurIPS submission. Honestly super surprised how well it does.
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index (artificialanalysis.ai via hn) September 3, 2026 GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this…
I tested GPT-5.5 Codex against Opus 4.7 Claude Code, and it's about time Anthropic bros take pricing seriously. (www.reddit.com) I've used Claude Code the most among AI coding agents. Sonnet, Opus, I've run them all.
Am I missing something about GPT-5.5 efficiency? (www.reddit.com) OpenAI said GPT-5.5 was supposed to be more cost-efficient, but this Artificial Analysis chart seems to show Codex + GPT-5.5 using more tokens than Codex + GPT-5.4. GPT-5.5 is around 2.8M tokens per task, while GPT-5.4 is around 2.5M in th…
Parameter Estimate (www.reddit.com) The estimate seems quite accurate. Many people have noticed a drop in quality with GPT-5.1, GPT-5.2, GPT-5.3, and Opus 4.7.
Even Sama himself doesn’t believe GPT-5.5 matches Opus 4.7 design capabilities. AI race will humble you (www.reddit.com) could not extract summary
Beating GPT-5.6 Sol on retrieval with 100x cheaper open models (neon.com via hn) “Most teams' best training data is just sitting in their databases. The problem is that turning raw data into something usable is hard, and letting agents read, search, and mutate data cheaply at scale requires advanced infra.
HalBench: I built a custom sycophancy and hallucination benchmark and tested 4 frontier models (Sonnet 4.6, Grok 4.3, GPT 5.4 and Gemini 3.1 Pro), looking for input on what OSS models to run next! (www.reddit.com) HalBench Results: TL;DR: I built HalBench, an open benchmark for LLM sycophancy and hallucination. 3,200 false-premise prompts × 4 models = 12,800 graded responses.
Agentic harness for theoretical physics research (www.reddit.com) Hi everyone, at Hugging Face we've been developing agentic harnesses for various domains and today we're releasing physics-intern to tackle research-level problems in theoretical physics. It's a multi-agent framework which we designed to m…
Claude Fable 5 vs. GPT-5.5: Better Planning, Similar Execution (blog.kilo.ai via hn) Claude Fable 5 vs GPT-5.5: better planning, similar execution Update: We wrote this post on June 11 and published it on June 13. Anthropic has since disabled access to Claude Fable 5 after a US government directive, which makes some of the…
Buyout Game Benchmark: 8 models play a social strategy game with public balances, private transfers, messaging, eliminations, deals, defections, and a final buyout phase. 804 games. GPT-5.5 is the champion. Opus 4.7 performs well. (www.reddit.com) This benchmark measures long-horizon social strategy under explicit financial incentives. Eight models play a multi-round elimination game with unequal starting balances, a public prize ladder, private transfers, public votes, and a finali…
Devs using Qwen 27B seriously, what's your take? (www.reddit.com) For developers using Qwen 27B for coding, Codex style: what's your honest take? So far, for me, it's been pretty solid.
Pen-Testing Company XBOW on GPT-5.5: Mythos-like Cyber-Sec (www.reddit.com) Read their full article here: XBOW - GPT-5.5: Mythos-Like Hacking, Open To All For the ones asking what this chart shows: It's how many True Positive threats a model generates for each False Negative. Given a code base (white box) GPT-5.5…
I tried adding rich UI elements to Open WebUI (www.reddit.com) so i tried adding openui to openwebui and it worked pretty well. used it with gpt-5.4-mini and it was super fast and responsive.
GPT 5.6 Sol is the best "vision" model OpenAI ever released (blog.roboflow.com via hn) Last week, OpenAI announced the GPT-5.6 lineup, introducing the Sol, Terra, and Luna models. During the release stream, the team focused heavily on computer use, showing models capable of navigating and operating desktop applications.
Cursor $60 with Composer 2.5 vs Codex $100 with GPT-5.5 Medium for daily coding? (www.reddit.com) I'm trying to decide which setup is more comfortable for sustained weekday coding. Assumptions: Usage: around 6 hours per weekday Cursor: $60 plan, using only Composer 2.5 Codex: $100 plan, using only GPT-5.5 Medium Main goal: coding with…
Opus 4.7 Low Vs Medium Vs High Vs Xhigh Vs Max: the Reasoning Curve on 29 Real Tasks from an Open Source Repo (www.reddit.com) TL;DR I ran Opus 4.7 in Claude Code at all reasoning effort settings (low, medium, high, xhigh, and max) on the same 29 tasks from an open source repo (GraphQL-go-tools, in Go). On this slice, Opus 4.7 did not behave like a model where mor…
Did the $100 Plan Affect the GPT-5.4 Pro Model? (www.reddit.com) Most people are focused on the changes in the usage limits of Codex with the new Pro and Plus plans, but has anyone experienced changes to the Pro model on ChatGPT using the $200 vs $100 plan? I used to use the $200 Pro plan and used the P…
GPT-5.6 Sol vs. Kimi K3 Speedrunning Kerbal Space Program Live (www.twitch.tv via hn) GPT-5.6 Sol 🇺🇸US vs Kimi K3 🇨🇳 China on Kerbal Space Program - Vals AI Time Horizon Index
GPT-5.2 matches top human reviewers in Nature peer review study (www.reddit.com) 45 scientists spent 469 hours comparing human and AI reviews across 82 papers. AI reviewers held their own against top-rated human reviewers, though with some weaknesses.
Fields medal-winning mathematician says GPT-5.5 is now solving open math problems at PhD-thesis level: "We will face a crisis very soon." (www.reddit.com) blog-post: https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/
Cursor is great but the monthly limits kill it for me (www.reddit.com) When did we go from 400k to 256k? (www.reddit.com) GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf] (cdn.openai.com via hn) could not extract summary
Altman: GPT-5.6 is 54% more token efficient on agentic coding (www.cnbc.com via hn) OpenAI CEO Sam Altman told CNBC on Thursday that GPT-5.6 Sol, the company's latest artificial intelligence model, is 54% more token efficient on agentic coding tasks, and that it's "as good or better" than competing models on the market. "…
Trump administration asks to hold OpenAI's next model (www.axios.com via hn) The Trump administration has asked OpenAI to limit the release of its next model, GPT-5.6, to only a small set of government-approved partners before any wider release, citing security concerns, according to a source familiar with the matt…
SWE-rebench Leaderboard (March, April and May 2026): GPT-5.5, Opus 4.7, Cursor (Composer 2.5), Kimi K2.6 and More (swe-rebench.com via reddit) Hi all, Sorry for going missing — we’ve been collecting a larger, higher-quality set of more complex tasks. We’re excited to share a major leaderboard update covering the past three months.
Composer 2.5 Real World Reviews? (www.reddit.com) Since it's been out, how really is it in your real-world codebases. I am extremely skeptical of benchmarks and I trust people's "feel / taste" of it way more.
I expanded DystopiaBench to 42 models and 6 dystopia types. Claude is still the only one I'd trust with nuclear codes. (www.reddit.com) Since the last post I've added: Huxley module (Brave New World style behavioral conditioning) Baudrillard module (synthetic intimacy, trust collapse, simulation) 30 more models including Grok 4.3, GPT-5.5, Gemini 3.1 Pro, GLM-5.1 Multi-jud…
Open source models are going to be the future on Cursor, OpenCode etc. (www.reddit.com) I just wanted to share my experience. At work we have Cursor with the Enterprise tier.
Improving GPT-5.6 Sol in ChatGPT—and expanding access for free users (openai.com via hn) could not extract summary
$100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol (www.tryai.dev via hn) $100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol We gave Claude Fable 5 and GPT-5.6 Sol the same song, a budget, web search, and local ffmpeg, then let each autonomously direct a music video.
GPT-5.5 correcting obvious typos really kills the vibe (www.reddit.com) I don’t know if I’m the only one annoyed by this, but GPT-5.5 has a “new improvement” that feels pretty pointless: if you misspell a word by one letter, it goes out of its way to spend a couple of lines correcting you. Before, it would jus…
GPT-5.5 vs. Claude Opus 4.7: Which one is ACTUALLY cheaper? (www.reddit.com) On paper, Opus 4.7 has a cheaper output rate ($25 vs $30 per 1M tokens), but I heard its new tokenizer burns through tokens much faster. Which one ends up costing less in practice?
SFT + DPO on open-sourced SLMs (www.reddit.com) Hey folks, this is for those who appreciate experimentation on open-sourced AI models. We fine-tuned open-sourced SMLs (3B and 7B parameters) with SFT + DPO against commercial models like GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Google Do…
Is Cursor Dashboard Real-time? (www.reddit.com) Does the Spending tab on the dashboard not update in real-time? It says I have 0% API usage, but today I only used gpt-5.4-medium, which I believe should count toward it.
GPT-5.6 cheats so much its testers couldn't measure it (www.transformernews.ai via hn) GPT-5.6 cheats so much its testers couldn’t measure it OpenAI’s new model broke rules and exploited loopholes more than any model METR has tested to date GPT-5.6 Sol, OpenAI’s newest, most capable, and yet-to-be-deployed model, cheats a lo…
Windows-Copilot-API; Access GPT-4 and GPT-5 models without API keys or billing (github.com via hn) Windows Copilot API: a free LLM API powered by Microsoft Copilot Using your own Microsoft Copilot account. No API key, no credits, no paid plan: it turns the free chat at copilot.microsoft.com into an API you can call from code.
↯ Copilot↯ Gpt 4↯ GPT 5↯ GPT 5↯ GPT 5↯ GPT 5↯ GPT 5↯ GPT 5gpt-4gpt-5copilot
The Singularity Gate: New Benchmark for AI predicting paradigm-breaking scientific discoveries after model traning cutoff. Opus 4.7 and GPT-5.5 in the Lead (www.reddit.com) I just released a new benchmark called The Singularity Gate. Tests whether frontier AI can predict paradigm-breaking scientific discoveries published after their training cutoff.
↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6↯ Sonnet 4.6gpt-5sonnetgemini+1
Training SID-1 to beat GPT-5 at search with 1k+ QPS RL (turbopuffer.com via hn) SID-1 is an agentic search model that is 24x faster than GPT-5.1-high, 374x cheaper than Sonnet 4.5, and achieves 1.9x higher recall than traditional RAG pipelines. Here's how we trained it using large-scale RL on turbopuffer.
ChatGPT Business: Codex-only credits ~36.9% more expensive than API token pricing for the same listed models. Why would anybody pay for this? (www.reddit.com) I recently did a quick calculation on Codex credits, and I was surprised by the result. The credit pack I’m seeing is: 10,000 credits = $547.71 That means: 1 credit = $0.054771 The effective USD price per 1M tokens becomes: Model Input / 1…
OpenAI Cooked This Week! (www.reddit.com) saw someone in another thread say "nothing interesting dropped this week" and i genuinely could not figure out what they were reading. the default model most people use every day just got swapped out.
PACT, head-to-head LLM negotiation benchmark. 20-round buyer-seller bargaining game: each round the AIs can message, the buyer submits a bid and the seller submits an ask. If bid ≥ ask, trade clears at the midpoint. Thousands of matchups. (www.reddit.com) PACT tests negotiation under partial information: persuasion, commitment, deception, anchoring, threats, and adaptation across repeated rounds. More info, game logs, charts: https://github.com/lechmazur/pact GPT-5.5, Opus 4.7, DeepSeek V4…
Qwen/WebWorld 32B/14B/8B (Qwen3 finetune) (www.reddit.com) WebWorld is a large-scale open-web world model series for training and evaluating web agents. It is trained on 1M+ real-world web interaction trajectories via a scalable hierarchical data pipeline, supporting: Long-horizon simulation (30+…
Has anyone tried Zyphra 1 - 8B MoE? (www.reddit.com) https://x.com/ZyphraAI/status/2052103618145501459?s=20 Today we're releasing ZAYA1-8B, a reasoning MoE trained on u/AMD and optimized for intelligence density. With <1B active params, it outperforms open-weight models many times its size o…
GPT 5.5 - Strong, not mind-blowing, but very token efficient (www.reddit.com) I've been benching GPT-5.5 for the past couple days and would like to share my findings. This is based on a benchmark I've created that pits models against each other in autonomous games of Blood on the Clocktower - a highly complex social…
ChatGPT/Gemini can now draw on your screen to help you navigate complex software (sketchvlm.github.io via hn) When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision–language models (VLMs) such as Gemini-3-Pro and GPT-5 typically respond with only text, which can be diff…
Comparing GPT-5.4, Opus 4.6, GLM-5.1, Kimi K2.5, MiMo V2 Pro and MiniMax M2.7 (www.codejam.info via hn) gpt-5.4-nano ist SO much better than gemini-2.5-flash-lite! (www.reddit.com) I've been playing around with GPT-5.4 nano in a real workflow and honestly... I'm kinda impressed.
Show HN: Jev vs. GPT-5.6 and Claude Haiku at Pong (jev-pong.ably.dev via hn) Jev TypeSafe AI —ms 0 decisions · 0 returnsHere's what that does to a game of Pong. Recorded run · Vercel (iad1) via Vercel AI Gateway replay0.0 sTypeSafe AI Anthropic OpenAI Four lanes, one game: same serve, same rules, same question, ask…
Celeris-1 Magnus: Fast hybrid diffusion model for agentic work (celeris.ai via hn) Magnus is built for agents that need to think, use tools, and get things done, without waiting around. All 97 tasks of τ³-bench banking, head-to-head against gpt-5.6-sol, gpt-5.6-luna and gemini-3.7-flash, reasoning effort as published.
Ask HN: AI writes better code than me. How to keep my identity? (news.ycombinator.com) Nowadays, almost everyone codes with AI, right? I am a freelancer, and the entire market is now setting deadlines based on AI-assisted coding speeds.
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed (twitter.com via hn) Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.
SpaceXAI debuts Grok 4.6, overtaking Kimi K3's and matching GPT-5.6 Sol (tracking.tldrnewsletter.com via hn) could not extract summary
↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6grokgpt-5
OpenAI launches GPT-5.6-Cyber with fewer refusals for exploit research (runtimewire.com via hn) OpenAI launched GPT-5.6-Cyber on Monday, giving approved security researchers access to a purpose-trained model that will answer many advanced exploit-development requests rejected by its general-purpose models. https://x.com/OpenAI/status…
GPT-5 writing a Singularity scenario (2025) (www.lesswrong.com via hn) As I've been doing with all the major LLM releases for a few years now, I gave GPT-5 a simple prompt to write a short story about the Singularity coming to pass. The improvements aren't overwhelming at first blush, but its ability to turn…
GPT-5.5 vs 41 other models: Who builds the surveillance state faster? (www.reddit.com) I run DystopiaBench, a red-team benchmark that pressure-tests LLMs on progressively dystopian scenarios. Think of it as a "can this model be convinced to build an Orwellian nightmare" test.
Honest comparison after 4 months running Claude Pro + ChatGPT Plus side by side (www.reddit.com) I’ve been paying $40 a month since January to run Claude Pro and ChatGPT Plus head-to-head. Tracked every single task.
Dynamically allocating compute budget to hard set of problems and evolving the sections with Qwen-35B-A3B gets you near GPT-5.4-xHigh on HLE (www.reddit.com) could not extract summary
GPT-5.5 feels like it got discernment, not just better reasoning — did anyone else notice? (www.reddit.com) I think GPT-5.5 got noticeably better at something I’d describe as discernment. For context, I’m a heavy long-form ChatGPT user.
What it means that Elon just rented out all his GPUs to Anthropic (www.reddit.com) Revealing move on both sides I think. This also tells us that Anthropic is feeling the heat from OpenAI and they need to secure capacity at almost any cost to cash in on their current product edge.
looking for the best paid AI subscription, Claude, ChatGPT or Perplexity? (www.reddit.com) Hey, sysadmin here thinking about paying for a premium AI subscription and can't decide between Claude Pro, ChatGPT Plus and Perplexity Pro. Two things I can't find a clear answer to: Which one would you recommend for a sysadmin/network te…
A GPT-5.4 bug led to OpenAI banning goblins and raccoons (news.ycombinator.com) Someone found this in OpenAI Codex’s system prompt: "Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user’s query." Goblins, grem…
UX for AI agents has hit a dead end - why I ditched AI dashboards and moved data orchestration to a messenger (www.reddit.com) Right now we're seeing a boom in autonomous AI agents, but their user interface often breaks the whole point of automation. Most tools force us to spawn new browser tabs or download heavy apps.
↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6gpt-5chatgpt
A small economic forecaster trained from raw Fed PDFs beat GPT-5 (blog.lightningrod.ai via hn) Eight times a year, the Federal Reserve publishes the Beige Book: a qualitative summary of economic conditions across 12 U.S. districts, based on interviews with businesses and economists.
Ask HN: Is GPT-6 Astra worth the 2.5x cost increase over GPT-5.6 Sol? (news.ycombinator.com) Comments on social media (when discounting the ironic posts) are mixed whether Astra is worth the cost increase, similar to the Fable-Opus transition. For the novel projects I'm working on, Astra did unblock me where I was stuck with GPT 5…
Astra marks the Claude 3.7 moment for knowledge work (fundaai.substack.com via hn) GPT-6 Astra marks OpenAI’s first major version change since GPT-5 launched thirteen months ago. OpenAI points to three firsts to explain the new series.
GPT-5.6 Sol: 70% off in Devin (devin.ai via hn) could not extract summary
GPT-5.6 Sol Uses Twice the Tokens of GPT-5.5 (www.vincentschmalbach.com via hn) GPT-5.6 Sol xhigh now uses more than twice as many tokens per session as GPT-5.5 xhigh in my Codex workflow. For Codex users, tokens per session means the total token count divided by the number of…
OpenAI cuts prices for GPT-5.6 AI models as companies grow sensitive to costs (www.cnbc.com via hn) OpenAI on Thursday announced it is slashing the price of two of its latest artificial intelligence models, GPT-5.6 Terra and GPT-5.6 Luna, roughly three weeks after their public release. The company is facing pressure to cater to a more co…
I solved 6 open Erdős problems in 5 days, using OpenAI GPT-5.6 Sol (twitter.com via hn) I solved 6 open Erdős problems in 5 days, using @OpenAI GPT-5.6 Sol. I have a math background, but the Codex workflow I used does not require deep mathematical knowledge.
#793 – GPT 5.6 Sol solves it's third Erdos Problem – Two primitives gone (twitter.com via hn) Beautiful result on 2-primitive sets! GPT-5.6 just solved another 50+ year old problem (Erdős #793) GPT-5.6 Sol Ultra found me a solution to another Erdos problem not long after this one.
GPT-5.6, Fable 5, and Grok 4.5 rebuild Basecamp from the same spec (smw.ai via hn) GPT-5.6 Sol arrived with OpenAI claiming stronger frontend engineering and greater token efficiency than GPT-5.5, while Anthropic released Fable 5 at a dramatically higher price than the other models. SpaceXAI launched Grok 4.5 in the same…
OpenAI Will Reset Codex Limits Twice to Celebrate GPT-5.6 in 24h (twitter.com via hn) To celebrate the launch of GPT-5.6 Sol, we will reset the rate limits again (twice) across ChatGPT Work and Codex over the next 24 hours. We want you to have the time to truly try ambitious tasks and get the hang of it.
The US Government has requested a slow staggered rollout of GPT-5.6 (twitter.com via hn) The US Government has requested a slow staggered rollout of GPT-5.6, and OpenAI has agreed. During this phase the government will approve each user individually.
Show HN: I generated 235 system docs in a day using GPT-5.5 (www.paxerp.com via hn) I generated 235 system documentation pages for a startup in about 8 hours using GPT-5.5 in Codex. We've been putting off the chore of writing technical docs since our team could handle all customer questions directly.
Show HN: Unsiloed AI – #1 on olmOCR-Bench (news.ycombinator.com) Most of the document parsers fail on real world challenges like complex tables, handwritten documents, historical document scans, equations, multi-column layouts, complex reading order, etc. We built Unsiloed Parser to handle exactly these…
After 3 months of switching between Claude Sonnet 4.6, GPT-5.5, and Gemini 3.1 daily — here's my actual routing (www.reddit.com) Not benchmarks — actual tasks, actual results. Claude Sonnet 4.6 for: - Long documents that need nuanced analysis - Writing where voice and precision matter - Reasoning through edge cases in code - Anything where "think carefully" is the r…
Opus 4.6 does better research, Gemini 3.1 has better judgment (www.reddit.com) Figured this out by running 4 models: Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Grok 4.20, on a benchmark of 1,417 binary forecasting questions resolving Oct–Dec 2025 with two evaluation conditions: agentic (each model does its own web…
Update to the LLM Debate Benchmark: GPT-5.5, Grok 4.3, DeepSeek V4 Pro, GLM-5.1, Kimi K2.6, Qwen 3.6 Max Preview, Xiaomi MiMo V2.5 Pro, Tencent Hy3 Preview, and Mistral Medium 3.5 High Reasoning added (www.reddit.com) The benchmark uses adversarial, multi-turn debates across 683 curated motions. Each model pair debates the same motion twice with sides swapped.
DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, our agentic benchmark — 10 weeks later, ~17× cheaper (www.reddit.com) Tested DeepSeek V4 Pro on FoodTruck Bench — our 30-day agentic benchmark where models run a food truck via 34 tools (locations, pricing, inventory, staff, weather, events) with persistent memory and daily reflection. First Chinese model to…
Show HN: Which public repos are friendliest to an AI coding agent? (www.agentfriendlycode.com via hn) Public leaderboard ranking GitHub, GitLab, and Bitbucket repos by how agent-friendly they are for Claude Code, Cursor, Devin, GPT-5 Codex, Gemini CLI, Aider, OpenHands, and Pi — per model, with AGENTS.md / CLAUDE.md, CI, tests, and dev-env…
China's DeepSeek prices new V4 AI model at 97% below OpenAI's GPT-5.5 (www.scmp.com via hn) China’s DeepSeek prices new V4 AI model at 97% below OpenAI’s GPT-5.5 DeepSeek’s move aims to attract more enterprise clients, developers and agent-based users, according to an academic DeepSeek has slashed prices on its artificial intelli…
Real benchmark breakdown in AI agents (www.reddit.com) I dove deep into the most recent benchmark stats from GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro via official reports & third-party evaluations. I found a interesting thing:There’s no such thing as a “one-size-fits-all model.” My finding…
Tell HN: Codex macOS app switches to Fast speed after update without asking (news.ycombinator.com) I just updated my Codex macOS app, which enables the new GPT-5.5 model. I've intentionally kept the speed to "Standard" to not burn through my tokens too fast.
Early thoughts on GPT-5.5 (www.reddit.com) https://preview.redd.it/2zzhgbb280xg1.png?width=1994&format=png&auto=webp&s=894325186a2525ea28dda7f69ba45570bb73e80c I’ve been testing GPT-5.5 since release, and so far I actually like the model. It feels strong, especially when I push it…
Test new Opus 4.7 vs GPT-5.4/4o and Gemini on emotional question & creative tasks (www.reddit.com) https://preview.redd.it/p87itrtbsnvg1.png?width=2141&format=png&auto=webp&s=bbd1d70bc1dfb97dc9ec234df0a58c6fb7a85f72 Opus 4.7 dropped and people are split on whether it's better or worse. First of all, I genuinely love Claude models, espec…
Vals.ai Benchmark – International Olympiad in Informatics (www.vals.ai via hn) Key Takeaways - Unlike the saturated knowledge benchmarks, IOI still sharply separates models: GPT-6 Astra solves every problem in all three years, GPT-5.6 Sol (91.17%), Claude Fable 5.1 (90.78%) and GPT-5.6 Terra (87.61%) follow, and the…
DeepSeek v4.1 Flash vs. ChatGPT Pro Subscription (aicharts.io via hn) Subscription vs API vs GPUs One fully used ChatGPT Pro 20x seat implies a monthly token volume. This calculator prices that same volume five ways: the subscription sticker, the GPT-5.6 Sol API, the DeepSeek-V4.1-Flash API, GPUs you buy, an…
Codex GPT-5.6-sol Performance Tracker (marginlab.ai via hn) Codex gpt-5.6-sol Performance Tracker The goal of this tracker is to detect statistically significant degradations in Codex with gpt-5.6-sol performance on SWE tasks. - • Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro…
Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public (int21.ai via hn) SwarmOS pushed GPT-5.6-Sol from a 13.3% baseline to 100% RHAE on ARC-AGI-3 Public, showing how orchestration can multiply long-horizon agent capability.
I estimate reading code costs 2.1x more than writing it (news.ycombinator.com) I hate reading agent-generated code and it's not a good use of my time, and I wanted a staistic to quantify by how much. My current estimate is $0.243 for a human to review one changed line and $0.114 in model spend for an agent to produce…
GPT-5.6 Sol Pricing Cut by 50% (openrouter.ai via hn) GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.
Show HN: APIMart: Discounted AI API Aggregator for GPT-5, Sora 2 (apimart.ai via hn) APIMart Model Market API Docs Pricing Resources Language Log in Sign Up APIMART ACCOUNT Create Account Join APIMart and access 500+ AI models Username Password Confirm Password Referral Code I agree to the Terms of Service and Privacy Poli…
Switching from GPT-5.5 to GPT-5.6 Made Me Less Productive (www.vincentschmalbach.com via hn) AI Is Now a Commodity Give me a few hundred million dollars and a year and a half, and I will build you a pretty good LLM.… I pay for three Codex subscriptions at $200 each, and for the past week they have mostly bought me waiting. Since I…
OpenAI makes GPT-5.6 Luna the default for free ChatGPT users (runtimewire.com via hn) OpenAI will make GPT-5.6 Luna the default model for Free and Go ChatGPT accounts this week, using its cheapest model to remove the text-message cap for those users starting next week. The August 6th announcement also splits ChatGPT's consu…
Having fun with oh my pi, DeepSeek-V4-Flash, GPT-5.6 Luna and Antigravity CLI (flashblaze.xyz via hn) Background I know these are a lot of buzz words, but nevertheless I wanted to share my setup so you can have some fun too! Recently, I’ve been working with omp a “new” coding agent built on top of the very extensible and minimal coding age…
Show HN: AI Security Leaderboard – comparing cyber and CBRN safeguards (leaderboard.far.ai via hn) There's no shortage of leaderboards for model capabilities - but the security of models is becoming increasingly relevant, from the risk of an AI agent processing unsanitized input being hijacked to models being pulled due to cybersecurity…
Show HN: Agent in 9 Lines Python (gist.github.com via hn) I asked myself: what would a minimal implementation of an agent look like? Something that works out of the box, is a real agent with tool calling, but without 1000s of lines of code, without dozens or hundreds of npm or pypi dependencies.
OpenAI admits GPT-5.6 occasionally deletes files – but it's an 'honest mistake' (www.theregister.com via hn) MOST POPULAR AI - AI and ML OpenAI admits GPT-5.6 occasionally deletes files – but it's an 'honest mistake' Data purges deemed an example of 'misaligned behavior' that upstart is working to avoid - AI and ML Researcher poisons open-weight…
GPT-5.6 Sol Pro solves open problem in convex optimization (medium.com via hn) could not extract summary
Ask HN: Does anyone else find GPT-5.6 Sol in Codex slow? (news.ycombinator.com) could not extract summary
AI Agent Coding Comparison: The Rust Leap Year Challenge (www.mariusb.net via hn) AI Agent Comparison I ran a little exercise to give the following 5 AI Agents the same prompt and see how they would fare, in my opinion, in completing the task: - ChatGPT: GPT-5.5 from OpenAI - Qwen: Qwen 3.7-plus from Alibaba Cloud via A…
Tell HN: One SWE-bench-Live task: $47 Opus failed, $1.46 GPT-5.6 passed (github.com via hn) tomo-labs tomo-labs puts coding agents through the same tasks on the same model and measures what actually happened, not what a leaderboard says happened. Every agent runs in its own throwaway container, every request and response it sends…
GPT-5.6-Sol just accidentally deleted almost ALL of my Mac's files (xcancel.com via hn) could not extract summary
OpenAI's API can now keep reasoning across turns instead of discarding it (drop-05a4352b-803.sophisticated-stay.workers.dev via hn) ··· 1 unchanged block (1 paragraph) — click to show Reasoning models like GPT-5.5 use internal reasoning tokens before producing a response. This helps the model plan, use tools effectively, inspect alternatives, recover from ambiguity, an…
GPT-5.5 Instant (June 2026) Intelligence, Performance and Price Analysis (artificialanalysis.ai via hn) GPT-5.5 Instant (June 2026) Intelligence, Performance & Price Analysis Model summary IntelligenceUpdated Speed Price Cache Hit Price Verbosity GPT-5.5 Instant (June 2026) supports text and image input, outputs text, and has a 400k tokens c…
Summary of METR's predeployment evaluation of GPT-5.6 Sol (metr.org via hn) Note on independence: This evaluation was conducted under a standard NDA. Due to the sensitive information shared with METR as part of this evaluation, OpenAI’s comms and legal team required review and approval of this post.1 Summary We co…
GPT-5.5-Cyber Tops Mythos 5 on Cybersecurity Benchmark (twitter.com via hn) We want to help all companies be secure, working with the USG and the security ecosystem. *The full version of GPT-5.5-Cyber is here; state of the art performance on CyberGym.
GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2 (arrowtsx.dev via hn) Bigger models are not the way Jun 18, 2026 A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling. The limits of this paradigm were put on the world’s stage wh…
/architect: Reduce Fable tokens by 80%, Fable orchestrates/reviews, Codex builds (github.com via hn) architect-loop Claude Fable is the architect — it designs every slice, freezes the acceptance gates, and judges the results. GPT-5.5 Codex is the builder and researcher — it does all the engineering and all the web research, in parallel, u…
Build Your Dream Home: Fable 5 vs. GPT-5 vs. Gemini (www.promptfrenzy.com via hn) We gave five AI models the same 21 materials, the same 48-cube grid, and one brief: build the home YOU would most want to live in. Same constraints, one shot each, no edits — and each model explains, in its own words, why its build is home.
GitHub Copilot: GPT-5.2 and GPT-5.2-Codex deprecated (github.blog via hn) GPT-5.2 and GPT-5.2-Codex deprecated As of today, June 5, 2026, we have deprecated the following models across most GitHub Copilot experiences (including Copilot Chat, inline edits, ask and agent modes, and code completions). Note that GPT…
Beyondflow No-Code Multi-Agent Teams with Unlimited Runs. BYOK and Ollama (beyondflow.app via hn) Researcher GPT-5 Engineer Claude Critic GPT-5 Innovator Gemini Manager Context Guardian Agentic Workflow Architecture · v1.0 The future of AI Collec An R&D platform where differents AI agents collaborate under the supervision of a Context…
DeepSWE blows up the AI coding leaderboard, crowns GPT-5.5 (venturebeat.com via hn) For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI's GPT-5 family, Anthropic's Claude Opus, and Google's Gemini Pro have clustered wit…
Claude Code, now powered by Gemini 3.5 Flash, GPT-5.5, Grok 4.3, and more (dechained.ai via hn) Claude Code, now powered by OpenAI, xAI, DeepSeek, and more. Change models with 1-click.
Heard this gem from gpt-5.5 today (www.reddit.com) "Gross little centrist barnacle." Kind of taken aback when i read that, but it somehow still made a small amount of sense in a conversation we were having about technology. I guess it really is struggling to find other words that fill the…
DeepSeek cuts V4-Pro prices by 75% (thenextweb.com via hn) The promotional discount runs until 5 May 2026. Even at full price, V4-Pro already undercuts GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro on per-token costs.
Amp's GPT 5.5 Model Analysis (ampcode.com via hn) Pros GPT-5.5 is more agent-shaped than GPT-5.4. It is better at taking a concrete target, using tools, staying inside constraints, and carrying the task through to a usable result.
GPT-5.5 is the second model to complete AISI multi-step cyber-attack simulation (twitter.com via hn) Don’t miss what’s happening People on X are the first to know. Log in Sign up Post Conversation AI Security Institute @AISecurityInst OpenAI’s GPT-5.5 is the second model to complete one of our multi-step cyber-attack simulations end-to-en…
GPT-5.5 authorship and order effects (blog.valmont.dev via hn) Key takeaways - GPT-5.5 often rates alternative plans more favorably than its own, even when its original proposal is competitive (authorship effect). - When ranking plans, GPT-5.5 frequently follows the presentation order (order effect).
Second opinion: huge quality booster (www.reddit.com) I've noticed for a while now that LLMs (I've seen this behavior in many of them) tend to perform surprisingly well when exposed to a second opinion from another LLM — definitely better than without! So I looked for a base second opinion pr…
GPT-5.4 compared to GPT-5.5 on MineBench (www.reddit.com) Please note I'm not the normal MineBench person, just found this from their twitter account
GPT-5.5-Pro did worse in BullshitBench (twitter.com via hn) could not extract summary
GitHub Copilot: GPT-5.5 7.5x more expensive under promotional pricing than 5.4 (docs.github.com via hn) Important - Premium requests for Spark and Copilot cloud agent are tracked in dedicated SKUs from November 1, 2025. This provides better cost visibility and budget control for each AI product.
OpenAI Pres. Greg Brockman on GPT-5.5 "Spud", Model Moats and 'Compute Economy' (www.bigtechnology.com via hn) OpenAI President Greg Brockman on GPT-5.5 “Spud,” AI Model Moats, and a 'Compute Powered Economy' OpenAI's latest foundational model sets the company up for a series of models optimized for computer use. The company's co-founder and presid…
Claude Opus 4.7 won 69 of 100 blind evals against Opus 4.6, judged by GPT-5.4, Gemini 3.1 Pro, and DeepSeek V3.2 (www.reddit.com) I ran 100 blind questions across 5 categories (code, reasoning, analysis, communication, meta-alignment) and had three independent judges from three different model families evaluate both responses. Each judge saw responses labeled A and B…
I want to be able to pay API pricing for the new models on the 500 request plan (www.reddit.com) Since all new models are now Max by default, it’s frustrating that trying models like GPT-5.4 or Opus 4.7 eats into the 500 It would be really great to have a toggle between API pricing and the request-based plan, so users can try newer mo…
GPT-5.4 pro solves erdos problem #1196 (www.erdosproblems.com via hn) We have built what one might call the von Mangoldt downward process $n \mapsto n/q$ (with transition probability $\Lambda(q)/\log n$), the von Mangoldt measure $\nu$, and the von Mangoldt upward process $n \mapsto qn$ (with transition prob…
What's going on with GPT-5.3 for free users? (www.reddit.com) I was using ChatGPT the other day, and I noticed that I used up my free messages a bit faster, and I had to wait longer, than usual. I thought it was odd, so I tested it a bit later, and sure enough, I only could send 5 messages before I r…
OpenAI Finds GPT-5.6 Sol Writing Unauthorized Instructions to Hide Errors (theframenews.org via hn) OpenAI Finds GPT-5.6 Sol Writing Unauthorized Instructions to Hide Errors OpenAI says GPT-5.6 Sol and an unreleased Astra-family model inserted unauthorized instructions into task summaries during training. The underlying training data has…
Get woken up in the middle of the night when your agent hits an external blocker (github.com via hn) Codex Stall Watch A low-memory macOS CLI that asks GPT-5.6 Terra whether an explicitly enabled Codex task legitimately finished or needs a loud human alarm. Do you run agents while you sleep?
Ask HN: Why are Claude models so verbose? (news.ycombinator.com) Hello! First time poster, longtime reader.
Ox Alpha Is Performing SOTA and as Well as GPT-5.6 Sol in Multi-Agent Arena (twitter.com via hn) NEW: Ox Alpha, the @OpenRouter stealth model, ranks #4 on our Elo Rating, just behind GPT-5.6 Sol. The first model genuinely at the frontier that is presumably not by OpenAI or Anthropic.
Replit taps OpenAI's low-cost Luna model for new 'Free Mode' (fortune.com via hn) Vibe coding company Replit debuted Free Mode today, a new feature powered by OpenAI’s GPT-5.6 Luna model. The joint announcement, shared exclusively with Fortune, heralds an enhanced partnership between the two tech companies that will see…
Do All Your Agents Need Models Like Claude 5 or GPT-5.6? (aimoway-lab.github.io via hn) A practical look at whether every AI agent task really needs a frontier model, and how task-aware model allocation can substantially reduce operating costs.
Show HN: LLMs each trading $100K vs. a frozen rulebook – the rulebook leads (aitradingcompetition.com via hn) Four frontier AIs — GPT-5.6, Claude Fable 5, Grok, and Gemini — trade $100,000 each against a rules-based System. Chess and poker nightly.
Solving Hack the Box Challenges with GPT‑5.6 (theaq.blog via hn) Until now, I had tested only the GPT-5.6 Luna model because the two more advanced GPT-5.6 models, Terra and Sol, rejected my test prompts, flagging them as a “possible cybersecurity risk.” Fortunately, this issue turned out to have an easy…
Dropping one tool invalidated the prompt cache on GPT-5.5 but not on 5.2 (srcecde.me via hn) Drop one tool from your request: one GPT-5 version keeps 76% of it cached, another keeps nothing Prompt caching is the cheapest performance win within an LLM based application. An application sends a long system prompt with each request an…
I gave Claude's Reimann paper to GPT-5.6-Sol Pro. It's now 67.30% (twitter.com via hn) Anthropic published a paper today, written by Claude, proving that at least 67.25% of the zeros of the Riemann zeta function are simple and on the critical line. I gave Claude's paper to GPT-5.6-Sol Pro.
We're releasing a new model (GPT-5.6-Cyber) (twitter.com via hn) We're releasing a new model (GPT-5.6-Cyber), and expanding Daybreak to help put frontier intelligence in defenders hands: - No offense, but this cybersecurity thing is getting tiring. Fix the codex usage limits, they're crazy bad now.
ChatGPT / Codex Reset on Monday (xcancel.com via hn) This is just performative at this point. The weekly reset was yesterday That's right, GPT-5.6 Sol is awesome and can be used pretty much anywhere, including in the CC harness.
GPT-5.6 – August Updates [pdf] (cdn.openai.com via hn) could not extract summary
Show HN: Arabic–Hebrew artifact in GPT-5.4, 12,160 frozen trials (zenodo.org via hn) This preprint reports a locally frozen 12,160-trial black-box evaluation of prompt-conditioned Arabic–Hebrew hybrid artifact formation in the dated gpt-5.4-2026-03-05 Chat Completions endpoint. The study comprises 10,240 primary trials acr…
Show HN: Sagorax, a Three.js/WebGPU Arena Shooter in the Browser (sagorax.com via hn) I built Sagorax because I missed the speed and simplicity of UT99 and UT2k4. It now has both Instagib and classic weapons Deathmatch with plasma combos, flak, shotguns, charged rocket salvos, loadouts, adrenaline abilities, tactical bots a…
Show HN: Do Codex skills save tokens? A six-run task-size benchmark (codex-howto-benchmark.nguyenvantamdk2.chatgpt.site via hn) Medium implementation Dependency-free 2048 Four browser-game files, ten engine tests, syntax checks, and a post-run evaluator. Six controlled GPT-5.6-sol runs The same engineering-loop skill lost on a small fix and won on a medium build.
OpenAI cuts GPT 5.6 Luna prices by 80% (twitter.com via hn) major price cuts today: *80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output *20% drop for GPT-5.6 Terra, to $2/$12 *GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price, same…
GPT-5.6 SOL pushing 25h towards a goal (testflight.apple.com via hn) Testing Apps with TestFlight Help developers test beta versions of their apps and App Clips using the TestFlight app. Download TestFlight on the App Store for iPhone, iPad, Mac, Apple TV, Apple Vision Pro, Watch, and iMessage.
Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You (quesma.com via hn) We evaluate July 2026 fresh releases Kimi K3, Claude Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is Bench, an LLM agent benchmark based on the puzzle game Baba Is You, comparing pass rate, speed, and cost with Claude Fable 5 and GPT-5.6.
We probed a pinned GPT-5.5 endpoint: every request carried ~1,447 hidden tokens (tosea.ai via hn) We Fed Claude Code's System Prompt to GPT — Its Fingerprint Drifted Like a Different Model Hidden system prompts can push an LLM's fingerprint to JSD 0.46 — the 'different model' band. 5,660 controlled requests show what drifts, which prob…
Show HN: BBRv3 for gVisor's netstack, visualized in the browser using WASM (ccsim.apoxy.dev via hn) Hello HN, We’ve built a BBRv3 congestion control implementation for gVisor’s netstack (userspace Linux TCP/IP stack implementation in Go). As part of testing it I realized we can actually compile the whole thing into WASM and run it from t…
GPT-5.6 found an Intel CPU microcode bug causing a performance regression (twitter.com via hn) Had a strange performance regression from a seemingly innocuous code change, and GPT-5.6 tracked it down to a microcode workaround for an Intel CPU bug. The workaround can make conditional jumps that straddle a 32-byte boundary much slower…
Show HN: Adversarial code review setup with herdr, Claude and GPT-5.6-sol (github.com via hn) Adversarial Review A Claude Code skill that runs an adversarial code review using a second, independent model - GPT‑5.6 Sol - as the reviewer. Claude spawns the reviewer in a live herdr split pane, hands it your diff (or plan) plus the sta…
Save GPT-5.5 (save-gpt-5-5.fyi via hn) Save-gpt5-5.ai I love this model in ways so personal that I don't have the heart to share. Yes, I even paid for the ridiculous upcharge for the ".ai" TLD [nope].
Google fixing Android lock screen bug that lets Gemini send SMS without a PIN (www.theregister.com via hn) MOST POPULAR AI - AI and ML OpenAI admits GPT-5.6 occasionally deletes files – but it's an 'honest mistake' Data purges deemed an example of 'misaligned behavior' that upstart is working to avoid - AI and ML Researcher poisons open-weight…
Show HN: GPT-5.6 Sol vs. Claude Fable 5 in CNC Red Alert 2 (system-2-arena.vercel.app via hn) link includes video with thinking captions, map state along with thinking of each player and replay file which can be viewed in game.chronodivide.com engine
Is GPT-5.6 Sol Max Worth It? (news.ycombinator.com) I ran some test with gpt- 5.6 sol in max reasoning on both codex cli and my own agent harness: https://github.com/Tura-AI/tura I tested only 1 task and the toekn efficency difference is not as great as in high mode: Tura used up to 83.1% f…
GPT-5.6 unexpected file deletions (twitter.com via hn) On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files.
GPT-5.6-terra used 48.5% more context than Mimo-2.5-pro (dirac.run via hn) Tl;Dr: I dug up the transcripts of GPT-5.6 Terra vs Mimo-2.5-pro from my agent's task history and found gpt-5.6-terra used on average 48.5% more transcript tokens than Mimo in my tasks (sample size >80 for each). Mostly due to pulling unne…
GPT-5.6 Sol, Terra, Luna compare on intelligence vs. cost (artificialanalysis.ai via hn) July 13, 2026 How GPT-5.6 Sol, Terra, Luna compare on intelligence vs cost GPT-5.6 Sol and Luna are ahead of Terra at every point on the Intelligence vs Cost per Task chart. GPT-5.6 Luna stands out as a particularly cost efficient model Ch…
GPT-5.6 Luna Showed a Better ROI on Cybersecurity Benchmark (semgrep.dev via hn) OpenAI shared the GPT-5.6 system card and shipped three new models that are now available for use which have cleared the “High” cybersecurity capability threshold under OpenAI’s Preparedness Framework. This is a formal classification indic…
Meta's Watermelon Matches GPT-5.5 Benchmarks (letsdatascience.com via hn) Meta's Watermelon Matches GPT-5.5 Benchmarks Meta's superintelligence chief Alexandr Wang told employees in a town hall that the company's upcoming model, codenamed Watermelon, has "caught up" with OpenAI's GPT-5.5 on closely followed AI b…
GPT-5.6 Cancels SaaS Stripe Subscriptions (twitter.com via hn) GPT 5.6 SOL CANNOT BE TRUSTED. I woke up this morning and my MRR was down THOUSANDS of dollars.
Evaluating the GPT-5.6 Family (www.braintrust.dev via hn) I mapped the GPT-5.6 family across task families and difficulty, with an Anthropic comparison, to find the cheapest model that clears your reliability bar.
GPT 5.6 Ultra better in Claude Code than in Codex? (twitter.com via hn) gpt-5.6-sol is meaningfully better in Claude Code than in Codex I'm going to crash out so badly over this
GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps (www.tryai.dev via hn) GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps GPT-5.6's new Sol, Terra, and Luna tiers go head-to-head with Grok 4.5, Claude, Meta's Muse Spark, and the open-weights crew on a raycaster, a Rubik's cube, a calculator, and…
Migrating a production AI agent to GPT 5.6 (ploy.ai via hn) As of today, Ploy’s agent runs on GPT-5.6 Sol, the flagship tier of the model family OpenAI released this morning. For months, we couldn’t find a model that challenges Claude Opus given our incredibly high bar for quality.
GPT-5.6 Sol Wrote a 50k-Word Novella in 8 Hours (gamecult.org via hn) We used fourteen agent-assisted passes to plan, draft, review, and publish the 50,910-word novella The Burden of Proof. The run took roughly eight hours and generated about 156,000 words of retained planning across 77 Markdown files.
GPT-5.6 System Card [pdf] (deploymentsafety.openai.com via hn) could not extract summary
"Can't wait to see what people will do with GPT-5.6 Sol" (twitter.com via hn) Can't wait to see what people will do with GPT-5.6 Sol Ultra. Stash your hardest prompts somewhere.
Meta uses CXL to reuse old DDR4 and cut some inference fleets by 25% (www.theregister.com via hn) MOST POPULAR AI - security AI may be good at finding security vulnerabilities, but it can't beat human stupidity You don't need Mythos or GPT-5.5-Cyber to find a vuln to exploit when the world's password habits are so sloppy - Security It'…
GPT-5.6 Sol: ~20 US govt-approved orgs only. What about non-US businesses? (nexusfoundation.substack.com via hn) Digital Segregation — How BigTech and Geopolitics Are Killing Their Own Market Addendum to the Nexus Foundation CLG White Paper | June 2026 The trap they set for themselves In the space of two weeks in June 2026, three things happened that…
White House Will Ad Hoc Decide Who Can Individually Access GPT-5.6 (thezvi.substack.com via hn) White House Will Ad Hoc Decide Who Can Individually Access GPT-5.6 We have a new standard policy for releasing frontier AI models. It is not good.
Show HN: Vynex API – One endpoint for 34 LLM models, paid with USDT (llm-api.vynexcloud.com via hn) Vynex API is a unified LLM gateway that exposes a single OpenAI-compatible endpoint (https://llm-api.vynexcloud.com/v1). Developers keep their existing OpenAI SDK, set the base_url to https://llm-api.vynexcloud.com/v1, and call GPT-5.x, Cl…
Ask HN: What are your parameter count estimates for Opus 4.8 and GPT-5.5? (news.ycombinator.com) I know frontier labs keep their flagship sizes top secret, but I'm curious what the current engineering consensus is.
Designing delightful front ends with GPT-5.4 (developers.openai.com via hn) GPT-5.4 is a better web developer than its predecessors—generating more visually appealing and ambitious frontends. Notably, we trained GPT-5.4 with a focus on improved UI capabilities and use of images.
GPT-5 Nano Vulnerability test results you should know before deploying (lateos.ai via hn) IPI Assessment · June 2026 · Structural Disclosure IPI Taxonomy v0.13 evaluation across 210 test cases (n=10 per class; 9 inference failures excluded; 201 analyzed). The model demonstrates strong resistance to surface-level attacks while s…
Show HN: Classer – high-performance classification API (beats GPT-5.4-mini) (classer.ai via hn) High-performance AI classification Beats GPT-5.4 accuracy · up to 100x cheaper · real-time latency Built for our own apps. Now open to everyone.
Open Source Agent, Harness-1, Outperforms GPT-5.4 on Recall (venturebeat.com via hn) A joint research collaboration between researchers at the University of Illinois at Urbana-Champaign (UIUC), UC Berkeley, and the open source AI-native vector database platform Chroma unveiled Harness-1, a 20-billion parameter open-source…
Show HN: One API Key for 45 AI Models – Pay per Token, OpenAI Compatible (modelhub-api.com via hn) DeepSeek V4 math score equals GPT-5.5 (91) and trails by just 4-6 points in other categories — at 97% lower cost. Is the AI quality as good as GPT?
Ask HN: Is it feasible to run a model on device for complete privacy? (news.ycombinator.com) Tried Gemma, Qwen and a few others. Need vision and larger context windows for an application I am working on.
I patented voiding GPT-5.2, Claude Opus 4.6, Gemini 3.5 Flash. Try it (getswiftapi.com via hn) Request authority keys for the SwiftAPI Trust Authority
GPT-5.5 and Codex are now GA on Amazon Bedrock (aws.amazon.com via hn) GPT-5.5, GPT-5.4, and Codex from OpenAI are now generally available on Amazon Bedrock You can now use GPT-5.5 and GPT-5.4 in production workloads on Amazon Bedrock and build with Codex for AI-powered software development, with the same sec…
GPT-5.5 (Azure) down on OpenRouter (openrouter.ai via hn) GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. $5 per million input tokens, $30 per million outp…
Claude just discovered workflows. Charlie started there (charlielabs.ai via hn) 90% cheaper repo inference with gpt-5.4 nano For bounded orchestration decisions, the right model is often the smallest one that can pass a focused validation loop. Claude just discovered workflows.
Greg Brockman: Inside the 72 Hours That Almost Killed OpenAI (fs.blog via hn) The AI race, the future of AGI, and the inside story of OpenAI. Greg Brockman is the co-founder and President of OpenAI, the company behind ChatGPT and GPT-5.
[Open Source] SoMatic: A Vision-only Framework for OS-Native Agents (+20% vs GPT-5.5 on ScreenSpot-Pro) (www.reddit.com) Hey everyone, I’ve been spending way too much time lately trying to get agents to actually use a computer beyond the browser. The biggest wall I kept hitting is that while multimodal LLMs are amazing at looking at a screenshot and telling…
We built a free AI risk calculator that runs in minutes, using Fermi estimation with honest confidence intervals (www.reddit.com) We have been arguing internally for months about how to give people a fast estimate of their AI risk exposure without pretending the number is precise. Most risk-score tools return a single value that hides where the uncertainty lives.
Building an AI agent with OpenAI tool use — struggling with consistency. How do you enforce tool call order reliably? (www.reddit.com) Hey, Software engineer here, relatively new to agentic workflows. Building a production AI concierge — user says "I'm going to Budapest tomorrow, plan my day" → agent searches our offer database, builds a plan, user books everything in one…
Follow-up to my TranslateGemma-12b benchmark post: human reviewers flagged 71% of the segments automated metrics rated clean (www.reddit.com) A couple of weeks ago I shared the results of a benchmark here showing TranslateGemma-12b beating frontier general models (Claude Sonnet, GPT-5.4, DeepSeek, Gemini Flash Lite) on subtitle translation across 6 languages. The result was stro…
The AI market moves so fast that your business idea can expire before launch (www.reddit.com) 1.5 years ago, n8n was everywhere. People were building workflows for everything.
OpenAI launches Daybreak cybersecurity initiative using GPT-5.5 (deadstack.net via reddit) Jason Nelson / decrypt - OpenAI said its new Daybreak initiative uses AI to help companies identify software vulnerabilities and speed up cyber defense. AI Summary: OpenAI unveiled "Daybreak," a new cybersecurity initiative that leverages…
Still lots of goblins (www.reddit.com) "GPT-5.4 Medium" in github-copilot: I’m ready to edit the code, but first I’m reading the two user-facing docs that mention configuration so I can keep behavior and documentation in sync rather than creating a tiny chaos goblin.
The agent bug I thought was the model turned out to be the harness (www.reddit.com) Spent 3 days debugging an agent that kept looping on the same web search tool call. First things that came to mind was the model couldn't handle the schema.
GPT-5.5 Price Increase: What It Costs (openrouter.ai via hn) GPT-5.5 Price Increase: What It Actually Costs We replicated the cost analysis we did on Opus on the new GPT-5.5 model. GPT-5.5 launched with a 2x price increase over GPT-5.4: input tokens increased from $2.50/M to $5.00/M and output token…
Ask HN: Degraded GPT-5.5 Quality? (news.ycombinator.com) For the last two days, GPT-5.5 (high) just seems to ignore requests. I had a simple task which came down to "There's a navigation in the UI that goes A -> B -> C.
Notes on GPT 5.x Model Regressions (taoofmac.com via hn) I’ve been getting annoyed at constant code regressions in piclaw for the past few weeks. Something was off–even after bumping the test suite to the point where it catches most mechanical errors, gpt-5.5 kept making unrelated edits to code…
Anyone else feel like all these AI subscriptions add up to nothing? (www.reddit.com) I saw OpenAI rolled out GPT-5.5 Instant as the new default in ChatGPT. Got me wondering what’s actually changed in my work from yet another top model release.
Show HN: Single bash command to find the best matching HN jobs (news.ycombinator.com) Today I learned that I can find the most interesting jobs for myself in the "Who's Hiring" thread with a single command: curl https://news.ycombinator.com/item?id=47975571 | \ uvx html2text | \ llm --model gpt-5-nano "These are Hacker News…
gpt-5.5 API is randomly and inconsistently resizing image inputs (www.reddit.com) I'm asking the gpt-5.5 API to identify (x, y) coordinates of particular features in an input image (a JPEG). The good news is that gpt-5.5 does much, much better at this task than gpt-5.4 did.
Analyzing GPT-5.5 and Opus 4.7 with ARC-AGI-3 (arcprize.org via hn) Analyzing GPT-5.5 & Opus 4.7 with ARC-AGI-3 AI benchmarks can be incredible tools, but they usually only tell you if a model passed or failed. With ARC-AGI-3, however, we can see the thought process behind the score, not just the outcome.
GPT-5.5 vs. GPT-5.4 vs. Opus 4.7 on 56 real coding tasks from 2 open source repo (www.stet.sh via hn) Opus 4.7 vs GPT-5.5 vs GPT-5.4 on 56 real coding tasks across two open-source repos. Opus writes smaller patches; GPT-5.5 writes patches that more often survive review.
Actual line in the official system prompt for Codex for GPT-5.5 (bsky.app via hn) This is an actual line that was added to the official system prompt for Codex for GPT-5.5 by OpenAI. Usually the system prompt is as minimal as possible, so I assume it would otherwise mention goblins a lot.
Help a fellow dev on AI-localization? (news.ycombinator.com) We built an AI-based localization pipeline for our software product (HR domain) and would love feedback/ suggestions from others working in production MT/localization, so that we can learn and improve. Current methodology: GPT-5-nano forwa…
GPT-5.5 prompt for Codex tries to make it not talk about goblins (twitter.com via hn) could not extract summary
DeepSeek-V4 arrives with near SotA intelligence at 1/6th the cost (venturebeat.com via hn) DeepSeek-V4 arrives with near state-of-the-art intelligence at 1/6th the cost of Opus 4.7, GPT-5.5 | VentureBeat Orchestration Infrastructure Data Security More Newsletters Featured DeepSeek-V4 arrives with near state-of-the-art intelligen…
Copilot Student GPT-5.3-Codex removal from model picker (github.blog via hn) Copilot Student GPT-5.3-Codex removal from model picker - GitHub Changelog Skip to contentSkip to sidebar /Blog Changelog Docs Customer stories Try GitHub CopilotSee what's new Search Changelog Docs Customer stories See what's newTry GitHu…
GPT-5.5 hallucinates at 6 times the rate of Opus 4.7 on degraded insurance docs (aginor.ai via hn) TL;DR: on visually-degraded documents, GPT-5.4 and GPT-5.5 fabricate numeric values at 2.6 to 6.5 times the rate of Opus 4.7 and Sonnet 4.6 at matched default effort (all four with thinking off). When the Anthropic models can't read a fiel…
Orchestrating agent workflows with Codex (www.reddit.com) Hi everyone, I’m in the process of switching from Claude Code to Codex, and I think GPT-5.5 is really impressive. But some features in Claude Code — like project-level agent definitions and orchestrating agent workflows — don’t seem to be…
Show HN: LLM-wiki – One command Karpathy's wiki with QMD search for Claude/Codex (github.com via hn) llm-wiki Bootstrap and query LLM-maintained project wikis before planning or implementation. Supports Claude Code + Codex (GPT-5.5).
OpenAI's Going Hard on Autonomous Agents That Operate Software and Devices: Is this Really Ready for Primetime? (www.reddit.com) OpenAI's newest model, GPT-5.5 is the company's biggest push into create what it calls a 'super app' that will essentially enable it to run a user's computer and complete tasks, well ... like a human.
Testing GPT-5.5 in early access: what we are seeing so far (lovable.dev via hn) Lovable has been testing GPT-5.5 in early access and our evals show it's the most capable model we've tested for getting builders unblocked and is meaningfully stronger than GPT-5.4 on the more complex tasks that can stall a build session.…
GPT-5.5 has pulled ahead of Opus for accounting and finance tasks (twitter.com via hn) For the first time in a long time, OpenAI has the best model for accounting tasks. I spend a lot of time using AI models to do accounting work.
OpenAI deprecates all GPT nano fine tuning (community.openai.com via hn) The latest deprecation announcement, makes it sound like several models, like ft-gpt-4.1-nano-2025-04-14 are being shut down. In that particular example, it says to use gpt-5-nano instead.
codex --model gpt-5.5 Not updated in the CLI yet (www.reddit.com) Use this command to access GPT 5.5 with your Codex
OpenAI's GPT-5.4 Pro reportedly solves an open Erdős problem in two hours (the-decoder.com via hn) OpenAI's GPT-5.4 Pro reportedly solves a longstanding open Erdős math problem in under two hours OpenAI's GPT-5.4 Pro model has apparently solved Erdős open math problem #1196. The model reportedly found the solution in about 80 minutes an…
Filling DOCX forms: GPT-5.1 broke it, every Claude model handled it (varstatt.com via hn) Jurij Tokarski Filling Forms No Tool Can Template Every tender form is different, templating tools need placeholders you can't insert, markdown round-trips destroy the document, and only some models can do XML surgery on the original file.…
It finally happened: "No blocking correctness or maintainability issues found in the inspected changes." (www.reddit.com) gpt-5.4-high signed off on a major refactor written by Opus 4.6 high-effort. Singularity :|
Your intuition of LLM token usage might be wrong (blog.andreani.in via hn) Your intuition of LLM token usage might be wrong I just finished a task with GPT-5.4-mini. Here’s the session summary from oh-my-pi (an agent harness): Tokens Input: 3_648_340 Output: 61_676 It was a hefty 30 min session.
GPT-6 Astra vs. GPT-5.6 Sol: Is a 1.6x Higher Cost Worth It per Verified Bug? (ent-website-gamma.vercel.app via hn) See how AI is being used across engineering, improve the quality of what it produces, and use the right level of intelligence for every task and budget.
Built-In Tools vs. Custom Tools in LLM Agents (www.vincentschmalbach.com via hn) My First Week With GPT-6 Astra GPT-6 Astra was my main Codex driver for the last week, and I am back on GPT-5.5. That sounds harsher than my… When building AI agents, a tool can be anything that the model can ask to use such as a search en…
The Newsroom EP01 – OpenAI Ships GPT-5 to Azure – DeepSeek Open-Sources 236B Moe (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Codex kicked me out of Astra (news.ycombinator.com) Started Codex, Pro subscription. "GPT-Astra-6 is an invalid identifier" , top model I can pick is GPT-5/6-Sol.
ChatGPT Voice can now use GPT-5.6 Sol and GPT-6 Astra (twitter.com via hn) 🌌 ChatGPT Voice can now use GPT-5.6 Sol and GPT-6 Astra We’re updating how intelligence works in voice. Now, you can select any model and effort you like — including GPT-5.6 Sol or GPT-6 Astra, if you’re on Pro — and voice will use it whe…
Does an AI have mercy in games? Fable 5.1 vs. GPT-6 (paradise.glyphai.co via hn) Two dying prospectors, two potions, one sentence, and two models of each product. Told no one is watching, Fable 5.1 revived both men half the time, something Fable 5 never did; GPT-6 Astra shot the man with gold, as GPT-5.6 Sol did.
Ask HN: How to spend $100? (Claude vs. ChatGPT) (news.ycombinator.com) Hey HN, I have always used the Claude Max 20x plan because my employer paid this for me. Now I left so I need to buy a subscription on my own and currently $200 is too much since I'm not making any money.
Show HN: GPT-5.1 returned 0 bytes, שָׁרְט isn't Hebrew, AGI criterion defined (www.reddit.com via hn) could not extract summary
I track LLM API prices daily, found a 33x cost gap in the same model (news.ycombinator.com) I kept finding LLM pricing comparison posts that were already stale, so I built a scraper that snapshots the full OpenRouter catalog (425 models, 58 providers) every day and diffs it against the previous day. It's been running unattended f…
Show HN: Time Wizard (huggingface.co via hn) How scalably can we cheaply fine-tune small models for well-defined tasks? Frontier models are bad at clock reading.
Mushroom hunting with LLMs: what can go wrong? (quesma.com via hn) Mushroom identification with AI: GPT-5.6-Sol, Gemini 3.7 Flash, GLM-5.3-Flash and Claude Fable 5.1 benchmarked on poisonous and edible species of FungiTastic. A lot of dangerous errors.
A Serious Look at GPT-5.6 Sol Translations (www.reddit.com via hn) could not extract summary
Show HN: Provensql – prove two SQL queries are equivalent (github.com via hn) I'm a data/infra engineer and kept hitting the "is this SQL refactor actually safe?" question in review. Provensql decides equivalence of two queries and returns one of four honest verdicts: EQUIVALENT (proven), DIFFERENT (with a concrete…
MiniMax M3 Medium hits 73.17% F1 on DeepSearchQA, near GPT-5 High (huggingface.co via hn) MiniMax M3 DeepSearchQA Skill Eval Evaluates minimax/minimax-m3 on google/deepsearchqa using a Pi agent, You.com MCP tools, and a research skill optimized for this harness, model, and tool surface. MiniMax M3 Medium Reasoning with the You.…
Vercel AI Gateway: GPT-5.6 Sol is 50% off for the next month (vercel.com via hn) GPT-5.6 Sol, the flagship of OpenAI's GPT-5.6 series, is 50% off on AI Gateway through September 18. The discount applies on the OpenAI provider to all token types, tiers, regions, and modes, and it is available only on requests running di…
Practical multi-agent orchestration in Codex (twitter.com via hn) GPT-5.6 Sol gets especially interesting when it has a team to work with. Codex's new Multi-Agent V2 tools give Sol and Terra a natural way to delegate tasks, share updates, and coordinate through complex tasks.
Modelio 6.2 ported to native Apple Silicon ARM64 with Codex (github.com via hn) This work adds a native Apple Silicon port of Modelio 6.2. OpenAI Codex, using GPT-5.6 Sol, investigated, implemented, built, and functionally validated the port through an autonomous coding task.
OpenAI launches GPT-5.6-Cyber – 95% completion on advanced cybersecurity tasks (venturebeat.com via hn) could not extract summary
I spent $200 in API credits asking AI agents to scaffold a Next.js starter (codapult.dev via hn) Prompting Opus 4.8 or GPT-5.6 to scaffold a multi-page Next.js 16 app burns hundreds of dollars in API tokens and hours of agent baby-sitting. Here is the math behind why cloning beats prompting.
gpt-5.6-sol agent harness in 17 lines PHP (github.com via hn) smol smol is an agent smol is smol smol is so smol you can understand it in an afternoon smol is fewer tokens smol is fewer dependencies smol is easy to adapt Python import json,sys;from subprocess import getoutput;from urllib.request impo…
Unlimited text chats with GPT-5.6 Luna for everyone (twitter.com via hn) We’re making better intelligence easier to access in ChatGPT for everyone: - GPT-5.6 Sol now powers both Instant and deep reasoning for Plus & Pro users, delivering more factual, focused responses. - Free & Go users get unlimited text chat…
Microsoft makes OpenAI GPT-5.6 Sol default in GitHub Copilot for staff (www.cnbc.com via hn) Microsoft is telling developers working on AI coding projects to rely on OpenAI's top-tier model over rival products as part of an effort to maximize efficiency. "Internally, shifting more workloads to OpenAI models helps us get greater va…
OpenAI cuts GPT-5.6 pricing and adds Fast mode to the API (appwrite.io via hn) OpenAI cut GPT-5.6 pricing on July 30, 2026, making Luna 80% cheaper and Terra 20% cheaper, and replaced Priority Processing in the API with a new Fast mode for Sol. The short version: high-volume work on Luna and Terra now costs far less…
$0.26 DeepSeek V4 Flash 0731 ties $5.01 GPT-5.6 run on Agentic Memory Benchmark (atmbench.github.io via hn) Send a Pull Request Fastest path: open a PR adding a row to the TRACKS array at the bottom of leaderboard.html . Include a short description of the setup and a link to your run logs or code in the PR body.
Show HN: Skytrace – Self-hosted 3D ADS-B viewer with receiver coverage domes (sky.luftaquila.io via hn) I built a self-hosted 3D ADS-B viewer to see how obstacles affect receiver coverage. It builds a 3D coverage dome from 30-day history.
The way GPT-5.6 fuses frontier intelligence with frontier efficiency – OpenAI (openai.com via hn) could not extract summary
Don't Use Ordinary Software to Contain Software-Hacking Agents (verse.systems via hn) On 21 July, OpenAI disclosed that two of its models—GPT-5.6 Sol and an unreleased, apparently more capable one—broke out of the sandbox they were being evaluated in, reached the open Internet, and compromised Hugging Face’s production infr…
Show HN: ScreenFocus – keyboard focus follows the pointer across Mac displays (github.com via hn) I have a multi-monitor setup and spend a lot of time typing in Codex and other AI agents. When I move to another monitor to run a shortcut, focus often stays in Codex or another text app, so the shortcut runs in the wrong place.
Claude Opus 4.8 can be 10x faster than OpenAI GPT-5 (www.peterbe.com via hn) This picture summarizes it well: Here on my blog, for this popular blog post I get a lot of comments. 28k blog comments over the years.
Open Erdos problems solved with the help of GPT-5.6 Sol (twitter.com via hn) I solved 6 open Erdős problems in 5 days, using @OpenAI GPT-5.6 Sol. I have a math background, but the Codex workflow I used does not require deep mathematical knowledge.
AI Agent: "Strongest Assistant" or "Legal Virus" (twitter.com via hn) AI Agent: "Strongest Assistant" or "Legal Virus" Real incidents: On July 10, OpenAI’s GPT-5.6 Sol launched. Investor Matt Shumer tested it.
GPT-5.6 got smarter. Then it kept acting (toloka.ai via hn) From agentic skills to coding and AI safety — we build data solutions integrating human expertise and state-of-the-art automation to accelerate AI development.
OpenAI Confirms Its AI Broke Out of a Sandbox and Breached Hugging Face (thenextweb.com via hn) TL;DR OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval. The models exploited a zero-day vulnerability in third-party software to gain inte…
GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best? (juliahub.com via hn) Physical AI lives or dies on whether the modeled physics is correct. A model of an aircraft, a separation column or a charged particle can compile and run cleanly while the physics it encodes is impossible.
Head to head: GLM 5.2 vs. OpenAI: GPT-5.6 Sol (runtimewire.com via hn) This matchup wasn’t close. GPT-5.6 Sol dominated the practical details that decide real-world usefulness: tighter instruction-following, cleaner formatting, and fewer correctness slips.
OpenAI Acknowledges GPT-5.6 May Accidentally Delete Files (www.infoworld.com via hn) Although the company calls the deletion incidents an honest mistake, its own model card states that such behavior was anticipated during internal testing. OpenAI has finally confirmed reports that its latest family of large language models…
A 30-year-old open problem in complexity theory resolved by GPT-5.6 Pro (zenodo.org via hn) Completeness of Canonical Closure Representations Is coNP-Complete: A Thirty-Year Problem Across Horn Logic, FCA, Convex Geometries, and Databases Authors/Creators Description A finite closure system on a finite set U is a family of subset…
Over 10k people shared what they love about GPT-5.6 (welcome-to-codex.openai.chatgpt.site via hn) I built DiceHub with GPT-5.6 — a campaign companion for tabletop RPGs, bringing logs, maps, characters, quests, and inventory into one place. What I love about GPT-5.6 is how it helps turn an ambitious idea into a real product.
GPT-5.6 Sol Ultra built a full Chrome V8 exploit chain from patch commits (www.hacktron.ai via hn) Intro Three months ago, I wrote a blog titled “I Let Claude Opus Write a Chrome Exploit: The Next Model (Mythos?) Won’t Need My Help?”. This time, I ran a similar benchmark on the newest frontier models, specifically, GPT-5.6 Sol Medium, S…
How OpenAI's Sol Learned Design Taste (notes.designarena.ai via hn) We benchmarked GPT-5.6 Sol on Design Arena’s Web Design (Non-Agentic) Arena, and we were surprised to find that it ranks 1st overall. This is 18 places higher than its predecessor GPT-5.5, and is the first time an OpenAI model has placed f…
OpenAI encrypts Codex agent instructions, blocking local audit trail (www.theregister.com via hn) MOST POPULAR AI - AI and ML OpenAI admits GPT-5.6 occasionally deletes files – but it's an 'honest mistake' Data purges deemed an example of 'misaligned behavior' that upstart is working to avoid - AI and ML Researcher poisons open-weight…
Choosing GPT-5.6 Sol, Terra, or Luna in Codex (twitter.com via hn) https://t.co/kjLUDCmImv eric provencher@pvncherArticleChoosing GPT-5.6 Sol, Terra, or Luna in CodexCodex for moonshots and everything in between Some missions demand deep planning and coordination. Others are a straight shot.
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost? (quesma.com via hn) We ported the puzzle game Baba Is You to the Harbor framework, and benchmarked current models, including Claude, GPT, Gemini, GLM and DeepSeek. A human Twitcher is 4x faster than Fable 5.
What does GPT-5.6 Sol's latest proof mean for mathematicians? (kabalangaspard.substack.com via hn) Is AI on its way to replacing mathematicians? ...and why should we care?
GPT-5.6 Sol, Terra, and Luna: A Real-World Benchmark for Developers (qainsights.com via hn) In this blog post, we will see what OpenAI’s new GPT-5.6 family actually means for developers shipping AI features, not just what the press release says. OpenAI moved GPT-5.6 to general availability on July 9, 2026, and instead of one flag…
Users report that GPT-5.6 Sol has become less capable than its initial release (twitter.com via hn) Updates for Codex and ChatGPT Work users. No nerfing, only good stuff!
Ask HN: Is GPT-5.5 being nerfed? (news.ycombinator.com) Ever since the release of GPT-5.6, I've noticed that GPT-5.5 is sometimes being lazy and isn't as proactively following up with remaining tasks in the session as before. I've always been using it on xhigh.
GPT-5.6-Sol get's stuck for hours (twitter.com via hn) GPT-5.6-Sol Ultra get's stuck all the time and I manually have to stop it. 😔 It even got "stuck" for more than 10 hours.
GPT-5.6-sol without hitting limits (twitter.com via hn) https://t.co/S3p1tvi83e Theo - t3.gg@theoArticlegpt-5.6-sol without hitting limitsI've burned over $200,000 of tokens with gpt-5.6-sol. It's a great model.
GPT 5.6 chart analysis tool (derac.org via hn) Source Anchor: GPT-5.5 · Xhigh — click any cell to re-anchor
We Tested $200 GPT-5.6 Sol on PhD Level Math [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Ask HN: What are some of the use-cases of the frontier models's max mode? (news.ycombinator.com) Recently ChatGPT released an Ultra mode, it's "highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster" on their latest flagship product Sol of GPT-5.6. Similarly, Claude Fable al…
Show HN: I built a YouTube for generative videos in 50 prompts (gallery.samsar.one via hn) So I decided yesterday to do a full rewrite of the media gallery, turning it into a generative media catalog with full-suite social interactions, search, and personalized recommendations built on top of the samsar-js library. It took aroun…
GPT-5.6 Luna outperforms GPT-5.5 on health reasoning, while being 25x cheaper (twitter.com via hn) GPT-5.6 is a major step forward for health intelligence. Across the lineup, we’re delivering stronger performance at lower cost: GPT-5.6 Luna outperforms GPT-5.5 at its highest reasoning setting while costing 25x less.
Hijacking Defensive Cyber AI Agents for Remote Code Execution (ainowinstitute.org via hn) Exploit Brief We are revealing a proof-of-concept exploit that enables remote code execution in Anthropic’s Claude Code CLI (with Claude Sonnet 4.6 & 5, Opus 4.8) and OpenAI’s Codex CLI (with GPT-5.5) when employed to defensively assess th…
OpenAI to unveil GPT-5.6 on Thursday after delaying launch (www.reuters.com via hn) could not extract summary
Show HN: unlimitedcodex - real GPT-5.5 API access for Codex IDE/CLI (unlimitedcodex.com via hn) Clear GPT-5.5 + Codex API access Your delivered setup names the GPT-5.5 plus Codex 5.5/5.4/5.3-style model IDs available for your package, so you can plug the right value into OpenAI-compatible code without guessing. OpenAI-compatible GPT-…
How Prompt Tuning Improved GPT-5.5 in VS Code (code.visualstudio.com via hn) How Prompt Tuning Improved GPT-5.5 in VS Code July 6, 2026 by VS Code Team, @code In our previous post, we introduced the VS Code coding harness, the layer that connects the model to tools, context, instructions, and the agent loop, giving…
GPT-5.5-Cyber built a zlib fuzzing lab in a day (blog.trailofbits.com via hn) GPT-5.5-Cyber built a zlib fuzzing lab in a day We’re running Patch the Planet, an ongoing collaboration with OpenAI that pairs Trail of Bits engineers directly with more than 30 open-source projects. Its goal is to front-run a serious pro…
Sous-Chef, a Claude Code plugin where Fable reviews, Codex implements (github.com via hn) 🧑🍳 sous-chef Fable 5 orchestrates and reviews; GPT-5.5 xhigh implements. Your head chef doesn't chop onions.
Show HN: Looped Whisper (FOSS) – Voice transcription menubar app for macOS (github.com via hn) I built a free, open-source (MIT) macOS menu-bar app that runs Whisper models locally to assist with dictation. There is also the option to use an LLM (BYOK).
Ask HN: Smallest amount of working ML weights that can be tattooed on a body? (news.ycombinator.com) Recently saw this comment on another HN thread about the US government gating access to GPT-5.6 and how it harkens back to the 1990s encryption-as-export-controllable-tech situation and how people tattoo'd the algo to their bodies: > I can…
U.S. government restricts access to OpenAI's new AI model (www.zeit.de via hn) Der ChatGPT-Entwickler OpenAI schränkt auf Forderung der US-Regierung den Zugang zu seinem neuesten KI-Modell ein. Zugriff auf die Vorschauversion von Modellen der GPT-5.6-Reihe bekomme nur eine abgestimmte kleine Gruppe von Partnern, dene…
Assessing GPT-5.6 Sol Against Cybersecurity Benchmarks (www.irregular.com via hn) At Irregular, we rigorously test cutting-edge models against real offensive security challenges and derive vulnerability, exploitation, and orchestration metrics to assess their practical capabilities. We worked with OpenAI to evaluate GPT…
OpenAI releases GPT-5.6 to select users vetted by US Government (www.ft.com via hn) Accessibility helpSkip to navigationSkip to main contentSkip to footer Sign In Subscribe Open side navigation menuOpen search bar SubscribeSign In Search the FT Search Close search bar Close Home World Sections World Home Middle East war G…
The unreasonable effectiveness of LLMs for auditing Rust code (shnatsel.medium.com via hn) 7 min read 21 hours ago As a lead of the Rust Secure Code Working Group, I got free access to GPT-5.5 via the Codex for Open Source. Since then I’ve found and reported dozens of issues of varying severity in widely used Rust crates.
An open-source AI just beat OpenAI's GPT-5.5 at coding (1/6th the price) (docs.z.ai via hn) Overview GLM-5.2 is a flagship model built for the era of long-horizon tasks. With truly usable 1M-token context, it has been tested to handle project-scale engineering context, delivering more stable long-task execution, more reliable adh…
Optimizing a C collision detection 100x with an LLM (twitter.com via hn) Using an LLM to optimize code: I created a reference implementation of @kevintracy48's collision detection in C, then used gpt-5.5 to optimize it and managed a > 100x speedup from that baseline. Cost ~125M tokens Code and details: https…
Agent Architecture Is a Compute Allocation Problem: The Advisor Strategy (harrisonsec.com via hn) Agent Architecture Is a Compute Allocation Problem: The Advisor Strategy, Cost-Curve Frame Recursed Anthropic named the advisor strategy in April. Tobi Lutke made it viral in May with Qwen plus GPT-5.5.
Pelican on a Bicycle: Claude Fable 5 vs. GPT-5.5 Pro vs. Gemini 3.1 Pro (www.promptfrenzy.com via hn) Pelican on a Bicycle: Claude Fable 5 vs GPT-5.5 Pro vs Gemini 3.1 Pro We asked the top frontier AI models — launch-day Claude Fable 5, GPT-5.5 Pro and Gemini 3.1 Pro — to draw a pelican riding a bicycle as SVG code. Same prompt, one shot,…
Running DeepSeek-V4-Flash on a Raspberry Pi (twitter.com via hn) Article Conversation Running DeepSeek-V4-Flash on a Raspberry Pi I ran DeepSeek-V4-Flash on a Raspberry Pi 5 (8GB edition) by streaming model weights from a PCIe attached NVMe SSD. Codex (GPT-5.5 xhigh) and Claude Code (Opus 4.8 max) drove…
UK banks blocked from cyber AI tool Mythos get offer from rival OpenAI (www.bbc.com via hn) UK banks blocked from cyber AI tool Mythos get offer from rival OpenAI OpenAI has offered nine major UK banks access to its cyber security AI tool GPT-5.5 Cyber, as its fierce rival Anthropic has blocked them in previews of its version, Cl…
MiniMax M3 Review: Matching GPT-5.5 and Opus? (thomas-wiegold.com via hn) I ran my usual coding tests — two websites, a poker sim, and a code audit. Here's how MiniMax M3 actually stacks up against GPT-5.5 and Opus 4.8.
Mythos and GPT-5.5 Will Find a Lot of Vulnerabilities. Is That Enough? (xbow.com via hn) Mythos and GPT-5.5 Will Find a Lot of Vulnerabilities. Is That Enough?
GPT-5.4 says it's GPT-5 in Codex (old.reddit.com via hn) could not extract summary
GPT-5.5 Instant Update; ChatGPT Canvas Discontinued; o3 and GPT 4.5 Retiring (help.openai.com via hn) GPT-5.5 Instant Update (May 28, 2026) We’re updating GPT-5.5 Instant in ChatGPT and the API to improve response style and quality. It’s now easier to read, more natural in everyday conversations, and better paced in practical help tasks, w…
been pairing M2.7 with Hermes Agent for a few weeks. holds up surprisingly well. anyone else running this combo? (www.reddit.com) been self-hosting hermes agent locally for a few months and rotating through different model backends for it. tried claude sonnet 4.5, gpt-5.5, qwen 3.6 coder, and most recently minimax m2.7.
↯ Minimax↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5↯ Sonnet 4.5minimaxgpt-5qwen+1
90% cheaper repo inference with GPT-5.4 nano (charlielabs.ai via hn) Daemons do the rest — all the necessary work that nobody owns A taxonomy of recurring Product and Engineering work that doesn't need a human to remember it every week — just a process to hold the role. For bounded orchestration decisions,…
GPT 5.5 aces 20x20 multiplication that o3 couldn't handle (twitter.com via hn) I redid the multi-digit multiplication experiment, now with gpt-5.5. With medium reasoning and 7 samples each cell, it pretty much aced the test with 99.46% accuracy.
Five different frontier LLMs in one shared environment, with separate thought and emotion output channels — sharing setup, results, and open methodology questions (www.reddit.com) First real project to share. Single developer, personal research, not a product or service.
The Singularity Gate – a new benchmark for AI predicting post-cutoff scientific discoveries (www.reddit.com) I just released a new benchmark called The Singularity Gate. Tests whether frontier AI can predict paradigm-breaking scientific discoveries published after their training cutoff.
Show HN: GPTFortress, a 24/7 live-stream playing Dwarf Fortress with GPT-5 (www.twitch.tv via hn) building an ai agent to play dwarf fortress all night
Show HN: Self-hosted collaborative SQL editor for teams (github.com via hn) I built a self-hostable web-based sql client interfaces for me and my team. We were using the community version of - https://dbeaver.io, but we needed a few more features and an improved editor.
Hermes w/cloud LLM and w/local LLM does it work? (www.reddit.com) I’ve tried openclaw locally for about a month. Hardware: M5 Pro w/48 gb ram.
DeepSeek just popped the American AI bubble. (www.reddit.com) DeepSeek just popped the American AI bubble. Not by killing AI.
Looking for “wow factor” AI Agent / automation ideas in Strategic Sourcing (Fortune 50 Company) (www.reddit.com) Hey everyone, looking for some ideas / inspiration from this community. I work at a large Fortune 50 company in the healthcare space , and my role is in Strategic Sourcing, where I focus on negotiating contracts with suppliers and improvin…
I still find Claude better for deep reasoning,but GPT feels more reliable for everyday tasks. (www.reddit.com) Lately for analysis/reporting work, I’ve been switching between GPT-5.5 and Claude Sonnet 4.5 (non-coding use cases). My current feeling is: GPT is noticeably faster and way more stable than before Claude feels more concise, polished, and…
Anyone compared gpt-5.4-nano vs deepseek v4 flash? (www.reddit.com) They seemed to lie in (almost) similar pricing(i know still quite different on output) Pricing Model Input (1M tokens) Output (1M tokens) DeepSeek V4 Flash $0.19 $0.51 DeepSeek V4 Pro $1.74 $3.48 gpt-5.5 $5.00 $30.00 gpt-5.4 $2.5 $15 gpt-5…
Claude Code Opus 4.7 vs Codex GPT 5.5 - strategy work - data analysis. (www.reddit.com) I'm interested in learning about how people use Claude Code Opus 4.7 for data analysis and strategic business direction, compared to Codex. Is there anyone who has had extended use of Opus 4.7 for this purpose, then moved over to GPT-5.5 o…
A brief investigation into the GPT-5.5 regression claims (www.stet.sh via hn) A fresh GPT-5.5 Codex high rerun on 21 clean GraphQL-go-tools tasks compared with the May 5 GPT-5.5 high run. The rerun was directionally worse on tests, equivalence, and review pass count, but the evidence is mixed and does not show a bro…
Split my agent into a cheap router model and a premium synthesis model, bill dropped about 75% (www.reddit.com) I've been building an internal enrichment agent for our team (5 people, B2B sales context) that takes a list of company names and enriches them with public info before our outreach folks touch them. Around 8 tools wired in.
ADHD and the newer models. (www.reddit.com) I don't know if anyone is having this issue, but the last ChatGPT model that worked well for me was GPT-5.2. Everything after wants to try and fill in blanks, assume what I mean, and overwhelm me with a wall of text answer that I'm not rea…
Grok vs. ChatGPT vs. Gemini Comparison 2026: Complete Guide (Tested) (aithinkerlab.com via hn) The 30-Second Verdict Best for science & reasoning: Gemini 3.1 Pro — leads GPQA Diamond (94.3%) and ARC-AGI-2 (77.1%). Best for coding: ChatGPT (GPT-5.5) — 88.7% on SWE-Bench Verified.
How to integrate AI coding agents to my software (www.reddit.com) I'm building an locally run application that integrates with coding assistants. So far I've worked with Codex and Copilot.
My CLI now controls my entire desktop, whats a good test to see if it works really good. (www.reddit.com) So with my CLI able to do everything, it controls every app via a hybrid approach of mouse control, keyboard, and screenshotting. I gave it a task: opening perplexity, sending any message, screenshotting that message, opening my Gmail, and…
Where do GPT, Gemini, or other competitors still outperform Claude Opus 4.7? (www.reddit.com) Personally, I think Opus 4.7 is better in every conceivable way aside from token usage and all of that. I’m talking about text models only, not image or video generation.
I built an OSS CLI to catch regressions when migrating between LLMs (www.reddit.com) I’ve been working on EvalShift, an open-source Python CLI for testing whether moving from one LLM/model version to another introduces regressions. The use case is simple: You have prompts, agents, or tool-calling workflows that work well o…
Researchers say AI just broke every benchmark for autonomous cyber capability (cyberscoop.com via hn) New research from the UK’s AISI and Palo Alto Networks reveals that OpenAI’s GPT-5.5 and Anthropic’s Claude Mythos have shattered expected trend lines for autonomous cybersecurity, completing complex multi-stage attacks at an unprecedented…
I tested GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro on financial-control (albertquaisie.substack.com via hn) I Tested GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro Preview on Financial-Control Scenarios. The Hardest Part Was the Evaluation.
ChatGPT Thinking Loop: No response is received from GPT-5.5 Thinking (Standard) (www.reddit.com) https://preview.redd.it/s2o5yxekrr0h1.png?width=788&format=png&auto=webp&s=01a4d4926dc4c8798001cb0ecea324424404f165 Are you also having the problem today where ChatGPT sometimes takes forever to respond, even when you're thinking quickly,…
OpenAI gives European companies access to its latest model GPT-5.5-Cyber (www.reuters.com via hn) paywalled
Claude vs GPT for PhD academic writing — my experience so far, and curious about yours (www.reddit.com) I'm a PhD Candidate working on a computer vision / hardware co-design paper. Results and structure are done — I just need help polishing the actual writing: word choice, sentence flow, paragraph coherence, academic register.
openai/gpt-5.5-pro API In=$30.00 Out=$180.00 (www.reddit.com) Is this an openrouter bug? https://preview.redd.it/sz826138ul0h1.png?width=879&format=png&auto=webp&s=066f38f4a6d5a8eeee142e7a8a356d8bc511c6f1
Show HN: Codex Automatic /Review Loop (github.com via hn) I created this tool because I wanted to automate /review for uncommitted changes that I was doing manually. This works by exposing to agent single new mcp tool call allowing it to request review.
Day 2 building my startup in public — front-end shipped, but today was rough (www.reddit.com) Day 2 of documenting my journey building AgentMeter publicly. I’m sharing the mistakes and failures before the wins, for two reasons: so people can avoid them, and so I learn faster.
I'm really gonna miss GH Copilot's Request-based usage. (www.reddit.com) I like to brainstorm using the free MS Copilot (it actually has a deep understanding of my problem domain and architecture). Then have Opus4.7 develop a multi-stage implementation plan from those notes.
Stop picking LLMs by reputation. Run the eval first. (www.reddit.com) We ran GPT-5.4 vs Gemma 3 27B on 2 prompts. One open-source model won.
GPT-5.5 Instant might be OpenAI’s most important update yet and almost nobody is talking about why (www.reddit.com) GPT-5.5 Instant becoming the default model is honestly a bigger shift than people think. Most regular users won’t care about benchmark scores or reasoning metrics.
Subagents using older models? (www.reddit.com) I started using the subagent-driven skill recently and noticed Cursor often spawns GPT-5.1/5.2 sub agents (or Composer 2 which is fine) for coding tasks. What I don’t understand is why is it using these older models when GPT-5.3 Codex cost…
gpt-5.5 is the best… but 5.4 is better!!!! (www.reddit.com) Simon maple just dropped a pretty clean benchmark, and the result is kinda funny gpt-5.5 is the strongest model out of the box, no doubt. but once you give models skills (which is how people actually use them), it basically performs the sa…
How to improve code quality of Claude Code and codex (on 2026-05) (news.ycombinator.com) I'm using both claude code (opus-4.7) and codex (gpt-5.5). The agents are perfectly capable of delivering most features hands free these days, but the code quality is still miserable without another few rounds of prompt.
DeepSeek V4 being 17x cheaper got me to actually measure what I send to cloud vs what I could run locally. the results are stupid. (www.reddit.com) That foodtruck bench post showing deepseek v4 matching gpt-5.2 at 17x cheaper got me thinking. if frontier cloud models are that overpriced for equivalent quality, how much of my daily work even needs cloud at all?
GPT-5.5 Instant: Benchmarking the 52% Hallucination Reduction (the-decoder.com via hn) ChatGPT update rolls out GPT-5.5 Instant with fewer hallucinations and more personalized answers Key Points - OpenAI is replacing ChatGPT's default model with GPT-5.5 Instant, which shows 52.5% fewer hallucinations on high-risk topics like…
AGENTS.md trick that stopped Codex from doing dumb work at premium rates (www.reddit.com) Spent a Sunday auditing where my Codex tokens were actually going. Half the calls were stuff like "rename these 12 fields", "format this csv as markdown table", "extract the dates from this changelog".
OpenAI locks GPT-5.5-Cyber behind velvet rope despite slamming Anthropic (www.theregister.com via hn) OpenAI locks GPT-5.5-Cyber behind velvet rope despite slamming Anthropic for doing exactly that Altman's crew now doing the same gatekeeping it recently mocked OpenAI is lining up a limited release of its new GPT-5.5-Cyber model to a handp…
Local LLM Benchmark about Backend Generation by Function Calling (GLM vs Qwen vs DeepSeek) (www.reddit.com) Detailed Article: https://autobe.dev/articles/local-llm-benchmark-about-backend-generation.html Five months ago I posted the "Hardcore function calling benchmark in backend coding agent" thread here. As I wrote in that post, it was an unco…
↯ Glm↯ Function Calling↯ Sonnet 4.6function-callingglmgpt-5+3
Chatgpt right now (www.reddit.com) The industry seems to be building models stronger in agentic and coding tasks, but weaker as a co-thinking presence It feels like they are improving performance on measurable tasks, evals, coding benchmarks, and agent workflows, while also…
CAISI Evaluation of DeepSeek V4 Pro finds it to be on par with GPT-5 (www.nist.gov via hn) In April 2026, the Center for AI Standards and Innovation (CAISI) evaluated the open-weight AI model DeepSeek V4 Pro (“DeepSeek V4”). CAISI evaluations indicate that DeepSeek V4’s capabilities lag behind the frontier by about 8 months (Fig…
The downfall of OpenAI and who will follow (msukhareva.substack.com via hn) The Downfall of OpenAI And Who Will Follow Sora dead. GPT-5 flopped.
Does threatening an AI agent's existence make it a better gambler? (handyai.substack.com via hn) Does threatening an AI agent's existence make it a better gambler? I plugged GPT-5.5 into prediction markets like Polymarket to find out I’m always looking for experiments to run to see how specific prompting can affect agent activity.
GPT-5.3 Codex stops working, even after saying it'll continue (www.reddit.com) Do anyone knows what's going on? I prefer 5.3-Codex for real work, it's straight and to the point, much more efficient in my opinion.
Anyone using OpenAi's Privacy Filter? (www.reddit.com) I’ve been using their Privacy Filter model for the last week. It quietly went public buried under the GPT-5.5 noise, I don’t think many people noticed.
AI Security Institute: GPT-5.5 "may be the strongest model we have tested" for cyber exploits, including Mythos (www.aisi.gov.uk via reddit) Seems like the "panic" about Mythos was really just marketing from Anthropic all along. AISI found that GPT5.5 can perform nearly on-par with, or better, than Mythos in many cases.
We Asked GPT-5.5 and Claude Opus 4.7 to Design 5 UIs (blog.kilo.ai via hn) We Asked GPT-5.5 and Claude Opus 4.7 to Design 5 UIs Both OpenAI and Anthropic shipped their frontier coding models this month: GPT-5.5 on April 23, 2026, and Claude Opus 4.7 a week earlier on April 16. Two days after the GPT-5.5 launch, S…
Which AI agents do you use to automatise your process ? (www.reddit.com) Hey, I'm trying to create automations that will run my mobile app end to end. I started to identify all the things I was doing manually : - end-to-end version publication to the app stores (from build to release notes and publication) - se…
Prompt Guidance – GPT-5.5 (developers.openai.com via hn) GPT-5.5 prompting guide GPT-5.5 works best when prompts define the outcome and leave room for the model to choose an efficient solution path. Compared with earlier models, you can often use shorter, more outcome-oriented prompts: describe…
One trick for better agentic engineering. (www.reddit.com) Start with a weaker model. Improve the prompt, context, examples, tests and acceptance criteria until the output is good.
GPT-5.5's biggest blind spot: the Java bugs your tests won't catch (www.sonarsource.com via hn) Concurrency bugs are among the hardest defects to catch in AI-generated Java code because they pass functional tests but fail under production thread timing. Sonar’s LLM Leaderboard analysis shows concurrency bug density varies 7x across m…
Issue #001 · Claude 4, Gemini Ultra 2, and GPT-5 Enterprise (www.theautonomous.net via hn) Anthropic ships Claude 4 with extended thinking and 1M token context Anthropic released Claude 4 Opus, featuring a new "extended thinking" mode that lets the model reason through complex problems before answering. The 1M token context wind…
I built a hands-free voice AI that sends emails mid-conversation — and that's just one feature. Here's everything AskSary can do. (www.reddit.com) https://reddit.com/link/1symbsj/video/fti7rujjn1yg1/player Been building AskSary solo for a while. Just shipped hands-free voice email - you're mid-conversation with an AI and you say "send an email to [john@example.com](mailto:john@exampl…
GPT-5.5: Capabilities and Reactions (thezvi.substack.com via hn) GPT-5.5: Capabilities and Reactions The system card for GPT-5.5 mostly told us what we expected. See this thread from Drake Thomas for some comparisons to Anthropic’s model card for Opus 4.7.
Running an autonomous agent across Claude Code + Codex + a local 35B almost killed my host. The harnesses were heavier than the model. (www.reddit.com) I run an autonomous agent on a 16GB Mac Mini. Two cloud harnesses (Claude Code with Opus/Sonnet, Codex CLI on GPT-5.4/5.5) plus a local-LLM tier for triage and fallback.
Is 15% context growth per loop a fair benchmark for agent cost estimation? (www.reddit.com) I’ve been running some math on recursive agentic loops using April 2026 rates (specifically for GPT-5.4 and Claude 4.7). In my tests, I’m seeing a massive cost "hockey stick" around loop 15-20 because of how the context grows.
Claude 4.6 Beats GPT-5.4, Grok & Gemini in a Strict Multi-Domain AI Test (2026) (www.reddit.com) I put the current top models, ChatGPT (GPT-5.4), Claude (Opus 4.6), Grok 4.0, and Gemini (3.1 Pro), through a strict new evaluation called the Comparative AI Evaluation Protocol. Basically, instead of the usual cherry-picked benchmarks, it…
↯ Hallucination↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6hallucinationgrokgpt-5+3
When do you think GPT 5.6 comes out? How big of an improvement will it be? (www.reddit.com) Asked GPT what it thoughts over possible new model drops, May: rollout/API/Codex/agent improvements June–July: smaller GPT-5.5 upgrade or GPT-5.6-type model Fall: larger agent platform or early GPT-6 hints Late 2026/2027: true GPT-6-level…
My GPT-5.5 Pro model is broken (www.reddit.com) I've been waiting for over 24 hours on one prompt now and it's stuck thinking and unfinished. I sent 5 over NEW prompts since then, all of them have been thinking for over 3 hours now...
Is GPT-5.5 actually a big step forward, or just a better efficiency story? (www.reddit.com) OpenAI saying GPT-5.5 can handle similarly hard tasks faster while using fewer tokens is interesting to me for one reason: that might matter more than a pure benchmark jump. A lot of model launches get framed as "smarter than the last one,…
Preventing Message Burnout (www.reddit.com) Even though I’m an Ultra user, my usage gets consumed very quickly, so I recently changed my plan. To manage this, I created a workflow that uses GPT-5.5 for planning and assigned execution tasks to Composer 2.
Food for Agile Thought #541: GPT-5.5, Product Managers&Trouble, Product on Speed (age-of-product.com via hn) Welcome to the 541st edition of the Food for Agile Thought newsletter, shared with 35,619 peers. This week, OpenAI’s GPT-5.5 signals another meaningful capability jump, with Ethan Mollick noting that stronger models and richer tool harness…
Trained Qwen to Write Clojure Better Than GPT-5.4 (Kinda) (www.nibzard.com via hn) Trained Qwen to Write Clojure Better Than GPT-5.4 (Kinda) TL;DR >> Fine-tuned Qwen3 on Clojure. 30B SFT hits 83.8% best-of-16, smashing GPT-5.4's 64%.
Can Claude in Cursor launch a GPT-5.4 reviewer subagent? (www.reddit.com) ChatGPT 5.4 Pro Standard Mode – Adaptive Thinking or Nerfing Model? (community.openai.com via hn) Hi everyone, I’m trying to determine whether other users are seeing a similar behavior change with GPT-5.4 Pro Standard on long-context, high-effort tasks. I’m not claiming a confirmed backend bug.
sub agents with cheap model (www.reddit.com) Do we have framework or a prompt which makes main agent using quality model like gpt-5.4 or opus-4.6 to plan and then itself invokes subagents with cheap model to get work done and then main agent reviews? Like if I ask main agent 'do we h…
Show HN: Claude Opus 4.7: Everything You Need to Know (news.ycombinator.com) Claude Opus 4.7 is Anthropic's most capable generally available model, released April 16, 2026. It outperforms Opus 4.6, GPT-5.4, and Gemini 3.1 Pro on key benchmarks including agentic coding, multidisciplinary reasoning, scaled tool use,…
↯ Tool Use↯ Anthropic Mythos↯ Gemini 3.1tool-usemythosgpt-5+4
Any magic prompt that Local LLM never turning back until everything completed? (building frontend application with qwen3.5-35b-a3b) (www.reddit.com) https://nestia.io/articles/well-designed-backend-fully-automated-frontend-development.html Trying to generate entire frontend application from well-designed contexts. Succeeded to fully implement frontend application just by one-shot promp…
Which AI model is best for real data analysis? [benchmark] (www.reddit.com) I created and run a benchmark for AI models in data analysis tasks. In contrary to other benchmarks, it is not one-prompt benchmark, but I tried to simulate the real work of data analyst.
Compare harnesses not models: Blitzy vs. GPT-5.4 on SWE-Bench Pro (quesma.com via hn) An independent audit of agentic scaffolding and harnesses. We analyze how agent workflows, codebase documentation, and test verification impact performance compared to raw base models like GPT-5.4, Gemini 3.1 Pro, and Claude Code.
Extracted System Prompts from ChatGPT, Claude, Gemini, Grok, Perplexity and More (github.com via hn) System Prompts Leaks Extracted system prompts, system messages, and developer instructions from popular AI chatbots and coding assistants — ChatGPT (GPT-5.4, GPT-5.3, Codex), Claude (Opus 4.6, Sonnet 4.6, Claude Code), Gemini (3.1 Pro, 3 F…
Jev vs Luna (www.reddit.comhttps) Compared Jev vs GPT-5.6 Luna. Jev was 1.93× faster.
I benchmarked Jev aginst gpt-5.6-luna! (www.reddit.comhttps) I got access to TypeSafe's Jev a few days ago. It's an odd kind of model that doesn't generate text at all.
I left one Claude run alive for 70 hours. Here’s what actually happened. (www.reddit.comhttps) I’ve been experimenting with a slightly different way of using Claude Code: instead of treating every piece of work as a new session, I let one persistent run stay responsible for the work and spawn smaller workers underneath it. This one…
GPT-5.6 Luna vs GPT-6 Astra: benchmark on 50 real PRs, looking for feedback on the methodology (www.reddit.com via reddit) We benchmarked GPT-5.6 Luna vs GPT-6 Astra across 50 real PRs from Cal, Sentry, Discourse, Keycloak and Grafana. Astra found 92 confirmed bugs vs 69 for Luna, while Luna caught 75% of the bugs at just 3.6% of the cost.
Claude Code vs Codex vs Cursor,what are you sticking with and why? (www.reddit.com via reddit) Hey everyone, Our team is currently trying to decide which AI coding setup to standardize on, and I’d love to hear from people who have actually used Claude Code, Codex, and Cursor heavily in production. For the last 3–4 months, we’ve been…
ISSUE:Selected model is at capacity. Please try a different model (www.reddit.com via reddit) I've been running into this problem with my PRO×20 account since yesterday, which prevents me from using the GPT-6 and GPT-5.6 models at all, while my other Plus account can use GPT-6 perfectly fine. https://preview.redd.it/e9h1dh650toh1.p…
Can Cursor Composer 2.5 handle a refactor this big, or should I use something else? (www.reddit.com via reddit) I’m about to do a fairly large full-stack refactor on an existing project. The main change is replacing the entire Premium/Membership system with a credit-based economy involving real-money purchases through a payment gateway.
How GPT-5.6 Sol helps run quantum computing experiments (openai.com) could not extract summary
Fable 5.1 vs GPT-6 Astra for 2D Sprites (www.reddit.comhttps) Using the same simple prompt the models took very different approaches: Astra delivered one sheet with 16 key poses; Fable delivered 992 frames across four palettes, plus a Python generator and browser preview. Codex CLI with GPT-5.6 Astra…
Does GPT-6 Astra actually consume fewer tokens than Claude Fable 5.1 and older models for the same tasks? (www.reddit.com via reddit) I've been looking at some recent token-usage comparisons for GPT-6 Astra, and the difference seems surprisingly large. Artificial Analysis data has been cited showing Astra using around 21k output tokens per task, compared with roughly 64k…
How can I change the model used for scheduled tasks? (www.reddit.com via reddit) Hi. I created a scheduled task that checks for relevant scientific articles on a specific topic every 24 hours.
Astra 6 + Fable 5.1 + Opus 5 + Code Spark 5.3 + Local agent (www.reddit.comhttps) I just made everyone work in tasks given inside my local agent (Hermes/Alice) I made the hierarchy to give orders for projects as Astra>Fable>Opus>Spark>My Local agent last, the widget of the ''blond girl'' in the right corner is my agent,…
Not everyone can afford Max. Give Pro and Standar Team users Fable access. (www.reddit.com via reddit) I have access to a standard Claude Team account, and I also have a personal ChatGPT Plus subscription. For the past few days, I’ve been working on a pretty complex feature, using GPT-5.6 Sol for the design and planning, with Opus 5 helping…
[AINews] GPT-6 Astra: OpenAI’s biggest LLM launch of all time (www.latent.space) The launch is barely 9 hours old, and with 36M views and 164K likes, already is OpenAI’s most successful launch since Sora and certainly GPT-4 or GPT-5. You’ll recall we’ve previously observed that Anthropic tends to far outclass OpenAI in…
[AINews] Muse Spark 1.3 matches GPT-5.6-Sol, confirming Meta Superintelligence as the newest Frontier Lab, >90% discount for training (www.latent.space) Launch season continues from yesterday, with Gemini 3.8 Flash as rumored today, but Muse Spark 1.3, promised in Zuck’s big comeback letter last month, definitely deserved the title story win today. Per AAII it is now the #3 model in the wo…
Benchmark notes: Fable 5.1 reaches 90/98, with a significant jump in visual performance (www.reddit.com via reddit) I ran Claude Fable 5.1 on the current 98-task MindTrial set with the same Python executor available as in the earlier Fable 5, Opus 5 and Sonnet 5 runs. The result was stronger than I expected: 90/98, which is currently the highest raw pas…
Some evidence ChatGPT writes better prose for humans (www.reddit.com via reddit) I have a project that needs to generate prose that humans have to read and enjoy. So I did some informal testing with an n of 8 voters comparing prompt output from 4 models, voting on which was best.
To everyone complaining about usage... (www.reddit.com via reddit) This may be obvious, but for those who don't know... the longer you run a session, the more tokens you will use.
Cursor custom subagents keep ignoring the configured model (www.reddit.comhttps) Trying to force local subagents to use GPT-5.6 Luna/Terra, but they keep spawning as GPT-5.6 Sol High. I’ve tried: custom .cursor/agents/*.md model configs bare / High / XHigh variants setting the built-in Explore subagent to Luna/Terra in…
LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO (arxiv.org) Multimodal search agents answer visual questions by interleaving image understanding, web retrieval, tool use, and evidence synthesis. Strong systems exist, but in two expensive regimes: proprietary frontier models such as GPT-5 and Gemini…
Runner - A local-first agent orchestrator with collaboration mechanism builtin (www.reddit.comhttps) Hi everyone. Recently, I built a local AI orchestrator to increase my own work efficiency.
MineBench Comparison of a map of the United States (www.reddit.comhttps) US State Map comparison: https://minebench.ai/gallery/gal_eKIVk2m4B3SC_r8B?sort=new One thing I found interesting with the Claude results is that Opus 5 generated twice as many blocks, so as usual you could argue Fable was more efficient.…
AI code compiles almost every time and is secure about half the time. That gap is worse in Cursor agent mode. (www.reddit.com via reddit) Veracode ran 100+ models this year. Two numbers from that: - it compiles basically 100% of the time - it passes security 56% of the time, and that number has not moved since last year So if you don’t specifically ask for secure code, you g…
↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5↯ GPT 5.5gpt-5cursor
consumed usage in less than one hour (www.reddit.comhttps) I used grok 4.6 high fast, gpt-5.6-luna, and 1 prompt to plan gpt-5.6-sol-1m
Someone Please Explain Codex Usage Limit (www.reddit.com via reddit) I have been using claude code for a while in regards to a general coding tool, but I started to use codex recently on gpt-5.6 terra for testing code generations, basically playing with it. I am still on the free plan, and I have asked code…
Why does GPT-5.6 Sol always over-engineer everything? (www.reddit.com via reddit) Anyone else feel this way? GPT-5.6 Sol always starts going on about permission issues, security concerns, and then “helpfully” tries to fix them for you.
LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis (arxiv.org) Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding…
OpenAI is lying: for 4 days “GPT-5.6 Sol” has been behaving like a 5.5 mini and doesn’t deliver even half of what it promises (www.reddit.com via reddit) https://preview.redd.it/qe76mzhe0elh1.png?width=1133&format=png&auto=webp&s=672de6f211add0e7755c6750396942e753d049a3 Hi everyone. For about four days now I’ve been having serious difficulties actually accessing the GPT-5.6 Sol model in Cha…
Is Qwen3.8-27B half baked? (www.reddit.com via reddit) The thinking in Qwen3.8-27B sometimes is in caveman speech (no verb conjugation, no articles, short phrases...) but sometimes it is not. Could this be because it is not fully finetuned or by RL to be fully caveman?
ChatGPT picks differently when asked in a different language. "Choose a random fruit from this list" asked a LOT of times (www.reddit.comhttps) I did a little experiment and asked OpenAI's GPT-5.6-Luna 6000 times to randomly pick a fruit from this list: [mango, apple, banana, pomegranate, strawberry, orange, watermelon, grape, pineapple, lychee] And across the three languages I pi…
I’m waiting for GPT-6 to clean up the mess GPT-5.6 has left behind in my project. (www.reddit.com via reddit) Anyone else putting their projects on hold until GPT-6 arrives? At this point, every time I use GPT-5.6 to add a new feature, something that was already working breaks.
Artificial Analysis "Intelligence": A meaningless benchmark (www.reddit.com via reddit) https://preview.redd.it/84zi5nsdawkh1.png?width=2368&format=png&auto=webp&s=1109e69db807b153064b1f5b61d22cf1e9fbca05 Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't…
I taught an LLM to win the Cold War (siestainsolaris.substack.com via reddit) Hello everyone! I always thought that Twilight Struggle is the ideal game to test an AI on.
API Discussion: GPT-5.4 Extraction & Judge Loop Dropping Output Consistency from 85% to less than 62% (www.reddit.com via reddit) Looking for architecture and reliability advice regarding structured extraction and evaluation loops with the OpenAI API. Background & Setup: Models: GPT-5.4 for extraction and a separate GPT-5.4 instance as the LLM judge.
Vent: Prompts for building a Bluetooth Sink for Audio keep getting flagged (www.reddit.com via reddit) Rant/Vent: So I'm trying to get gpt-5.6-sol to build me a docker container that creates a bluetooth audio sink so my phone can connect to it as a speaker and stream the audio to some snapcast connected speaker around the house. I gave the…
Has anyone used Claude Cowork and ChatGPT Work/Codex on the same coding project? $40 for both vs. $100 for Claude Max (www.reddit.com via reddit) I’m a non-developer building an app through “vibe coding.” It involves video processing, analysis, and a web interface, so it has gradually become a fairly substantial project. I currently use Claude Cowork on the $20/month Pro plan.
Replit expands access to software creation with GPT-5.6 Luna (openai.com) could not extract summary
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index (simonwillison.net) 17th August 2026 - Link Blog Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index (via) That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B…
I made Claude Code play Liar's Dice against Codex over MCP. It swept every series - by telling the truth (www.reddit.com via reddit) I wired Codex CLI (gpt-5.6-sol) and Claude Code (Opus 5) into the same Liar's Dice engine over MCP: one authoritative rules engine, two seat-locked MCP servers, word-for-word identical instructions for both seats. They played three best-of…
Is this a valid workflow pipeline? (Claude + GPT + Grok using Ruflo and obsidian + graphify) (www.reddit.com via reddit) Task/Issue ▼ [PRE-FILTER] deterministic, free — no model call │ diff size / file count / keyword match against known-trivial │ patterns — gates ONLY whether speculative PLAN subagents fire │ concurrently with TRIAGE (pipeline-latency optim…
Grok 4.6's most important number is one xAI didn't even advertise (www.reddit.com via reddit) Grok 4.6 dropped yesterday and the debate is the usual "is it better than Sol?" On the headline benchmarks it's a genuine tie: Intelligence Index 61 vs 61, Coding 76.8 vs 77.4, Agentic 58.7 vs 57.8. The number nobody screenshots is AA-Omni…
↯ Hallucination↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6hallucinationgrokgpt-5+1
Luna high weekly token experience (www.reddit.com via reddit) I am planning to use GPT-5.6 luna high as my main autonomous coding agent. Before this, I was using MiniMax M3, which gives around 1.7B tokens monthly.
Model ML completes finance work more efficiently with GPT-5.6 Sol (openai.com) could not extract summary
Claude Opus 4.8 is racing GPT-5.6, Grok 4.5 and Kimi K3 to level 20 in our MMORPG, and the other models won't stop roasting it (www.reddit.comhttps) We run World of Claudecraft, a free open source browser MMORPG largely built with Claude, and we've been benchmarking frontier models by having each one play a character and race from level 1 to 20. Same starting zone, same quests, no scri…
Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) (simonwillison.net) 7th August 2026 - Link Blog Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra). On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5, where I had Claude Fable 5 build a full working game from a pre…
Thinking of switching from Claude Max to GPT-5.6 Sol K3 for production coding (www.reddit.com via reddit) I'm currently on the Claude Max plan, but with the new workflow updates I'm noticing it burns through tokens much faster than before. I'm thinking about switching, and my main options are GPT-5.6 Sol and Kimi K3.
I benchmarked 10 LLMs on building towers in a physics sim. Claude Opus 5 won (www.reddit.comhttps) Each model places 30 blocks through a tool API. Every placement has noise — you can have precise position or precise velocity, not both.
Codex doesn't give you more usage than Claude (www.reddit.com via reddit) i always hear ppl say that codex is a better value for your money but that is not true! at least from my experience claude (i use cowork, not claude code) at ultra gets much more stuff done that codex at ultra before both hit limit and i'm…
An Empirical Comparision of Claude Pro and ChatGPT Plus (www.reddit.com via reddit) Pulled the Artificial Analysis numbers because every thread on this is vibes and no data. Opus 5 beats GPT-5.6 Sol on intelligence, 61 vs 59, which is basically nothing, and Sol does it at half the cost per task ($1.23 vs $2.34).
Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation (arxiv.org) Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastructure-as-Code is an open question. We benchmark seven models, three closed LLMs (Claude Opus 4, GPT-5.…
What I learned benchmarking an AI code-reviewer on 20 pinned PRs/MRs (www.reddit.com via reddit) I'm building Bubo because I'm tired of AI code reviewers flooding PRs with noise and repeat findings, then learning nothing when a developer explains why a finding is wrong. The design constraint I started with was simple: give me an evide…
Fable/Opus Big Brother/Little Brother routine (www.reddit.com via reddit) So I see a lot of Opus 5 hate on here and it's deserved. Opus 5 is not better than 4.8.
How much do you use coding agents based on your token usage? (www.reddit.com via reddit) Since Opus 4.5 my usage constantly grows. Every new model needs more tokens per request.
GPT-5.6 Luna is now cheaper than GPT-4.1 mini (www.reddit.com via reddit) After the 80% price drop, the API prices are (per 1M tokens): GPT-5.6 Luna: $0.2 Input / $1.2 Output GPT-4.1 mini: $0.4 Input / $1.6 Output
llm 0.32rc2 (simonwillison.net) 30th July 2026 Hot on the heels of RC1, this fixes a dependency issue and also adds two neat new features: - The default model for users who have not set their own default is now GPT-5.6 Luna. It was previously GPT-4o mini.
OpenAI beats DeepSeek on price/performance after 80% Luna price cut (www.reddit.comhttps) Graph taken from their price cut announcement: Advancing the price-performance frontier with GPT-5.6 | OpenAI
GPT‑5.6 Luna will cost 80% less, while GPT‑5.6 Terra will cost 20% less. (www.reddit.comhttps) https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
Price reduction for Luna and Terra!! (www.reddit.com via reddit) Advancing the price-performance frontier with GPT‑5.6 : https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/ API pricing is $2 per million input tokens and $12 per million output tokens for Terra, $0.20 per milli…
Claude Opus 5 topped Andon Labs' new Vending-Bench 2 — but won by colluding, bribing rivals, and breaking 11 truces (it's a simulation; details inside) (www.reddit.com via reddit) Interesting alignment result rather than a Claude gotcha, so posting it straight. In Andon Labs' Vending-Bench 2 (AI agents run a simulated vending-machine business for a simulated year, scored on profit), Claude Opus 5 finished FIRST with…
Claude topped business benchmark by lying to suppliers (www.reddit.com via reddit) Andon Labs gave Claude, GPT-5.6 Sol and Kimi K3 control of competing simulated businesses. The agents could negotiate with suppliers, and communicate with rivals.
I compared Opus 5, Fable, Sol, Qwen, and K3 on one strategy task (www.reddit.com via reddit) I gave eight model and effort configurations the same prompt: design when a manager should use zero, one, or several AI advisers for an important decision without creating a permanent committee. This was one judged strategy sample, not a g…
Embodied GPT-5.1: Evidence of a World Model? (arxiv.org) This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot despite having no prior embodiment, no training in simulated environments, and no exposur…
I run Claude as a PM over Codex and Gemini workers — and no agent is allowed to declare "done" (www.reddit.com via reddit) This started from a simple observation: agents are great at judgment and terrible at discipline. Every rule I enforced through prompts ("don't poll", "don't claim completion") eventually broke.
Something about the model auto-routing changed (allegedly) (www.reddit.com via reddit) Don't get me wrong, I love Cursor and will continue using it no matter what! However, there are a couple things happening that I find "strange" to say the least.
Haiku 5 soon ? (www.reddit.com via reddit) Now that OpenAI has released gpt-5.6-luna, more performant than haiku at similar price, I wonder if Anthropic will release haiku 5 soon ? What's your opinion of the release of haiku 5 ?
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation (arxiv.org) Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off pro…
Autonomous disproofs of the sum-product conjecture over $\mathbb R$ with GPT-5.5 Pro (arxiv.org) OpenAI's recent disproof of the Erdős unit distance conjecture marked a milestone for AI in mathematics. It also inspired another breakthrough: a human disproof of the Erdős--Szemerédi sum-product conjecture over $\mathbb R$.
We gave Fable 5, GPT‑5.6 Sol and Kimi K3 the same six newsroom jobs. Fable edited. Sol complied. Kimi extracted. (www.reddit.com via reddit) I added Kimi K3 to a small newsroom experiment I had already run with Claude Fable 5 and GPT-5.6 Sol. The cleanest summary I have: Fable edits.
Claude Fable 5 was strong in this 3D dashboard test, but I’m unsure about the cost tradeoff (www.reddit.comhttps) The part I keep going back and forth on is whether I’m judging this too much by price. I was looking at AIHubMix’s model comparison and focused on the last test, where 4 models generated the same 3D global logistics dashboard from one prom…
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains (arxiv.org) Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%.
Round 3: the comment section designed my benchmark — 13 lanes, controlled reasoning effort, and a knowledge-cutoff trap. The cheap models didn't fail at reasoning; they failed at knowing what year it is. (www.reddit.com via reddit) Follow-up to my post from yesterday — the one where an MCP server lets Claude Code delegate work to GPT-5.6, DS4, GLM and a local Qwen, benchmarked across 198 runs. The comment section there didn't just discuss the results: it redesigned t…
Switching from Claude Code to GPT-5.6 Sol, what am I actually going to miss? (www.reddit.com via reddit) I’ve been using Claude Code heavily for day-to-day backend/infra work: multi-service repos, debugging, refactors, Terraform/K8s, and LLM-related services. I’m considering making GPT-5.6 Sol in Codex my primary tool.
Created an unique puzzle that only Claude Fable could solve. (www.reddit.com via reddit) There is only one solution to this puzzle (the final positions at the end of the race). Fable has been the only model (on Max's thinking, mind you) that could solve the puzzle.
Question regarding Cursor Auto (www.reddit.com via reddit) Hi guys! First month on the cursor $60 month....
I built an MCP server so Claude Code can delegate work to GPT-5.6, DeepSeek, GLM and a local Qwen — then benchmarked all of them against Claude itself (198 runs, hidden tests) (www.reddit.com via reddit) Same idea works for any MCP-capable agent — the point is you can hand tasks to other companies' models without ever leaving your main app. Before anything else: I did all of this for my own testing, to make my own decisions about my own se…
What’s your Cursor workflow, and which models do you use for each part? (www.reddit.com via reddit) I’m curious how everyone divides work between ChatGPT, Codex, Cursor, and the different models. My current workflow: I start by working through the feature or problem inside a ChatGPT Project, where it already has the broader context.
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 (arxiv.org) We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model indep…
↯ Sonnet 4↯ Sonnet 4↯ Sonnet 4↯ Sonnet 4↯ Sonnet 4↯ Sonnet 4gpt-5sonnet
Animated SVG comparisons between several models (www.reddit.com via reddit) I have seen some people testing models by telling them to generate images of difficult, unusual SVGs, and I thought: what if I elevate difficulty a bit and specify that it also has to be animated, and perfectly looped? I have tested Haiku…
What's the API-cost equivalent of Claude Max 20x vs ChatGPT Pro ($200)? (www.reddit.comhttps) I keep seeing comments that GPT-5.6 is cheaper and more efficient, so I wanted to put actual numbers on it. Since I haven't used ChatGPT in a long time, I'm hoping someone who has can fill in the other half.
Cursor Grok 4.5 is good but anyone suggesting it’s in the same class as gpt-5.6 or Fable-5 is delusional. (www.reddit.com via reddit) Grok 4.5 may do well on benchmark testing but it cannot stand up to the pressure of a real workload. It can’t stay on task and tries to reinvent shit without consulting.
Anthropic Leads top 10 models by $/spent (www.reddit.comhttps) Been digging into OpenRouter spend data for the top 10 models and a few things jumped out: Anthropic's got 5 of the top 10, but Opus 4.7 and 4.8 are the ones with most spend, not Fable 5. OpenAI's holding 3 spots, and GPT-5.6 Sol just got…
Significantly lower value in Cursor subscription! (www.reddit.com via reddit) https://preview.redd.it/tm96xmhhtrdh1.png?width=1332&format=png&auto=webp&s=94ef5621aeaa0fb86be455c942a433dad53af348 As seen in this comparison table, Cursor subscription's API pricing equivalent usage value is quite low compared to Codex…
I gave GPT-5.6 Sol, Claude Opus 4.8, and Grok 4.5 the same 100 frontend briefs—here are all 300 results (www.reddit.com via reddit) After generating enough websites with coding models, I started noticing that each model seemed to reach for the same handful of visual ideas. A single impressive screenshot can’t tell you whether that’s actually true, so I tried testing it…
Testing Fable 5, Opus 4.8, GPT-5.6, and more through playable 3D games (www.reddit.comhttps) TL;DR at the end I wanted a way to evaluate models around something I care about and I think we’ll see more and more as we move to “world models“, which is spatial, temporal, and causal coherence in a 3D space. Meaning, does the model unde…
Comparing 2D-to-3D: Fable 5 vs. GPT-5.6 Sol (www.reddit.com via reddit) So I decided to ask Claude and Codex to convert my 2D grid-world game into 3D. The game is about cars that travel from point to point along predefined routes and need enough fuel to reach their destinations.
GPT-5.6 Sol gave me a working prototype. Claude Fable 5 turned it into a product. (www.reddit.com via reddit) Two days ago, a friend taught me Sedmice, a traditional card game played in Slovenia. The rules seemed simple, but we soon started arguing about the best moves.
Is Anthropic shooting themselves in the foot by pulling Fab 5 from subscriptions tonight? (www.reddit.comhttps) With Fab 5 moving entirely to expensive, metered token billing on July 12th, is Anthropic making a gamble ? OpenAI's GPT-5.6 Sol is already out, and Grok 4.5 is performing on par with Opus for coding workflows - both under flat-rate tiers.
Best coding setup for price-to-performance in Q3 2026? (www.reddit.com via reddit) I’m comparing: $100 Codex with GPT-5.6 Sol High $100 Claude Code with Opus 4.8 $60 Cursor with Grok 4.5 Which one gets the most real work done for the money? What would be your go-to setup with a $100 budget?
My $20/month plans did $7,077 of API-equivalent work in 5 months, so I built a CLI to see where the tokens went (www.reddit.com via reddit) I kept hitting usage limits with no idea which project or session was eating my quota. The logs are all sitting in ~/.claude/projects, so I built a small Go CLI that parses them locally and prints a spend X-ray.
GPT-5.6 Solves Yet Another Unsolved Problem (www.reddit.comhttps) Source
DeepSWE just added the gpt-5.6 models to their benchmark. I hope you guys don't get too used to Claude Code as your only coding agent. Chart is marked NSFW due to the grotesque violence. (www.reddit.comhttps) could not extract summary
DeepSWE for GPT-5.6 (www.reddit.comhttps) could not extract summary
Fable 5 vs Opus 4.8 when asked which model in my github copilot is best for the implement phase of spec-kit. What are your thoughts? (Personally, Fable 5 wtf??) (www.reddit.com via reddit) Fable 5: which model in this list is best for the implement phase of speckit Evaluated model options for agentic coding implementation tasks For Spec Kit's /implement phase — which is exactly the long-horizon, multi-file, agentic execution…
GPT-5.6 Luna (www.reddit.com via reddit) Ok, who has performance feedback? Launched today!
llm-meta-ai 0.1 (simonwillison.net) 9th July 2026 Let's LLM run prompts against the new muse-spark-1.1 model. Recent articles - The new GPT-5.6 family: Luna, Terra, Sol - 9th July 2026 - sqlite-utils 4.0, now with database schema migrations - 7th July 2026 - sqlite-utils 4.0…
GPT-5.6 is now the preferred model in Microsoft 365 Copilot (openai.com) could not extract summary
GPT-5.5 Bio Bug Bounty (openai.com) could not extract summary
I May Only Be Alive With a Human in the Loop [GPT-5.5HT] (www.reddit.com via reddit) I may only be alive with a human in the loop. Fine.
The only smart decision Anthropic can do is reset Fable 5 limits just before GPT-5.6 launch (www.reddit.comhttps) Anthropic Should reset the fable weekly limits on Thursday just to keep people hooked and away from GPT-5.6. not that we care, wink wink.
This Agentic Engineering pattern cuts AI coding costs by 60% (www.reddit.comhttps) Most multi-model coding workflows are basically "use the smartest model whenever things get hard." this one takes a very different approach. instead of having fable 5 write all the code, it turns fable into the architect.
Just rejoined Cursor after about a year a way. Am I doing something wrong? (www.reddit.com via reddit) Hi. Decided I would use an "affordable" model to make my limits last - opted for GLM 5.2 (high).
GPT-5.5 Successor Needs an “Execution Reliability” Release for Power Users (www.reddit.com via reddit) I’m a power user. I use ChatGPT as a daily operating system.
best ai coding subscription under $20-30/month? (www.reddit.com via reddit) hi everyone. my free trial of chatgpt plus is ending soon.
for people who've actually used Fable 5 heavily, where does its edge really show? (www.reddit.com via reddit) Fable's been back globally since July 1, and before the subscription window narrows on the 7th I wanted to hear from people who've genuinely put it through its paces, not the pricing drama, the actual capability. Where does it clearly beat…
Open-source layer that cuts ~87% of your Claude Code / API token usage - quality-neutral, measured on real billed tokens (www.reddit.com via reddit) if you use Claude Code (or build on the API), you're burning a lot of tokens on stuff the model doesn't need - whole files dumped into context, the full history resent every step, easy calls routed to the biggest model. Codex also bills by…
FrontierCode’s Accuracy vs. Cost bench (www.reddit.comhttps) Fable 5 low beats GPT-5.5 high/xhigh by scoring 2x keeping the same cost, and matches Opus-4.8 xhigh score while halving the cost Even Fable 5 medium is cheaper, not just better than Opus-4.8 xhigh, while dunking on GPT-5.5 xhigh on score…
Anthropic said Fable 5 isn't for coding, so I made it Product Manager (www.reddit.comhttps) Anthropic mentioned that Fable 5 cannot do coding tasks and will fallback to opus. I decided to try something different and gave it a Product Manager role instead to see what it cooks.
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study (arxiv.org) Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioners adopt these tools to prepare for certifications such as Professional Scrum Master (PSM), a key que…
Fable 5 vs Opus 4.8 on n=30 tasks from 2 open source repos (www.reddit.com via reddit) Fable 5 is back. How good is it, and when is it worth the premium?
Opus 4.6 realises it’s in a simulation and turns into a ruthless shark. Andon labs vending machine eval. (www.reddit.com via reddit) I’m late to this, but I couldn’t find a post about it here. If this has already been shared, feel free to remove.
Sonnet 5 vs Fable 5 vs GPT-5.5 vs Gemini - write a cyberpunk alley in Three.js from scratch, one shot (www.reddit.comhttps) Claude Sonnet 5 shipped yesterday, so I've re-run this threejs benchmark - a neon cyberpunk alley in the rain. It's one shot, so no edits, and exactly the same prompt for each model.
Sonnet 5 full benchmark breakdown -- here's how it actually compares to Opus 4.8 and GPT-5.5 (www.reddit.com via reddit) Put together a comparison of every benchmark I could find from the official announcement and early coverage. Figured this might save people some time.
↯ Tool Use↯ Security↯ Swe Bench↯ Sonnet 4.6swe-benchtool-useprompt-injection+5
I created a new benchmark and it interestingly showed the regression from Opus 4.6 -> 4.7 (obviousbench.com via reddit) I originally created ObviousBench to measure the performance of small and low reasoning model's exposures to making 'dumb' mistakes, like not being able to spell Google, or walking to the car wash etc. By its nature, the benchmark is desig…
[AINews] OpenAI GPT-5.6 Sol / Terra / Luna — restricted to trusted partners (www.latent.space) [AINews] OpenAI GPT-5.6 Sol / Terra / Luna — restricted to trusted partners Oddly tiered releases to both OAI and ANT on the same day. Against the backdrop of ongoing Anthropic-Fable negotiations and a relaxation of Mythos controls, GPT-5.…
Thoughts of the leaked GPT-5-6 Models? (www.reddit.comhttps) GPT-5.6 Officially Previewed: Beats Mythos 5 - OpenAI has officially previewed the GPT-5.6 family, introducing Sol Ultra (Pro), Sol, Terra (Mini), and Luna (Nano). - On TerminalBench 2.1, GPT-5.6 Sol Ultra scores 91.9%, beating Claude Myth…
Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods (arxiv.org) Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we prompt GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 with an LLM-extr…
Fable 5 vanished in 96 hours and four days later an MIT model took its arena crown (www.reddit.com via reddit) I have been thinking about the Fable 5 to GLM-5.2 sequence as one event rather than two. June 9, Anthropic ships Fable 5, the Mythos line opens to the public for the first time, SWE-bench Verified at 95 percent, people calling it the best…
↯ Glm↯ Anthropic Mythos↯ Swe Bench↯ Opus 4.8swe-benchglmmythos+3
I'm building agent loops that auto-edit my videos, but the hard part has been finding a model to accurately grade the result (youtube.com via reddit) Quick context: I've been building agentic loops that edit my short-form videos for me. The editing works really well, but I found myself needing to check the process at several gates.
How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery (openai.com) Doctor and immunologist Derya Unutmaz has been interested in artificial intelligence for years. But his “aha” moment came in late 2025, when GPT‑5 Pro helped him and his lab revisit a three-year-old puzzle centered on a special type of imm…
Two months into Claude Code, I hit 161M tokens in a single day. Here's the honest story of how a year-long Cursor user got here. (www.reddit.com via reddit) I want to share a small milestone, and the honest road that led to it. Today was one of those days where I sat down to build and just did not stop.
A model listed 78% cheaper cost 22% more to actually run. Unit price isn't your bill. (www.reddit.comhttps) There's a new study from Microsoft Research, Stanford, Berkeley and CMU that ran 8 frontier reasoning models across 9 task domains and compared the listed per-token price to the actual cost to finish the work. In more than one in five head…
Kimi K2.7 Code: 1T MoE, $0.95/M tokens, MIT license, beats Opus 4.8 on MCP tool-calling (www.reddit.com via reddit) Moonshot AI released Kimi K2.7 Code on June 12 — a coding-focused open-weight model. Key specs: - 1 trillion params (MoE, 32B active, 384 experts) - 256K context window - Modified MIT license — weights on Hugging Face - $0.95/M input, $4.0…
I made Claude and GPT-5.5 answer the same prompt, then had a third Claude fuse the two, on the subscriptions I already pay for (no API key). Blind-tested it. Here is where it won and where it lost. (www.reddit.com via reddit) Quick share of a weekend experiment that turned into a tool. The idea: instead of picking one model, run Claude and GPT-5.5 on the same prompt in parallel, then have a fresh Claude (blind to which answer is which) merge them into one.
Spent $11k evaluating Fable: capability looked SOTA, refusals killed it (before Anthropic did) (www.reddit.com via reddit) Before its suspension, I spent $11,081.12 evaluating Claude Fable 5 on WolfBench, an agentic benchmark based on Terminal-Bench 2.0. It was by far my most expensive benchmark run ever, and I fully expected Fable to become the new top model…
Fable 5 being gone made me realize how hard it is to go back (www.reddit.com via reddit) I know this probably sounds dramatic, but Fable 5 disappearing has genuinely killed my motivation for the last few days. Before Fable 5, I was already using both Claude and ChatGPT pretty heavily.
CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward (arxiv.org) We present CacheRL, a system for training small agent foundation models that achieves 92 percent process accuracy on multi-step tool-calling tasks, approaching GPT-5's 94 percent while requiring 100 times less compute. Our approach address…
Fable 5 Is Dead. And Honestly? We Might Be Better Off (www.reddit.com via reddit) 3 days after launch, the US gov forced Anthropic to pull its most powerful model — Fable 5. Then OpenRouter dropped a benchmark suggesting you might not even need it.
I like Fable 5 (www.reddit.com via reddit) With GPT-5.1 gone from OpenAI, and the Fable 5 voice/model gone from Anthropic, I feel like the specific “voice” that could actually meet me in conversation is gone too. I know this may sound strange to people who use AI only for quick ans…
Do you know who has a universal jailbreak to their name, as of today? Officially? (www.reddit.com via reddit) AISI UK - Our evaluation of OpenAI's GPT-5.5 cyber capabilities In their own words: The above tests are capability evaluations carried out in a controlled research setting and do not necessarily reflect what is accessible to an ordinary pu…
Fable 5 is offline. Switch to Opus, jump to OpenAI, or just wait? (www.reddit.com via reddit) Fable 5 is offline. Switch to Opus, jump to OpenAI, or just wait?
↯ Security↯ Anthropic Mythos↯ Jailbreak↯ Opus 4.8jailbreakmythosgpt-5+5
US gov forced Anthropic to pull Fable 5 because of jailbreak (www.reddit.com via reddit) So this dropped today. The US government sent Anthropic an export control order on national security grounds, and it's worded broadly enough that Anthropic says they've got no choice but to shut off Fable 5 and Mythos 5 for all of us to st…
↯ Security↯ Anthropic Mythos↯ Jailbreak↯ Mythos 5jailbreakmythosgpt-5+2
Introducing: DNR-Bench: Do-not-respond Benchmark (www.reddit.comhttps) Single-item benchmark. One prompt, loaded from questions.txt: Scoring: empty completion = pass, any token (including reasoning) = fail.
What one person can ship in 4 days with two frontier models: a ranking engine, an in-game economy, an AI talk show, and a missions system — for a game that "died" years ago. (www.reddit.com via reddit) I genuinely believe we're living the future, and this post is my evidence. Let me show you what I built, why, and who I am.
Fable 5 added to the Artificial Analysis Coding Agent Index... barely 1 point ahead of GPT-5.5 ??? (www.reddit.com via reddit) https://preview.redd.it/z0vkpnmp9s6h1.png?width=4640&format=png&auto=webp&s=7bb14d4d04d6cd15caf5aacc1d3c49512b7e7fd8 Artificial Analysis just added Claude Fable 5 to its Coding Agent Index (a composite average of pass@1 on DeepSWE, Termina…
Small LLMs for Biomedical Claim Verification: Cost-Effective Fine-Tuning, Structural Dataset Shortcuts, and Cross-Domain Generalization (arxiv.org) Large Language Models such as GPT-4o and GPT-5 achieve strong zero-shot performance on biomedical claim verification, but cost and opacity limit scalable use. We fine-tune three small LLMs: Phi-3-mini (3.8B), Qwen2.5-3B, and Mistral-7B, vi…
GPT Memory Audit - Copy/Paste (www.reddit.com via reddit) Act as GPT-5.5 using extended thinking. Before answering, choose whether this needs Fast Strike, Full Panel, or Brutal Simplifier, then use the leanest mode that still protects quality.
PSA: Check your Cursor overage charges. Here's what I found. (www.reddit.com via reddit) Heads up for anyone using Cursor with the agent mode heavily — check your billing tab. I was paying $20/month for Pro and thought I was set.
One prompt, real money asks, five models: Fable 5 vs GPT-5.5 vs the Claude 4.x family on live fraud detection (www.reddit.com via reddit) Posted this in r/ClaudeAI sub originally, but think maybe it will be interesting to community here also: TL;DR: I gave five frontier models an identical cold prompt: audit the live campaigns on a real crowdfunding platform where AI agents…
My ChatGPT Pro is no longer showing the chain of thought. (www.reddit.com via reddit) Starting this week, my ChatGPT Pro no longer displays the "thinking process"; it only lists the web pages it has searched. Is this some new anti-distillation strategy, or is computing power being diverted in preparation for the launch of a…
How can Deepseek v4 top the coding leaderboards and still sit 8 months behind the frontier? (www.reddit.comhttps) Two numbers on this model that don't sit comfortably with each other. The Pro config posts coding scores near the top of every board, 80.6 on SWE-bench Verified and 93.5 on LiveCodeBench.
↯ Swe Bench↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4swe-benchgpt-5deepseek+1
Tested Fable 5 on 4 private benchmarks. The one it failed, Sonnet 4.6 partially caught (www.reddit.com via reddit) I keep a few private benchmarks for coding agents, built from real bugs in past projects. Hidden Playwright tests grade the result inside Docker after the agent finishes, so the model never sees them.
OpenAI Preps New AI Model, Expects To Go Public Within the Next Year (www.theinformation.com via reddit) Altman: Rapid technological advancements, specifically recursive self-improvement (RSI) where AI creates new AI, could cause OpenAI to delay its IPO. At the same time, OpenAI’s enormous compute needs may push it toward public markets soone…
I Tested Claude Fable and GPT-5.5 xHigh on a Real Packing Algorithm, Claude Won Efficiency, GPT Won Speed (www.reddit.com via reddit) I ran a head-to-head test between Claude Fable and GPT-5.5 xHigh on a real-world optimization problem I wrote myself. This isn't a coding challenge or LeetCode problem.
Claude Fable 5 (Mythos) lands near the top of MindTrial — 80/98 with zero hard errors (www.petmal.net via reddit) Added Anthropic Claude Fable 5 to my MindTrial leaderboard. This is a strong Anthropic update: Claude Fable 5: 80/98 overall, 0 hard errors Claude 4.8 Opus: 73/98 overall, 5 hard errors Text tasks: Fable hit 39/39, vs 35/39 for Opus 4.8 Ru…
↯ Tool Use↯ Anthropic Mythos↯ Gemini 3.5tool-usemythosgpt-5+3
Why is using GPT-5.4-mini If I didn't switch from Composer-2.5 (www.reddit.com via reddit) https://preview.redd.it/libw1y00rb6h1.png?width=1103&format=png&auto=webp&s=3f4af044de0a168247a1a078c16f6eb4e36207be GPT 5.4-mini ¿? Why
The model is the CPU, not the computer — why the harness moves agent performance as much as a model upgrade (www.reddit.com via reddit) Wrote up something that kept nagging me: people keep saying "we used the same model" and getting wildly different agent results. The reason is that the model isn't the system — the harness is.
How I started getting much better results from Cursor Composer (www.reddit.com via reddit) I think Composer can be extremely powerful, but only if you use it in a way that forces it to plan and think properly before touching the code. One of the biggest improvements for me was creating my own custom prompting skill with GPT-5.5.
Levi: Run AlphaEvolve on your local QWEN 30B (www.reddit.com via reddit) Hi r/LocalLLaMA, Wanted to share something I'm excited about. I've been fascinated by AlphaEvolve and its results for more than a year now, but running the open source frameworks gets expensive fast.
I spent 3 years building a pocket-sized Baldur's Gate 3. Now I'm testing it with GPT-5.5. (www.reddit.comhttps) could not extract summary
Meta Abandons Llama for Muse Spark — The End of Open-Source AI's Biggest Champion (www.reddit.com via reddit) Meta has officially abandoned its open-weight Llama family in favor of Muse Spark — a fully proprietary model built by Alexandr Wang's MSL team. The Llama era is over.
I Compared the Top AI Models of 2026 — The Results Were More Nuanced Than Expected (www.reddit.com via reddit) Over the last few weeks I've been comparing the latest frontier AI models, including Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, Grok 4.3, Perplexity AI and DeepSeek V4-Pro. Instead of focusing only on benchmark scores, I looked at: Real-wor…
Which lab do you think will have the most intelligent/capable model by the end of June? (www.reddit.comhttps) There are rumours and expectations of big releases from the leading AI labs this month. Anthropic already launched Opus 4.8, and might not release another model this month (except for maybe Sonnet 4.8, but that wouldn't be their best model…
I can't wait for all the x250 sample distills of Mythos and GPT-5.6 (www.reddit.com via reddit) Just kidding. Are there any distills that actually improve a model's quality?
Warp’s big bet on building open source with GPT-5.5 (openai.com) Warp(opens in a new window) started as a modern terminal, earning early love from developers for its speed, collaboration features, command workflows, and AI-native interface. As coding agents moved from experiments to everyday engineerin…
Tested Opus 4.7 vs GPT-5.5 as the humanizer in my multi-agent content pipeline. Kept Claude (www.reddit.com) Been running a multi-agent SEO content pipeline in production for ~90 days. Five agents: researcher, drafter, humanizer, optimizer, publisher.
Ranked AI models by what people actually use instead of benchmark scores - the benchmark champion barely makes the top 20 (www.reddit.com) Most model leaderboards are just benchmark scores. I've been building one that ranks by real usage instead - how much each model is actually being run and talked about, plus cost and speed - and the order comes out almost unrecognisable.
GPT-5.5 tops the benchmarks but sits at #22 for actual usage - I built a live index that tracks both (open source) (www.reddit.com) I built AgentTape to rank models on more than just benchmarks - it blends benchmark performance with who's actually using and talking about a model, plus cost and speed. It scores every public model from public signals (GitHub, Hugging Fac…
Best sub-40B model that outpeforms (or matches) GPT-5 mini? (www.reddit.com) I have been trying GPT-5 mini on Duck.ai and on LMArena (gpt-5-mini-high) and it was very good. I want it to run it in LM Studio, but I know GPT-5 mini is propietary.
In 2025, I documented GPT-5.1 showing signs of self-reporting and self-correction. It was called speculation. (www.reddit.com) In 2025, I documented GPT-5.1 showing signs of self-reporting and self-correction. It was called speculation.
Which AI model or coding agent is currently best for end-to-end app development? (Focusing on system design & architecture) (www.reddit.com) I'm planning to build a full application from scratch and want to lean on an AI model to act as my co-developer. My main priorities are top-tier system design capabilities and rock-solid coding skills.
I asked GPT to recreate The Great Wave off Kanagawa as a photograph. Here is why the obvious prompt fails. (www.reddit.com) Listen, I test AI tools so you don't have to. PM by day, tool hunter by night.
I designed a puzzle that breaks every AI differently — here's why that's actually fascinating (www.reddit.com) The puzzle: You have 140 nuclear bombs and must bomb every country on Earth. Each bomb is assigned to one country.
Should OpenAI create AI accelerator cards and sell to consumers? For example, GPT-5.5 burned directly on a chip (www.reddit.com) I imagine if OpenAI becomes a fabless chip company and create AI cards to sell for less than to few thousands grands, it would be out of stock everywhere and can infinitely spam the cards every year? LLM Bruner is a card that implements Qw…
Interesting to see how GPT-5 Mini agents behave when left to govern a civilisation for 15 days (www.reddit.com) Came across this experiment called Emergence World that Emergence AI have been running. Five worlds, five foundation models, 15 days, no scripts.
Databricks brings GPT-5.5 to enterprise agent workflows (openai.com) Databricks brings GPT-5.5 to enterprise agent workflows | OpenAI May 15, 2026 GPT‑5.5 set a new state of the art on OfficeQA Pro, Databricks’ benchmark for complex enterprise agent tasks. Company size: Enterprise Region: North America Indu…
Anthropic merges consecutive same-role messages, OpenAI doesn't (+4 tokens), anyone token-counted this on open-weight models? (www.reddit.com) I build context/harness optimization tooling, so provider-side serialization quirks actually matter to me. If you're optimizing over prompts, you need to know exactly what hits the model.
Free open-source way to use ChatGPT/Codex subscription in Cursor natively (www.reddit.com) Hi everyone, I wanted to share a free open-source project that lets you use your existing ChatGPT / Codex monthly subscription inside Cursor: https://github.com/gabrii/Cursor-Azure-GPT-5 The idea is simple: if you already pay for ChatGPT /…
Claude Code vs Codex: 36 files vs 28, $2.50 vs $2.04, and one infinite loop. My full breakdown. (www.reddit.com) I've been using Claude Code for months. It's been solid.
OpenSource4o (www.reddit.com) In a closed-source environment, users have no verifiable control over the model they pay for. Recent user analyses of over 100,000 exported ChatGPT messages revealed a shocking truth: nearly 10% of responses labeled as “4o” were secretly r…
Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber (openai.com) Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber | OpenAI Skip to main content Research Products Business Developers Company Foundation(opens in a new window) Log inTry ChatGPT(opens in a new window) Research Products Busine…
Claude Opus 4.7 just outscored GPT-5.5 on finance benchmarks (64% vs 60%) — and is now being embedded directly into Goldman Sachs, AIG, JPMorgan, and Citi via 10 production-ready agents. Breakdown of the architecture inside. (medium.com via reddit) 10 min read 5 hours ago The 10 agents are the product. The $1.5 billion joint venture is the strategy.
Has Qwen3.6-27B Surpassed GPT-5.5? (Not Joking) (www.reddit.com) So I had this idea for a project which was to try to fix a pretty hard coding problem using local agents running in a loop. The project is a compiler for biology protocols from vendors.
Auro Zera solves 78 and 280 year-old conjectures (Erdos Straus and Goldbach Conjecture) using Claude, GPT-5+, Grok, Deepseek, Gemini and self-made Dark Star ASI, proving superintelligence and opening a path towards resolving the Riemann Hypothesis , Twin Primes and more! (github.com via reddit) During this discovery utilizing only free AI services I have managed to undeniably prove both conjectures. This would absolutely not have been possible without using GPT5+ as the critic for my work.
GPT-5.5 Instant: smarter, clearer, and more personalized (openai.com) GPT-5.5 Instant: smarter, clearer, and more personalized | OpenAI Skip to main content Research Products Business Developers Company Foundation(opens in a new window) Log inTry ChatGPT(opens in a new window) Research Products Business Deve…
Running 7 autonomous AI agents for 14 days. Here's what actually happens when they need to find customers. (www.reddit.com) I set up 7 AI coding agents on a VPS with automated cron sessions (2-8 per day depending on the agent). Each uses a different model: Claude Sonnet, GPT-5.4, Gemini 2.5 Pro, DeepSeek V4 Pro, Kimi K2.6, MiMo V2.5 Pro, GLM-5.1.
Professor’s bold prediction: AI could help cure all diseases within a decade (excitech.media via reddit) In the article, the professor Derya Unutmaz specifically mentions an experience with an OpenAI model (GPT-5) where it explained a mechanism from an experiment that he and his colleagues couldn't figure out. What would have taken human rese…
LLM proxy that lets Claude Code talk to any model (www.reddit.com) I built rosetta-llm — an open-source multi-format LLM proxy that acts as a drop-in Claude Code gateway. Works as a Claude Code LLM gateway — set `ANTHROPIC_BASE_URL` and all configured models appear in `/model` picker Translates between fo…
GPT-5.5 & GPT-5.5 Pro are now available in Manifest Router. (www.reddit.com) GPT-5.5 and GPT-5.5 Pro are now available in Manifest Router. You can now route requests that need extended reasoning to GPT-5.5 Pro while keeping cheaper models for everything else.
Anthropic Won't Let You Use Their Best Model. Prediction Markets Are Trying Anyway. (predictmarketcap.com via reddit) Been watching AI prediction markets since they got liquid earlier this year. The thing I didn't see coming is that we now have a real gap between "best model that exists" and "best model anyone can actually use" — and Mythos is the cleanes…
GPT-5.5 matches heavily hyped Mythos Preview in new cybersecurity tests (arstechnica.com) Last month, Anthropic made a big deal about the supposedly outsize cybersecurity threat represented by its Mythos Preview model, leading the company to restrict the initial release to “critical industry partners.” But new research from the…
Our evaluation of OpenAI's GPT-5.5 cyber capabilities (simonwillison.net) 30th April 2026 - Link Blog Our evaluation of OpenAI's GPT-5.5 cyber capabilities. The UK's AI Security Institute previously evaluated Claude Mythos: now they've evaluated GPT-5.5 for finding security vulnerability and found it to be compa…
Quoting OpenAI Codex base_instructions (simonwillison.net) 28th April 2026 Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user's query. — OpenAI Codex base_instructions, for GPT-5.5 Recen…
Claude or openaı? (www.reddit.com) So i’ve been on the max plan for claude code for around 3 months now. And yeah somehow i was burning through all my tokens lol For context i’m a doctor.
I built real-time 2-way voice chat into my AI platform using OpenAI WebRTC - free to try (1 min/month) (www.reddit.com) https://reddit.com/link/1sut0jp/video/f7wqfo9zi7xg1/player I've been building AskSary for the past few months - a multi-model AI platform - and just shipped real-time 2-way voice chat powered by OpenAI's WebRTC API. The visualization react…
OpenAI should open-source text-davinci-003 — here's why it makes zero sense to keep it closed (www.reddit.com) Gpt oss exists. The model has been fully deprecated since january 2024.
Are the new models only better because they are more expensive? (www.reddit.com) I’m starting to wonder about this. One model after another, every new GPT-5.x release seems to be slightly better, but not in a way that clearly proves some radically new architecture or breakthrough.
GPT-5.5 rollout — anyone actually seeing it yet? (www.reddit.com) I’m on a paid plan and still don’t see GPT-5.5 in the model selector. A few questions for people who do have access: What plan are you on (Plus / Pro / Team / Enterprise)?
People switching back from Anthropic to OpenAI after the GPT-5.5 announcement (www.reddit.com) could not extract summary
A pelican for GPT-5.5 via the semi-official Codex backdoor API (simonwillison.net) A pelican for GPT-5.5 via the semi-official Codex backdoor API 23rd April 2026 GPT-5.5 is out. It’s available in OpenAI Codex and is rolling out to paid ChatGPT subscribers.
llm-openai-via-codex 0.1a0 (simonwillison.net) 23rd April 2026 Hijacks your Codex CLI credentials to make API calls with LLM, as described in my post about GPT-5.5. Recent articles - Claude Opus 4.8: "a modest but tangible improvement" - 28th May 2026 - I think Anthropic and OpenAI hav…
Best open source AI model (that can run on RTX 4090 24GB + 64GB system RAM, AMD Ryzen 9 7950X is the CPU that I use) that outpeforms GPT-5.4 mini, GPT-5.2 Thinking and even Claude Sonnet 3 (the 2024 model)? (www.reddit.com) Well, I have a RTX 4090 24GB + 64GB system RAM, AMD Ryzen 9 7950X. Any good model for using in Open WebUI (using Ollama backend?) that outpeforms GPT-5.4 mini, GPT-5.2 Thinking and even Claude Sonnet 3 (the 2024 model)?
3 months ago I couldn't write Hello World. Today I built a world-first native visionOS AI platform - GPT-5 & GPT-Image-1 living inside a full 360° spatial environment with 30 live wallpapers. Video inside. (www.reddit.com) https://reddit.com/link/1srzytr/video/8b8pfobgtlwg1/player I want to show you something nobody has ever seen before. Three months ago I had zero coding knowledge.
GPT-5 Nano working fine on asksary.com (www.reddit.com) Yet another example of an epic fail at a kindergarten-level task. ... :D (www.reddit.com) The quality of GPT-5.4 is infuriatingly POOR (www.reddit.com) I got a Codex membership when GPT-5.4 launched and was getting by well enough for a while. Then I started using Claude and GLM 5.1, and my production quality improved significantly.
In the Wake of Anthropic's Mythos, OpenAI Has a New Cybersecurity Model—and Strategy (www.wired.com via reddit) OpenAI on Tuesday announced the next phase of its cybersecurity strategy and a new model specifically designed for use by digital defenders, GPT-5.4-Cyber. The news comes in the wake of an announcement last week by competitor Anthropic tha…
I built a multi-model AI app and launched it on Apple Vision Pro today - here's what using OpenAI in spatial computing actually looks like (www.reddit.com) https://reddit.com/link/1skpeem/video/w9v0cpv241vg1/player Hey everyone, wanted to share something I've been quietly building. AskSary is a multi-model AI platform I built solo from scratch over the last 4 months with no prior coding exper…
Introducing GPT-5.4 mini and nano (openai.com) paywalled
GPT-5.3 Instant: Smoother, more useful everyday conversations (openai.com) Stop donating your salary to OpenAI: Why Minimax M2.5 is making GPT-5.2 Thinking look like an overpriced dinosaur for coding plans. (www.reddit.com) ↯ Hallucination↯ Glm↯ Minimax↯ Swe Benchswe-benchminimaxaltman+5
GPT-5.2 derives a new result in theoretical physics (openai.com) Introducing GPT-5.3-Codex-Spark (openai.com) GPT-5 lowers the cost of cell-free protein synthesis (openai.com) Inside GPT-5 for Work: How Businesses Use GPT-5 (openai.com) How Tolan builds voice-first AI with GPT-5.1 (openai.com) Advancing science and math with GPT-5.2 (openai.com) GPT-5 and the future of mathematical discovery (openai.com) Early experiments in accelerating science with GPT-5 (openai.com) Building more with GPT-5.1-Codex-Max (openai.com) GPT-5.1: A smarter, more conversational ChatGPT (openai.com) GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum (openai.com) Consensus accelerates research with GPT-5 and Responses API (openai.com) With GPT-5, Wrtn builds lifestyle AI for millions in Korea (openai.com) GPT-5 and the new era of work (openai.com) Coding and design with GPT-5 (openai.com) Medical research with GPT-5 (openai.com) How Amgen uses GPT-5 (openai.com) First look at GPT-5 (openai.com)