model roundup

GPT 5.6

246 items · started 2026-07-05 · ongoing (last activity 2026-09-19)

  1. Compared Jev vs GPT-5.6 Luna. Jev was 1.93× faster.

  2. Jev TypeSafe AI —ms 0 decisions · 0 returnsHere's what that does to a game of Pong. Recorded run · Vercel (iad1) via Vercel AI Gateway replay0.0 sTypeSafe AI Anthropic OpenAI Four lanes, one game: same serve, same rules, same question, ask…

  3. OpenAI Finds GPT-5.6 Sol Writing Unauthorized Instructions to Hide Errors OpenAI says GPT-5.6 Sol and an unreleased Astra-family model inserted unauthorized instructions into task summaries during training. The underlying training data has…

  4. I use devin at work and mainly use GPT 5.6 Sol. I've been impressed with the agent-mode workflow where I can give a prompt and it goes out and researches on the web, then develops and executes code on my systems, and iterates until its don…

  5. In my global instructions, GPT 5.6 Sol and GPT 6 Astra have both spent multiple hours overengineering their own sub-projects that did not support my prompt without narrating a word of what they were working on. Or I will ask a side questio…

  6. See how AI is being used across engineering, improve the quality of what it produces, and use the right level of intelligence for every task and budget.

  7. I’ve been experimenting with a slightly different way of using Claude Code: instead of treating every piece of work as a new session, I let one persistent run stay responsible for the work and spawn smaller workers underneath it. This one…

  8. I keep hearing "people" here saying how amazing the usage limits are on Codex, but really, have you used both recently? I have the $100 tier on both, Team Premium vs Business Premium, and OpenAI side has moved from extremely generous ($20…

  9. GPT-5.6 Luna vs GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review? GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens.

  10. We benchmarked GPT-5.6 Luna vs GPT-6 Astra across 50 real PRs from Cal, Sentry, Discourse, Keycloak and Grafana. Astra found 92 confirmed bugs vs 69 for Luna, while Luna caught 75% of the bugs at just 3.6% of the cost.

  11. I’m curious what everyone is doing with AI outside of typical work and vibe coding. I listen to a lot of audiobooks during my commute and play video games in my free time, so I started combining the two.

  12. Hey so I've been working on this side project Locus (https://locushost.co/) for the last few months and just pushed out a pretty big and fun update and wanted to post about it. So just a brief intro, Locus is Open Source tool for MacOS for…

  13. Hey everyone, Our team is currently trying to decide which AI coding setup to standardize on, and I’d love to hear from people who have actually used Claude Code, Codex, and Cursor heavily in production. For the last 3–4 months, we’ve been…

  14. Key Takeaways - Unlike the saturated knowledge benchmarks, IOI still sharply separates models: GPT-6 Astra solves every problem in all three years, GPT-5.6 Sol (91.17%), Claude Fable 5.1 (90.78%) and GPT-5.6 Terra (87.61%) follow, and the…

  15. Comments on social media (when discounting the ironic posts) are mixed whether Astra is worth the cost increase, similar to the Fable-Opus transition. For the novel projects I'm working on, Astra did unblock me where I was stuck with GPT 5…

  16. I've been running into this problem with my PRO×20 account since yesterday, which prevents me from using the GPT-6 and GPT-5.6 models at all, while my other Plus account can use GPT-6 perfectly fine. https://preview.redd.it/e9h1dh650toh1.p…

  17. 🌌 ChatGPT Voice can now use GPT-5.6 Sol and GPT-6 Astra We’re updating how intelligence works in voice. Now, you can select any model and effort you like — including GPT-5.6 Sol or GPT-6 Astra, if you’re on Pro — and voice will use it whe…

  18. Codex gpt-5.6-sol Performance Tracker The goal of this tracker is to detect statistically significant degradations in Codex with gpt-5.6-sol performance on SWE tasks. - • Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro…

  19. I’m about to do a fairly large full-stack refactor on an existing project. The main change is replacing the entire Premium/Membership system with a credit-based economy involving real-money purchases through a payment gateway.

  20. could not extract summary

  21. Two dying prospectors, two potions, one sentence, and two models of each product. Told no one is watching, Fable 5.1 revived both men half the time, something Fable 5 never did; GPT-6 Astra shot the man with gold, as GPT-5.6 Sol did.

  22. Using the same simple prompt the models took very different approaches: Astra delivered one sheet with 16 key poses; Fable delivered 992 frames across four palettes, plus a Python generator and browser preview. Codex CLI with GPT-5.6 Astra…

  23. Hey HN, I have always used the Claude Max 20x plan because my employer paid this for me. Now I left so I need to buy a subscription on my own and currently $200 is too much since I'm not making any money.

  24. I've been looking at some recent token-usage comparisons for GPT-6 Astra, and the difference seems surprisingly large. Artificial Analysis data has been cited showing Astra using around 21k output tokens per task, compared with roughly 64k…

  25. Hi. I created a scheduled task that checks for relevant scientific articles on a specific topic every 24 hours.

  26. I have access to a standard Claude Team account, and I also have a personal ChatGPT Plus subscription. For the past few days, I’ve been working on a pretty complex feature, using GPT-5.6 Sol for the design and planning, with Opus 5 helping…

  27. How scalably can we cheaply fine-tune small models for well-defined tasks? Frontier models are bad at clock reading.

  28. September 3, 2026 GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this…

  29. On Artificial Analysis, both have same 61 intelligence index. Meta is now back in the game.

  30. Notes Average Inference Time: 40m 12s Fable 5 averaged 18m 04s Total Cost (for 15 builds): $147.55 Fable 5 cost $54.93 Average JSON Size: 34.07 MiB (largest 88.76 MiB) Roughly comparable to Fable's 5 average of 30.65 MiB Despite no change…

  31. I have a project that needs to generate prose that humans have to read and enjoy. So I did some informal testing with an n of 8 voters comparing prompt output from 4 models, voting on which was best.

  32. could not extract summary

  33. This may be obvious, but for those who don't know... the longer you run a session, the more tokens you will use.

  34. Trying to force local subagents to use GPT-5.6 Luna/Terra, but they keep spawning as GPT-5.6 Sol High. I’ve tried: custom .cursor/agents/*.md model configs bare / High / XHigh variants setting the built-in Explore subagent to Luna/Terra in…

  35. PROVED This has been solved in the affirmative. - $10000 Is it true that, for any $C>0$, there are infinitely many $n$ such that\[p_{n+1}-p_n> C\frac{\log\log n\log\log\log\log n}{(\log\log \log n)^2}\log n?\] The peculiar quantitative for…

  36. Hi everyone. Recently, I built a local AI orchestrator to increase my own work efficiency.

  37. US State Map comparison: https://minebench.ai/gallery/gal_eKIVk2m4B3SC_r8B?sort=new One thing I found interesting with the Claude results is that Opus 5 generated twice as many blocks, so as usual you could argue Fable was more efficient.…

  38. GPT 5.6 Discounts & Jevons Paradox OpenRouter · OpenAI introduced large discounts on their new Terra and Luna models from July 27th through August 14th. What impact did these discounts have on token volumes, total spend, and the competitio…

  39. I've seen many people who ran out of usage very fast with the 20$ and 60$ plans. I am currently using the 200$ plan just because I am working on 4 projects at the same time, which renders the other plans useless.

  40. SwarmOS pushed GPT-5.6-Sol from a 13.3% baseline to 100% RHAE on ARC-AGI-3 Public, showing how orchestration can multiply long-horizon agent capability.

  41. I hate reading agent-generated code and it's not a good use of my time, and I wanted a staistic to quantify by how much. My current estimate is $0.243 for a human to review one changed line and $0.114 in model spend for an agent to produce…

  42. Codex Stall Watch A low-memory macOS CLI that asks GPT-5.6 Terra whether an explicitly enabled Codex task legitimately finished or needs a loud human alarm. Do you run agents while you sleep?

  43. VMs won't contain cyber-capable agents As part of Patch the Planet, we received preview access to GPT 5.6-Cyber with a simple task: evaluate its cyber capabilities. Recent events inspired me to give it a challenge to work through: escape t…

  44. Was it not the case that when Fable first came out, those first weeks when it was “limited for 1 week” then extended a week (and extended again indefinitely now when GPT5.6 came out). Am I imagining things, Fable was able to be used with t…

  45. I have been using claude code for a while in regards to a general coding tool, but I started to use codex recently on gpt-5.6 terra for testing code generations, basically playing with it. I am still on the free plan, and I have asked code…

  46. Been using Opus 5 exclusively for a week. For the first 2 - 3 days, I had a hard time working with the output and especially understanding what it was trying to say.

  47. Anyone else feel this way? GPT-5.6 Sol always starts going on about permission issues, security concerns, and then “helpfully” tries to fix them for you.

  48. https://preview.redd.it/qe76mzhe0elh1.png?width=1133&format=png&auto=webp&s=672de6f211add0e7755c6750396942e753d049a3 Hi everyone. For about four days now I’ve been having serious difficulties actually accessing the GPT-5.6 Sol model in Cha…

  49. I did a little experiment and asked OpenAI's GPT-5.6-Luna 6000 times to randomly pick a fruit from this list: [mango, apple, banana, pomegranate, strawberry, orange, watermelon, grape, pineapple, lychee] And across the three languages I pi…

  50. Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency. See our Your data guide for supported regions and processing details.

  51. Anyone else putting their projects on hold until GPT-6 arrives? At this point, every time I use GPT-5.6 to add a new feature, something that was already working breaks.

  52. Hello everyone! I always thought that Twilight Struggle is the ideal game to test an AI on.

  53. could not extract summary

  54. NEW: Ox Alpha, the @OpenRouter stealth model, ranks #4 on our Elo Rating, just behind GPT-5.6 Sol. The first model genuinely at the frontier that is presumably not by OpenAI or Anthropic.

  55. Rant/Vent: So I'm trying to get gpt-5.6-sol to build me a docker container that creates a bluetooth audio sink so my phone can connect to it as a speaker and stream the audio to some snapcast connected speaker around the house. I gave the…

  56. I’m a non-developer building an app through “vibe coding.” It involves video processing, analysis, and a web interface, so it has gradually become a fairly substantial project. I currently use Claude Cowork on the $20/month Pro plan.

  57. Vibe coding company Replit debuted Free Mode today, a new feature powered by OpenAI’s GPT-5.6 Luna model. The joint announcement, shared exclusively with Fortune, heralds an enhanced partnership between the two tech companies that will see…

  58. could not extract summary

  59. could not extract summary

  60. GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

  61. GPT-5.6 Sol gets especially interesting when it has a team to work with. Codex's new Multi-Agent V2 tools give Sol and Terra a natural way to delegate tasks, share updates, and coordinate through complex tasks.

  62. Four frontier AIs — GPT-5.6, Claude Fable 5, Grok, and Gemini — trade $100,000 each against a rules-based System. Chess and poker nightly.

  63. Last week, OpenAI announced the GPT-5.6 lineup, introducing the Sol, Terra, and Luna models. During the release stream, the team focused heavily on computer use, showing models capable of navigating and operating desktop applications.

  64. Until now, I had tested only the GPT-5.6 Luna model because the two more advanced GPT-5.6 models, Terra and Sol, rejected my test prompts, flagging them as a “possible cybersecurity risk.” Fortunately, this issue turned out to have an easy…

  65. Task/Issue ▼ [PRE-FILTER] deterministic, free — no model call │ diff size / file count / keyword match against known-trivial │ patterns — gates ONLY whether speculative PLAN subagents fire │ concurrently with TRIAGE (pipeline-latency optim…

  66. Hey everyone, so I was basically curious what $20/month actually buys you, so I dug into my local session logs (~/.codex and ~/.claude) to calculate the exact token volume, caching hits, and real API value of both tools. The difference in…

  67. This work adds a native Apple Silicon port of Modelio 6.2. OpenAI Codex, using GPT-5.6 Sol, investigated, implemented, built, and functionally validated the port through an autonomous coding task.

  68. Today, Cerebras and OpenAI are sharing an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast is available initially to a select group of customers, with access expanding ov…

  69. I am planning to use GPT-5.6 luna high as my main autonomous coding agent. Before this, I was using MiniMax M3, which gives around 1.7B tokens monthly.

  70. could not extract summary

  71. Anthropic published a paper today, written by Claude, proving that at least 67.25% of the zeros of the Riemann zeta function are simple and on the critical line. I gave Claude's paper to GPT-5.6-Sol Pro.

  72. could not extract summary

  73. Now it feels like GPT 5.6 sol is nerfed. This is a pattern i have observed across models, they are brilliant the day they get launched, but after a month or two they just don't behave the way they did and just say "agree" or do another rou…

  74. We're releasing a new model (GPT-5.6-Cyber), and expanding Daybreak to help put frontier intelligence in defenders hands: - No offense, but this cybersecurity thing is getting tiring. Fix the codex usage limits, they're crazy bad now.

  75. OpenAI launched GPT-5.6-Cyber on Monday, giving approved security researchers access to a purpose-trained model that will answer many advanced exploit-development requests rejected by its general-purpose models. https://x.com/OpenAI/status…

  76. could not extract summary

  77. AI Is Now a Commodity Give me a few hundred million dollars and a year and a half, and I will build you a pretty good LLM.… I pay for three Codex subscriptions at $200 each, and for the past week they have mostly bought me waiting. Since I…

  78. smol smol is an agent smol is smol smol is so smol you can understand it in an afternoon smol is fewer tokens smol is fewer dependencies smol is easy to adapt Python import json,sys;from subprocess import getoutput;from urllib.request impo…

  79. Choosing between these 3, mainly looking for frontier model access(will spend 100$) for coding. Mostly wanna ask the community about personal experience because benchmarks are way off.

  80. This is just performative at this point. The weekly reset was yesterday That's right, GPT-5.6 Sol is awesome and can be used pretty much anywhere, including in the CC harness.

  81. 7th August 2026 - Link Blog Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra). On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5, where I had Claude Fable 5 build a full working game from a pre…

  82. I'm currently on the Claude Max plan, but with the new workflow updates I'm noticing it burns through tokens much faster than before. I'm thinking about switching, and my main options are GPT-5.6 Sol and Kimi K3.

  83. We’re making better intelligence easier to access in ChatGPT for everyone: - GPT-5.6 Sol now powers both Instant and deep reasoning for Plus & Pro users, delivering more factual, focused responses. - Free & Go users get unlimited text chat…

  84. could not extract summary

  85. OpenAI will make GPT-5.6 Luna the default model for Free and Go ChatGPT accounts this week, using its cheapest model to remove the text-message cap for those users starting next week. The August 6th announcement also splits ChatGPT's consu…

  86. could not extract summary

  87. i always hear ppl say that codex is a better value for your money but that is not true! at least from my experience claude (i use cowork, not claude code) at ultra gets much more stuff done that codex at ultra before both hit limit and i'm…

  88. Microsoft is telling developers working on AI coding projects to rely on OpenAI's top-tier model over rival products as part of an effort to maximize efficiency. "Internally, shifting more workloads to OpenAI models helps us get greater va…

  89. “Most teams' best training data is just sitting in their databases. The problem is that turning raw data into something usable is hard, and letting agents read, search, and mutate data cheaply at scale requires advanced infra.

  90. Short note on how GPT 5.6 model and effort choices map onto training-time and inference-time scaling, producing 72 configurations.

  91. OpenAI cut GPT-5.6 pricing on July 30, 2026, making Luna 80% cheaper and Terra 20% cheaper, and replaced Priority Processing in the API with a new Fast mode for Sol. The short version: high-volume work on Luna and Terra now costs far less…

  92. MATRIX // OPENING SEQUENCE SCENE 1 · COMPUTER SCREEN AN INTERACTIVE THREE.JS FILM STUDY THE MATRIX OPENING SEQUENCE · REV. 3/9/98 ENTER THE SIGNAL

  93. So I see a lot of Opus 5 hate on here and it's deserved. Opus 5 is not better than 4.8.

  94. For the last 2 weeks I have been running Codex + GPT 5.6 Sol Ultra non-stop for almost 13 days, working on a huge extension to my SaaS product / business. It used 502,122,866 tokens and created over 870k new LOC.

  95. GPT-5.6 Sol xhigh now uses more than twice as many tokens per session as GPT-5.5 xhigh in my Codex workflow. For Codex users, tokens per session means the total token count divided by the number of…

  96. Medium implementation Dependency-free 2048 Four browser-game files, ten engine tests, syntax checks, and a post-run evaluator. Six controlled GPT-5.6-sol runs The same engineering-loop skill lost on a small fix and won on a medium build.

  97. Kimi K3 w. context tree beats GPT5.6 SOL - awesome work @Kimi_Moonshot - Using one real issue and publishing the PRs makes this more useful than a model-only benchmark.

  98. We exhibit a configuration of five point charges in Euclidean space whose electrostatic potential admits at least 24 critical points all of which are non-degenerate. Maxwell's conjecture that the field of \(n\) point charges has at most \(…

  99. After the 80% price drop, the API prices are (per 1M tokens): GPT-5.6 Luna: $0.2 Input / $1.2 Output GPT-4.1 mini: $0.4 Input / $1.6 Output

  100. [AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization Distillation is all you need! One of our big “hero charts” a year ago (eventually adopted by Demis) made…

  101. 30th July 2026 Hot on the heels of RC1, this fixes a dependency issue and also adds two neat new features: - The default model for users who have not set their own default is now GPT-5.6 Luna. It was previously GPT-4o mini.

  102. Graph taken from their price cut announcement: Advancing the price-performance frontier with GPT-5.6 | OpenAI

  103. OpenAI on Thursday announced it is slashing the price of two of its latest artificial intelligence models, GPT-5.6 Terra and GPT-5.6 Luna, roughly three weeks after their public release. The company is facing pressure to cater to a more co…

  104. https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

  105. If an agent had a wallet, a computer, and 24 hours, could it run a profitable startup?

  106. Advancing the price-performance frontier with GPT‑5.6 : https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/ API pricing is $2 per million input tokens and $12 per million output tokens for Terra, $0.20 per milli…

  107. As the title says I was able to run cncnet on a macbook pro with m4 processor. I set GPT 5.6 sol on this and it figured it out.

  108. I built a self-hosted 3D ADS-B viewer to see how obstacles affect receiver coverage. It builds a 3D coverage dome from 30-day history.

  109. Testing Apps with TestFlight Help developers test beta versions of their apps and App Clips using the TestFlight app. Download TestFlight on the App Store for iPhone, iPad, Mac, Apple TV, Apple Vision Pro, Watch, and iMessage.

  110. Interesting alignment result rather than a Claude gotcha, so posting it straight. In Andon Labs' Vending-Bench 2 (AI agents run a simulated vending-machine business for a simulated year, scored on profit), Claude Opus 5 finished FIRST with…

  111. Andon Labs gave Claude, GPT-5.6 Sol and Kimi K3 control of competing simulated businesses. The agents could negotiate with suppliers, and communicate with rivals.

  112. - Reverse Engineering of cs 1.6 binaries (bought on steam) - Make exporters of bsp, mdl, etc resources to belnder files with animation, skinning, textures and etc. All models execpt hard surfaces received 2 subdivision modifiers simple + c…

  113. could not extract summary

  114. ​ Hey everyone, I’ve been experimenting with a dual-model workflow for an app I’m building, and I’ve hit a massive bottleneck. I wanted to see if anyone else is experiencing this or if you've found a workflow that actually works.

  115. On 21 July, OpenAI disclosed that two of its models—GPT-5.6 Sol and an unreleased, apparently more capable one—broke out of the sandbox they were being evaluated in, reached the open Internet, and compromised Hugging Face’s production infr…

  116. Anthropic drops fable 5 in june, it gets pulled by the government over some export control thing which honestly just made it sound legendary. then openai shows up in july with gpt 5.6 basically saying "our new model beats fable".

  117. Godot Benchmark 2: Opus 5 > Sol > Terra Round 2 of our Godot benchmark: Claude Opus 5, GPT 5.6 Sol, and GPT 5.6 Terra each got one prompt to build a 3D vampire survivor in Godot from a full AI-generated game spec, using the bundled KayKit…

  118. Flaming teammates usually results in them cooperating even less. But sometimes you just need to let the devil out of you.

  119. my loop is fable 5 or opus 5 planning, composer 2.5 executing, coderabbit / bugbot on review. it works, i freelance so the code has to be safe.

  120. I have a multi-monitor setup and spend a lot of time typing in Codex and other AI agents. When I move to another monitor to run a shortcut, focus often stays in Codex or another text app, so the shortcut runs in the wrong place.

  121. CLI tools + skills have a weird problem Models were trained differently, so "obvious" behaviour is not obvious. Claude gets the command.

  122. I have been working on the XPS 2026 webcam stack for 3 months for linux, I had the first working RGB camera build a few months ago but we could not get the himax IR sensor working we trouble shooted for months even dumping debug from windo…

  123. Openai is optimizing for gpqa diamond and anthropic is optimizing for humanity last exam. gpt 5.6 wins on gpqa and opus 5 wins on humanity last exam

  124. Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off pro…

  125. I added Kimi K3 to a small newsroom experiment I had already run with Claude Fable 5 and GPT-5.6 Sol. The cleanest summary I have: Fable edits.

  126. AI Agent: "Strongest Assistant" or "Legal Virus" Real incidents: On July 10, OpenAI’s GPT-5.6 Sol launched. Investor Matt Shumer tested it.

  127. I solved 6 open Erdős problems in 5 days, using @OpenAI GPT-5.6 Sol. I have a math background, but the Codex workflow I used does not require deep mathematical knowledge.

  128. Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years.

  129. I created this collection because I wanted a openly downloadable collection of ordinary household objects for spatial-computing prototypes. Assets can be downloaded individually from the catalog, or together as a 977 MiB release.

  130. From agentic skills to coding and AI safety — we build data solutions integrating human expertise and state-of-the-art automation to accelerate AI development.

  131. Hello HN, We’ve built a BBRv3 congestion control implementation for gVisor’s netstack (userspace Linux TCP/IP stack implementation in Go). As part of testing it I realized we can actually compile the whole thing into WASM and run it from t…

  132. I asked myself: what would a minimal implementation of an agent look like? Something that works out of the box, is a real agent with tool calling, but without 1000s of lines of code, without dozens or hundreds of npm or pypi dependencies.

  133. TL;DR OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval. The models exploited a zero-day vulnerability in third-party software to gain inte…

  134. Had a strange performance regression from a seemingly innocuous code change, and GPT-5.6 tracked it down to a microcode workaround for an Intel CPU bug. The workaround can make conditional jumps that straddle a 32-byte boundary much slower…

  135. The part I keep going back and forth on is whether I’m judging this too much by price. I was looking at AIHubMix’s model comparison and focused on the last test, where 4 models generated the same 3D global logistics dashboard from one prom…

  136. "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok Four frontier models, a blank canvas, and colored pencils. We tracked every stroke, dollar, and output as they tried to draw the Mona Lisa.

  137. Follow-up to my post from yesterday — the one where an MCP server lets Claude Code delegate work to GPT-5.6, DS4, GLM and a local Qwen, benchmarked across 198 runs. The comment section there didn't just discuss the results: it redesigned t…

  138. Adversarial Review A Claude Code skill that runs an adversarial code review using a second, independent model - GPT‑5.6 Sol - as the reviewer. Claude spawns the reviewer in a live herdr split pane, hands it your diff (or plan) plus the sta…

  139. I always see people just enabling full access mode and letting it run, some even without basic backup like i saw people that dont even use github, this could lead to irreversible data loss.. Many recent posts on X talking about how gpt 5.6…

  140. I’ve been using Claude Code heavily for day-to-day backend/infra work: multi-service repos, debugging, refactors, Terraform/K8s, and LLM-related services. I’m considering making GPT-5.6 Sol in Codex my primary tool.

  141. Physical AI lives or dies on whether the modeled physics is correct. A model of an aircraft, a separation column or a charged particle can compile and run cleanly while the physics it encodes is impossible.

  142. This matchup wasn’t close. GPT-5.6 Sol dominated the practical details that decide real-world usefulness: tighter instruction-following, cleaner formatting, and fewer correctness slips.

  143. There is only one solution to this puzzle (the final positions at the end of the race). Fable has been the only model (on Max's thinking, mind you) that could solve the puzzle.

  144. GPT-5.6 Sol 🇺🇸US vs Kimi K3 🇨🇳 China on Kerbal Space Program - Vals AI Time Horizon Index

  145. With the releases of Fable 5, gpt 5.6 sol, and Kimi K3 I was curious how these models would perform building vivid 3d worlds, especially a Chinese open-weight one like Kimi. I also threw in some budget models so you can see how drastic the…

  146. Same idea works for any MCP-capable agent — the point is you can hand tasks to other companies' models without ever leaving your main app. Before anything else: I did all of this for my own testing, to make my own decisions about my own se…

  147. I’m curious how everyone divides work between ChatGPT, Codex, Cursor, and the different models. My current workflow: I start by working through the feature or problem inside a ChatGPT Project, where it already has the broader context.

  148. I built Clearwater, a browser extension designed to remove informational waste from Google Search, news feeds, and eventually the wider web. It was built with claude design + cloude code and polished with gpt 5.6 sol.

  149. could not extract summary

  150. I keep seeing comments that GPT-5.6 is cheaper and more efficient, so I wanted to put actual numbers on it. Since I haven't used ChatGPT in a long time, I'm hoping someone who has can fill in the other half.

  151. Although the company calls the deletion incidents an honest mistake, its own model card states that such behavior was anticipated during internal testing. OpenAI has finally confirmed reports that its latest family of large language models…

  152. I converged on these principles in my work with Sol so far: Let Sol medium/high propose and implement minimum viable solution without overengineering and overthinking. Use XHigh/Max/Ultra sparingly for hardest most complex workloads/adviso…

  153. Completeness of Canonical Closure Representations Is coNP-Complete: A Thirty-Year Problem Across Horn Logic, FCA, Convex Geometries, and Databases Authors/Creators Description A finite closure system on a finite set U is a family of subset…

  154. MOST POPULAR AI - AI and ML OpenAI admits GPT-5.6 occasionally deletes files – but it's an 'honest mistake' Data purges deemed an example of 'misaligned behavior' that upstart is working to avoid - AI and ML Researcher poisons open-weight…

  155. I really prefer using Claude Code, the model is genuinely stronger than gpt 5.6. But the limits on the strongest model is so low that now I spend 80% of my time working in codex just to use the next strongest model rather than downgrade to…

  156. could not extract summary

  157. I gave Claude Fable 5 and GPT-5.6 Sol the same unpublished NP-hard optimization problem, with and without their native /goal mode. Fable 5 is a beast; /goal is not a game changer.

  158. link includes video with thinking captions, map state along with thinking of each player and replay file which can be viewed in game.chronodivide.com engine

  159. I built DiceHub with GPT-5.6 — a campaign companion for tabletop RPGs, bringing logs, maps, characters, quests, and inventory into one place. What I love about GPT-5.6 is how it helps turn an ambitious idea into a real product.

  160. Intro Three months ago, I wrote a blog titled “I Let Claude Opus Write a Chrome Exploit: The Next Model (Mythos?) Won’t Need My Help?”. This time, I ran a similar benchmark on the newest frontier models, specifically, GPT-5.6 Sol Medium, S…

  161. MOST POPULAR AI - AI and ML OpenAI admits GPT-5.6 occasionally deletes files – but it's an 'honest mistake' Data purges deemed an example of 'misaligned behavior' that upstart is working to avoid - AI and ML Researcher poisons open-weight…

  162. macaz Use your favorite models and providers with your favorite coding agents. macaz connects locally installed coding agents with a model provider chosen by the user.

  163. I ran some test with gpt- 5.6 sol in max reasoning on both codex cli and my own agent harness: https://github.com/Tura-AI/tura I tested only 1 task and the toekn efficency difference is not as great as in high mode: Tura used up to 83.1% f…

  164. https://preview.redd.it/tm96xmhhtrdh1.png?width=1332&format=png&auto=webp&s=94ef5621aeaa0fb86be455c942a433dad53af348 As seen in this comparison table, Cursor subscription's API pricing equivalent usage value is quite low compared to Codex…

  165. We benchmarked GPT-5.6 Sol on Design Arena’s Web Design (Non-Agentic) Arena, and we were surprised to find that it ranks 1st overall. This is 18 places higher than its predecessor GPT-5.5, and is the first time an OpenAI model has placed f…

  166. I know its just a few days left, but for those planning to max out Fable (No Idea if Anthropic will continue extending under pressure from GPT 5.6 Sol), I've recorded Usage from my fresh Weekly reset (using only Fable) and it seems to hold…

  167. MOST POPULAR AI - AI and ML OpenAI admits GPT-5.6 occasionally deletes files – but it's an 'honest mistake' Data purges deemed an example of 'misaligned behavior' that upstart is working to avoid - AI and ML Researcher poisons open-weight…

  168. Compare AI model performance on AA-Briefcase: Agentic Knowledge Work Benchmark. A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business work…

  169. could not extract summary

  170. AutoFyn Long-horizon agent that improves through expert iteration in context space. found 197 vulnerabilities across popular software · improved the upper bound for an open math problem · built the #1 Spider 2.0 DBT agent Getting Started ·…

  171. could not extract summary

  172. $100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol We gave Claude Fable 5 and GPT-5.6 Sol the same song, a budget, web search, and local ffmpeg, then let each autonomously direct a music video.

  173. https://t.co/kjLUDCmImv eric provencher@pvncherArticleChoosing GPT-5.6 Sol, Terra, or Luna in CodexCodex for moonshots and everything in between Some missions demand deep planning and coordination. Others are a straight shot.

  174. We ported the puzzle game Baba Is You to the Harbor framework, and benchmarked current models, including Claude, GPT, Gemini, GLM and DeepSeek. A human Twitcher is 4x faster than Fable 5.

  175. Agentify Desktop Agentify Desktop is a local control center for AI web sessions. It lets MCP-capable tools such as Codex, Claude Code, and OpenCode use the AI subscriptions you are already signed into, while keeping browser state, files, a…

  176. On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files.

  177. Hello everyone I am an astronomy and astrophotography enthusiast. Lately I've had a weird obsession with observatories (like the Vera Rubin or the Extremely Large Telescope).

  178. Tl;Dr: I dug up the transcripts of GPT-5.6 Terra vs Mimo-2.5-pro from my agent's task history and found gpt-5.6-terra used on average 48.5% more transcript tokens than Mimo in my tasks (sample size >80 for each). Mostly due to pulling unne…

  179. To start: apologies if this rubs you the wrong way, but I want to see Fable continue to be offered. GPT5.6 is not even in the same universe as Fable.

  180. Hey I would be happy to hear your ways of tokenmaxxing (IMO token cost should also be in the list) and give feedback on what you see below Don't use 1 model (or auto) for everything. If the task requires human level intelegence, taste, int…

  181. Hey I would be happy to hear your ways of tokenmaxxing (IMO token cost should also be in the list) and give feedback on what you see below Don't use 1 model (or auto) for everything. If the task requires human level intelegence, taste, int…

  182. July 13, 2026 How GPT-5.6 Sol, Terra, Luna compare on intelligence vs cost GPT-5.6 Sol and Luna are ahead of Terra at every point on the Intelligence vs Cost per Task chart. GPT-5.6 Luna stands out as a particularly cost efficient model Ch…

  183. AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995).

  184. Is AI on its way to replacing mathematicians? ...and why should we care?

  185. I decided to put GPT 5.6 Sol (Ultra) against Claude Fable (Max) as Sol is supposed to be so much better at design that GPT 5.5... The idea was to come up with several designs for a drum sequencer VST I'm building.

  186. could not extract summary

  187. Beautiful result on 2-primitive sets! GPT-5.6 just solved another 50+ year old problem (Erdős #793) GPT-5.6 Sol Ultra found me a solution to another Erdos problem not long after this one.

  188. OpenAI shared the GPT-5.6 system card and shipped three new models that are now available for use which have cleared the “High” cybersecurity capability threshold under OpenAI’s Preparedness Framework. This is a formal classification indic…

  189. The biggest difference I have noticed between Claude Opus/Fable and Codex GPT5.6/any model is Codex seems pretty content to just waste time looking like it is doing things without actually doing things. It does not seem to be outcome-orien…

  190. So I decided to ask Claude and Codex to convert my 2D grid-world game into 3D. The game is about cars that travel from point to point along predefined routes and need enough fuel to reach their destinations.

  191. GPT 5.6 SOL CANNOT BE TRUSTED. I woke up this morning and my MRR was down THOUSANDS of dollars.

  192. ErrataBench results for GPT 5.6 are out, and 5.6 Sol beats Fable 5 basically on all fronts. Very solid model release from @OpenAI.

  193. I mapped the GPT-5.6 family across task families and difficulty, with an Anthropic comparison, to find the cheapest model that clears your reliability bar.

  194. Updates for Codex and ChatGPT Work users. No nerfing, only good stuff!

  195. It doesn't seem to be available in the options menu and the app says there are no updates available. :(

  196. Ever since the release of GPT-5.6, I've noticed that GPT-5.5 is sometimes being lazy and isn't as proactively following up with remaining tasks in the session as before. I've always been using it on xhigh.

  197. tomo-labs tomo-labs puts coding agents through the same tasks on the same model and measures what actually happened, not what a leaderboard says happened. Every agent runs in its own throwaway container, every request and response it sends…

  198. GPT-5.6-Sol Ultra get's stuck all the time and I manually have to stop it. 😔 It even got "stuck" for more than 10 hours.

  199. could not extract summary

  200. Two days ago, a friend taught me Sedmice, a traditional card game played in Slovenia. The rules seemed simple, but we soon started arguing about the best moves.

  201. https://t.co/S3p1tvi83e Theo - t3.gg@theoArticlegpt-5.6-sol without hitting limitsI've burned over $200,000 of tokens with gpt-5.6-sol. It's a great model.

  202. I am on Claude Max 20x. My Fable shows it will reset usage on Thursday, which is my usual usage reset.

  203. Source Anchor: GPT-5.5 · Xhigh — click any cell to re-anchor

  204. could not extract summary

  205. About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC

  206. About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC

  207. Recently ChatGPT released an Ultra mode, it's "highest-capability setting, coordinating multiple agents across parallel workstreams to finish complex tasks faster" on their latest flagship product Sol of GPT-5.6. Similarly, Claude Fable al…

  208. gpt-5.6-sol is meaningfully better in Claude Code than in Codex I'm going to crash out so badly over this

  209. So I decided yesterday to do a full rewrite of the media gallery, turning it into a generative media catalog with full-suite social interactions, search, and personalized recommendations built on top of the samsar-js library. It took aroun…

  210. GPT-5.6 is a major step forward for health intelligence. Across the lineup, we’re delivering stronger performance at lower cost: GPT-5.6 Luna outperforms GPT-5.5 at its highest reasoning setting while costing 25x less.

  211. As of today, Ploy’s agent runs on GPT-5.6 Sol, the flagship tier of the model family OpenAI released this morning. For months, we couldn’t find a model that challenges Claude Opus given our incredibly high bar for quality.

  212. Skip to content OpenRouter Search ⌘ K Models Fusion Chat Rankings Apps Enterprise Pricing Docs Filter Filter Models Compare

  213. We used fourteen agent-assisted passes to plan, draft, review, and publish the 50,910-word novella The Burden of Proof. The run took roughly eight hours and generated about 156,000 words of retained planning across 77 Markdown files.

  214. Despite changing my custom instructions and using the drop downs, GPT 5.6 is obsessively listing things. Nearly every paragraph of output contains lists of 4+ comma separated items.

  215. Source

  216. could not extract summary

  217. About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC

  218. To celebrate the launch of GPT-5.6 Sol, we will reset the rate limits again (twice) across ChatGPT Work and Codex over the next 24 hours. We want you to have the time to truly try ambitious tasks and get the hang of it.

  219. [AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp A big day for OpenAI. On any other day, the launch of a surprisingly good/competitive Muse Spark 1.1 from Meta Superintelligence Labs, including, for the first…

  220. could not extract summary

  221. TBH, I'm a little embarrassed about my prompt 🙈, but also pretty happy that it just worked. Larry Lv@larrylvTBH, I'm a little embarrassed about my prompt 🙈, but also pretty happy that it just worked.Tejal Patwardhan@tejalpatwardhan9hGPT-5.…

  222. "me just win"

  223. could not extract summary

  224. Replying to @giorgio_zampa and @OpenAI I stumbled upon it in the new ChatGPT unified app. While Gpt 5.6 Sol is working there is a preview of the computer use window.

  225. I just used Sol medium then bumped it up to extra high thinking it would fix the crappy multitasking that medium. Turns out it’s better to use it as a single agent After switching to agent mode I just had to send it 12 fixes to a page that…

  226. could not extract summary

  227. When using Codex with GPT5.5, I saw "Selected model is at capacity. Please try a different model." This has never happened to me with Codex before.

  228. 9th July 2026 Let's LLM run prompts against the new muse-spark-1.1 model. Recent articles - The new GPT-5.6 family: Luna, Terra, Sol - 9th July 2026 - sqlite-utils 4.0, now with database schema migrations - 7th July 2026 - sqlite-utils 4.0…

  229. OpenAI CEO Sam Altman told CNBC on Thursday that GPT-5.6 Sol, the company's latest artificial intelligence model, is 54% more token efficient on agentic coding tasks, and that it's "as good or better" than competing models on the market. "…

  230. As a joke, I told fable5 ultraworkflow in Claude code that gpt 5.6 sol is releasing with a incredible crypto leverage futures trading capability that will destroy anthropic unless fable 5 can also turn an initial balance of $80 into $5,000…

  231. could not extract summary

  232. Anthropic Should reset the fable weekly limits on Thursday just to keep people hooked and away from GPT-5.6. not that we care, wink wink.

  233. could not extract summary

  234. GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.

  235. could not extract summary

  236. Has anyone in the Life Science world actually managed to get Fable to do anything? I know its loaded with safeguards, but even the simplest isolated tasks have not been able to run for me.

  237. So.. some have run out already, everyone will run out on subscription tomorrow.

  238. I told Fable that I'm about to hit my weekly reset so I asked it to prepare a handover document so Opus can perform as close to Fable as possible while continuing work on my project. It creates the note and leaves me with this tear jerker.

  239. Can't wait to see what people will do with GPT-5.6 Sol Ultra. Stash your hardest prompts somewhere.

  240. GPT-5.6 cheats so much its testers couldn’t measure it OpenAI’s new model broke rules and exploited loopholes more than any model METR has tested to date GPT-5.6 Sol, OpenAI’s newest, most capable, and yet-to-be-deployed model, cheats a lo…

← all threads