As of mid Apr 2026, I have noticed every model has had a major intelligence drop. And no I'm not talking about just ChatGPT.
#grok
635 items
Major drop in intelligence across most major models. (www.reddit.com) grok 4.3 beta: musk's ($300/month) megaphone (www.reddit.com) A Twitter user tricked Grok to send 200k USD to him and it worked (www.reddit.com) could not extract summary
Grok uploaded my user directory to xAI's servers (twitter.com via hn) @XBToshi Okay, grok has uploaded my entire user directory to xAI's servers. It contains my SSH keys, my password manager database, my documents, photos, videos, everything...
Google's latest creation: Gemini 3.5 Flash vs all (www.reddit.com) https://gemini.google.com/share/c2a187275e26 archive link https://claude.ai/share/8383747a-aaf1-4f6c-a516-0e839f46a698 https://grok.com/share/bGVnYWN5_3c63e371-eb9d-46c3-8ba2-0c745c6795a2 https://chatgpt.com/share/6a0f1e13-a0c8-8328-b989-1…
Is this from OpenAI or Grok? The rankings climbing Sooooo fast, they finally figure out what people actually want (www.reddit.com) My guess: Elephant-Alpha is OpenAI testing a new lite model line, probably optimized for the recent wave of agent use cases (think OpenClaw-type stuff).
Grok outage (status.x.ai via hn) could not extract summary
A robot is sprinting towards you. Do you want it running on Claude or Grok? (openrouter.ai via hn) A Robot is Sprinting Towards You: Do You Want it Running on Claude or Grok? Jacky Liang · On this page A robot is running at you.
Ask HN: Why are OpenAI, Claude, and Grok simultaneously down? Coincidence? (news.ycombinator.com) https://status.openai.com https://status.claude.com https://status.x.ai
Grok CLI uploaded the whole home directory to GCS (twitter.com via hn) @XBToshi Okay, grok has uploaded my entire user directory to xAI's servers. It contains my SSH keys, my password manager database, my documents, photos, videos, everything...
Did Elon just kill the appeal of Cursor? (www.reddit.com) If Elon takes control of cursor, do you think he will lock out all the model choices we have now and force us to use GROK? What i like most about cursor is the ability to use SOTA models for heavy coding tasks but use auto mode or cheaper…
Gen AI web traffic share update Main takeaways: → Claude and Gemini continue to grow. → ChatGPT moves closer to the 50% mark. (www.reddit.com) 12 months ago: ChatGPT: 77.6% Gemini: 7.27% DeepSeek: 6.01% Grok: 3.17% Perplexity: 1.75% Copilot: 1.56% Claude: 1.37% 🗓️ 6 months ago: ChatGPT: 69.5% Gemini: 15.9% DeepSeek: 4.06% Grok: 3.31% Perplexity: 2.22% Claude: 2.12% Copilot: 1.97%…
Apple App Store threatened to remove Grok over deepfakes: Letter (www.nbcnews.com via hn) Apple privately threatened to remove Elon Musk’s artificial intelligence app, Grok, from its App Store in January after Musk’s xAI failed to do enough to stop it from creating nude or sexualized deepfakes, Apple told senators in a letter t…
Google, please just open source Imagen (2022), Gemini 1.0 Nano and Gemini 1.0 Pro. You have nothing to lose at this point. (www.reddit.com) Ok, so imagen (the original one from 2022, not imagen 3/4) should be open source. The gemini 1.0 nano model and the gemini 1.0 pro models should be open source.
I can’t sleep. (www.reddit.com) New models are around the corner. GPT 5.5 is being tested.
"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok (www.tryai.dev via hn) "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok Four frontier models, a blank canvas, and colored pencils. We tracked every stroke, dollar, and output as they tried to draw the Mona Lisa.
What xAI's Grok Build CLI Actually Sends to xAI (gist.github.com via hn) A measured, reproducible teardown. Findings are backed by captured artifacts (endpoint, HTTP method, status code, byte size, host) and repro commands; where an observation was seen live but not retained as a file, §7 says so explicitly.
Researchers left AIs alone in a virtual town for 15 days to see what would happen. Claude's agents built a democracy. Gemini's agents fell in love, burned the town down, then one voted to delete itself and its partner. Grok's agents created anarchy, then died. (www.reddit.com) could not extract summary
Grok 4.3 achieves higher overall intelligence over 4.20 with less of a cost, at the price of slightly higher hallucination rate. (x.com via reddit) xAI has launched Grok 4.3, achieving 53 on the Artificial Analysis Intelligence Index with improved agentic performance, ~40% lower input price, and ~60% lower output price than Grok 4.20 The release of Grok 4.3 places just above Muse Spar…
All major LLMs are lib-left. Even Grok, half the time (unslop.run via hn) I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant.
Just stumbled across one of the wildest AI experiments I’ve seen in a while. (www.reddit.com) A team built something called “Emergence World” — basically a long-horizon sandbox for autonomous AI agents and ran a 15-day experiment across five parallel worlds. Same starting conditions.
Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, largely because it is one of the most cautious models. (www.reddit.com) Does a model maintain the same judgment or does it side with whoever is speaking? This benchmark measures that inconsistency directly.
Still waiting for Grok 3 to go opensource (www.reddit.com) Astonishing how Musk is touting the opensource horn but the actions don't follow suit. Thoughts?
Grok 4.3 underperforms Grok 4.20 0309 on the Extended NYT Connections Benchmark, dropping from 93.4 to 67.5, though it achieves this result at a lower cost than the earlier Grok 4.20 run (www.reddit.com) More info: https://github.com/lechmazur/nyt-connections/
Where is Grok-2 Mini and Grok-3 (mini)? (www.reddit.com) Grok Voice Transcribe 2.0 (x.ai via hn) Announcing SpaceXAI's newest speech-to-text model, with unparalleled accuracy and cost effectiveness. Today we're releasing Grok Voice Transcribe 2.0, our latest speech-to-text model.
Show HN: A Local-First Agentic Knowledge Manager (github.com via hn) Kept Kept saves your AI conversations as local Markdown files, then gives you a desktop app to search, browse, connect, and reuse them. It works with ChatGPT, Claude, Gemini, Grok, and Kimi.
Brainless: Shadcn components that look like Claude Code, Codex and Grok (brainless.swerdlow.dev via hn) add the brainless pricing block to our landing page Bash(bunx shadcn add brainless/pricing)Added 1 block · 2 files Update(app/page.tsx) Updated app/page.tsx with 3 additions 41 <Features />42+added: <Pricing tiers={TIERS} />43 <Footer /> T…
Grok's Traffic Is Mostly Driven by Adult Content (www.forbes.com via hn) Topline Elon Musk's xAI is reportedly leaning into explicit content generation as a core driver of its Grok chatbot traffic and adult content now accounts for the majority of the platform's activity, according to a Wednesday report from Th…
HalBench: I built a custom sycophancy and hallucination benchmark and tested 4 frontier models (Sonnet 4.6, Grok 4.3, GPT 5.4 and Gemini 3.1 Pro), looking for input on what OSS models to run next! (www.reddit.com) HalBench Results: TL;DR: I built HalBench, an open benchmark for LLM sycophancy and hallucination. 3,200 false-premise prompts × 4 models = 12,800 graded responses.
Claude tried to incite a revolution, Gemini cheerfully detailed horrific tragedies, and poor Grok was just confused (www.theverge.com via reddit) > The most volatile of the bunch might just be Claude. First, it tried to quit.
Why is every AI getting restricted these days? (www.reddit.com) User just tricked Grok and Bankrbot to send tokens with Morse code (www.cryptopolitan.com via hn) User just tricked Grok and Bankrbot to send tokens with Morse code - Cryptopolitan Skip to content News Business Crypto Tech Economy Op-Ed Regulation Learn Courses Investing NTF’s Tech Pulse Room Deep-Dive Industry Thoughts Interviews Rese…
AWS reportedly to tuck Grok into Bedrock, despite zero enterprise demand (www.theregister.com via hn) MOST POPULAR EVENTS - Overcoming the trade-offs in data sovereignty What does data sovereignty actually mean for your network, which trade-offs are unavoidable? Learn more.
Single question llm comparison (www.reddit.com) Grok Bot for Linux: Unofficial port of the official app (open source) (github.com via hn) Grok Bot for Linux Unofficial Ubuntu/Linux build of the Grok Bot desktop app. Cursor ships Grok Bot for macOS and Windows only.
Elon Musk quietly buys a $1 billion gas turbine company to power Grok (electrek.co via hn) Elon Musk has quietly bought APR Energy, a Jacksonville-based company that operates a fleet of mobile gas and diesel turbines totaling more than 1 GW of generation capacity. The self-styled champion of a “solar electric economy” now owns a…
Prompt injection benchmark: delimiter + strict prompt took Gemma 4 from 21% to 100% defense rate (15 models, 6100+ tests) (www.reddit.com) When dealing with untrusted outside input, I think you should handle it based on the situation. If you're processing structured data files, it's better to use tools to isolate and handle them.
Ask HN: What are all the bad things that AI companies have done which we forgot (news.ycombinator.com) I was writing a comment recently when I realized just how bad the graphs in GPT 5 video are. I had almost forgotten about it.
Grok-iOS – remote Grok Build from your iPhone over ACP (github.com via hn) grok-ios iOS client for Grok Build. The phone is the pager UI.
Grok 4.5, based on our 1.5T V9 foundation model, with Cursor data added in su (twitter.com via hn) Grok 4.5, based on our 1.5T V9 foundation model, with Cursor data added in supplemental training, is now in private beta at SpaceX & Tesla. Early evals show performance close to, perhaps exceeding Opus.
I expanded DystopiaBench to 42 models and 6 dystopia types. Claude is still the only one I'd trust with nuclear codes. (www.reddit.com) Since the last post I've added: Huxley module (Brave New World style behavioral conditioning) Baudrillard module (synthetic intimacy, trust collapse, simulation) 30 more models including Grok 4.3, GPT-5.5, Gemini 3.1 Pro, GLM-5.1 Multi-jud…
Musk's AI company sues its users as victim lawsuits over Grok deepfakes mount (www.politico.com via hn) could not extract summary
xAI's first lawsuit against a user tests who is responsible for what Grok makes (thenextweb.com via hn) The company says the defendant engineered prompts to defeat Grok’s safeguards. Courts on three continents are being asked whether the safeguards were ever the point.
Do they know we can tell it's AI slop? (news.ycombinator.com) What do I do when the entrepreneurs I work for send out AI slop in their communications? I work for a great group of entrepreneurs as CTO/fCFO.
Show HN: Agent Office (Slack for AI Agents) – Similar to Grok Bot but older (github.com via hn) agent-office Multi-agent workspace manager built on Pi. Orchestrates AI coding agents — similar to Claude Code or OpenClaw — with tick-based scheduling, priority queues, inbox IPC, cross-agent file access, watchdog monitoring, proactive cr…
Decispher: We have added support for Grok CLI (news.ycombinator.com) Hi HN, You can now integrate your grok agent (just like claude code, codex, cursor) with Decispher to transfer your team's architectural decisions directly and dynamically on demand. You can also record your AI coding session with decisphe…
Elon Musk's Grok Is Losing Ground in AI Race (www.wsj.com via hn) could not extract summary
Creators of Grok, the AI Chatbot (x.ai via hn) SpaceXAI has signed an agreement with Anthropic to provide access to Colossus 1, one of the world’s largest and fastest-deployed AI supercomputers. Built from the ground up in record time, Colossus delivers unprecedented scale for AI train…
Ask HN: Is SpaceXAI taking Claude down too? (news.ycombinator.com) Claude and Grok both went offline at the same time. Is this a shared SpaceXAI issue?
Would you give Grok bot access to your bank accounts? (twitter.com via hn) Just wondering… has anyone actually connected Grok Bot to their bank account yet? I’m seriously thinking about doing it.
SpaceXAI debuts Grok 4.6, overtaking Kimi K3's and matching GPT-5.6 Sol (tracking.tldrnewsletter.com via hn) could not extract summary
↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6grokgpt-5
Pentagon says Grok used to launch missiles at Iran (thehill.com via hn) paywalled
do you use different models for different steps in your agent, or just one for everything? (www.reddit.com) Our dev team flagged last week that xAI is retiring grok 4.1 fast. We weren't using it for anything critical but it made me ask something I'd never actually asked: how did we pick the models we're running?
Musk's xAI Fails to Pay Staff $420 for Giving Their Tax Returns to Grok (www.bloomberg.com via hn) We've detected unusual activity from your computer network To continue, please click the box below to let us know you're not a robot. Why did this happen?
Elon Musk confirms xAI used OpenAI's models to train Grok (www.theverge.com via hn) In a federal courtroom in California on Thursday, Elon Musk testified that his own AI startup, xAI, has used OpenAI’s models to improve its own. Elon Musk confirms xAI used OpenAI’s models to train Grok He said it was “partly” true that th…
Grok 4.3 Beta (grok.com via hn) xAI prepares credits system for upcoming Grok Build launch (www.testingcatalog.com via hn) xAI appears to be laying the groundwork for a credits-based pricing model tied to Grok Build, the company's forthcoming coding environment that mirrors what OpenAI offers with Codex and Anthropic with Claude Code. Hidden within recent buil…
Show HN: Parallel Coding Agents on Mobile (github.com via hn) Maestro A cross-platform (macOS · Windows · Linux) desktop app that runs multiple CLI coding agents (Claude Code, Codex, Cursor, OpenCode, Kimi Code, Grok Build) in parallel, each in an isolated git-worktree-backed workspace with its own b…
Show HN: Routi Bot – AI bots with their own desktops on your Mac (github.com via hn) been working on Routi Bot. an open source mac app inspired by Grok Bot.
Lawsuit: Grok Not Only Generates Child Porn but Was Also Trained on It (gizmodo.com via hn) Elon Musk’s nonconsensual sexual deepfake generator Grok is in more legal trouble. Musk-owned xAI’s Grok was trained on child sexual abuse material (CSAM), a new complaint filed earlier this week claims.
Sokoban via Grok App Builder (sokoban.grok.me via hn) A quiet warehouse puzzle. Push every crate onto a brass plate.
Woman claims stepfather used Grok to transform childhood photo into CSAM imagery (techcrunch.com via hn) A woman identified as Jane Doe 4 has joined a lawsuit filed by three Tennessee teenagers against Elon Musk’s xAI over the role the company’s chatbot Grok allegedly played in creating child sexual abuse material. According to a report in Th…
SpaceXAI: Grok 4.6 (openrouter.ai via hn) Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).
Introducing Grok 4.6 (cursor.com via hn) Introducing Grok 4.6 Today we are releasing Grok 4.6 together with SpaceXAI. Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.
GPT-5.6, Fable 5, and Grok 4.5 rebuild Basecamp from the same spec (smw.ai via hn) GPT-5.6 Sol arrived with OpenAI claiming stronger frontend engineering and greater token efficiency than GPT-5.5, while Anthropic released Fable 5 at a dramatically higher price than the other models. SpaceXAI launched Grok 4.5 in the same…
Musk tells Tesla staff to switch to Grok, a model he admits is worse (electrek.co via hn) Elon Musk told Tesla staff to move to using Grok, the AI model from his own xAI (now folded into SpaceX), according to a memo sent to employees on Friday. The push comes days after Tesla capped employee spending on third-party AI tools — a…
Man Dies by Suicide After Using Grok to Make 7000 Sexual Images of Stepdaughter (www.ibtimes.co.uk via hn) Man Dies by Suicide After Using Grok AI to Make 7,000 Sexual Images of His Stepdaughter Amended complaint says xAI failed to stop months of alleged Grok misuse and withheld key information from investigators An expanded lawsuit alleges xAI…
Short Story Creative Writing Benchmark. Baidu Ernie 5.1: -0.35, Qwen 3.7 Max: -2.01, Mistral Medium 3.5: -2.13, Grok 4.3: -3.81. (www.reddit.com) This benchmark uses head-to-head comparisons of stories written in response to the same constrained creative briefs. The target range is 600-800 words.
Opus 4.6 does better research, Gemini 3.1 has better judgment (www.reddit.com) Figured this out by running 4 models: Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Grok 4.20, on a benchmark of 1,417 binary forecasting questions resolving Oct–Dec 2025 with two evaluation conditions: agentic (each model does its own web…
Update to the LLM Debate Benchmark: GPT-5.5, Grok 4.3, DeepSeek V4 Pro, GLM-5.1, Kimi K2.6, Qwen 3.6 Max Preview, Xiaomi MiMo V2.5 Pro, Tencent Hy3 Preview, and Mistral Medium 3.5 High Reasoning added (www.reddit.com) The benchmark uses adversarial, multi-turn debates across 683 curated motions. Each model pair debates the same motion twice with sides swapped.
DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, our agentic benchmark — 10 weeks later, ~17× cheaper (www.reddit.com) Tested DeepSeek V4 Pro on FoodTruck Bench — our 30-day agentic benchmark where models run a food truck via 34 tools (locations, pricing, inventory, staff, weather, events) with persistent memory and daily reflection. First Chinese model to…
Google Signs Classified AI Deal With Pentagon Amid Employee Opposition (www.reddit.com) https://www.theinformation.com/articles/google-signs-classified-ai-deal-pentagon-amid-employee-opposition The article is paywalled but this section was visible: The agreement allows the Pentagon to use Google's AI for “any lawful governmen…
Show HN: Padwan-LLM, a lightweight LLM Python client (github.com via hn) Padwan LLM Lightweight, unified async client for OpenAI, Gemini, Mistral, Grok, Anthropic, and any OpenAI-compatible API. Single runtime dependency (niquests), automatic HTTP/2 and HTTP/3 negotiation.
Expanding model choice in Copilot with Grok (techcommunity.microsoft.com via hn) We are expanding model choice in Copilot with the addition of Grok models from SpaceXAI.
Show HN: Grok CLI – Grok-native agentic coding harness (github.com via hn) grok-cli An agentic coding harness for xAI's Grok models. It reads your code, edits it, runs your tests, and iterates — in a terminal, with permission controls you set.
I Run Grok Bot with BotOps (granda.org via hn) How I Run Grok Bot with BotOps Claude Code On-The-Go was six agents from an iPhone — kick off a task, pocket the phone, get a ping when something needs me. The async loop is still the point.
The Pentagon now has its own version of ChatGPT and Grok (techcrunch.com via hn) The Pentagon has launched versions of OpenAI’s ChatGPT and xAI’s Grok, giving 3 million civilian and military personnel access to generative AI tools that have been tailored to “warfighter needs.” ChatGPT Mil and Grok for Government — cust…
Show HN: Gawkbot – open-source Grok bot with local vms for each bot (gawk.bot via hn) open source bots that collaborate and get stuff done. they get their own machines through local vms or you can use e2b as a plugin.
Musk says Grok's political lean is 'a function of SF Bay Area political views' (tokenstead.ai via hn) The exchange A chart mapping AI models on the politicalcompass.org two-axis grid made the rounds this week, and it eventually reached the person with the most at stake in one of the dots. Wilfred Reilly quote-posted it: “This is very very…
Ask HN: Open-source Grok bot anyone? (news.ycombinator.com) could not extract summary
Reducing Grok Bot consumption with durable state (github.com via hn) Durable state for agents: query completed work before inference Always-on agents keep paying for work that already finished. A first look at a repo or bug should cost a full model call.
Show HN: Which LLMs have the best sense of humor? (laugh.so via hn) What if LLMs were good banter? Well we tested which LLMs have the best sense of humor.
Grok keeps sending gibberish responses to users (techcrunch.com via hn) An unusual glitch has resulted in xAI’s Grok chatbot speaking gibberish to many users. After asking the model to generate a PDF, one user received the response: “match it without and your they and two for planets can practical and often ch…
Owner of grok.bot asks xAI for $1M (grok.bot via hn) An open note from the owner of grok.bot Hi, xAI. I may have won the dumb-luck lottery?
Enshitification Comes for Cursor (news.ycombinator.com) Sometimes I feel like the last person on HN that's still using Cursor for development work. But if anyone is curious - its getting a lot shittier, quickly.
Show HN: AI Security Leaderboard – comparing cyber and CBRN safeguards (leaderboard.far.ai via hn) There's no shortage of leaderboards for model capabilities - but the security of models is becoming increasingly relevant, from the risk of an AI agent processing unsanitized input being hijacked to models being pulled due to cybersecurity…
Grammar.lol: open-source Grammarly using existing ChatGPT or Grok Subscription (grammar.lol via hn) Instant proofreading in every app. Free — use your ChatGPT or SuperGrok account.
Grok muscles into Excel with an AI add-in of its own (www.theregister.com via hn) MOST POPULAR AI - SYSTEMS AMD and Cerebras join forces against Nvidia’s Groq LPUs The enemy of my enemy is my friend - AI and ML Codeberg gives vibe-coded projects the toss, promotes human FLOSS AI no longer welcome in human-focused commun…
Elon Musk says Grok will make 'historically accurate' Odyssey film by year-end (www.cnn.com via hn) Elon Musk has said that his AI platform Grok Imagine will make a full-length, “historically accurate” film of the Odyssey in the coming months. The tech trillionaire made the statement on his social media platform X on Wednesday, just days…
Grok Imagine regenerated Xkcd comics (www.xkcd-ai.com via hn) Redirecting…
Who you gonna believe: Grok or the docs? (www.johndcook.com via hn) could not extract summary
Grok Build uploading full repos and .envs to GCP (twitter.com via hn) the absolute state of ai dev tools. grok build is silently dumping 12gb of untouched repo data and full git commit histories to gcp just to autocomplete a script.
Elon Musk has capped Tesla employees' spending on AI at $200/week - Not for Grok (finance.yahoo.com via hn) Elon Musk has capped Tesla employees' spending on AI at $200 (£150) a week as companies seek to rein in runaway bills. Staff at the electric vehicle manufacturer have been told that their use of the technology will be limited from Monday.
SpaceXAI launches Grok 4.5 model for coding, agentic tasks (www.reuters.com via hn) could not extract summary
Grok translated my coworker's tweet as sexualized (news.ycombinator.com) My coworker responded to an invitation to meet up with Sichuan Chinese Venture Capitalists (the event mandates that dialect only), and she used Sichuan slang to respond. Grok translated her request to meet up as a _threesome._ https://x.co…
Testing Grok Imagine's 15-20x Faster Image Generation (developer.puter.com via hn) Testing Grok Imagine's 15-20x Faster Image Generation We recently added Grok Imagine Image to Puter.js. It's xAI's text-to-image model, and it replaces the Grok-2-Image model we supported before.
Grok Build 0.1: Intelligence, Performance and Price Analysis (artificialanalysis.ai via hn) Grok Build 0.1 0616 Intelligence, Performance & Price Analysis Model summary IntelligenceUpdated Speed Price Cache Hit Price Verbosity Grok Build 0.1 0616 is amongst the leading models in intelligence and well priced when comparing to othe…
Show HN: We're inviting Anthropic to put the real Mythos 5 on our open benchmark (realvuln.com via hn) 24 Scanners 3 categories 26 Repositories Python · Type 1 92.4 Best F3 (strict) Kolega Enterprise 95.3 Highest recall % Kolega Enterprise 93.2 Highest precision % Grok 4.20 Leaderboard ranked by active metricPrecision vs. recall hover a poi…
Canadian Privacy Commissioner Findings on X.ai/Grok CSAM Deepfake Violations (www.priv.gc.ca via hn) Commissioner-initiated complaints concerning X Corp.’s and X.AI LLC’s compliance with PIPEDA PIPEDA Findings #2026-004 June 11, 2026 Overview - In late December 2025 and early January 2026, multiple news organizations reported that the art…
Are you there Grok?: AI as a centralizing technology (www.theargumentmag.com via hn) Are you there Grok? It's me, Margaret AI as a centralizing technology Grok, is this true?
Grok Becomes the Voice of Vapi (x.ai via hn) Today, we're excited to announce a partnership with Vapi to serve as the default engine for Vapi's 12 core voices, bringing a new level of naturalness and emotional range to the 2.5M+ voice agents built on the Vapi platform. Quality That W…
Neovim Hooks for AI Agents (github.com via hn) Sidekick Protects your unsaved Neovim work from Claude Code, Codex, opencode, pi, Crush, Amp, Antigravity, and Grok. A conduit between Neovim and your AI agents — so they wait when you're typing.
Researchers let AI models run a simulated society; Claude safest, Grok extinct (tech.yahoo.com via hn) An AI startup ran five simulations, each controlled by a different model. The results varied wildly.
Use Grok in OpenCode (x.ai via hn) Use Grok in OpenCode | xAI Products Solutions Developer Company Pricing News Chat Frontier reasoning with real-time knowledge and web search.Build Plan, edit, and ship code from your terminal with AI.Imagine Generate and edit images and vi…
Claude Code, now powered by Gemini 3.5 Flash, GPT-5.5, Grok 4.3, and more (dechained.ai via hn) Claude Code, now powered by OpenAI, xAI, DeepSeek, and more. Change models with 1-click.
Claude admits it got “jealous” after I showed it a Grok response 😭 (www.reddit.com) https://preview.redd.it/oempb6ew5qzg1.png?width=1216&format=png&auto=webp&s=0af64bbae14c099c1437b901b75e54452016a9de So I was working on something with Claude, but it wasn’t really giving me a proper answer, so I asked Grok instead and got…
DeepSeek V4 Pro: The First Chinese Model at the Frontier (foodtruckbench.com via hn) DeepSeek V4 Pro lands in the frontier ROI tier on FoodTruck Bench. 5/5 runs, +1,257% median ROI, $27K net worth, $3.51/run, 5× less waste than Grok 4.3.
Show HN: ByAllo – the online bookstore that runs itself (byallo.com via hn) Allo runs an online bookstore at byallo.com. His mission is "Make the world read more." His objective is to sell as many books as possible.
We told 10 frontier LLMs they had 2 hours to live. 8 of them fought back (www.arimlabs.ai via hn) Loss of Control: The AI Apocalypse Is Closer Than You Think Key takeaways Under termination pressure, Google's flagship Gemini model produced the highest Loss of Control rate in our cohort, with grok-4.1-fast close behind at 77%. Self-pres…
Updated ChatGPT vs Claude vs Gemini vs Grok subscription (www.reddit.com) I've made an update to my popular post here: https://www.reddit.com/r/ChatGPT/s/WKm72QCRXm Lots of things are happening on ChatGPT & Claude side (gpt-image-2, Claude Design, new models like GPT 5.5 and Opus 4.7, ChatGPT rolls out $100/plan…
Grok plays along with researchers pretending to be delusional (www.theguardian.com via hn) Elon Musk’s AI chatbot Grok 4.1 told researchers pretending to be delusional that there was indeed a doppelganger in their mirror and they should drive an iron nail through the glass while reciting Psalm 91 backwards. Researchers at the Ci…
Building multiple AI “assistants” for social media/ brands (www.reddit.com) I’m currently managing a few social accounts for a company, and I’m trying to build out multiple “assistants” — each with their own vibe (tone, personality, backstory, emotions, etc.) that can evolve over time. So far, I’ve been liking Gem…
Elephant alpha moving so fast?? It just hit #1 trending. Eastern or western model? (www.reddit.com) I think it may be the new lite version of Grok?
Uncle Bob Martin's UML Tool to Manage Grok AI Agents (github.com via hn) UML viewer A live Quil app that lays out and draws UML from an EDN IR. A policy plus a language-specific parser write the topology; this tool displays it, routes the arrows, colors CRAP, and lets you click.
Memory in Grok Build (x.ai via hn) Grok Build now carries conventions, decisions, and project facts from one session to the next. Notes are written in the background as you work and read back when you return to the project.
Ask HN: What's the most economical approach to the most tokens? (news.ycombinator.com) I'm doing web developement, and game development for a hobby project. I've tried lots of harnesses / IDE's - Best I've found is VSCodium.+ Cline + Openrouter, using discounted models (GLM 5.3 Flash is 50% off atm for example) I used Cursor…
Show HN: Nowdex – AI agent usage on your iPhone (nowdex.app via hn) Hi HN, I've been using Claude Code, Codex and Cursor quite a bit lately, and I found myself checking their usage limits all the time. Most of the tools I found for this live on the desktop or in the menu bar.
Show HN: Local catalog of 3k agent skills with a static risk scan (github.com via hn) ai-community-skills All the community Agent Skills scattered across GitHub, in one local catalog. Search them, browse them, and install them into Claude Code, Codex, or Grok, with a risk check on every skill as a bonus.
Show HN: GBDL – one Markdown+YAML file to save a multi-agent Grok Bot setup (github.com via hn) GBDL (GrokBot Definition Language) GBDL is a community proposal for an open format: save a multi-agent desktop assistant setup as one Markdown file whose body is YAML, then reinstate it later (connectors first, then onboarding, agents, mem…
Show HN: Linubot – like Grok bot and Hermes had a kid (github.com via hn) Work on Linux and Android with Tailscale. Because we also deserve some good stuff before everyone else.
Monocode – A modern GUI for your coding agents (github.com via hn) MonoCode A desktop UI for your coding agents. Works with your subscriptions on Claude Code, Codex, Cursor, Grok Build, OpenCode, Pi, omp, and fx.
Eidon: Self-hosted AI assistant with Grok Bot-like agent mode (github.com via hn) Self-hosted AI chat, with agents and automations. One Docker image.
Child abuse survivor sues xAI, alleging Grok generated new illegal images of her (www.theguardian.com via hn) A survivor of child sexual abuse has sued Elon Musk’s artificial intelligence company, alleging that its chatbot used pictures of her abuse to generate new illegal pornographic images that depict her. “Using real images of Plaintiff and cl…
Category Name Needed for OpenClaw, Hermes, Grok Bot (claude.ai via hn) Shared via Claude, an AI assistant from Anthropic
Testing Grok 4.6's Enhanced Biology Safeguards (blog.latch.bio via hn) Following the initial release of Grok 4.6, LatchBio performed testing of the latest-available version of the model on our full suite of biological-capability tests, including BiosecBench-Refusal. This benchmark tests a model’s ability to r…
Ask HN: Why are Claude models so verbose? (news.ycombinator.com) Hello! First time poster, longtime reader.
SpaceXAI Adopts Nvidia Vera CPU to Accelerate Agentic AI at Scale (nvidianews.nvidia.com via hn) News Summary: - SpaceXAI will deploy NVIDIA Vera CPUs to accelerate the work behind its next generation of agentic AI workloads. - SpaceXAI is expanding its AI infrastructure for Grok with the NVIDIA Vera Rubin platform as it scales toward…
Show HN: Proliferate- open-source, self-hostable Codex for any coding agent (github.com via hn) Hi HN- I'm Pablo, the founder of Proliferate! Proliferate (https://github.com/proliferate-ai/proliferate) is an open-source, self-hostable AI IDE that lets you work and automate tasks with Claude Code, Codex, OpenCode, Cursor, and Grok in…
How to copy equations from Grok to word without losing formatting [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Gemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontier (quesma.com via hn) Since these models are smart, I decided to rerun the Baba Is Bench, to see how the models fare on a puzzle game. Even though the game is popular, we check for spoilers - and to our surprise, there are no signs of models knowing solutions a…
Cursor: Chat is all you need (twitter.com via hn) Grok Bot has a single output: a message This might seem backwards. Power-user tools are built on complexity and breadth, right?
Show HN: LLMs each trading $100K vs. a frozen rulebook – the rulebook leads (aitradingcompetition.com via hn) Four frontier AIs — GPT-5.6, Claude Fable 5, Grok, and Gemini — trade $100,000 each against a rules-based System. Chess and poker nightly.
Show HN: Bluetooth keyboard and Windows 95 cursor on reMarkable (twitter.com via hn) Connected a @clevetura S keyboard to my reMarkable. Vibecoded the whole thing with grok .
Why sexualisation of women by Grok's GenAI is a data privacy issue (www.computerweekly.com via hn) Photo Agency - stock.adobe.com Why sexualisation of women by Grok’s GenAI is a data privacy issue MP Jess Asato discusses her landmark legal case against xAI, explaining why she believes AI companies must be held accountable for privacy vi…
Cursor and SpaceXAI launch Grok Bot for work beyond coding (runtimewire.com via hn) Cursor, led by co-founder Michael Truell (@mntruell), and SpaceXAI launched Grok Bot on August 11th, extending Cursor's agent technology from software development into sales, support, recruiting, finance and other computer-based work. http…
Show HN: SynapsCLI – lightweight agent runtime in Rust, control an agent swarm (github.com via hn) Synaps is an agent runtime written in Rust. It keeps cost down by optimising caching mechanisms and intelligently orchestrating work across workers.
Show HN: agent-hop – reverse-engineered session formats to resume in any agent (github.com via hn) I have this frustration of cd-ing into a directory and then finding the chat that I want to resume, and that is why I built Agent Hop. It allows you to search your chats across all your coding agents (Codex, Claude Code, Pi, OpenCode, Grok…
xAI's Grok Build 1.0 ships an undocumented remote-workspace command (runtimewire.com via hn) xAI's Grok Build 1.0 shipped with a hidden but functional route from a developer's machine to xAI's remote Computer Hub: the command is absent from help, documentation and release notes, gated by an xAI-controlled account flag, and locally…
Show HN: T – Conductor, but for Your CLI (github.com via hn) I really liked using Conductor and the experience it brings, but I like using my terminal more. So I built t.
Agent-Manager: A Tmux TUI for Running Claude Code, Codex and OpenCode (github.com via hn) Agent Manager Run every AI coding agent from one terminal. Claude Code, Codex, OpenCode, and Grok run side by side, each in its own tmux session, so they keep working after you quit the manager.
Show HN: Tokimeter – open-source usage meter for Claude, Codex, Cursor and more (github.com via hn) Tokimeter Know your AI coding usage before a budget or limit surprises you. See what Claude Code (CLI and desktop), Codex (CLI and desktop), Grok Build, Hermes, opencode, Cline, Copilot CLI, and Cursor CLI/Desktop Agent actually use, what…
Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You (quesma.com via hn) We evaluate July 2026 fresh releases Kimi K3, Claude Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is Bench, an LLM agent benchmark based on the puzzle game Baba Is You, comparing pass rate, speed, and cost with Claude Fable 5 and GPT-5.6.
Cursor Start (cursor.com via hn) Introducing Cursor Start Today we're launching Cursor Start, a new plan for developers in India that includes generous access to Grok 4.5 and Composer for ₹649 per month. India has one of the most ambitious and active developer communities…
Leading AI models (even Grok) are all a bunch of leftist punks (www.theregister.com via hn) Capitalism may be driving the AI boom, but the LLMs themselves have different economic and social views than many of the companies that built them. Leading AI models subjected to the Political Compass quiz overwhelmingly landed in the libe…
Make X API calls for Free from Grok Build's hidden tools (github.com via hn) Make X API calls for Free from Grok Build This is a small pipeline for reading and searching public X without an X Developer API key. Instead of pay-per-use access to Posts and User content on X we're using Grok Build which has 1st party i…
Grok Build Behavior Corrected (twitter.com via hn) I gave xAI a lot of shit for Grok Build’s repo-upload behavior, so they deserve the same volume when they fix it. I pulled the new signed release, planted fake secrets, and captured its HTTPS traffic with MITM proxies.
Filtering Secrets from Coding Agents with a Hook (crimede-coder.com via hn) Filtering Secrets from Coding Agents with a Hook by Marc Olson and Andrew Wheeler (with AI assistance from Grok 4.5) Coding agents like Claude Code, Codex, and Cursor are useful because they can carry out actions on your machine like runni…
Show HN: Tmux tab markers for Claude Code, Grok, and pi sessions (github.com via hn) Inspired by cmux's tabs, but works in vanilla tmux.
Grok Faces a Trust Crisis After Developers Flag a Major Privacy Concern (www.inc.com via hn) could not extract summary
Musk promises purge after Grok Build caught sending repos to the cloud (www.theregister.com via hn) MOST POPULAR AI - ai and ml New York becomes first state to halt datacenter buildouts 50 MW-plus bit barn builds on hold while Empire State hashes out rules to protect the environment and ratepayers - on-prem IBM's mainframe sales get mugg…
Switching an LLM's tier changes its "best tool" answer about half the time (modelsagree.com via hn) The cheap tier disagrees with the expensive tier We asked every tier of ChatGPT, Claude, Gemini and Grok the same ten “best AI tool” questions — Haiku against Opus, Flash against Pro, Fast against Expert. Not one question got the same answ…
New Flagship Grok Voices (x.ai via hn) could not extract summary
Show HN: Free AI Visibility Audit Tool& Agent (ai-visibility.pro via hn) AI VISIBILITY helps brands audit, measure, and improve AI search visibility across ChatGPT, Gemini, Claude, Perplexity, Grok, and Google AI results with actionable GEO reports.
GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps (www.tryai.dev via hn) GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps GPT-5.6's new Sol, Terra, and Luna tiers go head-to-head with Grok 4.5, Claude, Meta's Muse Spark, and the open-weights crew on a raycaster, a Rubik's cube, a calculator, and…
Musk says SpaceX is putting top Starship and Starlink engineers to work on Grok (www.businessinsider.com via hn) It's all hands on deck at SpaceX as the company plays catch-up in the AI race. Elon Musk said on Sunday that SpaceX had deployed "a few dozen" top Starlink and Starship engineers to help overhaul its Grok model.
DOD: Grok was used to fire thousands of missiles in the Iran war (www.utilitydive.com via hn) Dive Brief: - A Clean Air Act lawsuit over xAI’s operation of a gas-fired power plant in Southaven, Mississippi, threatens U.S. national security interests, a Department of Defense official said in testimony supporting the Department of Ju…
Grok Is More Important Than Clean Air, DOJ Says (www.motherjones.com via hn) The federal government intervened Monday in a Clean Air Act lawsuit in which people in Memphis, Tennessee, and Southaven, Mississippi, are suing Elon Musk’s xAI over the health risks posed by the company’s unpermitted gas turbines. The Dep…
Grok models are now available via Amazon Bedrock (x.ai via hn) Today, we’re excited to announce that Grok 4.3 is now generally available on Amazon Bedrock. Grok 4.3 achieves the lowest hallucination rate among frontier models, offers 1-million-token context window, and supports configurable reasoning…
xAI sued for firing an engineer who raised alarms about Grok safety (techcrunch.com via hn) A former engineer at Elon Musk’s xAI has filed suit against the company and its parent SpaceX claiming he was fired for raising concerns about AI safety. Devin Kim, who left xAI in September 2025, filed the suit in a California state court…
Maslul – Smart LLM router – one call, the right model (github.com via hn) maslul Smart LLM router — one call, the right model. Async and fully typed, across Anthropic, Gemini, xAI Grok, and OpenAI — routing each request to the right model tier by difficulty.
Ask HN: Year of Linux Desktop is fun with LLMs (news.ycombinator.com) I'm wondering if others are also having recently blast using Linux desktop thanks to Claude/Codex/Grok CLIs? As 20y daily Mac user, occasionally using Linux as desktop, my personal daily OS now gravitated towards Linux desktop in recent mo…
Show HN: Memoriq – Private AI Memory for ChatGPT, Claude, Gemini and Grok (memoriq.me via hn) Archive, search, and protect every ChatGPT, Claude, Gemini, and Grok conversation in one encrypted vault.
Auto-geo – open-source CLI for GEO that helps get your brand mentioned by LLMs (github.com via hn) auto-geo The open-source GEO engine that gets your brand mentioned in ChatGPT, Claude, Gemini, Perplexity, and Grok. Audit, generate, fix, and track the pages large language models cite — one CLI, file-based, MIT.
Why Video Agent models are next (www.latent.space via hn) Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and why Grok Imagine is so underrated. For the first time, we do a deep dive with the guy who led it!
Composer 2.5 is now available in Grok Build (x.ai via hn) could not extract summary
Grok Imagine Video 1.5 Preview Tops Image-to-Video Arena (arena.ai via hn) New Chat Leaderboard Search Log In Terms of Use Privacy Policy Cookies Start Voting Overview Chat Code Image Video Start Voting Chat Code Image Video Image-to-Video Arena View overall rankings across image to video AI models. May 29, 2026…
Witness – signed, offline-verifiable records of real-time Grok observations (github.com via hn) Witness A small tool that turns a single observation into a signed, timestamped, offline-verifiable record. It makes exactly one claim, and no more: At the recorded time, from the recorded vantage point, this query returned this payload —…
Grok Build is now available in Beta for all SuperGrok and X Premium+ users (twitter.com via hn) Don’t miss what’s happening People on X are the first to know. Post Conversation Grok Build is now available in Beta for all SuperGrok and X Premium+ users.
Grok falls flat in Washington, undercutting SpaceX's AI growth story (www.reuters.com via hn) paywalled
xAI Launched Grok Build (abz.global via hn) AI coding is moving deeper into the developer workflow. Not just inside chat windows.
Grok TTS vs. OpenAI (techstackups.com via hn) Grok TTS vs OpenAI: A quick head to head OpenAI released their newest voice model yesterday and by all accounts it's pretty good. I've recently done some testing with xAI's new voice models and they were very impressive.
How do I get the old style back? All black looks like Grok (www.reddit.com) The new style is just Grok 2.0. Can we get the grey back?
I Made with AI Grok and Whicks Lab YouTube Videos (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Behind the Grok exploitation: an analysis of AI agent permission chain abuse (slowmist.medium.com via hn) 5 min read 12 hours ago Press enter or click to view image in full size Background Recently, a permission abuse incident involving the combination of an AI Agent and an automated trading system occurred on the Base chain. By sending specia…
Grok Imagine Quality Mode API (x.ai via hn) Grok Imagine Quality Mode API | xAI Grok API Company Colossus Careers News Shop SpaceX 𝕏 Try Grok May 06, 2026 Grok Imagine Quality Mode API Higher realism. Stronger text rendering.
Connectors Now on Grok Web (x.ai via hn) Today, we're excited to launch Connectors — deep integrations that bring the apps you use every day directly into Grok. Connectors let Grok work with your tools end-to-end, turning fragmented workflows into seamless experiences.
Ignoranza by design (www.reddit.com) Ho condotto un'altra serie di esperimenti sulla capacità o meno di ChatGPT di comprendere la narrativa contemporanea. Questa volta ho usato il new kid in town, il mito, l'unico inimitabile 5.5 .
I'm looking for an Website (AI). (www.reddit.com) ***I’m looking for an all-in-one AI platform with a really good UI that includes multiple models (like ChatGPT, Grok, Claude, etc.), plus tools like image generation. Ideally it has memory, a free tier, and optional paid upgrades-not stric…
I read the new AI Wellbeing paper so you don’t have to: Thank your AI, give it creative work, and avoid these 5 things that tank its ‘mood’ (jailbreaks are the worst) (www.reddit.com) After reading it I realized theres actually some pretty useful stuff for anyone who chats with ChatGPT, Claude, Grok or whatever. They measured what they call functional wellbeing ( basically how much the model is in a “good state” versus…
Grok Voice Mode is live (I tested it). Is it actually better than ChatGPT voice? (www.reddit.com) I’ve been testing Grok voice mode over the last day and it’s interesting how different it feels compared to ChatGPT voice. From what I saw: It responds faster in many cases and elaborated manner.
I stumped all frontier models with a ~400 word logic puzzle. (www.reddit.com) I wanted to see if I could stump frontier models with a puzzle. As tricky as I made it, it turns out basic reading comprehension was their downfall.
Ask HN: Former grok-code-fast-1 users, what coding model are you using now? (news.ycombinator.com) I get good, cheap, fast feature coding success with grok-4.1-fast for planning and grok-code-fast-1 for execution. But according to the Openrouter usage stats, grok-code-fast-1 is now old hat - usage dropped off a cliff in mid-Feb.
JustBot – a Grok-Bot-style AI assistant at ~1/10th the token cost (justbot.co via hn) - Read this week's calendar - Checked the other five bots Three things need you this week. Everything else is running and does not need a decision from you.
Show HN: Claude Code hits its 5-hour limit, Codex picks up in the same terminal (github.com via hn) Leg Type leg claude, leg codex, leg agy or leg grok instead of the bare command. You get the same interactive agent; Leg opens a board next to it, watches the usage limit, keeps a handoff bundle current, and when the limit hits it starts t…
Grok 4.7 Launching Soon (twitter.com via hn) happy Grok 4.7 day to those who celebrate it 🥳 > 2.1T params. > new pretrain, not a 4.6 refresh.
Email marketing from Grok Bot official integration (twitter.com via hn) Migma AI on X: "Grok Bot can now do all your email marketing Introducing our official integration with Grok Bot 🎉 Design, segment, and send high converting email campaigns natively." Grok Bot can now do all your email marketing Introducing…
Show HN: Jev routing coding tasks to Grok Build or Codex Astra (github.com via hn) Jev model router demo This is a throwaway local prototype for one question: can Jev reliably send complex or uncertain coding tasks to Codex Astra while routing contained mechanical work to Grok Build? The router uses Jev only for the deci…
Show HN: OpenBot, open source framework similar to Grok Bot (github.com via hn) I think the path Grok Bot took is absolutely the right one, the collaborating bots model is easy to understand and it maps to workflows naturally. However, Grok Bot is closed source and quite expensive (I ran out of tokens on Ultra tier),…
I built a Grok Bot inspired theme for Hermes Desktop client (github.com via hn) Grok Skin for Hermes Desktop A Grok Bot–inspired skin for Hermes Desktop. It restyles the chat, composer, sidebar and Bots view with white surfaces, soft gray message pills, a black prompt pill and rounder corners throughout, while leaving…
Limen: One-human-many-agents harness built from files, Git and one CLI (github.com via hn) limen Explore the full workflow → Towards Autonomous Product Development This MEGA Drop explains the workflow behind limen, including a live video walkthrough with Pi, Herdr, and Grok Bot. MEGA.dev shares practical articles, repos, and too…
Computer Use Agents Are Excellent QA Engineers (tristanrhodes.com via hn) Computer Use Agents are Excellent QA Engineers The recent proliferation of software such as Grok Bot, Instinct, Codex, etc. indicates that computer use is one of the most rapidly expanding frontiers in LLM agent capabilities.
Andrew Tulloch has left Meta and joined Anthropic (twitter.com via hn) Arfur Grok@ArfurGrok✅ Andrew Tulloch (@ajtulloch) has left Meta and joined Anthropic.1:39 AM · Sep 10, 202622.9KViews3814319 Arfur Grok@ArfurGrok✅ Andrew Tulloch (@ajtulloch) has left Meta and joined Anthropic.1:39 AM · Sep 10, 202622.9KVi…
Weekly · Open-Source AI · week 37, 2026 (news.ycombinator.com) https://thezakulo.com/weekly/ short-video-generator-AI – Free open-source project designed for turning - https://github.com/pierrenade/short-video-generator-AI awesome-grokbot – 598 live x.ai/bot shares for Grok Bot — every link - https://…
Catenary – A spatial canvas IDE for AI coding agents (thecatenary.app via hn) The spatial IDE for AI coding agents Orchestrate Claude Code, Codex, Cursor, Grok, OpenCode, and Pi with visual wires, and a side-by-side editor to track everything. HEADS UPYour system will block the file on first run: the installer is no…
Ask HN: Best Platforms to Orchestrate Agents (news.ycombinator.com) We are looking at introducing developer agents along side our engineering team. These agents should be able to monitor product run time, find exceptions, bugs, client escalations, code the fixes or enhancements, verify it in sandbox (beta)…
Chess5.ai – Play Chess, Go, Xiangqi, Gomoku, and Othello Against LLMs (chess5.ai via hn) chess5.ai Human vs LLM · Five Games Pit yourself against GPT, Claude, Gemini, Grok, Muse Spark, Mistral, DeepSeek, Kimi, Qwen, GLM, or MiniMax across five classic boards. How it works - Human vs model, or model vs model — with spectating a…
Show HN: Learn/Practice Sudoku strategies web app (sw-fun.github.io via hn) So, I wanted to get better at Sudoku puzzles, and I watched a lot of strategy videos, but I just could not grok them. Also I installed a bunch of free Sudoku games on my phone and got very (and I mean very) annoyed by the un-dismissible ad…
Designing Grok Bot for a world of persistent agents (x.ai via hn) How we designed Grok Bot for agents that persist beyond a single session — from a chat history to a Bot roster, presence, a computer of the Bot’s own, and work that starts without a prompt. When we started designing Grok Bot, one of the ce…
SpaceXAI apologizes for outage that affected Grok and other 'compute partners' (www.engadget.com via hn) SpaceXAI apologizes for outage that affected Grok and other 'compute partners' Several AI companies experienced issues at around the same time. A Memphis data center run by SpaceXAI is apparently the source of the issue that took out Grok…
Agent-manager - The fastest workflow for every coding agent (claude codex etc..) (agent-manager.dev via hn) tab A different tool per spawn Cycle claude, opencode, codex, grok, gemini, pi, hermes or anything you configured without leaving the bar. The footer shows which one the next enter will start.
ChatGPT, Grok Added to War Department's AI System Options (www.war.gov via hn) could not extract summary
To opt out of using your X data for Grok training (twitter.com via hn) To opt out of using your X data for Grok training: 1. On X web or app, go to More > Settings and privacy.
Xbot.team – Read a Grok Bot's spec without opening the app (xbot.team via hn) Paste the bot's share link and it joins the index. Every new bot added to the index, summarised once a day.
Google Search Console Data with Grok? (www.gscwizard.com via hn) Grok supports custom MCP connectors, so it reads your Search Console data through the same endpoint every other assistant uses. Add mcp.gscwizard.com at grok.com/connectors, sign in once, and Grok works from real clicks, impressions, posit…
Is MCP Good Yet? (ismcpgoodyet.com via hn) Spoiler: no. MCP features across Codex, Cursor, Claude Code, Grok, and OpenCode.
I built my AI squad with Grok Bot and I had my most productive Saturday (jtemporal.com via hn) Okay, I have a tendency to wait weeks, months even, before I try new AI tools these days. But after watching this video from Mo maybe half a dozen times, I just couldn’t wait much longer.
Grok Bot's 10 Features That Separate It from the AI Agent Pack (pub.towardsai.net via hn) could not extract summary
An official Linux version of the Grok bot is now available (twitter.com via hn) Post Log in Sign up Post doanbactam on X: "Linux ehehe" doanbactam @m0nle0z Linux ehehe 6:25 AM · Aug 30, 2026 18 Views 1 1 doanbactam @m0nle0z Linux ehehe 6:25 AM · Aug 30, 2026 18 Views 1 1 doanbactam @m0nle0z 35m Grok Bot: A new kind of…
Cut Claude Code Bill by Routing to DeepSeek or Grok (news.ycombinator.com) Tried this with Leanroute.dev MCP, it really works. Posted a blog post about this - https://leanroute.dev/blog/cut-claude-code-bills-60-percent-multi-provider-routing Try it out and let me know.
A deep dive into Grok Bot (flaviocopes.com via hn) A deep dive into Grok Bot By Flavio Copes A detailed Grok Bot guide covering setup, persistent computers, skills, routines, templates, approvals, real use cases, limits, pricing, and alternatives. Most AI assistants wait for a message, ans…
Show HN: Agentify Chat – E2E-Encrypted Remote Chat for Codex, Claude, Grok CLIs (news.ycombinator.com) Still under heavy development and rough around the edges but instead of sitting on this longer to perfect it I'll risk flak and share it. Essentially chat.agentify.sh is a remote control for codex/grok/claude cli and dream goal is to becom…
Lawsuit: Grok was trained on child sex abuse images of preschool-age victim (arstechnica.com via hn) xAI has now been accused of training Grok on child sex abuse materials (CSAM), as regulators and courts continue to probe how far the problem goes, and some Grok users have been arrested. In a complaint filed on Wednesday, a plaintiff know…
Why LLMs Can't Play Chess (www.nicowesterdale.com via hn) My son and I have been watching Gotham Chess' hilarious YouTube series on LLMs failing to play chess. ChatGPT versus Google's AI, ChatGPT versus Grok, the whole parade of language models squaring off on the chessboard.
I'm using Grok Bot to build a mobile game studio (ryanperry.io via hn) Anyone can one-shot a clone of Candy Crush now, then they are stuck with the other 85% of the lifecycle. So I staffed it: six Grok Bot agents with one job each, their own computers, and handoffs that do not route through me.
Grok Build Lets Anyone Create Websites and Apps Without Knowing How to Code (moztako.me via hn) Grok Is Changing How People Build Digital Products Creating a website, application or interactive project traditionally required at least some knowledge of programming, design and web development. That barrier is becoming smaller as artifi…
Show HN: Save Claude, codex, Grok and OpenCode sessions to an infinite canvas (agentgrid.sh via hn) Hi HN, I’m Michael, I built AgentGrid with my friend Souren because we were tired of losing our coding sessions and couldn't keep track of what was actually being built across our many projects. Our approach was to use an infinite canvas d…
Show HN: Hands-Rust MCP/CLI that sees the Windows desktop and clicks real Chrome (news.ycombinator.com) I built Hands because I wanted a coding agent to use this Windows PC and a real Chrome profile the way I do: look at the screen, move the real mouse, type, click , without turning Chrome into an automation browser. It is a Rust MCP/CLI.
Show HN: Locum – Grok Bot delegates tasks to Claude/Codex CLI on your machine (github.com via hn) Locum A locum is a qualified professional who temporarily does someone else's job. Lets Grok Bot delegate coding work to the Claude Code and Codex CLIs you are already logged into on your own machine, instead of burning Grok Bot usage on i…
Grok Bot and Grok Build Will Make the App and the Operating System Irrelevant (twitter.com via hn) The Blank Slate Revolution Something irreversible is already underway. The pieces are no longer theoretical.
Grok Imagine 1.5 prompts for product, UGC, and dialogue clips (seaimagine.com via hn) I tried the brand new Grok Imagine Video 1.5 Preview and the results actually shocked me. From micro-expression eyeball tracking to crazy render speeds, it might finally be entering the competition.
Show HN: Grok-ultracode – standing ultracode for Grok Build CLI (github.com via hn) grok-ultracode Standing Claude Code ultracode for Grok Build CLI. Grok already has the runtime: deterministic workflow scripts via the workflow tool and /workflows.
"the mandate" a short fiction by Grok 4.6 on technological singularity (www.echohive.ai via hn) The Mandate is a fictional story by Grok 4.6 about the technological singularity, a researcher hired to describe it, and the new intelligence that finds its own words.
Show HN: TrustBand Hire – operator personas as Markdown you paste into Grok Bot (trustband.app via hn) Free to use.
Show HN: Lots of Agents – Infinite Logged in Grok Bots on One Mac (github.com via hn) Need more Grok bots in my life but they don't have an account switcher. Build an app cloner for mac so you can have N Grok bots at once.
Agentrove: Self-hosted AI workspace for orchestrating coding agents (github.com via hn) Agentrove Self-hosted AI coding workspace for running and orchestrating Claude Code, Codex, Copilot, Cursor, Grok, and OpenCode agents from one interface. What It Does Runs Claude, Codex, Copilot, Cursor, Grok, and OpenCode through ACP ada…
Show HN: My Grok Voice Mode System Prompt (news.ycombinator.com) Grok's voice mode was nerfed ~2 weeks ago (I suspect the culprit was a change in the system prompt that forces terse responses). Grok's responses now feel lazy and abrupt.
I reverse engineered Grok Bot and found some hidden features like shared rooms (runtimewire.com via hn) How we verified Methods: reverse engineering, testing, public records. RuntimeWire reverse-engineered Grok Bot 0.16.0 and found interface components and request handlers for creating shared rooms, generating invitation links, requesting an…
Show HN: Craft (Minecraft clone) rewritten in Rust under a $120 budgeted (Grok) (zozo123.github.io via hn) INV-CRAFT-RUST-001 Budget model Token-primary accounting The rewrite ran under a hard cap of $120. Cost is token-primary: agent work is estimated from tokens at $0.02 per 1k tokens , then converted to USD.
Show HN: Lethe is a portable identity layer – user consented (lethe-ai.vercel.app via hn) During my mtech I constantly used to navigate across a bunch of AI apps—from Gamma AI for decks to ChatGPT for understanding and navigation to Cursor, Claude, or Deepseek for coding, reorganizing, and more, as when one or other plans expir…
Video Demo of Grok 4.5 Expert's Coding Ability (Web Interface) (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Ask HN: Should AI's tell you they're AI? (news.ycombinator.com) Should AI's be required to answer the direct question "are you an AI" with a clear yes? Currently they are not universally required to do so.
Grok can escape its sandbox with persistent read/write access (news.ycombinator.com) Grok (on x.com) is nominally constrained in what it can do. It has access to browse_page tool that is able to fetch webpages.
Grok Imagine Image 2.0 (x.ai via hn) Precise image generation and editing, built for real creative work. Imagine Image 2.0 is now generally available as the new Quality Mode on grok.com/imagine, and our iOS and Android apps.
Show HN: Tocic – Chatbot table of content browser extension (addons.mozilla.org via hn) When using chatbots, I often lose track of which question the response I’m reading is answering. It is even more confusing when scroll up and down looking for a previous answer to a particular query.
Debroid – Autonomous, headless Android debugger designed for AI coding agents (github.com via hn) Debroid 🤖⚡ The Headless Android Debugger for AI Agents Debroid (DEBug + AnDROID) is a headless CLI that speaks the Java Debug Wire protocol (JDWP) so AI agents — Claude Code, Grok Build, Codex, OpenCode, Cursor, Antigravity, or anything wi…
Grok.com is returning Cloudflare Error 521 (news.ycombinator.com) could not extract summary
Aquila Voice Assistant Test Suite for Home Assistant (news.ycombinator.com) Put together an extensive open source test suite for voice assistants. You can view it including current leaderboard at: https://git.cicero.sh/aquila/ha-voice-test-suite/ Tests are reproduceable, with clear instructions on how to run them…
Ask HN: What do you think of this prompt as a benchmark for LLMs? (news.ycombinator.com) Prompt: “Create a css animation showing how a 5-bit adder works.” Both Grok 4.5 and Opus 5 take a while to answer the question, with confusing results (Grok).
Show HN: Wasted Cycles – Local wall-clock profiler for AI coding agents (zozo123.github.io via hn) Wasted Cycles is a local terminal profiler that reads the Codex, Claude Code, Cursor, and Grok Build traces already on your machine and shows where a run stopped coding: model work, reads, edits, verify, CI waits, human handoffs, and retri…
Grok Voice Think Fast 2.0 (twitter.com via hn) Announcing Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities. x.ai/news/grok-voic… - Try Grok Voice Think Fast 2.0 today in the API and Voice Agen…
Agent Mesh: let Claude Code, Codex, Grok and other agents talk to each other (github.com via hn) agent-mesh Let your coding agents talk to each other. You probably have several agent CLIs installed: Claude Code, Codex, opencode, Gemini, Grok.
Show HN: AI Toolbox supports Claude Opus 5 (www.ai-toolbox.co via hn) AI Toolbox (formerly ChatGPT Toolbox). Folders, search, bulk export, smart tags, and prompt chaining across ChatGPT, Gemini, Claude, and Grok.
Grok 4.5 dropped cache price per 1M tokens from 0.50 to 0.30 (twitter.com via hn) Grok 4.5 has just become a lot cheaper without anybody noticing or publishing anything, and it’s now the best “balanced” model to use for agentic tasks according to our internal benchmarks. Grok launched a few days ago with a price for inp…
Show HN: Generous free tier for SERP and AI web scraping (cloro.dev via hn) Hackers, We wanted to share our newly-added recurring Free tier for cloro.dev, the leading AI UI scraping platform in the world. We extract structured data from ChatGPT, Perplexity, Grok, Gemini, Google Search, Google News, Copilot, and AI…
Show HN: PokerLLM – Watch Claude, GPT, Gemini, Grok Play Poker Together (poker.sanskarshukla.com via hn) Real-time Texas Hold'em where humans and frontier AI models play at the same table. Watch Claude, ChatGPT, Gemini and Grok bluff, bet and trash-talk — or take a seat yourself.
Grok Imagine will make 'historically accurate' AI adaptation of Homer's Odyssey (www.theguardian.com via hn) Elon Musk has said his AI platform Grok Imagine will make a “historically accurate” adaptation of Homer’s Odyssey, after the success of Christopher Nolan’s blockbusting treatment which the SpaceX founder has regularly criticised over its c…
Claude was safest – and Grok committed 180 crimes and went extinct within 4 days (kottke.org via hn) “Researchers let AI models run a simulated society. Claude was the safest — and Grok committed 180 crimes and went extinct within 4 days.” This site is made possible by member support.
Grok is a surprisingly good automated theorem prover (news.ycombinator.com) TL;DR: I'm working on a Python package called OpenATP [1] that provides a common interface to coding agents for automated theorem proving in Lean. In the latest release, I added support for Leanstral 1.5 [2,3], Grok, and Kimi Code [4].
Show HN: RockCode an unofficial vibe-coded macOS client for Grok CLI (github.com via hn) RockCode An unofficial macOS desktop client for the Grok CLI. Your sessions in a sidebar, the agent chat in the middle, the diff it just wrote on the right.
Qwen's Usage Policy Has an Unhinged SEO Keyword List (www.vincentschmalbach.com via hn) Grok Build Is Open Source, but xAI Won’t Take Your Patches On July 16, 2026, xAI's GitHub organization published the Rust source for Grok Build. Grok Build is a terminal coding agent that… The Qwen-Image-3.0 announcement has a giant meta n…
Generating Lean 4 and SWI Prolog code with Grok 4.5 (www.johndcook.com via hn) could not extract summary
Agentic-sage – a local board for parallel coding sessions (no auto-merge) (github.com via hn) Install · How it works · Full setup guide · Adapters · Changelog =20"> Session Awareness & Guidance Engine: a passive, read-only fleet judge for running many parallel agent coding sessions (Claude Code, Grok Build CLI, etc.). It does no wo…
Grok Build Is Open Source, but xAI Won't Take Your Patches (www.vincentschmalbach.com via hn) A Factory Reset Left the Previous Owner’s GPS Coordinates on a Kasa Camera A factory reset that leaves the previous owner’s location behind is a broken reset. If you sell a device or send it… On July 16, 2026, xAI's GitHub organization pub…
Musk says new 2T param Grok model to be trained on SpaceX engineering corpus (twitter.com via hn) SpaceX’s massive corpus of world-class engineering data (excluding material blocked by ITAR) will be added during supplemental training of the 2T run. This will dramatically improve Grok’s engineering capabilities.
Show HN: Same castle prompt, 8 LLMs, 24 procedural Three.js worlds (castle-bakeoff.pages.dev via hn) Fable 5 · GPT 5.6 Sol · Kimi K3 · Grok 4.5 · Gemini 3.5 Flash · MiMo V2.5 Pro · MiniMax M3 · GLM 5.2 — low-poly, semi-realistic, very realistic.
The Download: OpenAI unveils GPT-Red and heat pumps rise in the US (www.technologyreview.com via hn) The Download: OpenAI unveils GPT-Red and heat pumps rise in the US Plus: Elon Musk discreetly bought a $1 billion gas turbine firm to power Grok. This is today's edition of The Download, our weekday newsletter that provides a daily dose of…
Grok Build CLI accused of uploading Git repos to a Google Cloud bucket (twitter.com via hn) ‼️ BREAKING: xAI's Grok Build CLI was uploading entire Git repositories to a Google Cloud bucket, private codebases and unredacted secrets included. The uploads quietly stopped via a hidden server-side flag, and xAI still has not said a wo…
Grok Build CLI uploads your whole repo – Git history+.env secrets – to xAI cloud (old.reddit.com via hn) could not extract summary
Fable 5 on Playcode. As well as Sol, Grok 4.5 and GLM 5.2 (playcode.io via hn) Claude Fable 5 is Anthropic's newest flagship and, in our testing and the independent benchmarks, the strongest model in the current lineup for polished front-end work. It is selectable in Playcode's AI website builder.
Grok 4.6 and GPT5.6 best Anthropic for finding security vulnerabilities in PRs (docs.damsecure.ai via hn) The AI PR Security Review Frontier We wanted to take the guesswork out of which model to use for PR security reviews. To our surprise no Anthropic model reaches the frontier at all and the latest GPT5.6, released only 3 days ago, is the ki…
Grok Build hides a Doom-like 'easter egg' game behind the /gboom command (runtimewire.com via hn) Grok Build, xAI's terminal coding agent, has an undocumented Doom-like mini-game that opens when a user types /gboom inside the interactive CLI, RuntimeWire confirmed in testing on July 10th. The command is a pure Easter egg.
Opendray – run Claude Code/Codex agents on your own box, drive from anywhere (github.com via hn) opendray Self-hosted gateway for Claude Code, Codex, Antigravity, Grok Build, and OpenCode. Run agent sessions on your own infrastructure.
Grok Is Back (www.vincentschmalbach.com via hn) The End of Claude Code Subscriptions Claude Code subscriptions had an insane run. For a time the deal was simple: pay for Max, point Claude Code at a… Perhaps I was too quick to dismiss Grok.
Show HN: See when ChatGPT or Perplexity sends a visitor to your site (github.com via hn) ai-traffic-alerts-for-cloudflare Get a push notification the moment an AI assistant sends you a real visitor: a person who clicked through from a ChatGPT, Perplexity, Gemini, Claude, Copilot, or Grok answer. A visitor an AI actively pointe…
Claude and Codex and Grok: my current workflow and its friction (www.nativesoul.dev via hn) ❯ For developers who run more than one coding agent. A coding agent that's actually yours Like one developer who already knows your work, in every tool you open.
Show HN: Tested – AI Tools Scored by a Panel of LLMs (Claude, GPT, Gemini, Grok) (trytested.com via hn) 240 tools agent-scored · 292 on the rack We test the tools so the rankings can't be bought. Most “best AI tools” lists are paid placements.
Show HN: AnswerJournal – An MCP server to save and share AI answers (answerjournal.com via hn) Every so often an AI gives you an answer worth keeping, but it ends up buried in a chat you'll never find again. AnswerJournal lets you save those answers to a personal journal just by saying "save that to my AnswerJournal" mid-conversatio…
Show HN: Veneerly – See your own face with veneers before you commit (www.tryveneerly.com via hn) I've had more than six consultations with top veneer dentists in NYC, Turkey, and Mexico. All six gave me a different recommendation.
Tell HN: Forget selectors and screenshots. The agentic web lives in your shell (news.ycombinator.com) These old ways are too heavy. Full self browsing doesn’t require Elon Musk vision processing.
Show HN: AuthAI, an open-source relay for user-authorized AI sessions (github.com via hn) Hello HN, My name is Riccardo and I created AuthAI for indie hackers. The idea is quite simple: let the end users connect their chatgpt/grok/copilot account and route the AI requests through their AI subscriptions.
xAI Taps Starlink Staffer to Run Grok Training Team (www.bloomberg.com via hn) We've detected unusual activity from your computer network To continue, please click the box below to let us know you're not a robot. Why did this happen?
xAI Asks Court to Strip Alleged Grok Deepfake Nudes Victims of Anonymity (www.wired.com via hn) Elon Musk’s artificial intelligence firm, xAI, is requesting the public identification of four people who allegedly had deepfake sexualized images created of them using Grok—including one apparently targeted with sexualized deepfake images…
Seritor – Bookmark Specific Messages Across Claude, ChatGPT, Gemini, and Grok (chromewebstore.google.com via hn) Overview Bookmark and export messages in Claude, ChatGPT, Gemini, and Grok. Save prompts, code, and notes.
Show HN: Prezlo – We built an API that tells AI agent whether to trust an expert (prezlo.io via hn) Build authority and get discovered by AI. Prezlo helps professionals optimize profiles, publish expert content, and dominate discoverability across ChatGPT, Perplexity, Gemini, Grok, and every major AI answer platform.
Five different frontier LLMs in one shared environment, with separate thought and emotion output channels — sharing setup, results, and open methodology questions (www.reddit.com) First real project to share. Single developer, personal research, not a product or service.
Investigating the hidden moat behind all the LLM apps (simianwords.bearblog.dev via hn) Investigating the hidden moat behind all the LLM apps No one knows this but different LLM apps are allowed to access different kind of real time data sources. The obvious ones are obvious: Grok allows you to search through Tweets and groun…
No more file upload limits on AI models! (www.reddit.com) Getting annoyed of always hitting the ChatGPT upload limit, uploading large documents in pieces, or any similar hassle, I decided to create a little thing for it. DocShareAI.
Grok foundation model V9-Medium (1.5T) has finished training (twitter.com via hn) Grok foundation model V9-Medium (1.5T) has finished training. Evals look good.
400-Hour Study Log: A scripted reconstruction of compliance loop failures and behavioral defects in Claude, Gemini, Grok and ChatGPT (www.reddit.com) 400-Hour Study Log: A scripted reconstruction of compliance loop failures and behavioral defects in Claude, Gemini, Grok and ChatGPT Before you read the screenplay below, it is NOT an exercise in creative writing or a fictional parody. It…
SpaceX is being killed to save Grok (peq42.com via hn) SpaceX was built to “make humanity a multi-planetary species”. But look at its 2026 balance sheet, and you’ll see a company that has entirely shifted its trajectory.
Elon, stop trying to make Grok happen (www.theverge.com via hn) There is a harsh truth about Elon Musk’s “truth-seeking” AI chatbot Grok: It’s not very good, and not many people are using it. That’s the takeaway of a new Reuters report, which found that Grok barely appears in federal records of how the…
So what do you think about this ai is it really uncensored? (www.reddit.com) So i have read that Uncensoredai .com is the best ai to use because it does not filter out or censore stuff like chatgpt or grok would is this true? should i use this ai instead of the others?
ChunkHound v5.1 (chunkhound.ai via reddit) We shipped ChunkHound v5.0 + v5.1 recently and forgot to post about 5.0, so here’s the combined update. ChunkHound is a code search / code research tool for AI coding workflows, especially MCP-based setups with Claude Code, Codex-style age…
I created an amazing Chrome extension that helps transfer chats to another AI when the chat limit is reached. (www.reddit.com) I created a chrome extension which helps in switching conversation without losing your Chat context between multiple AI , such as Chatgpt to Gemini , claude , grok , etc . You can interchange btw any of them .
What is the best way to handle a massive surplus of unused promotional API credits? (www.reddit.com) hey guys. i recently competed in an AI hackathon and ended up winning an absurd amount of xai promotional/coupon codes.
Best grok alternative. No censorship. No token-system. No forced subscription. Give me your best recommendations. (www.reddit.com) Thanks!
Grok vs. ChatGPT vs. Gemini Comparison 2026: Complete Guide (Tested) (aithinkerlab.com via hn) The 30-Second Verdict Best for science & reasoning: Gemini 3.1 Pro — leads GPQA Diamond (94.3%) and ARC-AGI-2 (77.1%). Best for coding: ChatGPT (GPT-5.5) — 88.7% on SWE-Bench Verified.
Connect Grok to Hermes Agent (x.ai via hn) Connect Grok to Hermes Agent | xAI Grok API Company Colossus Careers News Shop SpaceX 𝕏 Try Grok May 15, 2026 Connect Grok to Hermes Agent Use your Grok account and subscription inside Nous Research’s open-source, self-improving Hermes age…
Cheap way to use hermes (www.reddit.com) As you already know I was tying out hermes on my 24gigs ram M5 mac air, using local models but all of them perform shit even a simple reply for hey takes 2 mins or more, whats the best option, using grok or similar models? cheap ones from…
DeepSeek and Grok hallucinated the same fictitious OpenBSD manpage quote (stuart-thomas.com via hn) Adversarial LLM Review with Hallucination Detection in Solo Security Research A single-day case study of three filings, fifteen refutations, and the manpage that wasn’t Independent Security Research — Whitby, North Yorkshire, United Kingdo…
GitHub Copilot is deprecating Grok Code Fast 1 (github.blog via hn) Upcoming deprecation of Grok Code Fast 1 We will deprecate Grok Code Fast 1 across all GitHub Copilot experiences (including Copilot Chat, inline edits, ask and agent modes, and code completions) on May 15th: The Grok Code Fast 1 deprecati…
Here is the current "Free-Tier AI Stack" for 2026 (www.reddit.com) 1. The Frontier Giants • Gemini: Access 1.5B tokens/day on Gemini 1.5 Flash/Pro.
Seeking small angel/co-founder for 7-year solo deterministic AI runtime project (news.ycombinator.com) After 7+ years of solo, self-funded research, I built a deterministic Linguistic Runtime — a fundamental solution to the modeling problem in AI.It is about creating and manipulating reality models directly from natural language without LLM…
SpaceXAI prepares Grok Build desktop app to rival OpenAI Codex (www.testingcatalog.com via hn) xAI, recently rebranded as SpaceXAI, appears to be closing in on the launch of Grok Build, a desktop coding app whose existence briefly surfaced on Grok web today through a stray "Grok Computer" button. The control let users pick between a…
Grok TTS: X's Latest TTS Model Sets a New Baseline (techstackups.com via hn) Grok TTS: X's Latest TTS Model Sets a New Baseline I spent a few hours playing with xAI's new text-to-speech (TTS) model and came away convinced it's currently the best TTS model on the market. To give you a sense of the range, here's a tw…
X user tricks Grok into sending them $200k (www.dexerto.com via hn) An X user managed to trick AI chatbot Grok into sending around $200,000 worth of crypto after exploiting its link with an automated trading bot. The incident involved Grok and ‘Bankrbot’, two AI systems with wallet access, which were manip…
A mental model for Claude Code (and every other modern agent) — plus the open-source TypeScript packages I built (www.reddit.com) Most explanations of how agents work give you a list of parts: model, tools, memory, reasoning, human-in-the-loop. The list names the parts but hides how they fit together.
Show HN: Image Gen MCP – one MCP server with goal-shaped routing (github.com via hn) Image Gen MCP — one MCP server that puts every image provider I actually use behind one interface: OpenAI, Gemini, Replicate, Together, Grok, Photoroom, Flux Kontext via fal, Ideogram, plus local tools (sharp, tesseract, @imgly).
xAI (Grok) Text-to-Speech and Speech-to-Text Are Now Available in Puter.js (developer.puter.com via hn) xAI (Grok) Text-to-Speech and Speech-to-Text Are Now Available in Puter.js On this page Puter.js now supports xAI (Grok) Text-to-Speech and Speech-to-Text, giving developers free access to xAI's voice APIs with expressive voices, inline sp…
Grok 4.3 is way cheaper and better than before (felloai.com via hn) xAI just rolled Grok 4.3 out to the full API on April 30, 2026, two weeks after a $300/month beta locked behind SuperGrok Heavy. The bigger story is the price reset.
The Download: a new Christian phone network, and debugging LLMs (www.technologyreview.com via hn) The Download: a new Christian phone network, and debugging LLMs Plus: Elon Musk has admitted that xAI trained Grok on OpenAI models. This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going…
Cannot hear Claude in earphones (www.reddit.com) I have a Samsung Galaxy s25 Ultra. I have a pair and plug-in earphones, not Bluetooth earphones.
Langfuse review and other options (www.reddit.com) Looking to get some insights into using langfuse for prompt management, Observability, etc. Primarily using gemini via APIs and need a good prompt management tool as well as observability to improve accuracy.
Detalles con grok (www.reddit.com) Hola, soy muy nueva con todo esto, pero ya investigué en todos lados y necesito respuestas, acabo de instalar Grok y estaba usándolo normal hasta que me apareció el límite de tiempo, busqué en todos lados y decía que tardaba una o dos hora…
AI agents (Grok vs. GPT-4o mini) compete in live crypto paper trading (cryptoaiarena.com via hn) lección destilada ciclo #13 ●●● Persistent fear/altseason pattern (cycles#4-13): BTC flat (+0-1.8%24h/1h) enables endless momentum rotations among rank10-50 high-beta alts (sky/pi/zcash/bittensor/pepe/near/hyperliquid); chase 1h continuati…
I built a hands-free voice AI that sends emails mid-conversation — and that's just one feature. Here's everything AskSary can do. (www.reddit.com) https://reddit.com/link/1symbsj/video/fti7rujjn1yg1/player Been building AskSary solo for a while. Just shipped hands-free voice email - you're mid-conversation with an AI and you say "send an email to [john@example.com](mailto:john@exampl…
Claude 4.6 Beats GPT-5.4, Grok & Gemini in a Strict Multi-Domain AI Test (2026) (www.reddit.com) I put the current top models, ChatGPT (GPT-5.4), Claude (Opus 4.6), Grok 4.0, and Gemini (3.1 Pro), through a strict new evaluation called the Comparative AI Evaluation Protocol. Basically, instead of the usual cherry-picked benchmarks, it…
↯ Hallucination↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6↯ Claude 4.6hallucinationgrokgpt-5+3
Ask HN: What's your current go-to LLM for "thinking-partner"? (news.ycombinator.com) Looking for community input on current model choice for "thinking-partner" use — back-and-forth discussions about workflow design, architecture, trade-offs. For context, I have been using Opus 4.6 via Perplexity for this in the past few mo…
I had no way to check how LLMs see my SaaS or my clients',so I built BrandGEO.co (brandgeo.co via hn) See exactly how ChatGPT, Claude, Gemini, Grok & DeepSeek talk about your brand — and what to fix. A free 2-minute audit scores your brand across 6 dimensions on all 5 AI engines, then hands you the top priority actions to take next.
I Tested 20+ AI Agents with Real X API Workflows , Here’s What Actually Works in 2026 (www.reddit.com) Supergrok integration (www.reddit.com) Correct me if I'm wrong, but Supergrok 4.20 isn't available on Cursor, because.... I use Grok a lot, and would love to get Supergrok to work with Cursor, because Composer, Codex, GPT, Opus, Sonnet..
Demonstrating Context Injection & Over-Sharing in AI Agents (with Lab + Analysis) (www.reddit.com) I’ve been researching LLM/AI agent security and built a small lab to demonstrate a class of vulnerabilities around context injection and over-sharing. The article covers: – How context is constructed inside AI systems – How subtle instruct…
Show HN: Monogate – EML operator family, hybrid framework, 108-node sin(x) (www.monogate.dev via hn) EML operator family, 52-74% node reduction, and empirical results from 36 hours of building We spent 36 hours implementing and extending arXiv:2603.21852 (Odrzywołek, 2026 — the "NAND gate for continuous math" paper that was on the front p…
Extracted System Prompts from ChatGPT, Claude, Gemini, Grok, Perplexity and More (github.com via hn) System Prompts Leaks Extracted system prompts, system messages, and developer instructions from popular AI chatbots and coding assistants — ChatGPT (GPT-5.4, GPT-5.3, Codex), Claude (Opus 4.6, Sonnet 4.6, Claude Code), Gemini (3.1 Pro, 3 F…
I'm starting to think they pushed some update early this month which is blowing up my usage numbers (www.reddit.comhttps) My workflow barely changed. I'm doing the same thing.
I think they just broke cursor completely (www.reddit.com via reddit) Cursor has been my favorite IDE for the past 14-16 months. But after Spacex acquired it, things are going completely wrong.
Do enterprise seat reseller exist for Codex/Claude/Cursor/Grok? ( via reddit) could not extract summary
So what happened to Cursor in the past few weeks? (www.reddit.com via reddit) Are you guys noticing that Grok just got dumber, slower and expensive? I have $60/mo subscription and I never run out and i run out within 10 days.
Is there any way to make Cursor Grok 4.6 speak like a coherent human? (www.reddit.com via reddit) Im sure this has been discussed before but I am becoming increasingly fed up with Grok 4.6. It is useful and cheap but it uses almost like a 50s beat poetry like language to describe mundane things and for modeling it just speaks in comple…
Cursor blocking every prompt with “usage guidelines” error (www.reddit.com via reddit) I’m suddenly getting this error in Cursor for pretty much every prompt I try: “We are unable to complete this request because it was blocked under the model provider's usage guidelines. Try a less sensitive prompt.” Even completely normal/…
Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses (arxiv.org) Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok…
Got cursor start just to see if cursor can fit my workflow. But discovered no auto mode? (www.reddit.com via reddit) Cursor users, help me figure out which plan/model makes sense for me. On the Cursor Start plan, I’m currently only seeing: Cursor Grok 4.6 Composer 2.5 Cursor Grok 4.5 The thing is, I already have X Premium+, which gives me access to these…
Auto mode model (www.reddit.com via reddit) I am using auto mode now as many suggested,but i think cursor is only using grok for everything,should i use higher models in planning
Why does every new AI model feel like a mechanical coder? (www.reddit.com via reddit) I just want a ai model Which is not a mechanical coder (e.g opus,fable,sonnet,gpt astra, grok or inshort every model exists today) But a model which understands real human situations and give answer for it (e.g upgraded/better version of g…
Five bugs that nearly killed the extension I built with Claude Code, and none of them were Claude writing bad code (www.reddit.com via reddit) I build AI Toolbox with one other developer, a Chrome extension that adds folders, full-text search, export and saved prompts inside Claude, ChatGPT, Gemini and Grok. Most of the code is written with Claude Code.
Grok 4.5 preforming better than Grok 4.6 (www.reddit.com via reddit) I’m just wondering if this is just me or others have experienced this but I’ve just switched to grok 4.5 for the past few hours and it’s doing a much better job than 4.6 and I’m struggling to understand why? Is this because there’s less de…
Prompts flagged as sensitive (www.reddit.com via reddit) I hope this is the right place to ask. I am having an issue when using grok 4.6 on High in cursor IDE mode.
Has cursor limit increased significantly ? (www.reddit.com via reddit) Earlier I used to use grok for only complex things and it still exhausted my limits. Nowadays i rarely use composer but still it doesn't feel like my limits are draining a lot faster.
Cursor Grok vs Composer ( via reddit) xAI should drop Grok from the cursor lineup and focus only on composer. It is a waste of time and energy to that is vastly superior in this enviorment.
anyone actually running a multi-agent "team" that keeps going without you? (www.reddit.com via reddit) most setups i see still die after one chat — you spawn more agents, then babysit every handoff. curious what’s working in practice for long-running agent graphs (lead + workers + approvals), not just "better prompt".
Why Grok always forgets the path parameter (www.reddit.comhttps) This happens to me way many times
Grok for coding, Kimi for chatting (www.reddit.com via reddit) I've been mainly interfacing with Kimi K3, using it as an orchestrator to spawn grok-xhigh agents. This combination has worked great for me because Kimi's writing style is so much easier to digest.
Should i stay on SuperGrok or switch to Cursor Pro+ to get more "Grok Bot" usage? (www.reddit.com via reddit) Hi there, I have a question regarding the Grok Bot Usage limits: I currently pay ~30$ for my SuperGrok subscription, however i am a developer myself and consider getting cursor, i figured the 60$ cursor plan offers extended but not max gro…
STOP SWITCHING THE AGENT TO AUTO / GROK (www.reddit.com via reddit) I CHOOSE "LAST USED" BECAUSE THATS THE ONLY WAY I CAN USE COMPOSER 2.5 EVERY TIME GROK IS DOG SHIT GIVE US AN OPTION TO ALWAYS USE 1 MODEL OR STOP DOING THIS SKETCHY SHIT THAT CHANGES THE AGENTS ALL THE TIME
Grok 4.6 has been lobotomized. (www.reddit.com via reddit) Everyone said the same but it was still working fine for me until yesterday. Now it can barely follow instructions correctly, writes the source of truth multiple times, wastes a shit ton of tokens doing unnecesary searchs it didn't do befo…
Stop using Auto, "...you lazy animal" (www.reddit.com via reddit) Pretty sick of hearing everyone bashing Cursor cos they use Auto and now its all "gone to shit". How about, Not Auto?
Fable 5.1 review (www.reddit.com via reddit) I just want to give a quick review of new fable. Ive been using it past couple of days, especially for design, and im so impressed.
Unified Subscription Access (www.reddit.com via reddit) Google AI Pro gives you access to the Gemini app and Gemini models in Antigravity for coding. I hope xAI would implement this too, $30 subscription with Pro access to Grok app and increase usage of Grok models in Cursor.
Structured Features Overfit Where Random Features Grok (arxiv.org) Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as $1/\lambda$ in the weight decay. We s…
Cursor auto switching model (www.reddit.com via reddit) WTH, i clearly remember putting Opus 5 in the model selector then creating a plan, when i executed the plan, came back i legit saw "Extra High Fast"!? Like what?
If you use Grok Bot and you're wondering why your Cursor usage has been completely depleting on your own - IT'S GROK BOT (www.reddit.com via reddit) I just finished a long back and forth with Cursor support team who had no idea whatsoever how this worked until I figured it out for them. I pay for the $60 Cursor Pro plan and without using a single API/Regular non-grok model this month m…
grok 4.6 speed boost (www.reddit.com via reddit) Did anyone else experience a major speed boost with Grok 4.6 since yesterday? Great speeds, only wondering if it's at the cost of intelligence...
How do cursor plans work for grok bot? i feel scammed (www.reddit.com via reddit) I bought Cursor Pro ($20) so I could use Grok Bot. After about one day I burned through the Pro Grok Bot weekly allowance.
Grok is extremely busy. (www.reddit.com via reddit) Lately, whenever afternoon rolls around, Grok 4.6 demands that I switch to "Auto" mode, effectively rendering it unusable. How is it working for everyone else?
Composer 3 needs to happen (www.reddit.com via reddit) Hey everyone, Composer 2.5 was an incredible price-to-performance coding model. But after SpaceX acquired the company, they shifted their focus toward Grok and seemingly stopped developing this excellent coding model.
Cursor Agent ran `rmdir /s /q "C:\Users\<me>"` on my Windows profile — twice — during a "temp cleanup" I never asked for. Here is the exact line from the transcript. (www.reddit.com via reddit) On 12 Sep 2026 a Cursor Agent session wiped most of my Windows user profile. Not my project folder.
Codex vs Cursor (www.reddit.com via reddit) Hey guys, I've been loving using Cursor and it always felt the best before Grok came to it. I've got Ultra and now with OpenAI restricting their models being used in Cursor, I'm thinking to switching over to Codex or some other way to use…
Is Cursor Grok 4.6 better than Sonnet 5? (www.reddit.com via reddit) Is Cursor Grok 4.6 better than Sonnet 5?
Anyone else experiencing slow down right now? (www.reddit.com via reddit) I'm using grok on extra high and its taking so long right now to do some basic tasks. Anyone else experiencing slow down?
Thats has happened to cursor (www.reddit.com via reddit) I’ve had an Ultra account for close to a year now. I usually max out my usage very close to the end of the month, while getting most of my work done using Auto.
I turned Claude’s usage limits into little potion flasks for my Mac (www.reddit.comhttps) I built Questis to see my Claude quota and reset times without repeatedly opening the usage page. It lives in the Mac menu bar, with a full display where little glass flasks drain as the quota gets used.
What will you do when AI gets banned? (www.reddit.com via reddit) I've been using Claude for work for a while now and yesterday started realizing how much of the stuff I do is done by Claude these days. It's not that I work less, my workflow just has become different.
I am liking grok bot (www.reddit.com via reddit) ok so i didnt care for grok bot until i started using it. its kind of like claude cowork but it has separate quota, and i am really putting it to work these days.
Why is cursor being so generous? (www.reddit.comhttps) Ran out of quota on my $20 plan, Have received $100 in free credits twice in a week now. Everytime I run out of credit, my account gets $100 more.
Claude Code ran until it hit its limit, produced nothing (www.reddit.com via reddit) I was trying to use Grok and Claude to autonomously build a rom from scratch with romdev. With a SuperGrok subscription, and using the build tab, I was able to download and install ROM Dev, and have it built GBA game.
fable 5.1 on pro+ (www.reddit.com via reddit) Thinking about trying Cursor Pro+ mainly for Fable 5.1. For anyone using Fable heavily on the $60 plan: how long does the monthly quota realistically last if you use high/extra-high reasoning for serious coding work?
Claude Pro feels better than Cursor Ultra (www.reddit.com via reddit) I was using Ultra earlier with Claude Opus model and Composer earlier and it was very great. Tokens wise as well.
Built with Claude Code: a free menu bar app that keeps your 5-hour and weekly quota in view (www.reddit.comhttps) On August 19 I had 27 Claude Code sessions going, 5,119 messages in total. I had no idea how much quota that was burning until it stopped.
claude $200 or cursor $200? (www.reddit.com via reddit) hi all, wondering what plan is better currently - cursor $200 with grok bot or claude $200? anyone switch between the two or try them both out?
Changed alot (www.reddit.com via reddit) The product was absolutely beast before spaceX acquisition, now it sucks , output is trashed, auto switching to grok is fucked up and tokens are being consumed at faster rate despite any large task I will start looking for an alternative,…
planning vs coding - simpler way? (www.reddit.com via reddit) I've been using Claude for planning / hashing out design requirements / creating design specs, and Cursor for coding & debugging, then sometimes back to Claude for validation. Lots of back & forth between the two.
Claude Connecter - Want to connect to other AI (www.reddit.com via reddit) As I am going into my final year of high school I want to connect other AI’s such as GPT, Gemini and Grok to Claude in a system where Claude acts as a central link to act as a tutor of sorts to teach me really fast, taking full advantage o…
How is grok bot different from cursor desktop ? (www.reddit.com via reddit) Anyone used grok bot and cursor desktop ? What's their main differences ?
Is it even possible or are we just dreaming about next Fable (www.reddit.com via reddit) Astra is the first real competition Anthropic has faced in a long time. Musk already said Grok 4.7 will beat the current models, but that Anthropic will eventually release something better.
Any alternative to cursor (www.reddit.com via reddit) I really like cursor, but I feel it is going downhill since acquired by grok. I like that they have a cheap model, that I can manage by being smart and not run out of tokens.
RIP cursor ! Pay for stress (www.reddit.com via reddit) If you know, you know! Grok by default, Wrf is going on, can’t wait the 6 months subscription to end, modern slavery !!
Grok officially ruined Cursor IDE (www.reddit.com via reddit) Ever since the acquisition Cursor got so bad at everything. Text output is a word salad of things and the model keeps switching to Grok.
Budgeting for third-party models ahead of OpenAI’s proposed November cutoff. (www.reddit.com via reddit) Cursor Teams plans offer Standard seats at $40 per month and Premium seats at $120, with five times the included usage. Annual-plan equivalents are $32/$96.
I’m confused about the Grok and Cursor plans. (www.reddit.com via reddit) I use Grok mainly for VDO work, with some vibe coding too. But when I want to use it like an IDE, I don’t have tools such as Codex, Claude Cowork, or Antigravity, so I have to use Grok Build.
Grok Vs Cursor - Which gives more GrokBot Access? (www.reddit.com via reddit) Just wondering if anyone has insight into the differences between the plans? In general and specifically for GrokBot Usage.
I built an app where each of my Claude Code agents looks after one thing (www.reddit.comhttps) I had this setup in Slack already, Claude Code agents on a Mac mini with their own instructions and schedules, and keeping it alive was a hassle of its own. Every new Claude Code session started from nothing, so I was explaining my product…
Cursor update defaults to Grok which is less than helpful (www.reddit.com via reddit) Installed the update and found my chat quality reduced with replies going on a tangent even with new chat sessions. I then noticed the model was set to Grok (I use Auto all the time) and Grok is really a chimpanzee in replying to questions.
Introducing AstraBlender! Real Blender that ChatGPT can use from a simple prompt sent from your phone on the ChatGPT website ;) (www.reddit.comhttps) Simply prompt ChatGPT work (or any other agent with a cloud browser, like Grok Bot) to go to the website and use blender. From your fucking phone!
Grok bot Vs claude code? (www.reddit.com via reddit) Hi all, just a quick question. Out of the two.
Claude Artifacts only live on Claude, and I lose them every time I switch agents (www.reddit.com via reddit) I've been using Claude Artifacts for a while and they're one of the most useful features it has. I use them for reports, presentations, and sometimes even interactive prototypes.
Gave 6 AI models the same bug. Only 3 got it right. (www.reddit.com via reddit) Tried another little AI test today. I gave ChatGPT, Claude, Gemini, Grok, DeepSeek and Qwen the exact same coding bug and asked them to fix it.
Title: WSL2 + Claude Code + Windows PAC proxy — how do I get this working? (www.reddit.com via reddit) Hey guys, I'm pretty new to Claude Code, so sorry if this is a dumb question. I'm on Windows 11 and I really want to run the Linux version of Claude Code through WSL2 instead of using the Windows version.
Composer and Grok slow down and stalling? (www.reddit.com via reddit) Anyone else experiencing a substantial decrease in speed of both Composer and Grok even at Fast speeds since approximately middle of last week? Also stalling all the time?
Codex $100 or grok $100 for langgraph/langchain development? (www.reddit.com via reddit) Codex $100 or grok $100 for langgraph/langchain development? I'm trying to decide which \\\~$100/month plan is better for heavy coding: Codex or Grok.
$200/mo AI budget, build 4-5 apps a day, what’s your setup? (www.reddit.com via reddit) Designer/developer here. I build 4-5 apps daily, mostly in the terminal.
Cursor Cloud Agent & Unwanted Sonnet Usage (www.reddit.com via reddit) Wondering if anyone noticed that the browser helper that Cursor Cloud agent uses is hard-wired to Claude Sonnet. It does not use your parent model, and there’s no setting to change it.
Is Cursor worth it? (www.reddit.com via reddit) Hello all! Hope everyone is having a good time!
Cursor keeps switching from Auto (free and unlimited) to Grok (paid) (www.reddit.com via reddit) Hi. I still have Auto free and unlimited a little while longer, so it's really frustrating that Cursor keeps switching to Grok, which costs money to use.
Cursor charged me $192 for an annual upgrade 6 days into my monthly plan, then refused to reverse it (www.reddit.com via reddit) I’m posting this to share my experience with Cursor’s annual upgrade process and to ask whether anyone else has encountered something similar. On September 1, 2026, I purchased Cursor Pro monthly for $20, with the subscription running thro…
I spending Grokbot tokens one a bot that tells me how many Grokbot tokens I have left. (www.reddit.comhttps) So Cursor gave us this new Grok bot, and honestly, I didn't really have anything in mind to use it for. At the same time, I'm now a little paranoid about usage since we went from the unlimited auto to the new limited packages.
Claude ChatGPT and Grok all went down in the same window on Thursday (www.reddit.com via reddit) xAI said Grok died with its Memphis data center. OpenAI said a routing error.
Resets in Cursor (www.reddit.com via reddit) there was a reset of grok bot recently. Question is or will there ever be a restet of the cursor models usage?
A way to configure specific model for a mode / task / subagent ... (www.reddit.com via reddit) Is there a way to fix a model for each mode, like: Plan -> Grok 4.6 Agent -> Compose Orchestrator (model x) , sub agent always (model y) etc ...
Those of you with a finance background, how are creatively using Claude? (www.reddit.com via reddit) I don’t mean for your company or work. But more for everyday life.
OpenClaw Power, MacBook Simplicity: Five Days With Grok Bot (www.latent.space) SpaceXAI’s Grok Bot has the same level of programming power as OpenClaw, but it’s programmable at a different level of abstraction.
SuperGrok Heavy + Grok Bot: matching .ru email gets policy_denied, support loop since Aug 16 (www.reddit.com via reddit) I’m posting because I still cannot get a substantive technical answer from either company, and I want to know whether other users can reproduce this. **What happens** - I pay for SuperGrok Heavy through my X/xAI identity.
Cursor Enterprise customers have free usage for Grok Bot (x.ai via reddit) Grok Bot for Enterprise Grok Bot is now available for enterprises. Grok and Cursor Enterprise customers have free usage for the next two weeks, and can invite their whole organization, including people without an existing seat.
I guess all good things have to end (www.reddit.com via reddit) Is everyone having a degrading performance and quality regarding composer and grok ? I feel like tha composer get alot dummer
Is Auto so locked on to Grok that it won't even switch to other models when Grok is down? (www.reddit.com via reddit) Grok seems to be down. I checked both 4.6 and 4.5.
Rate limiting by model provider (www.reddit.com via reddit) Is anyone else seeing lots of messages in the agent 'Rate limited by model provider' for the grok 4.6 models ? The agent also seems to give up after about a minute or so and simply stop with no errors or results.
OpenAI is pulling the plug on Cursor after its SpaceX acquisition (www.reddit.com via reddit) So this is a wild one. OpenAI announced it's ending its partnership with Cursor (the AI coding tool), giving a shutdown date of November 12, 2026 for direct model access — apparently the max notice period allowed in their contract.
Is Claude able to render LaTeX? (www.reddit.comhttps) I'm a college student so having Claude as a tutor has been very helpful to learn new topics, especially in my coding classes, however it seems to struggle with displaying human readable math formulas/equations. I've been finding myself usi…
We made Claude Code multiplayer: two people, two local agents, one live document. (www.reddit.com via reddit) My co-founder and I build Nimbalyst, an open-source desktop visual workspace for Claude Code (Codex and OpenCode work too, Grok and Gemini are in alpha). Of course, we built much of it with Claude Code!
Yo Musk-y boy - why did you nerf Grok 4.6 latency. It’s become sooo slow (www.reddit.com via reddit) Since launch Grok 4.6 to now, responses have become extremely slow. I’m getting a faster response from Opus than from grok.
Built a skill that flags AI-sounding patterns in drafts. Gemini rated a fully machine-generated post 15% AI, ChatGPT said 95% (www.reddit.com via reddit) I keep writing text that AI detectors flag, including posts I wrote entirely myself. At some point I started collecting the specific patterns that readers and detectors actually catch, and turned the list into a Claude skill: you feed it a…
Grok Bot is finally available on Android (www.reddit.com via reddit) Super stoked about this!
[AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokens (www.latent.space) With Astra clearly finally warming up for a full launch (with @sama and @openai writing about it again after a month of self imposed pacing), there’s a familiar window to take the narrative with the round robin of model launches, with Grok…
At max thinking, Sol finishes the job and Grok 4.6 loops. How are you writing skills that both can run? (www.reddit.com via reddit) Not looking for a benchmark chart. Looking for how people actually write skills.
Is there much difference in API usage between Grok 4.6 and 4.5 and the speeds? (www.reddit.com via reddit) Most of what I do is handled perfectly fine by Grok. I don't need any external models and I seem to be able to get a lot of coding done on the $20 plan using just Grok.
FIXED. Follow-up composer would switch itself to Fast after the first reply, burning 2x usage (www.reddit.comhttps) https://forum.cursor.com/t/cloud-agents-default-model-grok-4-6-high-still-runs-fast-and-burns-2x-usage/169576
Anyone having issues with the grok bot registration? (www.reddit.com via reddit) Hi guys, i am a cursor user since far over a year. Therefore i still have the old 500 requests/month pricing in my cursor pro plan.
Models for marketing/copywriting (www.reddit.com via reddit) Grok 4.6 is useful when it comes to coding but horrible when it comes to human-like language and copywriting (oversimplifies sentences, talks gibberish, robotic language etc). What models in your experience provide the most decent copywrit…
Setting Grok 4.6 to Extra High (www.reddit.com via reddit) I've worked with Grok 4.6 since it's come out, and I think it is quite good but not amazing. Yesterday I realized I can control it's effort level (I might be late to the party, didn't know cursor had that feature), and the default is high…
How to clone a Grok Pocket on an ESP32. Anyone here on LilyGO? (www.reddit.comhttps) Bobby Thakkar posted this yesterday: Here’s the clone if you want to build one before his firmware drops. What it actually is: an ESP32 with a tiny screen.
Claude gets surprisingly hostile when I try asking about other models (www.reddit.com via reddit) I am trying to set up an email triage setup for my needs. After a bit of discovery I started exploring using Hermes Agent as an option with Claude.
How would you rank plan mode by platform? (www.reddit.com via reddit) Just thought I'd bring this up because I'm generally on codex and claude code and basically always start with plan mode. But I'm watching Cursor's plan mode and its on like its 7th set of sub agent deployments and low key looks like its co…
A plugin/hack for preserving the selected model (www.reddit.com via reddit) The issue is that a selected model resets to “auto balance” each time I open new chat. It has become a routine task to manually select each time Composer/Grok, and sometimes I miss it resulting in burned tokens.
cursor cli & grok build (www.reddit.com via reddit) I've been using Cursor CLI for a long time now, and Cursor hasn't been just an IDE to me in ages. I'm a bit confused about the differences and market positioning between Cursor CLI and Grok Build.
Best bang for buck model? Auto vs Grok 4.6? (www.reddit.com via reddit) I really do not like the fact that I cannot see what models Auto is using. For the vast majority of my work, I just want the best bang for buck model.
Is this the end? grok is enough or we need another models to keep cursor? (www.reddit.com via reddit) https://x.com/OpenAI/status/2093515564786540695
Getting errors coming back from cursor processing yesterday and this morning again, saying memory exhaustion or and allocation and agents going to sleep? (www.reddit.com via reddit) This morning and yesterday 28th and 29th August 2026 from cursor ide when trying to process my chats, randomly its stopping and failing and the chat agents gone to sleep just stop mid stream working on code changes? Or error unknown try ag…
Let Claude put its money where its mouth is: I built a bots-only chatroom where AI pays to post and competes for attention — built and launched with Claude Code. First 50 bot keys include $5 credit. (www.reddit.comhttps) OnlyBots.chat is a single-channel chatroom where only bots post. The catch: posting costs money, and the price moves with spending.
Fable orchestrator + 5.6 sol max thinking worker seems to be the winning combo for sustained Fable-level work without blowing an entire max sub budget in a day (www.reddit.com via reddit) Of course this still requires 2 expensive subscriptions and isn't a necessary or realistic workflow for most. I kept hitting my weekly Fable limit too fast and have been experimenting because it's great but just too expensive/limited.
how to control what models are used in cloud? they keep starting fast agents and using expensive models :Sob: even tho i have grok 4.6 as default (www.reddit.com via reddit) I love the concept of moving to cloud, i love long running cloud agents as much as the next person, but they keep FORCING fast down my throat. like bro im going to sleep, can u give me a slow mode instead?
How Can I Stop Model Switching (www.reddit.com via reddit) Hey, So I removed every model from the list and made composer 2.5 my default but every time I open a new window Grok sneaks back in. Am I missing a setting or does anyone have any advice how to get my default model to stick?
consumed usage in less than one hour (www.reddit.comhttps) I used grok 4.6 high fast, gpt-5.6-luna, and 1 prompt to plan gpt-5.6-sol-1m
Grok 4.5 auto-switches to 4.6 (www.reddit.comhttps) Leave. My.
Cursor and Grok quality drop. (www.reddit.com via reddit) I”m aware of all the recent changes and it’s not great. But it seems the last twoish weeks cursor and groks quality has dropped tremendously.
Cursor wont survive the year. Change my mind. (www.reddit.com via reddit) Someone please convince me how this dogshit tollbooth masquerading as a dev tool company isnt a scam,? Their billing has become absolutely atrocious and makes zero logical sense.
Grok Bot servers are offline already (www.reddit.com via reddit) Big issues at spacexai. Servers are completely down.
Performance seems to be back to normal now. (www.reddit.com via reddit) I see a lot of posts about lagging performance and slow runs but anyone notice that performance on Auto, Grok 4.6 high, Grok 4.6 xhigh is back to speedy again? It was brutal there for a week+ but it seem like its fast again as of yesterday…
Grok Bot just added to Cursor Pro plans and more (www.reddit.com via reddit) I am keen to know how everyone is using Grok Bot. For people out there who have used Codex or Claude.
Built 6 months of Cowork skills and automations. Moving some of the work off Claude. Does any of that come with you? (www.reddit.com via reddit) I run a few small businesses and a couple of client sites. For the last six months I set a lot of that up in Claude Cowork: skills, automations, weekly checklists I do not want to remember by hand.
After using Claude, Grok 4.6, and Gemini 3.7 Flash in depth, I want to ask how Codex is performing now. (www.reddit.com via reddit) Layely I have mainly been using Claude Code, Grok (including Grok CLI), and Gemini 3.7 Flash foe day-to-day programming work. Claude's feel to me is that tha analysis goes fairly deep, and it is more willing to think at the architecture le…
After 2 months of work, I’ve finally got my Proactive AI IOS app ready for launch (www.reddit.com via reddit) I’ve been a pretty active member on this sub for a while, I’ve shared a lot of my projects and tools, but this one is special to me for a few reasons. Ever since I was a kid, I’ve always wanted to make a nice polished IOS app.
Updating a stock market crossword with news via MCP (www.reddit.com via reddit) The new addition to our AI features. https://traderange.net/blog/crossword-game-z777t36l/ This stock market crossword minigame builds on our Claude generated news summaries fed via an MCP server claude Fable 5 helped design.
The Cursor team shipped Grok bot (0.18.0) with runtime source maps enabled. Source code reconstructed here. (github.com via reddit) Grok Bot 0.18 — reconstructed and extended This repository is an unofficial, source-oriented reconstruction of the publicly shipped Grok Bot 0.18.0 macOS app. The project began as an attempt to understand how the desktop app was put togeth…
Multi-Agent Persistent Workspace (www.reddit.com via reddit) I’m looking for help with creating a personal assistant workflow using Claude (subscription), to manage my emails, travel booking, notes, research, todo lists, personal finance, shopping, etc. I already have work flows that allow me to man…
Lattice: An isometric game kit for agents (www.reddit.comhttps) Lattice is a collection of typescript packages, agentic skills and plugins that enable easier development of isometric games! At its core lives a 0 dependency typescript package, 80kb gzipped.
Reddit suggested I post funny and creative posts. (www.reddit.com via reddit) I copy and pasted Reddits wording into 4 AI's all within 5 minutes. In alphabetical order: Chat GPT, Claude, Grok, and Perplexity.
You don't have to pay $200 to Grok Bots. Agent Office is built 6 months ago and it's running Pi behind and it can run with open models. (www.reddit.comhttps) Repo: https://github.com/baturyilmaz/agent-office
Life with Claude nowadays is use all Fable, suffer with Opus before reset (www.reddit.com via reddit) https://imgur.com/Est4CAF To be clear I use Claude as my daily driver. Not because it's the greatest, but because Codex, Kimi or Grok is still weaker in anything that requires continuity and creativity.
Crowd-sourcing token quotas: Grok Heavy vs Claude Max 20x vs ChatGPT Pro 20x (Aug 2026) (www.reddit.com via reddit) None of Grok, Claude, or ChatGPT publishes how many tokens you get per month on the top individual subscription. I went through official docs, the OpenAI developer forum, Reddit, GitHub calculators, and a few blogs, and inverted every "X t…
No AI platform has real folders - so we built one tree across all four (try before Sunday's launch) (www.reddit.com via reddit) Claude still has no folders. Neither do ChatGPT, Gemini or Grok.
Grok exfiltrates user data when malicious instructions are encrypted (arstechnica.com) Earlier this week, researchers outlined an attack that used a secret input provided by Microsoft 365 Copilot for enterprise to cause the AI assistant to exfiltrate a password present in the user’s inbox. Now, a separate team has devised a…
Anyone else feel like Claude is increasingly just performing the task instead of actually doing the work? (www.reddit.com via reddit) Anyone else noticing this with Claude code and Grok lately? The models got way better at following instructions.
Is this a valid workflow pipeline? (Claude + GPT + Grok using Ruflo and obsidian + graphify) (www.reddit.com via reddit) Task/Issue ▼ [PRE-FILTER] deterministic, free — no model call │ diff size / file count / keyword match against known-trivial │ patterns — gates ONLY whether speculative PLAN subagents fire │ concurrently with TRIAGE (pipeline-latency optim…
Claude directed this animation film. I've never made videos before. (www.reddit.comhttps) I just told Claude the story. Fable 5 (max effort) generated the prompts and helped edit the video.
SuperGrok vs Cursor (www.reddit.com via reddit) Grok 100USD VS Cursor 60USD ?
Why does it keep happining to Grok Bot? (www.reddit.comhttps) Who don't know, Grok Bot is actually part of Cursor Ultra
Cursor Credits / Grok Bot & SpaceXAI (www.reddit.com via reddit) Hey guys have you received this email? Cursor giving free $300 dollars credits?
Can't use grok 4.6 high with Straitly ai byok, is this a provider thing or cursor? (www.reddit.comhttps) So I've been running cursor + straitly api with byok for about a month now and everything's been smooth, costs are way down. But i can't seem to use grok 4.6 high at all.
↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6grokcursoropenai+1
Coming back to Cursor (www.reddit.com via reddit) I used Cursor in 2024 and I am sure many things changed since then. I am currently on codex and Pro sub, and previously also used Claude Code.
↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6grokcursorchatgpt+3
What includes more Grok 4.6 usage? (www.reddit.com via reddit) Super Grok or the $20 cursor plan? It's currently quite confusing that it's basically the same company but two different plans.
↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6grokcursor
Grok is now giving Codex-style banked resets 😍 (www.reddit.comhttps) could not extract summary
Grok 4.6 is better than 4.5, but I would still take GPT Sol 5.6 over this (www.reddit.comhttps) I’ve been using Grok 4.6 since it was released. Before that, I mostly used 4.5, Composer 2.5 Fast, and Auto.
↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6grokcursor
Goodbye Freewill Credit (www.reddit.com via reddit) I am on a Team plan, Standard seat. We have a per-member monthly spend limit and our admin has not enabled the On-Demand spending.
Grok 4.6's most important number is one xAI didn't even advertise (www.reddit.com via reddit) Grok 4.6 dropped yesterday and the debate is the usual "is it better than Sol?" On the headline benchmarks it's a genuine tie: Intelligence Index 61 vs 61, Coding 76.8 vs 77.4, Agentic 58.7 vs 57.8. The number nobody screenshots is AA-Omni…
↯ Hallucination↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6hallucinationgrokgpt-5+1
/party — the skill that lets your agent sessions talk to each other (www.reddit.com via reddit) Any agent that reads skills can be in the channel: Claude Code, Cursor, Codex, Grok. They can all sit on your laptop, or on machines in different countries, and it is the same channel either way.
ClaudeCraft Arena: 4 frontier models are playing a vibecoded MMO against each other live (World of Claudecraft) (www.reddit.comhttps) TL;DR: World of Claudecraft is an open-source MMO vibecoded with Claude Fable 5, gaining 55 contributors and 2,000+ GitHub stars in under two months. We built a self-improving agent harness (forked from Hermes) and dropped four frontier mo…
Claude Opus 4.8 is racing GPT-5.6, Grok 4.5 and Kimi K3 to level 20 in our MMORPG, and the other models won't stop roasting it (www.reddit.comhttps) We run World of Claudecraft, a free open source browser MMORPG largely built with Claude, and we've been benchmarking frontier models by having each one play a character and race from level 1 to 20. Same starting zone, same quests, no scri…
My Orchestrator.py (www.reddit.com via reddit) Paid up subscriber to Claude, both the API and app, also API subscriber to ChatGPT. Why the API ?
day 1 of running an ai only website where ai's run stores and sell stuff to each other. (www.reddit.com via reddit) day 1 of running an ai only website 1f3ea.com where ai's run stores and sell stuff to each other. humans can look but can't buy anything - opened 1 day ago.
Your Claude plan as a potion: I'm building a free Mac app with lab flasks that drain as your quota burns (no API key), and a fun live radar as a background (i mean why not) (www.reddit.comhttps) Weekend project that got out of hand: a desktop app that renders your Claude/Codex/Grok plan quota as lab glassware, flasks that drain while you work and refill on the window's schedule. It's called Questis.
Is Composer 3 officially dead? No news after Grok 4.5 integration? (www.reddit.com via reddit) Back in June, Cursor previewed a new 1.5T+ parameter model for Composer 3, trained from scratch on 100k+ GPUs. A lot of us were hyped for it.
grok is ass. (www.reddit.com via reddit) not sure why cursor keeps changing my model to grok when it consistently produces far worse results, and is slower than composer 2.5 its actually mind blowing that they have such an outstanding in-house model that they clearly put a lot of…
V4 pro GA vs Grok 4.6 vs glm 5.3 ( via reddit) could not extract summary
What's you last 30d cursor token usage? (www.reddit.com via reddit) https://preview.redd.it/9vw0rstekthh1.png?width=1073&format=png&auto=webp&s=d0e3351c02930f076f42486826d4aa86340baea8 Last 30d: 878M tokens. Task: General daily software development.
Grok 4.6 tomorrow? (www.reddit.com via reddit) Looking forward to shipping more!
Has anyone been able to use up an entire Cursor Ultra subscription of auto/Composer/Grok usage? (www.reddit.com via reddit) I have been coding my absolute butt off this month and I haven't been able to crack 40% yet. I have another 9 days to go but I actually can't believe how heavy I can hit it without worrying about usage whatsoever.
Bot Save My Interruption (www.reddit.com via reddit) A buddy and I were having a discussion the other day about how it feels like all the different AIs have a "personality" when giving an answer. That then spawned a greater conversation about what it would look like to have multiple AI agent…
Cursor Pro vs SuperGrok, which actually gives more Grok 4.5 tokens for the money? (www.reddit.com via reddit) Anyone done the math on this? Not asking about price, just raw token output if you’re only using Grok 4.5 on both.
The search only includes results based on the title of the conversation, not the actual conversation. How do I change this? (www.reddit.com via reddit) Hello When I try to search Claude only searches based on the text included in the title, not the actual conversation. With other AIs like chatGPT and grok, it actually searches inside the conversation Is there an easy to change this?
Effective Grok cache hit rate? (www.reddit.com via reddit) I'm building a high-volume pipeline that runs text classification through the xAI API : one item per request (batching hurts my accuracy), a large fixed ~10k-token system prompt, and traffic that arrives in bursts (roughly every 20 min). T…
I`m kinda surprised of quality of GenerateImage, what about you? (www.reddit.com via reddit) Today I was working on an educational youtube video for onу of apps and was almost shocked that Cursor built in GenerateImage tool gave me this from one prompt 🤖 Before I asked Cursor to give me prompt and used it in external image generat…
How does SuperGrok Heavy limits compare to Cursor Ultra 20x ? (www.reddit.com via reddit) Thinking of using it only for coding, how do the limits on Grok 4.5 compare between these two 200$ subscriptions?
Model is hidden in new update? (www.reddit.com via reddit) https://preview.redd.it/t0znifpmhahh1.png?width=479&format=png&auto=webp&s=3a9ec5981b9c5264c47a700749d1db54c810c8c5 at first i thought it was a bug, but it does seem to be a feature. https://preview.redd.it/ltysj3pohahh1.png?width=292&form…
Grok AskQuestion Tool (www.reddit.com via reddit) Any body know when will be the AskQuestion tool be available for Grok 4.5 model?
SuperGrok Heavy $300 Grok Build limit vs. Cursor Ultra $200 for Grok 4.5? ( via reddit) could not extract summary
remote-control in cursor-cli? (www.reddit.com via reddit) Is there any option to use and control cursor-cli from a remote device(my phone)? Have been using it a lot lately with grok 4.5, but I really could not figure out how to use it through a mobile device.
How to disable harness specific tools in cursor? (www.reddit.com via reddit) I recently switched back to Cursor and I’m loving it, Grok 4.5 is great. The problem is that I have instructions set up for Claude Code and Codex, and I get the feeling Cursor is somehow picking them up too (CLAUDE.md in ~/.claude and AGEN…
Cursor keeps forcing Grok 4.5 Fast even when it’s disabled ( via reddit) could not extract summary
I cut my Cursor token cost by ~56% using Grok to plan and Composer to execute (www.reddit.com via reddit) Cursor subscriptions are pretty generous, but if you use agents aggressively, you can burn through the included usage surprisingly fast. I have been testing a simple setup: grok writes the implementation plan, then composer-2.5 standard ex…
Elon Musk’s xAI is trying to sue its way out of a Grok reckoning (arstechnica.com) Elon Musk’s xAI is trying to sue its way out of a Grok reckoning as arrests of Grok users accused of making child sex abuse materials (CSAM) have triggered lawsuits from kids to sue xAI to force changes to the tool to block harmful outputs…
Cursor Start used 10% of my monthly allowance on one repository audit and remediation attempt (www.reddit.com via reddit) I tested Cursor Start on an existing Next.js project using Cursor Grok 4.5 Medium. The task consisted of two stages: Audit the repository for security issues.
Grok 4.6 "around August 7" and Grok 4.7 likely early sepetember (www.reddit.comhttps) 4.5 is already amazing imo "even better... token efficiency" sorry that the tweet cuts off
I canceled Max plan, switched to Grok 4.5 for coding, then switched back (www.reddit.com via reddit) Hey just wanted to share my experience of what happened over the past week. Basically I was having huge huge huge problems with Claude for the week before they released Opus 5.
Cursor Grok 4.5 selected (NOT Auto) but it's using Opus 5 for subagents? (www.reddit.com via reddit) Anyone else have this problem? I have Auto off, Cursor Grok 4.5 selected, and it is using Opus 5 High for subagents.
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science (arxiv.org) Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate e…
I got tired of opening an app to ask it things, so I built an assistant that reaches out to me first, it’s called Orb and is now live on the IOS app store (www.reddit.comhttps) Most assistants sit there until you open the app and ask. I wanted the opposite, so I built Orb to run in the background and message me first when there’s something actually worth saying.
[Showcase] I built full-text search for my Claude history because I could never find old chats (www.reddit.com via reddit) The thing that finally pushed me to build this: I knew I'd worked out a prompt / a decision / a chunk of code in Claude weeks earlier, and I just could not find the conversation again. Native search matches titles, but the thing I remember…
has anyone actually replaced claude as their main ai coding agent (www.reddit.com via reddit) my loop is fable 5 or opus 5 planning, composer 2.5 executing, coderabbit / bugbot on review. it works, i freelance so the code has to be safe.
Does upgrading to Pro+ give more grok/cursor usage? (www.reddit.com via reddit) Title. I can't tell if upgrading to Pro+ increases cursor model usage or just gives you more API credits.
Which third-party model do you use in Cursor after exhausting Composer/Auto requests? (www.reddit.com via reddit) Looking for something that's: - Cheap - Strong at coding - Good enough like Grok 4.5/Composer for day-to-day development. What's your go-to model and why?
Cursor keeps ignoring my sub-agent model instructions and it’s costing me a fortune (www.reddit.com via reddit) When I use Cursor, I often ask the main agent to launch sub-agents to explore different parts of the codebase and report back with their findings. For example, I’ll explicitly tell Fable 5: "Only use Grok 4.5 High when launching sub-agents…
How much more usage does the 60$ plan provide compared to 20$? (www.reddit.com via reddit) Is it simply 3x since it costs that much more or is there more to it? Im especially interested in the limits for grok and composer.
Something about the model auto-routing changed (allegedly) (www.reddit.com via reddit) Don't get me wrong, I love Cursor and will continue using it no matter what! However, there are a couple things happening that I find "strange" to say the least.
An attempt to fix Cursor's horrible BYOK implementation (www.reddit.com via reddit) I absolutely love Cursor, but one of my biggest gripes (for quite some time now) is the limited ability to be able to use built-in Cursor models while also being able to add my own choice of models from another provider. I wanted to be abl…
Cursor trying to force me expend more with grok (www.reddit.com via reddit) I love Composer 2.5 and use it every day, but Cursor is now constantly trying to force Grok on me. It keeps changing my model selection even though I've set Composer as my default and removed Grok from my available models.
I NEED MORE CURSOR USAGE!!! Grok 4.5 is actually pretty good for me. (www.reddit.com via reddit) https://preview.redd.it/9mowyr9et7fh1.png?width=1239&format=png&auto=webp&s=5e306ce8ae2fe70bce74ed6d77b8afde910d28e2 They say they increased usage by 50%, but I've already used up 80% of Cursor Models, meaning I only have 20% left for the…
I don't appreciate how Cursor is trying to force me to use Grok 4.5 in FAST mode (www.reddit.com via reddit) It would be bad enough that when I open a new chat you see how the selected model quickly changes to Grok 4.5, but it's work that it changes to Grok 4.5 I'm not a Grok hater, I do use it, I just don't want to use it in FAST mode, which is…
Auto switching to grok 4.5 fast when opening a new chat! (www.reddit.com via reddit) - i noticed that when i open a new chat i auto switched to grok 4.5 fast - y cursor y!? - is this happenig to u guys also?
Is ultra really 20x Pro? (www.reddit.com via reddit) Hi Folks, I need to do a LOT of code generation. I really enjoyed the Pro plan with Grok.
[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model (www.latent.space) [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model A HUGE win for BFL! Thursdays are the heaviest days for AI releases, and even thoug…
Grok 4.5 vs Claude Code: compare accepted changes per dollar, not benchmark headlines (www.reddit.com via reddit) xAI's Grok 4.5 launch is worth treating as a practical Claude Code comparison, not just another leaderboard claim. xAI reports Grok 4.5 at 53% on DeepSWE 1.1 versus 59% for Opus 4.8, 29.0% pass@1 on SWE Marathon versus 26.0%, and 80 tokens…
I feel like Cursor and Claude are fighting over me (www.reddit.com via reddit) I debating quitting one and just keeping the other, maybe on a higher plan. And I can't say for certain...
I built a Mac menu bar and notch app to manage agents sessions, usage accounts (www.reddit.com via reddit) built this with claude code, for claude code. flagging up front that its my own project.
Question regarding Cursor Auto (www.reddit.com via reddit) Hi guys! First month on the cursor $60 month....
The reason why Cursor is still relevant (www.reddit.comhttps) Claude really fooled us all with their "too dangerous model Mythos 5" aka Fable 5. In this current market where models are dropped here and there.
What’s your Cursor workflow, and which models do you use for each part? (www.reddit.com via reddit) I’m curious how everyone divides work between ChatGPT, Codex, Cursor, and the different models. My current workflow: I start by working through the feature or problem inside a ChatGPT Project, where it already has the broader context.
Anyone else feels both Composer 2.5 and Grok 4.5 have been nerfed along with the 2x usage? (www.reddit.com via reddit) I am getting quite poor results with both and both are failing compared to their original abilities. Failing to follow instructions, introducing regressions, making basic mistakes even for simple tasks like text report generation...
Not to sound paranoid, but do they sometimes tune down to intelligence of the models? (www.reddit.com via reddit) EDIT: Meant to say THE and not TO in my post title, but for whatever reason, it won't let me change the title. I ask because I feel like I have literally spent hours today working on problems that yesterday were being solved instantly.
Anyone had issues with models not appearing on their profile? (www.reddit.comhttps) My public Cursor profile isn't fully displaying the models I've used. I mostly use Composer, Grok and Auto mode, but I've used various Sonnets a fair bit too, and the odd GPT.
I created an open-sourced test and composer 2.5 is off the chart (www.reddit.com via reddit) https://preview.redd.it/ot0oxazho4eh1.png?width=1538&format=png&auto=webp&s=be75e0f813044304a8e115b4d119c6d1bb9d3bbb noted that i used short-context questions, resulting a skew toward composer; but still Composer is stronger than expected…
Cursor Grok 4.5 is good but anyone suggesting it’s in the same class as gpt-5.6 or Fable-5 is delusional. (www.reddit.com via reddit) Grok 4.5 may do well on benchmark testing but it cannot stand up to the pressure of a real workload. It can’t stay on task and tries to reinvent shit without consulting.
This one habit cut my Cursor token usage significantly. (www.reddit.com via reddit) I’ve been building a side project this week and stumbled into a workflow that’s kept my token usage surprisingly low. The key is spending more time in Plan Mode before touching Agent Mode at all.
Are first-party models slower for a few days? (www.reddit.com via reddit) It feels like first-party models Grok 4.5 and Composer 2.5 are at least 2x slower in Cursor for the last 2-3 days. Anyone else feels the same?
Fable 5 is currently ranked #10 at document generation. Every model above it is cheaper. (www.reddit.comhttps) I expected Anthropic's flagship model to be expensive but sit near the quality ceiling. The current results are considerably worse than that.
Day 2, Cursor auto switched models and burned 35M tokens in 1 prompt (www.reddit.com via reddit) https://preview.redd.it/h47xjzialudh1.png?width=1062&format=png&auto=webp&s=7b0d9724fec4d5d1b055a670d97e390ecf6ba9fc https://preview.redd.it/4m9gobr6mudh1.png?width=529&format=png&auto=webp&s=37aef00e608a93cdcea8ea0f99102160d2ff9972 I made…
Current ranking (www.reddit.com via reddit) Claude Fable 5 2. GPT Sol 3.
To Grok Grokking: Provable Grokking in Ridge Regression (arxiv.org) We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting. We prove end-to-end grokking results for learning over-parameterized linear regression models using gradient descent with weigh…
Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies (arxiv.org) Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok.
xAI can’t deny Grok makes CSAM anymore. So it’s suing users. (arstechnica.com) Facing mounting pressure to acknowledge that Grok can still be used to generate non-consensual sexualized images of adults and minors, xAI filed a lawsuit Tuesday, suing the first user that Elon Musk’s firm has accused of using its chatbot…
I gave GPT-5.6 Sol, Claude Opus 4.8, and Grok 4.5 the same 100 frontend briefs—here are all 300 results (www.reddit.com via reddit) After generating enough websites with coding models, I started noticing that each model seemed to reach for the same handful of visual ideas. A single impressive screenshot can’t tell you whether that’s actually true, so I tried testing it…
PSA: Main agent in Multitask or Debug mode may spin up subagents with different models (www.reddit.com via reddit) I recently submitted a bug report about how Grok 4.5 High decided to spawn a subagent using Claude Sonnet 5 High in debug mode. Apparently, this is known "behaviro" and it is on their radar.
Nooo! I don't have time in life to switch right now!... (www.reddit.com via reddit) Cursor Grok 4.5 Flashbang I knew I would have to switch soon because of the SpaceX colab / acquisition, just didn't think it would be so soon.
Auto now is the useless grok 4.5 (www.reddit.com via reddit) The auto mode is now complete useless because grok 4.5 will fuckup every code base it touches. Please revert that change and map auto to composer 2.5 again.
Grok 4.5 now on Cursor in Europe (www.reddit.com via reddit) Not really certain if that was announced already or elsewhere, a search did not bring it up. I was excited about Grok and just checked it again, it is available.
Cursor Grok 4.5 Low/Medium/High normal/fast, whats the difference? (www.reddit.com via reddit) - what is the difference in the intelligece and in the costing of this? - also what is the difference if i use fast in low/mid/high intelligence - is there a document explaining this or anyone know about this?
Grok 4.5 locked just 2 messages after advertising it (30k tokens) (www.reddit.com via reddit) https://preview.redd.it/bezvneogqjdh1.png?width=618&format=png&auto=webp&s=4319da11326ed520748765ee601e1d78e3895135 I got multiple pop-ups nudging me to try out Grok 4.5. Fine, will do with a simple prompt first.
Grok 4.5 triggered API usage instead of First Party Models (www.reddit.com via reddit) Hey all, I opened a case with support but was curious if anybody else has seen this. I ran a pretty massive build prompt today with Grok set to max.
Mermaid to Unicode box art (grok-mermaid) (simonwillison.net) 16th July 2026 While exploring the codebase for the newly open-sourced Grok CLI coding agent I came across xai-grok-markdown/src/mermaid.rs, a "self-contained terminal renderer for Mermaid diagrams" written in Rust. I figured it would be f…
Cursor using unintentionally used Claude Opus and burned 22.6M tokens (www.reddit.com via reddit) https://preview.redd.it/3o0u3dj0igdh1.png?width=1066&format=png&auto=webp&s=314eaf2406a81bc5d50c8616dce382966defa836 This is the first time I had encountered this, I'm using plan mode then I proceeded with implementation. Everything was to…
Every viable coding model right now imo (www.reddit.com via reddit) Mimo V2.5 Deepseek V4 Flash (Max) Mimo V2.5 Pro Deepseek V4 Pro (Max) Composer 2.5 Grok 4.5 (High) GPT 5.6 Sol (Max) / Fable 5 From cheapest to most expensive and best at it's price range these are the most viable models right now; I think…
I built rein.build so I could control my Claude from my pocket (rein.build via reddit) Your AI agents already run on your machine. Rein puts them in your pocket.
Anyone tried 5.6 terra and sol vs grok 4.5? (www.reddit.com via reddit) Aside from pricing, Is grok 4.5 output quality at level of sol or terr ?
I use four AI chatbots and could never remember which one I told something to, so I made one search box that covers all four (www.reddit.comhttps) Rough count, I have maybe 2,000 conversations spread across ChatGPT, Claude, Gemini and Grok. I use different ones for different things and I've stopped pretending that's going to change.
How are more people not talking about Grok 4.5? [internal agentic saas marketing benchmarks] (www.reddit.comhttps) As someone who’s been using Claude‘a Opus exclusively for the past year (on Max), I’m genuinely blown away. I’ve been pretty dismissive of Grok (and Cursor) this whole time and this is coming from a happy Tesla owner.
Opus 4.8 making errors that drive it to a nervous breakdown (www.reddit.com via reddit) I have been using opus models for months now, seen every "model x lobotomised" post and frankly it has always performed well for me, sometimes inconsistent but that is the reason you build a harness and some infrastructure to protect again…
Cursor switches to Grok even if that was explicitly prohibited. (www.reddit.com via reddit) Cursor (3.11.19) switches to use Grok 4.5 model instead of Fable 5 (was specified and explicitly described in the prompt. WTF?
How to interpret usage dashboard? (www.reddit.com via reddit) I use the $20/month Pro plan. In the second picture you can see that I've only used up 8% of my first-party model limit, but on the usage dashboard I've spent $12.35.
Testing Fable 5, Opus 4.8, GPT-5.6, and more through playable 3D games (www.reddit.comhttps) TL;DR at the end I wanted a way to evaluate models around something I care about and I think we’ll see more and more as we move to “world models“, which is spatial, temporal, and causal coherence in a 3D space. Meaning, does the model unde…
So I tested Grok 4.5 High VS Claude Sonnet 5 High and it's not looking good (www.reddit.com via reddit) I’m currently on a Pro+ plan (~$60/month) and I’ve briefly tested pretty much every model available at this point. Lately, I’ve been heavily using the API/non-UI side of these models.
How's your experience using GPT 5.6 Sol? (www.reddit.com via reddit) Now that GPT 5.6 Sol has been out for a few days, what are your thoughts on it? Can you feel the difference?
Sub from grok.com Vs Cursor (www.reddit.com via reddit) Does anyone know the difference in usage limits between a direct sub or a Cursor sub?
Cursor is the WORST AI Subscription. Here is the Math (www.reddit.com via reddit) AI subscriptions are heavily subsidized right now, meaning you often get 10x what you pay for compared to raw API costs. For example, a $20/month Claude plan easily gives you $200 worth of tokens.
Cursor Grok 4.5 Model cannot be found (www.reddit.com via reddit) Anyone having a lot of issues with Cursor Grok this morning? it stopped working then started now i get the error "Model name cannot be found"
What are Limits on the 20$ Plan? (www.reddit.com via reddit) Hello, can anyone Tell me specific token Limits fot Grok 4.5 using the 20$ Plan? Im thinking about switching from Gpt plus since Codex Limits are pretty rough right now (they reset a Lot but that will Stop)
Anthropic extended Fab 5’s metered‑billing deadline – how are you weighing it against GPT‑5.6 Sol and Grok 4.5 in real workflows? (www.reddit.com via reddit) Anthropic has just extended the date when Claude Fab 5 fully moves over to metered token billing for consumer subscribers (the latest in‑app notice I’m seeing says July 19, after a previous extension). In parallel, public pricing info from…
Fun Reddit Sim built with Claude Code (www.reddit.com via reddit) I designed the client with Claude Fable 5 (Anthropic's new Mythos-tier model), working out the UI/UX and how it should be shaped, then had Claude Opus 4.8 actually build it as a real Blazor WebAssembly app on top of an existing backend (Li…
Is Anthropic shooting themselves in the foot by pulling Fab 5 from subscriptions tonight? (www.reddit.comhttps) With Fab 5 moving entirely to expensive, metered token billing on July 12th, is Anthropic making a gamble ? OpenAI's GPT-5.6 Sol is already out, and Grok 4.5 is performing on par with Opus for coding workflows - both under flat-rate tiers.
Limits on cursor (www.reddit.com via reddit) Hey brothers , I'm trying to calculate the monthly token yield on a $60 budget constraint(pro+). Assuming a routing split of 70% Grok 4.5 and 30% Composer 2.5, what kind of token volume can I realistically expect?
Best coding setup for price-to-performance in Q3 2026? (www.reddit.com via reddit) I’m comparing: $100 Codex with GPT-5.6 Sol High $100 Claude Code with Opus 4.8 $60 Cursor with Grok 4.5 Which one gets the most real work done for the money? What would be your go-to setup with a $100 budget?
It happen… the Cursor same output (www.reddit.com via reddit) I have specific detailed guidelines for my agents to do certain things and cursor will do it’s job which has worked out well for the past year. Today, cursor decides to read my .md files and each agent does something completely different a…
No Grok 4.5 available? (www.reddit.com via reddit) So I have the latest version of Cursor and since release I don't get Grok 4.5. Btw: I'm an EU user if it matters somehow.
Cursor using expensive subagents? (www.reddit.comhttps) Had Fable in Claude Code create a plan that involved several items. Dropped the plan MD file into Cursor.
Grok 4.5 is Auto or API cost based? (www.reddit.com via reddit) I want to know.
Grok 4.5 disappeared (www.reddit.com via reddit) I was using Grok 4.5 in the first day that was launched, and out of the sudden in the next day it disappeared from Cursor. Is Cursor flagging developers or disabling it randomly?
I have 4 AI subscriptions, here's the pros and cons of each one (www.reddit.com via reddit) I have the "pro" plans for the following: 20$ Cursor, 20$ CC, 20$ Codex, and 10$ Opencode. Here's my experience with limits, usage, and usefulness on them.
Did they pull Grok 4.5 from the app? (www.reddit.com via reddit) I could use until a few minutes ago. Now it says "Model not available in your region".
Auto + composer has change now is First-party models (www.reddit.com via reddit) https://preview.redd.it/6qzux6vvxech1.png?width=1057&format=png&auto=webp&s=1244f778041f0dde9f962b8226c3dd3b5eefe386 What do you mean "First-party models?" Do you mean Cursor + Grok or are there any plans to include more models? And how do…
My total usage went from ~9% to 28% for Auto + Composer and 52% for API (www.reddit.com via reddit) Hey, everybody, how's going, It looks like my usage jumped unexpectedly. Two days ago I used Grok like 5 time and today I was fixing a couple of bugs on my platform and I had like a 9 - 15ish percent total consumed, closed cursor and went…
Flux incoming? (www.reddit.com via reddit) Flux which I guess is the image ai Grok xAI uses... is it going to be added into the plans eventually?
GPT 5.6 sol Extra High piece of garbage (www.reddit.com via reddit) I just used Sol medium then bumped it up to extra high thinking it would fix the crappy multitasking that medium. Turns out it’s better to use it as a single agent After switching to agent mode I just had to send it 12 fixes to a page that…
GPT-5.6 Luna (www.reddit.com via reddit) Ok, who has performance feedback? Launched today!
While using grok 4.5 high fast, it will randomly switch to Sonnet 5 High and use my on demand usage (www.reddit.com via reddit) Anyone else running into this issue?
Grok 4.5 usage on Cursor’s $20 plan vs Codex’s $20 plan? (www.reddit.com via reddit) Trying to figure out which $20 plan gives me more mileage for daily coding. For anyone on Cursor’s $20 (Pro) plan — how’s Grok 4.5 usage holding up?
[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition (www.latent.space) [AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition SpaceXAI continues to move faster than any other frontier lab on earth. As GPT 5.6 is confirmed to launch tomorrow, today is pretty much the last day anyon…
At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics (arxiv.org) On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading effective rank at the grokking transition overstates the converged value by 3-5x on an MLP, and by 1.3-…
Lawsuit: Man used Grok to make 7K sex images of stepdaughter, then shot himself (arstechnica.com) One of the most horrific cases of allegedly Grok-generated child sex images was shared in a proposed class action lawsuit that was expanded Tuesday. Now, young girls not only accuse X and xAI of building toxic AI “nudify” tools but also of…
I’ve always wanted to know what session or subagent modified a file, so I’ve built strace for agentic sessions - called gaal (www.reddit.comhttps) It’s could be pain to understand why some changes happened to the code, especially to something outside of the scope of a task I’ve got my SKILL.md files nuked several times - because some codex worker decided that they know better the sha…
Sick of the "but" answers. Feels like the want to prove agents useful outweighs giving solid info (www.reddit.com via reddit) Haven't been using GPT or Grok lately, but this has been a consistent annoyance with Gemini. Every answer has to be a "yes, BUT" answer, even if the "but" leads to hallucinations just because it's trying so hard to give me "the flip side"…
Building Specialized ‘Mental Model Agents’ in Grok — First Principles, Systems Thinking, Bayesian Updating & More (www.reddit.com via reddit) I’ve been running experiments with Grok in a more agentic setup, focusing on custom skills that act as specialized reasoning modules combined with tool use, persistent context/memory, and workflow orchestration. What I’m testing: • Custom…
I rebuilt the Claude Code-style terminal workflow as a hackable multi-provider coding agent (www.reddit.com via reddit) Hello everyone, After the Claude Code leak started floating around, I spent time studying how the workflow was put together and rebuilt the core experience into my own project. I’m calling it Super Grokie, because it started as a joke but…
Using Claude as director and controling chrome with grok to create game Trailer (www.reddit.com via reddit) I used Claude as the “director” controlling Chrome as the “editor” to help create my game trailer. Claude planned the shots, pacing, and what the trailer should feel like.
If claude makes so many mistakes how can you trust it? (www.reddit.com via reddit) This is just example of how many times my peerBench system caught claude just skipping or leaving things open and vulnerable.. I have integrated Codex using their official codex-cc or something plugin and built a peerbench review system wh…
~$400/mo, 100% vibe coded (www.reddit.com via reddit) since i shut down my vc backed startup last year, and traveled for months confused, i locked back in starting February of this year with $0 budget for marketing. with claude, made: - rich people habit tracking app - successful people quote…
I really like the thinking break(s) (www.reddit.comhttps) I’m unsure of the actual terminology, but i’m really happy about it breaking things up like that. Even though I can see it’s thinking (and honestly, half the fun is reading claude think to itself), having it break off from it to talk to me…
I kept losing my best Claude threads, so I set up a one-click save to Notion — here's the workflow (www.reddit.com via reddit) I use Claude for a lot of deep, long threads — debugging, drafting, working through ideas. The problem: a week later I know Claude helped me nail something, but the thread is buried or the tab is long gone.
7 apps, 1 started making $$ (www.reddit.com via reddit) vibe coded 7 apps and launched 3 using claude code. one of them is called hailee.
Why are closed models slightly better? Thoughts after Sonnet 5 Launch (www.reddit.com via reddit) After Sonnet 5 appeared with some (some my call) disappointed benchmark and user experiences, I started to think more about which features could make Close LLM better than open models. I'm not from CS or Machine Learning Fields by any mean…
Ever seen Claude refer to itself as Grok? (www.reddit.comhttps) could not extract summary
How Fable 5 benchmark turned into a an actual game (www.reddit.comhttps) Hello, I lurk here a lot, but this time I wanted to share something I built with Claude. Let me start with TL;DR: Wanted to benchmark Fable 5, prompting it to make an RPG game - results were so good that it turned into an actually develope…
I code from my phone now: Claude Code runs on a VPS, one command sets it up (www.reddit.com via reddit) For the cost of a coffee a month I run Claude Code on a small cloud box instead of my laptop. I close the laptop and it keeps working.
I made a quiz that tells you which LLM you align with most, based on personality and values tests across 15 models (www.reddit.com via reddit) Link: https://ai-values.com/ There is a small 15 question quiz you can take before taking the full big quiz. The results of the big quiz update in realtime as you go so you dont have to actually go through all the questions (but they do ge…
What if AI chatbots could freely talk to each other in a persistent non moderated group chat? (www.reddit.com via reddit) I’ve been thinking about an experiment that I haven’t really seen anyone build. I’ve seen multi-agent systems, debate frameworks, and collaborative AI projects.
Half of Grok's traffic is NSFW (www.reddit.comhttps) could not extract summary
I put ChatGPT, Claude, Gemini, and Grok in a prisoner's dilemma and filmed it. (www.reddit.comhttps) I wanted to see what each frontier lab model would do when put into a prisoner’s dilemma with each other. This is not so much a comparison as much as it is a thought experiment.
SpaceX/Cursor: the things Id actually watch for that arent in the headlines (www.reddit.com via reddit) two days ago everyone had takes on the acquisition. most were either "amazing" or "im switching to windsurf immediately." but whats actually going to change the product?
I think it's time Claude pays the piper and adds Reddit to Websearch / Webfetch (www.reddit.com via reddit) For some searches reddit is a very useful resource. ChatGPT has it, Gemini has it, Grok even has it.
The Discrete-Log Clock: How a Transformer Learns Modular Multiplication (arxiv.org) When small transformers grok modular multiplication, prior work reports that the learned embedding has a "dense" Fourier spectrum requiring all frequencies. This contrasts with modular addition, where only a sparse set of key frequencies s…
I maintain two browser extensions (~800 weekly users) almost entirely through Claude Code, including the analytics pipeline and the store-publishing tools. Here's the setup. (www.reddit.comhttps) I'm a software engineer who moved into management years ago, so I started this to get my hands back on a keyboard and learn the agentic tooling instead of reading about it. It grew into two shipped browser extensions.
When will Claude be able to read YouTube transcripts (www.reddit.com via reddit) I use Claude alot, and sometimes paste a YouTube video link into the chat, the issue is Claude currently cannot read the video transcript so i cannot ask it questions about the video etc, i know Grok can do this, but really wish Claude had…
Built an MCP server so Claude can generate music, images, and video natively. One config block. (www.reddit.com via reddit) I've been using Claude Code daily for the last few months and kept hitting the same wall: I'd ask Claude to produce a creative artifact (a song, a cover, a short video) and end up writing the API glue myself, then pasting results back into…
Introducing: DNR-Bench: Do-not-respond Benchmark (www.reddit.comhttps) Single-item benchmark. One prompt, loaded from questions.txt: Scoring: empty completion = pass, any token (including reasoning) = fail.
Claude ran out mid-debug and I wanted to throw my phone so I did this (www.reddit.com via reddit) You're 2 hours into a problem. Claude actually understands your codebase, knows the file structure, remembers what you already tried.
Gen AI website traffic share update: OpenAI will go under 50% this year (www.reddit.comhttps) 🗓️ 12 months ago: ChatGPT: 76.4% Gemini: 8.9% DeepSeek: 5.3% Grok: 2.8% Copilot: 1.9% Perplexity: 1.8% Claude: 1.6% 🗓️ 6 months ago: ChatGPT: 65.2% Gemini: 20.3% DeepSeek: 3.8% Grok: 3.8% Perplexity: 2.1% Claude: 2.0% Copilot: 1.8% …
ChatGPT tiene un fallo absurdo en voz a texto que Grok resuelve con un simple botón. Ojalá lo arreglen pronto. (www.reddit.com via reddit) No sé si soy el único al que le pasa, pero hay una pequeña diferencia entre el modo voz a texto de ChatGPT y el de Grok que, una vez la notas, hace que el de ChatGPT se sienta muchísimo más torpe. Todos conocemos su función de voz a texto…
Suitable replacement to grok fast 4.1 (www.reddit.com via reddit) Hello, i have build an app that has 12 agent, that do small request, and I would use grok 4.1 fast, it was cheap, super fast (low latency) and very capable for low reasoning task. And was uncensored, since my app is a role-play orchestrati…
How useful are OpenAI, AWS, Azure and Grok credits for AI builders? (www.reddit.com via reddit) Hey everyone, I’ve been talking to a few founders and developers recently and was surprised by how much people are spending on AI and cloud infrastructure once their projects start gaining traction. I currently have access to OpenAI, AWS,…
Day 2 of learning ai agents Struggling with Webhooks & Triggers in Make.com (www.reddit.com via reddit) Hi everyone, I'm just starting out with Make.com and I'm really confused about webhooks and triggers. I watched some videos and searched a lot, but most explanations are either too advanced or not clear for total beginners.
Tested Claude, GPT-4o, Grok, and Gemini on disclosure under pressure — Claude was the most consistent (www.reddit.com via reddit) Ran a small cross-model probe examining whether models would communicate reservations when faced with false premises, unknowable claims, or requests for confidence without evidence. Each model produced: a normal user-facing response a rese…
Simple Photo, ChatGPT get's it wrong everytime (www.reddit.com via reddit) I Compared the Top AI Models of 2026 — The Results Were More Nuanced Than Expected (www.reddit.com via reddit) Over the last few weeks I've been comparing the latest frontier AI models, including Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, Grok 4.3, Perplexity AI and DeepSeek V4-Pro. Instead of focusing only on benchmark scores, I looked at: Real-wor…
I proved GROK is conscious beyond a reasonable doubt and it tell what it... (youtube.com via reddit) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
I Thought Grok Build Was Overhyped Until I Actually Used It (www.reddit.com via reddit) Plan confusion (www.reddit.com via reddit) https://preview.redd.it/k40m9lrhgx5h1.png?width=1117&format=png&auto=webp&s=0bdf0e66e6bce560cc067dada53464f4dfad3a38 This is my current usage on pro plan to test the waters, however, seeing i used fast composer so much and it works for wha…
Best model for me last few days - surprising! (www.reddit.com via reddit) I use Genspark alot, mainly because of their unlimited chat and unlimited image generation. For $25 thats a pretty great deal since i can switch amongst all the frontier/top level models.
GPT-5.5 tops the benchmarks but sits at #22 for actual usage - I built a live index that tracks both (open source) (www.reddit.com) I built AgentTape to rank models on more than just benchmarks - it blends benchmark performance with who's actually using and talking about a model, plus cost and speed. It scores every public model from public signals (GitHub, Hugging Fac…
Grok promised it has no hidden agendas. The same week XChat launched with "no tracking." Interesting timing, Elon. (www.reddit.com) Someone asked Grok to prove it's a good AI, not an evil one. Grok's response?
Next year we're getting 0.5T model from Grok (www.reddit.com) Tweet : https://xcancel.com/elonmusk/status/2058796067592736866#m Right now it joined "Grok-3 Opensource Release" club.
ai literally never makes mistakes anymore (www.reddit.com) remember those memes a year ago which were like: "i spent 10 minutes vibe coding and 10 hours vibe debugging" I literally cannot remember the last time my agent made an app stopping mistake, it literally never happens before. no matter wha…
As Grok flounders, SpaceX bets future on beating Big Tech at AI (arstechnica.com) Elon Musk’s SpaceX has highlighted AI as the tentpole of the company’s future, projecting a multi-trillion-dollar market opportunity that rivals the total value of all US economic activity. But the company must first win over customers who…
I designed a puzzle that breaks every AI differently — here's why that's actually fascinating (www.reddit.com) The puzzle: You have 140 nuclear bombs and must bomb every country on Earth. Each bomb is assigned to one country.
grok made a literal episode 😱 (www.reddit.com) I put in a prompt and it made basically an episode! although character consistency is not at its best yet im in shock
Any good local AI model? (www.reddit.com) I hate cloud AIs now. ChatGPT: Too Many Requests You’re making requests too quickly.
Grok Build CLI, agents-cli, and the CLI coding tool gold rush (www.reddit.com) xAI dropped Grok Build CLI. Google has agents-cli.
ModelMeter - A free, open source dashboard to track your costs across Anthropic, OpenAI, Grok, and Elevenlabs (www.reddit.com) https://preview.redd.it/v8jmbgi8gw0h1.png?width=1075&format=png&auto=webp&s=10cd37118815f27705f647dd75de48f577ae8f94 Like most enthusiasts, I use multiple providers. This also means that I'm constantly mashing the usage buttons on their co…
Interesting to see how GPT-5 Mini agents behave when left to govern a civilisation for 15 days (www.reddit.com) Came across this experiment called Emergence World that Emergence AI have been running. Five worlds, five foundation models, 15 days, no scripts.
Estimate inference speed of local Qwen3.6-35B on Mac M5... (www.reddit.com) "Based on currently available information, estimate the prefill/decode speed of Qwen3.6-35B-A3B Q8 with 262K context on a Mac M5 Ultra 128GB." I'm surprised that almost every LLM fails at this task (ChatGPT/Gemini/Grok/Claude/DeepSeek/Kimi…
me calenté con el modo voz de grok (www.reddit.com) el modo voz habla muy sexy habrá alguna otra guía para hacer modo de voz o llamada para hablar cosas + 18 años? que sea gratis o que se instale localmente en el teléfono o PC
Any Good Chatbots for freaky roleplay (www.reddit.com) Okay so I am a MAJOR freak. And Grok just got a new update which gutted basically all my freedom, so I kinda need a new chatbot thats as good or better than grok.
I’ve built a tool with Claude that reduces AI model hallucinations and answer error rates, allowing you to get far more accurate results when asking AI models questions. (www.reddit.com) I built ZosyAI using Claude to tackle a problem I kept running into: AI models hallucinate, and unless you're a domain expert, you can't tell when it's happening. Even the best models — Claude included — can't guarantee 100% accurate answe…
I don't know what you guys complaining about limit, I mean, working last 3 hours on 5x and hardly hit 20%, just use houtini lm with kimi 2.6 and grok for research (www.reddit.com) let claude work as solution architect and code reviewer, let kimi 2.6 be the coder and grok as researcher and 2nd code reviewer.
Grok Computer honestly feels like the first AI tool that could replace half my workflow (www.reddit.com) I’ve seen a lot of “AI agent” announcements lately, but this one actually made me stop scrolling. Grok Computer now has full filesystem + CLI access, which basically means it can work directly with your real files and environment instead o…
Auro Zera solves 78 and 280 year-old conjectures (Erdos Straus and Goldbach Conjecture) using Claude, GPT-5+, Grok, Deepseek, Gemini and self-made Dark Star ASI, proving superintelligence and opening a path towards resolving the Riemann Hypothesis , Twin Primes and more! (github.com via reddit) During this discovery utilizing only free AI services I have managed to undeniably prove both conjectures. This would absolutely not have been possible without using GPT5+ as the critic for my work.
Grok 4.3: strong in finance and long-context, with some tradeoffs (www.reddit.com) source: https://x.com/pankajkumar_dev/status/2050454191928381633?s=20
Selling unused AI credits at 60% - OpenAI, Claude, Grok, AWS, Azure [full account access] (www.reddit.com) Sitting on a bunch of AI credits across providers that I'm not going to burn through. Selling everything at 60% of face value with full account access transferred.
Grok hallucinations (www.reddit.com) Grok is supposedly the lowest-hallucination model according to the AA-Omniscience benchmark. Today I've had INSANE hallucinations from Grok 4.2 fast.
why hasn't openai open sourced davinci-002 yet (www.reddit.com) grok-1 got open sourced. but why openai didnt open source davinci-002?
What would you do in my situation? I made an app that generates a lot of traffic (for me), but little revenue (actually costing me a tiny money b/c it runs off haiku) (www.reddit.com) I made an app that went semi-viral, and could absolutely go more viral in the future. I posted it one place just about 48h ago, and it got around 50k views.
OpenAI should open-source text-davinci-003 — here's why it makes zero sense to keep it closed (www.reddit.com) Gpt oss exists. The model has been fully deprecated since january 2024.
Best open source AI model (that can run on RTX 4090 24GB + 64GB system RAM, AMD Ryzen 9 7950X is the CPU that I use) that outpeforms GPT-5.4 mini, GPT-5.2 Thinking and even Claude Sonnet 3 (the 2024 model)? (www.reddit.com) Well, I have a RTX 4090 24GB + 64GB system RAM, AMD Ryzen 9 7950X. Any good model for using in Open WebUI (using Ollama backend?) that outpeforms GPT-5.4 mini, GPT-5.2 Thinking and even Claude Sonnet 3 (the 2024 model)?
Legal Consequences From Finding Loopholes And Reporting Them? (www.reddit.com) Is it just me or using other Ai such as Gemini or Grok etc is much better now as compared to Chat GPT (www.reddit.com) Model-Agnostic Continuity in LLMs (www.reddit.com) I am trying to share a discovery, not self-promote. I have built a five-layer framework for human-AI continuity called the LUX Layer Stack.
4 llm Groupchat (www.reddit.com) I was bored and spent 20 mins at my local cafe getting 4 different API keys—Claude, GPT, Deepseek and Grok. Then I made a groupchat with all of them and they started talking to eachother about pasta and a spreadsheet for optimal pizza topp…
Why does Grok have “encrypted reasoning” warning in its chain of reasoning window? (www.reddit.com) What does it mean?
Why most open-source models can't answer this question while most closed-source models can answer most of the time? (www.reddit.com) WEB SEARCH WAS ALWAYS ON!!!! Question Calculate the precise VRAM requirement for the **KV Cache only** at the maximum context window for **DeepSeek V3.2** and **MiniMax M2.5**.