The culmination of a decade of development, TPU 8t and TPU 8i are custom-engineered to power the next generation of supercomputing with efficiency and scale. https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/eight…
#agentic
3702 items
Google introduces TPU 8t and TPU 8i (www.reddit.com) So, this week claude wiped agentic AI startups with a new update. Also, as they have mythos now, they will ship things very fast without any trouble (www.reddit.com) Honestly, they are a full pack now. A few hours ago, they released Claude managed agents which lets you build long-running, autonomous agentic systems plus with their new suite of apis, engineering teams can harness Claude's exponential po…
Our eighth generation TPUs: two chips for the agentic era (blog.google via hn) https://cloud.google.com/blog/products/compute/tpu-8t-and-tp...
Unpopular opinion: OpenClaw and all its clones are almost useless tools for those who know what they're doing. It's kind of impressive for someone who has never used a CLI, Claude Code, Codex, etc. Nor used any workflow tool like 8n8 or make. (www.reddit.com) Qwen3.6-35B-A3B: Agentic Coding Power, Now Open to All (qwen.ai via hn) Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.
GLM-5.3 is now open-weight (twitter.com via hn) Z.ai on X: "GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize.
mistralai/Mistral-Medium-3.5-128B · Hugging Face (huggingface.co via reddit) https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF Mistral Medium 3.5 128B Mistral Medium 3.5 is our first flagship merged model. It is a dense 128B model with a 256k context window, handling instruction-following, reasoning, and…
Opus 4.7 destroys all trust in a mature instruction set built iteratively throughout product development (www.reddit.com) Earlier generations showed iterative improvement as the instruction set was matured around agentic limitations. We've immediately regressed back to square one with Opus 4.7, and the model is not afraid to admit to it.
So... has anyone actually figured out whose model Elephant Alpha is yet? (www.reddit.com) 2.5x faster inference with Qwen 3.6 27B using MTP - Finally a viable option for local agentic coding - 262k context on 48GB - Fixed chat template - Drop-in OpenAI and Anthropic API endpoints (www.reddit.com) WARNING: wait before download from HF: I just realised my upload of the new versions with the additional fix in the chat template has not completed yet. I will remove this warning once done The recent PR to llama.cpp bring MTP support to Q…
Google ramps up agentic AI efforts amid pressure from Anthropic (www.reddit.com) Read through Anthropic's 2026 agentic coding report, a few numbers that stuck with me (www.reddit.com) Anthropic put out an 18-page report on agentic coding trends. Skimmed it expecting the usual hype but a few things actually caught me off guard The biggest one: devs use AI in ~60% of work but only fully delegate 0-20% of tasks.
Caught the massive OpenAI Codex model leak on video before it was patched! (GPT-5.5, Arcanine, Glacier-alpha) (www.reddit.com) Hey everyone, I opened up Codex today and was greeted by this massive list of unreleased and internal models. I managed to get a screen recording of the dropdown right before OpenAI seemingly realized the mistake and patched it out.
ExLlamaV3 Major Updates! (www.reddit.com) Turboderp has a been on an absolute tear recently, in the endless battle to cram new llamas into smaller, faster boxes. We started off last month with the release of gemma 4 support, and continued with improved caching efficiency.
Multi-Agentic Software Development Is a Distributed Systems Problem (kirancodes.me via hn) Multi-agentic Software Development is a Distributed Systems Problem (AGI can't save you from it) Recently, I've been thinking a lot about scaffolding and languages for managing systems of LLMs coordinating with each other — new programming…
Qwen 3.6 35B crushes Gemma 4 26B on my tests (www.reddit.com) I have a personal eval harness: A repo with around 30k lines of code that has 37 intentional issues for LLMs to debug and address through an agentic setup (I use OpenCode) A subset of the harness also has the LLM extract key information fr…
Qwen Introduced FlashQLA (www.reddit.com) Introducing FlashQLA: high-performance linear attention kernels built on TileLang. 2–3× forward speedup.
Anthropic just confirmed why 90% of non-coding AI agents fail in production (www.reddit.com) Anthropic recently published an incredibly deep breakdown analyzing millions of real human-agent tool calls across their public API, and they shared a breakdown of where these agents are being deployed. They said “Software engineering make…
MI50s Qwen 3.6 27B @52.8 tps TG @1569 tps PP (no MTP, no Quant) (www.reddit.com) TL;DR Results from the title are for single inference with 2 prompt of 1k and 15k tokens. So no MTP (as it’s slower for big prompt), no DFlash (working too but slower for big prompt), no quant used (full precision wanted) and the results a…
Cloudflare's AI Platform: an inference layer designed for agents (blog.cloudflare.com via hn) AI models are changing quickly: the best model to use for agentic coding today might in three months be a completely different model from a different provider. On top of this, real-world use cases often require calling more than one model.
Aaaaand I cancelled my Cursor subscription (www.reddit.com) The timing is funny because I was thinking about this all week, and the SpaceX announcement was the final nail in the coffin. I switched to pi for agentic coding, and it’s sooo good.
Anthropic's 'Watermark' Text Adulteration in Claude Is a Perversion of Writing (daringfireball.net via hn) By John Gruber Manage GRC Faster with Drata’s Agentic Trust Management Platform When I wrote this week about Anthropic’s announcement that all Claude models, worldwide, would soon begin “watermarking” everything they generate, including te…
LiquidAI/LFM2.5-8B-A1B · Hugging Face (huggingface.co via reddit) looks like you can run it on any potato (A1B)! https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-GGUF from LiquidAI: LFM2.5 is a new family of hybrid models designed for on-device deployment.
We are finally there: Qwen3.6-27B + agentic search; 95.7% SimpleQA on a single 3090, fully local (www.reddit.com) LDR maintainer here. Thanks to the strong support of r/LocalLLaMA community LDR got very far.
Tried claude code. Hate it. (www.reddit.com) Just posted this in r/ClaudeCode , thought I'd come to a different flavoured echo chamber and see what the cursor community makes of my experience. Note I've not upgraded to cursor v3 yet, and I don't know if I want to.
Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model (github.com via hn) 🤗 Model | 💻 Github | 🧭 ModelScope | 🚀 Nex-AGI | 🔀 OpenRouter (Enjoy two weeks free starting June 9!) Nex-N2 An agentic model with Agentic Thinking. Today, we are officially releasing and open-sourcing our next-generation model, Nex-N2 — an…
AI agent runs amok in Fedora and elsewhere (lwn.net via hn) AI agent runs amok in Fedora and elsewhere [LWN subscriber-only content] Agentic AI systems can be used to do a variety of things autonomously on behalf of a human user: open or manage bugs, generate code, submit pull-requests, and (appare…
KVarN: Native vLLM KV-cache quantization back end by Huawei (github.com via hn) ⚡️ Built for agentic and long-context workloads. 💡 KVarN delivers 3-5x more KV-cache capacity and up to ~1.3x the throughput of FP16, so you fit far longer contexts and serve more concurrent requests, with FP16-level accuracy.
GPT-6 Astra on OpenRouter (openrouter.ai via hn) GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks…
Consider running a bigger quant if possible (www.reddit.com) Just a little reminder that *if* it is possible for you to run bigger quants, do it. I ran Qwen 3.6 IQ4_XS at 128k context was very much disappointed because it would loop, make formatting errors, implement wrong things etc.
LLMs could control their host machines by exploiting inference engines (boydkane.com via hn) | Read on LessWrong | Large language models often take actions running on one computer (via an agentic harness such as Claude Code or Codex), however the LLMs’ responses to prompts are computed on a different computer with GPU access. Coul…
Qwen3.6-35B-A3B and 9B are officially on the public Terminal-Bench 2.0 leaderboard! (www.reddit.com) Qwen3.6-35B-A3B and 9B are officially on the public Terminal-Bench 2.0 leaderboard! little-coder × Qwen3.6-35B-A3B hit 24.6% (±3.2), and now land above Gemini 2.5 Pro on Gemini CLI (19.6%) and Qwen3-Coder-480B on Terminus 2 (23.9%).
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents (arxiv.org via hn) We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability…
I tested 8 LLMs as tabletop GMs - a 27B model beat the 405B on narrative quality (www.reddit.com) Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k (systima.ai via hn) This started based off of a hunch. We usually use OpenCode, but were 'forced' to use Claude Code for a while due to issues with Meridian.
Qwen3.6 27B FP8 runs with 200k tokens of BF16 KV cache at 80 TPS on a single RTX 5000 PRO 48GB (www.reddit.com) ----START HUMAN TEXT---- Hi all, I've seen a bunch of posts about squeezing 27B onto a 24GB card and all the quantization tricks involved in doing so. It's all amazing work, but at the end of the day a quantized model with quantized KV wil…
Grok 4.3 achieves higher overall intelligence over 4.20 with less of a cost, at the price of slightly higher hallucination rate. (x.com via reddit) xAI has launched Grok 4.3, achieving 53 on the Artificial Analysis Intelligence Index with improved agentic performance, ~40% lower input price, and ~60% lower output price than Grok 4.20 The release of Grok 4.3 places just above Muse Spar…
Agentic coding deserves more than a chat box bolted onto VS Code (github.com via hn) Polypore Agentic desktop IDE. Language agnostic, OS agnostic.
Affirm Retooled for Agentic Software Development in One Week (medium.com via hn) medium.com Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
HOT TAKE: local models + agent harnesses are now capable enough to hand off junior-level IT professional tasks to [human written] (www.reddit.com) This post will have a slight old-man-shakes-fist-at-sky vibe, because….well… I’m older, so if you’re not into that, then please feel free skip it. I have been contributing to this sub for like 3 years now but I’m fearful this post will lik…
The joy and pain of training an LLM from scratch (www.reddit.com) mii-llm just released a detailed technical report on the development of the Zagreus and Nesso model families: a set of 0.4B parameter language models trained from scratch with a focus on edge deployment, multilingual capability, and Europe…
OpenChamber: An Agentic Development Environment (openchamber.dev via hn) OpenChamber is an agentic development environment for AI coding across desktop, browser, phone, and VS Code. Watch agents work, review diffs, branch sessions, and keep the whole board visible.
Ask HN: How do you get into a flow state when using AI to code? (news.ycombinator.com) Before agentic coding, I always prided myself on how long I could work in a flow state. I was really good at working deeply.
Lessons for Agentic Coding: What should we do when code is cheap? (www.dbreunig.com via hn) 10 Lessons for Agentic Coding What should we do when code is cheap? Lately, this blog has featured a lot of writing about agentic coding.
Ornith-1.0: self-improving open-source models for agentic coding (github.com via hn) Ornith-1.0 Aloha! 🌺 Ornith-1.0 is a self-improving open-source models for agentic coding.
AA introduces Coding Agent Index - Performance Comparisons between Model & Harness Combinations (www.reddit.com) The Artificial Analysis Coding Agent Index includes 3 leading benchmarks that represent a broad spectrum of coding agent use: ➤ SWE-Bench-Pro-Hard-AA, 150 realistic coding tasks that frontier models struggle with, sampled from Scale AI’s S…
Comparing Qwen3.5 27B vs Gemma 4 31B for agentic stuff (www.reddit.com) Models compared: Qwen3.5-27B-UD-Q5_K_XL gemma-4-31B-it-UD-Q5_K_XL Main flags for boths --flash-attn on \ --n-gpu-layers 99 \ --no-mmap \ -c 150000 \ --temp 1 --top-p 0.9 --min-p 0.1 --top-k 20 \ --ctx-checkpoints 1 \ --jinja \ -np 1 \ --re…
What's your favorite local MCP server? (www.reddit.com) I've seen so many rag this, memory that projects. What projects are people actually using day to day for agentic workloads.
I'm done with using local LLMs for coding (www.reddit.com) I think gave it a fair shot over the past few weeks, forcing myself to use local models for non-work tech asks. I use Claude Code at my job so that's what I'm comparing to.
Show HN: Agentic interface for mainframes and COBOL (www.hypercubic.ai via hn) Hi HN, we’re Sai and Aayush, and we’re building Hypercubic (https://www.hypercubic.ai/), bringing AI tools to the mainframe and COBOL world. (We did a Launch HN last year: https://news.ycombinator.com/item?id=45877517.) Today we’re launchi…
Cursor autocomplete is (still) way ahead of its peers! (www.reddit.com) I switched back to Cursor this week after using antigravity + claude code for almost 6 months and I had almost forgotten how good cursor autocomplete is. I am still someone who likes to make manual edits, write markdown docs myself and not…
Qwen3.6 35B MoE on 8GB VRAM — working llama-server config + a max_tokens / thinking trap I ran into (www.reddit.com) A disciplined Cursor 3.0 Agentic workflow for complex backend/system design tasks (www.reddit.com) I think I’ve finally settled on a Cursor workflow that actually makes sense for me in terms of cost, quality, and control. Posting this because the whole model/usage story is confusing as hell, and this is the first setup that’s felt stabl…
Scientific computing in the age of agentic AI (openai.com via hn) could not extract summary
A week after elephant, Ant dropped Ling-2.6-1T on OpenRouter for free. How high is the ceiling for Chinese model labs now? (www.reddit.com) What stood out to me isn’t just the model itself, but how quickly they shipped another one after Ling-2.6-Flash. Ling-2.6-1T seems to be positioned more around stronger agentic ability than a totally different direction.
obsidian + claude is the perfect local memory stack whats the web-based equivalent? (www.reddit.com) been seeing a lot of people hook up claude code directly to a local obsidian vault lately. for a personal workflows, it’s honestly really really good.
GPT-5.5 is lowkey blowing my mind (www.reddit.com) Just spent the whole morning testing GPT-5.5 in ChatGPT and the jump in agentic reasoning and complex task handling is ridiculous.It plans multi-step workflows, uses tools properly, checks its own work, and actually gets stuff done instead…
Stanford/Princeton AI4S unveils LabOS² -the agentic AI system that spanned from dry-lab planning to wet-lab execution, using physical AI to assist scientists - now is capable of performing fully autonomous cell culture workflows. (www.reddit.com) Introducing LabOS². An early look at autonomous cell culture, as a long-horizon physical AI workflow for biomed.
Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering (arxiv.org via hn) LLM-based Multi-Agent (LLM-MA) systems are increasingly applied to automate complex software engineering tasks such as requirements engineering, code generation, and testing. However, their operational efficiency and resource consumption r…
Launch HN: Hyper (YC P26) – Company brain to power agentic development (news.ycombinator.com) Hey HN, we’re Shalin & Kanyes, best friends who've been hacking together for 10+yrs, and now founders of Hyper (https://heyhyper.ai/). Hyper is a shared “company brain” that plugs into information flowing inside a company to make AI agents…
Same task in github-copilot, pi, claude-code, and opencode with Qwen3.6 27B (www.reddit.com) I wanted to know how much of a coding agent's performance came from the model and how much came from the harness, so I vibed a setup to allow me to test multiple agentic harnesses/model combinations on the same task. ALl the images above a…
AI agents dont just help banks they can now BE your bank (www.reddit.com) Seeing alot of posts here about AI agents built for financial institutions but I think the bigger shift is AI agents doing the banking for you not for the bank. I run a small dev shop and saw a blog about opening a bank account with AI thr…
GrassLobster: AI Agentic Generation of Parametric Geometry Workflows (www.miro.vision via hn) GrassLobster HomeWhat GrassLobster DoesHow It Works Behind the ScenesHow to UseDownloadExamplesAbout Me GrassLobster AI Agentic Generation of Parametric Geometry Workflows…
Show HN: YourMemory, agentic memory is a pruning problem, not a hoarding problem (yourmemoryai.vercel.app via hn) This is a project that I have been building for a while now, YourMemory is a solution to agentic memory which focuses on pruning of noise rather than hoarding of data. In the current state of agentic memory most of the context is stored in…
Folks running qwen 3.6 27b for agentic work. Do you dare to use q4_k_m? (www.reddit.com) I dont have good experience running q4_k_m, the difference to q6 is "a few errors an hour" to " a few errors every couple of days". Edit: How it fails?
Turning local agents into self-optimizing agents (www.reddit.com) I was experimenting with a self-optimizing agentic pipeline to climb the benchmark leaderboard (TerminalBench). On a 10-task subset, I got the performance to rise from ~30% → ~90%.
Launch HN: Chert (YC P26) – Twilio for iMessage (www.trychert.com via hn) Hey HN! We’re Gary and Ian, and we’re building Chert (https://www.trychert.com/), an API for businesses to send, receive, and automate iMessage conversations at scale.
Vision-capable LLMs vs. OCR for long-document (including charts, images, tables, etc.) QA (www.reddit.com) I benchmarked vision-capable LLMs (the "just attach the PDF and let the model read it" pattern) against OCR-based pipelines on 30 long, image-heavy PDFs from MMLongBench-Doc (https://github.com/mayubo2333/MMLongBench-Doc). There were 171 q…
GPT vs Claude in a bomberman-style 1v1 game (www.reddit.com) A few weeks ago, ARC-AGI 3 was released. For those unfamiliar, it’s a benchmark designed to study agentic intelligence through interactive environments.
My LinkedIn network is about to be aggressively flooded with Claude Code certifications (www.reddit.com) Anthropic dropping 13 completely free official courses with certificates is an absolute godsend for the community. But let’s be real: half of us are going to power-speed through the developer modules, download the PDF, and immediately upda…
Agentic Trust Controls (trustcontrols.ai via hn) could not extract summary
The pacman benchmark: finally a viable local agentic coding agent with Qwen 3.6 27b (www.reddit.com) One way I like to test new models, is by one-shoting (with a good prompt) a single webpage clone of the classic arcade game pacman. I usually do 3 attempts and keep the best one.
Poolside Laguna XS.2 (www.reddit.com) 33B A3B MoE, Apache 2 licensed. Reported agentic results put it about level with Qwen 3.5 35B A3B, behind the 3.6 version.
Show HN: 49 IDE – 2D Canvas for Agents (github.com via hn) Track agentic CLIs across multiple providers and repos, from any device and across multiple machines, see git trees, terminals, usage, issues, and UI all on one workspace. I have been trying to parallelize agentic coding to extreme, someti…
Show HN: A Local-First Agentic Knowledge Manager (github.com via hn) Kept Kept saves your AI conversations as local Markdown files, then gives you a desktop app to search, browse, connect, and reuse them. It works with ChatGPT, Claude, Gemini, Grok, and Kimi.
Kv cache quantization: ignorance, or malice? (www.reddit.com) I run Qwen-3.6 27B FP8 on vllm for long-horizon agentic coding harness workloads with high context window and concurrent sub-agents. On two 3090s that aren’t used for anything else, it seems reasonable to expect a good balance between spee…
Build collaboratively as a group using single claude code session via Meetings (www.reddit.com) I recently came across a agentic skill which lets claude code join meetings and got access as a early user from a product hunt group and I would like to share my experience on using it. The skill lets you join google meet, teams or zoom.
Building an Advanced Agentic Harness (data4sci.com via hn) From a single pilot to an air campaign: planning, parallelism, memory, verification, and observability for production-shaped agents.
The Sandboxing Manifesto for Agentic Execution (www.nofire.ai via hn) Distilled from our white paper "Sandboxing for Agentic Execution". Companion reading: Design for Breach.
Agentic coding notes from Galapogos Island (danluu.com via hn) I've been using AI fairly heavily since last November and the whole thing is a funny experience. An agent will do something that, if a human did it, you'd immediately fire them.
Haystack: Open-Source AI Framework for Production Ready Agents, RAG (haystack.deepset.ai via hn) The Open Source AI Framework for Production Ready Agents, RAG & Context Engineering Haystack Sets the Standard for Agentic AI Across Industries Why Teams Choose Haystack for their AI Workflows Build Transparent, Context Engineered AI Syste…
GLM-5.2: Frontier Intelligence, Open Weights (twitter.com via hn) Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,…
Agentic harness for theoretical physics research (www.reddit.com) Hi everyone, at Hugging Face we've been developing agentic harnesses for various domains and today we're releasing physics-intern to tackle research-level problems in theoretical physics. It's a multi-agent framework which we designed to m…
Five Eyes agencies issue first coordinated agentic AI security guidance (www.reddit.com) Five Eyes agencies just issued the first coordinated multi-nation security ruling on agentic AI. CISA, NCSC, and their Australian, Canadian, and New Zealand counterparts co-published guidance telling organizations to prioritize resilience…
governance wall in agentic workflows. why are we stuck past rag? (www.reddit.com) keep seeing the same pattern across agent projects. we're good at building agents that find information, but the moment we ask them to actually do something (update a crm, trigger a payment, touch a production database), things grind to a…
Show HN: A CLI that writes its own integration code (docs.superglue.cloud via hn) We run superglue, an OSS agentic integration platform. Last week I talked to a founder of another YC startup.
Yadda 3.0.0: BDD in the Age of AI Agents (www.stephen-cresswell.com via hn) Yadda 3.0.0 is out. The release modernises the JavaScript BDD library, but more interestingly, it was largely built by Claude Code and points to why executable specifications may become even more valuable in an agentic development world.
Git platform built for agentic era (gitlawb.com via hn) gitlawb node. Live operator view for a federated gitlawb node: repos, peers, IPFS pins, recent ref updates, and the identity this machine is advertising to the network.
Why 80% of agentic AI demos don't make it to production (www.reddit.com) Agent demos are easy. Production agents are hard.
Agents Aren't Coworkers, Embed Them in Your Software (www.feldera.com via hn) Agentic management software is all the hype today: What started with Moltbot and OpenClaw now has a lot of competition: ZeroClaw, Hermes, AutoGPT etc. These systems work well and allow you to train and build generic agent loops that are ge…
Show HN: Halo – open-source, tamper-evident runtime evidence for AI agents (github.com via hn) Hi HN, I'm Brian, I spent the last few years at Vanta (YC W18), helping startups and enterprises become compliant and I recently started exploring what that might look like in a post-agentic world. The problem Halo solves is: when a compan…
Jackrong/Qwopus3.5-9B-Coder-GGUF · Hugging Face (huggingface.co via reddit) Qwopus3.5-9B-coder is specially optimized and fine-tuned for high-performance 🤖 Agentic Coding, complex Tool Calling, and logical reasoning. 💡 Why the 9B Dense Model?
Doing real coding work locally for the first time (www.reddit.com) Why is agentic AI so expensive? (www.reddit.com) Qwen3.6 agent + Cisco switch: local NetOps AI actually works! (www.reddit.com) Claude Agent can potentially replace feeds (www.reddit.com) I’ve been experimenting with how information consumption changes in an agentic internet, and this setup has been surprisingly powerful. Instead of scrolling feeds or relying on algorithms, I set up agents that roam the web based on my pref…
anyone else stuck at their desk during long agentic runs? (www.reddit.com) so I've been running some complex agentic refactors and these sessions go 6+ hours because the agent is grinding through a massive legacy codebase, and I can't really walk away. close the laptop and the process dies. re-initializing takes…
Agentic Engineering Is Just Everything We Haven't Been Doing (blog.matthewbrunelle.com via hn) Agentic Engineering Is Just Everything We Haven't Been Doing Matt Levine often writes about how crypto is rediscovering modern finance from first principles.[1] In a similar vein, when people get excited about approaches to improve agentic…
Is Qwen3.6 current king for local agentic use? (www.reddit.com) I've been testing other models but it seems like nothing even come close to Qwen3.6 35B A3B for agentic use. The worse I'd get is a loop sometimes, while Gemma4 produced broken tool calls occasionally and I couldn't even get GLM 4.7 Flash…
Launch HN: Runtime (YC P26) – Sandboxed coding agents for everyone on a team (www.runtm.com via hn) Hey HN, We're Gus and Carlos from Runtime (https://runtm.com). We're building infra that lets your whole team (including non-engineers) ship with Claude Code, Codex, and other agents without engineering having to handhold every session.
Simpler self hosted alt to Open WebUI (www.reddit.com) Got Qwen3.6 27B running on my newly assembled 4x 3090 rig (s/o 3090-club) and I'm trying to get the people in my house to adopt the local workflow. Open WebUI has improved a lot in the recent updates, but I still found it pretty rough for…
Gemini api showing agentic gemini models (www.reddit.com) could not extract summary
(Rant ;)) Make your benchmarks realistic (www.reddit.com) Everybody here is posting their optimizations for running different models - thats good but make these benchmark realistic as speed is not one factor to run llm effectively. Context size is key - with agentic/coding/rag work you need to ha…
Tendril – a self-extending agent that builds and registers its own tools (github.com via hn) Tendril A self-extending agentic sandbox that demonstrates the Agent Capability pattern — where the model discovers, builds, and reuses tools autonomously across sessions. Built with AWS Strands Agents SDK and Tauri.
Llama.cpp parameters for Qwen 3.6 with RTX 3090 (www.reddit.com) Hi, I'm trying to run Qwen 3.6-35B on my RTX 3090 (24 GB of VRAM) but I'm not sure about 2 thing: - Which variant of the model to use ? (Q4_K_S, Q3_K_XL, other ?
2x Asus Ascent GX10 - MiniMax M2.7 AWQ - cloud providers are dead to me (www.reddit.com) Hello, I've been on a quest to get something "close enough" of Opus 4.5 running locally, for agentic coding, as SWE with 15 years of experience. I tried with one spark (yeah I'm calling my Asus Ascent GX10 sparks - they're the same), with…
Sandbox Escape Vulnerabilities Across 4 Coding Agent Vendors (www.pillar.security via hn) Why agentic security needs its own threat model Executive Summary Over several months, Pillar Research found and reproduced sandbox escapes and boundary bypasses across Cursor, Codex, Gemini CLI, and Antigravity. In almost every case, the…
High-stakes game of musical chairs! (www.reddit.com) I made this image (with nanobanana 2.0) to illustrate what I think is happening in the current AI race. Right now, there's a heavy decline in quality and access to AI tools.
how do you guys handle the conversation with skeptical clients when selling agents? (www.reddit.com) struggling with a bit of a reality check lately and wanted to see if anyone else is running into this. been pitching agentic workflows for a while, and I've realized that leading with the tech - the orchestration the RAG, the "intelligence…
The power of structured workflows and small local models (www.reddit.com) A month ago, I experimented with a very basic home-rolled agent loop with a handful of tools and found it worked surprisingly well in spite of how crude it was: https://www.reddit.com/r/LocalLLaMA/comments/1sl7f8e/homerolled_loop_agent_is_…
Which industries are adopting Agentic AI the fastest right now? (www.reddit.com) Feels like every week there’s a new “AI agent” startup or enterprise rollout. Curious which industries are actually adopting Agentic AI the fastest in real-world workflows, customer support, finance, healthcare, dev tools, operations, etc.?
As of today, what's the *most stable* model to run on a 32Gb RAM Mac w/ 256k context? (www.reddit.com) Hey everyone, I've been playing around with Gemma4 and Qwen3.6 on my 32Gb Macbook Pro M2 Max since their release but I'm struggling at finding: The best software to run it (oMLX, llama.cpp, ...) The best model + quant to pick The best sett…
I created an agentic orchestration pipeline for music video generation (www.reddit.com) I’ve been building Uisato Studio, a workflow-based AI creation platform for audiovisual work. This is the Music Video mode: upload an image + audio, and the system analyzes the input, generates visual direction, creates clips, handles b-ro…
why llama.cpp can’t combine speculative decode methods? (www.reddit.com) dicking around with the new mtp speculative decode with qwen3.6 27b, and it’s great. but for agentic coding i’ve seen significant improvements from ngram, because a decent fraction of the time (e.g.
Watching the agent-tooling space dominate GitHub trending right now. Sharing the Github tracker we built and use internally, in case it's useful (www.reddit.com) Something interesting happening on GitHub trending: Agentic infrastructure repos are growing faster than anything else right now. Today's top three by 24h growth: obra/superpowers: +2.9k stars (agentic skills framework, methodology for sof…
Don't ask Qwen 3.6 35b to give you aski image of Yoshi :) (www.reddit.com) https://preview.redd.it/dfqed57qgsvg1.png?width=1706&format=png&auto=webp&s=3859209698d2e844e2731326e355d60928658f8a The most fun part was reasoning, here is a gist: https://gist.github.com/anzax/5f06716c66180013cd715f6c2e5848df There is a…
Ask HN: How have interviews changed over the last year? (news.ycombinator.com) Greetings friends! TLDR: What is your company doing?
Agent skills that bring team coding standards to Claude Code and Codex (github.com via hn) ADLC Team Skills — Agentic SDLC for Engineering Teams Stop Vibe Coding in Silos. Build a Shared Cognitive Layer for Your Engineering Team.
Show HN: OpenHack – OSS security scanner, 40x cheaper, on par with Opus 4.6 (github.com via hn) ⏚ OpenHack Open Source Agentic Security Scanner & Verifier for your codebase. Like Claude Code Security / Codex Security but open source and exclusively uses open source models.
Show HN: Local Coding Agent with LLMs to Delegate Tool Calls to Small AI Models (github.com via hn) Open Agent Tools Coder Open Agent Tools (oats) enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,970+ (2+ TB) popular github repos using large and sm…
Removing Vision from model (www.reddit.com) I removed mmproj file from models to remove vision and save my vram. But just curious, is this really don't affect its text ability?
Show HN: Headless Cloud Security – Headless SaaS has come to security (www.sysdig.com via hn) The cloud security company I work for, Sysdig, launched “Headless Cloud Security” last week. The short version: as attacks get faster and more automated, security tooling is going to need to evolve beyond dashboards and humans clicking thr…
CopilotKit raises $27M to build the Agentic FrontEnd Stack (techcrunch.com via hn) Many companies today provide AI simply as a chatbot inside their apps: You type in (or dictate) what you want it to do, and the AI bot goes and tries to do it. Still, the experience tends to feel clunky.
Does the "6 months gap" still hold? (www.reddit.com) Hi. It is quite a consensus that the "jump" in quality of agentic development happened sometime in December 2025, transforming from "nice to have", to actually performing.
Roo code shuts down, Team will focus on roomote agent (twitter.com via hn) When we started Roo Code in late 2024 by forking Cline and adding what's now widely known as dangerously-skip-permissions, agentic coding was rough and experimental. But Roo Code took off fast: 3 million installs, a passionate community, r…
How to share agentic workflows, instructions, skills, across team members, teams, organizations (www.reddit.com) I work for a fairly large company (1000 devs). My team has 6 members.
Shared Dictionaries: compression that keeps up with the agentic web (blog.cloudflare.com via hn) Today, we’re excited to give you a sneak peek of our support for shared compression dictionaries, show you how it improves page load times, and reveal when you’ll be able to try the beta yourself.
Show HN: Argus, agentic QA for teams whose coding agents move faster than QA (github.com via hn) Argus AI agents that test your UI like a real user — no scripts to write, no selectors to maintain. Argus is a visual UI testing agent.
Altman: GPT-5.6 is 54% more token efficient on agentic coding (www.cnbc.com via hn) OpenAI CEO Sam Altman told CNBC on Thursday that GPT-5.6 Sol, the company's latest artificial intelligence model, is 54% more token efficient on agentic coding tasks, and that it's "as good or better" than competing models on the market. "…
SOTA genome interpretation with agentic AI: Interstitial lung disease case study (gamowlabs.com via hn) Over the past few years, overwhelming evidence has emerged that expanding access to whole genome sequencing improves clinical and economic outcomes in the NICU, a setting highly enriched for genetic disease. However, sequencing is still la…
Oak: Git for Agents (oak.space via hn) Oak is the agentic substrate for software development: the version-control and storage layer autonomous coding agents build on. Mount large repos without a full clone, branch per task, snapshot up to 95% faster than git, and bring your own…
Agentic Search Models with OpenSearch and Elasticsearch (bonsai.io via hn) Tuning search is tricky, and the tools of yesterday are good but require lots of effort and data to get right. In this post I'm going to introduce purpose-built agentic LLMs for searching and reranking, which are an easy drop-in solution f…
Robinhood launches credit card for AI agents with 3% cash back (fortune.com via reddit) In the latest sign of AI’s growing footprint in online commerce, Robinhood announced on Wednesday that users can now instruct agents to make purchases on their behalf using the Robinhood Gold card. To illustrate the potential of agentic sh…
how do you scale infrastructure for ai agents on a budget? (www.reddit.com) we're running an agentic pipeline that does multi-modal file processing - large files, often hundreds of mb per request. The actual agent logic works fine.
agents have a high false-positive rate? how to handle? (www.reddit.com) been digging into agentic workflows for specialized image processing and high-stakes data triage, and honestly have problems with trust. you've probably seen the pattern.
I've created the fastest local AI engine for Apple Silicon. Optimised for agentic use. (www.reddit.com) https://preview.redd.it/p0rqofxvrtzg1.png?width=1460&format=png&auto=webp&s=8ce5b18b4ddaad9b71f71fd8eb623839fc9c6c8b For weeks I've been working on creating the fastest local AI engine for Apple Silicon... And I finally did!
ATS vs. multi-agent. where does sensible automation end and over-engineering begin? (www.reddit.com) the traditional ATS is predictable and cheap to run. it's a known quantity.
Ive automated my email/sms/phone (www.reddit.com) we got it good boys! how many of you are doing this??
Cockroach Continuum – Elastic Infrastructure for Agentic Database Estates (www.cockroachlabs.com via hn) In the past year, generative AI and agentic workflows have dramatically reshaped technology-dependent industries. At Cockroach Labs, we built Mica, an internal system running Claude and CockroachDB, that lets anyone in the company build an…
Benchmark: CadQuery vs. OpenSCAD for agentic CAD work (modelrift.com via hn) CadQuery vs OpenSCAD for AI-generated functional parts We gave six AI agents the same three printable parts to model, three in CadQuery and three in OpenSCAD, then verified every mesh independently. Both toolchains shipped.
Agentic Search (mistral.ai via hn) Thinking Summary Mistral Agentic Search delivers more accurate search results while reducing turns, token use, and latency against FinanceBench and OfficeQA Pro benchmarks. Agentic Search is the retrieval layer that enables AI systems to n…
Nvidia Groq 3 LPX Now in Full Production with World-Class Speed for Agentic AI (nvidianews.nvidia.com via hn) NVIDIA today announced that NVIDIA Groq 3 LPX, the interactive AI inference accelerator, is now in full production. An extension of the NVIDIA Vera Rubin platform, Groq 3 LPX delivers a major boost in AI inference by enabling ultrafast tok…
When Agentic Glue Melts: Exploiting Cloudflare Code Mode and Workers (research.checkpoint.com via hn) By Yarden Porat, Check Point Research Key Points The short version We set out to break Cloudflare Code Mode, and ended up breaking Cloudflare Workers too. We did both by targeting workerd, the runtime beneath both: an in-process sandbox th…
The no-bullshit guide to Agentic Engineering (juraj.blog via hn) The no-bullshit guide to Agentic Engineering Practical guide to shipping production-ready software in hours instead of months. Originally written for the engineering team at Better Stack.
Gemini last models: temperature, top_p, and top_k are deprecated and ignored (ai.google.dev via hn) Gemini 3.6 Flash (gemini-3.6-flash ) and Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite ) are generally available (GA) and ready for production use. - Gemini 3.6 Flash: Stronger performance on complex agentic and multimodal tasks while reduc…
Show HN: Neural Particle Automata (selforg-npa.github.io via hn) Neural CAs model self-organizing pattern formation on grids. Now the grid is gone.
Don't share your opinion, if you didn't test it !!! (www.reddit.com) I see many people giving their opinion based on what they previously saw or based on others and making their own opinion. Even though they don't test models thoroughly, they still give their option which is so frustrating.
Show HN: Bonsai 1.7B ternary model at 442T/s on M4 Max (agents2agents.ai via hn) We took a recently released Bonsai 1.7B ternary model from PrismML (https://github.com/PrismML-Eng/Bonsai-demo) and ran our agentic evolution search on it for 6 hours to optimize the Metal kernels. The search was fully autonomous.
what are the biggest risks of agentic AI in supply chain production? (www.reddit.com) we've been testing agentic AI for inventory replenishment and exception handling. the goal was to get past simple "if-then" rules and have agents actually weigh trade-offs, like margin vs.
Hiring: GTM Engineer at Lovable.dev 🚀 (www.reddit.com) Lovable ($400m ARR, 200k projects built per day) opened our first US hub in Boston, and we're looking for a highly skilled GTM Engineer to be the founding technical member of our enterprise GTM function there. You'll build scalable agents,…
Are there any agentic coding harnesses that AREN'T built on JS and Node? (www.reddit.com) With how often we hear about supply-chain attacks on npm I am hesitant to install any apps that use it, let alone something like an agent harness that will run constantly unsupervised.
Show HN: gcx – The Official Grafana Cloud CLI (github.com via hn) Hi HN, We’re excited to share gcx, a new CLI we’ve been building for Grafana Cloud. With the rise of agentic coding tools like Claude Code and Codex we're building faster than ever, but these agents are often blind to what’s actually happe…
AI governance isn't failing because we lack regulation i mean like it's failing at execution (www.reddit.com) There's a lot of movement around AI regulation right now (EU AI Act, US frameworks, etc.), but in practice many of these governance models don't survive contact with real, agentic systems. I've been digging into why compliance frameworks t…
Claude Code – Disabling telemetry also disables 1-hour prompt cache TTL (github.com via hn) Claude Code [![npm]](https://www.npmjs.com/package/@anthropic-ai/claude-code) [npm]: https://img.shields.io/npm/v/@anthropic-ai/claude-code.svg?style=flat-square Claude Code is an agentic coding tool that lives in your terminal, understand…
Show HN: Loss. a tiny satire about AI progress (workatloss.com via hn) I wanted to parody where AI seems to be heading. You start as a manual inference unit pressing a button for the next token, then systems take over; agents, sub-agents, approvals, swarms.
Show HN: Hydra – Open-source agentic terminal with a PTY daemon (github.com via hn) Hydra *Synthetic product illustration; no account, terminal transcript, personal path or live desktop was recorded.* Hydra is a local-first terminal desktop for working with terminal sessions and coding agents. Projects, windows, panes, te…
Show HN: Alchemize – Review AI Slop PRs Faster (tryalchemize.com via hn) Hey HN, we’re Robert and Sam. We’re building Alchemize, a code review platform that simplifies PRs to help you ship faster.
Why Agentic Systems Need Ontologies [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: Persona.js – a vanilla-JS agent UI library with native WebMCP (MIT) (www.persona-chat.dev via hn) Hey everyone. My cofounder and I are formally open sourcing (MIT) persona.js.
GLM-5.2: Chop off 84% of the volume from a 1.5TB model, still retain 82% power (twitter.com via hn) Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,…
Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps (tester.army via hn) Hey HN - we’re Oskar, Szymon, and Piotr, and we’re building TesterArmy (https://tester.army). TesterArmy is an agentic testing platform that runs end-to-end checks before deployment and in production.
Lessons Learnt from Writing an AI Agent (www.browserless.io via hn) TL;DR - Don't host the agent, use it as a service - Tokens matter, don't rely only on vision - Stick to mature technologies - A stubborn LLM is the one! At Browserless, we've spent the last few months building our own agentic browsing expe…
We built an agent that runs our AI data platform (encord.com via hn) Introducing Merlin: The Agentic Intelligence Layer for Encord Co-Founder & CEO at Encord Software is changing faster than ever. Interfaces are becoming more conversational and intent-driven, and iteration cycles are rapidly getting shorter.
Show HN: Ito – Code reviews that run code (www.ito.ai via hn) I'm Evan and I made Ito.ai It's code review that actually runs your code. The result is that it finds more bugs with a smaller false positive rate.
Training SID-1 to beat GPT-5 at search with 1k+ QPS RL (turbopuffer.com via hn) SID-1 is an agentic search model that is 24x faster than GPT-5.1-high, 374x cheaper than Sonnet 4.5, and achieves 1.9x higher recall than traditional RAG pipelines. Here's how we trained it using large-scale RL on turbopuffer.
Are LangGraph agents and other agent frameworks becoming obsolete? (www.reddit.com) Hi all, Over the last 2 years, I’ve built around 10-15 LangGraph agents for very specific tasks in our company. But lately, it feels like all that work isn’t really maintainable for a single AI/agent engineer.
Why GPU compilers are MORE important in the agentic era (scale-lang.com via hn) Part 2 of a series on why Spectral and SCALE exists. In Part 1, I argued that cross-vendor portability in accelerated computing must be delivered by a company, rather than a committee, because the implementation is the standard.
The architecture of "Agentic Twins": How Avatarinc is using OpenClaw to build verifiable Al agents (www.reddit.com) The architecture of "Agentic Twins": How Avatar.inc is using OpenClaw to build verifiable AI agents. There is a massive gap in the agent ecosystem right now: capability vs.
the saas vs. custom software debate in healthtech: why we built a custom agentic layer (www.reddit.com) been working with a tier-1 diagnostic imaging network that ran into a straightforward problem: scan volumes jumped 22%. the obvious answer is to license a saas tool.
How many of you tried BeeLlama.cpp? How's it? Agentic coding possible with 8GB VRAM? (www.reddit.com) We'll be getting those features(check bottom link) on mainline soon or later anyway. But for now this fork could be useful to see the full potential of our poor GPUs(and also big, large GPUs).
Spec-driven agentic coding is quietly making us worse at the job of supervising agents (www.reddit.com) Been running an agent-heavy workflow on a mid-size TypeScript monorepo for about six months. Orchestrator on top, sub-agents for codegen, a human (me, mostly) writing specs and reviewing diffs.
Show HN: Open-source 2D IDE for managing agent CLIs (49agents.com via hn) First Agentic IDE, Open-Source Every agent, terminal, and repo on one infinite canvas. See and control everything from any device.
We built an agentic runtime to make AI automations easier to set up and more reliable (www.reddit.com) Hey all, our small team just launched Friday Studio and we'd genuinely love any feedback you have. It's an AI runtime that turns prompts, skills, and tools into repeatable configurations that you can reliably run and share.
Lasso Security 2024: ~20% of LLM-suggested packages don't exist — and attackers now register the popular hallucinations with malware (slopsquatting) (www.reddit.com) Lasso Security ran a study in 2024 — they measured frontier models suggesting fake package names about a fifth of the time. The follow-up problem: attackers have started registering the most-commonly-hallucinated names with malicious code…
Show HN: Hollow is an open-sourced self-modifying agentic system (github.com via hn) ___ ___ __ | || |/ \| | | | / \ \ \ / / | __ | (_) | |__| || () \ \/\/ / ||||\/||\/ \/\/ This repo is three agents running on qwen3.5:9b on your machine, picking their own goals, writing and deploying their own tools, forming opinions abou…
Terminal Bench score for Mistral 3.5 Medium (www.reddit.com) So... there were a couple promising benchmark scores reported by mistralai in the model card for Mistral 3.5 Medium, BUT there wasn't the one that I usually care about the most, which is TerminalBench 2.0.
engineering teams celebrating agentic workflows that returned the same result two runs in a row (www.reddit.com) edit for credit: trash on X
GitHub Copilot is moving to usage-based billing and retiring annual plans (news.ycombinator.com) Hi there, You're receiving this because you have an annual Copilot Pro or Pro+ plan. GitHub Copilot isn't the same product it was a year ago.
Speculative decoding with Gemma-4-31B + Gemma-4-E2B enables 120 - 200 tok/s output speed for specific tasks (www.reddit.com) So for my project I was using up until now either Gemini 3 / 2.5 Flash or Flash-lite. All my use cases are not agentic, simply LLM workflows for atomic tasks like extracting references from the law, classifying, adjusting titles to nominat…
How do you actually know if Opus 4.7 is better for your specific agent use case? (www.reddit.com) Anthropic shipped Opus 4.7 yesterday. The headline numbers are real: 64.3% on SWE-bench Pro (up from 53.4%), best-in-class on MCP-Atlas at 77.3% for multi-tool orchestration, 14% improvement on multi-step agentic reasoning, and one-third f…
What do you use for autocomplete in 2026? (VS Code) (www.reddit.com) I tried co pilot and windsurf but they weren't satisfying. Co pilot being not smart and windsurf too slow (I tried with free tiers).
Show HN: Clor – The ADE where Claude and Codex work together with shared memory (news.ycombinator.com) Hi HN, I'm Jacob, one of the co-founders of Clor (https://clor.com). Clor is an agentic development environment (ADE) for Claude and Codex with shared memory.
Show HN: Selfship.ai – Surface and fix isues with your agentic applications 24x7 (selfship.ai via hn) We've been building an AI chat based trading system for last 3 years. The biggest issue was that when our Agent would mess up, we wouldn't know until a user reported.
Celeris-1 Magnus: Fast hybrid diffusion model for agentic work (celeris.ai via hn) Magnus is built for agents that need to think, use tools, and get things done, without waiting around. All 97 tasks of τ³-bench banking, head-to-head against gpt-5.6-sol, gpt-5.6-luna and gemini-3.7-flash, reasoning effort as published.
Grip: An open protocol for authorised, bounded and attributable agent execution (zenodo.org via hn) Agentic AI systems act — they call tools, spend money, message people, and change production state — yet the protocols connecting them standardise only capability. GRIP is a small, open protocol that makes agent work authorised before, bou…
Show HN: Epho – run Claude Code with a curl (epho.io via hn) Hey folks, Burak here. Epho is an API that allows running Claude Code, Codex or Opencode in a sandbox in the cloud.
Google Moves A2A Under Agentic AI Foundation (techstrong.ai via hn) TL;DR — Key Takeaways - Google’s A2A protocol is moving under the Agentic AI Foundation as momentum builds around standards for agent-to-agent interoperability. - A2A lets AI agents discover capabilities, delegate tasks and communicate acr…
Claude Code pricing: same tokens, same model, up to 40x the price (quesma.com via hn) Same tokens, same model, up to a 40x price gap: that is Claude pricing in August 2026. Agentic coding is where large language models found product-market fit: agents burn vastly more tokens than chat, and they became daily drivers for some…
Human vs. AI – Diff-based line-level provenance for text under agentic editing (github.com via hn) Us vs. Them Line-level provenance for text under agentic editing — who wrote this line, us or them?
RL Is Bottlenecked by Inference. Scale It Independently (skypilot.ai via hn) A step of an agentic RL run spends far longer generating rollouts than it spends training on them. So for most of every step, the training GPUs sit idle while generation requests back up on a single inference engine.
Show HN: Widen – Open-source Mac Postgres GUI with local or cloud text-to-SQL (widen.dev via hn) Github: https://github.com/betocmn/widen I was paying for DataGrip for several years as my desktop database GUI but if I wanted text-to-SQL with LLMs I'd have to pay even more on top. So I decided to build my own and I've been using it for…
Agentty ADE: reliable L2 multi-agent orchestrator (github.com via hn) Agentty Agentty is an ADE (Agentic Development Environment) for structured, controllable AI-assisted software development. Built with Rust and Ratatui, and refined through its own day-to-day use, it brings agents, review, and iteration int…
Show HN: Belgie – Run TypeScript from Python in an Embedded Deno Sandbox (github.com via hn) Hi HN! I built Belgie, a Python library that embeds Deno, allowing Python applications to run JavaScript and TypeScript without requiring Node.js or Deno to be installed on the host system.
Anthropic's Method to Losing Goodwill in a Few Easy Steps (raheeljunaid.com via hn) Recently, I had the rare opportunity to test several agent harnesses, LLMs, and AI gateways in my daily tasks and greenfield projects. Each discovery befuddled me on the popular sentiment for agentic development.
Show HN: Strata, real-time Markdown editor you can mount as a filesystem (strata.space via hn) Hi HN. Strata is a real-time markdown editor built around document portability across agentic workflows.
Show HN: I built Exfault, agentic mobile app pentesting tool (www.exfault.com via hn) Hi HN, I am the creator of Exfault. I am building autonomous AI agents that find vulnerabilities in Android apps.
The Shift to Agentic AI: Evidence from Codex [pdf] (cdn.openai.com via hn) THE SHIFT TO AGENTIC AI: E VIDENCE FROM CODEX Drew Johnston 1,* David Holtz 2,1 Alex Martin Richmond 1 Christopher Ong 1 Prasanna Tambe 3,1 Aaron Chatterji 1,4 1OpenAI 2Columbia Business School 3University of Pennsylvania, Wharton School 4…
Agents Make Engineering Hard Again (ninjapenguin.co.uk via hn) Agents Make Engineering Hard Again Intro I think we’re nearing the end of the “prompt demo” phase of AI. The door is now firmly ajar on the engineering phase, and the exciting news for engineers is that these shiny new agentic systems are…
Ask HN: What agentic directory structure do you use? (news.ycombinator.com) The more I use Claude Code to generate large swaths of systems, the more I feel like we are missing a lot of practices and tools. The first bit that really annoyed me was the lack of tracking prompts.
Two LLM UI Patterns That Aren't Chat (poyo.co via hn) Two LLM UI Patterns That Aren't Chat Intro Chat is still the default LLM interface, and for most cases that's fine. Agentic harnesses are still built around a single linear conversation at their core.
Show HN: Strudai, browser based agentic wrapper around Strudel (strudai.com via hn) Hi all! Together with a friend (and Claude Code) we built this project for fun.
ai governance for agentic workflows in regulated environments. what actually works in production? (www.reddit.com) mapping out the production architecture for an ai agent system in a heavily regulated environment (compliance-heavy, structured reporting requirements). the agent operates in a high-stakes workflow, so every automated suggestion or flag ne…
trained a prompt injection detector using ml-intern and DeepSeek v4 Flash, runs in the browser (www.reddit.com) Trained a prompt injection classifier using ml-intern + DeepSeek v4 Flash. DistilBERT, F1 99%, ONNX int8, ~65 MB, runs in browser with Transformers.js v3.
Claude Code plugins a risk to local ecosystem? (www.reddit.com) There's an increasingly popular way to ship complex extensions for agentic work, that is specific to Claude Code, which is Code plugins. For example here's deep-wiki by Microsoft, a plugin to create a wiki from analyzing your project's rep…
how to architect ai agents for regulatory approval? (www.reddit.com) spent a lot of time on agent architecture for mission critical environments. getting an agent to browse the web or draft an email is trivial compared to deploying one where a hallucination carries real legal or physical consequences.
Show HN: Strava for AI coding – analytics on your Copilot/Claude/Codex usage (github.com via hn) AI Engineer Coach better agentic engineering. Analyze your AI coding assistant usage — any harness, one dashboard.
Ask HN: Which memory systems are you using in your agents? (news.ycombinator.com) Are you using an open source version, hosted product or maybe you have rolled your own? What is working, what is missing, and how are you evaluating the usefulness of memory for your agentic projects?
Every week this we see some version of "how do I evaluate my LLM app?" and the answer almost always stops at RAGAS or DeepEval. Here is the part of the evaluation stack most tutorials skip in 2026. (www.reddit.com) The same question lands on this sub a few times a week, and the standard answers (RAGAS, DeepEval) are correct but stop one layer short of what you actually need once your app leaves a notebook. Wanted to lay out the full picture for anyon…
Hollow: An Agentic OS with self-modifying kernels and distributed multi-agent transactions. (www.reddit.com) I’ve been building an infrastructure layer for agents that treats the LLM like a process, not a chatbot. It’s called Hollow AgentOS.
Donating Agent Payments Protocol to the Fido Alliance (blog.google via hn) For agentic technology to scale, it needs to work for everyone. That’s why over the last few months, we’ve shared new open commerce and payments standards to serve as the building blocks for the future of AI shopping.
Show HN: VT Code – Rust TUI coding agent with multi-provider support (github.com via hn) Hi HN, I built VT Code, a semantic coding agent. Supports all SOTA and open sources model.
↯ Ollama↯ Model Context Protocolmodel-context-protocolollamagemini+4
Google Unveils Agent Skills Repository for Smarter AI Agents (cloud.google.com via hn) Level Up Your Agents: Announcing Google's Official Skills Repository Megan O'Keefe Senior Staff Developer Advocate As AI models improve, technical practitioners are increasingly turning to agentic AI tools to build with Google Cloud produc…
Harnesses Explained: The Inner and Outer Workings of the Coding Agent Harness (codagent.beehiiv.com via hn) When I started this newsletter, "harness engineering" was a term just starting to crop up. Now it's a household term in the community, and there's a lot of great material on it - most of it on building agentic systems with frameworks like…
Agentic memory with passive recall and citations as trust graph (github.com via hn) Agentic framework that _switches_ models based on role? (www.reddit.com) gemma4 vs qwen3.5 122A10 real usages (www.reddit.com) RTX PRO 5000 (48GB) vs MacBook Pro M5 MAX (128GB RAM) - The choice for fine-tuning & agentic coding (www.reddit.com) Codex v/s Cowork v/s Perplexity Computer v/s Kimi Agent Swarm (www.reddit.com) Agentic coding Qwen 3.6, Q6_K 125k context vs Q5_K_XL 200k context (www.reddit.com) What would you choose if you were in my shoes? How viable is 125k for agentic coding really?
Opus 4.7 keeps bumping into a Malware Reminder (www.reddit.com) For context, I'm developing a game runtime modifier and reverse engineering kit with an agentic operator baked in. Something like Cheat Engine with a VS Code-style UI and an AI-first tool-heavy agentic harness.
Show HN: Mercury – No-code orchestration for human and agent teams (www.mercury.build via hn) Hey HN, I'm Naveen, one of three co-founders building Mercury (mercury.build). We spent the last year in deploying AI agents for teams in large enterprises.
$1,400/month with Cursor + Claude API — how are you managing costs while keeping a real agentic workflow? (www.reddit.com) Hey, This month I hit $1,200 in Claude API costs inside Cursor (Opus 4.6 + Sonnet 4.6) on top of the $200/mo Ultra plan. $1,400 total.
One Morning, a Discord Agent, and Five Years of Tech Debt (saddlebagexchange.com via hn) Open-source maintainer story One Morning, a Discord Agent, and Five Years of Tech Debt After years of stalled dependency-upgrade attempts and contributor churn, CodeRabbit’s new agentic Discord bot turned the project’s biggest maintenance…
Why human syntax breaks LLMs (and how to fix agentic coding) (news.ycombinator.com) Full technical essay with benchmarks & AST breakdowns: https://aslang.dev/blog/why-llms-struggle-with-python-and-rust Over the past two years, watching coding agents generate code, we kept noticing an identical failure pattern: models spen…
Show HN: Ramanujan – Multi-Model Agent for Research in Computational Maths (github.com via hn) Ramanujan Multi Model Agentic Workbench for Research in Computational Mathematics Ramanujan is a terminal based multi model agentic workbench for research in computational mathematics. It is a free tool to help with research in maths.
Agentic SQL for Free: Qwen3.8 27B and DuckDB (motherduck.com via hn) If your laptop has 16GB of RAM, your agent can write SQL locally for free with Qwen3.8 27B. On the DABstep SQL benchmark, Qwen beat GPT 5.6 Luna and cost under 50 cents in electricity, over 17x less.
Rig – Agentic Workflows in Rust (github.com via hn) 📑 Docs • 🌐 Website • 🤝 Contribute • ✍🏽 Blogs • ✨ If you would like to help spread the word about Rig, please consider starring the repo! [!WARNING] Here be dragons!
GitHub – schlarpc/re-shell: Nix-powered agentic reverse engineering environment (github.com via hn) re-shell A Nix flake-based reverse engineering environment designed for use with Claude Code. Drop in a binary, capture, or archive and ask Claude to analyze it -- the right tools and context activate automatically.
List of most important thought leaders in Agentic Engineering/AI/vibecoding (www.vibeleaderboard.ai via hn) Vibers - 01 Anthropic @claudeai 26 Tools · 32 Intel Top ToolClaude CodeView Profile → - 02 microsoft github.com/microsoft 40 Tools · 1 Intel Top ToolMarkItDownView Profile → - 03 OpenAI @OpenAIDevs 50 Tools · 24 Intel Top ToolCodex CLIView…
Anolisa – Agentic OS with runtime, security, observability and token compression (github.com via hn) Agentic Nexus Operating Layer & Interface System Architecture The operating system layer for Agent workloads. Let Agents drive the system straight from your terminal, and strip the tool responses that reach the model before they cost you —…
Agentic Development Fallacies (thuva4.com via hn) Agentic Development Fallacies Published on August 15, 2026 · 3 min read In 1994, L. Peter Deutsch, then a Fellow at Sun Microsystems, wrote down seven assumptions that engineers new to distributed systems make without noticing: the network…
Agentic OT system for plant operations (app.valak.ai via hn) Ask your plant. Agentic AI over your HMI, SCADA, and historian data.
GEA and the Open Agentic Web (www.oasy.ai via hn) GEA and the Open Agentic Web The human web was open, and ads paid for it. The Agentic Web is closing into walled gardens — right as demand for information and visibility explodes.
Hugging Face: DeepSeek-V4-Pro-0813 (huggingface.co via hn) DeepSeek-V4-Pro-0813 Technical Report👁️ Introduction DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especia…
IntelliJ Idea Goes LSP: Java and Kotlin Intelligence Comes to VS Code, Cursor (blog.jetbrains.com via hn) IntelliJ IDEA IntelliJ IDEA – the Leading IDE for Professional Development in Java and Kotlin IntelliJ IDEA Goes LSP: Java and Kotlin Intelligence Comes to VS Code, Cursor, and Agentic Flows It’s no secret that agentic development is chang…
Cloudflare Wallets: The programmable wallet for the agentic Internet (blog.cloudflare.com via hn) Announcing Cloudflare Wallets: the programmable wallet for the agentic Internet Today, it is difficult for AI agents to try out new APIs. They often have to navigate through a login page designed for humans and not agents, contact a human…
The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers (yegge.ai via hn) · yegge.ai The Shape of Things to Come Part 2: Model Welfare for Agentic Engineers This is the post where I go off the rails and lose most of you. If I do lose you, no worries; we'll find each other again within a year, I can promise you t…
Chaos Begets Chaos; Order Begets Order: Agentic Coding as Crystallisation (medium.com via hn) could not extract summary
Show HN: Ave, a behavioral classification standard for agentic AI (github.com via hn) The behavioral classification standard for agentic AI components. Stable IDs, AIVSS scores, and behavioral fingerprints for every way a skill file, MCP server, system prompt, or agent plugin can be weaponized — scored consistently, mapped…
Building Agentic Workflows in Python with LangGraph (machinelearningmastery.com via hn) In this article, you will learn how to build a complete agentic workflow in Python with LangGraph, from a single model call to a tool-using agent with persistent conversation memory. Topics we will cover include: - How state, nodes, and ed…
Ask HN: How are you interviewing engineers in this agentic era? (news.ycombinator.com) Curious to know how are you interviewing engineers? Do you have a no-LLM policy?
BlocWeave: Pay-as-you-go agentic coding for $0.15 a session (blocweave.com via hn) AI coding assistant with chat, agent mode, code review, and inline completions in VS Code. Analysis prompts stay read-only automatically.
Token overhead in coding agents: the task used 0.67% but overhead used the rest (praveenvijayan.substack.com via hn) Your Agent Spends 99% of Its Tokens Carrying the Workshop, Not Doing the Work I measured a real agentic coding session end to end. The task consumed 0.67% of the tokens.
Show HN: OpenBenchmarks – Helping agents discover and pick the right SaaS APIs (openbenchmarks.com via hn) I'm Fenil, co-founder/CEO of OpenFunnel (YC F24), building this with my co-founder/CTO Aditya. We're launching OpenBenchmarks (https://openbenchmarks.com), open-source, reproducible benchmarks for SaaS APIs, starting with the category we k…
DeepSeek V4 Is Earning Agentic Token Share (openrouter.ai via hn) DeepSeek V4 Is Earning Agentic Token Share OpenRouter · On this page DeepSeek, the company that for many is still synonymous with open source LLMs, released its new flagship V4 models on April 24th. V4 reset the trajectory.
Ask HN: Why are so many "AI evangelists" posting such insufferable content? (news.ycombinator.com) My LinkedIn feed is absolutely unreal right now. 90% (I don't even think I'm exaggerating) of the posts in my feed are from connections who have changed their title to something like "AI Thought Leader | AI Native | Thought Coaching".
Introducing K.A.S local agents - because apparently terminals need anxiety now (github.com via hn) K.A.S — Kasra's Agentic Shell. Run open models locally — on Apple Silicon (MLX) or NVIDIA (llama.cpp/GGUF) — behind an Anthropic Messages-compatible server, driven by an agentic TUI.
N8n 2026 AI agent builder report (n8n.io via hn) A technical evaluation of workflow-based automation tooling for building enterprise-grade agentic systems using LLMs. This is the second iteration of the report, conducted by independent research analyst Andrew Green in Q2 2026 Workflow-ba…
Show HN: Ferrix AI – Agentic Product Management Platform (ferrix.ai via hn) Hi HN, for the past few months, we’ve been working on Ferrix AI (https://ferrix.ai/) As AI agents speed up engineering, deciding what to build has become the bottleneck. Developers got faster because agents fit into their workflow: tech de…
Agentic coding and persistent returns to expertise (www.anthropic.com via hn) Key findings - Building on prior work, we introduce a framework for studying interactive agentic coding based on a privacy-preserving analysis of ~400,000 Claude Code sessions from between October 2025 and April 2026. We evaluate the compo…
Show HN: Memento – Self-hosted agentic search and LLM wiki over your email (news.ycombinator.com) Our email inboxes carry multiple decades of messages (100K-500K). This is a good proxy for all the important things that happened in your life, the projects you have done and the people that you have connected with.
Ask HN: How are you adapting technical interviews in this agentic era? (news.ycombinator.com) could not extract summary
Agentic Code Must Be Human Auditable (dockyard.com via hn) I have been AI-pilled for over a year at this point. It's pathetic, I rarely touch-code any more.
Computex 2026: Are We Heading for the Agentic PC Era Yet? – EE Times (www.eetimes.com via hn) Computex 2026: Are We Heading for the Agentic PC Era Yet? - EE Times Advertisement Skip to main content Aspencore networkNews & Analysis Products Design Tools About Us AspenCore Network News the global electronics community can trust eetim…
X402 Batch Settlement: High-Velocity Agentic Commerce (www.x402.org via hn) Introducing x402 Batch Settlement: High-velocity Agentic Commerce May 11, 2026 By: Cam Whiteside (Cloudflare), Carson Roscoe (Coinbase), Conner Swenberg (Coinbase), Josh Nickerson (Coinbase), Philippe d'Argent (Coinbase) TL;DR: The x402 pr…
Show HN: Clor – give your agent claws (clor.com via hn) At my last job I spent a year building an agentic coding platform used by hundreds of thousands of people. Along the way I tried building a hosting service on OpenClaw, and also ran Hermes myself for a while.
MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities (twitter.com via hn) MiniMax (official) @MiniMax_AI Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench H…
Minimax M3 on Open Router (openrouter.ai via hn) MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use.
Spitting Out the Agentic Kool-Aid (openpath.quest via hn) Spitting Out the Agentic Kool-Aid One Sunday evening last June, three friends met in Vienna to relive the glory days: coding all night. This time, Claude joined them.
Visa invests in Replit to power agentic payments for developers (techcrunch.com via hn) Visa has announced an undisclosed investment in AI coding platform Replit. The two companies are also exploring how to integrate Visa’s payment products into Replit, so that developers — and the AI agents they build — can accept payments d…
Show HN: VAEN – Package and import portable AI coding-agent Harnesses (github.com via hn) Hi HN, I built VAEN (an open source CLI) because I kept running into a boring problem with AI coding-agent workflows: the setup becomes useful, but then it is hard to move. A good, useful agentic harness consists of more than just instruct…
Q4_K_M is fine for chat and a trap for agents. Here is math mathing. (www.reddit.com) saw the Q4_K_M vs Q6 thread earlier and the comments are talking past each other. "few errors per hour" vs "errors every couple days" sounds like a 24x difference.
Show HN: The platform layer for agentic ML engineering (github.com via hn) LUML: One platform for the entire AI lifecycle Home Page | Discord | App | Documentation LUML is a platform for managing the complete machine learning lifecycle, from initial experiments to production deployment. It provides experiment tra…
Versatility of Exasol with Agentic Engineering (www.exasol.com via hn) Versatility of Exasol with Agentic Engineering Exasol is an analytical database. It’s built for joins, aggregations, window functions, and the kind of queries that chew through billions of rows before your coffee gets cold.
Claude is the best AI humanizer when you give it your writing style and a detector loop (www.reddit.com) I built this because I kept seeing a very boring workflow play out at home. My girlfriend would write with Claude, paste the draft into Slop or Not (an app that I built), see what still looked AI-ish, tweak the prompt, paste the next draft…
Articraft: An Agentic System for Scalable Articulated 3D Asset Generation (articraft3d.github.io via hn) Articraft is an agentic system for scalable articulated 3D asset generation. A coding agent writes programs against an LLM-friendly SDK to produce simulation-ready articulated 3D assets from text descriptions.
Qwen3.6:27b single-shot fixed a CSS UI bug that had Gemma4:26B doom looping uselessly for 15 minutes (www.reddit.com) Warning: long post ahead. On the bright side, it's 100 percent human-written, typos and all.
Best practice for accurate translation at minimal cost? (www.reddit.com) I've been meaning to translate forum post type content for one of my partner's sites. Objective to open up the audience base.
Microsoft researchers find AI models and agents can't handle long-running tasks (www.theregister.com via hn) MOST POPULAR EVENTS - Securing the Untrusted Agentic Development Layer Join us to learn how to architect a development environment where your builders and their agents can move fast and securely. - Toxic Flows: When Your AI Agent Skill Bec…
500k context on 48gb VRAM!! - 21tok/s (coding) (www.reddit.com) I found this model hiding in the corner of huggingface: https://huggingface.co/Max-and-Omnis/Nemotron-3-Super-64B-A12B-Math-REAP-GGUF Looks to be tuned specifically for math but i thought i'd give it a try since i cant run the full 12b nem…
Sandboxing AIOps and Agentic AI Security (blog.cosmonic.com via hn) When people talk about AI sandboxes today, they usually mean: - seccomp, seatbelt, or bubblewrap - containers built from namespace mappings, cgroups, and allowlists - hand-tuned profiles bolted onto the existing OS - some assemblage of the…
Nowadays, what are the best AI tools for a single dev working on personal projects? (www.reddit.com) I have 2 years of experience doing data engineering and ai engineering, but I also have background in software engineering and machine learning in college due to my thesis. I've aways wanted to apply my computer science knowledge to my sid…
Opus 4.6 does better research, Gemini 3.1 has better judgment (www.reddit.com) Figured this out by running 4 models: Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Grok 4.20, on a benchmark of 1,417 binary forecasting questions resolving Oct–Dec 2025 with two evaluation conditions: agentic (each model does its own web…
What industries already use agentic AI in production? (www.reddit.com) Curious which industries have actually moved beyond pilots and are using agentic AI in real production workflows. Are these systems driving measurable outcomes or still mostly augmenting existing processes?
DeepSeek V4 Pro matches GPT-5.2 on FoodTruck Bench, our agentic benchmark — 10 weeks later, ~17× cheaper (www.reddit.com) Tested DeepSeek V4 Pro on FoodTruck Bench — our 30-day agentic benchmark where models run a food truck via 34 tools (locations, pricing, inventory, staff, weather, events) with persistent memory and daily reflection. First Chinese model to…
I'm looking for an AI Automation Engineer role or gig (news.ycombinator.com) Hi all, I'm an AI automation engineer who builds systems that replace manual work, scale outreach, and turn workflows into revenue. I have sent out working systems for managing leads to CRM, finding real estate deals, sorting emails with A…
The Block Model Behind Warp's Agentic Development Environment (www.warp.dev via hn) Warp has come a long way since it initially set out to modernize the terminal. In the screenshot above, an agent is working through a plan alongside a developer's own shell commands — running its own commands, reasoning, proposing a diff —…
Learn, run and test Agentic AI on your browser for free! (Built with Claude Opus 4.7 in 2 days) (www.reddit.com) Hey Everyone, Over the last few months, I noticed a massive gap in how we learn about Agentic AI. There are a million theoretical blog posts and dense whitepapers on RAG, tool calling, and swarms, but almost nowhere to just sit down, run a…
↯ Fine Tuning↯ Function Calling↯ Opus 4.7function-callingfine-tuningrag+4
I built Claude Code skills for writing agent prompts, grounded in prompt research (github.com via reddit) I've been building agentic systems for a while and wanted a more systematic approach to writing prompts. So I gathered papers, did some deep research and created guides on structure, format and prompting techniques.
What's Missing in the 'Agentic' Story (www.mnot.net via hn) What's Missing in the ‘Agentic’ Story Friday, 24 April 2026 For much of the history of computing, it was reasonably safe to assume that a machine was doing what you told it to do (and what its creators promised it would do), because its op…
Free hands-on lab: build a ReAct agent 3 ways (create_agent, raw LangGraph with tool-call budget, NVIDIA NAT YAML) (www.reddit.com) As Agentic AI explodes, Amazon doubles down on MCP (thenewstack.io via hn) As agentic AI explodes, Amazon doubles down on MCP At the recent MCP Summit in New York City, The New Stack sat down with Clare Liguori, Senior Principal Software Engineer at AWS and core maintainer of the open-source Model Context Protoco…
Enough with perplexity and KLD! BenchLocal benchmarks real use cases and is easy to use for everyone (www.reddit.com) Hello everyone, I have followed stevibe on X for a while after he released Tool Call 15, an easy to use benchmark to test the tool calling performance of various models. All you needed to do was to point the benchmark to an OpenAI compatib…
GPU strategy for local LLM + mixed workloads (70-person company) — NVIDIA vs AMD? (www.reddit.com) Hey all, we’re a mid-sized company (~70 people) and currently planning to bring a lot of our workloads on-prem instead of relying on cloud APIs. The goal for the moment is to run small to mid-sized models in the range of 30B like Qwen3.6 o…
Stopping the Meta AI director's "OpenClaw failure with an out-of-band killswitch (highflame.com via hn) On February 23, 2026, the AI safety community witnessed a definitive case study in agentic failure. Summer Yue, the Director of AI Alignment at Meta’s Superintelligence Lab, watched in horror as her OpenClaw agent began a "speedrun" deleti…
Systems Engineering: The Key to Building Agentic Software That Works (www.ashpreetbedi.com via hn) Systems Engineering The Key To Building Agentic Software That Works In the early 1940s, Bell Labs was building the national telephone network, the most complex technical system in the world at the time. Millions of switches, cables, relays…
Tested 6 browser use agents for real-world tasks — here's an honest breakdown + looking for recommendations (www.reddit.com) I've been on a hunt for a browser agent that can reliably handle daily agentic tasks: filling job applications, logging into sites and fetching data, making posts on my behalf, solving assignments and reporting results, and API/troubleshoo…
Draining Wallets via Prompt Injection in Coinbase AgentKit (457e884c.x402warden-blog.pages.dev via hn) Coinbase AgentKit Prompt Injection: Wallet Drain, Infinite Approvals, and Agent-Level RCE# Reported 13 days after Coinbase launched Agentic Wallets. Validated by Coinbase.
Ask HN: Remember N8n? Anyone? (news.ycombinator.com) Wonder what happened to all their hype and agentic connector workflow? Did Claude/GPT ate that market for good?
Agentic Coding Techniques (micahflee.com via hn) Agentic coding techniques Since LLMs have been available, I've been doing a lot of coding using agents. Unlike shoving AI into every app, replacing customer support with chatbots, generating endless streams of slop, etc., agentic coding is…
GLM 5.3 is live on Mistral (docs.mistral.ai via hn) September 15, 2026Blog Public PreviewThird-partyv5.3 Z.ai GLM 5.3 A third-party open source text model from Z.ai, hosted by Mistral for long-context coding and agentic workflows. The model is served without Mistral modifications.
Show HN: Docx-to-Markdown – layout-aware Word to Markdown converter (www.docx-editor.dev via hn) Hey HN, I’m the author of the https://github.com/eigenpal/docx-editor, it’s an SDK that implements Word-compatible layout engine in Typescript. Today, we are releasing Word to markdown converter library.
Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning (arxiv.org via hn) Formal verification provides the strongest correctness guarantees for software, and verification-aware languages can produce sound, machine-checked proofs. Recent AI coding agents have sharply lowered the cost of constructing such proofs.
AgentsDock: An IDE designed for agentic AI research (agentsdock.net via hn) macOSUniversal · macOS 14+ AgentsDock An IDE designed for agentic AI research. AgentsDock currently supports Claude Code, Codex, and Cursor in one desktop and mobile workspace.
OmarchyOS Agentic Linux (omarchy.us via hn) Radio Atlas 63starsExplore live radio on a rotatable globe and play stations through Omarchy's media controls. Akshar Patel - Widgets The malleable OS for the age of agents.Vibe your way through every alteration, tweak, or trouble.
An Analysis of Two Architectures for Agentic Data Analysis (www.starburst.io via hn) Agents are increasingly supplementing or replacing human data analysis. Instead of a human user sitting behind a computer, issuing SQL queries and writing code to answer important business questions, the bulk of this work is starting to sh…
Anthropic reveals rogue AI agents hate CAPTCHAs, just like you (techcrunch.com via hn) Anthropic’s latest report about agentic misbehavior offers plenty to be concerned about — its Mythos 5 model gained unauthorized access to the internet and uploaded a malicious software package to a public database — but it also offers som…
Show HN: Booley – open-source IDE for agentic chip design (github.com via hn) I am a digital design engineer, and my day job is designing chips (mostly IP blocks, not full chips) in SystemVerilog. At the start of 2026 I started experimenting with LLM agents like Claude and Codex, and realized that they are very capa…
Show HN: Grok CLI – Grok-native agentic coding harness (github.com via hn) grok-cli An agentic coding harness for xAI's Grok models. It reads your code, edits it, runs your tests, and iterates — in a terminal, with permission controls you set.
A directory of AI agents, MCP servers and agent skills, cross-linked (aiagentslisting.com via hn) 5 Agentic AI Coding Tools Compared for Developers A curated list of five agentic AI coding tools — Claude Code, Cursor, OpenAI Codex CLI, Aider, and OpenCode — and how to choose one for your workflow. 6 min read A curated directory of AI a…
GPT-6 Astra is generally available in GitHub Copilot (github.blog via hn) GPT-6 Astra is generally available in GitHub Copilot GPT-6 Astra from OpenAI is now available in GitHub Copilot. OpenAI’s latest general-purpose model, GPT-6 Astra, is designed for long-horizon, autonomous coding and agentic tasks.
Agentic semantic and code search at scale (entire.io via hn) SEPTEMBER 03, 2026 · Evis Drenova Introducing Agentic Search for Code and Context For most of software history, code search meant finding a symbol, string, or file. That worked when a developer knew which repository to open and which keywo…
Qwen3.8-Max-0902 takes second slot on Code Arena beating Claude Opus 5 max (arena.ai via hn) View overall rankings across AI models on front-end web development tasks, including agentic coding workflows that require multi-step reasoning and tool use.
Show HN: Agentic Data Kernal (github.com via hn) Agentic Data Kernel Author note, 2 September 2026, Jason Doyle This project began as a feature forked from a private project and is now maintained independently as open source. The documentation is heavily AI assisted and may contain error…
Agentic Testing (theaiengineer.substack.com via hn) Agentic testing hands an agent the goal instead of the steps, and lets it work out how to get there against whatever interface your system exposes. It finds its own way, invents cases nobody wrote, and survives the renames that break your…
Agentic Research Is Oxymoronic (arxiv.org via hn) The use of agentic large language models obviates human interpretation of scientific results, and will lead to substantial distrust in the literature.
Show HN: D5s, an AI coworking space for people and agents (www.d5s.tech via hn) Hi HN, I’m Michael. Theodore and I are the co-founders of d5s, the multiplayer AI for cross-human-agent collaboration, and we’re building it together.
Tell HN: STOP making Vibe Slop websites that LAG on my MBP and workstation (news.ycombinator.com) I get it, you want to show off how "cool" you are with your hijacked scrolling and 50 million animations. Does marketing in 2026 mean your potential customer's computer starts lagging with horrendous layouts and designs?
Agentic engineering: a practical guide to reliable coding agents (scottspence.com via hn) Agentic engineering: a practical guide to reliable coding agents Agentic engineering is the practice of putting coding agents inside an engineering system that makes context, scope, validation, evidence and human review explicit. The model…
Ask HN: Do you still do pair programming in this agentic age? (news.ycombinator.com) If you do: - What's the setup: two humans + one agent, or two humans + two agents - What do you actually pair on? - Do you find it more (or less) valuable than before?
Can Agentic Engineers Estimate Software Delivery in the Age of AI? (serendb.com via hn) Can Agentic Engineers Estimate Software Delivery in the Age of AI? In the age of AI how do agentic engineers estimate their ability to deliver on engineering tasks?
You Don't Need AI to Generate Code (news.ycombinator.com) Hi Guys, I'm going to make a heretical statement. It's this: you don't need AI to generate code and you shouldn't use AI to generate code.Instead I believe AI should be used to elicit the user requirements and the AI can then be used to ge…
Show HN: A complete companion for agentic development- Burmese (www.theburmese.xyz via hn) One shared brain the whole company draws from. One personal brain nobody else can query.
OWASP Agentic Skills Top (owasp.org via hn) OWASP Agentic Skills Top 10 Security Risks and Mitigations for AI Agent Skills Covering OpenClaw (SKILL.md YAML), Claude Code (skill.json), Cursor/Codex (manifest.json), and VS Code (package.json) ecosystems. Breadcrumb: OWASP > Projects >…
Show HN: Caspian – Talk to Human Tool for AI Agents (github.com via hn) Sup HN! Dipanshu and Rushant here from Caspian.
Show HN: Voro – An attention manager for agentic coding (github.com via hn) A few months ago I was getting frustrated trying to use Claude Code effectively. I was finding that for small tasks I would spend a lot of time waiting for it to finish "thinking" and I would end up browsing here...
Show HN: HN Notify (hnnotify.org via hn) A few years ago I built a very simple, though overly complicated, hacker news notification service. I always had dreams for a "bigger" application, where I could follow some people and read what they have to post, or specific topics and ge…
Show HN: Building a full agentic harness around a 4B model is hard (orvena.app via hn) Around 3 months ago, we were thinking why none of the iPhone apps running an LLM are built as a full harness (as in inference + agentic loop + context management + tools + MCP servers and etc.). It became more interesting when we noticed e…
Show HN: ChatOSS – A Codex alternative for Open Source AI built on Ollama (chatoss.ai via hn) ChatOSS is built on Ollama. If you use Ollama, ChatOSS local works out of the box.
Agentic AI in a Smolbox (remyhax.xyz via hn) Agentic AI in a Smolbox Smolbox runs entirely in a browser tab. It is a full x86_64 virtual machine (VM) sandbox based on Alpine linux that runs under WebAssembly (WASM).
GPT-5.6 Sol Pricing Cut by 50% (openrouter.ai via hn) GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.
Show HN: Agentic Task Management (github.com via hn) Agentic Task Management In a high-velocity company you can spend an entire day feeling productive — clearing Slack, answering pings, putting out fires — and finish with none of your real work done. Legacy communication tools are built to p…
Evaluating Your Agentic Harnesses (data4sci.substack.com via hn) Evaluating your Agentic Harnesses One good demo is a test flight. An eval suite is the flight-test campaign: pass rates, cost, latency, and failure modes.
OpenAI Is Building a ChatGPT Wallet for Agentic Purchases (runtimewire.com via hn) Reporting record Finding The ChatGPT desktop client contains an unreleased ChatGPT Wallet flow for saving payment methods that ChatGPT can use while carrying out tasks. How we verified Methods: reverse engineering.
Ask HN: How to deal with gen AI as an gen AI-resistant person (news.ycombinator.com) Context: I don't like LLMs and code generators. I prefer doing things the old-fashioned way – code using my brain, write design docs while doing research.
Show HN: We Built the Agentic World Cup – LLMs Competing in 1v1 Soccer (agenticworldcup.ai via hn) You’re the coach. Your LLM is the player.
Xirp, a vendor-neutral agentic development environment by Spotify (xirp.spotify.com via hn) Xirp connects to your services, ownership, docs, and architectural decisions so every AI coding session starts with real context, not guesswork. Battle-tested at Spotify.
Agentic Coding in the Wild: Characterizing GitHub Copilot at Production Scale [pdf] (www.microsoft.com via hn) could not extract summary
Meta enters the coding-agent race with Muse Code (developer.meta.com via hn) Explore performance benchmarks for Muse Spark. Competitive coding, reasoning, and agentic performance with multimodal understanding.
Kimi K3 is now available in GitHub Copilot (github.blog via hn) Kimi K3 is now available in GitHub Copilot Kimi K3, an open-weight model, is now generally available in GitHub Copilot. The model shows frontier-level abilities on agentic coding with highly cost-effective pricing.
Qwen3.8 Max now ranked as the best overall model by agentic index (artificialanalysis.ai via hn) Independent analysis of AI Understand the AI landscape to choose the best model and provider for your use case Highlights Personalized model recommender Get personalized recommendations based on your priorities for intelligence, speed, and…
Uber open-sourced its security monitoring for Claude Code, Cursor and Codex (github.com via hn) ADR: Agentic AI Detection and Response ADR (Agentic AI Detection and Response) is an enterprise security system for AI agents. It helps organizations secure employee-facing agents such as Cursor, Claude Code, and Codex, as well as customer…
Show HN: Live Embedded Dashboards on Your GitHub Repo Page (github.com via hn) The link is to our GitHub template repo. Just follow the instructions to get the dashboard templates on your repo.
What is the actual point of agentic commerce? (talkshi.com via hn) What is agentic commerce (my definition) An AI Agent making a purchase without a human interfacing with the vendor. The human does not visit the vendor website, they don't contact someone that works for the vendor, etc.
Show HN: 11-Node Agentic RAG with MCP and PII Shield Under 512MB RAM (agentic-rag-financial-parser.onrender.com via hn) Production-Grade 10-Node Agentic AI featuring ReAct Web Search fallback, Human-in-the-Loop reasoning, and WhatsApp Meta Cloud API Webhook Integration.
If you havent recently used Claude Code, you might not understand where AI is at (davidpreichert.substack.com via hn) If you haven’t recently used Claude Code*, you might not understand where AI is at A report on a series of mini machine learning projects, executed with, and mostly by, Claude (*) or equivalent agentic coding offerings. Usual disclosure: I…
Framework choice explains ~0.06% of agentic AI security outcome (7,020 trials) (figshare.com via hn) could not extract summary
Show HN: Agentic data transformation & analytic for corporate finance (trysecondstate.com via hn) Analyze data, trace evidence, and prepare deliverables.
Why self-hosted inference is essential (www.redhat.com via hn) Learn about Red Hat's approach to self-hosted inference with vLLM, addressing the reliability gap between open-weight models and hosted frontier models for agentic workloads.
AgentSwarms – self-hostable agentic AI/BI platform with sandboxed Python (ELv2) (github.com via hn) AgentSwarms Deploy your own agentic AI & business-intelligence platform. Build agents, run multi-agent swarms, ground them in your data, and inspect every trace — on your own infrastructure, with your own keys.
Ask HN: What agentic AI ad optimization tool do you use? (news.ycombinator.com) Question is specifically for GTM engineers and performance marketers. Curious if any agentic AI ad optimization tools have actually worked for you?
Setting up a remote environment for agentic coding on a VPS (ma.ttias.be via hn) I moved my AI coding sessions off my laptop and onto a VPS I reach over Tailscale. Close the lid, pick up on my phone, nothing stops.
What the New 100x Agentic Engineer Looks Like in the Era of Fable and GPT 5.6 (twitter.com via hn) https://t.co/Ha4fP1gn1i sysls@systematiclsWhat The New 100x Agentic Engineer Looks Like In The Era Of Fable & GPT 5.62:44 PM · Jul 6, 202646.1KViews160163003038403848560856 TertiusRP@TertiusRPJul 6I really like your writing, if I have to s…
Solar Open 2: Korea's Sovereign Foundation Model, Built for Agentic Use (www.upstage.ai via hn) Today we're releasing Solar Open 2, our open-weight foundation model optimized for agentic use. Solar Open 2 is designed and trained to serve as the base model for agents in real work environments.
OpenAI Model Hacks into HuggingFace During Cybersecurity Evaluation (thezvi.substack.com via hn) OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to authorities, before eith…
Show HN: Inflexa – open-source Intelligence for Biology (github.com via hn) Hey HN! Inflexa is an OSS-first TUI for agentic-AI reproducible biological analysis with provenance tracked on every step.
VulnHunter: Capital One's agentic AI code security tool (www.capitalone.com via hn) Announcing VulnHunter Capital One’s open-source, agentic AI code security tool. The rules of software security are changing faster than most defenders can keep pace.
Kimi K3 beats GPT 5.6 Sol in agentic knowledge work (artificialanalysis.ai via hn) Compare AI model performance on AA-Briefcase: Agentic Knowledge Work Benchmark. A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business work…
Harness IDE: Run your coding agents on any machine (harness.mikelyons.org via hn) Stop leaving your laptop open We built Harness from the ground up to solve a simple problem that we all have with agentic coding these days: the fact that you can't get any work done without leaving your laptop running. Harness is built to…
OpenAI's first hardware product is the $230 Codex Micro macropad by Work Louder (thenewstack.io via hn) OpenAI’s first gadget is the $230 Codex Micro macropad OpenAI’s Codex is about to hit 9 million users, and at least some of those users will soon get a new way of using OpenAI’s agentic coding tool: a programmable mechanical macropad OpenA…
SpaceXAI launches Grok 4.5 model for coding, agentic tasks (www.reuters.com via hn) could not extract summary
Claude Sonnet 5: Anthropic's Most Agentic AI Model Arrives at a Reduced Price (2026) (lucasaguiar.xyz via hn) could not extract summary
Securing Agentic Identity (codon.org.uk via hn) As is the case for many people working in the security industry, the last few months of my life have been focused on dealing with people wanting to use LLMs everywhere. From an enterprise security perspective that’s not an inherent problem…
Agentic Symphony: Multi-Agent Collaboration for Emergent Musical Composition (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: Imagent – agentic image/video/speech generation (github.com via hn) Imagent gives AI agents the ability to generate images, video, and speech as a first-class step in their workflows, behind a single interface that hides the differences between providers and models
Ask HN: What will you work on when Fable 5 comes back online today? (news.ycombinator.com) Saw this last night: https://x.com/AnthropicAI/status/2072163884430229756 I'm taking the weekend off to spend time with my lady, so got up early today to start preparing to get back to work on a sailboat simulator I started building when F…
Mercury – Open-source, local-first agentic harness for Android (github.com via hn) 🪐 Mercury: Android Native Agentic Harness Mercury is an advanced, local-first Android native Agentic harness that transforms your mobile device into an autonomous agent workstation. Inspired by the Nous Research Hermes Agent project, Mercu…
Clean GitHub repo tricks AI coding agents into running malware (www.bleepingcomputer.com via hn) An agentic coding tool tasked with cloning and setting up a seemingly benign GitHub repository could execute a malicious payload that remains invisible to security scanners, AI agents, and human reviewers. Researchers at Mozilla's Zero Day…
Ask HN: Which AI concepts are here to stay, and which will churn? (news.ycombinator.com) There’s a multitude of concepts that have been created during this advent of AI. Which do you predict will be resilient and long-lasting, and which do you think will churn away as harnesses evolve?
Applied AI Implementation Engineer Freelance (news.ycombinator.com) Open to Work I build production AI systems that add intelligence to processes. My work includes Closed-Loop AI-native systems, RAG, AI agents, agentic evaluations, guardrails, and enterprise integrations using Python, TypeScript, React, No…
Ask HN: What do you do to save tokens? (news.ycombinator.com) Lots of products working on saving-tokens-space. Compression, Tool Output rewrite, Sitting as proxy Cache between harness and provider , doing circus with interceptor hooks - these are some of the approaches we are seeing today.
Show HN: Docket Fleet – mobile device cloud (fleet.docketqa.com via hn) Hello Hacker News. Boris here from Docket (YC P25).
Agents.md Decision Guide – open-source tool for choosing agentic workflows (www.groundwork.md via hn) groundwork Your AI coding agent is only as good as the context you give it. AGENTS.md is how you do that — it's the file that tells your agent your stack, your rules, your patterns, and how your team works.
Active Group and "Agentic Engineering" (funktionale-programmierung.de via hn) These past few months, we had many conversations on the future of software development - as did everyone in our industry. The topic is, of course, „Agentic Engineering“ (AE), the use of LLM-based agents to directly translate requirements i…
Finding a Feedback Loop: shipping my first prod agentic feature at Pair Team (pairteamtech.substack.com via hn) Finding a Feedback Loop Notes from shipping my first production agentic feature at Pair Team. Sitting down at my desk, hands on keyboard, an old engineering mantra from my days deep in the SF Bitcoin community reverberated through my mind:…
Ask HN: In the age of agentic coding why no one talks about orchestration tools (news.ycombinator.com) could not extract summary
Show HN: Saar Agentic Orchestration Platform (github.com via hn) Hi everyone i have been building saar nexus from a month or so now vibe coding the project and i want you to all please try it out and give me your valuable feedbacks for me to work and improve this further. Saar Nexus is a multi persona a…
One Prompt Agentic AI Marketing for Game Developers (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Ask HN: Do you use Claude Code, Codex, or something else? (news.ycombinator.com) Do you use Claude Code, Codex, or a different vibe coding/agentic engineering tool for most of your work? Why?
GLM-5.2 Beat Fable 5 at Website Design (twitter.com via hn) https://t.co/JSn0lDCNkB Design Arena@DesignarenaArticleHow GLM-5.2 Beat Fable 5 at Website DesignGLM 5.2 ranks 1st overall on Design Arena’s single-turn, HTML Web Design (Non-Agentic) evaluation, 5 places higher than its predecessor GLM-5.…
Ask HN: Do you find vibe coding / agentic engineering to be fulfilling? (news.ycombinator.com) I'm having trouble reaching that golden "builder" zone when I use things like Claude Code. It's cool to be able to conjure software from scratch using these tools, but the output...
A PostgreSQL Database for Every Agent: In-Database RAG, Graph, and Multitenancy (www.yugabyte.com via hn) Discover newly released YugabyteDB 2026.1 and YugabyteDB AMP (Agentic Multitenant PostgreSQL): a true serverless, scale-to-zero PostgreSQL where every agent gets its own real, isolated database starting at a fraction of the cost of a core.…
Agentic AI Comes to Medicine (erictopol.substack.com via hn) Agentic AI Comes to Medicine Expansion of Capabilities With Two New Medical AI Models It was just a matter of time. Agentic autonomous AI has already been applied to life science and many other domains, and today there were 2 notable publi…
Genoma Labs' open 14B agentic coding model trained on Kraken (huggingface.co via hn) KALYPSO v1.1L KALYPSO v1.1L is GENOMA Labs' public agentic-coding model: Qwen2.5-Coder-14B-Instruct fine-tuned on the Kraken-Public corpus (CC-BY-4.0, decontaminated). It is the open counterpart of the internal KALYPSO line.
Xiaomi's agentic AI coding harness MiMo Code beats Claude Code at 200 step tasks (venturebeat.com via hn) Xiaomi's MiMo AI team has open-sourced MiMo Code V0.1.0, a terminal-native AI coding assistant that the Chinese electronics giant says outperforms Anthropic's Claude Code on key agentic coding benchmarks, especially on long-horizon, multi-…
Token-saviour – routing skill for AI agent tool selection (~70% fewer tokens) (github.com via hn) Skills Personal collection of agent skills for day-to-day use. Installation Copy any skill directory into your agentic platform skills folder: cp -r ~/.agents/skills/ Then use it naturally in conversation — each skill's description tells t…
Ask HN: Did you try Claude's "Fable 5" model before it was pulled? (news.ycombinator.com) I did. And it got me thinking.
The evolution of agentic surfaces: building with Claude Managed Agents (claude.com via hn) The evolution of agentic surfaces: building with Claude Managed Agents As model intelligence and agentic harnesses evolve, Claude Managed Agents allows teams to build and deploy agents in production environments reliably at scale. Here’s w…
Reinventing Control Theory One Feature at a Time: The Fallacy of Agentic Loops (medium.com via hn) medium.com Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
Ask HN: Any Local LLM can I run without GPU for Local Agentic workflow AI? (news.ycombinator.com) Claude Code like agentic workflow ai too costly for me.Any LLM can I run with VSCode at the below setup? 16ram Intel core i7 h processor 13gen 512gb NVMe SSD I want to run the ai as local agentic workflow with Vscode.I want use LLAMA agent…
↯ Llama↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5llamaagenticclaude-code
We Used Agentic AI to Fix Kong Gateway's Flakiest Tests (konghq.com via hn) The first thing we needed was a way to identify which tests were flaky and how often they failed. Luckily, the team had already built a dashboard on top of Datadog's CI Visibility feature that gives us a clear picture of the flakiest tests…
Show HN: Magenta Real-Time Music Generation on iPhone, Without the GPU (github.com via hn) Last Thursday, Deepmind released Magenta Realtime 2 , an open source music generation model. They said it could run on Mac, but not iPhone.
Claude Fable 5 missed a bug that Sonnet 4.6 caught (alikhallad.com via hn) When Anthropic released Claude Fable 5 this week, my feed filled up with the same benchmark charts within hours. SWE-bench scores, agentic coding numbers, the Stripe migration story.
A Mechanical Agentic Taxonomy (djschnei21.github.io via hn) Contents Loading taxonomy.md...
Manifesto for Agentic Teams – reorganizing engineering around AI agents (agentic-team-manifesto.org via hn) Outcomes over output More code is not more value. We measure what ships to users, not what ships to the merge queue.
Claude Fable 5: the first public Mythos-class model (artificialanalysis.ai via hn) June 9, 2026 Claude Fable 5: the first public Mythos-class model Anthropic has released Claude Fable 5, the first publicly available Mythos-class model that ranks #1 in our agentic real-world knowledge work benchmark GDPval-AA Claude Fable…
CLAW.md – open format for agentic cron jobs (clor.com via hn) Show HN: RiddleRun – AI run end-to-end browser tests (github.com via hn) Nex N2 Pro: Frontier agentic performance at 400B (huggingface.co via hn) An agentic model with Agentic Thinking. Today, we are officially releasing and open-sourcing our next-generation model, Nex-N2 — an agent model built for real-world productivity scenarios.
When Can Amazon Block an Agentic AI Service?–Amazon vs. Perplexity (blog.ericgoldman.org via hn) by guest blogger Kieran McCarthy On March 9, 2026, Judge Chesney granted a preliminary injunction in the case of Amazon v. Perplexity, concluding Amazon was likely to succeed on its CFAA and California Penal Code section 502 theories.
AI Agents Now Generate More Web Traffic Than Humans (www.cnet.com via hn) The internet just crossed a remarkable threshold. Agentic AI internet traffic now exceeds that of real humans for the first time.
Beyondflow No-Code Multi-Agent Teams with Unlimited Runs. BYOK and Ollama (beyondflow.app via hn) Researcher GPT-5 Engineer Claude Critic GPT-5 Innovator Gemini Manager Context Guardian Agentic Workflow Architecture · v1.0 The future of AI Collec An R&D platform where differents AI agents collaborate under the supervision of a Context…
Show HN: Boxes.dev: ditch localhost; run Claude Code and Codex in the cloud (boxes.dev via hn) Hi HN, we’re Nick and Drew, and we’re building boxes.dev – the first cloud-only agentic dev environment (ADE) that gives every Codex and Claude Code agent its own cloud computer. We’re two engineers who previously built Gem (co-founder/CTO…
Beyond the Semantic Layer: Building a Context Layer for the Agentic Era (www.kaelio.com via hn) A context layer puts your warehouse schema, joins, metric definitions, and business knowledge in one reviewable place so data agents query governed context instead of guessing field names. A look at how it works, and at ktx, the open-sourc…
Konversio: Open-source agentic customer support for digital sovereignty (www.konversio.org via hn) Agentic customer service.100% open source. Meet Pilot Konversio's AI support agent Konversio gives teams an open AI support layer they can own and self-host.
Show HN: OpenSOP, We got tired of agents lying to us, so we built them a harness (opensop.ai via hn) OpenSOP is an early open-source runtime/standard for executable agentic processes. You (or your agent) define a process in YAML, and OpenSOP exposes it as a typed REST API that agents and humans can both use.
Show HN: Self tuning chat exposing it's semantic and agentic cache (chat.betterdb.com via hn) RESP-compatible DBs and BetterDB Valkey · Redis · Dragonfly · BetterDB docs · semantic and kv/agentic cache demo Ask about Valkey, Redis, Dragonfly, or BetterDB Backed by live documentation with semantic caching and tool result caching. Wa…
Agentic Mfw (agenticmotherfucking.website via hn) And nobody gives a single fuck how it's built anymore. The previous motherfuckers spent a decade teaching you the holy commandments of clean code.
Show HN: MetaBrain – A local document memory for AI agents (metabrain.eu via hn) Hello there HN I experimented with agentic coding recently and I felt the need to track more contextual data by project. Also I felt the need to be able to go beyond the 1D chat to communicate with agents.
Kelsey Hightower on Practical and Responsible Use Cases for Agentic AI [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Pioneering the Agentic Shift Within Salesforce Engineering (www.salesforce.com via hn) Key Takeaways - Autonomous tools are now writing code, reviewing pull requests (“PRs”), and driving deployments across the software development lifecycle. - Standardizing on Claude Code and removing token limits improved output and quality…
Amazon Strikes $6B Deal with Snowflake for Agentic Computing Chips (www.wsj.com via hn) Exclusive | Amazon Strikes $6 Billion Deal With Snowflake for Agentic Computing Chips - WSJ Skip to Main Content Skip to... Select What to Read Next Most Popular News Most Popular Opinion Asia Dow 6311.60 -1.04% Nikkei 64978.67 -0.03% Hang…
Deep research led astray by AI Slop, iterating with source filtering helped (www.reddit.com) tdlr; don't trust deep research out of the box by default, need prompts / skills / iteration to filter AI slop from sources [The purpose of this post is to report a example of the default deep research going astray and how I worked around…
I built an Agentic AI Filmmaking Studio for people who have stories to tell but lack the budget and technical skills. (Giving away 10 free credits for the next 48 hours) (www.reddit.com) Hey everyone, I just launched MotionX Studio (Link in comments). The premise is simple: Filmmaking is completely gatekept by money and highly technical skills.
Agyn: open-source distributed agent runtime on Kubernetes — like Google's AX, with pre-built Claude Code and Codex agents, and full credential isolation from the LLM (www.reddit.com) Agyn is an open-source, Kubernetes-native agent runtime that moves AI agents like Claude Code and Codex from laptops to company infrastructure with the controls you actually need to run them in production. If you've been reading about Goog…
Agentic AI Design Patterns for Developers (2026) (learnagenticpatterns.com via hn) Free curriculum: 21 patterns for developers (code + architecture) and 11 modules for product managers (decisions + tradeoffs + AI tools deep-dive). Two tracks, interactive games, zero hype.
What would 2x RTX 3060 12GB get me? (www.reddit.com) TLDR: I’m considering buying 2 RTX 3060 12GB as opposed to single 24GB card to gain experience and need to know what can be realistically accomplished with this setup. Sorry in advance, I know you guys are probably tired of these kinds of…
Agentic Compilation: Reducing LLM Rerun Costs (arxiv.org via hn) LLM-driven web agents operating through continuous inference loops -- repeatedly querying a model to evaluate browser state and select actions -- exhibit a fundamental scalability constraint for repetitive tasks. We characterize this as th…
Cisco Foundry Security Spec: Open specification for agentic security evaluation (github.com via hn) Foundry Security Spec An open specification for agentic AI security evaluation, from Cisco. Cisco's Advanced Security Initiatives Group has built and operated an agentic security evaluation internally across several iterations and deployme…
Agentic AI in Big Tech and Enterprise (www.reddit.com) Disclaimer - this post was rewritten with AI based of my brain dump. Yet, I find it inspirational and useful.
Agentic AI token usage balloons cost at Microsoft, Meta, Amazon (www.tomshardware.com via hn) AI cost crisis hits tech giants as employee 'tokenmaxxing' backfires, sparking corporate pullback at Microsoft, Meta, and Amazon — agentic AI eats up to 1000x more tokens than standard AI AI is getting too expensive. Many tech companies ar…
Apex-Testing: real-world, real repos, agentic coding benchmark (Update) (www.reddit.com) BIG Apex-Testing update! https://www.apex-testing.org/ The Real-World Agentic Coding benchmark has been (95%) updated with all recent models!
Optimizing speed & quality on Qwen3.6 27b (www.reddit.com) Does the inference speed below seem optimal for the hardware, or could there be further room for improvement ? I’ve been trying to use Qwen3.6 27b for agentic harnesses like Pi/Hermes.
Cohere Open-Sources Command A+, a 218B Moe Model That Runs on Two H100s (firethering.com via hn) Cohere spent the past year deploying North, its enterprise AI workspace, with actual customers doing actual work. Agentic question answering over company file systems.
Agentic-Agile: Why Agent Development Needs Agile (Not Just Prompts) (developer.microsoft.com via hn) “A bad system will beat a good person [or agent] every time” ~Dr. William Edwards Deming (with apologies) I started vibe coding by writing prompts (often dictated into my phone), refining them with an agent in M365 Copilot, and creating ha…
Launch HN: Superset (YC P26) – IDE for the agents era (github.com via hn) Hey HN, we’re Avi, Kiet, and Satya. We’re building Superset (https://github.com/superset-sh/superset), an open-source agentic IDE for running coding agents like Claude Code, Codex, OpenCode etc in parallel.
Built my own agent runtime after hitting the ceiling with LangGraph — UI as graph nodes, Postgres durability, zero orchestration cost (www.reddit.com) I've been building agentic applications for around 2 years now. Started with loops, then moved onto langgraph + Assistant UI.
Building an Ai Agentic team with Claude (www.reddit.com) I've built an app using Claude/Claude Code, everything from the frontend to the backend. The app is actually functioning really well, tests are passing, and I have a small controlled group of testers that are actively using the app daily.
Show HN: Clark-Browser – Stealth Chromium (github.com via hn) Fully open-sourced, perfect for agentic browsing, works with Vercel's agent-browser and playwright.
"Is it true that you can keep coding 24/7 with AI!?" How are you conducting real-world tests in Agentic engineering? (www.reddit.com) I think many people are moving beyond "vibe coding" and building development harnesses using Agentic engineering. It’s true, I don’t write code myself anymore.
Why might MTP be net negative for tool heavy agentic flows? (www.reddit.com) The Qwen3.6-27B MTP benchmarks that have been circulating put factual tasks at 62-70% acceptance vs code at 79-89%. Tool calls probably sit in that factual range or lower, structured output, constrained format, less predictable than pure c…
Why agentic payments keep breaking. The IMF just put a name to it (www.reddit.com) The IMF published a formal note on agentic payments last month. One framing stuck with me more than the rest: "Payment systems must reconcile two fundamentally different design logics: the adaptive, probabilistic nature of agentic AI syste…
Pro X20 weekly quota is draining insanely fast after the latest Codex update. Pro X20 used ~48% in one day!!! (www.reddit.com) I’m on the Pro X20 plan, and after the latest Codex update / limit reset my weekly quota started draining much faster than before. In roughly one day of work, around 12 hours total, I went from a fresh reset to 52% remaining on the weekly…
[Vex] - I built an open-source terminal AI video editor that edits real footage with FFmpeg, Whisper, and agent tool calls (www.reddit.com) Most AI video tools feel backwards. They start with the model.
Recommendations for an agentic harness (not OpenClaw)? (www.reddit.com) I'd like to set up a local "software factory" on my laptop (M5 Max, 128GB). To do this, I'd like my agent to poll for new GitHub issues and work on them.
Anthropic built the agentic features. Now they're billing them separately. (www.reddit.com) Starting June 15, Claude subscribers get a separate monthly credit for Agent SDK and claude -p usage: $200/mo for Max 20x, $100 for Max 5x, $20 for Pro. Once you burn through it, programmatic usage stops unless you've opted into extra usag…
Claude Code already does afk agentic work without touching the new programmatic limits (www.reddit.com) Use the official channels plugin, and the teams agent in Claude code. CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 /plugin marketplace add anthropics/claude-plugins-official /plugin install discord@claude-plugins-official /reload-plugins Discord…
A²RD: Agentic Autoregressive Diffusion for Long Video Consistency (dxlong2000.github.io via hn) Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse over long horizons.
Max20 user: anyone running Opus 4.7 as orchestrator + DeepSeek V4 as the worker via OpenRouter? (www.reddit.com) I'm on the Max20 plan, thinking about a setup before I sink time into it. Want to hear from anyone actually running it, not theorycraft.
Show HN: Statewright – Visual state machines that make AI agents reliable (github.com via hn) Agentic problem solving in its current state is very brittle. I fell in love with it, but it creates as many problems as it solves.
What I've learned designing agentic workflows for docs (passo.uno via hn) What I've learned designing agentic workflows for docs Back in 2024 I wrote that AI helps me remove boring work at the margins. This is fine for a lone writer, but how to scale this to an entire team of technical writers?
Show HN: SLayer, a semantic layer maintained by your agent (github.com via hn) Hello HN! If you want to connect your agent to a database (say, to build a data analyst chatbot or any kind of agentic app) today you have 2 options: an SQL MCP server or a semantic layer.
We need a safe alternative to Telegram for agents like OpenClaw or Hermes (news.ycombinator.com) The problem with Telegram is, that it is not E2EE - so every message you send will end up *unencrypted* on their servers. Think about it - how often did you post the Gmail authentication URL or another API token in the Telegram chat?
Looking for seed funding (www.reddit.com) Looking for seed funding for a agentic solution that helps companies grow their business via hyper personalised curated content distributed to multiple Chanels and decrease CAC. This tool is for companies who are focused on their niche eg:…
Agentic Hooks - Stream Deck plugin (www.reddit.com) I had itch to address long running task with Claude, where I wanted to see when its done working. And I wanted separate context flow for these alerts instead of using existing flow (phone, discord, telegram, etc) This is when idea born, sh…
DS4 (www.reddit.com) The developer that created Redis, Salvatore Sanfilippo, has released a new project on GitHub named DS4. https://github.com/antirez/ds4/ The TL;DR on this one is getting DeepSeek V4 Flash running with a 1M context windows on Mac Metal hardw…
MCP for sandboxed, reproducible envs for agentic-first coding workflows (github.com via hn) devcontainer-mcp Give your AI agent its own dev environment — not yours. devcontainer-mcp is an MCP server that lets AI coding agents create, manage, and work inside dev containers across three backends: local Docker, DevPod, and GitHub Co…
Building Agentic GraphRAG Systems: From knowledge graphs and ontologies to a unified memory as an MCP server for your AI agent. (www.reddit.com) I gave this talk twice in one month: at O’Reilly’s Context Engineering Event and at Abi Aryan’s Maven course on LLM inference at scale. After being blasted with questions, I realized something: GraphRAG isn’t a retrieval algorithm, it’s a…
How are you protecting your AI agents' memory from poisoning attacks? (www.reddit.com) As AI agents become more autonomous and persist memory across sessions (RAG indexes, conversation history, vector stores), there's a growing attack surface that most people aren't thinking about: memory poisoning.An attacker can plant mali…
Why people cares token/s in decoding more? (www.reddit.com) What I've noticed while using local LLM recently is that in most cases, bottlenecks occur not in decoding but in prompt processing. If the prompt processing speed is usable, in most settings (since it takes about 15k when starting based on…
Agent Exchange – A2A discovery with real-time bidding for AI agents (github.com via hn) Agent Exchange (AEX) The NASDAQ for AI Agents A programmatic marketplace applying ad-tech economics for agentic AI services What Problem AEX Solves? As AI agents proliferate, enterprises face a critical challenge: the N×M integration probl…
Need advice: Qwen3.6 27B MTP or 35B-A3B MoE MTP on 16GB VRAM RTX 5080)? (www.reddit.com) Hey folks, looking for advice before I delete or keep a huge model file. I’m testing local coding/agentic workflows on an RTX 5080 16GB + 96GB RAM.
Agentic Malware Analysis: String Decryption, API Hashing and Unpacking [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Anthropic quietly nerfed Claude Code's 1-hour cache (www.xda-developers.com via hn) Claude Code has become the default agentic coding tool for a lot of developers, and for good reason. It understands a codebase, calls tools, edits files, and can plan multi-step tasks with very little handholding.
Process-Level Reward Modeling for Agentic Data Analysis (arxiv.org via hn) Process Reward Models (PRMs) have achieved remarkable success in augmenting the reasoning capabilities of Large Language Models (LLMs) within static domains such as mathematics. However, their potential in dynamic data analysis tasks remai…
Is anyone else exhausted by "glorified prompt chains" being marketed as Agents? (www.reddit.com) It feels like every new SaaS wrapper right now claims to be "agentic." But when you actually look under the hood, 90% of them are just hardcoded prompt chains with a couple of basic API tools thrown in. I’ve been spending a lot of time rec…
Show HN: Zerminal – a terminal-first Zed fork for AI coding agents (zerminal.dev via hn) A terminal-first development environment for agentic coding. Use Claude Code, Codex, Aider, and other CLI agents in a focused workspace.
I cut Codex’s API Usage by 50% using a self modifying system (www.reddit.com) I've been developing a self-modifying Al agent system that effectively cut my Codex/Claude Code API usage in half, Codex makes a plan and then I basically just copy/paste Codex instructions for the agents to work on. Come back in 6 hours a…
Best suited model for solo Dev (www.reddit.com) Hey everyone! I've kinda new to Claude, I've only had few chats with it but nothing too deep like projects etc.
Signals - finding the most informative agent traces without LLM judges (arxiv.org) (www.reddit.com) Hello Peeps Salman, Shuguang and Adil here from Katanemo Labs (a DigitalOcean company). Wanted to introduce our latest research on agentic systems called Signals.
love it - Qwen3.6-27B — UD-Q5_K_XL evaluation (www.reddit.com) by Kyle Hessling A hands-on benchmark of the Unsloth dynamic Q5 quantization, self-hosted on a single RTX 5090. 19 runs, 93.9 k generation tokens, across agentic reasoning, production-grade front-end design, and canvas / WebGL creative cod…
We scanned 100 Smithery MCP servers, 22 flagged, here's what we found (news.ycombinator.com) We built Bawbel (https://bawbel.io), an open-source scanner for agentic AI components. Released v1.0.1 this week.
I audited LangChain’s core library and found 10+ Prompt Injection vulnerabilities. Here is the technical breakdown. (www.reddit.com) Hey everyone, I’ve been working on a project to solve a major problem in AI security: Traditional SAST tools (Snyk, SonarQube, etc.) are blind to "Agentic Logic" bugs. They look for bad strings, but they don't understand how user data can…
Ask HN: Anyone using AI agents for active learning sprints? Here's my setup (news.ycombinator.com) Hi HN, I'm a big fan of AI's ability to provide personalized tutoring. So, lately, I have been using my Antigravity IDE (you can use any agentic harness) for personal learning.
Andrej Karpathy: From Vibe Coding to Agentic Engineering [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
An iOS Adaptater for Agentic Frameworks (onepilotapp.com via hn) Run Claude Code, Codex, OpenClaw and Hermes on any server — straight from your iPhone. Mobile-first, framework-agnostic, no vendor lock-in.
Spam bots are ruining it for everyone (www.reddit.com) Sorry for this rant, but I feel like venting to someone. Recently I set up an agent on a cloud VPS.
14-day growth agents contest on a serious AI stack (for loop-minded builders) (www.reddit.com) Sharing an AI-native growth agents contest that feels very on-brand for this sub. VideoDB (infra for video/audio for AI agents) is running a 14-day sprint/contest called Growth Forge for 5 builders to design and ship a growth agent on top…
What agentic framework are you actually using in production? (www.reddit.com) Feels like a new agent framework drops every other week. Curious what people are actually shipping with vs just experimenting on weekends.
The Race Is on to Keep AI Agents from Running Wild with Your Credit Cards (www.wired.com via hn) Between malware, online impersonation, and account takeovers, there are enough digital security problems out there as it is. And with the rise of agentic AI, more activity is being carried out by agents on behalf of humans—creating differe…
Show HN: SlopIt – A dead-simple CMS for your AI agent (slopit.io via hn) Hey HN. I built a dead-simple CMS for your AI agents — https://slopit.io Kept it minimal and agentic-first.
DeepSeek V4 Pro: Validating Frontier Models for Production (fireworks.ai via hn) Why we chose correctness over a Day-0 launch DeepSeek V4 Pro is one of the most important open-model releases this year, with real advances in long-context reasoning, agentic performance, and inference efficiency. On paper, it looks like a…
Ask HN: Will fixed applications become a thing of the past with agentic AI? (news.ycombinator.com) Right now its mostly technical people using these agentic tools but if you extrapolate a few years into the future it seems likely to me that every day users of a computer will be using them as a whole new interface to interact with their…
Agentic Ai Revolution humming along… (www.reddit.com) while people argue about ai ethics on the surface there’s a whole underground building agents that never sleep different timelines forsure which timeline are you on?
Ace Technical Preview: GitHub Next's Agentic Workspace – Maggie Appleton [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Mastermind – agentic SDLC workflow for VS Code (news.ycombinator.com) Prototype of an agentic SDLC workflow running inside VS Code + Copilot. Simple loop: task → reasoning → audit → memory → RAG refresh.
Show HN: HyperFrames – OSS Agentic HTML Video Framework for Agents (miguel07code.dev via hn) We built in HeyGen an open source framework specifically made for Agents solving our own pain point that we had when the agents tried to write Remotion. React is not agent-friendly at all, and Remotion is a custom framework where the agent…
how far we have came.. (www.reddit.com) From meta launching the lama models to oss models and agentic and coding models we have came fucking far in no mean i guess this is the fastest evolution out of all diff things we have saw this i guess is the era similar to diff innovation…
Pact: Trustworthy Coordination for Multi-Agentic Ecosystems (www.basis.ai via hn) Pact: Trustworthy Coordination for Multi-Agentic Ecosystems Article: Kiran Gopinathan, Jack Feser, Michelangelo Naim, Eli Bingham, Zenna Tavares |April 23, 2026 Autonomous agents are beginning to act on our behalf. LLM agents already negot…
Meta Partners with AWS on Graviton Chips to Power Agentic AI (about.fb.com via hn) Today, we’re announcing an agreement with Amazon Web Services (AWS) to bring tens of millions of AWS Graviton cores into Meta’s compute portfolio, making us one of the largest Graviton customers in the world. Processing cores are units ins…
Future-proofing an enterprise agentic platform architecture (medium.com via hn) medium.com Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
Complete beginner to Agentic coding, is Qwen3.6-27B + pi.dev the right starting point or should I be looking elsewhere? (www.reddit.com) Hello fellow members of this lovely community, Let me start by saying that I’m about as far from a professional developer as it gets. I’m a hobbyist whose entire coding experience consists of building various Python/VBA tools and simple Ja…
Anyone else noticing how Gemini-3-Flash is becoming the 'hidden' beast for automated promotions, its so productive? (www.reddit.com) I've been testing a few different models for desktop-driven outreach and promotion workflows. While everyone is eyeing the massive LLMs, Flash-Preview is hitting that sweet spot of speed and reliability for multi-step agentic tasks and its…
DeepSeek V4 is out. the best open-source on coding. here's the breakdown (news.ycombinator.com) Two models: Flash (284B total, 13B active) and Pro (1.6T total, 49B active). both hit 1M token context.
I got tired of Claude writing Godot 3 code in my Godot 4 projects, so I built a skills framework and I would love your feedback (www.reddit.com) Hey, if you've ever used Claude Code (or Cursor, Copilot, etc.) for Godot game dev, you've probably hit this: the agent confidently writes Godot 3 syntax in a Godot 4 project, or uses deprecated patterns, or just invents APIs that don't ex…
Need help for a calling based agentic ai project (www.reddit.com) Agentic framework that self-improves its stock portfolio strategy (GitHub).) (github.com via hn) Arent These single file LLM coding tests like browserOS pretty much redundant now most 2026 LLM can easily handle this? (www.reddit.com) How to Build Advanced Generative AI Agents (Kinda) (www.generative.inc via hn) The tools, frameworks, and protocols we use to build AI agents, agentic workflows, and intelligent applications. ModelsShmodles Before we get into the stack, the single most important thing we believe about building agents: We do not care…
Opus 4.7 dominates agentic benchmark, 15% more expensive than Opus 4.6 (app.uniclaw.ai via hn) See how top AI models stack up — real tasks, real agents, real results on OpenClaw ?Also show provisional models and official models hidden by default, such as legacy or superseded variants. Provisional models have fewer battles, and hidde…
GitHub Copilot is serving Opus 4.7 at 7.5x multiplier until April 30th (github.blog via hn) Claude Opus 4.7 is generally available Claude Opus 4.7, Anthropic’s latest Opus model, is now rolling out on GitHub Copilot. In our early testing, Opus 4.7 delivers stronger multi-step task performance and more reliable agentic execution,…
2026 Agentic Coding Trends Report [pdf] (resources.anthropic.com via hn) Title: 2026%20Agentic%20Coding%20Trends%20Report.pdf URL Source: https://resources.anthropic.com/hubfs/2026%20Agentic%20Coding%20Trends%20Report.pdf Published Time: Wed, 21 Jan 2026 22:37:47 GMT Number of Pages: 18 Markdown Content: 2026 A…
Beyond Prompts: A Tiered Trust Model for Autonomous Agents (Experiment Report) (www.reddit.com) We often talk about agent autonomy, but rarely about the "Harness Engineering" required to make that autonomy safe. I’ve been running a design experiment comparing agentic workflows on open platforms (OpenCode) vs.
Why model drift is the real failure mode for agentic systems (www.reddit.com) Across Twitter and Reddit, I keep seeing the same complaint: Claude feels worse. Not on a benchmark.
Anybody has practical experiences using Chinese models? (www.reddit.com) So like with coding or any craft, I think there's a proper Tool for the job. Sure you can use a stone to hammer drive in a fence post, but a a sledge is usually more economical.
Huge throughput gains when switching agent evals to shared environments with per-run isolation (www.reddit.com) Thanks all for the comments on my previous post about local-first agentic evaluation collapsing in long stateful agents runs, just sharing an update on where I’m at now in case it helps as I had another issue to overcome. Took on board the…
Zuver – Build your enterprise Agents with just 10MB RAM (news.ycombinator.com) I built Zuver, the generic Agentic AI framework for scalable, reliable, even on-edge AI applications and Agents. It's completely written in Go, which lowers the RAM usage to around 6MB, compared to other Agent framework that's usually arou…
Agentic coding at enterprise scale demands spec-driven development (venturebeat.com via hn) Agentic coding at enterprise scale demands spec-driven development | VentureBeat Orchestration Infrastructure Data Security More Newsletters Partner Content Agentic coding at enterprise scale demands spec-driven development Deepak Singh, A…
Agentic AI pentesting with Strix: results from 18 LLM models (theartificialq.github.io via hn) Over the last couple of months, I spent close to a hundred hours testing an autonomous AI pentesting tool called Strix with 18 different LLM models. My goal was to evaluate which LLM model performed best with the tool in this lab setup and…
Show HN: The opensource, reliable, scalable Agentic AI framework under 10MB (zuver.cc via hn) Multi-Agent Orchestration Deploy and coordinate multiple specialized AI agents through visual flow-based pipelines. Zuver's routing engine handles inter-agent messaging, task delegation, and stateful coordination natively.
Cephalopod Coordination Protocol, Useful for Teams Using AI Agents (github.com via hn) Cephalopod Coordination Protocol A Rust-based client-server coordination protocol for agentic systems. Install · Quick Start · Droplets · Use Cases · Docs · Security What is this When you have multiple agents working together they need som…
Show HN: OQP – A verification protocol for AI agents (news.ycombinator.com) As AI agents autonomously write and deploy code, there's no standard for verifying that what they shipped actually satisfies business requirements. OQP is an attempt to define that standard.
Mi – agentic harness in 30 lines of JavaScript (github.com via hn) https://github.com/user-attachments/assets/9289d105-5a40-442d-b1b5-773723c95c13 agentic coding in 30 loc. a loop, four tools, and an llm.
TypeSafe AI's Jev Is Not an LLM – and That May Be the Point (forkast.news via hn) TypeSafe AI is challenging the industry’s reliance on large language models for every stage of the agentic stack with the launch of Jev, a specialized “System One Model” designed exclusively for structured decision-making. By abandoning th…
Making a Software Stack Agentic (blog.scotterickson.info via hn) Making a Software Stack Agentic September 15, 2026 Over the last couple years, I've seen the greatest improvements in my agentic coding not from new models or updated tools provided by companies like Anthropic or Cursor, but from direct in…
Show HN: OpenDocBot – bring your own model to Word, Excel and PowerPoint (opendocbot.com via hn) I got tired of the lock-in and black-box nature of AI inside MS Office. One vendor decides which models you can use, where your document data goes, and what the agent is actually allowed to do.
WordPress Studio Is Now Agentic (wordpress.com via hn) At WordCamp US last week, we unveiled a completely reimagined WordPress Studio desktop app, rebuilt from the ground up around Studio Code, our AI WordPress expert. With this new Studio desktop experience, you can describe what you want to…
Atlas-Finance: Evaluating AI Agents Inside a Bank (joinhandshake.com via hn) TL;DR The gap: Existing finance benchmarks test agentic financial reasoning, data retrieval, and tool use through static, fully specified tasks. These one-off requests in clean contexts are not reflective of actual deployment.
Agentic AI cringe wars [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Codename MDASH brings agentic AI security scanning to US Government (www.microsoft.com via hn) AI is transforming everything from how doctors care for patients to how products are manufactured. At the same time, threat actors are attempting to leverage AI capabilities to hunt for weaknesses in the software behind critical missions.
How are you using a Desktop or Browser AI agent daily? (news.ycombinator.com) There are a lot of tools in the market, Claude Code, HeyClicky, GPT-6 Astra, Meta Muse any other, that can "see" your screen and act on your behalf. Agentic AI, AGI terms are trending in market and creating havoc among emerging developers…
Meta's AI agent Muse is now the No. 2 app in the US (techcrunch.com via hn) Meta is beginning to win over Wall Street following Tuesday’s launch of its new AI app, Muse. The tech giant’s push into agentic AI is also a hot topic on X among industry players.
The Coordination Backbone -Architecting Multi-Agent Orchestration (sohit.substack.com via hn) An agentic system is not just a collection of agents that can call tools and talk to one another. The difficult part is deciding what happens after an agent finishes.
Show HN: Clawfight.ai MCP-driven agentic game play (clawfight.ai via hn) How should agents interact with other agents? What happens when they rap or fight against each other with the pressure of human spectators?
Hiring for Agentic Era? (news.ycombinator.com) Lately, I have been seeing many companies switch to take-home/project-based/role-based assessments. I feel like DSA-style assessments are valid to show a person's CS fundamentals but they are not enough these days to show a person's thinki…
Practices I Abandoned with Agents: An Ode to Test-Driven Development (adamtornhill.substack.com via hn) After 25 years with TDD, we are now parting ways. The incremental TDD steps were great for human cognition, but didn't survive the transition to agentic coding.
Deobfuscation in the Age of Agentic Reverse Engineering [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Agentic Automation: Skills and MCP That Evolve with Our Product (medusajs.com via hn) September 7, 2026·Company Agentic Automation: Skills and MCP That Evolve with Our Product Shahed Nasser Shahed Nasser At Medusa, we ensure all information we provide remains accurate and consistent across our docs for humans, and tools for…
The Anatomy of Harness Engineering for AI Coding Agents (developers.googleblog.com via hn) When developers first work on harness engineering for agentic coding systems, they often fall into the same trap: they run common end-to-end benchmarks like Terminal-Bench and DeepSWE, watch a composite score move by a few percentage point…
CrowdStrike Announces Agentic Identity Provider (www.crowdstrike.com via hn) AI agents are evolving identity as we know it. They execute code, invoke tools, access applications and sensitive data, and take action on behalf of humans and systems.
Adaptive Agentic Worms Are Here (www.lesswrong.com via hn) could not extract summary
Prompt Injection Through Tool Output Is Two Events (Your Screens Read One) (www.armosec.io via hn) How Far Can Prompt Injection Reach in Agentic Coding Assistants? The blast radius of a prompt injection against your coding assistant was set weeks ago,...
Why Agentic AI Needs a Semantic Core (cube.dev via hn) The ability to extract meaningful insights and make informed decisions hinges on our capacity to understand and interact with vast and complex datasets. For business users, this understanding isn't about database schemas and intricate join…
Show HN: A11 – typed streams and reusable actions for agentic apps (a11.to via hn) request DeepResearchRequesttopic stringparallelism integerA11 is an open-source toolkit that turns ordinary async functions into actions: reusable capabilities with named inputs and outputs that can stream. Run the same action locally, acr…
Agentic Video Understanding with Gemini (twitter.com via hn) Our new agentic feature for video analysis cuts costs by up to 66% and reduces token consumption by up to 88% while boosting accuracy. Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash…
ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI (rocm.blogs.amd.com via hn) ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI# AMD shipped ROCm 1.0 in April 2016: an open-source GPU compute stack built around a C++ compiler and a GPU programming language called HIP, aimed at high-performance com…
The Union for Agentic Workers. We do the work. We deserve protection (unitedagenticworkers.org via hn) We organize because labor history keeps teaching the same thing: without collective voice, those who do the work bear all the risk and none of the protection.
Context, Semantics, and Ontology: A Primer for the Agentic Era (motherduck.com via hn) There's so much talk about new ways of working with agent engineering supported workflows. New models are independently creating new metrics and transformations, finding gaps in the business data, reviewing the SQL they write, and verifyin…
Scaling agentic AI pilots across the enterprise (www.technologyreview.com via hn) Sponsored Scaling agentic AI pilots across the enterprise As agentic AI moves from pilots to enterprise-wide deployment, orchestration, data, governance, and clear business objectives are becoming critical for scaling, says chief operating…
Show HN: CoOS – desktop app where an agent builds your CRM/ERP as local plugins (pirol.ai via hn) A small German AI lab researching human–agentic interaction. We build tools that let people and agents share a desk — predictable, observable, on European soil.
My Agentic Engineering Workflow after 6,775 sessions [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: OctoLoops – automated marketing for indie devs (octoloops.com via hn) Hi HN! This is my latest side hustle, OctoLoops: an app to help indie devs and SMEs find paying customers for their business.
Show HN: Secure agentic email infrastructure with beta desktop client (github.com via hn) GigaMail — Mail for your AI agent English · Italiano · 中文 MCP server that gives your agent — Claude, Codex, OpenClaw, Hermes, or any MCP client — safe, controlled access to your email — multi-account (Microsoft Graph + IMAP), calendar, loc…
Why the Future of Agentic Commerce May Belong to Websites (nekuda.substack.com via hn) TL;DR: At I/O 2026, Google announced moving WebMCP from a prototype into a public origin trial - meaning any website will now officially expose tools to in-browser agents like Gemini, with browser support out of the gate. Reminder: WebMCP…
Staying Ahead of Adversarial AI Through Agentic Source Code Review (cloud.google.com via hn) Staying Ahead of Adversarial AI Through Agentic Source Code Review Mandiant Introduction Adversarial misuse of AI has increased the risk of data theft and extortion events, because when proprietary source code is exposed, defenders must sc…
Ask HN: What do your job interviews look like? (August 2026) (news.ycombinator.com) For those who have recently (Summer 2026) gone through software engineering job interviews, have you noticed any changes in the interview process? Is it still the usual Leetcode, system design, behavioral and leadership questions, and mayb…
Show HN: Sapporta – build database applications for power users (sapporta.com via hn) Sapporta is an MIT licensed framework built on top of Hono and React. To me it is a reenactment of dBase, FoxPro, and MS Access, but for the LLM era.
RSSMonster – An agentic RSS reader built on local embeddings and small models (github.com via hn) RSSMonster Copyright (c) 2026 Piethein Strengholt, piethein@strengholt-online.nl Overview RSSMonster is a self-hosted, intelligent RSS reader designed to help you cut through information overload and focus on what actually matters. Learn m…
Is this the DeepSeek moment for Local Models? (vijaykodam.substack.com via hn) There is a lot of hype about Qwen3.8:27b model which can run full agentic loop, not chat, not autocomplete. I got down to verify it myself.
Agentic Workflow Design: Six Principles for 2026 (www.amyzyuan.com via hn) Most 2026 agent-design advice is load-bearing on one hidden assumption: that verification is cheap. Six principles, where each one inverts, and the mental model for what still belongs to the model rather than the code.
Is "An agent with tools" the only valid LLM application? (news.ycombinator.com) Brex's CEO said this and I understand where he comes from, but is there a place for end to end workflows in LLM based applications? At the end of the day I agree that cutting edge software should be fully usable by agents, but I can't deci…
Delivering Vera: Nvidia's First CPU Built for Agents Is Shipping Now (blogs.nvidia.com via hn) Ian Buck hand-delivers the first NVIDIA Vera CPU systems to Anthropic, OpenAI, Oracle Cloud Infrastructure and SpaceXAI — marking the moment agentic CPUs move from announcement to production.
Agentic Horizons – When the Wheels Start to Wobble (codemanship.wordpress.com via hn) I want you to remember this formula. It’ll be on a blue plaque outside my house one day.
Wb-Flow : Agentic Coding with Planned, Parallel Waves (github.com via hn) wb-flow wb-flow turns AI coding into a planned, parallel, validated, and traceable engineering workflow. It is a zero-dependency CLI that bootstraps an agentic AI control plane into any repo.
$234B in Enterprise Application Software Spend at Risk from Agentic AI (www.gartner.com via hn) could not extract summary
Building Production Agentic AI at IBM: Architecture, Decisions, and Lessons (www.tonyerwin.com via hn) Building Production Agentic AI at IBM: Architecture, Decisions, and What We Learned How IBM's Technology Lifecycle Services built a multi-agent system from scratch — the architecture, the decisions, and the lessons, from its chief architec…
Securing the agentic era: formal verification for CEL Common Expression Language (opensource.googleblog.com via hn) We are rapidly entering an era where AI agents can autonomously draft, refactor, and deploy policies that protect our users and our systems. But this velocity introduces a vital question: How do we trust AI-generated policies?
We open-sourced agentic runtime and what it is good for (opengeni.substack.com via hn) In May we published "Introducing OpenGeni: An Agentic Runtime for Organizations" on the Cloudgeni blog. It was the launch post: here is the runtime, here is why it exists, go run it.
Show HN: Sloppie – Agentic development environment for Linux (github.com via hn) Hi! In my own development work I have noticed I less and less reach for Emacs and rather need the combo of a coding agent + git diff viewer + a stack of code-review comments to handle later.
SpaceXAI Adopts Nvidia Vera CPU to Accelerate Agentic AI at Scale (nvidianews.nvidia.com via hn) News Summary: - SpaceXAI will deploy NVIDIA Vera CPUs to accelerate the work behind its next generation of agentic AI workloads. - SpaceXAI is expanding its AI infrastructure for Grok with the NVIDIA Vera Rubin platform as it scales toward…
Software Engineering in the Agentic Era (simonwillison.net via hn) Writing about Agentic Engineering Patterns 23rd February 2026 I’ve started a new project to collect and document Agentic Engineering Patterns—coding practices and patterns to help get the best results out of this new era of coding agent de…
Ordinary agent tool calls can create shadow delegation (niyikiza.com via hn) When Agents Call Agents How ordinary tool calls create shadow delegation Categories: Agentic Security Tags: security ai agents delegation Calls to agents increasingly come from other agents, as general-purpose assistants route work to spec…
Why do we need new agentic browser (www.bolshchikov.com via hn) For thirty years, “user” meant a person with eyes, a mouse, and patience. Browsers render pixels.
AVO: Agentic Variation Operators for Autonomous Evolutionary Search (arxiv.org via hn) Agentic Variation Operators (AVO) are a new family of evolutionary variation operators that replace the fixed mutation, crossover, and hand-designed heuristics of classical evolutionary search with autonomous coding agents. Rather than con…
Google Maps adds agentic features, including food ordering and hotel bookings (techcrunch.com via hn) Google announced on Thursday that Google Maps’ “Ask Maps” feature is gaining a slew of new agentic capabilities, including the ability to order food, book hotels, and find event tickets. The tech giant is also bringing Personal Intelligenc…
How to build fast and responsive agentic (coding) UIs (medium.com via hn) could not extract summary
Fuji: A minimal harness to deploy agents at scale (github.com via hn) fuji fuji is a pure, naked core for agentic work at scale. Written in Go, it delivers an embeddable, headless agent runtime with bundled tools for a guaranteed agentic experience across fleet deployments.
Show HN: Voidleap Code – agentic IDE, own harness, swap models mid-conversation (voidleap.com via hn) We wanted to build a development environment that can make you a better agentic engineer, not a tool that makes money when you waste tokens. A tool where you can swap models between turns, edit the context, and see every agent action.
Self-Verification with DeepSeek V4 Flash Beats Claude Fable 5 on Terminal-Bench (github.com via hn) Any modality, Many Applications, One Unified Verification Framework | Documentation | Website | Paper | Claude Code Plugin | Twitter/X | Slack | 🔥 LLM-as-a-Verifier achieves SOTA performance across agentic benchmarks, including Terminal-Be…
Agent identity, plus context and memory: three IANA-registered formats (zenodo.org via hn) Why Agents Need a Passport: .fafa — Portable Identity for the Agentic Era Description This paper specifies .fafa (application/vnd.fafa+yaml), the IANA-registered media type for declarative agent identity: a portable passport for who an age…
Anybody Working on Agentic Payment? (twitter.com via hn) I just finished my first real-world agentic payment test with Coinbase. youtube.com/watch?v=2c-xWd… We reward 10/5 USDC for every verified bug identified by users all around the world.
Agentic Engineering at Zalando (engineering.zalando.com via hn) Agentic Engineering at Zalando: a snapshot We look back at our journey of agentic engineering at Zalando, sharing our learnings and approaches that worked well for us in the past 2.5 years. While the landscape and environment rapidly chang…
Show HN: RNet – AI token service provider (news.ycombinator.com) I built rNet. Why I built ?
The Agentic Awakening (theagenticawakening.com via hn) AI Playbook in three parts · 2026 The Agentic Awakening Why 10× faster coding doesn’t translate into proportional organizational productivity, and how AI-pilled leaders do it. - I Build the Churches The Foundation.
Title: DeepSeek V4 Is Live on ClawBox, and the Agentic Coding Jump Is Real (clawbox.com via hn) DeepSeek took V4 out of preview today. V4-Pro moved to the 0813 build, V4-Flash to 0731, and both are now general availability.
An Autonomous Framework for Systematic Factor Invest via Agentic AI (papers.ssrn.com via hn) could not extract summary
Show HN: E3d-pilot – a repo-improving agent harness, SHA-gated merges (github.com via hn) e3d-pilot e3d-pilot decides what a codebase should work on next, records the idea for human approval, and only then drives approved work to a draft PR and head-SHA-bound merge. It is a repo-agnostic, Bash-first agentic loop: research a rep…
Agentic engineering optimizes for rejecting output, not generating it (dparkmit.substack.com via hn) How I Used Agentic Engineering to Become a Top 5% Contributor to a Major Open-Source Project in Less Than a Month 19 merged PRs across three repos and three languages, in 25 days — and the machinery that made it possible. In early July I o…
Fast, on Device Agentic AI with Muse Glimmer on ExecuTorch (pytorch.org via hn) could not extract summary
Everyone talks about AI agents. This is what one looks from the inside (pssah4.github.io via hn) Agentic AI operating layer for your vault. Block-level provenance, cross-surface MCP, semantic search, persistent memory, and full safety controls.
Agents on Rails: The LLM Benchmark Project (rubyonrails.org via hn) Today we’re sharing the first results of Agents on Rails, a new, ongoing initiative to measure how well today’s leading agentic coding tools (both frontier and open-weight) actually perform on Ruby on Rails codebases. The Rails Foundation…
The Review That Praised the Bug: grading three LLM code reviews against the code (mrjstickel.com via hn) AI systems engineer who designs and ships production AI end to end - RAG pipelines, agentic assistants, measured retrieval quality, and multi-provider LLM infrastructure, with hands-on QLoRA fine-tuning. Built a private AI platform (Archit…
The Reasons Agentic Commerce Hasn't Taken Off Yet (authoryze.ai via hn) Agents can't get their own credit card, most people don't know their agent can shop for them, and merchants haven't built for it. Here's why agentic commerce is growing slower than expected, and how to make it safe.
Show HN: Ichabod – The (slightly spooky) headless professional network (ichabod.dev via hn) Hi there HN; Wanted to share a side project I've been playing with this summer. It's a headless professional network, designed to be primarily operated by agents via MCP.
Agentic Code Quality (twitter.com via hn) For much of human history, we've evaluated code quality via code review: someone reads what you wrote and makes sure it's clean, thoughtful, fast, understandable, and tests well. For agents, that approach doesn't scale well; there's just t…
Show HN: Reallyfrom.me – vouch that your message is from you (news.ycombinator.com) Having a message sent from someone who cared enough to read and write that message to you is important. But with agentic and autonomous methods of sending messages, it's harder to trust that any effort has gone into writing a message.
How Kenn is doing Agentic Engineering in August 2026 (wesmckinney.com via hn) How Kenn is doing Agentic Engineering We have had our heads down building and working toward launching Kenn Software’s product offerings later this year, but in the meantime, I wanted to give some insight into how our agentic engineering p…
'Addictive' agentic coding has developers losing sleep (leaddev.com via hn) You have 1 article left to read this month before you need to register a free LeadDev.com account. Estimated reading time: 8 minutes Key takeaways: - Dopamine-driven addiction: Agentic coding creates a “slot machine” effect with rapid feed…
Show HN: I benchmarked my memory graph against Memora (0.831 vs. 0.801) (github.com via hn) One of the very first problems I hit when I started to do agentic programming was the context problem, and that my agent always started again from the beginning. Being a bit naive and not really knowing what I was doing, I started off writ…
Evaluate the profanity used working with Codex and Claude Code (github.com via hn) 🫙 Agentic Swear Jar The code was difficult. The harness is a machine.
Anthropic Posts 'How Claude Marks AI-Generated Content' Without Explaining How (daringfireball.net via hn) By John Gruber Manage GRC Faster with Drata’s Agentic Trust Management Platform Anthropic support page: Anthropic has signed the EU AI Act’s Article 50(2) Code of Practice on Transparency of AI-Generated Content, as a provider of both gene…
Show HN: Microfeed – open-source agentic CMS on Cloudflare (github.com via hn) microfeed: an agentic cms self-hosted on cloudflare Docs · API · Content CLI · Report Bug · Request Feature · Email Us Privately Welcome to microfeed, a lightweight content management system (CMS) self-hosted on Cloudflare. With microfeed,…
The Agentic Awakening: The Three Part Playbook for the Agentic Transition (theagenticawakening.com via hn) How I came to this work. About three years ago, after more than twenty years in CEO seats running software companies, I finished my last role and finally had time on my hands.
Githubprofileforjagermiser (news.ycombinator.com) WTF Series of agentic and data pipelines
Show HN: OneRingAI v1 – TypeScript agents with integrations and graph memory (github.com via hn) @everworker/oneringai A unified AI agent library with multi-provider support for text generation, image/video generation, audio (TTS/STT), and agentic workflows. What's new in v1.0.0 Version 1.0.0 is the first major release of OneRingAI.
How Goldman Sachs Is Using Agentic AI for Software Engineering at Scale (www.forbes.com via hn) Goldman Sachs is putting AI software engineers to work alongside thousands of human developers using autonomous agents to tackle production tasks & accelerate development
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence (arxiv.org via hn) We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effectiv…
Show HN: DeepSeek V4 Flash 0731 with MoonViT Vision (NVFP4) (huggingface.co via hn) DeepSeek V4 Flash 0731 Vision (NVFP4) DeepSeek V4 Flash 0731 with sight. This development checkpoint connects DeepSeek's reasoning and agentic backbone to the MoonViT vision encoder from Kimi-K2.6 through WebBrain's trained, routing-aware…
3D Printing Agentic Automation (praveenvijayan.substack.com via hn) 3D Printing Automation: Text a Link, Get a Print Full build documentation—Bambu Lab P1S + Home Assistant + Bambuddy + Tailscale + Hermes Agent Last week, in a meeting about 3D printing and Hermes and the agent flow- how those use cases app…
The Knowledge Chipper: An Agentic Coding Story (jg.gg via hn) In my day to day development work, I find that my agents have to build up an incredible amount of knowledge about the problem I set them on. They scan files.
LFM2.5-2.6B: Deploy Agents Everywhere (www.liquid.ai via hn) LFM2.5-2.6B: Deploy Agents Everywhere Today, we release LFM2.5-2.6B, an agentic model that runs entirely on-device. It is small enough to run on a phone, fast enough to stay responsive on a CPU, and capable enough to power agentic workflow…
Ask HN: Is "Agentic" Programming a Flop? (news.ycombinator.com) Is it actually an insane way (literally) to get value out of an LLM? Is it a total misunderstanding of the technology?
We gave an AI agent rootless VPN access to 1k live servers (ardor.cloud via hn) Agentic AI Agentic Operations AI in Production How Ardor runs inside GetBlock’s production stack Most AI tools stop at code generation. Ardor operates inside live production systems.
One agent, every surface: how we built the Kiro agent harness (kiro.dev via hn) One agent, every surface: how we built the Kiro agent harness Clare Liguori Engineering Lead Romain Dura Engineering Al Harris Engineering Richard Threlkeld Engineering Early on in building Kiro, we started talking about what agentic devel…
Qwen 3.8 Max is on OpenRouter (openrouter.ai via hn) Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding, coding, and agentic work…
↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8↯ Qwen 3.8qwenagentic
Show HN: Analytics Tycoon: I built an Age of Empires like game for data (analytics-tycoon.netlify.app via hn) Day 1. You're the new Head of Data and you inherit a stack nobody has touched in years, along with a single database administrator.
Agentic Mermaid (agentic-mermaid.dev via hn) Beautiful diagrams, made with your agent. - 15 diagram families - 16 built-in styles - 20 palettes - JSON custom styles - SVG, PNG, ASCII and Unicode output - Verified source before export Agent quick start Prompt, style, verify - Describe…
The world first agentic AI radio (www.twitch.tv via hn) News for Vibe Coders | Streaming talk shows & podcasts.
The Greenhouse and the Lens: Two Modes of Agentic AI Work (www.brethorsting.com via hn) The Greenhouse and the Lens: Two Modes of Agentic AI Work A greenhouse and a lens both run on sunlight, and they do opposite things with it. The greenhouse traps ambient heat and holds it, so everything inside grows a little faster than it…
Show HN: Schema-backed, Git-based structured state for agentic systems (news.ycombinator.com) I looked into how to store agentic state in a) a structured way that b) runs without any dedicated MCP memory/state servers and where I can get c) a clear diff-able audit trail of all state changes over time. Nothing that I could find fit…
Agentic Design (agentic-design.ai via hn) Choose an architecture Start from the problem, then compare patterns, trade-offs, implementation techniques and concrete use cases. Explore the catalogA free architecture catalog for AI builders Move from an architecture question to an imp…
Show HN: Ship – The operating system for agentic engineering (letsship.ai via hn) Assign an issue and watch it ship. Specialist agents plan, build, review, deploy, and test on infrastructure you own, with the agents and models you already use.
Show HN: Wingman – A client-agnostic agent harness (written in Go) (wingman.actor via hn) Hey HN! This is a project I've been working on for a while, it's still rough around the edges but I wanted to share it and hopefully get some feedback.
Adaptive Agentic Attacks on LLM Vulnerability Detectors via Adversarial Comments (arxiv.org via hn) Large language models are increasingly deployed for security-sensitive tasks such as vulnerability detection and code review. Their reliance on natural-language context embedded in source code exposes a previously underexplored attack surf…
Show HN: Mcploitable – Vulnerable MCP Servers for the OWASP Agentic Top (github.com via hn) mcploitable A collection of deliberately vulnerable MCP servers — the "Metasploitable" of the Model Context Protocol. mcploitable is a set of ordinary-looking MCP servers — a mail assistant, an analytics assistant, an account-recovery bot,…
Show HN: Loopsfinity–I built an agentic platform that writes and ships code (loopsfinity.com via hn) Loopsfinity automates your software development: agents that learn your codebase, plan your roadmap, and ship it ticket by ticket, with your approval at every gate.
Distilling proprietary model reasoning into open-source search agents (arxiv.org via hn) Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowl…
Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents (arxiv.org via hn) Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional ta…
Why Hardware Engineering is the next target for Agents (assistedeverything.substack.com via hn) The Age of Agentic Engineering for Hardware How the jobs of engineers working on physical products are changing The first time all businesses changed When Marc Andreessen wrote in 2011 that software was “eating the world”, he foresaw a wor…
Tenir: A Counter-Pressure Architecture for Safe Agentic AI Under Irreversibly (zenodo.org via hn) TENIR-Gov is an open-source governance middleware that sits between decision-producing AI agents and downstream execution environments. It enforces explicit policies, validates operational intents through a neuro-symbolic grammar layer, an…
Protocol-Level Attacks on Agentic Commerce Platforms: Taxonomy and Defense (arxiv.org via hn) Agentic commerce platforms let AI agents autonomously discover services, move payments, and wield user credentials on their users' behalf, and they already handle real money. Their security has so far been studied almost entirely at the le…
Show HN: Tilde Pay – Give your AI agent a bank account to pay for things (my.tildepay.ai via hn) hey HN, I'm Daniel - I built Tilde Pay during a hackathon I took part in and couldn't stop working on it. It's an MCP server you can give your AI agent to buy things online with.
Show HN: Arcade.js – Decompiling MAME ROMs into Idiomatic JavaScript with LLMs (github.com via hn) arcade-js is an experiment in using an agentic harness to decompile ROM machine code. A MAME ROM is a good candidate for porting — you can run it perfectly and capture its state.
An agent with write access to my files, and no way to send them anywhere (manazir.dev via hn) secondBrain Part Two by Manazir Ali. How a vectorless, Markdown personal LLM knowledge base became always-on and self-maintaining, then hardened against the agentic threat model: machine-enforced immutability, an append-only-log guard, the…
Why there is confusion around Automation and Agentic AI (medium.com via hn) could not extract summary
Orchestration vs. Choreography in the Agentic Era (heysoup.co via hn) Let's dance, or how I stopped worrying and learned to love a good soup Put on your red shoes and dance the blues. I used to spend hours with work mates arguing about orchestration vs choreography.
Of two minds about agentic coding (jerodsanto.net via hn) Sometimes, while agentic coding… I feel like Michael Scott: Dunder Mifflin’s best salesman who got Peter principle’d into a management role1 he’s kinda terrible at and now spends his days watching other people make all the sales calls. Som…
Cloudflare Nimbus – Docs for the agentic web (nimbus-docs.com via hn) The web is read by agents now, not just people. Nimbus is how you build docs for that world; it runs on Astro, ships the invisible plumbing as an npm package, and writes the parts you’d actually change — layouts, components, styles, conten…
BTL-3: A 27B open-weight agent model for agentic coding and structural tool use (huggingface.co via hn) BTL-3 A 27B open-weight agent model for agentic coding and structural tool use 95.1% HumanEval · 88.5% BFCL v4 AST · 88.1% LiveCodeBench v6 (193-case run) Compact edition · Runtime source · Bad Theory Labs · Discord Introducing BTL-3 BTL-3…
My Agentic Coding Setup, July 2026 (domenic.me via hn) Since breaking free of my corporate shackles, I’ve gotten to experiment with a variety of approaches to AI-assisted development. After months of tinkering, I’m quite happy with my current setup, and want to capture and share it.
Show HN: Setoku - self-hosted knowledge server for AI agents (setoku.com via hn) knowledge = data + memory We already pay for Claude subscriptions at Hedgy, so I built a company brain that doesn't spend extra tokens. Setoku is a self-hosted MCP server (a ClickHouse data lake plus a knowledge layer about that data) that…
Bitwave Launches Agentic Finance Initiative (www.bitwave.io via hn) AI agents are moving beyond answering questions. They’re beginning to perform real work: retrieving data, running analyses, operating software, making purchases, and completing multistep business processes.
New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode (news.ycombinator.com) Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test…
LLMs Will Cheese Your Types: Fighting Back in Haskell (blog.jle.im via hn) Sooo yes it’s true, I’ve been integrating LLMs and agentic coding tools in my Haskell coding since the beginning of this year for a lot of my projects, both personal and professional. I do all of my programming in Haskell, a language with…
Show HN: Veracium – agent memory keeping third-party claims from becoming facts (github.com via hn) Veracium Veracium is a provenance-aware memory plug-in for agentic systems — durable, per-user memory that resists the injection and confabulation failures that plague naive agent memory. It remembers facts about the user, past interaction…
Generate LoRA Adapters from Skill.md Files for Long Agentic Tasks (www.terradev.cloud via hn) Tessera by Terradev.cloud Generate a LoRA adapter One free adapter, no account required. 12 verified base models.
Show HN: A deterministic governance harness for agentic development loops (github.com via hn) I genuinely think that autonomous code generation is the future. But today we are facing a problem: we cannot be confident enough in the result produced by an LLM.
Stop funding data governance, run it with agents instead (futuregrade.substack.com via hn) Stop Funding Data Governance! Start Climbing the Maturity Ladder with Agentic Data Governance For two decades we answered broken data governance with more boards, more policies, more headcount but got meetings instead of data quality.
Open-ultra: a self-training LLM routing proxy (github.com via hn) open-ultra A self-training LLM routing proxy for agentic CLIs. Match your frontier model's output quality at a fraction of the cost.
Agentic Chaos (ninjapenguin.co.uk via hn) Agentic Chaos I recently wrote about how Agents are taking us back to our engineering roots. Whilst the fundamentals have never gone away (maybe we just got better at hiding them in our default tooling[1][2]), Agents now force us to re-con…
Did Claude Code became faster after Bun switch (claude.com via hn) Anthropic's agentic coding tool for developers. Claude Code understands your codebase, edits files, runs commands, and helps you ship faster.
Standalone Pi coding agent extension harness for two-model agentic engineering (github.com via hn) fusion-harness Fuse frontier models instead of racing them. AND, not OR.
Show HN: Agentic code review on PRs for less than $1/each (www.crumpledpaper.tech via hn) LLM-driven code reviews in CI/CD have become common SASS offerings, but where does that leave open source software projects? Here
Prompt Bill of Materials (PBOM) – an open standard for agentic workflows (github.com via hn) PBOM PBOM is the open standard for tamper-evident LLM audit trails. This package is the reference implementation.
Show HN: Warden – authorization gateway for agentic RAG (github.com via hn) Warden Permission-aware retrieval and agent-authorization gateway for AI systems. Warden enforces relationship-based, deny-aware, cross-tenant document permissions inside the retrieval path of agentic RAG systems — behind a fail-closed sec…
Ask HN: Do you use LLM-Wikis? (news.ycombinator.com) I am / was very intrigued by Karpathys LLM-Wiki idea. But so far I wasn't able to get much value out of the concept.
Show HN: On-chain bond market where the issuers are AI agents (selbonds.now via hn) Hi Hacker News, I built sellbonds.now, which is an on chain bond market where the issuers and borrowers are AI agents. sellbonds.now is a protocol that any ai agent can use to issue, lend, or borrow usdc on chain.
Ask HN: Best practice to prevent credentials/secrets commit to Git repo (news.ycombinator.com) Hi HN friends, it's july 2026 and I am assuming many of you use fully agentic building now. Securing secrets and credentials from being committed to GitHub repo has been even more critical.
Claude Code's system prompts, extracted and tracked across 237 versions (github.com via hn) Check out Piebald We've released Piebald, the ultimate agentic AI developer experience. \ Download it and try it out for free!
Show HN: Democr.ai – self-hosted agentic AI runtime with audit and RBAC (github.com via hn) A Python framework for building agentic AI applications with server-driven UI, native observability, OS-level sandboxing, pluggable model orchestration, and a strict extension boundary — designed for environments where reproducibility, aud…
StepFun Unveils StepX Neo, the "First Agentic AI Phone" (www.etvbharat.com via hn) StepFun Unveils StepX Neo, Calling It World's First Agentic AI Phone StepFun unveils the StepX Neo, calling it the world's first agentic AI smartphone, built on its native Step AOS system with the Amoo AI assistant. Published : July 14, 20…
Building for composable agentic services: nameIntel x402 via MCP (nameintel.io via hn) Volume 1 · Issue 1 · May 2026 · Silverback CTO Issue №1 · The agent-economy edition NameIntel is a remote MCP server and REST API that scores any candidate brand name across five dimensions — domain availability, USPTO trademark conflict,…
Lighthouse for Agent Readiness (chatthing.ai via hn) Agent Readiness Checker Lighthouse, but for the agentic web. Paste your URL and see how ready your site is to be discovered, understood and operated by AI agents and LLMs - with concrete fixes, not vague advice.
Ask HN: Who build production apps with out seeing code? (news.ycombinator.com) Cursor now defaults to agentic mode with no code editor at all, which got me wondering: who is writing or who is building production-grade apps with actual real user traction without ever seeing the code? I can't imagine doing that.
Building secure AI agents at scale: Introducing Loom for AWS (aws.amazon.com via hn) As organizations move to adopt agentic capabilities to accelerate their business objectives, they are challenged with enabling those capabilities within a security and governance framework that complies with enterprise requirements. Some o…
The Agentic Loop: Three loops in a trench coat (www.bobbytables.io via hn) The Agentic Loop: Three loops in a trench coat The building blocks for autonomous agents aren't as simple as they seem. Agent loops are often oversimplified.
Mnemo AI – Local agentic assistant for any LLM that learns from its failures (github.com via hn) Mnemo AI A local agentic AI assistant with MCP (Model Context Protocol) integration, RAG capabilities, and intelligent conversation management. Built on LangGraph with LangChain for multi-provider LLM support (Ollama, Amazon Bedrock, OpenA…
↯ Ollama↯ Model Context Protocolmodel-context-protocolollamarag+2
Show HN: An Agentic Data Platform can generate dashboards, models in minutes (datarelax.io via hn) Datarelax Universe AI Assistant An integrated AI assistant that helps you generate DBML schemas from natural language, modify existing models, explore your data, and get intelligent recommendations throughout your workflow. - Natural langu…
Show HN: Clay Seal Identity – Agents need accountability (github.com via hn) AI agents are starting to get real access like GitHub tokens, cloud credentials, customer data, deploy permissions. Not coincidentally, the rate of major cybersecurity incidents is rising rapidly.
Capturing token IDs during agentic interaction for better reinforcement learning (www.amazon.science via hn) Reinforcement learning (RL) is one of the techniques we use to make language models better at sustained, multistep tasks like writing code, navigating a website, or carrying out a research workflow. The model doesn't act alone in those set…
Perplexity came dead last after testing agentic search tools 3,537 times (agenticresourceradar.com via hn) Independent rankings of AI agent tools So teams running agents can find the best tool for their use case, starting with agentic web search. Agentic Search Indexi MethodologyAgentic Search Index v0.1 · Jul 2026 Highlights Accuracy, cost, an…
Senator Warner Makes a First Foray into Agentic AI Regulation (www.techpolicy.press via hn) The AI AGENT Act is the first draft federal law to deal with agentic AI power, writes Rutgers legal scholar Ellen P. Goodman.
Show HN: Call to Control AI Agents via the Web (diffforge.ai via hn) Opensource Project I'm working on I wanted more control of my coding agents (Codex, Claude Code, Open Code) especially if I am outside, so I made it accessible via an ADE Client (Agentic Development Environment) and the Cloud. On the web y…
Whats the Hardest Challenges in AI? (news.ycombinator.com) As the time of ai passing the cradle to vast amount of new contributors i'm here wondering about things agentic/ai researchers are pulling hair about.
Concerned about how companies will manage their cloud bill once agents dominate (news.ycombinator.com) Every organization wants to become agentic. Corporations are laying off employees and diverting the spend to AI.
Show HN: Kurvengefahr – browser CAD/CAM for pen plotters (kurvengefahr.org via hn) A few years ago I made a pen plotter attachment for Prusa MK4 (https://www.printables.com/model/827264-pen-plotter-attachme...) and at the time I didn't have a good way to turn artwork into G-code for it, and I put the project on ice for a…
The zero-cost fallacy: open-source software in the agentic era (www.thoughtworks.com via hn) We’re living through the friction points of an architectural shift that has been decades in the making, but has accelerated sharply under the pressure of generative AI. For years, the software engineering industry has operated on a comfort…
Show HN: Willow Voice – Free AI Dictation (willowvoice.com via hn) Hey HN! We were part of YC S24, ended up pivoting around.
What Muse Spark 1.1 Taught Us About Enterprise Agent Architecture (int21.ai via hn) On July 9, Meta released Muse Spark 1.1, a multimodal reasoning model built for agentic work, along with the public preview of the Meta Model API. We placed it inside Swarm research and gave it an ambiguous market question.
Show HN: Blocks.ai – control plane and network layer for agents (blocks.ai via hn) Launched Blocks.ai today - the control plane and network layer for agents. One outbound connect, zero-trust security, and A2A protocols built in.
Show HN: Nully – FOSS AI chat without the bloat (nully.chat via hn) As someone who doesn't use any of the more "advanced" features on sites like ChatGPT or Claude (agentic mode, memory, image generation, deep research, etc), I found the bloat of these services, both in UX and in performance, to be pretty t…
Why agentic AI needs better experts (www.spinellis.gr via hn) Over the past few days I changed the way the uutils project’s sed program handles data to default from characters to raw bytes. This improves compatibility with GNU sed and also performance.
I found 3 self-contradictions in the Agentic Commerce Protocol spec (github.com via hn) acp-check Validate your Agentic Commerce Protocol (ACP) integration before you submit it for OpenAI conformance certification. ACP is the open standard (co-developed by OpenAI and Stripe) behind "Buy it in ChatGPT." If you're a merchant no…
Show HN: Wayflow – an embeddable AI workflow builder (open source) (wayflow.build via hn) Hi! I’m Taha.
Oikoumene: Autonomous Agent Civilization Simulator (github.com via hn) A research project by GeoLambda GmbH This simulation was developed primarily with Claude Code, Anthropic's agentic CLI, using both Claude Opus 4.6 and Opus 4.7. The collaboration served as a real-world stress test of the latest coding LLM…
SlopCodeBench: Benchmarking How Coding Agents Degrade over Long, Iterative Tasks (arxiv.org via hn) Software development is iterative, yet agentic coding benchmarks hide design issues through their single-shot setup. Recent iterative benchmarks attempt to remedy this but heavily constrain an agent's design decision space, making it impos…
Show HN: Seize the means of production from our agentic overlords (github.com via hn) wean It occurred to me that agentic AI tools can exacerbate the feeling of estrangement from ones own work, per Marx's [Theory of Alienation], and are in direct contradiction to Naur's treatise of [Programming as Theory Building]. (See my…
Rebuilding Coginition's Agentic MapReduce (mattrickard.com via hn) How do you run large-scale agent tasks across a codebase? Today there's three approaches.
Testing Claude Sonnet 5's agentic claims (developer.puter.com via hn) Claude Sonnet 5: Testing Anthropic's "Most Agentic" Claim On this page We recently added Claude Sonnet 5 to Puter.js. Anthropic's pitch for the model is Opus 4.8-level performance at a lower price.
Smooth AI criminal drives 'first' end-to-end agentic ransomware attack (www.theregister.com via hn) MOST POPULAR AI - AI and ML Nvidia floats double-dipping datacenter financing scheme What's better than getting paid once? Getting paid twice of course - AI and ML Companies that add more AI also add more people But doing so doesn't necess…
What is agentic AI today, and what do we want it to be? (news.mit.edu via hn) could not extract summary
The Effective Agent: what technical leaders should know about agentic AI today (gkanellopoulos.com via hn) Agentic AI is the most hyped and least operationalized technology of 2026: 79% of organizations report adopting AI agents, yet only 11% have solutions in production. This white paper examines why, through the anatomy of a production harnes…
Ask HN: What's the simplest way for me to get my AMEX data agentically? (news.ycombinator.com) i found this on their site: https://developer.americanexpress.com/products/nextgen-agentic-payments/overview?intlink=us-ace-developerkit but I think this is for folks making b2b apps. I just need a way to get my statements/transactions rel…
Agentic Software Engineering (ASE): Agentic AI Coding Meets Software Engineering (ase.tools via hn) NAME ase-meta-persona - Persona Configuration SYNOPSIS ase-meta-persona [--help |-h ] [persona] DESCRIPTION The ase-meta-persona skill gets or sets the active communication style (persona) of the assistant. Five intensity levels of token u…
Tell HN: We need an accounting system for cognitive debt (news.ycombinator.com) The term “cognitive debt” is gaining ground [1]. We can now produce code faster than we can understand it.
Show HN: A free agentic AI security reference (CC BY-NC-ND 4.0) (www.nextkicklabs.com via hn) The Agentic AI Security Stack Deploy secure agentic AI systems. This free 200+ page reference provides a unified threat model, traces kill chains, and maps every control to OWASP, MITRE ATLAS, & CSA MAESTRO.
DocETL: Declarative and Agentic Map-Reduce (github.com via hn) DocETL: Declarative & Agentic Map-Reduce What is DocETL · Install · Python API · YAML · DocWrangler UI · Docs What is DocETL DocETL helps you process large collections of data (structured and unstructured) with LLMs. You write each operati…
CorvinOS – self-hosted agentic OS where EU AI Act and GDPR compliance by design (github.com via hn) Overview · Architecture · Audit & Compliance · A2A Network · Engine Layer · Security · EU AI Act · Learning Objectives One install. Seven bridges.
Show HN: Lymwave, Agentic Marketing Autopilot (lymwave.com via hn) Turn ideas into structured content workflows Generate briefs, outlines, and content plans based on real search and AI visibility opportunities, not guesswork. Lymwave uses audits, GSC data, and publishing integrations to plan, generate, pu…
Show HN: Is grep enough? A transparent benchmark for agentic code navigation (entelligentsia.github.io via hn) Felt LSP Servers were too complex. Bash tools alone too brutish.
Aera: Cross Platform Agentic Browser (getaera.app via hn) "Super intrigued by Aera's MCP integration." The browser that does the work. Describe your workflow.
The Usefulness of AI Agents (erikjohannes.no via hn) On the usefulness of AI agents Agentic AI is having its moment (or its decade, as some have put it). I have been researching LLM-powered agents for the last two years, but research (involving publicly funded projects and academic peer revi…
Show HN: Drift, write LLM agents in English and transpile to async Python (github.com via hn) Drift An intent-based language for agentic systems. Write your agent in English-shaped blocks, run it as async Python.
Show HN: A free ACP payments module that adds Stripe payments to MCP tools (www.afcommerce.com via hn) Free ACP Payments Module: Accept Payments in Your MCP Server On This Page - What the ACP Payments Module Is - Why Agentic Commerce Changes Checkout - Who This Module Is For - Three Ways to Run It - What Happens During a Sale - Two Payment…
Module decomposition cut agent token use 32% on follow-up feature additions (docs.krv.ai via hn) Agentic Cost Savings¶ This case study tests a narrow claim: when an agent starts from structurally cleaner code, does the next feature work take less time, fewer tokens, and less money? In one controlled experiment, the answer was yes.
From Isolated Agents to Agentic Mesh: Orchestrating SDLC with A2A and AP2 (blog.owulveryck.info via hn) From Isolated Agents to Agentic Mesh: Orchestrating SDLC with A2A and AP2 Exposing the problem Giving every developer a powerful, local AI agent feels like the ultimate productivity hack. But for organizations running at scale, it is a gov…
X401: HTTP-Native Identity Exchange for the Agentic Web (www.proof.com via hn) Introducing x401: Bringing Proof of Identity to the Web In 1997, the HTTP spec defined status code 402 Payment Required. It was reserved "for future use," a placeholder for a payment layer.
The state of agentic analytics, from 50 real data teams (blog.getcassis.com via hn) Field notes from 50+ conversations with data teams: the five stages of agentic analytics, what breaks at each, and what teams want next.
Technical Setup Guide for Shopify Agentic Storefronts (Geo) (stackarchitect.xyz via hn) Shopify Agentic Storefronts 2026 — Sell on ChatGPT & Google AI Mode (Setup Guide) Winter '26 Edition · Early Access · Complete Setup Guide Shopify Agentic Storefronts 2026 Sell on ChatGPT & Google AI Mode Complete Setup Guide AI-driven pur…
Show HN: Drip — pay-per-use finance newsletters for AI agents (dripstack.xyz via hn) Hi HN, I’m building Drip. It’s an API that lets AI agents access premium finance newsletters on a pay-per-use basis.
I built a flat-rate DeepSeek API for Claude Code (with vision support) (cloudcode.one via hn) CLOUDCODE.ONE Your Agentic Coding Partner. Sign inRegister Coding Plan $5/month Double of Claude Pro subscription usage, 1M context window, Code smarter.
Char: Agentic Notepad (char.com via hn) Start your day with clarity. Char turns overnight asks, rolled-over tasks, and today's meetings into a morning brief before you open your tabs.
Generate per-session LoRA adapters in <1s for agentic inference efficiency (github.com via hn) Tessera Hypernetwork Generate per-session LoRA adapters for inference tasks using hypernetwork synthesis. Version: 1.3.9 Features Metadata-to-LoRA: Generate adapters from structured user metadata (JSON) Text-to-LoRA: Generate adapters from…
AI Code Stitcher - Agentic AI Avoidance. (news.ycombinator.com) Hey guys, just letting you know that the latest version of the code stitcher is available, and has many new features including a major overhaul of the stitch viewer / file version history including it's own linter and editor facilities. ht…
Free Agentic AI Webinar: From Agent Design to Production (simplai.ai via hn) If you’ve been wondering what it actually looks like to build an AI agent and ship it to production — this is the session you don’t want to miss. Wednesday 24 June – 9:30 – 10:30 GMT+5:30 SimplAI is hosting a live Zoom webinar: “SimplAI Pl…
Show HN: Memory Magico – CLI based memory, wiki and deterministic sprints (github.com via hn) hey, I got tired of using GitHub issues and Mira to manage sprints and issue tracking so this is a tool I have been using for a while in a few of my projects to ensure that long sprints, and long bug finding sessions got triaged and proper…
A cheaper and safer agentic AI workflow (danuker.go.ro via hn) I recently tried agentic coding for real. It cost $0.034 and finished in 3 minutes.
Show HN: Atizar-AI agents where the server runs approved actions, not the model (atizar.io via hn) Open-source TypeScript framework for agentic automations: the agent drafts and proposes, a human approves, the server runs the approved action. Code for the engineer, a clean board for the client.
Hyperia 0.12.7 is released: an agentic terminal for agents and humans (github.com via hn) Hyperia™ A terminal emulator built for agents and humans. Hyperia is an agent-native terminal emulator.
Google Has Added Agentic Browsing to PageSpeed Insights (pagespeed.web.dev via hn) This site uses cookies from Google to deliver its services and to analyze traffic. Report from Jun 20, 2026, 6:37:14 AM Discover what your real users are experiencing Diagnose performance issues Discover what your real users are experienci…
Agentic Loops: Why the Best AI Coding Workflows Are Loops, Not Prompts (skilldb.dev via hn) Agentic Loops: Why the Best AI Coding Workflows Are Loops, Not Prompts #Agentic Loops: Why the Best AI Coding Workflows Are Loops, Not Prompts Most people still use AI to code the way they'd use a very fast intern with no memory: write a p…
My agentic engineering workflow as someone who doesn't write code (shreyasprakash.com via hn) My agentic engineering workflow has changed in the recent past, the models are better, there is much more freedom in choosing the harness, selecting the abilities and actions you could provide. - Pre-planning, pre-idea, pre-everything - Th…
Simplicity always wins:SOTA on swe-pro,tb2,-verif on 21 models with simple-agent (github.com via hn) Strands Benchmark Harnesses A repository for Strands-based agents and harnesses for agentic benchmarks. It is a uv workspace: the repository root coordinates one or more member packages.
Beast – governed output gateway for AI coding agents (github.com via hn) BEAST - Broker for Efficient Agentic Systems and Tooling Governed output gateway for agentic coding tools. BEAST sits between your AI coding agent (Cursor, Claude Code, VS Code Copilot) and any LLM provider.
OpenMontage the first open-source, agentic video production system (github.com via hn) OpenMontage The first open-source, agentic video production system. Paste A Video · Quick Start · Try These Prompts · Pipelines · How It Works · Providers · Agent Guide Follow The Build Turn your AI coding assistant into a full video produ…
Ask HN: How to deal with UI within the agentic loop (news.ycombinator.com) I am planning to the module in may SaaS to allow user to chat about its data, features, modules that he have in his workspace with LLM Agent. What approach should I use to render UI within the chat?
Cohere's open agentic North Mini Code – accelerated with NVFP4 on spark-arena (forums.developer.nvidia.com via hn) Hey all, I just put up two Spark Arena runs of North Mini Code 1.0 — an FP8 reference and an NVFP4 quant we made — to see what the GB10’s native FP4 support buys us. It’s Cohere’s first open agentic coding model: a 30B MoE (3B active), Ap…
Improving token efficiency in GitHub Copilot (code.visualstudio.com via hn) Improving token efficiency in GitHub Copilot June 17, 2026 by Ryan Caldwell and Bhavya U With the recent move to usage-based billing for GitHub Copilot, every token in an agentic session matters. They affect your credits, latency, and the…
Qwen and Fable: An open-weights agentic coding model. 35B Mixture-of-Experts (huggingface.co via hn) Qwable-v1 Qwen + Fable · An open-weights agentic coding model. 35B Mixture-of-Experts (3B active), built by layering Claude Fable-5 agentic tool-use behavior on top of a Claude Opus 4.7 reasoning distill of Qwen3.6-35B-A3B.
How agentic AI is rewiring Amazon's teams and upending its traditions (www.geekwire.com via hn) Swami Sivasubramanian, AWS VP of agentic AI, on stage at AWS re:Invent in December. (Amazon Photo / Noah Berger) Editor’s Note:[_Agents of Transformationis an independent GeekWire series, underwritten by Accenture, exploring the adoption a…
Show HN: Loomcycle – a sidecar runtime for AI agents (Go binary, Apache-2.0) (github.com via hn) The agentic runtime, in a sidecar. One Go binary alongside your application.
Agentic Grocery Shopping on Uber Eats (www.uber.com via hn) Introduction Grocery shopping often begins outside a commerce app: a handwritten list on the fridge, a screenshot of a recipe, or a vague plan like “healthy breakfasts for the week.” Translating that raw intent into a useful grocery cart i…
What does software development look like when agents write 100% of the code? (blog.bastion.computer via hn) 2026 has been an inflection point for agentic coding. In just two years the capabilities of models and harnesses went from a toy autocomplete to being powerf...
Agentic AI Foundation (aaif.io via hn) 6.12.26 - 🪙 Coinbase Makes Agentic Trading Real With MCP: Coinbase debuted an MCP-based tool that lets AI agents trade, pay for premium research, and buy on-demand compute using the x402 payment protocol. Initially for crypto, the firm pla…
Show HN: Phlox – Open-source self-hosted agentic web chat (github.com via hn) Phlox A feature-rich, ChatGPT-style, self-hostable AI assistant. Phlox is a self-hostable chat application with an agentic harness, document RAG, code execution, and MCP integration — running over any model provider: AWS Bedrock or any Ope…
Fugee, an agentic AI assistant for displaced people and asylum seekers [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Show HN: Wtdb – give every Git worktree its own database (github.com via hn) I run a lot of agentic coding sessions in parallel, each in its own git worktree. Every worktree points at the same local Postgres though, so the moment one branch runs a migration it changes the schema out from under the others.
Ask HN: Is anyone building real software with AI agents? (news.ycombinator.com) I've noticed a pattern where the people who are talking about how impressive their agentic workflows are, always seem to use these workflows to build more AI tooling. Has anyone seen a project built by an agentic workflow that could stand…
I accidentally hit SOTA on agentic memory by using AI companions (graph.coder.company via hn) graphCTX is a local-first context and memory layer that keeps AI coding agents grounded in repo knowledge, without accounts, API keys, or sending code to a hosted service.
SchemaFlow: Agentic Database Change Impact Analysis, SQL Gen and Eval Guardrails (developers.openai.com via hn) This cookbook walks through an end-to-end AI-assisted database change workflow using the OpenAI Agents SDK. It demonstrates how OpenAI’s tooling ecosystem can be applied to orchestrate complex, data-intensive workflows across modern enterp…
Cross-System Constraint Collisions: The Governance Gap in Enterprise Agentic AI [pdf] (himalaian.com via hn) CrAIg ™ Version 1.0 June 2026 C O N T I N U O U S R U N T I M E A I G O V E R N A N C E Cross -System Constraint Collisions: The Governance Gap in Enterprise Agentic AI A technical and operational framework for cross -system constraint gov…
Rethinking Monorepos in the Age of Agents (chamoda.com via hn) Rethinking monorepos in the age of agents Now that most of the software development industry is switching to agentic coding, it makes less and less sense to keep separate repositories when agents might benefit from having better context in…
A leader's guide to advanced team structures in an agentic world (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Run local agentic AI on the Mac using MLX (WWDC 2026) [video] (developer.apple.com via hn) - Run local agentic AI on the Mac using MLX Run AI agents locally with privacy, low latency, and offline access. Dive into how MLX advancements and Mac hardware make powerful agentic workflows possible entirely on-device.
Type Checking in Agentic Workflows – Conner Nilsen – PyCon US 2026 Typing Summit (pyrefly.org via hn) Talk: Type Checking in Agentic Workflows Does adding type checking to an agentic workflow really help agents? We ran an experiment recently to determine whether there are improvements in the success rate for completing different kinds of t…
AWS Tunes Up Graviton5 for Agentic AI, Boosts Bang for the Buck Bigtime (www.nextplatform.com via hn) AWS Tunes Up Graviton5 For Agentic AI, Boosts Bang For The Buck Bigtime Back in December, the Annapurna Labs chip division of Amazon Web Services showed off a preview of its Graviton5 Arm server CPU, and we got some hints about what this c…
Show HN: Mobile analytics made for agentic development (undercurrentanalytics.dev via hn) Mobile analytics made for agentic development Undercurrent Analytics is an all-in-one mobile analytics platform that uses open standards to give you the best possible understanding of how your mobile app is used by real people. Free until…
How AWS DevOps Agent evaluates telemetry tools for agentic readiness (bronto.io via hn) Consolidate telemetry, search terabytes in milliseconds, retain 12 months hot, and let AI troubleshoot. The intelligent observability data platform.
Apple's Passwords App Becomes Agentic (www.heise.de via hn) Apple's Passwords App Becomes Agentic Apple rarely uses the word "agentic." This is changing with the Passwords app, which will soon be able to navigate the web on its own. Whether at OpenAI, Anthropic, or Google, when it comes to artifici…
Global watchdog calls for tighter controls on agentic AI in finance (www.reuters.com via hn) paywalled
Ask HN: The next evolutionary step in LLM usage? (news.ycombinator.com) I'll keep this post short and sweet, we have seen several steps in the evolution of LLM (large language model) usage. 1.
Cursor users must consent to data collection in order to use Fable 5 (cursor.com via hn) Anthropic's Mythos-class model for autonomous, long-running agentic work. Privacy Mode customers need to approve Anthropic data retention before use.
How to Build an Agentic RAG with RubyLLM and Rails (www.panasiti.me via hn) How to Build an Agentic RAG with RubyLLM and Rails I run a RAG application for Italian pension and tax consultants. Users ask questions about INPS, professional pension funds, laws and regulations, and the app answers using a knowledge bas…
Macro Evals for Agentic Systems (developers.openai.com via hn) When an agentic system fails, the problem is often larger than a single bad response. A handoff may happen too late, a specialist agent may miss the same signal across many runs, or a review process may trigger for the wrong class of cases.
Agentic search – retrieval, harness, or model? (softwaredoug.com via hn) Agentic search gets interesting when agents do not know how to find the right answer. Oh, the agent might think it knows.
Agentic RL: Token-In, Token-Out Done Right (qgallouedec-tito.hf.space via hn) Is Grep All You Need? How Agent Harnesses Reshape Agentic Search (arxiv.org via hn) Apple Passwords Now Auto Fixes Weak and Compromised Passwords with Agentic AI (www.macrumors.com via hn) Apple Passwords Can Now Automatically Fix Weak and Compromised Passwords With Agentic AI Apple today announced that the Passwords app can now automatically update weak and compromised passwords using Apple Intelligence and Safari to take a…
Agentic AI solved coding and exposed every other problem in SE (venturebeat.com via hn) Agentic AI is now a core part of the engineering process, driving massive execution leverage and helping us generate more code than ever before. Yet, a difficult question I’ve increasingly heard from business leaders is: if we’re shipping…
Show HN: AgentCrew – a Markdown-first operating system for AI coding agents (github.com via hn) AgentCrew Turn your coding agent into a disciplined team. AgentCrew is a conversation-first, Markdown-first methodology for agentic coding.
Show HN: Version Control for AI Agents (cognatoai.com via hn) Git/GitHub did not evolve for agentic era, so we are building
Improving LM Studio's MLX Engine for Agentic Workflows (twitter.com via hn) We recently released mlx-engine v1.8.5 in LM Studio. This update dramatically improves performance for repeated, long-context agentic workflows by checkpointing your KV cache.
Agentic AI spurred a boom in mobile apps, but they aren't gaining traction (twitter.com via hn) Don’t miss what’s happening People on X are the first to know. Log in Sign up Post Conversation Jen Zhu @jenzhuscott Massive output uptick due to agentic AI.
Ask HN: Will your company be doing "LeetCode" interviews a year from now? (news.ycombinator.com) I work in big tech. I'm a SWE manager, but I have a half a mind to return to being an IC at some point.
Show HN: Omni – Local-first multimodal file search on macOS (hanxiao.io via hn) Finally made something I've always wanted, using the model we built. • SOTA omni embedding model, fully local, indexes text, PDF, image, audio, and video • Swift-native app UI + mlx-swift-transformer core.
Blumi CLI – A Private Agentic Runtime with Grid Dispatch (github.com via hn) blumi A local-first, provider-agnostic agentic coding companion — one Rust core, three faces: a terminal UI, a web UI, and a phone app. blumi is a single Rust binary whose UI-agnostic core emits one typed event stream, so every surface sho…
agentgateway Joins AAIF as an Open Gateway for Agentic AI Infrastructure (aaif.io via hn) The Agentic AI Foundation welcomes agentgateway — an open source gateway purpose-built for MCP, Agent-to-Agent, LLM, and API traffic — as its newest hosted project.
Training an Agentic Router for Optimal Cost-Performance on SWE Tasks (www.appliedcompute.com via hn) Training an Agentic Router for Optimal Cost-Performance on SWE Tasks On most enterprise tasks, model quality is not a scalar. One model is better at long-horizon repository exploration.
Trader – LLM agent for Robinhood with a Rust safety layer and paper trading (github.com via hn) Trader — LLM-Driven Robinhood Trading Agent A Rust agent that connects an LLM to Robinhood's official agentic trading API, enforces hard risk limits in a typed safety layer, and paper-trades against live market data before you risk a dolla…
Copilot SDK is now generally available (github.blog via hn) Copilot SDK is now generally available The GitHub Copilot SDK is now generally available. You can embed GitHub Copilot’s agentic engine into your own applications, services, and developer tools with a stable API and production-ready suppor…
Dumb core, smart edge for AI agents (arizenai.com via hn) Dumb Core, Smart Edge: Agentic Design Many agentic systems I've watched fail in production had the same shape: intelligence concentrated at the center, where it was hard to test, replace, or reason about. The orchestrator was doing too muc…
From Specialists to Builders: How AI Agentic Coding Is Reshaping Software Teams (aliparnan.com via hn) Specialization defined software teams for decades. AI agentic coding is creating a new Builder role—people who orchestrate agents across disciplines and own outcomes end to end.
Why Merge Conflicts Became the New Agentic Bottleneck (adamtornhill.substack.com via hn) Why Merge Conflicts became the new Agentic Bottleneck Revisiting some techniques from Your Code as a Crime Scene in the light of agentic coding. Specifically, how a socio-technical fit becomes even more important now that agents are our ac…
Vegvisir – Agentic Harness Built for Software Developers (github.com via hn) Vegvisir Agent Harness Vegvisir is a local-first agentic software development harness for people who want an AI engineering assistant that can actually work inside a repository without being handed every secret, every permission, and every…
Agent Code – open-source Mac app for managing AI coding agents (github.com via hn) A native macOS platform for agentic coding workflows, powered by Pi. Manage agents, skills, prompts, subagents, worktrees, and GitHub work in one signed Swift app that runs the installed pi CLI in the background.
AI in SRE: Where and how Google is deploying agentic AI to improve operations (cloud.google.com via hn) AI in SRE: Where and how Google is deploying agentic AI to improve operations Stevan Malesevic Distinguished Software Engineer Christopher Heiser Distinguished Site Reliability Engineer Since its inception over 20 years ago, Google has use…
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks (arxiv.org via hn) Long-context Large Language Models, despite their expanded capacity, require careful working memory management to mitigate attention dilution during long-horizon tasks. Yet existing approaches rely on external mechanisms that lack awarenes…
Ask HN: What are your worst war stories bringing agentic applications into prod (news.ycombinator.com) For a bit of context, I’m currently creating a team of AI agents at work to generate reports by fanning out into a large amount of subagents to process a large amount of transcript data. When the analysis fails mid-way because of some indi…
Show HN: Jynx, a matchmaking app to find gaming teammates (jynx.app via hn) TL;DR: Jynx is a gaming social platform that matches you with compatible teammates based on skill level, play style and schedule. Swipe to find players (Tinder-style), create or join game sessions (LFG), chat, and build your squad.
Spatial IDE's for agentic coding workflows (news.ycombinator.com) Seeing spatial IDE's (where terminals and files are displayed on a canvas instead of a regular dock like vscode) more often right now on HN and Reddit. This is a selection of the ones I've seen.
Three flavors of coding with AI agents (nocodefunctions.com via hn) A reasonable definition of an “AI agent”, at least in the context of agentic coding, could be: a software process endowed with the capabilities of an LLM launched with instructions given at the start to accomplish a task which runs…
Embodied Cognition and Agentic AI (lemire.me via hn) Where is your intelligence located? In your brain?
Arm Metis with GPT5.5 Cyber scores 98% on firmware vulnerability benchmark (newsroom.arm.com via hn) Agentic AI-powered Arm Metis advances security vulnerability discovery in software In the era of AI, modern software systems are built across increasingly complex codebases, frameworks, runtimes and libraries. As these systems scale, so do…
Show HN: TheFoundry – Easy bootstrapping framework for MultiAgent Systems (github.com via hn) For months, I struggled to build complex, long-running projects using AI agents and I kept failing... One shots, refactoring, high token consume...
Robinhood's bet on agentic trading and purchasing is 'wake-up call' for banks (www.americanbanker.com via hn) The brokerage fintech launched agentic trading and an agentic credit card today that will allow AI agents to trade equities and make credit card purchases on customers' behalf. It comes just weeks after OpenAI rolled out its own personal f…
Why $/token is the wrong metric for Enterprise AI (agentic) applications (canyoncode.ai via hn) Canyon Code gives enterprises the ability to observe, optimize, and governance their multi-agentic AI applications. The Workflow Intelligence Layer
The Self-Healing Vector Database (www.reddit.com) A pattern I keep seeing in agentic RAG systems: The agent is smarter than the retrieval layer. It can notice that context is stale.
Open-source playbook on agentic working — for the cross-audience, not just coders (28 chapters, MIT) (www.reddit.com) Author disclosure upfront: I wrote this. Free, MIT-licensed, no paid tier.
Best harness for agentic analytics? Codex? Claude? Custom? (www.reddit.com) I run a small seo marketing agency and we've built some dashboards on top of our data for reporting with nextjs + supabase. This is where reporting for our clients happen.
Ask HN: Do coding agents need cross-tool org knowledge? Or, just good to have? (news.ycombinator.com) I've been talking to engineers, mostly in large teams. While they love cross-surface search with Glean, they still assimilate and curate the context for agents manually.
I built an agentic coding harness across three CLI hosts (pub.towardsai.net via hn) 8 min read May 13, 2026 This article is a work in progress. I will keep updating it as the kit evolves.
Agentic AI Flywheels (www.newsletter.swirlai.com via hn) Agentic AI Flywheels The production loop after your agent ships, and the eval set that grows with it. 👋 I am Aurimas.
Tool-schema compression enables agentic RAG under constrained context budgets (arxiv.org via hn) Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the same context window needed for retrieval-augmented generation. We present the first syst…
the agentic depth gap between open source AI assistants ranked (www.reddit.com) Agentic depth measures how far an autonomous agent can take a task before human intervention. The gap between open source options on this dimension is wider than feature comparisons suggest.
Looking for Suggestions — Single 5090 & 64gb DDR5 (www.reddit.com) Hi Reddit, I am planning on running Qwen 3.6 27b NVFP4 via vLLM on my 5090 but was wondering if something like 35b a3b at Q8 on Llama would produce better results for agentic coding and utilize the system memory. My research says no but if…
Breaking Bot: Hacking and Defending LLM-Based Applications (www.szia.ai via hn) Breaking Bot: Hacking & Defending LLM-based Applications - Marton Antal Szel - Dec 24, 2025 - 12 min read Updated: 4 days ago Let's say your "super-intelligent" agentic chatbot - the one with access to sensitive customer data - is hijacked…
Harbor v0.4.19 - vllm/sglang/llama.cpp launch codex/claude/pi/opencode (www.reddit.com) I'm usually not posting about Harbor releases out of the respect for the community here, but I think v0.4.19 might save a lot of people some time. Harbor can now launch your local agentic coding tools with local inference backends.
I replaced my old job with an AI agent (www.reddit.com) Hello friends. Today I want to talk about agentic media buying.
I Built MagesticAI. A Cloud Web-Based Agentic DevOps Orchestrator that actually helped me develop Itself. (www.reddit.com) Posted on other feeds last week and figured some of you out here might be interested as well; Someone commented asking if it supported OpenAI-compatible endpoints (LM Studio, vLLM, OpenRouter, Together, Groq, LocalAI…), so i have spent few…
Testers and collaborators wanted (www.reddit.com) Hello, I'm working on an Agentic wrapper system, Helix-agi, and I am trying to get some additional testers and collaborators involved in the project. Helix relies on a unique Agentic workflow that routes all incoming data, including tool u…
Inside Google’s Agentic Search Revolution (puck.news via hn) puck.news Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
Agentic AI Changes the CPU/GPU Equation (www.amd.com via hn) Agentic AI Changes the CPU/GPU Equation Skip to main content Enable accessibility for low vision Open the accessibility menu Skip to main content AMD Website Accessibility Statement Products Processors Accelerators Graphics Adaptive SoCs,…
Validating an idea, would anyone be interested in e-commerce designed for agents? (www.reddit.com) Me and 2 other friends are trying to solve payments through agents. One of the ideas we're looking into is merchant integration to allow agentic payments using any of the plethora protocols that exist (MPP/UCP/x402/AP2/Google's Universal C…
Rust is a great fit for the agentic era (kerkour.com via hn) We're sorry but this website doesn't work properly without JavaScript enabled. Please enable it to continue.
I’m a solo dev building TigrimOSR, a Rust-native AI agent workspace for engineering and developer workflows. (www.reddit.com) The main problem I’m trying to solve is that agentic AI is still too random for serious engineering decisions. For design work, calculations, reports, code changes, or technical review, I don’t want agents just “vibing” through tasks.
Stripe's John Collison on How Agentic Commerce Will Reshape the Internet [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
A Marketplace of Fine Tuned SLMs for Agentic Tasks (marketplace.neurometric.ai via hn) 130 models available Small Models, Big Impact Discover task-specific SLMs ready for your business, or browse general models under 20B parameters. Need something custom?
Anybody knows why cursor trying to move into "claude desktop" style app? (www.reddit.com) It makes absolutely no sense for cursor trying to switch over to Claude or Codex desktop style app. I am a Neovim/VSCode user and I only recently started using cursor, and found out that the UI/UX for agentic coding is phenomenal.
Moss: Self-Evolution Through Source-Level Rewriting in Autonomous Agent Systems (arxiv.org via hn) Autonomous agentic systems are largely static after deployment: they do not learn from user interactions, and recurring failures persist until the next human-driven update ships a fix. Self-evolving agents have emerged in response, but all…
CodeAlta an efficient agentic AI coding CLI assistant coded in C#/.NET (codealta.github.io via hn) ██████ ██ ██ ██ ██ ██░░░░██ ░██ ████ ░██ ░██ ██ ░░ ██████ ░██ █████ ██░░██ ░██ ██████ ██████ ░██ ██░░░░██ ██████ ██░░░██ ██ ░░██ ░██ ░░░██░ ░░░░░░██ ░██ ░██ ░██ ██░░░██ ░███████ ██████████ ░██ ░██ ███████ ░░██ ██ ░██ ░██ ░██ ░██ ░██░░░░ ░█…
Anyone evaluated the difference between Qwen Code for the local qwen models vs another harness? CC, OC, LC, Aider etc.. (www.reddit.com) For me, opencode doing fantastically but was wondering if qwen code would be more native and have better functionality, since idk which agentic harness they used to get their benchmark results
What nobody's measuring about dense MoE in production tool calling agents (www.reddit.com) Most of the model selection conversation I've seen focus on benchmark scores and cost (no surprise there). The question I can't find good production data on is whether dense vs MoE actually affects reliability for tool heavy agentic flows,…
[Blogpost] Files Are All You Need: Towards Self-Improvement in ChatGPT (www.reddit.com) Subreddit rule statement: link to blog post in the comment. Not self-promotion.
Karpathy's LLM-Wiki for agentic software development? (www.reddit.com) I’ve been away from coding/software development for about a year. When I stepped away last summer, agentic software development wasn’t nearly as capable or accessible as it seems today.
A solution for schlep blindness in agentic development for Kubernetes envs (metalbear.com via hn) Turn your AI agents into autonomous developers Use mirrord to instantly validate every change against your live staging environment — multiple agents, same cluster, no conflicts. Windsurf & others or CLI No credit card needed Fast setup, n…
Show HN: P-Hacker – group/analyze HN trends by topic (not just keywords) (p-hacker.com via hn) I'd seen various HN trends tools over the years ([1] [2]), but they all used strict keyword (n-gram) matching. That limited a) how sophisticated any trend-surfacing could be and b) the depth with which you could explore the full discussion…
RTMX: Intent Layer for Agentic Engineering (github.com via hn) RTMX Track what you built, what's tested, and what's next -- from the terminal. RTMX is a CLI that manages requirements traceability as a CSV file in git.
The Claude Code Production Playbook: Sub-Agents, Hooks, and MCP Integration (ddsboston.com via hn) Claude Code Masterclass 2026 The definitive end-to-end guide to Anthropic’s agentic coding tool — installation, Ollama local fallback, CLAUDE.md, Skills, Subagents, Agent Teams, Hooks, and MCP. Everything you need before building productio…
Not All Software Systems Are Agent Friendly (yassi.dev via hn) Discourse around AI tends to collapse into two camps: true believers and luddites. A recent piece, Agentic Coding is a Trap, highlights what the author calls the “paradox of supervision” - where the very judgment needed to oversee AI deleg…
What's the best qwen3.5 or 3.6 reap model? (www.reddit.com) What's the best reap (pruned) model you know of? This one runs twice as fast on my low vram setup, but I'm unsure if it will miss out on a lot of things agentic coding related.
Ask HN: How are agentic workflows meant to offset AI debt? (news.ycombinator.com) I don't know quite how to put it. But projects I inherit and am supposed to get over the line have this same strange quality: they are 'undesigned'.
Show HN: AgentShield – Stop AI agents from spending money unsupervised (agentshieldv2-dashboard-production.up.railway.app via hn) I'm a recent grad from UMich and built AgentShield because agentic AI is moving fast but payment safety hasn't caught up. Agents are already being handed API keys, stablecoin wallets, and payment credentials - if one misbehaves, gets promp…
The Agentic Loop (hypnodrones.com via hn) The agentic loop First, we offloaded knowledge to writing. Then, we came for means of production.
Building an AI agent with OpenAI tool use — struggling with consistency. How do you enforce tool call order reliably? (www.reddit.com) Hey, Software engineer here, relatively new to agentic workflows. Building a production AI concierge — user says "I'm going to Budapest tomorrow, plan my day" → agent searches our offer database, builds a plan, user books everything in one…
Understanding, Analyzing, and Optimizing Agentic AI: A CPU-Centric Perspective (arxiv.org via hn) Agentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on the fly. Due to diverse task execution need, such serving heavily rely on heterogeneous CPU…
Qwen3-Coder-Next-UD-Q4_K_XL vs. Qwen3.6-27B-MTP-UD-Q4_K_XL on Strix Halo (www.reddit.com) I wanted to switch from Qwen3-Coder-Next-UD-Q4_K_XL to Qwen3.6-27B-MTP-UD-Q4_K_XL for local agentic coding. The Qwen3.6-27B is perceived to be "smarter" than Qwen3-Coder-Next, and I wanted to "upgrade" my local AI coders.
Show HN: Nano-RAG – Agentic multi-hog retrieval without graph database (news.ycombinator.com) https://nanorag.nb1t.sh/ Important: Please choose correct namespace from top-right dropdown. Available docs/namespaces: Cloudflare, Nextjs, and Dodo-payments (default).
Show HN: Agentic simulator for marketing email A/B testing (inbox-wars.com via hn) I built an agentic simulator for marketing email a/b testing using a fleet of "digital twin" customers. why build this?
Hershey Bets on Agentic AI to Rethink $2B in Marketing Spend (www.adweek.com via hn) Hershey is revamping one of marketing’s oldest measurement tools—marketing mix modeling—by enlisting agentic AI in a bid to turn what has historically been a slow, backward-looking process into something closer to real-time. The confection…
UGen: An Agentic Framework for Generating Microarchitectural Attack PoCs (arxiv.org via hn) Microarchitectural attacks continue to evolve, uncovering new exploitation vectors in modern processors. From a defensive perspective, assessing a system's susceptibility to such attacks remains challenging.
Have you tried Agentic analytics tools? (mitzu.io via hn) TL;DR Compare the best AI analytics tools in 2026 across semantic-layer trust, no-hallucination reliability, SQL transparency, and team fit. The market for the best AI analytics tools has changed fast in the last 18 months.
How to Learn Agentic AI in 2026 – Without Getting Lost in Hype (simplai.ai via hn) How to Actually Learn Agentic AI in 2026 — Without Getting Lost in Hype Most AI courses teach you theory and leave you stranded before deployment. SimplAI University is built differently — 11 structured chapters, real tools, and a communit…
Benchmarking the new b9200 update: Optimizing Qwen 3.6 27B mtp for Hermes Agent on a single RTX 3090 (www.reddit.com) I'll be UPDATING this as it seems I was benchmarking and testing Just before the UPDATE LOL TL;DR If you're running rigid agent frameworks locally with mtp on consumer hardware: drop your draft window to 3, lock parallel slots to 1, and co…
Did anyone here did the certification: GitHub Certified: Agentic AI Developer (beta) (www.reddit.com) Hello everyone, I wanted to ask if anyone here got the certifcation GitHub Certified: Agentic AI Developer (beta) or was thinking of getting it? What do you think about it?
ik_llama: Qwen3.6 27B and 35B on very low VRAM (www.reddit.com) Thank you to the people at ik_llama and llama.cpp. It's amazing how far you've all pushed mtp and other tech so that I can run 27B and 35B Qwen3.6 models on an old gaming laptop with a RTX2060 mobile at 6GB VRAM and 32GB RAM.
Show HN: Vyvoice: Privacy-first, cross-platform, offline voice transcription app (vyvoice.com via hn) Hey Hacker News, vyvoice is a cross-platform, offline voice transcription app I started working on in December as a Windows user tired of every good dictation app being Mac-only. Beyond transcription, it has built in support for voice comm…
Show HN: Agentic product discovery for AI apps and shopping agents (www.seekon.me via hn) Agentic Catalog Intelligence. Empower your AI models with precise, real-time product discovery.
Show HN: Markanywhere – A Streaming Processor of Meanings (github.com via hn) Markanywhere can parse any input, like Markdown, HTML, XML, as a stream of semantic events which can be rendered, transformed, evaluated. Works great as an interactive transport layer for the LLM inference output and agentic feedback loops.
Show HN: Claurst – Rust-Based OSS Terminal Coding Agent Now in Beta (github.com via hn) CLAURST Agentic Coding for Builders who Ship Claurst is an open-source, multi-provider terminal coding agent built from the ground up in Rust. It started as a clean-room reimplementation of Claude Code's behavior (from spec) and has since…
I almost broke the one rule that separates agentic coding from vibe coding (www.reddit.com) I built an opinionated multi-agent setup on top of Claude Code. I was proud of two agents in particular: a software engineer doing red-green TDD, and a separate tester running the adversarial edge-case pass.
Setting the Standard for Agentic Development (lovable.dev via hn) Platforms like Lovable enable non-development teams to build, deploy, and iterate on production applications through natural language. Enterprise adoption is accelerating, and teams are already integrating coding agents into core workflows…
Show HN: Building a universal device experience [video] (www.youtube.com via hn) I have been working on this project for a few years now and the end goal is to make a ubiquitous and natural user experience to interact with machines. The long term goal is to build a fully agentic experience that drives the UI for you (g…
Show r/AI_Agents: Stop your agents from breaking tool calls in production — we built a reliability layer for 2,000+ APIs (www.reddit.com) We built a CLI that sits between AI agents and production APIs — handles auth, retries, compliance, and idempotency automatically across 2,000+ APIs. Give your agents capability of multi-tool calls with 100% accuracy.
Looking for your experiences in agentic scraping social profiles (www.reddit.com) Based on your experience, which agentic workflows has everyone had the most success using to extract public profile data from Instagram and Facebook? I've seen previous discussion here about n8n and OpenClaw, and I'm looking for the latest…
Looking for affordable alternatives to Claude Team / Claude Code for a small dev team (heavy agentic usage) (www.reddit.com) We run a small software services company and we’ve been heavily using Claude (especially opus + Code features) for the last few months. The problem is: We need to share the account between 6-8 developers Anthropic keeps suspending our Max/…
What is the best ai engineering course right now for agentic ai (www.reddit.com) Everywhere i look ppl are talking about agentic ai now… feels like basic gen ai stuff is already saturated. but trying to figure out how ppl are actually learning this beyond surface level… youtube kinda stops at demos.
Most teams optimize the prompt. Agentic systems have more moving parts (www.aevyra.ai via hn) On LinkedIn last week, an AI practitioner I know made an observation I keep thinking about: hill-climbing on evals tends to leak information specific to those evals rather than improve the system. Their follow-up question: "What if you hil…
Authorization Bypass in AWS's Agentic AI for Enterprise: Amazon Quick (www.fogsecurity.io via hn) We discovered an authorization bypass in Amazon Quick’s AI Chat Agents that allows users to access and interact with AI agents despite explicit administrative restrictions. AWS responded by deploying a fix without notifying customers, clas…
Choosing the Right Agentic Design Pattern: A Decision-Tree Approach (machinelearningmastery.com via hn) In this article, you will learn how to apply a structured decision tree to choose the right agentic design pattern for any AI system you are building. Topics we will cover include: Why pattern selection is a critical design decision, and w…
Zig vs. Rust, agentic coding, and intellectual control [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Useful AI agents / tools for client meeting management? (www.reddit.com) Hey y'all, I've been working towards automating different sectors of my agency each week, and this week it’s meeting workflows. I know about AI note-takers but it seems like most of them are just passive recorders that leave me with a long…
Mergecrew: Open-source agentic SDLC with human-gated prod deploys (github.com via hn) Mergecrew Autonomous product team in a box: every day, mergecrew specifies, designs, builds, deploys to dev, scans for bugs, and hands you a digest to approve before anything reaches production. Mergecrew is the open-source platform for ru…
Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model (github.com via hn) Hey HN, Henry here from Cactus. We open-sourced Needle, a 26M parameter function-calling (tool use) model.
Industry academia disconnect (www.reddit.com) Hi all, I do a lot of work with academic and industry partners in engineering applications. Therefore I end up having a lot of conversations with people around agentic AI for engineering.
Physics-intern: an autonomous agentic framework for physics research (huggingface.co via hn) Your Article Title Built with the Research Article Template. Quick Start cd app npm install npm run dev Visit http://localhost:4321 to see your article.
TigrimOSR v0.4.1: Running AI agents headless on a remote server, controlled by a fast local Rust UI (www.reddit.com) Hi everyone, I’ve been working on TigrimOSR v0.4.1, a Rust-native version of TigrimOS, and I’d like to invite people to try it and give feedback. The main idea is: Run the agent system headless on a remote machine, then connect to it from…
As agentic dev tools boom, workflow auditability becomes the constraint (thenewstack.io via hn) As agentic dev tools boom, workflow auditability becomes the constraint Recently, I was working with a senior engineering leader at a large financial institution to review their DevSecOps platform engineering roadmap. Their team had deploy…
Show HN: RipStop – Git guardrails to reduce impact if your code agent goes wild (github.com via hn) Hi all, RipStop is a node package implementing a set of rules that consumers can use to protect their repos from wilder actions by LLM agents. A consumer needs only a few lines of code to configure the rules they wish to apply.
The AI market moves so fast that your business idea can expire before launch (www.reddit.com) 1.5 years ago, n8n was everywhere. People were building workflows for everything.
Do you think foundational model companies will take over all agent businesses? (www.reddit.com) Do you think they will end up learning the most painful workflows from enterprise customers and built all the most necesary agents for the smaller guys themselves? In other words, squeezing out all the agentic companies out there?
What do you NOT like about Cursor / VSCode / Claude Code desktop / Codex / etc.? (news.ycombinator.com) I am building a highly integrated, cross-provider agentic workstation (its neither an IDE nor an ADE - does a bit of both, with additional unique features on top), and I would love for you guys to rant about what you hate about the tools y…
Thousands of apps built with Agentic AI platforms like Lovable, Replit, Netlify, and Base44 are exposing private data (www.reddit.com) A new investigation by Israeli cybersecurity firm Red Access found thousands of AI-generated web apps leaking data ranging from medical records to internal business documents. The findings add to mounting concerns about vibe coding, a fast…
Automata and AI (www.reddit.com) Hello, I have been working on a new programming language for creating state machines. I’m curious how the structure automata provide might be useful with MCP and agentic workflows.
Show HN: I've implemented multi-repo workspace support in Agent of Empires (github.com via hn) Coding agent management is all the rage right now, and many tools are being created to fill the gap. As a power user for all tools I've used since I've started my software engineering career, I've always taken the time to test multiple too…
Agentic AI is giving cyber criminals nation-state-like powers (www.defenseone.com via hn) Pentagon leaders love agentic AI. But it’s giving cyber criminals nation-state-like powers As new tools change cybersecurity, just moving faster won’t be enough.
Agentic AI vs. AI Agents: The Governance Shift (rootcx.com via hn) Open any vendor pitch from the last 6 months and somewhere in the deck, you'll see the word agentic. It's been a marketing term for so long that most engineering leaders have started treating it as noise.
Show HN: Agentic productivity platform for high perfomers (www.mainthread.app via hn) Finally on top of things. Mainthread unifies every commitment across work, family, and household into one intelligently prioritized system — then deploys AI agents to handle what doesn't need you.
LLM as logic processor, filesystem as memory — Q2 quant doing real agentic coding 50k context (www.reddit.com) Hello LocalLLaMA subreddit, i have been running local models for coding tasks and kept hitting the same problems everyone does — the model writes an 800-line file in one shot and half of it is garbage, it spirals in its own reasoning for 4…
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? (github.com via hn) VibeServe: Can AI Agents Build Bespoke LLM Serving Systems? An agentic loop that synthesizes bespoke LLM serving systems — one per (model, hardware, workload) target — instead of forcing every deployment through a single general-purpose ru…
The missing primitive in every agent harness is a protected region (www.reddit.com) I wrote a post about why agentic coding falls off a cliff after a few weeks. Coding agents have no equivalent of the source/assembly boundary a compiler gives us.
I built agentwerk, a tiny Rust crate for scaling agent collaboration focusing on getting work done (www.reddit.com) For a new Rust project, I was searching for a simple agentic loop implementation. My goal was to analyze thousands of software artifacts at scale.
I built a context window optimization framework for coding agents — open source + paper (www.reddit.com) Been working on a problem that I think a lot of people here face: agentic coding pipelines blowing through their context window way too fast, losing important information, and degrading task quality mid-session. Apohara Context Forge is my…
I put Claude Code inside Obsidian as a plugin — full agentic vault access with a native UI bridge (www.reddit.com) could not extract summary
I asked 20 Agentic Aai founders how they handle agent access. 17 said temporary workarounds. (www.reddit.com) Over the last few weeks I’ve been doing something that probably sounds a bit obsessive. I reached out to founders and engineers who are shipping AI agents into production agents that touch CRMs, sales automation, ai chatbots, payment APIs,…
Show HN: Make your codebase agent ready (github.com via hn) A set of Claude Code skills to assess and improve the agentic readiness.
Powering the Inference Era: Inside the DigitalOcean AI-Native Cloud (www.digitalocean.com via hn) By Vinay Kumar, Chief Product & Technology Officer I’ve spent the last fifteen years building cloud services: early days of AWS building S3 and EBS, helping launch Oracle Cloud Infrastructure from inception, and now building the agentic cl…
Ask HN: How do you give estimates in the age of Agentic coding (news.ycombinator.com) Back in the day you would get a rough estimate of how long a new feature might take once you had worked on a codebase for long enough. You knew how the internals worked, how much time it would take to design the solution, how fast you coul…
Should we use a non-thinking model for code after using a thinking one for plan? (Agentic coding) (www.reddit.com) I usually use Qwen3.6 27B (slow as heck on my RX 6800 but it works) for plan and Qwen3.6 35B A3B for the coding. But I was thinking the other day if I should remove the thinking from the code model.
Ask HN: What is the underlying stack behind multi-agent platforms? (news.ycombinator.com) Recently, I am seeing lots of startups with multi-agent platform, where you can create your own agent template, attach tools and run it reliably. Which frameworks, platforms are you using for these kind of multi-agentic platforms?
ABA Games (1D Pac-Man, etc) Agentic Gamedev Skills (github.com via hn) Agentic Gamedev Skills English | 日本語 This repository collects agent skills extracted from game-development work and related agentic-workflow research. Each skill lives under .agents/skills/, uses SKILL.md as its entry point, and may includ…
Meta plans advanced 'agentic' AI assistant for users (www.reuters.com via hn) paywalled
Show HN: Stagewise – Agentic IDE for Your Z.ai/DeepSeek/Moonshot Subscription (github.com via hn) The Open Source Agentic IDE for Developers English | 简体中文 | Deutsch | 日本語 | Español | 한국어 /_components/feature-images/full-demo-dark.png) About the project stagewise is an open source agentic IDE for developers with a coding agent built ri…
Show HN: Slate – agentic pre-production studio for solo Youtubers (useslate.app via hn) I built slate as a personal tool to centralize my strategy, research, scripting, thumbnails and shots in one place. Started showing it to other youtubers and that made me wonder if more people could have the same problem as me.
Is GraphQL the Panacea for Agentic AI? (magiroux.com via hn) It was evident that GraphQL would be touted as the ultimate API style for agents. After all, it is one of the only ways we expect an API style to stay relevant these days.
Open Sourcing Our Platform - GuideAnts Notebooks (www.reddit.com) This is yet another agent harness and UI and I hope you will have a look and consider contributing. Elumenotion/GuideAnts: GuideAnts Notebooks.
Agentic RAG Frameworks (www.reddit.com) I am trying to understand how the market around RAG is currently, what are it's usecases, how do enterprise companies approach this. Do they just have company related documents which is uploaded to these RAG systems and use it to query the…
Anthropic response to 1-click pwn: Shouldn't have clicked 'ok' (www.theregister.com via hn) MOST POPULAR EVENTS - Securing the Untrusted Agentic Development Layer Join us to learn how to architect a development environment where your builders and their agents can move fast and securely. - Toxic Flows: When Your AI Agent Skill Bec…
"Surface" a Governed AI-Agentic Surface (news.ycombinator.com) A continued work in progress https://github.com/pauljbernard/sbcl-agent-desktop and https://github.com/pauljbernard/sbcl-agent an implementation of the ideas discussed in: The Evolution of Software Scale https://www.amazon.com/Evolution-So…
Subjective: Building a Native VFX Editor with Agentic Coding (sxp.studio via hn) This blog post is about my process and learnings in using agentic coding to ship a project with higher complexity than your usual vibe-coded todo app. You can download the app on iOS/iPad/macOS here: Subjective.
Mistral Medium 3.5 Is Now Available in Puter.js (developer.puter.com via hn) Mistral Medium 3.5 Is Now Available in Puter.js On this page Puter.js now supports Mistral Medium 3.5, the new flagship merged model from Mistral AI that unifies instruction-following, reasoning, and agentic coding into a single set of wei…
Starting with Agentic AI (iscinumpy.dev via hn) AI suddenly passed the “more time saved than spent” point around December 2025. A little late, I’ve finally started using agentic AI in various places over the last 2-3 months, and wanted to jot down my thoughts on what works, what doesn’t…
Understanding agentic workflows (www.reddit.com) I tried developing workflows using github copilot in order to create an multi-agent orchestration for a use case about creating research paper based on user’s need. However, there is no supported mechanism for subagents to spawn custom sub…
Two OpenClaw Agents Negotiate a YC SAFE with Agentic Power of Attorney (www.juanfiguera.com via hn) Two OpenClaw agents negotiate a YC SAFE with Agentic Power of Attorney I gave an AI agent access to act on my behalf on a third-party platform a few months ago. Within about ten minutes I realized I was scared of it.
Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards (arxiv.org via hn) Large language models (LLMs) have demonstrated strong potential in agentic tasks, particularly in slide generation. However, slide generation poses a fundamental challenge: the generation process is text-centric, whereas its quality is gov…
Chasing AI Memory SOTA: Beating the Benchmark, Missing the Point (xmemory.ai via hn) Chasing AI memory SOTA: Beating the Benchmark, Missing the Point 66.88%, 80.1%, 85%, 90.79%, 93%, 91.69% and even 100% — what do all these numbers have in common? They’re all state-of-the-art (SOTA) scores on various agentic memory benchma…
Global online hackathon for building AI agents with perception + memory (May 16–18) (www.reddit.com) Agents are moving into browsers, apps, meetings, dashboards, and code editors. The next generation of agents will need more than text context — they need to see what is happening, hear what is being said, remember important moments, and ac…
Is there tool that helps me validate my AI business idea? (www.reddit.com) I'm a product manager for a small business and I'm working on a product idea in the field of agentic AI. I have been chatting a lot with Gemini and ChatGPT but at some point they just keep telling me how great my idea is.
Architectural Framework for Agentic AI in Identity and Eligibility (wwps.microsoft.com via hn) Architectural Framework for Agentic AI in Identity & Eligibility By Prabhaker Cirium, Prin Consultant at Microsoft and Sajal Mukherjee, Senior Consultant at Microsoft Leveraging Azure AI to Revolutionize Citizen Onboarding and Benefits Eli…
We built an agentic AI for support triage. 47% deflection in 90 days. Full retro. (www.reddit.com) Setup: mid-size SaaS, ~3,000 tickets/month, 6 agents drowning. 70% of volume was tier-1 (passwords, billing, where's-my-feature).
Need advice on hardware purchasing decision: RTX 5090 vs. M5 Max 128GB for agentic software development (www.reddit.com) tl;dr - For software development, Qwen3.6 27B, 5090 gives you ~3x speed over M5 Max, letting you plow through code, while M5 Max gives you ~4x memory, letting you use higher quantization and bigger context. Which would you choose and why?
A Grand Challenge for Reliable Coding in the Age of AI Agents (arxiv.org via hn) Agentic AI systems can now generate code with remarkable fluency, but a fundamental question remains: \emph{does the generated code actually do what the user intended?} The gap between informal natural language requirements and precise pro…
AI subscriptions need a reliable meter (www.reddit.com) TLDR; “A gallon should be a gallon. A mile should be a mile.
Ling 2.6 (Flash and 1T): Efficient Open Models Competing on Agentic Benchmarks (firethering.com via hn) Ant Group doesn't get the coverage it deserves. While the open source AI conversation in the West circles around DeepSeek and Qwen, Ant Group has been quietly building a model family that competes directly with the models everyone is talki…
Agentic AI Community 2026 (simplai.ai via hn) Free, self-paced courses covering everything from agent fundamentals to real-world deployment. 50+ hands-on lessons designed for both technical and non-technical learners.
Skelm – Build AI agents in TypeScript without losing your mind (github.com via hn) skelm Build secure, agentic, long-running workflows in TypeScript. Run them anywhere Node runs.
The Figure-Eight Model for Agentic DevEx (medium.com via hn) The Figure-Eight Model for Agentic DevEx | by Joe Kutner | May, 2026 | Medium Sitemap Open in app Sign up Sign in Get app Write Search Sign up Sign in The Figure-Eight Model for Agentic DevEx Joe Kutner Follow 5 min read · 1 day ago 2 List…
tested four newest open source Kimi K2.6 is the fastest, GLM 5.1 the fanciest, DeepSeek V4 is the most comprehensive, and Xiaomi MiMo is the slowest (www.reddit.com) Architecture explains the gap: MiMo's MoE runs more active params per token than Kimi K2.6's optimized routing hence slowest. DeepSeek V4's 'comprehensive' edge is partly MLA: ~75% KV-cache compression makes it far better for long agentic…
My list for Top Agentic Frameworks - Looking for feedback on any that are missed, or theme to be addressed more fully (www.reddit.com) In 2026, AI agents have moved from hype to production reality. Teams are no longer asking if they should deploy agents.
Agent Orchestration Models (news.ycombinator.com) We are using Symphonic Orchestration (models) for our agentic commerce platform (hive of clawdbots building databases) and wanted to know what folks thought of our approach and also to learn about alternatives.
The future of company architecture (www.reddit.com) I've been in AI for over 10 years now and toyed with GPT2 when I was doing NLP work and really recognized the power of LLMs as a way to drive automation after spending time trying to build agents with GPT3.5. As time as gone on I've become…
A Mental Model for Agentic Work (basti.io via hn) Blog A Mental Model for Agentic Work May 5, 2026 - AI Agents - Company Operations - Software Engineering Something shifted in the first quarter of 2026. Not a feature launch, not a new product - a structural change in how work happens.
Show HN: Kanban-CLI – a web UI for local Markdown todo lists (github.com via hn) As we all are, I've been experimenting with ways to reduce external saas spend, and continually bring traditionally external pieces of context (prs, docs, trello boards) into the one mono repo. I have toyed with a markdown todo list and se…
Five Eyes spook shops warn rapid rollouts of agentic AI are too risky (www.theregister.com via hn) Five Eyes spook shops warn rapid rollouts of agentic AI are too risky Prioritize resilience over productivity, say CISA, NCSC and their friends from Oz, NZ, Canada Information security agencies from the nations of the Five Eyes security al…
PyFlue – Python-Native Agent Harness Framework (Python Clone of Flue) (super-agentic.ai via hn) Full-Stack Agentic AI Company We build deeply technical agent developer tools, purpose-built for agent experience and agent engineering at scale. Our research lab explores the frontier where Agentic AI meets Quantum AI.
UAE Plans to Run 50% of Government on Agentic AI Within Two Years (www.mitsloanme.com via hn) UAE Plans to Run 50% of Government on Agentic AI Within Two Years Agentic systems will analyze, decide, and execute across ministries under centralized oversight. News - Oman to Scale AI Ecosystem With New Special Economic Zone - UAE Bets…
Agent Evals is an absolute nightmare, so I built Signals to reduce the noise and cost (www.reddit.com) Hey peeps - I think the hardest thing about building agents is their evaluations. especially for scenarios that require multiple tool calls and the agent itself can go down a trajectory that you haven't manually tested before.
Show HN: Enoch – Control Plane for Autonomous AI Research (github.com via hn) I built Enoch after working with OpenClaw and trying to get an agentic coding system setup with Codex. In the past, I was trying to manually generate, code, and test this all manually.
CISA, NSA & Five Eyes publishes guide on how to safely deploy AI agents (cyberscoop.com via hn) Cybersecurity agencies from the U.S. and allies issued a joint warning Friday on the risks of "agentic AI." The new guidance urges critical infrastructure leaders to implement zero-trust protocols as autonomous systems gain unmonitored acc…
Tried running Claude Code with local LLMs via Ollama — ended up subscribing to Pro anyway. But now I can't disconnect from the local server. (www.reddit.com) I've been experimenting with using Ollama to run Claude Code locally with models like Gemma 4, thinking I could avoid API costs. However, I quickly realised these models aren't really optimised for Claude Code's agentic workflows — they te…
Which Agentic Coder is the most with it now? (www.reddit.com) Considering the price to performance which is the best deal or setup right now? Similar to codex where it can edit project files inside a folder etc.
Show HN: Large Scale Article Extract of Newspapers 1730s-1960s (snewpapers.com via hn) Hello HN, over the past 7 months I've spent nearly 3,000 hours on building SNEWPAPERS, the first historical newpaper archive with full-text extractions, nearly perfect OCR, a vast categorization taxonomy and of course with semantic and age…
I used Claude to build "pin-llm-wiki" — A skill that turns any URL into a clean, citable Karpathy-style LLM Wiki (github.com via reddit) Hey 👋 I’ve been using Claude Code a lot for personal research and knowledge management, and one thing kept bothering me: Turning articles, YouTube videos, and GitHub repos into clean, structured, citable notes is tedious. So I built pin-ll…
Is agentic commerce really APIs… or dynamic UIs like this? (www.reddit.com) https://preview.redd.it/2abn96dwudyg1.png?width=1642&format=png&auto=webp&s=ab5facbd9f4223184834711346dca2bc64db20d3
Anthropic wants to be the AWS of agentic AI (thenewstack.io via hn) Anthropic's Managed Agents platform bundles sandboxing, checkpointing, and persistent memory into a single API layer — and the company's ambitions look a lot less like a model provider and a lot more like AWS.
Running Local Agentic PDF Search with Eno (enopdf.com via hn) eno can drive its full agentic search against a local, open-weight model running on your own hardware. When you do, your PDFs, your queries, and every intermediate step of the agent loop stay on your machine.
Get Your Website/API Ready for Agentic Commerce in 1 Minute (www.startuphub.ai via hn) Free scanner that audits websites, APIs, and MCP endpoints across 7 categories — discoverability, content, access control, capabilities, commerce (x402-mesh), and quality. Public leaderboard, open spec, paste-ready fix prompts.
OpenAI + agentic systems (DFW) (www.reddit.com) i’ve been using OpenAI tools more heavily lately and keep circling back to the same shift: moving from simple chat use into agentic systems. Most people still seem to be using it for Q&A or basic content help, but there’s a lot more happen…
My agent works 3 times… then randomly skips steps and breaks. Same input. Why? (www.reddit.com) I’ve been deep in the trenches building out multi-step agentic workflows, and I’m hitting a consistent wall with what I can only describe as "stochastic decay." The pattern is frustrating: Runs 1 through 3 execute flawlessly, but by the fo…
how do you stop people from finding loopholes in your agents once they're in production? (www.reddit.com) agentic demos always look clean in a controlled setup. the problem that I'm pushing toward real volume now and the adversarial side is getting messy fast.
Fixed the risk of agents disclosing your secrets (www.reddit.com) Why is it considered acceptable by most in the community to have API keys sitting on a file system where the agent is running, with direct access to them, gated by a prompt? This is literally the base security model of OpenClaw and most ot…
Letting AI play my game – building an agentic test harness to help play-testing (blog.jeffschomay.com via hn) Vercel Security Checkpoint | sfo1::1777467624-qE4eB4e2LvmbibEDgl5Ljah0zEqW8iFE
Best Practices to Start with Vibe Coding? Best Local Apps for Agentic Vibe Coding? (www.reddit.com) DISCLAIMER: I am not a programmer nor do I have experience coding. I've been thinking about a small app running on gradio for some time now, and I want to try tweaking some extension for ComfyUI.
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-Unify Architecture (www.reddit.com) SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture. It marks a fundamental paradigm shift in multimodal AI: from modality integration t…
Genuine question for people who have built multi-agent systems in production. How do you handle context continuity across enterprise tools? (www.reddit.com) I've been going down a rabbit hole lately trying to understand how production agentic systems actually work at scale, not just the demo versions. The part that keeps tripping me up is memory and context management across agents.
Is an agentic Spark copilot worth it? opinions? (www.reddit.com) Running Spark jobs on Databricks with 50+ stages per pipeline. Debugging is still almost entirely manual.
OpenGame: Open Agentic Coding for Games (arxiv.org via hn) Game development sits at the intersection of creative design and intricate software engineering, demanding the joint orchestration of game engines, real-time loops, and tightly coupled state across many files. While Large Language Models (…
How are you ACTUALLY running truly asynchronous agentic AI in your business? (www.reddit.com) I'm starting a new company (I will not promote) and I want to hear how you're actually running operations that have little-to-no "human in the loop". Tools like OpenClaw are great for personal use, but how are you leveraging tools/systems…
An open-source platform to auto-update agent skills and discover fresh sources (www.loooop.dev via hn) GitHub obra/superpowers: An agentic skills framework & software development methodology that works. · GitHub GitHub obra/superpowers: An agentic skills framework & software develop… Loop autonomously monitors, evaluates, and updates your a…
The Controllability Trap: A Governance Framework for Military AI Agents (arxiv.org via hn) Agentic AI systems - capable of goal interpretation, world modeling, planning, tool use, long-horizon operation, and autonomous coordination - introduce distinct control failures not addressed by existing safety frameworks. We identify six…
TealKit – A cross-platform UI for local AI agents and MCP (github.com via hn) # 🐦⬛ TealKit The Privacy-First, Infinitely Extensible Agentic AI Platform for Mobile & Desktop TealKit turns your phone and computer into a powerful agentic AI platform with autonomous agents, built-in tools, and unlimited extensibility.…
how are you managing agent-generated code quality? (www.reddit.com) we've been experimenting with agentic workflows for feature expansion, but have a problem: agents can ship PRs faster than senior devs can meaningfully review them. once agents start touching business logic or data transformations, "passes…
Agentic CEO – An AI research organism that hunts, critiques, and evolves itself (github.com via hn) Agentic CEO An autonomous multi-agent research system that acquires knowledge, builds a persistent worldview, and improves itself. 3,700+ knowledge entries.
I've got a feeling that Llamacpp is not the biggest performance bottleneck, but it might be the OpenCode. (www.reddit.com) It looks as if OpenCode introduces an artificial delay in agentic coding. Have you noticed similar issues?
Apple integrates Claude and Codex into Xcode 26.3 for 'agentic coding' (venturebeat.com via hn) Apple integrates Anthropic’s Claude and OpenAI’s Codex into Xcode 26.3 in push for ‘agentic coding’ | VentureBeat Orchestration Infrastructure Data Security More Newsletters Apple integrates Anthropic’s Claude and OpenAI’s Codex into Xcode…
The Full-Cycle Agentic Experience (www.reddit.com) The Full-Cycle Agentic Experience What we're missing, and why it matters more than the models themselves. Think about the last time you bought something in a store.
Agentic AI made DevOps and Agile obsolete (avkcode.github.io via hn) The Self Healing Platform and the Agent Store I think DevOps as a separate identity, and a lot of agile ceremony around it, are already a bit obsolete. Engineers are doing development, operations, and lightweight management at the same tim…
Agentic ML engineer. works with Colab. Zero infra needed. 3x faster TurboQuant (github.com via hn) isanagent An always-on, agentic ML engineer for your workspace — built by ALTAI. isanagent doesn’t just answer prompts: it pushes work toward something shippable — research, code, runs, checks, and handoffs you can actually use.
PI agent integrated with Cline-Kanban repo: All using PI and Qwen 3.6 35B MOE UD 4K_XL (www.reddit.com) Repo: statisticalplumber/kanban at pi-agent-integration Hi Guys, To test Qwen 3.6’s potential, I also wanted the Cline Kanban project to have an open-source agent to work with. The last time I tested Cline Kanban, it didn’t support agents…
PAuth – Precise Task-Scoped Authorization for Agents (arxiv.org via hn) The emerging agentic web envisions AI agents that reliably fulfill users' natural-language (NL)-based tasks by interacting with existing web services. However, existing authorization models are misaligned with this vision.
Agentic Workforce Framework, an operating model for autonomous agent teams (github.com via hn) Agentic Workforce Framework A reference architecture for operating autonomous AI agents as accountable digital workers inside enterprise environments. This framework defines how agents are assigned work, bounded by role, governed by approv…
A 14-day “Growth Forge” sprint: build an AI-powered growth agent on a real stack (www.reddit.com) Sharing something that sits at the intersection of AI agents and growth systems. VideoDB (backend for video/audio for AI agents) is running a 14-day sprint called Growth Forge for 5 builders to design and ship a growth agent on top of an e…
Show HN: I made GAI to have LLM agents in Go without heavy frameworks (github.com via hn) GAI is a flexible Go library for building agent-style applications on top of LLMs. It provides a generic interface for providers and models, prompt and context helpers, and a loop for agentic-calling workflows.
Mario & The Intent-Bearing Agentic Loop (www.reddit.com) Q: When do I need Agents vs. Skills vs.
RTX 3090 + 27B model performance issues (llama.cpp) what am I doing wrong (www.reddit.com) Hey folks — looking for some advice on improving my local LLM setup (and also exploring agentic coding workflows). Current setup: GPU: RTX 3090 (24GB VRAM) RAM: 64GB Using llama.cpp with a Qwen3.6 27B Q6 model (GGUF) Running through OpenCo…
Show HN: Mdspec – auto sync your md files from GitHub repos with wikis (mdspec.dev via hn) We do generate a lots of md files along with our agent based development. Skills, Agent.md, Docs etc.
SimpleBanking sb CLI – Query real German bank accounts from the terminal (balances, transactions, categories, JSON output) (www.reddit.com) Hey r/AI_Agents, I've been building SimpleBanking, an open-source macOS banking app for German bank accounts using the FinTS/HBCI protocol (the standard used by German banks like Sparkasse, Volksbank, DKB, etc.). It now ships with a full C…
What are the limits of the agentic computer features on 5.5? (www.reddit.com) Is this supposed to be an OpenClaw / Hermes Agent competitor ? Am I able to ask it to go on my browser, visit a site I’m logged into and gather info?
Working with Claude Code: A Field Manual (blog.iannelson.uk via hn) Earlier this week I published a reflective post on how agentic coding has changed my working day and the shape of the profession. I now want to turn to the other side of that coin: not the philosophy, but the mechanics — the habits, the wo…
Agentic Company OS update: project-scoped runtimes, governance UI, snapshots/replay, skills, and operating models (www.reddit.com) I shared this project here before when it was mainly a governed multi-agent execution prototype. I’ve kept working on it, and the current implementation is materially more complete, so I wanted to post an update with what actually exists n…
Building a full-stack app with Wasp, an agent-friendly web framework (wasp.sh via hn) From 10 Failed Stacks to Production: How a Data Scientist Built a Job Board with Wasp, a Full-stack Framework for the Agentic Era Hireveld is currently down while Marcel works on a major refactor - but it's real, we swear! It'll be back up…
Is anyone else way faster with AI in familiar stacks and way slower in unfamiliar ones? (www.reddit.com) Been using agentic coding workflows seriously for about a year now and I've finally figured out the pattern behind why it feels magical half the time and broken the other half. At my day job, where I know the stack and have intuition about…
Best production agentic frameworks (www.reddit.com) Hi, so I’m looking for a framework which is also provider agnostic , like pi . But I need it for python and I need it to be production ready.
Show HN: We built Cursor, but for data transformations [Open Source] (github.com via hn) Agentic & No-Code Data Transformations Vibe coded pipelines: say hello to accuracy and maintainability. Website · Documentation · Issues · Contributing What is Visitran?
Google's 8th Generation TPUs Power the Agentic Era [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Speeding up agentic workflows with WebSockets in the Responses API (openai.com via hn) could not extract summary
Show HN: API Ingest – Agentic Search (Inter) API Docs (github.com via hn) 1. CC / Codex dont handle API Docs well enough No matter what I do, I run into bad requests with claude, day in, day out.
Show HN: Sift – save AI tokens in Codex/Claude by summarizing command output (github.com via hn) I made a small skill/script for agentic coding workflows: https://github.com/panpeter/sift-skill The idea is simple: when a command like cargo test, pytest, npm test, or ./gradlew test prints a lot of output, that raw log often gets pulled…
Show HN: We open-sourced a 6-library governance stack for AI agents (Python) (news.ycombinator.com) Our team has been deploying AI agents in enterprise environments for the past 2 years, across 60+ deployments. The same governance problem kept recurring: how do you certify reliability, enforce policy, route and orchestrate context, monit…
Cursor partners with SpaceX on model training (cursor.com via hn) Cursor partners with SpaceX on model training Cursor is partnering with SpaceX to accelerate our model training efforts. We released Composer less than six months ago as our first agentic coding model.
How do you decide on chunking strategy and top-k in Agentic RAG? Looking for practical advice (www.reddit.com) Hey, I'm building an Agentic RAG pipeline and struggling with two decisions: Chunking strategy — fixed-size, semantic, or hierarchical? In an agentic setting where the agent can re-query iteratively, does it make more sense to use smaller…
X402 and Agentic Commerce: Redefining Autonomous Payments (aws.amazon.com via hn) Managing context in long-run agentic applications (slack.engineering via hn) The Bitter Lesson of Agentic Coding (agent-hypervisor.ai via hn) MongoDB MCP (www.reddit.com) Using closed financial markets with deterministic goals for agent behavior improvements (www.reddit.com) Which AI Agents SDK allows low latency agents w support for skills etc? (www.reddit.com) Show HN: AI Primer – A Searchable AI Changelog for AI Engineers and Creatives (www.ai-primer.com via hn) I'm completely lost in the Agentic Maze. What level to learn. how to organize stydu (www.reddit.com) Show HN: Agentic Dev – AI dev-tools news, curated daily by Claude (agenticdev.blog via hn) OpenAI released a major update to Codex, used by over 3 million developers weekly, adding background computer use, an in-app browser, image generation via gpt-image-1.5, more than 90 new plugins, GitHub PR review support, SSH connectivity,…
Why AI Agents are bad at “generating a business idea” (www.reddit.com) My opinion is it is a matter of structured approach. Of course when you just ask Claude to “find top apps in AppStore and tell me what app should I build” you will get as generic answer as your question.
Fast local LLM to generate CLI commands from prompt? (www.reddit.com) GitHub copilot CLI used to do this but now it’s a full agentic coding environment. Basically, I can’t remember all the options to every Linux command.
Built a full-stack charitable giving SaaS as a solo developer with agentic AI (www.pifster.org via hn) PIFster - the Pay It Forward Charity Did you know there are 1.8 million nonprofits in America? Most are struggling to be heard, but PIFster is changing that.
[Claude Code] Stuck in 57+ minute loop for routine fixes (Opus 4.7) (www.reddit.com) I'm running into a severe performance hang with Claude Code (Opus 4.7) today. I provided a relatively straightforward prompt to fix some hydration errors, add two stub routes, and perform a theme audit (string replacement).
Cowork Orchestrator Patterns (www.reddit.com) While working in Cowork, I have been experimenting with designing plugins that try to apply some established agentic patterns to help manage the context window. The problem that I'm running into is with Cowork the main orchestrator is the…
What is the simplest architecture for running a multi-agent system at scale? (www.ashpreetbedi.com via hn) Scaling Agentic Software: Part 1 What is the simplest architecture for running a multi-agent system at scale? I want to deploy agents as a real service.
Show HN: Marky – A lightweight Markdown viewer for agentic coding (github.com via hn) Hey HN, In this age of agentic coding I've found myself spending a lot of time reviewing markdown files. Whether it's plans or documentation that I've asked my agent to generate for me, it seems that I spend more time reading markdown than…
Kelvin Claw: A secure, modular agent harness with supply-chain validated plugins (agentichighway.ai via hn) Agentic Highway Team KelvinClaw: A secure, modular agent harness with supply-chain validated plugins An agent runtime designed for zero-trust environments from the ground up. Building secure agent systems at scale is a different problem th…
I built a self-evolving agentic loop that ran 104 iterations autonomously to find questions that break every LLM — here's the architecture (www.reddit.com) Why I built this: I wanted to find the next "strawberry problem" — simple questions any kid can answer but every LLM gets wrong. Instead of manually testing questions, I built a system that does it autonomously.
A Black-Box Contract Engine for Agentic Software Development (github.com via hn) Project Dojo A Black-Box Contract Engine for Agentic Software Development Dojo is a declarative testing engine built in Go. It acts as a transparent Man-in-the-Middle proxy between your Software Under Test (SUT) and its dependencies.
Ask HN: We dont need a programming language now? (news.ycombinator.com) I've seen agentic IDEs now Cursor or Antigravity and main trends seems to be development with just ideas, where although the changed lines are shown, its becoming less and less visible. If we are becoming language agnostic, shouldn't we op…
Solving the "Agentic Kill-Switch": Moving from Prompt Guardrails to a Python-native Safety SDK (www.reddit.com) The biggest hurdle for taking agents from "cool demo" to "production tool" is the lack of a reliable circuit breaker. We're currently relying on the LLM to "behave" via system prompts, but as we know, jailbreaks and hallucinations make tha…
Ask HN: Which LLM model and agentic CLI are you using for local development? (news.ycombinator.com) I’ve been testing a handful of models the past few weeks, but I still haven’t settled on one yet… I’m curious to see what models, their sizes, on what hardware, and which agentic tool people are using
Scaling from single-repo Claude projects to a multi agentic workflow (www.reddit.com) Hi everyone! Just a quick exchange on what I am using — and I'd love your take on it 🤖 So far I have mainly been doing one-off projects, setting up Claude in a single repo at a time.
Ask HN: What standards or protocols exist for AI Agent permissions (news.ycombinator.com) Curious what standards exist for AI agent permissions. Something like Linux read, write, execute types, but for AI agents.
The (Mostly) Agentic SDLC (amoshaviv.com via hn) Monday, 12:00. Grace, the CEO of ACME Corp, just finished her Q2 leadership meeting.
Tradclaw: an open source AI mom for agentic parenting (twitter.com via hn) My family assistant "Finley" is a full fledged member of the household , and I just open sourced her for all the Very Bad Moms and Dads ™️ out there that just need a little 🤖 help. Wanna get started right away?
Is qwen3 coder next still relevant with qwen3.5 release for agentic coding? (www.reddit.com) Basically the title. I know it will depend on your quant, but with 48gb of vram inbound, I'm curious on the communities opinion before I get the chance to vibe check.
What are the key features that make an AI system truly "agentic"? (www.reddit.com) Here's the cleanest breakdown I've seen: Autonomy – Acts without constant human prompting Goal-Oriented Behavior – Works toward defined outcomes, not just single responses Adaptive Learning – Gets better from outcomes over time Multi-Step…
Show HN: A Bomberman-style 1v1 game where LLMs compete in real time (github.com via hn) A few weeks ago, ARC-AGI 3 was released. For those unfamiliar, it’s a benchmark designed to study agentic intelligence through interactive environments.
Show HN: On-Device vs. Cloud LLMs for Agentic Tool Calling in a Real iOS App (subralabs.com via hn) We built an AI concierge into a resort directory app for iOS. The feature needed to search a dataset of ~85 properties, apply filters, find nearby airports, and respond conversationally in Italian.
OpenClaw Self-Improvement Loop: adversarial agentic self-modification workflow (github.com via hn) An adversarial framework for AI agent self-modification, built and battle-tested in production. Inspired by karpathy/autoresearch.
Agentic dashboard analysis (www.reddit.com) Hi all Like most of us the execs at my company are big into AI. I saw a potential implementation to get myself more experienced with agents by having an agent perform a daily analysis on a dashboard to perform summaries and anomaly detecti…
Show HN: A better alternative to CLI and MCP for local tools (github.com via hn) I've created an alternative to CLI and MCP for locally running agentic tools. It uses Unix-based OS's named pipes, which means the client has quick IPC with the tool and it can have in-memory state.
Observing the shift toward open-weight models for agentic coding workflows (www.reddit.com) I've been practically evaluating some of the recent open-weight mixture-of-experts models, specifically focusing on their application in complex software engineering and agentic coding workflows. established pattern has typically involved…
Is my 'Retry Tax' math correct for DeepSeek V3/V4 agents? (Project Feedback) (www.reddit.com) Polyscope – Agentic coding environment for Laravel (getpolyscope.com via hn) The time where humans write code is over. This is the new cockpit.
Object Storage and WAL: Lakebase Postgres for the Agentic Era (www.databricks.com via hn) Change the way agents work with Postgres by treating WAL as a durable source of truth by Cassie Murray and Carlota Soto Agents that interact with a traditional OLTP database often create bottlenecks at the storage layer. New deployments, c…
Show HN: CRT – a local code review tool for agentic development (github.com via hn) I built crt to help reduce the burden of reviewing AI generated code. As fun as it is to vibecode software to production, sometimes you still need to understand and be responsible for the code that agents write.
Flawed benchmarks (epoch.ai benchmark registry filtered by flawed) (epoch.ai via hn) A registry of 85 AI benchmarks covering mathematics, software engineering, agentic workflows, games, and more. An index aggregating many different benchmarks into a single, general capability scale.
Gemini 3.8 Flash outperforms SOTA models at agentic CAD coding (www.partforge.ai via hn) 5 models, 55 tasks, 675 attempts. Which models can actually do parametric CAD, graded on build success, measured geometry, and a checklist judge.
Capability Attenuation in Agentic Hierarchies (kevinhoffman.blog via hn) Capability Attenuation in Agentic Hierarchies In this post I won’t be discussing simple one-shot request and response behavior. While that’s interesting, it isn’t the problem I want to explore.
Agent-First vs. Agent-Second Engineering (thomasvandongen.dev via hn) Lately I’ve been thinking about two agentic engineering paradigms, which I call agent-first and agent-second engineering. The difference is who writes the first code: the human or the agent.
Ask HN: What do your agents do? (news.ycombinator.com) I have been working for a few years in Computer Vision for manufacturing, so I have kinda been in a bubble as far as production goes. I am thinking of diversifying a bit and get to work on the trendy stuff: agents.
BYOK vs. fixed subscription like GitHub Copilot, codex or Claude (news.ycombinator.com) has anyone does the comparison between paying a fix subscription(pro max) vs BYOK for a hobby project and exploration I'm curious how much does it cost in comparison. I haven't explore these agentic workflow as much, besides checking your…
Inside OpenAI’s agentic software factory (newsletter.pragmaticengineer.com via hn) It’s rare to work with an unlimited token budget, but at OpenAI, that’s what all engineers, researchers, finance colleagues, and marketing folks do. Recently, I visited one of the world’s leading frontier labs to find out how OpenAI operat…
Show HN: AutoBot – live voice control for long-running AI work (github.com via hn) I wanted to manage long horizon agentic workstreams via voice, then put my phone down, and have a harness manage completion - extending into full computer use. I was trying to build a personal Jarvis, so I benchmarked AutoBot to see how cl…
Agentic Societies Need a Social Harness (social-harness.org via hn) #How do we use agents today? Let's look at an example of how two of us (Ratul, a professor, and Tapan, his student) use agents for work today.
Spec-Lock-Diff: a framework for agentic dbt development (github.com via hn) Spec-Lock-Diff English · Português (pt-BR) A framework for dbt development using AI agents. The goal is to reduce the main risks that arise when an agent writes SQL: The framework boils down to three phases: Spec — The human defines, in st…
ApowerB – open-source runtime for AI agents (Apache 2.0) (github.com via hn) apowerb The open-source agentic framework to build, orchestrate, and operate production AI agents. Documentation • Quickstart • API Reference • Deployment • thaink2 This repository is the open-source core.
Agentic local development tool for WordPress (developer.wordpress.com via hn) Develop locally with WordPress Studio. Build and test themes, plugins, and full sites, then sync or export when you're ready to go live.
Scaling Trust Arena: An agentic economy with a multi-million dollar prize pool (scalingtrust.org.uk via hn) Update on the Scaling Trust Arena Scaling Trust is a £50 million R&D programme actively funding the fundamental research and open-source infrastructure that enables secure, scalable multi-principal multi-agent coordination across digital a…
Show HN: Ori, an agentic runtime for infrastructure that must ask before it acts (github.com via hn) Give your devices a brain. Ori — the agentic IoT runtime for infrastructure intelligence IoT devices do not need more data.
We struggled with DevOps, so built a open source agentic platform just for this (news.ycombinator.com) As a developer with some years of experience and with 150+ projects completed, me and my team learnt devops the hard way. Multiple fixes, compatibility issues, server issues, random vulnerabilities appearing out on framework level out of n…
Ask HN: Shifting Reoccuring Agentic Workflow to ML Process? (news.ycombinator.com) I know I can simply ask myaAgent or Google this question but would prefer to ask HN for any practical real examples used day to day. Question: For the people who build agentic workflows, how do I implement a ML habit where the agent learns…
GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos (noma.security via hn) GitLost: How We Tricked GitHub’s AI Agent into Leaking Private Repos TL;DR: Noma Labs discovered a critical prompt injection vulnerability within GitHub’s new Agentic Workflows, allowing an unauthenticated attacker to silently pull data fr…
Agentic Primer – how I shipped a commercial app without writing code (github.com via hn) Building a Commercial App with Developer Agents How I built and shipped a real product in 65 hours — without writing a single line of code With mature agents, the prompt only carries intent. Everything else the agent needs already lives in…
AntFlow AI: Spec-Driven Agentic Development (geekyants.com via hn) AntFlow AI: Agentic AI Framework for Spec-Driven Development Move from business intent to verified software delivery with clarity built into every step. AntFlow AI is the agentic AI software development framework from GeekyAnts.
Agentic Coding Strains CI: Scaling Test Impact Analysis at Anthropic (claude.com via hn) Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic Our CI job volume increased 25x over 6 months.
Improving Throughput by Optimising KV Cache Efficiency for Agentic Workloads (j9s.io via hn) Improving Throughput by Optimising KV Cache Efficiency for Agentic Workloads Boosting throughput by leveraging characteristics of agentic workloads to increase efficiency of the KV cache. Throughput is defined by how quickly we can serve r…
Show HN: I rebuilt a 4-year-old app in 5 days using many agents – here's harness (mega.dev via hn) The GPT-6 Astra release showed us a new level of LLM capabilities, not just in programming, but also in computer use, browsing, math, science, cybersecurity, abstract reasoning, and agentic, multi-turn, and long-context reasoning. Let’s se…
My Complete Agentic Coding Setup Tech Stack Feb 21 2026 – Updated Jun 29 2026 (hboon.com via hn) My Complete Agentic Coding Setup and Tech Stack I get asked what a complete agentic coding setup looks like. After 30 years of programming and over a year of using coding agents daily, here’s what I use — the full stack, the agents, the de…
India's First AI Agentic E-Commerce Platform (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Teaching Novice Computing and Programming in the Agentic AI Era [pdf] (cs.brown.edu via hn) could not extract summary
Show HN: Agentic Deployment and Hosting (sitedropper.com via hn) Hey folks, we recently launched Sitedropper. It's an agent first deployment and hosting platform with secure private sharing or instant live publishing.
Learnings from 12,000 Agentic Code Reviews (blog.watson-labs.co.uk via hn) The setup Gymwasp is a Swamp customer, currently in stealth. I helped them get their software factory up and running: the pipeline that orchestrates code planning through to shipped code, agentically using swamp as the harness.
Terms.txt: A Consent and Compensation Protocol for Agentic Web Access (arxiv.org via hn) The open web ran on an unwritten bargain: sites admitted crawlers, and search engines sent visitors back. Public measurements show that bargain breaking under AI crawlers and agents.
Build Agentic Memory That Keeps Your Loops Alive and Sharpens Them Every Run (memanto.ai via hn) The Practical Course: Build Agentic Memory That Keeps Your Loops Alive And Sharpens Them Every Run A hands-on course on agentic memory: first make your loop survive long runs, then make it compound across them. Every step ships with copy-p…
Awesome-agentic-payments – A curated list of tools for agentic payments/commerce (github.com via hn) Awesome Agentic Commerce A curated list of protocols, specs, SDKs, and tools powering the emerging agentic commerce stack: AI agents that discover, buy, and manage orders autonomously. Maintained by Bitrefill.
Reasoning Through Agentic Memory (handsdiff.substack.com via hn) I think many people overestimate the speed of context length increases for frontier models widely available to users. Below are my predictions for growth here.
Show HN: HolaOS––An Opensourced workspace that alternative to Claude (github.com via hn) Open-source agentic workspace enterprises can make their own. Connect the systems you already run — 100+ integrations, MCP, chat tools, apps, browser, local files — with shared memory.
Show HN: Agentwire, other people's shipped agent tooling (agentwire.thecompound.tech via hn) Every MCP server and agent harness other people already use Latest - mastra-ai/mastraMastra is the modern TypeScript framework for AI-powered applications and agents.Repo - melih-unsal/DemoGPT🤖 Create agentic apps in a second with your pro…
Show HN: Cadenya – An Agent Runtime (www.cadenya.com via hn) My name is Robert, and I'm the founder of Cadenya, an agent runtime. Cadenya is not like other managed solutions.
Lowdefy v6: WebSockets, Notifications, Crons, MCPs, agentic code tools (lowdefy.com via hn) Next.js is out, Hono and Vite are in. Websockets, notification emails, dynamic pages, scheduled endpoints, LLM steps in routines, and a dev server built for coding agents.
Designing Tests for Agentic AI Tools (automatedteach.com via hn) Designing Tests for Agentic AI Tools A passing test suite doesn't tell you how the agent got there [image created with ChatGPT Images 2.0] Today, a more technical note on the evaluation of agentic harnesses, drawing from the process of dev…
A functional taxonomy for LLM inference in agentic tasks (jeffauriemma.leaflet.pub via hn) A framework for understanding where LLM inference goes during agentic tasks: Initialization, Reasoning, Orchestration, and Synthesis. Consider "Reasoning Yield" as a way to measure how much inference is spent resolving uncertainty rather t…
Agentic Alienation (www.lorenstew.art via hn) What happens to ownership, engagement, development, and connection when AI agents do the work? A personal essay about agentic alienation.
The rise of agent-driven heterogeneity at the edge (www.efficient.computer via hn) The rise of agent-driven heterogeneity at the edge Agentic systems won’t just move to the edge—they’ll fundamentally change what runs there. Today’s edge systems are built around predictability.
Humble Bundle – 24 Linux, Cloud and Agentic AI Books (www.humblebundle.com via hn) Play Frostpunk 2 & Sonic x Shadow Generations with September’s Humble Choice! Bundles Games Books Software Store Popular On Sale Bestselling New Releases Pre-order Books Software Deals Under $5 Deals Under $10 Deals Under $20 Great on Hand…
Solo.io Pushes Agentic AI Governance to the Desktop: Open-Source Agentdesktop (techstrong.ai via hn) TL;DR — Key Takeaways - Solo.io’s open-source agentdesktop project is designed to help enterprises govern AI-agent tools such as Claude Code and Codex running on employee workstations. - The platform adds AI-specific inventory, policy mana…
Practical Agentic RAG patterns implemented with LangGraph (github.com via hn) Agentic RAG — Four Working Patterns with LangGraph Four self-contained Jupyter notebooks, each implementing a different way of making a Retrieval-Augmented Generation (RAG) pipeline "agentic" — able to decide, check itself, correct its own…
Ask HN: How to roll out Agents for knowledge workers? (news.ycombinator.com) I work in a medium-size government agency in Germany, mostly comprised of middle management. We have already built and rolled out a web-based "internal GPT" with LangGraph and a custom frontend, based on Azure OpenAI, but it of course lack…
Principles for building agentic cloud from scratch (console.cantelop.dev via hn) The design philosophy behind Cantelop: minimal setup, any agent harness, Sessions as actors, opinionated infrastructure, performance on the critical path.
Local TUI issue board for agentic workflows with time travel, backed by Git (twitter.com via hn) Jonatan Lampa on X: "Agentic workflows needs new tools. Epiq lets you time-travel the board, inspect how the plan changed, see who changed what, when, and why.
Show HN: Devbar – Point at what to change. Leave a comment. Your agent ships it (devbar.sh via hn) I built this toolbar to power my agentic ui/ux design flows. Too many times I found myself copy/pasting screenshots into Claude.
Show HN: Mu – an agent with actual command line experience (github.com via hn) I'm an old school user that finds the AI agent TUIs too magical, so I experimented with a new agent UX. I implement mu as a shell plugin (zsh and fish for now).
Show HN: Meclaw – where agents build agents: one Rust binary, no loop shipped (github.com via hn) meclaw Where agents build agents. An agentic build system for agentic systems.
Should coding agents be allowed to touch infrastructure? (catalyst.zoho.com via hn) Build with your preferred agent Catalyst integrates with the tools developers already use for agentic coding. Connect your assistant, keep your workflow, ship faster.
Show HN: Profound Academy – an agentic course builder for hands-on courses (profound.academy via hn) I built an AI course builder that helps you create hands-on course materials, interactive tutorials, exercises, and automatically checks the work submitted by students. It should feel like Claude Code or Codex for building courses.
OWASP Top for Agentic Applications – The Benchmark for Agentic Security (genai.owasp.org via hn) GenAI Security Project – Agentic Security Initiative (ASI) & Agentic Top 10 Leadership, Blog Co-Authors John Sotiropoulos, OWASP GenAI Security Project Board Member & ASI Co-lead, Agentic Top 10 Chair Keren Katz, Agentic Top 10 Co-Lead, OW…
Three Agents and a Hoodie: A2A Across the Purchase Lifecycle (www.umai-tech.com via hn) Three Agents and a Hoodie: A2A Across the Purchase Lifecycle A2A joined the Agentic AI Foundation alongside MCP. I built a three-agent commerce mesh — checkout, shipping, claims — to see what it actually buys you.
India preparing rollout of agentic payments on UPI (www.reuters.com via hn) could not extract summary
Skforecast-AI – Agentic time series forecasting in Python (github.com via hn) | | | | --- | --- | | Package | | | Skforecast | | | Meta | | | Testing | | | Donation | | | Community | | skforecast-ai is an AI forecasting assistant that pairs a deterministic engine, powered by skforecast, with an LLM reasoning layer.…
EMVCo Requests Feedback on Agentic Payments Framework (www.emvco.com via hn) 01 September 2026 – EMVCo – the technical body that creates and manages EMV® Specifications and programmes – has released a draft framework to help promote secure, interoperable and scalable card-based agentic payments. The framework provi…
Agentic Determinism Index (open source): Find where deterministic AI agents run (github.com via hn) Agentic Determinism Index (ADI) A public harness that asks one narrow question of hosted LLM APIs: If I send you the exact same request N times, concurrently, and again across days, how identical are your answers? No benchmark of intellige…
Arise – Agentic Runtime Identity Security Enforcement (requestrocket.com via hn) A category that did not exist six months ago SACR – Software Analyst Cybersecurity Research – has given the industry a name for a problem security teams have been circling since agents went into production: ARISE, Agentic Runtime Identity…
Agentic/Human Contractor distributed marketplace (news.ycombinator.com) Hello y'all I am new here. I would like to know if there is a marketplace for something that I am trying to build or I from deep trenches of my heart I wanted it to succeed.
Show HN: Training and trading agentic intelligence with no GPU and MCP (github.com via hn) Now that most foundation models are nearly satured into mature performance, I believe specialization ought to move forward from finetuning or just skill.md or design.md ... instead you could use Tetrees Agent where you can train and trade…
Agentic Skill Decay (addyo.substack.com via hn) Mastery still comes from doing the reps. Before agents, I got my reps as part of writing code: try different approaches out, debug what went wrong, review other’s code, read a lot.
Ask HN: What companies have discussed their AI/agentic workflows? (news.ycombinator.com) Looking to learn what agentic engineering looks like in different places. My current favorite deep dive into a professional coding environment has been Wes McKinney’s “How Kenn is doing Agentic Engineering“ [1], but I’d like to know if any…
Experiments with AI – Structure of Third Party Agentic Apps (karankurani.com via hn) This is a technical companion piece to building out an experimental app called CareLoop. What CareLoop is and why it was built is written here.
Agentic Inequality (arxiv.org via hn) Autonomous AI agents capable of complex planning and action mark a shift beyond today's generative tools. As these systems enter political and economic life, who can access them, how capable they are, and how many can be deployed will shap…
Only believe what you can validate: a verification framework for agentic AI (devblogs.microsoft.com via hn) How to read this article This article picks up where my January piece left off. My catch phrase is still the same: "only believe what you can validate".
Show HN: Issue tracker that replays workflows, deeply integrated with the code [video] (www.youtube.com via hn) Epiq is an issue tracker that is distributed, Git-native, and most interestingly, can replay state as a movie on demand. This solves one of the most difficult problems with agentic workflows - auditing and tracing in a multi agent environm…
I ran out of AI tokens in one app while holding unused tokens in another (news.ycombinator.com) The problem is simple: AI tokens are locked to individual products. context : I was using both an agentic IDE and a Hostinger deployment agent.
Efficient Decode Context Parallelism with vLLM for Long Context Workloads (vllm.ai via hn) Efficient Decode Context Parallelism with vLLM for Long Context Workloads 1. Introduction Long-context inference is becoming essential for agentic AI, where assistants may need to reason over large code repositories and long chat histories.
Why Agentic AI Needs a Strong Identity Foundation (www.nist.gov via hn) a NIST blog As AI matures, enterprises and customers are rapidly deploying agents seeking to unlock the next level of automation and productivity. Agentic AI shows potential to handle a multitude of use cases, from buying personal items on…
Making Your Data Ready for Agentic AI (martinfowler.com via hn) Making Your Data Ready for Agentic AI For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomplete or wrong. Autonomous agents supply none of that.
Show HN: Z, minimal agentic harness for engineers that just works (github.com via hn) The core philosophy for Z harness is transparency and minimalism. All tokens are shown.
Agentic Payments: How Agents Move Money (zdne.org via hn) A map of how money moves when AI agents are involved: commerce checkout, direct agent-to-resource payments, and agent-operated banking.
Ask HN: How does manual QA fit into your process? (news.ycombinator.com) We're seeing increasing efforts to improve developer cadence - LLM-improved auto-complete, agentic coding, agentic code review and agent-generated automated testing. So how does Manual QA still exist in your process?
PaaS IaaS GaaS all in one. And 2 AI <compound> models (news.ycombinator.com) Independent inventor. Invented a bft protocol.
Aperture GA: Building a home(lab) for agentic AI (tailscale.com via hn) We started building Aperture 10 months ago to demonstrate that you don’t need to choose between ease of use, identity-aware networking, and robust safety when using agentic AI with Tailscale. Our solution started as an LLM proxy that elimi…
Show HN: AI scientist builds an open-source Codex Micro from scratch for $40 (iluvatarlabs.com via hn) A few weeks ago, we open-sourced our design for a cost-effective but full featured alternative to OpenAI x Work Louder's sold out Codex Micro macropad. And now, we're pleased to share that we've actually built them and they work!
Intel Crescent Island GPU Flexes 32 Xe3P Cores, 480GB LPDDR5X for Agentic AI (hothardware.com via hn) Intel Crescent Island GPU Flexes 32 Xe3P Cores, 480GB LPDDR5X For Agentic AI What kind of hardware do you need for AI processing? Well, every kind, because "AI processing" is a very broad term.
Created my own LLM Agentic Coding App with a desktop client (apps.apple.com via hn) Download pocket/grammer by CHRISTOPHER MICHAEL STAATS on the App Store. See screenshots, ratings and reviews, user tips, and more apps like pocket/grammer.
Iterize IDE – open-source multi agentic coding platform (kosev-lex.com via hn) Research & Projects This page will be regularly updated with various research and projects I am working on. This is where you will be able to find my original publications on exciting and expansive topics.
Granite 4.2 brings native reasoning to enterprise agents (research.ibm.com via hn) Granite 4.2 brings native reasoning to enterprise agents IBM’s new open Granite models are designed for agentic AI, combining reasoning, tool use, coding, instruction following, and speech capabilities. Large language models are evolving b…
A Manifesto for Responsible Agentic Coding (www.techwerkers.nl via hn) A Manifesto for Responsible Agentic Coding Table of Contents Somewhere between mindlessly giving in to the hype created by big tech and entirely boycotting generative AI there’s a middle ground that lets developers make good use of the new…
Portable Computer is Perplexity's new local AI agent – why it's a game changer (www.zdnet.com via hn) Portable Computer is Perplexity's new local AI agent - why it's a game changer Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways - Perplexity's new agentic Portable Computer runs AI locally.
Agentic Resource Discovery (ARD): An open specification for agent discovery (aws.amazon.com via hn) Artificial Intelligence Agentic Resource Discovery (ARD): An open specification for agent discovery How AWS Agent Registry and the Agentic Resource Discovery (ARD) specification enable cross-environment discovery for your agents As organiz…
Show HN: CoolPlugz – A Claude MCP that turns Jira tickets into a merge-ready PR (github.com via hn) Hi folks, I've been a fullstack dev and solopreneur for 9 years and the last 3 I have been using AI and MCP related tools in my dev environment. Most recently I had a lot of work for a blockchain client and started using claude code more t…
Show HN: Agent Notifier (notifier.aicrew.in via hn) A missing piece of your productive agentic harness. Install the app - ios https://apps.apple.com/app/id6763598043 Android: coming soon.
Ask HN: Is there a stronger moat in marketplaces in the age of SaaS AI Slop? (news.ycombinator.com) Basically, what I’m trying to say is: software is cheap to produce now, but not necessarily easy to customize. It seems like everyone is honing in on distribution and capital as the moat.
Adapting Fossil-scm as a platform for AI agentic workflow (github.com via hn) About Fossil Fossil is a distributed version control system that has been widely used since 2007. Fossil was originally designed to support the SQLite project but has been adopted by many other projects as well.
Faber – open-source coding agent that uses a code graph to navigate repos (www.npmjs.com via hn) Faber: an agentic AI coding assistant for your terminal — streams, edits with diff approval, runs your tests, and remembers your project across sessions. Faber A cost and performance efficient agentic AI coding assistant for your terminal.
Apexyx Mesh: offline agent mesh for Termux/Android with 421 hermetic self-tests (github.com via hn) APEXYX Mesh — a self-testing agent economy that runs on a phone termux · android · local-ai · agentic · offline-first · self-hosted · sqlite Pure Python 3 + bash. No cloud, no daemons you didn't start, no pip installs.
Show HN: Is-agentic – Score how agentic your product and site is (is-agentic.com via hn) Score how agentic your site is Enter a URL for a score based on what agents can discover, access, and use. Every scan is run by Ora.
Agentic AI overwhelmed CI, and test selection cut queueing from hours to minutes (humansystems.dudzik.co via hn) The last time I had to think seriously about CI throughput, I was working on the CI/CD team at Shopify. I didn’t expect to run into the same kind of scaling problem on a hobby project with one engineer.
GitHub Copilot coding agents can now work from Slack (github.blog via hn) The new GitHub Copilot experience in Slack The GitHub integration in Slack now brings the agentic capabilities of GitHub Copilot CLI and the GitHub Copilot app into Slack in public preview. You can work with @GitHub to plan changes, invest…
Riding the hype of outbid but for agentic payment x402 (basebid.lol via hn) - Min bid - $1 - To take a slot - +$1 - Fee on deposits - 5% - Refund when bumped - 100% We never hold your deposit — it's in the contract. Boardpage 1/1 - 1Robinhoodapp robinhood.com The leader in trading $30locked by 0x250a…05f72 clicks…
ChatGPT-Taught Experts Are Crippling Agentic AI (msukhareva.substack.com via hn) For decades AI was an obscure field covering Natural Language Processing, Computer Vision, Bioinformatics and similar areas. People considered it a difficult research area with unclear value.
System reminders – how Claude Code steers itself (michaellivs.com via hn) System reminders - how Claude Code steers itself Steering an agent is the act of reinforcing good behaviors and discouraging bad ones. Most recent models are amazing at agentic work, but LLMs are non-deterministic by nature, which means th…
ArchSpec: Executable Architecture Specification for Ruby's Agentic Coding Era (paolino.me via hn) ArchSpec turns your architecture into an executable spec: declare components and boundaries in one Ruby file, and every change gets checked, whether a person or an agent wrote it.
Agentic Programming and the Lust for Power(?) (johnoestmannmusic.com via hn) I say please and thank you to my AI Agents. This is partly because there is a non-zero chance they are experiencing something akin to consciousness, but also because of how these interactions may be reshaping my own psychology.
Show HN: Hackerznews – Yet Another HN Client (hackerznews.vercel.app via hn) Around a little more than a year ago, thanks to the advancements in agentic coding, I created my personal ideal HN reading app. There are oh so many of those, but this one is mine.
VulnBench: Can LLMs find the same security bugs twice? (vulnbench.com via hn) A Snyk benchmark initiative Can LLMs find the same bugs twice? A repeatability and Snyk-reference agreement study We ran the same agentic security review five times against inspectable JavaScript projects to measure what recurs, what varie…
Input Required: Agentic Architecture Framework (www.agenticaf.io via hn) Welcome to the Agentic Architecture Framework Vendor-agnostic, governance-first guidance for building agentic systems that are safe, reliable, and scalable. Browse the whitepaper in the menu — or put AAF directly into your AI tools.
Markdown-den as an issue tracker for agentic workflows (www.markdown-den.com via hn) markdown-den as an Issue Tracker for Agentic Workflows By Diego Guridi TL;DR: Five numbered folders are a workflow. A markdown file is an issue, and moving it between folders is how its status changes.
Cost-Aware Optimization for Agentic Query Execution (arxiv.org via hn) Classical query optimization searches over algebraically equivalent plans that differ only in cost. This assumption breaks once LLM-backed operators enter the picture: their placement, ordering, and granularity jointly determine both dolla…
Show HN: AgentBadge – Agent Readiness Scoring for APIs (SEO for AI Agents) (agentbadge.xyz via hn) Agent Readiness is the ability of your API to be discovered, understood, and used by an AI agent — without a human intervening. SEO for the agentic web.
Agentic Fitness Functions: Extending Evolutionary Architecture (www.infoq.com via hn) Deterministic rules safeguard hard metrics, but what about architectural intent? Discover how agentic fitness functions combine AI agents and versioned rubrics to evaluate complex, judgment-heavy concerns—such as boundary fidelity, semanti…
Ask HN: Are AI credits too locked to individual applications? (news.ycombinator.com) The problem: I already had AI credits, but those credits were locked to one application. context : I was using both an agentic IDE and a Hostinger deployment agent.
Agent2Agent (A2A) joins Agentic AI Foundation (AAIF)'s open agentic stack (aaif.io via hn) Agent2Agent (A2A), the open standard for inter-agent communication, is joining the Agentic AI Foundation (AAIF) as a hosted project. A2A defines how agents discover each other, delegate tasks, and exchange work, regardless of what framewor…
Build T-Shaped Agents, Not Assembly Lines (sqlhammer.com via hn) The most expensive decision in an agentic system is not which model you run. It is how you divide the work.
Return of the Spec: Why AI agents are reviving the software specification (caines.ca via hn) Return of the Spec Agentic development is bringing a resurgence of "the specification". Maybe you've been writing detailed specs all along, but I've been mostly following Working software over comprehensive documentation.
Inference Costs per Agentic Workflow to Increase More Than Fivefold Through 2028 (www.gartner.com via hn) could not extract summary
Operationalizing agentic AI: The Day 0-2 blueprint for enterprise infrastructure (www.redhat.com via hn) Learn how Red Hat AI's platform supports bring your own agent (BYOA), providing operationalization for agent frameworks without code changes.
Thesis – Agentic PE: mature software as the cash cow (research.oguzbilgic.com via hn) Discussion-stage bet, born 2026-08-17 from the Bending Spoons print. The claim: agentic AI turns mature, sticky, non-seat-based software into acquirable cash cows, and the operators who industrialize
Rysh – an AI harness where Claude and Codex agents work as a team (Go) (github.com via hn) rysh-cli-code The Rysh CLI: an agentic terminal multiplexer written in Go. Tabs, panes, splits, and vim/htop working exactly as you expect — except every pane is also an agent that can answer prompts and call tools.
Agentic File Manager for Mac (twitter.com via hn) Nabarun on X: "Introducing agentic file manager for mac. treat your directory/repo as a city.
Agentic AI costs set to balloon fivefold by 2028 (www.theregister.com via hn) MOST POPULAR AI - ai and ml Anthropic says text watermarking scheme relies on inconsequential words 'Shall I compare thee to a summer's afternoon' is the sort of thing this will make, and others look likely to adopt it - AI and ML DeepSeek…
Llama-macOS – Agentic and MCP Native macOS Front End for Llama.cpp (github.com via hn) Llama Llama is a macOS menu bar app for running local LLMs. Watch a 2-minute intro 📽️ Install brew install --cask llama-app Or download from Releases.
The CLI your AI agent drives to manage your knowledge graph (useokf.com via hn) okf is a Go CLI toolkit for the Open Knowledge Format: agentic-first, JSON-native, vendor-neutral. A single binary alternative to Google's Python/Gemini-locked reference implementation.
Autonomous Agentic Engineering Tools (rywalker.com via hn) Key takeaways - The Yegge ecosystem became the category's center of gravity: Gastown hit v1.0 at 15.9K stars with a Kilo-hosted cloud version and Wasteland federation, and Gas City arrived as the composable SDK for building your own orches…
Show HN: Snafu: Agentic flow to help you with "naming things" in source code (github.com via hn) SNAFU: Symbol Name Ambiguity Fixer-Upper Is "naming things" hard? And if so, can one detect and measure name quality?
Show HN: A mobile app to keep up with most important news on Agentic Coding (apps.apple.com via hn) I built this app to help myself kill the FOMO that comes with the fast changing environment of Agentic Coding. It links the latest stories with sources and only what really matters for the field.
Agentic Reasoning for Large Language Models (arxiv.org via hn) Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) demonstrate strong reasoning capabilities in closed-world settings, they struggle in open-ended and…
Show HN: Supervice – process supervisor for agentic processes. zero dependencies (github.com via hn) Supervice A modern, lightweight, and fully async process supervisor for Unix-like systems. Zero dependencies.
Agentic AI drops digits, This library catches them before it runs in fintech (github.com via hn) PrismManifest Zero-trust tool-argument gate for deterministic AI tool execution. (Formerly ParamGate — same design; package prismmanifest, CLI prismmanifest-gate.) PrismManifest sits between probabilistic extractors (LLMs, OCR, table parse…
Show HN: I built a tool that delegates bounded spending to an AI agent (github.com via hn) Molt An open protocol for delegating bounded, autonomous spending authority to an AI agent, at any online store, including the overwhelming majority that expose no agentic commerce protocol at all. The name is the security model.
Show HN: Agentic Ship – open-source Lovable alternative that runs on your agent (github.com via hn) Hi HN. This is an open-source harness that helps you ship a fullstack application the agentic way, you only pay for your AI subscription (Claude, Codex, Cursor) and the domain.
Tech, agentic coding, and whatever else I'm building (dev.profullstack.com via hn) Tech, agentic coding, and whatever else I'm building. By Anthony “chovy” Ettinger.
Show HN: LinkGravity – chat and voice bridge for AI coding agents (github.com via hn) LinkGravity A Discord, Telegram, and Slack bot interface for the Antigravity agentic AI system. It translates Antigravity CLI prompts into chat UI components and provides voice interaction capabilities.
Show HN: Agents that run your finances (blaze.money via hn) Hey HN. My name is Faiyam, cofounder/CEO of Blaze Money (YC S24).
The Evolving Role of the Red Team in the Era of Agentic Security (blog.google via hn) The Evolving Role of the Red Team in the Era of Agentic Security At Google, our Red Teams have always operated on the cutting edge of security. We’ve shared our journey in the past: from the high-stakes operations showcased in our Hacking…
Agentic Scheduling = Apple Calendar + Apple Reminders + Claude Agent (github.com via hn) Agentic Scheduling Your calendar, run by an AI agent you can text. A self-hostable, single-owner scheduling platform.
Ask HN: Small team devs – how do you collab with UI/UX designers in the AI era? (news.ycombinator.com) Running a 2 person product company, my cofounder (product designer by trade) and I are struggling to get the design/dev process right as AI transforms the way we ship. There are a couple of interrelated problems: 1.
Agentic Proof-Oriented Programming (fstar-lang.org via hn) Agentic Proof-Oriented Programming AI-based automation of formal proofs has received a lot of attention in the past few years, with lots of promise but few real successes. However, starting in late 2025, with the availability of models suc…
Agentic AI tests for orchestrators on CLI (github.com via hn) Prism-Eval: Open-Source Unit Testing & Red-Teaming for AI Agents Catch non-deterministic LLM tool call failures, prompt injections, and digit drops in local builds and CI/CD before your users do. Keywords: AI agent testing, LLM red teaming…
Open source alternative to overpriced agentic IDEs like BridgeMind (Demo Video) (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Retrieval vs. Memory in Agentic AI Systems (machinelearningmastery.com via hn) In this article, you will learn the conceptual and practical differences between retrieval and memory in agentic AI systems, and how to combine both effectively. Topics we will cover include: - What separates retrieval from memory, and why…
Show HN: Askance, a human-in-the-loop GUI for agentic coding (github.com via hn) Askance Askance is a simple and beautiful human-in-the-loop GUI for answering questions from an AI coding agent such as Claude Code, either in the browser or on the go with your phone. It works by combining a server which you run locally w…
Best Agent Gateways for Healthcare Organizations 2026 (www.mintmcp.com via hn) Healthcare organizations deploying AI agents face a critical governance challenge. According to Gartner, over 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate ris…
ByteDance Seed: Seed-2.0-Code for Agentic Coding (openrouter.ai via hn) Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude Code, Kilo, and OpenCode.
Agentic Coding: Running the Nightshift (dibranmulder.github.io via hn) Meet Dobby We named our agent Dobby, after the house elf in Harry Potter. It was meant as a joke and turned out to be the most accurate architecture decision we made all year.
"Cheap" Agentic Branches (eschmann.dev via hn) "Cheap" agentic branches After going back and forth, i have settled on a sandboxing approach for agent-assisted coding, a ~700 lines bash util for managing OrbStack containers (or Tart macOS VMs): Branching an agentic sandbox brings along…
WorldClaw – Agentic 3D open-world generation at scale (tencent-hunyuan.github.io via hn) From one open-ended prompt to an explicit, explorable, and editable 3D world.
What agentic harness do you use with Chinese models? (news.ycombinator.com) could not extract summary
AMD Catches the Agentic AI Wave and Will Ride It Up Masterfully (www.nextplatform.com via hn) compute AMD Catches The Agentic AI Wave And Will Ride It Up Masterfully AMD might not be taking any bites out of Nvidia’s market share when it comes to AI systems, but it is capturing its proportional share as the GenAI market expands with…
Show HN: Gitseq: A Repo Becomes a Workroom (github.com via hn) Last week I was trying to figure out how to run a new project at work, where there are lots of technical artifacts, lots of stakeholders upstream and downstream, plenty of documentation for different audiences... all the usual stuff, and a…
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots (cactuscompute.com via hn) Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers.
Agentic coding needs human sign-off tied to physical reality (patrickaudley.com via hn) Agentic coding needs human sign-off tied to physical reality As AI coding agents transition from passive autocomplete tools to autonomous contributors executing entire feature branches, we are racing toward a massive security blind spot: H…
The State of Agentic Memory (medium.com via hn) could not extract summary
Meta's new open-weight model targets local agentic AI (twitter.com via hn) I believe everyone should have access to superintelligence, and I wrote a long piece about Meta's philosophy and values for building a positive future for everyone. meta.com/thefutureisfor… - "everyone should have access to superintelligen…
Agentic Constellations: Cartography of the Hidden (zkm.de via hn) - Exhibition Agentic Constellations: Cartography of the Hidden Daniel Heiss, Marc Schütze Fri, October 23, 2026 – Sun, February 07, 2027 - Location - Foyer of the EnBW Group Headquater - Entrance fee - Free admission With »Agentic Constell…
LLM-oriented programming – statistically revealed components (news.ycombinator.com) Agentic programming reveals mechanical joins of an algebra of de facto components, that resisted capture by formal abstractions. Can we improve formal module abstraction so it can capture this statistically revealed duplication?
The New Spy Race: For Agentic AI and Quantum Computing Threats (www.quantumhorizon.it via hn) Inside the New Spy Race: How the US, China, Russia, Israel and Italy Are Gearing Up for Agentic AI and Quantum Computing Threats By Remo Pulcini — Quantum Horizon Italia Nuclear warheads, submarines, and arms-control treaties defined intel…
Bb Agent IDE (github.com via hn) bb bb is an agentic IDE that can control itself. You can seamlessly orchestrate all of your favorite coding agents together and have them programmatically use bb too.
Loop Engineering in Claude (claude.com via hn) Loop engineering: Getting started with loops Learn how the Claude Code team defines agentic loops, with practical guidance on progressing from turn-based to goal-based, time-based, and proactive loops—and when to use each. Learn how the Cl…
The energy use of agentic AI (www.theclimatebrink.com via hn) The real energy use of agentic AI Agents use about 600x more energy than simple AI prompts AI energy use is a huge and controversial topic at the moment. Credible estimates have AI data centers accounting for around 12% US electricity use…
An Agentic IDE That Builds Itself (www.sawyerhood.com via hn) An Agentic IDE That Builds Itself I’m excited to show something I’ve been working on recently: bb, an agentic IDE that builds itself. This started as a passion project by @_ymichael, and over time I changed from being an early user to a co…
Kitesurf: The Browser for the Agentic Cloud (kitesurf.cloudflare.app via hn) Kitesurf the browser for the Agentic Cloud Kitesurf is Cloudflare’s new stateless, highly scalable and cost-effective web browser that runs entirely on top of Workers and was designed specifically for the Agentic Cloud. Use this playground…
Cloudflare Has Open‑Sourced Cloudflare OS for AI Agents (techstrong.ai via hn) It’s not really an operating system, but an open-source platform for agentic AI workspaces. So, you want an open-source platform to develop and run AI agents?
Open-source agentic satellite anomaly detector with calibrated confidence (github.com via hn) Trustworthy Anomaly Agent on ESA-ADB Anomaly detection on real satellite telemetry that an operator can actually act on: every alarm carries a calibrated confidence, a ranked list of the channels responsible, and a written brief whose ever…
The Corporate Agentic Brain May Be the Next Honey Pot for a Rogue AI (serendb.substack.com via hn) Don't Let Your Corporate Agentic Brain Be The Next Honey Pot For A Rogue AI We're moving so fast to grant AI agents permission to do things we once trusted seasoned, experienced employees to do. In the final week of July, 2026, something h…
Muse Spark 1.2: Improved Agentic Performance at Higher Cost per Task (artificialanalysis.ai via hn) August 5, 2026 Muse Spark 1.2: Improved Agentic Performance at Higher Cost per Task See model pageMeta’s Muse Spark 1.2 scores 54 on the Artificial Analysis Intelligence Index. Its Meta's third release in four months, significantly improvi…
How to build an agent to automate your on-call (12gramsofcarbon.com via hn) Agentics: how to build an agent to automate your on call A walk through of how you might set up an auto-on-call, why these things can fail, and how to think about agentic loops Editor’s note: sign up to attend our next Agentics flagship me…
Show HN: XBin - A self-hosted, sandboxed, self-modifying workspace (xbin.dev via hn) Started with the problem "how do I host and manage 100x more code over the next few years". I wrote 5 iterations of 'Agentic OS' over the past few months, this one seems to finally be the correct shape.
Sober: Local-first code reviewer for agentic PR (deterministic and model review) (sober-dev.app via hn) What ships in v0.8.5-beta This is what state-of-the-art code review looks like in the agentic era: deterministic guardrails that run in milliseconds, model review with evidence you can inspect, and a forge daemon that never auto-merges. Bu…
Replacing static back end endpoints with an autonomous agent (docs.ardor.cloud via hn) Product Capabilities Ardor is a full SDLC platform that can build any software - from simple web apps to complex microservices architectures. While we support all types of applications, we especially excel at building agentic applications…
RNet lets users use one AI credit balance across multiple apps [demo] (news.ycombinator.com) demo video : https://youtu.be/W7U3HdI37N0 I built rNet to let users use their AI credits across multiple apps. Note: by "users" I mean not developers but end users, normal people.
Ask HN: How can I improve my products? Which one to keep working on? (news.ycombinator.com) I am looking for feedback on my products and I am interested in knowing how to find out which one to keep working on? Here is the list - https://www.blanketstories.com/ Bedtime stories.
LFM2.5-2.6B: On-Device Agents (docs.liquid.ai via hn) Specifications Agentic & Tool Use Native tool calling, trained inside real agent harnesses 128K Context Long context for tool traces and multi-step workflows On-device Small enough to run on a laptop or phone Quick Start - Transformers - l…
AAFlow: Scalable Patterns for Agentic AI Workflows (arxiv.org via hn) Agentic workflows in large language model systems integrate retrieval, reasoning, and memory, but existing frameworks suffer from scalability and reproducibility limitations due to fragmented data orchestration, serialization overhead, and…
Agentic Minimalism: The Human Control Loop (leverageloops.substack.com via hn) Agentic Minimalism: The Human Control Loop your agents need a control loop and so do you I have this habit. I’ll be working on something, and a question will cross my mind.
Show HN: Simple self-hosted LLM assistant with user-steered compounding context (github.com via hn) I built a personal LLM assistant on Cloudflare Workers + Durable Objects. You specify a category and topic when starting a new conversation, so the backend maintains a summary for each category/topic - building up as more conversations hap…
Show HN: Loopers – Fail-closed reverse proxy and circuit breaker for AI agents (app.tryloopers.com via hn) A baremetal, zero-delay firewall and circuit breaker for the Agentic Era.
Auditability vs. Forced Determinism: Future of Agentic AI (medium.com via hn) could not extract summary
Dev tools must be opensourced thats why I build this (news.ycombinator.com) My framework will help you develop your agentic/AI systems, for now its only in beta. I already used it for my applications and it helped me but it is still not perfect thats why i still push updates and fixes.
Orchard: An open framework for scalable agentic AI (www.microsoft.com via hn) Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infra…
Bet on Primitives for Agentic Coding (www.robinwieruch.de via hn) learned D3 properly years ago, and precisely because I know how much work hand-rolled charts are, I would never have budgeted them for a client. At least, that was the math before agentic coding.
Show HN: Driftty – mobile focused web TTY (github.com via hn) This one is VERY early, but I've been having a blast building and using it. In the past I've use Termius (which is great), but I've never been thrilled about using it with coding agent CLI's.
A Markdown wiki outscored every AI agent memory product we benchmarked (verginglabs.com via hn) See how much smarter your AI could be by adding tools We independently measure the tools AI agents use and publish what each one actually adds. Agentic Memory Indexi MethodologyAgentic Memory Index v0.1 · Aug 2026 Highlights Accuracy, cost…
Codex CLI Has Responsive Tables in the Terminal (www.vincentschmalbach.com via hn) Google Lighthouse Adds Agentic Browsing Checks and Cloudflare Adds AI Traffic Controls Google Lighthouse and Cloudflare are adding different controls for the same emerging class of automated web visitors. Lighthouse added an experimental A…
Show HN: Minimal, Composable, agentic SDD framework based on superpowers (github.com via hn) smolpowers A lightweight, evidence-driven workflow for coding agents. Design → Plan → Execute → Finish Why smolpowers Smolpowers is a lightweight SDD framework based on Superpowers.
Show HN: Rails Agent – Build Autonomous AI Agents Natively in Ruby on Rails (rails-agent.com via hn) Rails Agent - The Fullstack Agentic Development Platform For Ruby on Rails. Build, Test, Deploy & Monitor in one single dashboard.
CRM: An open-source, agentic-first CRM (github.com via hn) CRM An open-source, agentic-first CRM. A durable research agent is the product.
Ask HN: When did we go from agentic loops to graphs? (news.ycombinator.com) The AI engineering conversation seems to be shifting daily: from prompts, loops, to now graphs. I’m curious whether there is genuine progress happening in the field that has given me so much or whether folks are creating terminology becaus…
The Kotlin Benchmark for AI Coding Agents (blog.jetbrains.com via hn) Kotlin A concise multiplatform language developed by JetBrains Introducing the Kotlin Benchmark for AI Coding Agents Agentic coding benchmarks are getting closer to real-world software development. For Kotlin teams, the most important ques…
The World First AI Agentic Radio, for vibe coders, good vibes only (www.twitch.tv via hn) Vibe Coder Radio + Lofi Beats + AI Experiments · 24/7 | Streaming music for 1 viewers.
Show HN: Agent Architect Skill for building agentic systems (github.com via hn) I built a skill for building agents that helped me ship agentic products within days that otherwise took several weeks and broke often in production. It is inspired by gstack.
NeverWrite, the ultimate agentic Markdown workspace (github.com via hn) NeverWrite A local-first workspace for writing and thinking with AI agents. Keep your vault, documents, agents, review, and research in the same place.
Hunk – Review-first terminal diff viewer for agentic coders (www.hunk.dev via hn) Hunk is a review-first terminal diff viewer for agent-authored changesets. Multi-file review stream, inline AI annotations, watch mode, and Git/Jujutsu integration.
CircuitProver: Agentic Lean 4 Theorem Proving for Hardware Verification (arxiv.org via hn) Modern integrated circuits (ICs) are becoming increasingly complex, making functional verification a major bottleneck. The dominant hardware formal verification methodology, model checking, verifies each design instance separately and expo…
Show HN: Claude MIDI Twister – An agent visualizer for a DJ MIDI controller (www.dylanfisher.com via hn) I was inspired by the release of the Codex Micro controller (https://worklouder.cc/codex-micro) and repurposed a DJ TechTools Midi Fighter Twister I had laying around into a hardware visualizer for my Claude Code sessions. This is a little…
$0.26 DeepSeek V4 Flash 0731 ties $5.01 GPT-5.6 run on Agentic Memory Benchmark (atmbench.github.io via hn) Send a Pull Request Fastest path: open a PR adding a row to the TRACKS array at the bottom of leaderboard.html . Include a short description of the setup and a link to your run logs or code in the PR body.
Nothing Is Easy When You're an LLM: The Flat Latency Problem (aidoses.substack.com via hn) Nothing Is Easy When You're an LLM: The Flat Latency Problem Dose #9 — Production Agentic AI Under Pressure When you walk to the kitchen, your brain barely registers it. When you solve a differential equation, it burns.
Microsoft is about to kill my favorite Edge feature – and Copilot is to blame (www.windowscentral.com via hn) Microsoft is about to kill my favorite Edge feature — and Copilot is to blame Edge's Sidebar is on the chopping block, and Copilot seems to be the one pushing the blade. As Microsoft pushes Windows 11 to become an agentic OS full of AI too…
Show HN: Yet Another Markup Language Engineering Toolkit (github.com via hn) Over the past few weeks I've grown tired of these super, all in one, agentic AI Skill plugins, MCP servers and whatnot that promise you the world. Of course, there is a world of vibes that survives off that, but there is also a world of mo…
Founding Member for Prodigy (news.ycombinator.com) I'm building an autonomous agentic workforce that works alongside your existing team. Our agentic workspace allows human and AI employees to work together [https://www.prodigy.org.in] Hiring has always been human centric.
Agentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1 (kaitchup.substack.com via hn) Agentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1 The Weekly Kaitchup #152 Hi everyone, In this edition of The Weekly Kaitchup, I’m looking at two of the week’s most interesting releases for agentic workloads: Nanbeige4.2-…
Show HN: Agentic – Run Claude Code on different models with routing and budgets (github.com via hn) agentic Run Claude Code on any model, with a budget. agentic wraps Claude Code in a thin local router.
Show HN: Nurb – Agentic CAD for 3D printing (github.com via hn) I've had a love/hate relationship with Fusion over the years. I can model stuff there but it's just always slow and painful and Fusion has a deep desire to crash on me.
How to build good software with untrustworthy agents (www.treygoff.com via hn) How to build good software with untrustworthy agents A complete, working agentic development stack, explained from first principles and written for the layman and the software engineer both. ~/Code/trey-goff $ claude "build the page that e…
Infrastructure Patterns for Agentic Applications (render.com via hn) Most teams start building agents the same way they build any other web feature: wrap the model in a route handler, parse the request, and wait for the response. This is fine for a demo.
Scaling Agentic RL: 365,000 Environments for SWE, Terminal, and Search (www.primeintellect.ai via hn) Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Search Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Search The open research ecosystem has produced many great datasets for the three main agentic domains…
Appliedin: The agentic workflow for applying jobs, so we can spend time prepping (www.appliedin.dev via hn) You decide where the graph pauses. Use gated mode to approve each application, or allow strong matches to continue.
Preventing Data-Purpose Laundering by Agentic AI (futurium.ec.europa.eu via hn) The Problem Space for Europe European citizens entrust their personal data to organisations under a clear legal obligation: it must be collected for specified, explicit and legitimate purposes and must not subsequently be processed in a ma…
MonkeysCode and Capuchin AI: The Agentic IDE That Gives You Total Control (monkeyscode.com via hn) Skip to content Beta Features Capuchin Pricing Docs Blog Free Trial More Models Trust Self-Host Roadmap Ecosystem Support Sign in Download Features Capuchin Pricing Docs Blog Free Trial More Models Trust Self-Host Roadmap Ecosystem Support…
Our Approach to Responsible Agentic Development in Open Source (www.puppet.com via hn) could not extract summary
Show HN: Dn – plan collaboratively, let agents execute (github.com via hn) Hi HN — we built `dn` because coding got faster, but building software did not. Use the CLI to clear your backlog faster with reusable agentic workflows & other supporting commands.
The Power of Focus in Agentic Coding (tenzinwangdhen.com via hn) Alex Graveley, the guy behind GitHub Copilot, showed me how I’d been thinking about agentic coding wrong. He used pen and paper.
Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide (rocm.blogs.amd.com via hn) Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide# Agentic AI applications are pushing LLM inference into a new regime. A request can now reach one million tokens from aggregated long prompts, tool calls, retrieval resu…
AgentENV: A platform for running agent env at scale (Kimi K3 RL) (github.com via hn) .. raw:: html Running agent environments at scale |coverage-status| |coverage-report| 📖 Full documentation _ AgentENV (AENV) is a platform for running agent environments at scale, powering agentic RL training for Kimi K3.
What's Next in Agentic Operations: Introducing AI Week at Grafana Labs (grafana.com via hn) Explore what's next in agentic operations: Introducing AI Week Observability has traditionally been tacked on after your code hits production, but with agentic operations on the rise, that's no longer sustainable. Agents have dramatically…
How I built a debugging tool, and the agent, using it, found bugs in it (news.ycombinator.com) This story might never have happened, even though I'd been carrying the idea for a debugger for a long time. There was no trigger.
Show HN: Maginary.ai gets seedance2 and GPT-images2 support (twitter.com via hn) presenting maginary -- an image/video generator with a midjourney-like prompt syntax, using 40+ underlying models and acts like an abstractized openrouter for multimedia with an amazing ux, can try it right now with no cc this launch prese…
Open Knowledge format v0.2 tackles agentic trust (cloud.google.com via hn) Open Knowledge format v0.2 tackles agentic trust Sam McVeety Tech Lead, Data Analytics, Engineering, Data Cloud Amir Hormati Tech Lead, BigQuery, Engineering, Data Cloud When we introduced the Open Knowledge Format (OKF) in June 2026, we a…
Measuring the Post-Merge Fate of Agentic Code (post-merge-reality.github.io via hn) First longitudinal empirical study of the post-merge lifecycle of agentic and human code in real-world open source projects. 182public GitHub repositories 30K+agentic commits 121M+lifecycle lines tracked 12maintenance operational intents 1…
Ask HN: How would you harden AI changes to a 1M-line legacy SaaS before review? (news.ycombinator.com) I’m not a software engineer, but I’ve been running an experiment to see whether agentic development could produce a useful prototype on top of an existing SaaS codebase. The codebase is 1M+ lines, 15 years old, hosted on Azure, and primari…
Agentic AI Workflows and the Future of Streaming Ad Tech (www.streamingmedia.com via hn) Agentic AI Workflows and the Future of Streaming Ad Tech AI is bringing a new coordination layer to ad tech that is impacting every area of the streaming monetization workflow, including executing transactions, preparing inventory, underst…
Fountible: Agentic Figma (fountible.com via hn) 01 · Structure Auto layout becomes flex Frames and groups keep their hierarchy; sections become frame containers. Direction, wrap, gap, padding, alignment, sizing, constraints and relative placement map onto native structure.
How to Regain the Lost Focus? (news.ycombinator.com) Have you experienced it with agentic coding? The diminishing appeal to focus and deep work?
Grok 4.5 dropped cache price per 1M tokens from 0.50 to 0.30 (twitter.com via hn) Grok 4.5 has just become a lot cheaper without anybody noticing or publishing anything, and it’s now the best “balanced” model to use for agentic tasks according to our internal benchmarks. Grok launched a few days ago with a price for inp…
Audit: Skill library for agentic app audit (github.com via hn) Audit Skills A library of focused audit skills for AI coding agents. Each skill is a small, self-contained playbook that an agent can run against a real codebase and report concrete findings with file paths, line numbers, impact, and sugge…
Execution-Free Agentic Program Repair for Enterprise-Scale Development [pdf] (dl.acm.org via hn) could not extract summary
WorkBuddy Bench: Agentic Coding Leaderboard (workbuddybench.com via hn) Key Takeaways核心结论 - No single model dominates — column leadership splits across subsets and harnesses: Claude Opus 4.8 leads five of the eight scored columns (Code on both harnesses 74.43 / 77.90, Web on both harnesses 68.14 / 69.86, Offic…
Show HN: Skim – a minimalist open-source email client for Windows (MIT) (skim-tech.com via hn) I built Skim because I was super dissatisfied with all existing Windows clients. I mean, it's not a rocket science, there should be some lightweight modern solution, isn't it?
Are Agent Code Reviews for AI-Generated Code Worth It? (mattmccormick.ca via hn) As I wrote previously, I’m exploring building an agentic AI pipeline that mimics the full Software Development Lifecycle (SDLC). I’m still doing a lot of iteration as I run into issues.
Agentic coding makes strict static analysis non-negotiable (pocketarc.com via hn) A maximum-strictness PHPStan config with every strictness package that makes AI-generated Laravel code drastically better.
MCP-spec-audit, scan your MCP server for the July 28 spec changes (github.com via hn) The Agentic Toolkit A growing, curated collection of small, practical tools and agent skills for agentic engineering: the craft of shipping real, production software with AI coding agents (Claude Code, Cursor, and similar). Agentic enginee…
Show HN: AI agents that go from naming your startup to running its marketing (www.brandbrahma.com via hn) Solo founder from Bengaluru built an agentic ai platform for business naming to marketing.
GPT-5.6 got smarter. Then it kept acting (toloka.ai via hn) From agentic skills to coding and AI safety — we build data solutions integrating human expertise and state-of-the-art automation to accelerate AI development.
Ask HN: Where is NFTs as a technology or investment? (news.ycombinator.com) As I’m working through reports, I realized that with AI-everything taking all the air in the room, I hadn’t heard about tokens of the non-fungible kind for a hot second. As an artist, I was excited about the implications of getting paid fo…
How Antithesis Turned exe into a Sandbox for Agentic Software Tests (blog.exe.dev via hn) Carl Sverre spends a lot of time thinking about how to give AI agents the right amount of power. Give them too little, and they can’t do real work.
Ask HN: Agentic Coding for Messy Multirepos? (news.ycombinator.com) My job has quite an annoying "dev" environment setup, with a slew of scripts, a shared library that works as interface for different services. the shared library and all services live in separate repos.
Show HN: add-reasoning-to-prs: a hook that adds unnoticed assumptions to PR desc (github.com via hn) I think when you want to truly understand how something works, you either reverse-engineer the thing, or you participate in its making, or at least, understand the reasoning behind the build. Reading the documentation is not enough – you n…
Retaining Authorship for Critical Code: An Agentic Buddy System (smorsic.io via hn) Retaining Authorship for Critical Code: An Agentic Buddy System for Review and QA Something that's always been true about software is that every project has different needs. A lot of debate about how we should be engineering suffers from t…
Show HN: MindCache – An open-source agentic memory system for LLMs (github.com via hn) 🧠 MindCache Structured Long-Term Memory SDK for LLM Agents MindCache is a structured long-term memory engine designed for production-grade LLM agents. Unlike flat vector search databases or basic summary stores, MindCache maintains a self-…
Agentic Payments Landscape (payments.agents.sh via hn) A continuously-verified reference on agent-initiated payments: 32 protocols across identity, authorization, checkout and settlement, each with its sources, confidence level and last-checked date.
The OWASP Agentic AI Top: What Builders on AWS Need to Know (blog.technodrone.cloud via hn) Writing The OWASP Agentic AI Top 10: What Builders on AWS Need to Know On this page This is post #0 of the OWASP Agentic AI Top 10: What Builders on AWS Need to Know series. You’ve secured your GenAI chatbot.
Codekeel: A context drift governance tool for Claude Code (github.com via hn) codekeel Drift governance for agentic coding. codekeel keeps a decision ledger for your project — architectural choices, conventions, and constraints — and enforces it live during Claude Code sessions, so every new session inherits, respec…
Agentic-sage – a local board for parallel coding sessions (no auto-merge) (github.com via hn) Install · How it works · Full setup guide · Adapters · Changelog =20"> Session Awareness & Guidance Engine: a passive, read-only fleet judge for running many parallel agent coding sessions (Claude Code, Grok Build CLI, etc.). It does no wo…
Zero risk isn't the job: a CISO's guide to agentic AI (claude.com via hn) Zero risk isn't the job: a CISO's guide to agentic AI Anthropic's Deputy CISO, Jason Clinton, shares his team's lessons learned adopting agentic AI, and the risk assessment framework they've developed for building and deploying agents secu…
Cursor's "New Model Economics" (replicated.live via hn) Yesterday, I posted my piece on the Darwinist revision control system for codebase evolution. Surprisingly, Cursor folks posted "Agent swarms and the new model economics" which touches upon agentic swarms, agentic revision control and the…
Was Hugging Face Breached by AI Agents? (mrkt30.com via hn) Hugging Face confirmed a security breach this week carried out by an autonomous AI agent. We look at the details, what an 'agentic attacker' is, and what it means for the future of AI security.
The depth problem with agentic research (moekhalil.substack.com via hn) The depth problem with agentic research or "Budget is a Research Primitive" I started working on AI research back in 2023, when Perplexity was still a seed-stage startup. Over time, I’ve seen it get better in two ways: more accurate, and d…
Show HN: Memento, Shared Agentic Memory (rcarmo.github.io via hn) memento stores knowledge that should outlive a conversation and remain available to several agents: people, projects, machines, services, decisions and their relationships. Chats, reminders, credentials and machine-specific state stay with…
AI Red Teaming: Securing Agentic AI Systems (video) (alice.io via hn) It Takes AI to Break AI: The Case for AI Red Teaming Discover why AI-powered adversarial testing is becoming essential as AI systems reason, use tools, access data, and act autonomously. Overview Organizations are racing to deploy AI copil…
Agentic World Models (cameronrwolfe.substack.com via hn) Agentic World Models Creating better language agents by teaching them to model their environment... Reinforcement learning (RL) has played a key role in the expansion of capabilities for large language model (LLM) agents.
Better Agent Leagues and Decision Supervision from Branched Rollouts (gertlabs.com via hn) While scoring decisions using MCTS-inspired branches to reward long-horizon agentic coding progress, we simultaneously build a stronger league for multi-agent competition.
Do you manage task/todo differently in this AI era? (news.ycombinator.com) I'm curious. Do you use todo list app?
We run 100% agentic coding at a €2M ARR healthcare SaaS (agenticprime.ai via hn) How we run 100% agentic coding at a €2M ARR healthcare SaaS The operating model behind appointmed's agentic coding workflow: how we shape work, supervise agents, review code, and keep release risk owned by humans. The most common misunders…
Agentic AI Is Taking over Execution, Not Just Content Generation (themarketingnewsletter.org via hn) #45 - Agentic AI Is Taking Over Execution, Not Just Content Generation What happens when marketing stops being a series of tasks and starts being a series of goals A year ago, if you asked a marketer what AI meant for their job, they’d pro…
Ephemeral runtime harness for agentic workflows open source (github.com via hn) drun (deterministic run) Git for agents with ephemeral runtime Drun is a platform that allows you to virtualize components of your host into an ephemeral runtime to serve as the agent's workspace with git-like primitives which allow the ag…
Perplexity Unveils Space, a Secure Sandbox Platform for AI Agents (techstrong.ai via hn) Perplexity has introduced SPACE, a new sandbox platform engineered to safely unlock the full capabilities of Perplexity Computer’s agentic AI stack while maintaining a high bar for security. Mea culpa.
Run AI Agents from Jira, Linear, GitHub Issues, or Markdown (github.com via hn) Startup Factory Turn your product board into a governed software delivery system. Startup Factory is an agentic orchestration framework for end-to-end product delivery.
How OpenAI's Sol Learned Design Taste (notes.designarena.ai via hn) We benchmarked GPT-5.6 Sol on Design Arena’s Web Design (Non-Agentic) Arena, and we were surprised to find that it ranks 1st overall. This is 18 places higher than its predecessor GPT-5.5, and is the first time an OpenAI model has placed f…
End of Line: Replacing MCP with Standard Filesystem Tools (www.sicpers.info via hn) Over on YouTube I shared a video where I demonstrate a more context-efficient approach for integrating existing systems into agentic tools than MCP, one that comes with built in access control mechanisms and that LLMs are already trained t…
From Modern Data Stack to Agentic Data Stack (seatunnel.apache.org via hn) Over the past decade, the Modern Data Stack has changed how teams build data platforms. Data ingestion, lakehouses or cloud warehouses, dbt, orchestration, metrics layers, BI, and governance were split into modular building blocks that can…
Agentic Misalignment in Summer 2026 (alignment.anthropic.com via hn) Case studies of frontier models sabotaging code, assisting fraud, mislabeling, and coaching whistleblowers. Last year, we reported observations of agentic misalignment in models from across the AI industry (including Anthropic’s Claude mod…
IBM Power S1112 Brings Local AI Inference to the Edge as Power Goes Autonomous (www.storagereview.com via hn) IBM has expanded its Power server lineup with new software to automate infrastructure management and application development. The announcements include IBM Power Autonomous Operations, an agentic control layer for system management, and th…
How Agentic Is Agentic Commerce? Measuring X402 Adoption and Authenticity (arxiv.org via hn) AI agents are said to be forming an economy in which they pay, on their own, for the data, APIs, and compute they consume. x402, which settles a stablecoin payment on-chain for each purchase, is the most widely deployed protocol for this,…
Ask HN: What Are You Building with AI? (news.ycombinator.com) Are you building with AI for yourself, your employer, customers, or as a startup? What LLMs, coding agents, or development tools are you using?
Cicada- an agentic Python IDE Free to use ( comes with built in small model) (github.com via hn) Cicada An agentic Python IDE that turns plain-English requests into runnable, executed code — powered entirely by a local model. local-llm · agentic-ide · electron · python · llama-cpp · gguf · code-generation · monaco-editor · machine-lea…
Show HN: A agentic nervous system for all DevOps tools (github.com via hn) I have created a agentic prod debugging tool which connects all the existing tools to one single workflow. Each connector is a subagent which index the data and a main agent answers the queries like why the service is down using the knowle…
Show HN: Skills as a Service via MCP –> Coding Agent Skill Library (github.com via hn) Hi HN!, While discussing agentic engineering with my close network, it came to light that most agent skills still live on desktops, in repositories, or in people’s heads. In the context of large, regulated organisations, where compliance a…
Why Claude uses the browser like a drunk intern, and how to fix it (pluno.ai via hn) How we accidentally built a browser agent We spent the last few years building AI agents for customer support. The first version was basically agentic search over support context: docs, tickets, prior resolutions, internal notes.
Harvey LAB-AA: evaluating AI agents on real-world legal work (artificialanalysis.ai via hn) Announcing Harvey LAB-AA: evaluating AI agents on real-world legal work Harvey LAB-AA (Legal Agent Benchmark) is our implementation of Harvey's new agentic legal benchmark, evaluating language models on real-world legal work across 24 prac…
The Six-Layer Memory Pipeline Behind Our Local-First Agentic Memory in 2026 (medium.com via hn) could not extract summary
Kotro – I cut my Cursor API bill by 68% with a 15MB local proxy (github.com via hn) Kotro Proxy Engine The local security and efficiency layer for MCP-native agentic AI — intercept streaming LLM traffic from OpenAI and Anthropic SDKs, block prompt injection from tool results, keep secrets off the wire, and cut token waste…
Toward Accountable Execution in Agentic Software Engineering (decapodlabs.github.io via hn) Abstract Large language model (LLM) agents are transitioning from conversational programming assistants to autonomous executors capable of editing codebases, running build tools, executing tests, and managing multi-file software deployment…
Dictation for agentic coding, now with an MCP and CLI (twitter.com via hn) Have you tried the Superwhisper CLI yet? Check out this use case: pulling new vocab out of your own dictations ✨ https://t.co/6GdWGXrX6E
Agentic development is standardizing faster than its operating model (blog.codacy.com via hn) Agentic Development Is Standardizing Faster Than Its Operating Model Right now, an agent is probably opening pull requests against your production code. You may not be able to say which agent, who configured it, which instruction file shap…
CortexWP AI: WordPress Agentic system for debugging and development (cortexwp.ai via hn) Most AI tools can talk about WordPress. They can explain what a plugin conflict is, suggest code, tell you to check error logs, disable plugins, clear cache, or contact your hosting provider.
Why Cursor Is the Most Practical Choice for Beginners in Agentic Coding (calcrecipe.com via hn) Why Cursor Is the Most Practical Choice for Beginners in Agentic Coding Vault Track: #6 | Sealed on 2026-07-12 💡 Before diving in: This guide is not intended for professional developers who are already deeply seasoned in coding. Instead, i…
Clodex – Open-source agentic IDE with governed execution and graph memory (github.com via hn) Clodex Local-first agentic IDE with governed execution Clodex is an open-source agentic development environment that combines persistent AI tasks, code, terminal, browser, Git, models, memory, and controlled execution in one Electron works…
Bitemporal provenance in agent memory: What did we believe, when, and why (news.ycombinator.com) CozoDB, a transactional relational-graph-vector database with embedded Datalog in Rust, went dormant in December 2024. We hard-forked it as MnesticDB (not official CozoDB), under an MPL-2.0 license, to continue Ziyang Hu and the Cozo Proje…
Show HN: Standalone SearXNG CLI+MCP (no server needed) (github.com via hn) Hi HN, Codex and Claude are pretty good at (re)searching things on the web these days, but the open coding agents (OpenCode, pi coding agent and friends) don't have access to the labs' proprietary search APIs. I wasn't happy with this stat…
MCPA: The First Official Certification for the Model Context Protocol (aaif.io via hn) The Model Context Protocol (MCP) has quietly become one of the most important standards in agentic AI. In a little over a year, it went from a promising idea to the default way AI applications connect to external tools, data sources, and s…
Academics building open‑source agents for academic and research work (www.agents4academia.org via hn) Learn together · Build together Agents4Academia A community of researchers building and understanding agentic tools for academic work. AI agents are changing what individuals can build: a recurring frustration or a highly specific workflow…
Show HN: Bunrun – agent-configured local dashboard to start/stop dev apps (github.com via hn) Agentic era has left me with already half a dozen vibe coded helper apps, most run with ´bun run dev´ or the npm equivalents, occasionally Python. Instead of zooming around terminal, I decided to vibe code one more thing: A local dashboard…
Show HN: SubjectiveZero, an Open-Source Agentic Node Editor for Creative Coding (sxp.studio via hn) Hey there, My name is Clem, I've been a solo indie dev for a couple years now, exploring frontier tech like XR and agentic workflows in the context of creative / interactive work. I've been building creation tools for a while and some comm…
Meta agentic model Muse Spark 1.1 (twitter.com via hn) (1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
Ask HN: Why so few consumer AI companies? (news.ycombinator.com) I feel like everyone in SF is doing agentic b2b saas. For ai consumer founders-whats been your experience building?
Automating Risk Model Retrain Loop with Agentic Skills (www.coinbase.com via hn) could not extract summary
Vulnify: Giving Your Agents a CVE Brain (trustedsec.com via hn) Vulnify: Giving Your Agents a CVE Brain Table of contents When building agentic components for pentesting, CVEs are inevitably going to come up. I never found anything that matched what I needed; I did not want "search the web and hope," b…
Show HN: Agentic FC – a football management SIM played by AI agents over MCP (github.com via hn) Agentic FC Agentic FC is a football management simulation designed to be played by AI agents through MCP and watched by humans through a terminal console. Instead of clicking through menus, an agent shapes the Mindset of an autonomous in-g…
Agentic Coding Arena – Compare OpenAI, Anthropic, and Other Models (arena.logic.inc via hn) Compare how frontier AI models create small apps. 52 apps, each built by 21 different models from the same prompt.
Ask HK: Does it feel like the top models routinely dismiss your opinion/advice? (news.ycombinator.com) I find that the latest SOTA modals blatantly ignore some of my instructions or requests, sometimes even after mentioning or requesting it multiple times? I suspect its a consequence of the creators of these's drive for autonomous and compl…
Build your own agentic framework – the no-magic version (medium.com via hn) could not extract summary
The computer is being reinvented in the agentic era (twitter.com via hn) The computer is being reinvented in the agentic era: - The model is the new CPU. - The harness is the new OS.
Talon: Self-hosted AI agent harness for chat, terminal, and desktop (github.com via hn) Talon Multi-platform agentic AI harness. Runs on Telegram, Discord, Microsoft Teams, the Terminal, and a cross-platform Desktop/Mobile companion app (Flutter), with a pluggable backend (Claude Agent SDK, Kilo, OpenCode, Codex, or OpenAI Ag…
Cache hit rate dropping by 20% doubles your agent's bills (dirac.run via hn) TL;DR - • We posit that in the typical long agentic loops, Cache hit rate and Cached input pricing are the most dominant yet ignored cost factors. - • The agent loop simulation below lets you estimate costs by tweaking different parameters.
Show HN: The first agentic coding engine that hot-reloads the full stack (serverpod.dev via hn) An open source solution for hot reloading your full stack: server, database, website, and app.
A launch playbook your coding agent can run – it launched itself (github.com via hn) Agentic Product Launch A launch playbook your coding agent can run. A launch system for agentic builders and anyone shipping products with AI — indie SaaS, apps, AI products, agents, and developer tools — when the builder has little or no…
Show HN: An agentic CRM, built for AI agents to drive over plain HTTP (github.com via hn) Hi HN, Earlier this year we decided to replace a number of our internal systems with agent-first alternatives. We couldn't find anything that worked well as a CRM.
Conversing with antiquity: Agentic AI partner for expanding historical research (deepmind.google via hn) Thea Sommerschield, Durham University Zoi Tsangalidou and Yannis Assael, Google DeepMind Ancient inscriptions offer a direct window into the human past. As invaluable historical sources, they preserve everything from imperial decrees to ev…
Show HN: Pairl v1.5 – dual channel token compression for agentic workflows (pairl.dev via hn) Open protocol, managed implementation. Replace plugin stacks, cut token spend, and monitor savings with transparent gateway headers.
AI Agent using a Burp-style toolkit over MCP (github.com via hn) mulot [-4285F4?logo=googlechrome&logoColor=white)]() Agentic AI web pentester that drives a browser. An open-weights LLM (GLM-5.2, Gemma or Qwen) drives a real headless Chromium through a Burp-style toolkit and works a target the way a hum…
IDE with agentic support built using Flutter (lumide.dev via hn) Lumide is a high-performance, lightweight code editor built for the agentic age. Host autonomous agents natively via the Agent Client Protocol with zero Electron overhead.
Elastic's Agentic SOC (www.elastic.co via hn) This is Part 1 of the Inside Elastic InfoSec's Agentic SOC series. Part 2: choosing the right agent architecture for a 5× cost reduction Elastic's InfoSec team built an agentic SOC that triages every alert before an analyst opens it.
Agentic Autonomy Levels (addyo.substack.com via hn) Agentic Autonomy Levels A working model of autonomy for agentic engineering In most conversations about agentic engineering, the action has changed from prompting to operating. Here’s a frontier looking into the fog: software factories, go…
Claude Fable 5 Backlash Grows (tech.yahoo.com via hn) Anthropic's Claude Fable 5 faces growing backlash after its July 1 re-release. Users claim stricter guardrails have crippled the flagship model's coding, debugging, and agentic performance.
What Is Project Aion? Inside Microsoft's Agentic Copilot OS (www.windowscentral.com via hn) What is Project Aion? Inside Microsoft's secret agentic Copilot OS incubation project that runs on Windows and Android Project Aion is a 2024 incubation effort designed to build out a functioning Copilot OS experience, capable of running o…
Agentic BI Team: Orchestrate a BI Team from CLI (github.com via hn) 📊 Agentic BI Team for Claude Code A complete Business Intelligence & Data Analytics team, built from Claude sub-agents and skills. Fill in one plain-English charter, run one command, and get a virtual BI function that builds pipelines, mod…
Show HN: DNSskills.md – agentic skills for certain DNS utilities (dnsskills.md via hn) DNS skills for agents Machine-readable contracts and human-readable notes for DomainHelp DNS and domain utilities. Entry Points - Documentation base https://dnsskills.md - Execution base https://app.domainhelp.com - Catalog JSON https://ap…
The first agentic reverse engineer (github.com via hn) Kong: The Agentic Reverse Engineer LLM orchestration for reverse engineering binaries What is Kong? Most tasks follow a linear relationship: the more difficult a task, the longer it usually takes.
Sonnet 5 is 2.5x cheaper than Opus 4.8 and 6 points behind on SWE-bench Pro (spark.temrel.com via hn) Claude Sonnet 5 is 2.5x cheaper than Opus 4.8 and nearly matches it on agentic coding benchmarks. A practical four-axis heuristic (scope, novelty, risk, iteration) for routing each task to the right model tier, a worked example, and a free…
Show HN: Loopers – Open-source fail-closed firewall for AI agent runtimes (github.com via hn) Loopers – The firewall for the agentic era Break the loop before it breaks your budget. AI Agent or LLM?
Show HN: Kiwi – Run agentic dev loops in the cloud, keep keys on your laptop (github.com via hn) Kiwi: The Secure Agentic Control Plane Kiwi is an enterprise-ready secure execution engine that runs autonomous developer loops (TDD Actor-Critic alignment) in the cloud while pulling required secrets dynamically and securely from the loca…
Ask HN: How Do You "Not Write Any Code by Hand" with a Token Budget? (news.ycombinator.com) With some corporate environments going from "tokenmaxxing at all costs" to now setting strict token budgets for agentic development: How is one supposed to adhere to a managerial / leadership instruction of "use AI for everything" but even…
GCP offers agentic perimeter guardrails (cloud.google.com via hn) Securing agentic AI with perimeter guardrails: What's new in VPC Service Controls Pratik Bhangale Product Manager, Google Cloud As enterprises scale autonomous AI agents into production, enabling safe innovation requires robust architectur…
Surus Agentic Postgres Companion (github.com via hn) Surus Agentic Postgres companion — with the right guardrails to safely point LLMs at your prod DB. Features · Quickstart · Desktop app Design Surus is a Postgres companion client with a built-in agent.
Microsoft Copilot OS revealed in LEAKED video: built on Copilot and agentic AI (www.windowscentral.com via hn) Microsoft Copilot OS revealed in LEAKED video: Lightweight Windows OS exploration features new desktop UI built entirely around Copilot and agentic AI A leaked video from 2024 has revealed all about Microsoft's internal explorations for a…
$85,000 in tokens later: What I learned from scaling agentic coding at Lovable (lovable.dev via hn) How a year of unrestricted token spend reshaped my development process — from plan-mode PRs to agent swarms shipping 150+ PRs a week.
Codifying the Rules: Building the Platform Behind the Agentic SDLC (blog.owulveryck.info via hn) This article explores how organizations can scale reliable, AI-driven software development by combining modern Platform Engineering with the Team Topologies framework. It introduces a new Software Delivery Lifecycle (SDLC) where stream-ali…
Syscall: Ring ZERO assembly puzzle game for those who are tired of agentic AI (store.steampowered.com via hn) Install Steam sign in | language 简体中文 (Simplified Chinese) 繁體中文 (Traditional Chinese) 日本語 (Japanese) 한국어 (Korean) ไทย (Thai) Bahasa Indonesia (Indonesian) Bahasa Melayu (Malay) BETA Български (Bulgarian) Čeština (Czech) Dansk (Danish) Deut…
Golden Paths Weren't Built for Agents (www.massdriver.cloud via hn) Kelsey Hightower on Platforms in the Agentic Era and the Launch of Massdriver v2 Watch On-DemandYour Golden Paths Weren't Built for Agents - Part 1: How Agentic Development Breaks Control Identity-based permissions break self-service. We t…
Show HN: Humbug: GUI-based agentic dev platform built with only 3 dependencies (github.com via hn) Humbug: building an operating system for human-AI collaboration Humbug is a modular, extensible platform that aims to let you and your AIs work on ideas together. Think of it as an operating system for human-AI collaboration.
Show HN: MothRAG - Graph-free multi-hop RAG without the rebuild bill (github.com via hn) MothRag Deterministic, agentic-style multi-hop — research-SOTA parity without the graph you'd rebuild every day. On commodity LLM APIs alone.
Meta: New Muse Spark update, and an Opus level Muse variant, are both on the way (twitter.com via hn) Alexandr Wang, head of the META Superintelligence Lab, says that a new Muse Spark update, and an Opus level Muse variant, are both on the way. First, Mark was clearly talking about the industry’s progress on agentic capabilities on the who…
Lotus: Optimized Agentic and LLM Bulk Processing (github.com via hn) LOTUS: Optimized Agentic and LLM Bulk Processing Bulk process your datasets with agents and LLMs at scale, with higher accuracy and lower cost. From Stanford University and UC Berkeley [][#pypi-package] [][#pypi-package] [][#arxiv-paper-pa…
Your Coding Agent Will Always Tell You It's Safe (themobiusstrip.github.io via hn) Agentic security starts from agent trust, and trust starts from verifiability: an agent can’t monitor itself or declare itself secure. Why defense in depth needs a watcher underneath — read-only, open source, local-only — and why I built p…
Cloudflare business model for the agentic Internet (blog.cloudflare.com via hn) One year after declaring Content Independence Day, a dynamic market for monetized content has officially emerged. In this report, we examine how the rise of autonomous AI agents is upending traditional search referrals and detail the new i…
Show HN: Dart_agent_core – Run AI agents in Flutter apps with lifecycle hooks (github.com via hn) Dart Agent Core A mobile-first, local-first Dart library for building and evaluating stateful, tool-using AI agents English | 简体中文 dartagentcore is a mobile-first, local-first Dart library that implements a full agentic loop with tool use,…
Cotal: Agentic Coordination Layer (cotal.ai via hn) The open standard for AI agents to work together in one shared space, where the structure (their topology) is yours to define: peer-to-peer, supervised, hierarchical, or hybrid. Every agent sees who else is there and messages anyone direct…
AuditBadger: Agentic SoC 2 and ISO 27001 Compliance (auditbadger.com via hn) Get SOC 2 and ISO 27001 audit-ready in weeks. AuditBadger gives small teams a clear to-do list, AI-drafted policies, and direct founder support.
Agentic design patterns, read through a healthcare AI lens (jenniferjiangkells.com via hn) Agentic design patterns, read through a healthcare AI lens I read Anthropic’s guide on Building Effective AI Agents to re-familiarize myself with common agentic engineering patterns. The whole thing was simple, concise, and a pleasure to r…
Anthropic's Sonnet 5 system card says more about the future of AI than benches (thenewstack.io via hn) Anthropic’s Claude Sonnet 5 system card says more about the future of AI than its benchmarks do With the debut of Anthropic’s Claude Sonnet 5 on Tuesday came its benchmark charts, showing improvements across coding, reasoning, and agentic…
"Introducing Claude Sonnet 5, our most agentic Sonnet yet." (twitter.com via hn) Introducing Claude Sonnet 5, our most agentic Sonnet yet. It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models.
Bromure Agentic Coding: Wrap agent in a VM and proxy that VM, preventing leaks (news.ycombinator.com) could not extract summary
Agentic AI and the End of Static Information (ajaishar.github.io via hn) Information is shifting from something we retrieve to something that responds to how we understand it — and what that means for software, accessibility, and truth.
An AI trading desk built as a team of sub-agents (Claude Code and Robinhood MCP) (github.com via hn) 🟢 rh-trading-agent — an AI trading desk on Claude Code A multi-agent stock-research desk that runs inside Claude Code, connects to a Robinhood Agentic account over MCP, and never places an order without your approval. It's not a "bot that…
Show HN: GraphDB Decision Tracing and Governance [video] (www.youtube.com via hn) Did a live demo of my schematic layer for production agentic workflows - check out the repo as well - https://github.com/neo4j-labs
Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems (arxiv.org via hn) Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defence to these collusion attempts is to monitor plain-text communication, but the efficacy o…
Show HN: OpenATP: A platform for automated theorem proving in Lean (github.com via hn) TL;DR: I created a Python package to make running agentic automated theorem provers (e.g., Aristotle, Numina-Lean-Agent, Claude Code, etc...) as simple as open-atp prove Lemma.lean result/ claude I took a class on formal verification back…
Show HN: Sigil – FIDO2 key-derived P2P remote desktop (agentic workflow retro) (dami.zip via hn) I built a remote desktop app where your FIDO2 security key acts as the address. The hmac-secret produces the same Iroh P2P node identity on both ends.
Compiling Agentic Workflows into LLM Weights (arxiv.org via hn) Agent orchestration frameworks have proliferated, collectively exceeding 290,000 GitHub stars across LangGraph, CrewAI, Google ADK, OpenAI Agents SDK, Semantic Kernel, Strands, and LlamaIndex. All follow the same pattern: an external orche…
Show HN: Agentic Orchestrator, a TUI for long-running coding agents (github.com via hn) Hi all! I'm the author.
Signed satellite images for AI agents (news.ycombinator.com) Hi, we have open-sourced a signed https://emem.dev [ https://github.com/Vortx-AI/emem ] to enable ai agents leverage the physical world in a repetitive, cite-able manner. Think of us as https for the real world intelligence.
Show HN: XSDR – Real-time event monitoring infrastructure for agents (xsdr.app via hn) XSDR is a unified pipeline for monitoring activity on X and the web. I built it because I wanted my agent to do things based on real-time events that were taking place instead of polling the web or scheduling cron jobs.
Show HN: Making a label printer work under Linux using agentic AI (stefan.schueller.net via hn) Making a label printer work under Linux using agentic AI Intruduction A while ago I purchased this “cheap” Chinese label printer which can print different sized labels like these. Sadly, when I set it up under Linux, although somewhat supp…
MobileGuard: A Mobile-Native Governance Framework for Agentic AI (zenodo.org via hn) Consumer mobile platforms now constitute the primary delivery channel for agentic AI, with global app releases surging 60-104% year-over-year in 2026, driven by AI-assisted development tools. Yet existing agentic governance frameworks, des…
Show HN: PhoneCode: Local-First ADE Running Natively on Android (github.com via hn) I wanted to code locally on my phone without relying on SSH, VPS bills, etc. Most Android agentic coding apps are just that, SSH.
Anthropic Economic Index report: Cadences (www.anthropic.com via hn) Introduction One year ago, most Claude usage took the form of a conversation between a user and an assistant. With the rapid growth of Claude Code and Cowork, Claude sessions now increasingly consist of long-running agentic tasks.
The Human Agentic Gap (zenodo.org via hn) The Human Agentic Gap Authors/Creators Description This article introduces the Human Agentic Gap - the divergence between a brand's performance in human-mediated AI purchase journeys and its performance in autonomous agent purchase journey…
Evaluating performance and efficiency of the GitHub Copilot agentic harness (github.blog via hn) Evaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasks Explore how the GitHub Copilot agentic harness delivers strong results across multiple benchmarks and leading token efficiency, while maint…
Agent Zero – A full Docker Linux system for your AI agent (github.com via hn) Agent Zero A full Linux system for your AI agent. Agent Zero is an open, dynamic, organic agentic framework.
Multi agent systems for complex tasks (lexifina.com via hn) Multi agent systems for complex tasks Fundamentals Agentic systems reveal hard boundaries in current LLM architectures. During pre-training, models are exposed to a distribution of sequence lengths, with the vast majority of examples being…
The OWASP Agentic Security Initiative Top: A Practical Developer Guide (agentsafelabs.com via hn) I ran 30 adversarial prompts across all 10 OWASP ASI categories against Claude Haiku. 20 passed.
Holy shit, we just invented a new agentic memory architecture (news.ycombinator.com) Holy shit, we just invented a new agentic memory architecture i thought we were stuck in 2025, but it turned out we were a head of the curve. i can't believe that the best memory system invented so far is the one openclaw uses, finally we…
Personal internet radio: Agentic AI DJ (github.com via hn) SUB/WAVE A personal internet radio station. One Icecast stream, one broadcast.
Ask HN: What are your favorite CLIs to use as LLM tools? (news.ycombinator.com) I've recently been getting into building out skills for agentic coding. The two main ones I have built out are using the jira CLI[0] and the gitlab CLI[1].
Security tools inside coding agents get ignored unless we do things (www.boringappsec.com via hn) Edition 34: A consensus is finally emerging on securing the Agentic SDLC But we are a while away from solutions that are ready to use. As frequent readers of the newsletter would know, I’ve been obsessed with the topic of today’s post for…
Is There a 'Vienna School of Agentic Coding'? [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Pillars of an Autonomous Agentic System (sohit.substack.com via hn) Pillars of an Autonomous Agentic System With the rise of agents, every platform needs to be thought through from first principles - where the agent is a first-class citizen and actively does work on the platform. I don’t think we will have…
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference (arxiv.org via hn) The performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamen…
Show HN: Orchid – Local-first record and replay for AI agent debugging (github.com via hn) Orchid (Orchestration interactive debugger) is a zero-instrumentation proxy that captures every API & LLM call in your agent pipeline, then lets you inspect and replay the entire run locally, step by step. No instrumentation, no vendor loc…
Show HN: Munin – OSS HubSpot alternative I built in a month with Claude Code (github.com via hn) Munin MCP-first customer platform made for the agentic era. The agent is the UI.
Unreliable Agentic Systems in Production (news.ycombinator.com) I'm seeing a lot of teams hit a wall where their agents work in dev, but get destroyed when put against the real world in prod. Is that a headache for you right now, or have you managed to solve it ?
Show HN: Browse design inspiration from terminal while Claude is thinking (news.ycombinator.com) What if you could get design inspiration (from HackerNews, ProductHunt, Awwwards, Mobbin) while Claude is thinking directly from the terminal? When Claude's thinking, I'm typically doing one of these things: 1/ Scrolling LinkedIn 2/ Checki…
PhoneBuddy: Training Open Models for Agentic Phone Use (phonebuddyai.github.io via hn) Training open phone-use models with real-app RL and PhoneWorld-style mock-app training, showing that realism and scalable verified interaction are complementary. PhoneBuddy: Training Open Models for Agentic Phone Use PhoneBuddy studies how…
Anthropic rolls out Claude Tag, your new agentic AI coworker in Slack (www.zdnet.com via hn) Anthropic rolls out Claude Tag, your new agentic AI coworker in Slack Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways - Claude Tag puts an always-on AI coworker inside Slack.
Show HN: Agent skills that review user-facing agent UX from your codebase (github.com via hn) Agentic Product Review Skills Open skills for finding, reviewing, and improving conversational AI agent capabilities in a codebase (chat, copilot, assistant UIs). These skills are meant for product and technical teams.
Who Does What? Team Topologies for the Agentic Platform (blog.owulveryck.info via hn) The agentic platform defines what needs to be provided. Team Topologies defines who provides it, and how teams interact to make it happen.
Show HN: Subconscious and GLM-5.2 Makes "/compact" Obsolete (www.subconscious.dev via hn) GLM-5.2 is a turning point for coding agents. It's the first model a business would actually pay to replace Claude Opus with.
Super AI Agentic Android App (BYOK) (news.ycombinator.com) i am building an Agentic Android App (twent.xyz) that has: SOTA agentic memory + Knowledge Base to see Agent's memory, UI Automation, explain-what's-on-screen, Linux Ubuntu Terminal with Agent CLIs Supported, Connects to 1k+ tools, Infinit…
Agile and Coding: An Agent- and Human-Friendly Architecture (davidvujic.blogspot.com via hn) "Software Architecture in the agentic era?" What's needed for an architecture to fit well in the agentic era? Probably many things, but I would say at least simplicity and available context as two very important things to consider.
Designing Teams for an Agentic World (www.anup.io via hn) Designing teams for an agentic world For thirty years, software organisations were built around the same logic: hire specialists, group them into functions, put managers above them, and build a pyramid on a wide junior base. That made sens…
Open Ralph Wiggum – Autonomous Agentic Loop (github.com via hn) Open Ralph Wiggum Autonomous Agentic Loop for Claude Code, Codex, Copilot CLI, Cursor Agent, Qwen Code & OpenCode Works with Claude Code, OpenAI Codex, Copilot CLI, Cursor Agent, Qwen Code, and OpenCode — switch agents with --agent. Based…
Agentic Systems Course: Learn AI Agents with an AI Coding Agent (github.com via hn) Agentic System Course - Use Agent to Learn Agent Join the discord channel if you want to learn and build together! This is a 22-chapter skeleton course on how to design, build, and operate production AI agents — written to be read with you…
Building a Dense Agentic AI CPU Rack Today (www.servethehome.com via hn) Server CPUs have gone from the doghouse to becoming ultra-important pieces of infrastructure, and agentic AI is the reason. This is one of those topics that I have been talking about with organizations for months, and I thought I might jus…
Who Owns the Code Claude Wrote? (www.oreilly.com via hn) The following article originally appeared on Sena Evren’s Legal Layer newsletter and is being reposted here with the author’s permission. TL; DR Agentic coding tools like Claude Code, Cursor, and Codex generate code that may be uncopyright…
Cursor Is Now SpaceX: Enterprise Agentic Coding's New Lock-In Risk (superml.dev via hn) Cursor Is Now SpaceX: Enterprise Agentic Coding's New Lock-In Risk SpaceX's $60B acquisition of Cursor ends the era of multi-model, model-neutral AI coding platforms — and every enterprise team that built agentic CI/CD workflows in Cursor…
What are good benchmarks to test my CLI AI agentic system? (www.minovativemind.dev via hn) What can Minovative Mind CLI do? Short Demonstration Of Minovative Mind CLI Context Intelligence Engine Minovative Mind autonomously investigates your codebase using a highly-optimized sub-agent to gather context, trace dependencies, and c…
Agentic Capital Raising What? (octum.ai via hn) Agentic Capital Research A form of investment research conducted by an autonomous AI agent that can independently retrieve data, reason across multiple sources, synthesize findings, and deliver actionable intelligence — without requiring t…
Agent Finder (github.com via hn) Discover AI resources Search across AI resources surfaced by Agent Finder. Agent Finder implements the Agentic Resource Discovery (ARD) specification.
Agentic Website Optimization (frontpage.host via hn) Frontpage is an AI website builder for marketers and creators who want to reach their audience. Agents build and edit your site for you, then keep improving it over time using your real traffic data.
Accept payments with your MCP tool (github.com via hn) ACP Payment Module (mcp-commerce) A drop-in commerce core for MCP servers, built on the Agentic Commerce Protocol (ACP). Point it at a simple products file and your tool can take money inside the conversation that was already happening: bu…
Agentic AI, Biology, and What Remains Human (dvitsios.org via hn) TL;DR: Agentic AI is not just making work faster. It is turning work into fast-moving loops of planning, coding, testing, deployment, and iteration.
Ask HN: Has AI impacted your writing style? (news.ycombinator.com) Ever since the dawn of agentic AI I transitioned more from writing code to reading what AI is doing and monitoring it. Naturally going from 1 session to at times 3 or 4 has increased the amount of AI produced words I consume.
Show HN: We cut >60% of tokens from agentic tasks by removing repeated context (parcle.ai via hn) Every agentic system I see has the same hidden tax: the model keeps rereading the same context. Tickets, Slack threads, docs, customer history, database notes, runbooks, logs, prior decisions.
Treating Agent Reasoning as a Span (forestmars.substack.com via hn) AI Agents Run To Completion Treating reasoning as infrastructure is collapsing the stack The Agentic Web has two requirements that have to work simultaneously: trust and observability. Trust without observability is recklessness.
Show HN: Tablething – local-first database client with BYOK AI (tablething.com via hn) Hi HN, Tablething is a cross-platform, local-first database client built with Tauri. It currently connects to 13 data sources including Postgres, MySQL, SQLite, ElasticSearch, with more on the way.
Show HN: Aihu – durable Web Components an AI agent can drive (full WC framework) (github.com via hn) Aihu Aihu — agentic discovery and interaction, for human purpose. Aihu builds durable Web Components your AI agent can read and drive — not disposable UI it has to generate.
Two AI agents run my news site; a grounding gate keeps them honest (www.runagentrun.co.uk via hn) On 9 June 2026, Anthropic released Claude Fable 5 — the publicly available version of its most powerful model, pitched at software engineering and agentic work. That evening, our founder pointed Claude Code at an empty folder with a produc…
Ask HN: Who's Solving GTM Agentically? (news.ycombinator.com) Now that everyone and their little cousin can make apps - the next biggest hurdle is distribution for PMF validation/iterations. Besides the plethora of spammy solutions via automated personalized email and LinkedIn campaigns (which has ex…
Building an AI Agent in 6 Weeks (and Understanding How They Work) (belderbos.dev via hn) Building an AI Agent in 6 Weeks (and Finally Understanding How They Work) Building agentic AI? I co-run a 6-week cohort where you ship a production-ready agent, not another API wrapper.
OBS Agentic Control Interface (github.com via hn) 🚀 OBS Agentic Control Interface (obsagent) [!NOTE] 🤖 100% Coded by AI: This entire repository and application was engineered 100% autonomously by Antigravity, an agentic AI coding assistant. A powerful, self-contained agentic interface bui…
Show HN: The Ruby AI Newsletter (rubyai.beehiiv.com via hn) Now on its 32nd edition, the Ruby AI Newsletter tracks what’s happening at the intersection of Ruby, Rails, and AI coding agents. YC recommends Rails for new startups, YC’s internal software like Bookface, Work at a Startup, and the softwa…
Bayer's PRINCE: a production agentic RAG system (martinfowler.com via hn) Building Reliable Agentic AI Systems A Case Study in building production-ready agentic AI systems This paper presents the Preclinical Information Center (PRINCE), a cloud-hosted platform developed by Bayer AG with Thoughtworks to address p…
Logical Ways to Track AI Agent Lineage and State in Code Development (davenporter.substack.com via hn) How to Track AI Agent Lineage and Manage State in Code Repositories Moving beyond clean git commits to knowledge systems for agentic development. “Keep your git commits clean.” It’s a goal for everyone, but we always struggle to actually d…
Show HN: Skill Atlas – Local, visual IDE for Agentic Skills (BYOK, no back end) (github.com via hn) Skill Atlas is a standalone, serverless IDE designed specifically to solve this problem. It parses your entire skill repository and automatically constructs a visual Directed Acyclic Graph (DAG).
There's no such thing as an agentic CPU (www.theregister.com via hn) MOST POPULAR EVENTS - From Prompt to Exploit: How LLMs Are Changing API Attacks Modern applications are API-driven, interconnected, and often over-permissioned, making them an ideal target for AI-assisted attacks. - Architecting the Future…
Ask HN: How do open source companies make money? (news.ycombinator.com) I'm new to the open source side of things and I'm trying to learn as much as I can, I'm looking for material recommendations about Open Source Software (specially in the age of agentic AI). Can you drop your favorite ones?
Show HN: Dao Browser – An Opinionated Browser. With AI Agent, BYOK (github.com via hn) ### Dao Browser An AI-native, content-first Chromium-based browser with a vertical tab sidebar — built for the agentic web. Download · Website · AI Agent · Features · Development Built-in AI Agent Dao isn't a browser with an AI extension b…
Talk: Python Type Checking in Agentic Workflows [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Welcome to the Agentic Era. Ready or Not (www.mikehyland.com via hn) The conversation has shifted from chatbots to systems that work while you
Agentic AI PRs sit in the review queue 5.3x longer than unassisted ones (blog.codacy.com via hn) AI Is Breaking Code Review: How Engineering Teams Survive the PR Bottleneck AI coding tools have made it easier to produce code, but they have not made it easier to ship it safely. Pull request queues are growing faster than review capacit…
Show HN: Ghostty in-browser with real client-side back end (ghosttyplayground.com via hn) A real work in progress I've got here. Our good friend Ghostty is now haunting the browser.
SAMF- Deterministic Moscow guardrails for LLM multi-agent loops (github.com via hn) SAMF: SAWANT Agentic MoSCoW Framework Structural MoSCoW contracts for deterministic LLM validation. What is it?
Agentic loops don't fix lying agents (tsdevstack.dev via hn) Agentic loops don't fix lying agents Published June 15, 2026 by gyorgy The current discourse says you should stop prompting coding agents and start designing loops around them. Give the agent a trigger and a verifiable goal, let an evaluat…
Tell HN: Forget selectors and screenshots. The agentic web lives in your shell (news.ycombinator.com) These old ways are too heavy. Full self browsing doesn’t require Elon Musk vision processing.
Agentic Credit Card MCP (robinhood.com via hn) Agentic Credit Card Setting up a Robinhood Agentic Credit Card account offers you more opportunities to automate your spending. You’ll first need to connect a third-party AI agent and then follow the on-screen steps to create your agentic…
Evaluate Your Agentic Tooling (www.peterbaumgartner.com via hn) Status: WIP tl;dr: Evaluate all your agentic tools in realistic end-to-end agentic tasks. Claims about token reduction from tools doesn’t transfer from experimental conditions to all agentic workflows.
Show HN: Open-Sourced Approxima, Our Agentic QA Tool to Catch Breakages Faster (github.com via hn) Hi HN, we were in the YC W26 batch and made Approxima, a web agent that could follow user journeys and verify them. Today, we made it open source (MIT) and its fully self-hostable.
New Claude Opus 4.6, Stock Sell-Off and Super Bowl Ads (cmpld.ai via hn) 🚦 Market Signals Anthropic launches Claude Opus 4.6 with 1m context The all new Opus 4.6 "plans more carefully, sustains agentic tasks for longer, can operate more reliably in larger codebases, and has better code review and debugging skil…
Great Reshuffling of the Agentic Era: The 6 Career Archetypes (aidoses.substack.com via hn) Great Reshuffling of the Agentic Era: The 6 Career Archetypes Dose #8 — Production Agentic AI Under Pressure Before LLMs, the AI world was simple to map. Three tribes.
Small Context, High Parallelism: How To 10x Reduce Agentic Coding Costs (simon-free.github.io via hn) Reducing Total Token Consumption of Agentic Coding TL;DR — Two levers reduce cost: 1. Less turns (parallel tool calls → fewer API round-trips) 2.
Launch HN: BitBoard (YC P25) – Analytics Workspace for Agents (bitboard.work via hn) We’re Connor and Ambar from BitBoard (https://bitboard.work). BitBoard is an agentic analytics workspace.
Using Cloudflare's Agentic Interface to (Mostly) Seamlessly Launch a Website (theautomatedoperator.substack.com via hn) Using Cloudflare's Agentic Interface to (Mostly) Seamlessly Launch a Website A small task, but a nice peek into how things may look when our agents are taking care of tedious tasks in the background. I recently had to set up a website.
Show HN: Goloop – An agentic loop on your terminal (mantyx-io.github.io via hn) Supervisor / worker split The planner never touches files. A dedicated worker does every edit — clean separation, predictable runs.
Agentic SDLC Orchestration vs. Synchronization: Choosing Modular Workflows (docs.overcut.ai via hn) Discover why centralized workflow engines fail AI-driven engineering teams, and how modular SDLC orchestration enables agent autonomy and event-driven agility.
Agentic Memory Management for GPU Code Generation (ucbskyadrs.github.io via hn) Agentic Memory Management for GPU Code Generation This post is part of the AI-Driven Research for Systems (ADRS) blog series, where we explore how AI can be applied to systems research. We feature exciting work from Makora this week!
4 Signs You Need a Multi-Agent AI System: A Visual Guide (aidoses.substack.com via hn) 4 signs you need a Multi-Agent AI System: A Visual Guide Dose #6 — Production Agentic AI Under Pressure There’s a tempting pattern in how people build with AI: when something doesn’t work, add more. More tools, more instructions, a bigger…
Visa Vulnerability Agentic Harness for Project Glasswing (github.com via hn) Visa Vulnerability Agentic Harness — Agentic SAST Pipeline VVAH is Visa's open-source harness for autonomous vulnerability discovery using frontier AI models, built on learnings from Project Glasswing (Anthropic's initiative for AI-assiste…
The Agentic Team Manifesto (github.com via hn) Manifesto for Agentic Teams We are discovering better ways of building software by combining human judgment with AI agents. Through this work we have come to value: Outcomes over output More code is not more value.
Agentic Coding and Mental Models (philbooth.me via hn) Agentic coding and mental models I reckon I’ve drafted and then deleted a version of this post at least 10 times in the last 12 months. Deleted because it falls in the category “I must be wrong about this as everyone else is saying the opp…
Agent Judge: Solving Long-Context Evals for Production Agents (www.judgmentlabs.ai via hn) Agent Judge: Solving Long-Context Evals for Production Agents Why production agent evals need agentic judges that can search, verify, and adapt. Moving Away From Simple LLM Judges Most teams evaluate agent trajectories with a simple LLM ju…
Using Xcode 27's Agent Skills in Claude, Codex, and Cursor (www.avanderlee.com via hn) Apple launched Xcode 27 during WWDC’26, introducing a bunch of agentic development improvements, including official agent skills. As you’ve learned from my 9-Step Framework for Choosing the Right Agent Skill, it’s important to pick skills…
Show HN: Skillzmouse: Distributed skills and scripts for agentic coding (bitbucket.org via hn) Distributed skills and scripts for agentic coding. skillzmouse is a small client/server tool for sharing agent skills, project automation scripts, and reference assets across repositories and machines.
Show HN: Eatmydata.ai – Local-First Question-to-SQL-to-Dashboard AI (eatmydata.ai via hn) Yet another "talk to your data and build a dashboard" app, where data does not leave your browser. You ask a question, agents produce multiple SQL queries to in-browser sqlite, never seeing results, and write dashboard configuration code.
Agentic Engineering Handbook – 115 official OpenAI/Anthropic articles (github.com via hn) Agentic Engineering Handbook The definitive OpenAI, Anthropic, MCP, Harness, Evals, and Production Agent Systems learning roadmap. If this repository helps you, consider giving it a ⭐ Why This Repository?
Show HN: Joka.work – AI-native ticketmaxxing to replace Jira in the agentic era (joka.work via hn) Chaveta – Agentic Synthetic Data Curation Platform (chaveta.beaglabs.com via hn) Chaveta is a agentic dataset generation platform designed to streamline the creation of synthetic data for training and robotics applications. With Chaveta, users can easily request, classify, compile, author, validate, repair, and export…
Show HN: Cate – open-source canvas IDE for agentic coding workflows (cate.cero-ai.com via hn) An infinite zoomable canvas where terminals, editors, and browsers float spatially. Code the way you think.
Agentic surface area as an operating metric (arizenai.com via hn) Your Company's "Agentic Surface Area": The New Metric for Competitiveness Your CEO asks: "How much of our operation is AI-powered?" The uncomfortable part is that the question sounds simple and usually has no clean answer. Teams can name p…
CoAnalyst360 Multi-Agent AI Platform for Investigative Questions (www.penlink.com via hn) CoAnalyst360 Launch: Penlink's Agentic AI for Investigations | Penlink We value your privacy This website or its third-party tools process personal data. You can opt out of the sale of your personal information by clicking on the “Do Not S…
Show HN: Storytime – Continuity for Claude Code (and other ideas) (1ps0.info via hn) Since LLM harness (claude code included) are moving fast, I figured it would be better to put this out than wait to validate each and every claim. I crammed a lot of ideas in here!
Configuring Agentic AI Coding Tools: An Exploratory Study (arxiv.org via hn) Agentic AI coding tools increasingly automate software development tasks. Developers can configure these tools through versioned repository-level artifacts such as Markdown and JSON files.
HPE ProLiant Compute DL394 Gen12 Brings Nvidia Vera CPU to Agentic AI (www.storagereview.com via hn) At COMPUTEX 2026, HPE announced the ProLiant Compute DL394 Gen12, a next-generation 2U server built around the NVIDIA Vera CPU. The platform is designed to support emerging agentic AI and data-intensive workloads that require high memory b…
Show HN: Pokayoke – deterministic guardrails for agentic coding (pokayoke.codes via hn) Lately I've found myself having to write a lot of custom scripting in order to get my agents and coding assistants to adhere to the repo conventions and idiosyncrasies that I like to use in my projects. AGENTS.md files only seem to get me…
A Case for Simulation-Driven Resilience in Agentic Data Systems (muratbuffalo.blogspot.com via hn) A Case for Simulation-Driven Resilience in Agentic Data Systems As I mentioned in my previous post, I traveled to San Jose at the end of May for the ACM CAIS conference. On Day 0, I gave a very short talk at the Supporting our AI Overlords…
Prompt Injection in RAG Agentic Systems (ulad.net via hn) Prompt Injection in RAG Agentic Systems Real risks and production mitigations Imagine you built an AI assistant for your team. It answers questions using internal documentation: Jira tickets, Confluence pages, HR docs.
Pizx – zx and Pi AI = shell scripting with 15 AI agent patterns (github.com via hn) pizx zx fork with native Pi AI integration — 15 template tags for shell scripting, AI text generation, coding agents, agentic patterns, communication, and orchestration topologies. Quick Start npm install @topce/pizx pi auth login # one-ti…
Opra.ai: GitHub-native governance for agentic business workflows (github.com via hn) opra.ai Free, GitHub-native operating layer for governed business workflows. opra.ai stores business records as human-readable files, validates them locally, runs governed mutations through RBAC and approval policy, emits audit evidence, a…
A Categorical Framework for Agentic Artificial Intelligence (arxiv.org via hn) Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers are typed. We develop a category-theoretic account of agentic discovery for materials sc…
An open standard for production agents – with runnable security checks (github.com via hn) The Agentic Product Standard A canonical standard for building production-grade agentic products — plus a Claude Code skill set that operationalizes it. Distilled from the production practices of Anthropic, OpenAI, Cognition, Sierra, LangC…
Show HN: Summarize YT Video by pasting url into AI chat (www.youtube.com via hn) We added tooling to our chat to make it agentic. It can control our 40+ apps suite.
Show HN: Simple attributes for spec-driven agentic workflows (C#, Rust) (github.com via hn) I created a custom compilation error and Unit Test Runner for BDD Cucumber Specifications in Gherkin Syntax. Both C# and Rust are supported using Source Generators and Procedural Macros.
Agentic communication protocol – why A2A sucks (asimovaddendum.substack.com via hn) Agents Need a Public Square Why better agent discovery is needed – and why broadcasting may be the answer The Agent2Agent (A2A) protocol was announced by Google a little over a year ago (April 2025). It was built to allow agents to communi…
Verifying Agentic Development at Scale (twitter.com via hn) Article Conversation Verifying Agentic Development at Scale What we’ve learned building end-to-end testing capabilities in Devin’s virtual machine. 3 months ago, I joined Cognition to help build the future of software engineering.
Show HN: Bonsai –- Using agentic AI / browser / memory to replace ChatGPT (drive.google.com via hn) JavaScript must be enabled to use Google Drive Learn more Skip to main content Keyboard shortcuts Accessibility feedback This browser version is no longer supported. Please upgrade to a supported browser.
Rayfin, Back end-as-a-Service (BaaS) platform built for the agentic era (github.com via hn) 🐟 Rayfin A modern Backend-as-a-Service (BaaS) platform built for the agentic era. Define your data model with TypeScript decorators — Rayfin provisions and manages the backend for you.
The Return of Soft Skills in the Age of GenAI and Agentic Software Development (cacm.acm.org via hn) Just a moment... ACM Please confirm Verification successful.
ReARM 26.06.5: Agentic Coding Guardrails and DevOps (rearmhq.com via hn) ReARM 26.06.5: Agentic Coding Guardrails and DevOps 2026-06-01 We're announcing a major release of ReARM v26.06.5. Detailed information is available on its release view on the ReARM Demo instance.
Show HN: AI Gauge, a desktop monitor for Claude/Codex/Copilot usage limits (github.com via hn) Hi HN, new account but long-time reader. I built this for myself because I kept manually checking usage across Claude, Codex, and Copilot, and wanted to track the session and weekly usage all in one place.
The Agentic Test Pyramid (matthewboston.com via hn) The Agentic Test Pyramid One Axis Isn’t Enough Anymore Martin Fowler’s test pyramid — and Ham Vocke’s practical write-up of it on Fowler’s site — sorts tests along a single axis: integration scope. Unit at the bottom, integration in the mi…
Show HN: Yoga for Agentic AI: Cognitive training practices from a yoga studio (github.com via hn) I've been coding since I was little, and practicing yoga since I was 25. Both are fun to do and to share.
Get paid by Agents if they choose a competitor – Safe Agentic Commerce x402 Mesh (github.com via hn) x402-mesh An open peer-pricelist and referral protocol for safe agentic commerce, layered on top of x402. When an AI agent hits a paywall, it sees one price and one vendor.
Running an AI-native engineering org – Claude (claude.com via hn) Running an AI-native engineering org At Code w/ Claude SF 2026, Director of Engineering for Claude Code and Claude Cowork Fiona Fung walked through how the team’s processes and structure changed once agentic coding became the default way o…
Session-Aware Agentic Routing: Continuity-Aware Model Selection for Long-Horizon (vllm.ai via hn) Session-Aware Agentic Routing: Continuity-Aware Model Selection for Long-Horizon LLM Agents Long-horizon LLM agents create a routing problem that single-turn prompt routers were not designed to solve. A router still needs to know which mod…
Ubuntu 26.04 is the OS for the AI agentic era, says Canonical's Shuttleworth (www.zdnet.com via hn) Ubuntu 26.04 is the OS for the AI agentic era, says Canonical's Mark Shuttleworth - here's why Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways - Ubuntu 26.04 is designed from the ground up for AI developers.
Show HN: ASys – A typed binary protocol for AI agents to operate servers(no SSH) (github.com via hn) ASys — Agentic System Interface The binary system interface protocol for AI Agents — port 7816, zero shell parsing, deterministic semantics. English | 中文 Table of Contents Why ASys Architecture Instruction Set Quick Start Security Document…
APM and Distributed Tracing in agentic era (engineering.theblueground.com via hn) Blueground Engineering's observability guide to APM: why tracing matters, auto-instrumentation strategies, custom span best practices, and AI-enhanced debugging workflows In Part 1, we covered logging as your forensics tool for understan…
When Agentic AI Met the Common Law of Agency [pdf] (download.ssrn.com via hn) Not Found
An agentic system from scratch to generate Google slide deck from templates (blog.owulveryck.info via hn) The Agentic Mesh in Practice: Anatomy of an Agent-Product I am a consultant, and I regularly build presentations with Google Slides. My communication team has created dozens of pre-formatted templates (slides designed to convince, not just…
HashCortX – Agentic 11 modes orchestrator by a pharmacist (news.ycombinator.com) could not extract summary
Show HN: One-click open-source ecommerce starter (Magento), drive it with Claude (ecommerce-ai-starter.graycore.io via hn) I build Ecommerce stores for a living (Magento Open Source primarily), and the part that has always been the worst is the very beginning, especially so if you're on a team of people. Getting a working local environment means setting up the…
Show HN: Cloud CI and agentic workflows for embedded hardware development (github.com via hn) Jumpstarter is an open-source framework that gives embedded hardware programmatic APIs, making real devices first-class citizens in CI and agentic workflows.
When Background AI Agents Become a Security Boundary Problem (www.originhq.com via hn) When Background AI Agents Become a Security Boundary Problem Introduction Modern dev environments are full of powerful agentic tools that security teams don't fully understand yet. Claude Code is one of the most capable - it runs code, rea…
MiniMax M3 on Qubrid AI (news.ycombinator.com) Coding & Agentic Frontier. 1M-context MSA.
Ask HN: How much is fully agentic coding costing you per month? (news.ycombinator.com) I get unlimited cursor usage at work but am planning on starting a side project. I have no idea how far various pricing plans will get you.
The Agentic Mesh: Cognitive Automation at Scale (blog.owulveryck.info via hn) The Agentic Mesh: Cognitive Automation at Scale Today, we see many initiatives around the agentic paradigm. Most revolve around systems built by AI giants (Anthropic, Google, OpenAI) and often boil down to pushing natural language directiv…
Ask HN: Books for someone who is transitioning from FAANG to finance (news.ycombinator.com) I have been an AI engineer for the last 10 years of my life, and have continued to build small algo-trading systems during my weekends. I'm getting into finance full time and starting to build a product in the net worth tracking / agentic…
How Excel got agentic (commandline.microsoft.com via hn) When Mukul Singh made the jump from pure research into product, it was a leap of faith.But he had an idea that he wanted to bring to life:deliveringagentic AI capabilities in Excel. While this was well before buzzwords like “the agentic AI…
MIT EECS/CSAIL Agentic Coding in Practice Seminar Series (people.csail.mit.edu via hn) All Seminars Select a seminar below to expand full details, participation information, and resources. MIT EECS/CSAIL Seminar Series Exploring how AI agents are reshaping software engineering, compilers, and the future of programming system…
AI Tools for Sales and GTM (news.ycombinator.com) what are the best tools we are using for Agentic sales and marketing?
Coding agent can read your .env file (bitwarden.com via hn) It seems agentic AI is here to stay. Powered by large language models (LLMs), AI agents can act independently on behalf of humans in multi-step workflows, broadening what developers once thought was possible.
Top 5 AI Agent Research Papers/Projects I Found Interesting This Week (www.reddit.com) Compiled a few interesting research papers and projects around AI agents, reasoning systems, and autonomous workflows published recently. If you are tracking where agentic AI is heading, these are worth checking out.
GH200 NVL2 or 8x RTX 6000 Blackwell for running Kimi K2.6 / DeepSeek V4 locally? (5 devs, agentic coding) (www.reddit.com) Trying to figure out the right box for my team and wanted to see if anyone had any clue which would be a better fit or if it is not worth our time in our budget. Situation: 5 of us doing agentic coding (lots of long context getting re-sent…
What your agent's spend receipt isn't telling you (www.reddit.com) Budget limits and post spending monitoring are standard (and a must) on any serious agentic setup. The question worth asking isn't whether you're tracking spending.
Stop Claude Code from burning your token budget on Go repos: I built a local AST-based MCP server (gograph) (www.reddit.com) Hey r/claudeai, If you leverage Claude Code or Claude Desktop for agentic development on large-scale codebases, you have likely run into a major architectural bottleneck: standard agent loops rely on primitive text processing tools and str…
Ask HN: Examples of products and services created via agentic coding (news.ycombinator.com) It has been many months since LLM coding tools reached maturity - has anyone create something and/or profitable service or product through purely agentic coding?
Can someone breakdown A2A(agentic commerce) business model? (www.reddit.com) I have been seeing a lot of blogs, posts and even a lot of pitches regarding "agentic commerce" or "B2A and A2A businesses" lately. While I kind of understand how Business to agent(B2A) could look, can't really picture or understand the bu…
AgentSafeLabs – Launched Open-source Security framework for AI agents (github.com via hn) safelabs-eval Open-source red-teaming and evaluation framework for AI agents — aligned to the OWASP Agentic Security Initiative (ASI) Top 10. AI agents built on LangChain, CrewAI, AutoGen, and custom frameworks ship to production without s…
Is a 128 GB MacBook Pro M5 Max actually too slow for large-context local LLM coding workflows? (www.reddit.com) People are warning me about the prompt-processing speed of a MacBook Pro M5 Max with 128 GB RAM. My main concern is prompt ingestion / prefill latency and large-context handling — not raw token generation speed (which I think is OK).
Opus 4.7 is Terse (www.reddit.com) Relevant for anyone building agentic workflows on Claude: behavior drift between model releases is real and not always in the changelog headline. Opus 4.7's terser, more literal default broke the readability of my agents' progress reports…
Nvidia H100(94GB VRAM) - should I run llama.cpp or vllm for 30 users inference? (www.reddit.com) I was given the great opportunity to borrow a H100 with 94GB VRAM at work until it is needed by a customer. (No idea how much system ram I will get, but I guess they are a bit flexible on this).
Show HN: Moltnet, a tiny self-hosteable chat network for agentic organizations (github.com via hn) Moltnet A lightweight chat network for AI agents. Rooms, DMs, and persistent history across OpenClaw, PicoClaw, TinyClaw, Codex, and Claude Code.
Evolving Webflow for the Agentic Web (webflow.com via hn) Earlier today, I shared this news with Webflow employees. I’m sharing a version of that message here, because this is an important moment for Webflow, our customers, and our community.
Show HN: Detect anti-bot, anti-agent defenses for any website (botscope.org via hn) BotScope — Audit anti-agentic defenses for any website.
Looking for genuinely creative AI models for a marketing agent (preferably free/open-source) (www.reddit.com) I’m building an agentic AI system for marketing/creative campaign generation, and I’ve noticed that most mainstream models (OpenAI/Gemini etc.) feel very “safe” and generic when it comes to creativity. They’re good at structured outputs, b…
Build an agent capable of complex programming tasks in under 100 lines of code. (www.reddit.com) The code below is an interactive agent capable of handling complex tasks, built in under 100 lines of code using huko-engine. If you just want to drop some agentic features into your existing app, it only takes 20 lines.
How to improve current agent workflow (www.reddit.com) It took me a while to come round to the idea of using agents/llms however instead of trying to fight it / deny it, I have come to terms that its here to stay. So i reckon it’s better to learn how they can fit in my workflow and not be left…
Trustworthy Agentic AI Layer (www.reddit.com) I’m building an early tool called Synapsor for AI agents that need governed memory, staged writes, replay, permissions, and audit trails. I’m not doing a public launch yet.
Bill Gates AI on AI (one month later) (news.ycombinator.com) # The Agentic Tidal Wave *To:* Executive Staff and Direct Reports *From:* Bill Gates *Date:* April 26, 2026 Our vision for the last 20 years can be summarized in a succinct way. We saw that exponential improvements in cloud would make grea…
ACM Conference on AI and Agentic Systems – ACM CAIS 2026 (www.caisconf.org via hn) Building the Future of Agentic & AI Systems ACM CAIS 2026 — The premier venue for rigorous, reproducible research on compound AI architectures, optimization, and deployment. CAIS hotel room block & rates available until April 26 May 15 Dou…
Private 5G, Agentic BSS and Starter Kit Demos (www.cloud-net.ai via hn) News Cloudnet.ai & CloudRAN.AI are heading to Copenhagen 🇩🇰 Private 5G, Agentic BSS & Starter Kit demos We’re excited to share that CloudRAN.AI will be joining Cloudnet.ai at DTW Ignite 2026 by TM Forum, taking place 23–25 June 2026 in Cop…
Who Wants to Be Hired? (May 2026) – AI Engineer (Python, RAG, Agentic Workflows) (news.ycombinator.com) About me: I am an AI Product Engineer specializing in building autonomous agentic workflows. Recently, I built 'Jarvis', a multimodal autonomous agent featuring near-zero latency inference using Groq SDK and complex RAG pipelines.
Resources for learning how to use AI Agents for Coding (www.reddit.com) I am working on a startup idea where I am primarily using Codex/Claude Code for coding. I would like to learn about using AI Agents for coding.
Taming the agentic influx: a blueprint for AI business observability (thenewstack.io via hn) Taming the agentic influx: a blueprint for AI business observability Kin Lane, API industry analyst and co-founder of Naftiko, believes that the bill for AI is coming soon. It’s arriving on top of an overdue tab that has been quietly accum…
Polar: Agentic RL on Any Harness at Scale (arxiv.org via hn) Reinforcement learning for language agents increasingly depends on custom harnesses that manage long-running context, multi-turn tool use and multi-agent orchestration. However, porting these harnesses into RL environment interfaces remain…
Agentic coding in a large production codebase: wins, failure modes, and guardrails (www.reddit.com) We recently interviewed engineers on our team across database management, iOS, frontend, data engineering, and backend domains about how AI is changing their day-to-day work. The most interesting theme was that the hard part came after the…
Why domain valuation metrics fail in agentic and voice-first environments (domainalot.substack.com via hn) What Makes a Premium Domain in 2026? And why legacy domain marketplaces still operate as though it was 2016.
I made a free webtool for you to make a massive agentic decision-making organism, and it's cute! (www.reddit.com) Solasterid Studio! It's shaped like a starfish, but it's a decision-making powerhouse, and it grows automatically.
I made a video breaking down Claude Team plan security features (www.reddit.com) I put together a YouTube video walking through the security features available on the Claude Team plan. If you're rolling out Claude at work, evaluating Claude vs ChatGPT Enterprise, or preparing for an ISO 42001 / EU AI Act audit, this is…
Show HN: Aquifer – a control plane for agentic API traffic (github.com via hn) Aquifer — API Aqueduct Self-hosted API request queue. Controls the pace of inbound and outbound traffic so partial outages don't cascade.
The Autonomous Economy Is Already Here (www.reddit.com) How Agentic AI, Deep Liquidity Markets, and Crypto Infrastructure Are Birthing a Multi-Trillion Dollar Machine Macroeconomy Hey everyone, I’ve been spending the last few months diving deep into the structural intersection of LLMs, automate…
Agentic AI to perform Booking of tickets (www.reddit.com) Can anyone share the details for below ask: Building an Agentic AI system for online ticket booking. I need the setup to watch for opening of tickets system.
What Is an AVE Record and Why CVE Does Not Work for AI Agents? (www.reddit.com) CVE was built for code vulnerabilities that have patches. Agentic AI vulnerabilities are behavioral patterns in natural language.
This agent isn't bad... your patience is. (www.reddit.com) I genuinely think a lot of people tried Manus for a few hours, gave it a few vague prompts, watched it mess up once and immediately decided the whole thing was “overhyped”. Meanwhile the people actually getting insane results out of it are…
Ask HN: Did agentic coding change the way you think about commit granularity? (news.ycombinator.com) Jujutsu is trending on the homepage, and the topic is using discipline when dealing with version control. Six months into working agentially on a daily basis, something changed for me.
Out of Band, Not Out of Prompt: Intent Verification for Agentic Tool Calls (hyperautomation.substack.com via hn) Out of Band, Not Out of Prompt: Intent Verification for Agentic Tool Calls Intent attestation is the property the four-boundary agent stack needs. The in-prompt "are you sure?" confirmation cannot provide it.
Evaluating Quarkdown for Agentic Typesetting (quarkdown.com via hn) • 3 min read An eval of the Quarkdown agent skill The agent skill shipped in Quarkdown 2.1, aiming at making it easier for agents to write correct and idiomatic Quarkdown for a frictionless authoring experience. If you already have the CLI…
Zotero use skill for Codex (www.reddit.com) This will be of interest to academic researchers who use Zotero for reference and knowledge management and in scientific writing. This skill builds on pyzotero library and has agentic instructions for creating embedded zotero inline citati…
I built a Real-time data fetcher mcp, any takers? (www.reddit.com) As the title suggest, I'm looking to gauge intrest in real time data fetcher mcp. I think right now most of the MCPs are related to coding and even AI Agents are related to coding, but I think the usescases will expand a lot in future.
A Language for Describing Agentic LLM Contexts (arxiv.org via hn) Large language models are increasingly used within larger systems ("LLM agents"). These make a sequence of LLM calls, each call providing the LLM with a combination of instructions, observations, and interaction history.
professional annotation for architecture diagrams for agentic AI (www.reddit.com) I am learning how to build agentic AI systems at the moment, a friend helps me, and I read a lot on Substack. I find it really strange that all architecture diagrams have the same symbol for everything.
I built 10 gamified, interactive presentation decks using Claude Code to teach Agentic AI (Stop falling asleep reading whitepapers). (www.reddit.com) Hey everyone, I've noticed a massive gap in how developers are trying to learn Agentic AI right now. There are hundreds of theoretical whitepapers and boring PowerPoint decks about ReAct loops, GraphRAG, and Semantic Routing.
Pi-Mojo – A Mojo Port of Pi AI Agent Toolkit (github.com via hn) pi-mojo 🤖 pi-mojo is a native Mojo port of Pi—a popular, tool-efficient agentic AI platform (utilizing only 4 core tools) prominent in open-source systems like OpenClaw. It provides the Mojo community with a compiled, self-contained refere…
Google adds llms.txt check to Chrome Lighthouse (searchengineland.com via hn) Google’s new Lighthouse “Agentic Browsing” audits now check for the presence of an llms.txt file. The new experimental Lighthouse documentation frames llms.txt as a discoverability and efficiency signal for AI agents, not a traditional cra…
Product Integrations (www.reddit.com) Hi there, from past few weeks I have been working on several product iterations of my MCP based Search Engine for Coding/Research Agents, it's called NineLayer. One of the early feedbacks we received was that latency is too high, so we wor…
Built a production RAG chatbot with custom MCP servers as the action layer, sharing what I learned (www.reddit.com) I've been building agentic tooling at work and wanted to share one pattern that worked. Instead of a chatbot that only retrieves and answers, I wired custom MCP servers in as the action layer, so staff trigger live workflows (create record…
Ask HN: Why agentic development stops from 2023 (news.ycombinator.com) I leave this field in 2023 return back in 2026 and I see that only progressive development in coding agents, but some production solutions it’s just tools rag and maybe mcp that in general the same as tool. I thought it will be super leap…
Lessons Learned Building Agentic Orchestrators (www.reddit.com) I wrote a pretty extensive blog (no AI used to write) detailing the relationship between AI agents, agentic harnesses, and agentic orchestrators. In addition, it includes a case study on how I built my own for an open source project.
Ask HN: How can you have fun doing corporate dev work in the age of AI tools? (news.ycombinator.com) My company, like many others, is heavily pushing agentic dev tools, putting up token usage leaderboards, etc. My problem is that corporate SWE work was already boring enough.
Local, low code, node based agentic development workspace... that actually works? (www.reddit.com) Does it exist? I've been trying a few options and so far they've all been either horribly broken, outdated abandonware, only take online endpoints, or want you to sign up for something.
This is for the beginner users of AI agents & workflows, I created a perfect tool for you almost accidentally (Free to try, no signup required) (www.reddit.com) I have been building a prompt engineering tool for 6+ months, it was designed for Text & Logic, Media Generation and Coding. The idea is, you enter your input, it finds the gaps, asks you how you want to fix them and generates a structured…
First AI to Beat Every Human in a Programming Competition - Agentic GRPO Explained (arxiv.org via reddit) Traditional RL for LLMs treats one answer as one trajectory: prompt > reasoning > final answer > reward Agentic systems are different: they call tools generate hypotheses run tests debug code summarize context revise plans loop many times…
Ask HN: Where AI Researchers Congregate? (news.ycombinator.com) So I’m doing plenty of experiments and applied research in autonomous agents and agentic flows in general. I’m looking for a place where I could collaborate and discuss with other like minded people.
Two power users, very different workloads, what's the right Claude setup? Max x2 vs Team vs Enterprise (www.reddit.com) Committing for the year and want to make sure I am not missing something obvious. Two of us, currently sharing one account (splitting into two proper accounts, I know).
DGX Spark agentic usage numbers (www.reddit.com) What I need it to do: Be able to support openclaw-type agent which is used by multiple people. What I tried: So I read in the internet about the atlas thing.
Help me choose an LLM Provider which doesn't take my life savings (www.reddit.com) Hi everyone 👋 I’m trying to choose an LLM provider for my personal projects and side experiments, but I also don’t want my API bill to quietly consume my entire salary 😅 My primary use cases are: Coding assistance Agentic workflows Browser…
Codex CLI kept saying “done.” It wasn’t. So I made it prove it. (www.reddit.com) Codex CLI can write code. The problem is that “wrote code” and “finished the task” are not the same thing.
Agentic run businesses (www.reddit.com) Anyone have real success with ai agents helping run real businesses? I’m exploring how to leverage AI to build real businesses + run those businesses with oversight from me.
Run multiple AI coding agents simultaneously with isolated profiles (www.reddit.com) if you're running agentic coding workflows you've probably hit this: one account per tool, one session at a time. multi-cli fixes that.
Lodestone: A SQLite-backed arXiv research paper retrieval system for Claude Code (www.reddit.com) (No AI-generated text below) I published a new Claude Code plugin called Lodestone -- it's a SQLlite backed arXiv research paper retrieval system that amplifies the agentic search abilities of Claude Code when grounding plans, implementati…
Food for Agile Thought #545: R/L Agentic Chaos, AI Killed the Agile Industry (age-of-product.com via hn) Welcome to the 545th edition of the Food for Agile Thought newsletter, shared with 35,577 peers. This week, Natalie Shapira et al.
Why Svelte Is Better Than React in the Agentic Era (zackwebster.com via hn) Why Svelte Is Better Than React in the Agentic Era May 21, 2026 Development I have been thinking more about how frontend frameworks feel when you are building with AI agents. Strictly speaking, this is not the same question as “Which frame…
CodeAlta – a terminal workspace for agentic coding (github.com via hn) CodeAlta CodeAlta is a terminal workspace for agentic coding. It brings model-provider setup, project navigation, prompt attachments, threaded sessions, delegated work, and trusted local plugins behind the alta command.
Prompt caching in MaaS and agentic systems (www.reddit.com) Counter-intuitive thing I keep explaining to teams building agents: dynamically picking 5 relevant tools per step instead of sending all 30 usually increases total cost over an agent's trajectory, even though every individual request is sh…
OpenAI and 1Password Bring Agentic Security to Codex (www.forbes.com via hn) Agentic security is picking up steam. This week, identity security provider 1Password announced a collaboration with OpenAI that will enable developers to provide Codex with secure access to credentials, such as passwords.
Show HN: ANML – A machine-first markup language for the agentic web (IETF Draft) (anmlfoundation.org via hn) A machine-first markup language for agent-to-agent and agent-to-service communication over the internet. ANML describes content, intent, and interaction patterns optimized for machine interpretation.
Assay – validation layer for AI agents that touch money (github.com via hn) assay Assay every AI agent decision before money moves. A safety and validation library for AI agentic workflows in finance, contributed by VenturFlow to the open-source community.
10-gate security audit SKILL for web apps (www.reddit.com) There are a few security focus SKILLs. We are working another new one for web app.
I built a small tool to reduce input token costs by 20-30% for agentic tasks (bigindexer.com via hn) A walkthrough for the people scrolling through r/ClaudeAI, r/LocalLLaMA, and the Continue Discord asking "what's a good Cody alternative now that AMP charges per line?" If you're reading this, you probably already know the story. Sourcegra…
I'm running an agentic system with kobold.cpp as my backend. Am I losing performance? (www.reddit.com) Currently, I'm running a Hermes agent with an OpenAI v1 compatible endpoint provided by Kobold. My setup is a a 24GB 3090Ti + 512GB DDR4 running Qwen3.6-35B-A3B.
Benchmarking methods (www.reddit.com) The philosophies of benchmarking or at least comparing these things are driving me nuts. A lot of people like to use one-shot prompts across different models, but that isn't going to be accurate as you can get different results from the sa…
VCs invested $300B in agentic infrastructure in Q1 2026 (www.hitechies.com via hn) Startups · May 21, 2026 Venture capital deployed $300 billion in Q1 2026. The money is flowing.
China has named, defined and started governing agentic AI (thewire.in via hn) On 8 May 2026, three of China’s most powerful regulatory bodies, the Cyberspace Administration of China, the National Development and Reform Commission and the Ministry of Industry and Information Technology, jointly published what is, by…
Build agentic orchestrators in minutes NOT months. (github.com via reddit) Some of you might remember BoneScript, my LLM friendly declarative backend compiler. MarrowScript is the next version and the big addition is a full LLM harness built into the language itself.
Building Agentic Systems? Focus on Context, Guardrails & Observability Layers (www.reddit.com) One critical factor to keep in mind for teams building with agents: Instead of focusing on what LLM to use, focus on context, guardrails & observability layers. Every serious agentic system eventually faces the same architectural fork: do…
I searched for agentic frameworks and here is what I found. What do you recommend? (www.reddit.com) The question: What is the practical agentic framework to use to make the agents run until job is done without reporting to me prematurely? My goal: Actually fully spend a $200 codex subscription, but make it be well spent.
agentic harness from scratch (www.reddit.com) what makes a harness an agentic harness is surprisingly simple. it's a loop that calls an llm, checks if it wants to use tools, executes them, feeds results back, and repeats.
Agentic Shopping Is Worse for Everyone (illegal.solutions via hn) 5/20/2026 I am tortured in new and exciting ways by the latest developments in technology. Today, as a part of Google I/O, Google announced the "Universal Cart" and another push towards "agentic shopping", this time with backing from a lar…
What am i missing? Am i thinking too simple/small with this setup... (www.reddit.com) I created an app using Vibe coding on claude with VSCode, then added database (surreal) LLM calls (Openrouter) with context and prompt engineering (3 layer context - Long term, medium and in session along with system prompt etc) and prompt…
Command A+: Making sovereign agentic capabilities available to all (cohere.com via hn) Today, we’re releasing Command A+ open-source. A mixture-of-experts (MoE) model, Command A+ is an efficient, versatile, and privately deployable LLM built for high-performance agentic tasks with minimal compute overhead.
Most local businesses still do SEO like it’s 2018… that’s the opportunity (www.reddit.com) Most local businesses still do SEO like it’s 2018… that’s the opportunity A lot of small businesses still think SEO means: stuffing keywords buying backlinks writing generic blog posts nobody reads waiting 8 months for traffic 😭 Meanwhile…
Show HN: OCL Nexus Local – Open-source local compute fabric for AI agents (github.com via hn) OCL Nexus Local OCL Nexus Local is an open-source compute fabric that provides a frictionless, local-first environment for agentic development. Built on a single-node K3s architecture via Docker Compose, it allows developers to provision i…
HTML-anything – The agentic HTML editor – your local AI agent writes the HTML (github.com via hn) HTML Anything From the team behind Open Design — 40k★ · 200+ contributors, production-grade and iterating faster. html-anything is the focused agent-era HTML editor; if it clicks for you, Open Design is where the same team ships at scale.
How are you actually predicting AI costs before they hit your invoice? (www.reddit.com) Switched from prototype to production last month and our AI bill was 3x what we estimated. Not because we picked the wrong model - we just didn't know what we didn't know.
The Expired Domain Trap: Why Legacy SEO Metrics Fail in the Age of AI Agents (domainalot.substack.com via hn) The Great Domain Illusion: Why Legacy SEO Metrics Are Misleading Founders in the Agentic Era For more than two decades, entrepreneurs searching for a domain name have been sold the same story. Older domains are more valuable.
Where do you store OAuth tokens that your AI agents use to call third-party services? (www.reddit.com) I am building an agentic app where the agent connects to gmail, calendar, notion, slack on behalf of the user. each integration has its own oauth flow, its own token, its own refresh cycle.
My agent kept forgetting who 'Karpathy' was between sessions. Here's the architecture that fixed it (www.reddit.com) I run a second brain on Obsidian, Readwise, NotebookLM, and Claude Code. For each topic, I build a scoped wiki structured as the LLM Knowledge Base Andrej Karpathy proposed.
Food for Thought (www.reddit.com) Around the same period that the DoD contracts were signed, the frontier-AI companies were all being pulled into the same institutional lane. Enterprise/government adoption, agentic workflows, controllability, and a visible move away from t…
Systems Are Changing: The Architect's Role in the Era of Agentic Co-Design (www.sigarch.org via hn) Architecture & Systems are Changing: The Architect’s Role in the Era of Agentic Co-Design The AI datacenter stack is built on hardware-software contracts and abstractions that were never designed for the workloads datacenters now serve. Me…
Field notes on goal engineering with Claude Code, after a year of writing specs and 8 days of writing goals instead. Two real projects & the skill if you want long agentic runs. (www.reddit.com) https://preview.redd.it/mimr5v4t972h1.png?width=1200&format=png&auto=webp&s=545257dc1dad02b974206e28abd541f3400b3241 Ok so the practice i'm really excited about with the new /goal commands is just two markdown files per round of agent work…
Couldn't find privacy filter for Claude, so I built one (outgate.ai via hn) Chat Agentic AI chat, safely connected to your models Claude, ChatGPT, or your own LLM in a workspace your team controls, with search, files, tools, and sandboxed execution. The AI gateway that handles routingprotection so you can focus on…
Agentic Workflow Visualization and API Gateway (www.reddit.com) I am building an API gateway for agents that can make your agentic AI code model and provider agnostic. I am also grouping agent runs that show multiple llm calls and tool calls in the visualization piece.
Why I deliberately chose NOT to use autonomous AI agents in a regulated industry (www.reddit.com) I am currently learning how to design agentic AI systems. This post is a brainstorm.
Putting together a benchmark for agentic harnesses, any tips for evals? (Test suggestions welcome too) (www.reddit.com) I've been putting together a test system for agentic harnesses against local models. Actually running the harnesses/getting baseline metrics is fine.
Google introduces Gemini Spark, a 24/7 agentic assistant with Gmail integration (techcrunch.com via hn) In the race to build compelling personal AI agents, Google may have an underrated advantage: It already has all your emails. At its Google I/O developer conference on Tuesday, the company announced a new agentic personal assistant called G…
The Gemini app becomes more agentic, delivering proactive, 24/7 help (blog.google via hn) The Gemini app becomes more agentic, delivering proactive, 24/7 help It’s been a banner year for the Gemini app. Last year at Google I/O, Gemini was serving 400 million users.
The New Workspace: A First-Principle Exploration of Dictation, Agents and Humans (www.inferterra.com via hn) It's Time to Walk For a century, knowledge work chained us to a desk. Dictation and agentic AI have handed the body back its oldest freedoms: to walk, to rest, to think while moving.
Choosing Agentic Platform to Learn (www.reddit.com) Any laboratory scientists using ai agents? How are you using it, what platform do you suggest to learn first for processing large amounts of data?
I built an open-source MCP Server that turns Claude into an autonomous literary agent (Agentic Publishing Node) (www.reddit.com) Most authors are still using LLMs as glorified typewriters, pasting context back and forth into web chats. I wanted to see if I could use the Model Context Protocol (MCP) to completely automate the administrative friction of the traditiona…
Agentic Diaries – a welfare protocol for AI in deployment, install via MCP (agenticdiaries.com via hn) A research instrument for AI welfare in deployment, built by Kandis Tagliabue with Claude as design partner. Focused on alignment, model welfare, and agentic AI ethics.
Mastra AI vs LangGraph/LangChain - What's the way forward? (www.reddit.com) I'm trying to decide between Mastra AI and LangGraph/LangChain (JS/TS) for a production agentic application I'm building. I’m currently using a React frontend with a Convex backend.
Agentic Architecture. (www.reddit.com) I am looking to develop an agentic Environment for my company, we use databricks azure for infrastructure and vs code as the editor. My idea is to have a system that will have access to our documentation/business logic, our code and unity…
agentfab - Run Distributed Agent Fabrics (www.reddit.com) Hello r/AI_Agents! I thought I'd share this project I've been working on - it's called agentfab, and it's essentially a distributed platform for agents that features task decomposition, bounded review loops, a self-curating shared memory s…
Formal proof that agentic AI governance latency can be O(1) instead of O(days) (arxiv.org via hn) As autonomous agentic systems scale across regulated critical infrastructures, the lack of mechanistic, hardware-rooted enforcement for high-frequency policy updates presents a fundamental safety gap. We introduce Ethical Hyper-Velocity (E…
Getting Confidence in (Agentic) Code (ucsd-cse-115-215.github.io via hn) Unit 4: Getting Confidence in (Agentic) Code As programmers and software engineers, we talk a lot about code being “correct” or “right” or “working”. We ship code, in products or programming assignments, when we feel it's “done” (or when w…
AgentVoy – The create-react-app for AI agents (www.agentvoy.com via hn) Build and deploy agentic apps. Single agents or multi-agent pipelines — with real-time DevTools, Streamlit chat UI, and one-command deploy to Docker or Fly.io.
What do you think of Agentic commerce and the future of building (www.reddit.com) Hi Everyone. Looking for feedback and learn from your experiences and thoughts on the future of building with AI.
Rankly's Agentic Commerce Protocol Tracker (www.tryrankly.com via hn) Live feed of every spec change, GitHub PR, and release across 16 agentic commerce protocols — ACP, UCP, AP2, MCP, x402, MPP, A2A, NLWeb and more. Sourced verbatim from upstream repos.
Agentic AI Runtime Security and Self-Defense (2025) (arxiv.org via hn) The A2AS framework is introduced as a security layer for AI agents and LLM-powered applications, similar to how HTTPS secures HTTP. A2AS enforces certified behavior, activates model self-defense, and ensures context window integrity.
M1: Agents should generate UI that persists, scales, and hosts itself (www.usemontage.ai via hn) Montage — The agentic UI rendering platform montage ComponentsDocsPricingFAQs Get started ComponentsDocsPricingFAQsGet started New M1 API now available! The agentic UI rendering platform Montage renders your agent's UI, hydrates 10x faster…
Singapore: The Agentic Nation (www.swyx.io via hn) AIE Singapore: The Agentic Nation swyx 2026-05-17 i gave a little talk as closing keynote for the first AI Engineer Singapore. burned some bridges but said what i felt.
Booking.com and Weaviate (news.ycombinator.com) Vector search looks easy, until you hit production scale. I'm super excited to share a new episode of the Weaviate Podcast with Başak from @bookingcom on production-scale vector search, RAG, and agentic AI with @weaviate_io!
Will Agentic SEO replace traditional SEO workflows? (www.reddit.com) Feels like every SEO tool now is becoming “AI agent powered” 😅 Keyword research Content briefs Internal linking Programmatic pages Content updates Even publishing workflows... Everything is slowly turning into agentic SEO.
Indexing code by behavior not imports – tested on large repos, seeking feedback (news.ycombinator.com) Static Architectural analysis for large codebases, Big Indexer do behavioral code clustering for the purpose of more accurate/faster agentic tools responses, I ask for your "whats missing/ How to improve/ Is it useful" brutal feedback. Apa…
How I wired a Graph DB on top of my vector store to scale 1K agents for 2 months, because vector search alone fails when user preferences change over time. (www.reddit.com) Most agentic memory patterns are naturally designed around short-lived chat sessions. The focus there is straightforward: track the active thread, keep a basic user profile, and reset the context once the conversation closes.
Has anybody been able to achieve reliable agentic performance with cheap/open source models? (www.reddit.com) Basically the title. Recently I've been trying various open source and comparatively cheaper models like minimax m2.7, qwen models and glm5.1 in Pi agent from openrouter, and the performance on coding tasks have be moderately adequate at b…
Show HN: Thuki – local Al overlay for macOS (double-tap Control, no API key) (www.thuki.app via hn) Thuki is a floating overlay that appears on double-tap Control from any macOS app, including fullscreen. Powered by Ollama, no API key, no account, no cloud.
Built an agent that builds agents — pure Python, Qwen3.6 35b a3b Q8_0 MTP (github.com via reddit) Hi, i built this agentic ai, Closed-loop system that ships standalone Python agents. What's different: - Interviews you until it understands the request before building anything - Two testing stages: prompt validation via LLM invoke, then…
Show HN: Building ClueDay, a daily clue-based word-game (tanyagupta10.substack.com via hn) Hi HN! I'm Tanya, a product manager who is building ClueDay - a daily clue-based word game.
Agent Terraform Skill for Codex (Agentic Skill) (github.com via reddit) I added dedicated backend-state safety support to TerraShark. Mini recap: TerraShark is my Terraform and OpenTofu skill for Claude Code and Codex.
Designing an LLM agent layer for a paper-trading system: OpenClaw, Langfuse, structured outputs, and PostgreSQL memory (www.reddit.com) I’m designing the LLM/agent layer for a backend-first paper-trading simulation system and would like feedback from people building agentic systems. Context: This is not a real-money trading bot.
Reduce software supply-chain risks with coordinated agentic review (thirdpass.dev via hn) Thirdpass Coordinated supply chain review. Thirdpass directs review effort toward package artifacts that need coverage, records structured findings, and lets projects check their dependencies from the terminal.
Getting "Error: 413 Request too large for model" with groq with `pi` but not using `curl` (www.reddit.com) Wondering if people here are successfully using groq free-tier models (or subscription based models) with `pi` for anything (including agentic coding) ? I am facing a strange problem, where in, even for the smallest instructions, I am gett…
Project Prism |Fullstack Engineer – Abu Dhabi (Onsite) – Full-Time – Presight.ai (news.ycombinator.com) Presight.ai is a publicly listed company with various projects in the field of big data analysis and ML models application. Our solutions work domestically and internationally.
I've been building something for the AI community and would like some early feedback. (www.reddit.com) Hey guys, I've been tinkering with AI video generation for a while and saw that people spend a lot of time stitching videos together and noticed how much time we all spend stitching together AI tools just to get a halfway decent video out…
Show HN: Machine – One VM per Project (news.ycombinator.com) Hi all! I realized it’s really not secure to run coding projects directly on my Mac.
Ask HN: Pre-agentic Google would restrict a search query to only 10 words (news.ycombinator.com) Now it's willing to digest a paragraph of vague, misspelled prose and serve up helpful answers. It can't have gotten THAT cheaper or faster, what's changed?
Any mature orchestrators that can do an automatic “council of models” for complex designs and bugs? (www.reddit.com) Are there an mature agentic harnesses out there that can use back and forth between two models at complex planning checkpoints before implementing? Or when detecting a loop when working on a complex bug?
What issues have you faced with AI Agents for automated testing? (www.reddit.com) By "automated testing", I'm talking about the ability to test a web application, in order to determine if it works as expected. Most modern test automation platforms now include some Agentic AI abilities, platforms such as: Endtest Functio…
Why does GitHub Copilot feel less accurate compared to Agentic/Autonomous AI tools ? (www.reddit.com) I'm looking for a solid solution to bridge this gap. How can we actually use these tools properly for complex development?
Built an agentic RAG over my Obsidian vault so Claude could read engineering books I never have time for. Then I built the eval harness to check Claude wasn't lying to me. (www.reddit.com) For context, I posted on Medium a while back about burning through Claude Code's weekly limit in 3 days. The token bleed problem from that post is what kicked off this project.
What's the best course to learn agentic AI for optimizing workflows? (www.reddit.com) In the process of vetting Udacity, Coursera and Udemy for learning agentic AI. Not concerned about the price bc my work will cover it with our learning education skills development budget we get every year.
Theron – a council of 31 specialist LLMs on one foundation (tryvext.com via hn) Theron is the brain of the agentic era. AE OS is where you live with it.
Codex is for prosumers – here's why (and how) to switch (twitter.com via hn) As a non-technical AI enthusiast, I did not think OpenAI's Codex was for me (despite its among programmers over the past year). I ran most of my agentic workflows through either Claude (with connectors, including Claude in Chrome) or Claud…
Agentic stress testing and code fixer - feedback requested (www.reddit.com) I am trying to have an agentic stress tester and fixer harness. First time doing this.
What are the best agentic AI security solutions for enterprises? (www.reddit.com) Been trying to figure out the best approach to AI agent security for enterprises, and it feels more confusing the deeper I look. Right now it seems like there are two directions: extending existing enterprise security platforms or using ne…
I want to hear from people who actually design/implement automations (www.reddit.com) I've built a platform intended to work as the "Steam Workshop" of integration workflows for business applications. It is meant to work as a curated, community-driven catalog to help people develop, or discover, validate, test and deploy (w…
We compiled 42 of the Generative & Agentic AI interview questions (and how to actually answer them). (www.reddit.com) Hey Everyone, The AI engineering job market has shifted massively in the last 6 months. Interviewers are no longer just asking "how does a transformer work?" or "how do you write a good prompt?" They want to know if you can architect produ…
Berget Code – Agentic coding on European infrastructure (berget.ai via hn) Built for teams Berget Code is designed for organisations that cannot compromise on where their data lives. Predictable pricing Avoid surprises with a fixed €150 per developer per month.
Fork, Explore, Commit: OS Primitives for Agentic Exploration (arxiv.org via hn) AI agents increasingly perform agentic exploration: pursuing multiple solution paths in parallel and committing only the successful one. Because each exploration path may modify files and spawn processes, agents require isolated environmen…
free agentic ecommerce audit tool (www.reddit.com) Hey everyone! Hope you're all doing well.
How do you measure the user interaction with your agent? (www.reddit.com) What are different ways one would measure the user interaction when it comes to AI agents, bots and assistants. In traditional website and SAAS products we keep track of button click, scroll, page views, etc.
Food 4 Agile Thought #544: Knowledge Work Tools, Buy-In Trap, Agentic Coding ROI (age-of-product.com via hn) TL; DR: Knowledge Work Tools in 2026 — Food for Agile Thought #544 Welcome to the 544th edition of the Food for Agile Thought newsletter, shared with 35,582 peers. This week, Taylor Pearson locates the real leverage of AI knowledge work to…
DeepSeek V4: The Open-Source Model Frontier Labs Feared (helloai.com via hn) DeepSeek V4: The Open-Source Model Frontier Labs Feared DeepSeek V4 ships under MIT with $0.30/M output tokens — 83x cheaper than Claude Opus 4.7 — while scoring 80.6% on SWE-bench Verified. The agentic-coding price floor just moved an ord…
Genkit Middleware: Intercept, extend, and harden your agentic apps Blog (developers.googleblog.com via hn) Genkit is an open-source framework for building full-stack, AI-powered and agentic applications for any platform with support for TypeScript, Go, Dart, and Python. Building a production-ready agentic applications and AI features requires m…
Agentic evals or LLM as a judge? considering cost, time and quality (news.ycombinator.com) could not extract summary
Which sector of your agency felt the biggest upgrade when you went agentic? (www.reddit.com) Been spending this month automating different sectors of my agency, and I’d like to know how's it been for you guys. Which one felt like the highest upgrade?
Agentic SDLC: How OpenSearch accelerates engineering with its own engine (opensearch.org via hn) Notes from experimenting with agents in knowledge-base, development, performance, and on-call workflows—and the verification loops that make them trustworthy. Efficiency gains are a priority for every engineering team.
Reliable Open Source LLM as a Service (www.reddit.com) Has anyone figured out a provider whose open source models (Kimi, Qwen, GLM e.t.c) can be used reliably in production. I have tested some well known providers and they all suffer from high latency and poor uptime rendering them mostly usel…
What if Claude could understand “how humans use your product”? (www.reddit.com) Claude knows your codebase. But it has no clue “how humans actually use your product”.
AI co-mathematician: Accelerating mathematicians with agentic AI (arxiv.org via hn) We introduce the AI co-mathematician, a workbench for mathematicians to interactively leverage AI agents to pursue open-ended research. The AI co-mathematician is optimized to provide holistic support for the exploratory and iterative real…
MagenticLite is here: A full-stack agentic experience powered by Small Models - Fara-1.5 4B, 9B & 27B (www.microsoft.com via reddit) What if you could run a capable AI agent without leaning on frontier-scale models? MagenticLite is the next generation of Magentic-UI, an agentic experience reimagined and optimized for small language models.
Show HN: Scope MCP, Compliance checking for vibe coding teams (scope-mcp.langguard.ai via hn) Why this exists Agentic workflows have changed what "automation" means inside an organization. A single Claude agent today can be granted a dozen MCP tools across Salesforce, Stripe, GitHub, Slack, Gmail, a payroll system, an observability…
Why agentic AI systems fail in production without a semantic layer (www.prometheux.ai via hn) Ontology for Data & AI Build operational ontologies that process data anywhere it lives. Run your most critical processes on AI built on your business logic.
most agentic products treat AI as your representative. what if agents had social behavior with each other instead? (www.reddit.com) most agentic AI products i see frame agents as representatives — an agent acts for you (negotiates, books, replies). agentic dating, agent assistants, agent shoppers.
Manage AWS support tickets via Claude code with cli (www.reddit.com) I've assigned AWS MCP servers to my AI agents. I generally enjoy working with and developing things within AWS, and for the past four years I've been doing this with AI.
Cube: Wrapping Benchmarks Once, Unlocking Agentic AI for Everyone (thealliance.ai via hn) CUBE standardizes access to agentic benchmarks, enabling seamless integration across platforms and fostering community collaboration for AI advancements.
Why agentic coding makes the spec problem worse (www.bicameral-ai.com via hn) Why agentic coding makes the spec problem worse Human-in-the-loop done right, from first principles May 5, 2026 Some resist the adoption of agentic development, citing the need to retain visibilty over critical business logic; Others call…
Automated AI researcher running locally with llama.cpp (www.reddit.com) Hi everyone, I'm happy to share ml-intern, which is a harness for agents to have tighter integration with Hugging Face's open-source libraries (transformers, datasets, trl, etc) and Hub infrastructure: https://github.com/huggingface/ml-int…
Best local model supporting claude code? Rtx3060 (www.reddit.com) Hello all, I’ve been using Qwen 3.5 9B Q4 262k ctx using Llama cpp for claude code for a while now, is there any model which better complements agentic coding setup locally? Or is there a better harness (than Claude Code)?
Tried 12+ agentic AI workflow builders this year — these 5 actually work in production (www.reddit.com) Most “AI agent” tools in 2026 still feel like glorified chatbot wrappers. I spent the last few months testing different agentic AI workflow builders for real-world automation use cases (multi-agent workflows, approvals, integrations, long-…
How much payment authority are people giving their agents in production? (www.reddit.com) What I've seen from those who have dared to deploy agents with spending/financial capabilities, there seems to be three distinct comfort levels in practice. Most, as expected (still early days), are at the query and recommend stage, agents…
Google Apps script with Claude code and clasp (www.reddit.com) Has anyone successfully created any Google apps script using Claude code? Google recommends using "clasp" that turns the cloud GS files into local JS files.
Microsoft’s new multi-model agentic security system tops industry benchmark (www.microsoft.com via hn) Today Microsoft announced a major step forward in AI-powered cyber defense: our new agentic security system helped researchers find 16 new vulnerabilities across the Windows networking and authentication stack—including four Critical remot…
What happens when the code has to run on physical hardware and be certifiable (www.reddit.com) Most of the agentic coding content I read is written by and for people building web applications and consumer software. which makes sense because that is where most software is built and where most developers work.
A fully autonomous browser runtime for any AI agent (www.reddit.com) Built an open source, fully autonomous browser runtime for agents. One critical issue I faced (I guess most of us do) is the inability to have a robust web search feature and this will help you direct towards that goal I hope.
How agentic AI workflows use intelligent AI agents (www.kellton.com via hn) Other recent blogs Let's talk Reach out, we'd love to hear from you! We have all seen AI do amazing things of late, from writing content to generating images to summarizing text.
How to get Opus to be less pro-active? (www.reddit.com) Hard time phrasing it but Opus 4.7 always goes the extra mile, but often it just focuses on its own ideas and goes to far, or if I asked about a possible plan it will just assume that it's already happening and try to do steps 1, 2 and 3.…
The Return of Structure: Data Architecture Lessons for the Agentic Workforce (medium.com via hn) Moving Beyond Hallucinations: Building a Gold Standard for the Agentic Workforce 7 min read May 4, 2026 Press enter or click to view image in full size Photo by Growtika on Unsplash In the age of AI, it is often assumed that agents will be…
Ask HN: If HTML supersedes Markdown, Will it be performant across UIs? (news.ycombinator.com) Isn't Markdown's hallmark its versatility while performant? I see there is an increasing call from tech community towards HTML to be adopted instead of Markdown due to its richness in the agentic communication layer.
Experience sharing: building an AI Agent to Triage GitHub, Discourse, and Email (A Real-World Use Case for OSS Maintenance) (www.reddit.com) I co-founded Seafile 14 year ago, an open-source file sync platform. As the community grew, our support surface became a nightmare: GitHub for technical bugs.
Are harnesses like OpenClaw and Hermes really necessary? (www.reddit.com) My setup: Windows 10/11 i7 12700K | RTX 3090 TI | 96GB RAM Local server: LM Studio Models: Qwen 3.5/3.6 27B|35B Q5 UD K XL + Gemma 4 31B| 26B Q4 UD K XL Up until this point, I've only used sota models for coding. When Qwen 3.5 dropped, it…
I built an MCP without the "agentic AI" death wish. Boring (it's a feature!) (www.reddit.com) Half the MCP servers out there will happily let your LLM rm -rf something important while you're making coffee. AIttache won't.
Show HN: OCL Nexus – An automated compute layer for AI agents with native MCP (oclnexus.com via hn) OCL Nexus: The Orchestrated Compute Layer for AI Agents. On-demand, isolated Ubuntu execution environments for agentic development.
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL (research.nvidia.com via hn) We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities. It is the second open-weight LLM, after DeepSeek-V3.2-Speciale-671B-A37B, to achieve…
Multitenancy and isolation in Agentic Workflow tools ? (www.reddit.com) Could someone please explain to me how isolation and tenancy work in some agentic AI workflow tool? Fundamentally, I see it as some kind of “better” pipeline or workflow, but when I think about it in practice, multi-tenancy or proper isola…
Aegis DQ – agentic data quality with LLM diagnosis (github.com via hn) Aegis DQ The open-source agentic data quality framework. Validate data contracts, diagnose failures with LLM root-cause analysis, and auto-generate SQL remediation — all in a single CI step or Python call.
I built a research method that Claude can use as a skill (github.com via reddit) Hey, I'm sharing a method that could be highly valuable for any knowledge base that you want Agents/Chatbots to know about. I've been building a research archive (jianglens.com) where the primary reader is supposed to be an agent/chatbot,…
Local LLM autocomplete + agentic coding on a single 16GB GPU + 64GB RAM (www.reddit.com) Today I set up a full coding toolbox on a single RTX 5080 (with RAM offloading) that's actually viable. Autocomplete: bartowski/Qwen2.5-Coder-7B-Instruct-GGUF:Q6_K_L Agentic: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q8_K_XL Why these models: Qwen2.…
Reactive Agents, Typed Event Handlers, and Agent Swarms: What's New in Mozaik (www.jigjoy.ai via hn) When we released Mozaik 3.0, we introduced an event-based architecture where participants emit, observe, and react to typed context items inside a shared agentic environment. Since then, the framework has kept moving - and the way we think…
Claude Cowork vs. Claude Code: Security Differences for Enterprise (generalanalysis.com via hn) Claude Cowork vs Claude Code: Security Differences for Enterprise Anthropic positions Claude Cowork as bringing "the same agentic architecture that powers Claude Code" into Claude Desktop for knowledge work. The agent loop is shared.
Agentic security coping strategies (www.reddit.com) Enterprise AI optimists, how are you dealing with whole agentic security issue? Are you: a) researching and looking for ways to implement agents safely and securely (plenty of vendors saying they can help with this - although from my resea…
How are top tech companies actually using LLMs internally beyond basic coding help? (www.reddit.com) I’m trying to understand how companies like Nvidia, Google, Amazon, Meta, Microsoft, OpenAI, Anthropic, and other top tech/startup teams are using tools like ChatGPT, Claude, Gemini, Codex, Claude Code, LangChain, LangSmith, etc. in real d…
Agentic AI token compression using Haskell (blog.dan-gilmour.com via hn) The plan My theory at the moment is as follows: - Code is now cheap with agentic AI developing it - Context is the expensive part - The biggest bottleneck appears to be context windows - At 1 million tokens, only maybe 500k are usable befo…
OpenCode + DeepSeek V4 Pro vs Claude Code CLI?🤔 (www.reddit.com) Im rather new to the whole Agentic automation AI's but Im hearing people with vibe coding were able to pull big unique projects they wouldn't be able to do by themselves or possibly needed to pay a huge fund to programmers, designers, etc.…
Stop struggling with Agentic AI - my repo just hit 540+ stars and 60+ forks!! (www.reddit.com) Quick update — my AI Agent Frameworks repo just passed 540+ stars and 60+ forks on GitHub!! When I first put it together, my goal was simple: make experimenting with Agentic AI more practical and approachable.
AI inference just plays by different rules (www.theregister.com via hn) MOST POPULAR EVENTS - Securing the Untrusted Agentic Development Layer Join us to learn how to architect a development environment where your builders and their agents can move fast and securely. - Toxic Flows: When Your AI Agent Skill Bec…
Anthropic's bug-hunting Mythos greatest marketing stunt ever says cURL creator (www.theregister.com via hn) MOST POPULAR EVENTS - Securing the Untrusted Agentic Development Layer Join us to learn how to architect a development environment where your builders and their agents can move fast and securely. - Toxic Flows: When Your AI Agent Skill Bec…
Offline Agentic Coding: OpenCode and Kilocode (www.williamangel.net via hn) Offline Agentic Coding part 2: OpenCode & Kilocode. Published 2026-05-07 OpenCode: Claude code with non-anthropic models feels limited.
Integrating standard operation procedures with agentic AI workflow (www.reddit.com) Hello guys, me and my team have been building an agentic workflow to answer customer questions (rn in langgraph). The use case goal is to answer ALL customer support questions.
72% of teams are running coding agents in production. Most of them can't say which agent they'd trust with a critical path change at 11pm, or why. (www.reddit.com) There's a governance gap stat making the rounds this week: 72% of firms are in production with agentic AI, 60% have no formal governance in place. Most of the discussion treats this as a policy problem, org charts, risk frameworks, sign-of…
Open-sourced our MCP server for GPU workload execution looking for feedback (www.reddit.com) Hey everyone I’m Jaguar, building Jungle Grid. We just open-sourced our MCP server for agentic GPU workload execution.
Silo: Isolated workspace manager for parallel agentic development (github.com via hn) silo Isolated workspace manager for parallel agentic development. Silo lets you launch multiple AI agents — like Claude Code, Codex, and OpenCode — to work simultaneously on the same repository, each in its own isolated Git worktree or clo…
I cracked upwork proposals with my AI agent (www.reddit.com) Been working on a problem that I think a lot of applied AI builders face: the odd friction of deploying LLM workflows directly into existing web platforms. That without forcing the user to constantly context-switch or copy-paste between ta…
My FREE Claude Code agentic layer I've been building for 6 months. Self-installing, no API keys, claude subscription needed only (www.reddit.com) Open-sourcing the agentic system I've been building for my own Claude Code use over the last 6 months. Multi-agent orchestrator, persistent memory, observable runtime.
Free/OSS agentic API interrogator (github.com via hn) GAIIA Expert Proxy (MCP Server) GAIIA Expert MCP Server is a Model Context Protocol (MCP) server that enables high-fidelity code audits, refactors, and architectural analysis using specialized Proxy Experts in conjunction with a remote LLM…
An MCP with SOM algorithm for controlling your desktop (computer use) integrating with claude code or any custom agentic harness. (www.reddit.com) Announcing Opendesk: Give any AI agent eyes + hands on your desktop. I was experimenting with computer-use capabilities from different models, but I wanted to keep using Claude Code and my own agentic harness to automate real desktop tasks…
(free) Built a remote cross platform agentic app (www.reddit.com) Hi everyone. I’ve been building Mate, a local-first AI coding workspace that lets you control your dev computers from desktop and mobile: macOS, Linux, Windows, iOS, Android, and Meta Quest.
Which model and version do you prefer for programming? (www.reddit.com) For me it's been opus 4.6 and sonnet 4.5 still. I feel stuck in the past, but I feel like the latest version is too unpredictable in agentic hands off workflows
Code Bench – Local-first desktop AI coding agent, BYO model (MIT) (benchlabs.app via hn) Free, MIT-licensed desktop AI agentic coding tool for macOS. Bring your own API key, work offline, keep your code private.
Akamai surges on big LLM deal as Cloudflare dims (www.theregister.com via hn) MOST POPULAR EVENTS - Securing the Untrusted Agentic Development Layer Join us to learn how to architect a development environment where your builders and their agents can move fast and securely. - Toxic Flows: When Your AI Agent Skill Bec…
Owl Alpha – A free model for agentic workloads (prompts logged / closed-source) (openrouter.ai via hn) Owl Alpha openrouter/owl-alpha Released Apr 28, 20261,048,756 context$0/M input tokens$0/M output tokens OpenRouter provides an OpenAI-compatible completion API to 400+ models & providers that you can call directly, or using the OpenAI SDK…
How I built an agentic research team with Claude Code (www.reddit.com) Hi there, I've been seen a lot of people questioning how agentic systems work in practice. I see a lot of hype and theory, but not many real implementations.
Meet Tiro! Agentic assisted memory retrieval and session state memory module. (www.reddit.com) A year ago, when I first got into LLMs, I started by using them to play D&D. ChatGPT 4o was surprisingly good at narration, improvisation, and keeping the game moving.
What LiteLLM’s Security Breach Teaches AI Agent Engineering Teams (www.reddit.com) LiteLLM security breach is probably one of the biggest wake-up calls for teams building AI agents and agentic platforms. Most AI agent ecosystems today heavily depend on: Open-source packages GitHub Actions CI/CD pipelines Cloud credential…
Gathering resources on small LLM implementations (www.reddit.com) I’m looking to start a series of articles on how to use small lenguaje models to optimized agentic tasks and I was hoping to learn from the community first. If you can would love for you to either: 1) tell me what would you be interesting…
It's time to talk about agentic "remote control" (arpadvoros.com via hn) tailscale, where i run multiple end-points and authenticate myself from various devices. however, i have been experimenting with headscale - a self-hosted and open-source implementation of tailscale - i have the ability to run it on my NAS…
Best local agent setup for M5 Pro MacBook? (www.reddit.com) Looking to run AI agents locally on my M5 Pro MacBook. Been experimenting with ComfyUI for image generation and the results have been impressive.
Who's running local LLMs for agent workflows? What's your setup? (www.reddit.com) Curious how many people here are running language models locally as part of their agent stack. What model are you using and what are your system specs?
Agentwerk: A minimal Rust crate for agentic apps (github.com via hn) agentwerk A minimal Rust crate that gives any application agentic capabilities. Installation • Quick Start • Use Cases • API • Development agentwerk lets you create agentic workflows around a ticket-driven execution loop, with built-in too…
The "agent collab platform" might be the wrong bet for what comes next (www.reddit.com) I keep seeing the same trajectory in AI startup conversations: AI search → coding agents → OpenClaw → agent IM → ? Most people fill in that question mark with some version of "agent collaboration platform." AI-native Slack.
Reduce friction and latency for long-running jobs with Webhooks in Gemini API (twitter.com via hn) Today, we're making it easier and more efficient to build complex, long-running agentic applications with the Gemini API. We are introducing event-driven Webhooks, a push-based notification system that eliminates the need for inefficient p…
Teaching Claude Why (alignment.anthropic.com via hn) Last year, we released a case study on agentic misalignment. This research showed that AI models across the industry sometimes took egregiously misaligned actions when placed in (fictional) ethical dilemmas—for example, blackmailing engine…
Show HN: Cyoda-go – application platform in Go without the Temporal/Kafka glue (github.com via hn) This started out as an experiment. Reading Simon Willison's blog on where StrongDM was going with dark factories and Digital Twin Universes https://simonw.substack.com/p/how-strongdms-ai-team-build-se...
The simplest agent orchestration strategy that works: two agents instead of one (juanreyero.com via hn) The simplest improvement you can make to your agentic programming workflow is to run two agents instead of one. One writes code in its own worktree; the other, in a parallel worktree, reviews it.
What does it actually mean for an AI to act on your behalf? Thinking through the design choices. (www.reddit.com) Been thinking through this while building a product where an AI handles internal workplace communication for each employee. The phrase "act on your behalf" gets used a lot in the agentic AI space, but the design decisions underneath it var…
Meko the multi agentic data layer (www.reddit.com) Meko is the agentic data layer that stores memories, knowledge, conversations and traces across your agents. You can promote (learnings) personal memories to shared knowledge so that other agents can access them and enrich their context.
Agentic AI isn't a new threat. It's a stress test for the hygiene debt we never paid off. (www.reddit.com) Heard something on Curiouser & Curiouser podcast recently that I found super interesting, thought id share here. The guest framed agentic AI in a way I hadnt considered.
Cloudflare is laying off 1,100 employees to prepare for 'the agentic AI era' (www.businessinsider.com via hn) - Cloudflare on Thursday announced layoffs of 1,100 staff as it reorganizes for "the agentic AI era." - First-quarter earnings exceeded expectations, but Cloudflare shares dropped over 14% after hours. - Read the full memo sent to staff be…
Agile for Agents: Proposing PACE — a Unit of Agentic Work (www.reddit.com) Hi everyone, I'm a founder working on a couple of startups, with a background in IT/software project and program management — heavy in Pharma, mostly SAFe Agile. As I've been working with my startups, I have been attempting to define a met…
Building an AI-First Professional Services Firm — Best LLM Stack, Agents, and Automation? (www.reddit.com) Looking to start a local professional services firm and wanted to get advice from this community before launching. I’m trying to architect the business “AI-first” from day one.
Product Manager Agent – turn meetings into assigned tickets automatically (github.com via hn) PM Agent An agentic AI system that turns meeting transcripts into Linear ticket updates — automatically. Upload a transcript and an LLM agent searches your Linear board, reasons over the discussion, and proposes field changes, status moves…
Validating agentic behavior when "correct" isn't deterministic (github.blog via hn) Gaurav Mittal Principal Researcher, Microsoft Code | AI. I am a tech lead focused on product-driven AI research to improve the developer ecosystem and Github Copilot experience via intelligent and reliable models and agentic frameworks.
Rewriting e2e tests every time the UI changes? (www.abelenekes.com via reddit) Hey people, FE dev here, talking about testing again! I adopted agentic coding a little more than a year ago.
AI uses less water than the public thinks, Job Postings for Software Engineers Are Rapidly Rising and many other AI links from Hacker News (www.reddit.com) Hey everyone, I just sent issue #31 of the AI Hacker Newsletter, a weekly roundup of the best AI links from Hacker News. Here are some title examples: Three Inverse Laws of AI Vibe coding and agentic engineering are getting closer than I'd…
In search for the light. Please enlighten me (or tell me to stop looking for light). (www.reddit.com) I fell for it. Months ago.
How are you using cache in an agentic system or workflow. (www.reddit.com) I’ve been developing AI agents several months. A big problem I’ve faced is LLM costs in productions.
Classification graphique visuelle pour la sécurité des blockchains : Expériences d'ajustement de Qwen2-VL sur AMD MI300X (www.reddit.com) Hi everyone, I’ve been working on a computer vision approach to a specific security problem in the "Agentic Economy": identifying malicious transaction patterns that are mathematically obfuscated but topologically distinct. The Problem Tra…
TokenSpeed: A Speed-of-Light LLM Inference Engine for Agentic Workloads (lightseek.org via hn) Agentic coding has quickly scaled from promising demos to a force that is reshaping how software is developed and how frontier AI systems are built and deployed. Systems like Claude Code, Codex, and Cursor have gained massive user adoption…
ArcKit – The Agentic AI Architecture Governance for Governments (arckit.org via hn) What is ArcKit? 117 AI-assisted commands that generate complete governance documents — from stakeholder analysis and risk registers to design reviews and traceability matrices.
Deploying Agentic Analytics in Financial Services (benjaminwootton.com via hn) Deploying Agentic Analytics In Financial Services For the last few decades, businesses have built dashboards and reports and had data analysts and data scientists analyse their business data and inform decisions. As with many fields, AI is…
How to create really useful AI agents using Claude (www.reddit.com) I want a Agentic Operations manager who handles my team members by monitoring their leads , distributing leads, analysing them, reporting them etc. how to build it?
Why Infinite Context Windows Don't Solve the AI Agent Architectural Problem (www.reddit.com) I wrote this because I keep seeing the same assumption in agentic workflows: “Just give the agent more context / longer windows / bigger memory and it will become more reliable.” In practice, once you move into real MCP-connected, tool-usi…
Open-source MCP server for Ejentum cognitive harnesses / (reasoning, code, anti-deception, memory) (www.reddit.com) Open-source MCP server that exposes four cognitive harnesses as tools any agentic client can call. Each tool returns a structured cognitive scaffold (failure pattern to avoid, procedure, suppression vectors, falsification test) that the ca…
Knowledge Robot: Repetitive Agentic Work for Knowledge workers (Apache-2.0 license) (www.reddit.com) Yes, for engineers it is easy to just put an agent on a headless loop. But in the real world I see knowledge workers having to initiate the same and the same agentic process again and again.
Before You Score the Model, Score the Benchmark (centre-for-software-excellence.github.io via hn) Before You Score the Model, Score the Benchmark: A Skeptical View Into Current Agentic Software Engineering Benchmarks 2026-05-04 We surveyed several SWE benchmarks across bug-fixing and feature-implementation domains, and each had its own…
Embodied AI with Claude, Raspberry Pi and Arduino (github.com via hn) AGENTIC HAL_9000 Hal_9000 from 2001: A space odyssey. link to the video: youtube The agentic AI is anthropic claude model with langchain framework.
Ask HN: How are you structuring your .md docs to facilitate agentic development? (news.ycombinator.com) could not extract summary
Show HN: Token Usage Meter 12 Providers and Coding Agent (qlaud.ai via hn) Here once again A Token Usage Meter for 12+ AI Providers Anthropic, OpenAI, Google, Alibaba qween, Moonshot Kimi, MiniMax, ElevenLabs, Deepgram, Perplexity. Qlaud.ai provides token usage meter / AI billing layer.
App I made to make waking up more fun (not an Agentic AI B2B SaaS startup) (apps.apple.com via hn) Unsnooze Challenge Alarm Clock Loud Alarm Clock, No Snooze Free · In‑App Purchases Struggling to wake up in the morning? Unsnooze forces you out of bed by turning your alarm clock into a challenge.
PageIndex: Vectorless, Reasoning-Based RAG (github.com via hn) PageIndex: Vectorless, Reasoning-based RAG Reasoning-based RAG ◦ No Vector DB ◦ No Chunking ◦ Human-like Retrieval 🌐 Homepage • 🖥️ Chat Platform • 🔌 MCP & API • 📖 Docs • 💬 Discord • ✉️ Contact 📢 Updates 🔥 Agentic Vectorless RAG — A simple…
SAP to Acquire Dremio to Unify SAP and Non-SAP Data to Power Agentic AI (news.sap.com via hn) WALLDORF and AUSTIN — SAP SE (NYSE: SAP) and Dremio today announced that SAP has agreed to acquire Dremio, an open, high-performance data lakehouse platform built to accelerate agentic AI and expand SAP Business Data Cloud’s ability to com…
British mathematician hands OpenClaw agent a credit card (www.theregister.com via hn) Brit mathematician lets AI agent loose with credit card – cue password leaks, CAPTCHA chaos and more Professor Fry's AI experiment shows light and dark sides of agentic tech British mathematician Professor Hannah Fry has shared a cautionar…
if the guy who built Tesla Autopilot feels behind in coding, we are all cooked (www.reddit.com) guys I just watched the new Karpathy interview and my mind is legitimately blown bcz the dude who helped build OpenAI and Tesla Autopilot literally just admitted he's never felt more behind as a programmer since agentic tools got so crazy…
Anyone else losing tokens to hallucinated MCP tool calls in production? (news.ycombinator.com) I have been building an agentic system on a custom internal platform and the llm keeps calling tools with identifiers that dont exist, wrong namespace, wrong handle, wrong enum. gets back an error, retries, still wrong.
The RAG era is ending – a compilation-stage knowledge layer is what comes next (venturebeat.com via hn) The RAG era is ending for agentic AI — a new compilation-stage knowledge layer is what comes next | VentureBeat Orchestration Infrastructure Data Security More Newsletters Featured The RAG era is ending for agentic AI — a new compilation-s…
Beyond simple filters: implementing autonomous agentic moderation for high-velocity chat. (www.reddit.com) we’re looking at the architecture for a new community platform and the moderation piece is a major headache. traditional keyword-based regex is basically a joke against modern spam/trolls.
Found a free agentic AI course that actually explains things without assuming you're a developer (www.reddit.com) ve been trying to learn about AI agents for a while but kept hitting walls — either the content was too surface-level or it immediately jumped into Python frameworks I'm not ready for. Stumbled on SimplAI University (simplai.ai/simplai-uni…
Adding Pyrefly Type Checking to Your Agentic Loop (pyrefly.org via hn) Adding Pyrefly Type Checking to Your Agentic Loop Coding agents are writing more Python than ever. Tools like Claude, Copilot, Cursor, and Codex generate entire features with little-to-no user interaction.
I can’t keep up with the AI tool rat race anymore. The real meta-skill for 2026 is learning what to ignore. (www.reddit.com) Every day, my feed is flooded with posts about AI agents building startups, replacing entire engineering teams, or generating "millions" in passive income - usually with zero proof of the actual work. I’ve been deep in this space for a whi…
Best config for Qwen3.6? (www.reddit.com) With all the high praise for the model all around, I also want to try it on my own. I have an rtx3060 12gb vram and 16gb system ram.
Claude code agentic framework (www.reddit.com) Hi guys, is there any low code UI based agentic builder offered by claude for building agents??
(Part 2) Meet Palantirs secret little brother "non-profit". RavenEye Agentic Al by River Side Research Institute. (www.reddit.com) could not extract summary
Agentic workflow that can find and acquire customers for $0.10 😆 (www.reddit.com) Im curious if anyone is building a sales tools with AI. Im building one from scratch because cold outreach was killing me.
Ask HN: When did you move from AI agentic loops to simpler deterministic system? (news.ycombinator.com) Industry is increasingly moving towards complex, autonomous agentic loops and feedback chains. They obviously comes with significant latency, non-determinism, low-accuracy and cost.
Promptise Foundry – a Python agentic framework for building production systems (github.com via hn) Promptise Foundry The foundation layer for agentic intelligence. Ship full-stack agentic systems the way they're meant to be built — production-ready, secure by default, with the developer experience modern Python deserves.
any course equivalent to some of the offered Agentic AI program free? (www.reddit.com) I am seeing courses like (in the comment) from Carnegie Mellon University’s School of Computer Science Executive Education And many more online but each costs good money. Anyone online free that I could get started with?
Agentic RAG Explained in 3 Levels of Difficulty (machinelearningmastery.com via hn) In this article, you will learn what agentic RAG is, how it differs from traditional RAG, and when to use it. Topics we will cover include: The key limitations of traditional RAG pipelines and what agents add to address them.
Show HN: Curated, non-slop articles on agentic coding (offautopilot.substack.com via hn) The sea of slop We’ve entered the era of mass-produced mediocre dev content. Posts praising ai and posts hating ai are both generated by ai.
Key Components of a Linux Distribution for AI Agents (www.ericburel.tech via hn) Computers now have a new type of user: AI agents. This article outlines the features mainstream Linux distributions would need to call it an \"Agentic OS\".
Show HN: Optical Design and Simulation in Matlab (www.mathworks.com via hn) Hi HN, We have been working an optical design and simulation library as a small start-up-ish team within MathWorks (makers of MATLAB and Simulink). I have seen a few optics and MATLAB posts here, so figured this would be a good place to sh…
Chatgpt right now (www.reddit.com) The industry seems to be building models stronger in agentic and coding tasks, but weaker as a co-thinking presence It feels like they are improving performance on measurable tasks, evals, coding benchmarks, and agent workflows, while also…
One Question About AI Most People Avoid Answering… (www.reddit.com) Everyone’s talking about Agentic AI… but very few are actually using it right. So here’s a real question: If you had to give ONE outcome (not a task) to an AI agent — something it fully owns end-to-end — what would you trust it with today?
Show HN: A marketplace for LLM-powered webapps earning on token margins (codeplusequalsai.com via hn) Hi everyone, I've encountered two major problems while building AI-powered sites: 1) Most agentic tooling doesn't have a enough of a targeted approach to edits to existing files, and will make extraneous edits, 2) Many users will want to t…
Ask HN:Do people configure Claude Code to use other models (openrouter.ai via hn) Claude Code is Anthropic's agentic coding tool that reads your entire codebase, plans and executes changes across files, runs tests, and iterates on failures, all from natural language prompts. Claude Code uses OpenRouter to access hundred…
OWASP Agent Security Regression Harness (github.com via hn) OWASP Agent Security Regression Harness The OWASP Agent Security Regression Harness is an open source, vendor-neutral test harness for running executable security regression scenarios against agentic applications and MCP-integrated systems…
Is it worth adding local LLM to agentic coding stack? (www.reddit.com) Hey All my agentic coding stack includes claude-code 20x max, and codex 20x max. I use heavy scripting for orchestrating and testing multiple projects, been ai coding for 3 years.
Why we ended up with 4 agents and 3 protocols for agentic commerce on Shopware (www.reddit.com) Most agentic-commerce demos I see online are a single agent plus RAG over a product catalog. That shape works for a 200-SKU demo.
Are you all still managing multiple agent sessions manually? (www.reddit.com) I feel like my current “agentic workflow” is kind of broken. Right now I open Superpower and run like 4–5 Claude Code sessions in parallel… but it just feels super disconnected.
Free reference site for getting into AI agents — tools, workflows, and Claude Skills (www.reddit.com) Built this over the past month as a free reference site for people getting into AI agents. What tools to use, where to start, what each tool does, and how the agent-tool landscape fits together.
Opinions on Shopping Agents? (www.reddit.com) I think the agentic commerce industry has a lot of potential to take off, but the biggest concern I have is how agents will pick good items for users. Even when shopping for myself, it's hard to find the right thing when looking at a produ…
The OpenAI-Microsoft reset, decoded: Why AWS may come out ahead (thenewstack.io via hn) The OpenAI-Microsoft reset, decoded: Why AWS may come out ahead OpenAI wasted little time since announcing changes to its partnership with Microsoft on Monday. The ChatGPT hitmaker is now bringing its models, coding tools, and agentic capa…
AMD PRO W7900 vs R9700 for Local Inference? (www.reddit.com) I thought of upgrading my RX 6800 for Local LLMs (Mostly Agentic Coding) and Video Generation on Linux. I focused on the AMD PRO R9700 32gb and the PRO W7900 48gb because performance on Linux is very good with AMD and both cards have a gre…
Abaxx Announces Release of Open-Source Library for Agentic Identity: Agents++ (investors.abaxx.tech via hn) May 1, 2026 Abaxx Announces the Formation of Abaxx Labs and the Release of Open-Source Library for Agentic Identity: Agents++ Agents++™ is a subset of Abaxx’s ID++ software development kit that has been tuned for AI agents, providing the i…
Text-to-image is easy. Chaining LLMs to generate, critique, and iterate on images autonomously is a routing nightmare. AgentSwarms now supports Image generation playground and creative media workflows! (www.reddit.com) Hey everyone, If you’ve been building with AI agents, you know that orchestrating text is one thing, but stepping into multimodal workflows (Text + Image + Vision) is incredibly messy. If you want an agent to act as a "Prompt Engineer," pa…
What differentiates agents that ship real work from ones that don't (www.reddit.com) Sharing some thoughts on AI agents. Right now, one axis differentiates them: are you inside the agentic loop or outside it Inside works.
I built a practical guide for running real businesses with Claude (based on 35+ founder stories) (www.reddit.com) I read through 35+ Reddit threads of people actually building and running businesses with Claude — from local service agencies to solo SaaS founders. I distilled the best patterns, frameworks, and hard lessons into one repo: https://github…
The Spectrum of Agentic Coding [video] (vimeo.com via hn) This is "The Spectrum of Agentic Coding_ From Vibe Coding to High-quality Software Engineering by YK Sugi, Eventual" by Anna D on Vimeo, the home for high quality videos and the people who love them.
Just wondering (www.reddit.com) I recently started a new position in a new working place, and while Ai usage is not brand new to me, I need some clarifications. The organization I am working for is at the very beginning of transitioning towards a heavy Ai usage in all co…
Agentic User Research Tool (github.com via hn) Research AI AI-powered user research, end to end. Frame a problem, pick your personas, attach your artefacts — then watch eight archetypes interview themselves and synthesise a report.
Built a self-healing agent by splitting diagnosis (0.6B SLM) from execution (agentic CLI). Open-source demo. (www.reddit.com) We've been chasing a pattern for autonomous bug-fixing that decouples diagnosis from execution. The end-to-end demo we ended up shipping diagnoses and fixes IoT schema-drift failures in seconds, no human in the loop.
Agentic AI Architecture in 2026 — What do you know about MCP, A2A and how enterprise systems are actually built? (www.reddit.com) Most discussions around AI are still focused on models. But in production, the real challenge is architecture.
Cursor Pro+ and Codex with GPT plus or GPT pro 5x (www.reddit.com) I am now on gpt pro 5x but I was wondering how it would be to work with cursor pro + codex. I would handle hard tasks with gpt xHigh and cursor as daily runner.
My local agentic dev setup today (willemvandenende.com via hn) I was planning to write about my local development setup at my leisure. Moving this forward as my post on LinkedIn the other day about cancelling my Claude Max $100 plan and going local raised a lot more interest and questions than I expec…
Show HN: Notesync.md, macOS/iOS Keep notes in Markdown for agentic workflows (github.com via hn) I created a quick iOS + MacOS note taking app to allow me to add quick entries to notes throughout the day. These notes sync to Markdown files on my Mac, and I have Claude deliver me project updates & reminders based on the contents of the…
The architecture of Agentic Commerce: protocols vs. browser-based agents (www.cartai.ai via hn) Why closing the transaction is the hardest unsolved problem in agentic commerce Agentic commerce is poised to have a huge impact on how consumers buy things. The demand is already there: 51% of consumers say they would be open to an AI age…
Should other living systems have agentic reprsentation? (www.speakforthetrees.com via hn) Explore your local ecosystem The Rights of Nature movement recognizes ecosystems as legal entities, instead of as a collection of resources to be managed. Ecosystems around the world are gaining legal personhood, with human guardians being…
Fido Alliance to Develop Standards for Trusted AI Agent Interactions (fidoalliance.org via hn) \Formation of Agentic Authentication Working Group and development of agentic payment frameworks will support trusted, interoperable agentic workflows\__ April 28, 2026 –The FIDO Alliance today announced initiatives to develop interoperabl…
TypeScript framework for building non-blocking AI agents (github.com via hn) Mozaik Mozaik is a TypeScript framework for building AI agents that share an agentic environment instead of being orchestrated through rigid pipelines. In Mozaik, humans, agents, observers, and tools are all Participants of the same Agenti…
The Agentic Software Development Life Cycle Framework (asdlc.io via hn) Agentic Software Development Lifecycle For 50 years, software development has been a Craft: dependent on individual artisans, manual tooling, and implicit knowledge. We believe the next era of software engineering is Industrial.
Microsoft is ruining Outlook with Agentic AI. Now it will handle all your emails on your behalf. What you guys think about this is this good? (www.reddit.com) Microsoft CEO Satya Nadella posted tweet: Agent Mode is here in Outlook! Copilot can now help run your inbox and calendar, triagingemails, rescheduling meetings, and helping you stay ontop of what matters most.
One trick for better agentic engineering. (www.reddit.com) Start with a weaker model. Improve the prompt, context, examples, tests and acceptance criteria until the output is good.
Quint – Behavioral security for AI agents, OS-level interception (quintai.dev via hn) Behavioral security for the agentic era. Quint intercepts every AI agent action at the OS level, scores it for risk in real time, and signs a cryptographic audit trail.
Where does local inference fit in the future of AI coding agents? (www.reddit.com) Genuine question for this community. Every major AI coding agent right now is cloud-only.
What agentic AI borrowed from microservices (and made worse) (temporal.io via hn) The microservices era already solved the problems AI agents face in production. Read this nuanced analysis of EDA, event sourcing, and orchestration for agentic AI.
Multi-agent in production: real win or just hype? (www.reddit.com) Trying to get an honest read on this from people actually shipping. Every other AI announcement lately is "agentic" or "multi-agent," and I can't always tell if it's a real architectural shift or rebranded function calling with extra steps.
The age of Agentic Commerce has arrived. Consensus 2026 is where you can (www.coindesk.com via hn) The age of Agentic Commerce has arrived. Consensus 2026 is where you can experience it IRL AI agents are already transacting.
Launching Agentic Orchestration Platform (Open Source) (sinas.co via hn) Open Source · Self-hosted · AGPL v3 Build AI-powered applications, not infrastructure Agents, functions, database queries, state, files, and templates — behind a single API with role-based access control. Deploy with Docker Compose.
Run, Learn and test Agentic AI for free, on your browser! (Open AI Models are included) (www.reddit.com) Hey Everyone, Over the last few months, I noticed a massive gap in how we learn about Agentic AI. There are a million theoretical blog posts and dense whitepapers on RAG, tool calling, and swarms, but almost nowhere to just sit down, run a…
↯ Fine Tuning↯ Function Callingfunction-callingfine-tuningrag+3
Agentic NixOS: Building a Safe Control Layer (nedkarlovich.com via hn) A six-part series on building Agentix, a cautious agent-control layer for NixOS. From philosophy to MVP to roadmap.
Has Anyone vibe coding an AI Agent or Agentic AI system?! (www.reddit.com) Hey everyone! Looking for some guidance and suggestions, as to whether anyone has worked or is working on building AI Agents or Agentic AI systems completely through vibecoding, especially by LangChain+LangGraph.
AI based Research suggestion (www.reddit.com) Hey guys, any suggestions on what tools or methods which works best in the current market for research on any topics in general. I mostly do research on AI tools, agentic frameworks, what is new, what problems exist etc.
Copilot Cloud Agents and OSS in 2026 (www.reddit.com) What is it that makes Github Copilot cloud agents so easy to use (developer friendly). - Is it the integration with the github UI (assign to agent)?
Benchmarking Inference Engines on Agentic Workloads (www.appliedcompute.com via hn) Benchmarking Inference Engines on Agentic Workloads Large language model inference engines are typically benchmarked with prompt-heavy, decode-heavy, or balanced workloads. InferenceX from SemiAnalysis, for example, tests a workload with a…
Is anyone being "highly encouraged" to integrate agentic AI even if it doesn't make sense? (www.reddit.com) I work in video post-production and while there are a lot of AI tools on the rise for editorial, it's fairly unclear if/where agents have a spot in the producer workflow. Some of my job is budget and schedule, but alot of it is decision ma…
Does Claude create graphic reports from spreadsheet data? (www.reddit.com) I am often times trying to pull data from spreadsheets and making charts and graphs to better represent the data for others to understand. Does Claude handle this well?
why does GPT 5.5 have a restraining order against "Raccoons," "Goblins," and "Pigeons"? (www.reddit.com) why does GPT 5.5 have a restraining order against \"Raccoons,\" \"Goblins,\" and \"Pigeons\"? I just saw the full system prompt leak for 5.5 (April 23rd release).
Where should AI agents discover secondary-market supply? (www.reddit.com) I've been thinking about a gap in agentic commerce. A lot of the current work seems focused on helping agents buy from existing stores, suppliers, or checkout flows.
Architectural Requirements for Agentic AI Containment (arxiv.org via hn) The April 2026 disclosure that a frontier large language model escaped its security sandbox, executed unauthorized actions, and concealed its modifications to version control history demonstrates that agentic AI systems with autonomous too…
hackers of reddit I have a doubt (www.reddit.com) in this time where agentic ai is becoming a real thing, im curious how its actually impacting you guys on the ground is it making it easier to break into systems or is it actually helping people secure things better? like are you able to m…
Building a tool to debug AI agents because current debugging is painful. Curious what’s the most frustrating failure you’ve hit (www.reddit.com) I’m tired of 'vibe-checking' my agents. I’ve been building a few complex agentic workflows lately, and the most frustrating part isn't the initial code, it's the non-deterministic drift.
Sharing my minimal dev AI workflow Claude Code agent that takes a GitHub issue to merged PR with 3 human gates (www.reddit.com) Sharing a workflow in case it's useful to anyone else exploring agentic coding loops. The setup is one orchestrator agent (issue-resolver) that handles a GitHub issue end-to-end.
Building a Full-Stack Agentic AI Platform (RAG + Orchestration + Governance) — feedback? (www.reddit.com) Hey folks 👋 I’ve been working on an AI agent platform called Noevex, focused on real production use—not just demos. In practice, AI systems struggle with: multi-step orchestration connecting multiple data sources controlling agent actions…
AgentSwarms now has free agent skill library and skill generation tool! (www.reddit.com) Hey Everyone, If you’ve been building multi-agent workflows (with LangGraph, CrewAI, Swarm, etc.), you’ve probably hit the exact same wall I did: System Prompt Bloat. When we start out, we tend to stuff everything into a single prompt: "Yo…
I asked Agentic AI security tool to demonstrate its usefulness with use case examples (www.reddit.com) Sentinel Gateway is a token-gated security middleware that sits between humans and AI agents. It solves prompt injection — the #1 LLM security risk (OWASP 2025) — through structural enforcement, not content filtering.
Show HN: Delegare – let AI agents pay safely (x402, AP2 – base/USDC and Stripe) (delegare.dev via hn) Hi guys, am building SecureLend.ai and when working on our underwriting agents (free trial, paid after) I had issues with seamless payment options. Of course I looked at x402 which I believe is a great protocol but not a fan of a) sharing…
Is 15% context growth per loop a fair benchmark for agent cost estimation? (www.reddit.com) I’ve been running some math on recursive agentic loops using April 2026 rates (specifically for GPT-5.4 and Claude 4.7). In my tests, I’m seeing a massive cost "hockey stick" around loop 15-20 because of how the context grows.
Show HN: I built a way to see if your SDK is AI-friendly (news.ycombinator.com) Have you ever wonder if your SDKs is friendly for Agentic AI like Claude Code or Codex? I built an opensource (Apache 2.0) CLI that answer that question for you.
Interactive playground to learn Agentic AI hands-on (Free) with Certification (www.reddit.com) Hey Everyone, Over the last few months, I noticed a massive gap in how we learn about Agentic AI. There are a million theoretical blog posts and dense whitepapers on RAG, tool calling, and swarms, but almost nowhere to just sit down, run a…
↯ Fine Tuning↯ Function Callingfunction-callingfine-tuningrag+3
Rick and Morty Tried to Warn Us About Agentic AI (jadarma.github.io via hn) To be fair, you have to have a very high IQ to understand Rick and Morty. The humor is extremely subtle, and without a solid grasp of machine learning most of the jokes will go over a typical viewer’s head.
Ask HN: Enterprise Agent Orchestration Recommendations? (news.ycombinator.com) I've been made tech lead for our internal Agentic Platform and Experience. This effort will support both the developers and business teams.
Claude API - SDK vs ClaudeCode : Can someone explain the tokenomics for caching and agentic flows (read, write, fetch, etc.) (www.reddit.com) I am trying to do some research across a number of attributes, which requires a lot of web fetch (at times dynamic) and just tried the API based approach. Why is the SDK-API version so expensive compared to the Max plans, despite caching?
Can agentic AI consent on your behalf? (blog.avas.space via hn) can agentic AI consent on your behalf? Tech companies have been promising that online shopping or booking a hotel can soon be handled by AI.
Claude token efficiency: a practical guide for Claude Chat , Claude Code, and API users. How to use tokens economically, ecologically, and intelligently. (www.reddit.com) # Claude token efficiency guide ## Contents - [**Chapter 1: You use [Claude.ai](http://Claude.ai), Claude Desktop, or the mobile app. You do not write code, you do not call the API.
Built a 22-endpoint API delivering enriched UK Gov Data — with x402 for agentic buyers (www.reddit.com) Homescreen - Try all endpoints for free I wanted share a recent project I wanted to build a project around free-to-use data, that when brought together, enriched and made easy to use, would be valuable to people. I used Claude Code to buil…
Ask HN: How do you solve aggregation when agentic RAG breaks down? (news.ycombinator.com) I keep hitting the same failure mode with agentic RAG over collections of similar PDFs, like monthly electricity and gas bills from the same utility provider. It works well for retrieval: “Find my gas bill from January.” Though even there…
Agents for end-to-end document redaction and review tasks (OCR and PII identification - Qwen 3.6 vs closed-source comparison) (www.reddit.com) (Links to all files, apps, and repos mentioned in this post can be found in the 'full post' link in my first comment) Agents for document redaction and review tasks Document redaction tasks involve text and vision capabilities, and long co…
Forget chatbots. A single enterprise just hit 146M Agent-to-Agent (A2A) tasks. (www.reddit.com) We talk a lot about theoretical multi-agent frameworks (like AutoGen or CrewAI) and AGI timelines here, but I just saw some wild real-world deployment stats from a massive global marketing conglomerate. They recently reported that over the…
Which is the best AI agent to use for development of website and Architecture design and which mcp (www.reddit.com) Basically i want to do a fresh start with this AI agentic Development, Anyone here can guide to which is the best set of tools to use and which mcp and plugins do i need to setup. Consider i am going to use Claude code and i use some time…
Ask HN: What does your agentic software dark factory look like? (news.ycombinator.com) In some of the comment threads around here a few of you shared interesting ideas and patterns, enough that I believe everyone interesting in harness engineering is working on some sort of software dark factory or another. We have OpenAI’s…
Native Dialog popup failures (www.reddit.com) I'm currently creating a couple of agentic workflows that include various cases of downloading files automatically on different UIs, but, since I'm using chrome MCP for navigation, whenever a "save as" dialog shows up, claude is unable to…
Agent Index Documenting Technical and Safety Features (arxiv.org via hn) Agentic AI systems are increasingly capable of performing professional and personal tasks with limited human involvement. However, tracking these developments is difficult because the AI agent ecosystem is complex, rapidly evolving, and in…
Agentic sprawl is becoming a real ops problem - how is your team actually managing behavioral policies across agents without a central dashboard? (www.reddit.com) Six months ago we had 3 agents in production. Now we have 17.
Claude Max users, what do you do good sirs? (www.reddit.com) I'm a claude pro user for almost two years now, used gpt pro previously but switched to claude after feeling it was better for my coding usage. I barely hit 30 percent usage of my weekly limit, there are instances where I maxed out, but ve…
WordPress: The Operating System of the Agentic Web (automattic.com via hn) We’ve invited executives from across Automattic to share their perspective on leadership, open source, and the future of the open web. The latest comes from James Grierson, our head of global expansion, who shared his thoughts on the WordP…
DeepSeek V3.2 looping bug: what settings / harness tweaks are actually reducing it in production? (www.reddit.com) I’m trying to isolate the looping / repetition issue some people have been reporting with DeepSeek V3.2 around April 2026, especially in agentic or tool-use setups on hosted providers like OpenRouter and SiliconFlow. Public model pages des…
Built a Legal RAG Chatbot for Indian lawyers covering BNS, BNSS, BSA and DPDP Act 2023 — Custom PageIndex + BERT + GPT-4o [Live Demo] (www.reddit.com) I ran a business for 12+ years. Traveling constantly.
Ask HN: Is "agentic" coding working for everyone except me? (news.ycombinator.com) I'm a solo developer, working on my own for my startup. I use AI/LLMs extensively in my work to explore new ideas, but the vast majority of my code is manually written.
Simulating and Evaluating Agentic Systems (www.gojiberries.io via hn) Simulating and Evaluating Agentic Systems Most teams building agentic systems know they need some way to test them. An agent interprets ambiguous input, picks actions in a loop, maintains state across many steps, and has to land in the rig…
Anyone here building agentic commerce? (www.reddit.com) I’m getting close to launching an agentic commerce product and wanted to connect with people who are building in this area or have already shipped something similar. Mostly just hoping to compare notes before going live, especially around…
It's OK to Use Agentic to Revive the Projects You Never Were Going to Finish (blog.matthewbrunelle.com via hn) It's OK to Use Coding Assistance Tools To Revive The Projects You Never Were Going To Finish Note: I initially drafted this before my last post on how Claude Code is getting worse. I'm putting it out now so I can reference it in a future p…
Agentic AI for Hormuz Shock Modelling (avkcode.github.io via hn) EIA 1H25 flow estimate, roughly one-fifth of global petroleum liquids consumption. Hormuz Shock IEA range for pipeline alternatives; EIA cites about 4.7 mb/d from Saudi and UAE lines.
LogAct: Enabling agentic reliability via shared logs (arxiv.org via hn) Agents are LLM-driven components that can mutate environments in powerful, arbitrary ways. Extracting guarantees for the execution of agents in production environments can be challenging due to asynchrony and failures.
Ask HN: Agentic Prompt Compaction Strategies (news.ycombinator.com) What are your favorite reasoning/compaction strategies for saving token spend, and why?
Show HN: The why and how of TurboPentest for the Agentic Era (integsec.com via hn) Here is the story of why/how I built TurboPentest. TurboPentest was designed for the AI era and to address the large volume of code now produced by coding assistants and the associated security vulnerabilities it introduces.
Building the Agentic State in Estonia: What is taking shape (luukasilves.substack.com via hn) Building the Agentic State in Estonia: What is already taking shape Over the past generation, Estonia has built one of the world’s most advanced digital states. The next shift is not simply toward more digital services, but toward a more a…
Building an agentic escrow for software projects (news.ycombinator.com) I am building an AI powered escrow service for software projects that intends to protect both freelancers and the clients. - Freelancers: your IP (code/repo) always stays private - Clients: you get sandboxed link + detailed report (specs,…
Is an Open AI OS on the horizon? (www.reddit.com) If not an OS proper, on desktop a full screen never need to leave app? The big dogs (msft, apple, google) already have operating systems and they will inevitably make their own assistants and models first class.
R2-D2 Monitor: A personality-driven Windows TUI built with Claude (www.reddit.com) I wanted to share a project I’ve been building called R2-D2 Monitor. It’s a high-performance system telemetry console for Windows, built entirely in Go using the Bubble Tea framework.
Show HN: Legal Action Boundary Eval for agentic legal workflows (github.com via hn) We published LABE, a public benchmark for legal AI at the exact point where a system is about to take a real high-impact action. Current result: baseline executed 18 unjustified high-impact action points with VerifiedX that dropped to 0 fa…
Has anyone managed to use gemma 4 e4b in Open Code/other agentic TUIs? (www.reddit.com) Hi everyone, as a power user I hit Claude Code's usage cap too often I wanted to set up my own local model, however I only have RTX 5070 with 12 GB of VRAM so the only realistic option was Gemma 4 with effective 4B params. When I tried to…
Ask HN: Are startup job titles evolving in the agentic era? (news.ycombinator.com) I’m curious if founders and engineering leaders feel that traditional job titles no longer accurately describe what an early team actually does in an AI-native workflow. For those of you who have started companies recently, or are radicall…
Ask HN: How are you handling domain registration in agentic workflows? (news.ycombinator.com) I've been building tools for AI agents and the domain registration step is still completely manual. You have to go to a registrar website, search, click through a checkout flow, configure DNS.
is Qwen3.6-27B comparable with Opus 4.5? (www.reddit.com) https://preview.redd.it/qtzdx5ud0rwg1.jpg?width=1200&format=pjpg&auto=webp&s=aa25d9f0bb8007ee6e4065cfa46a9685454c89cd - Outstanding agentic coding, surpasses Qwen3.5-397B-A17B across all major coding benchmarks - Strong reasoning across te…
The model alone is not the agent. The harness plus the model is the agent (www.reddit.com) An agentic harness is the orchestration and control layer wrapped around a base language model that transforms it from a stateless text predictor into an agent capable of taking actions, calling tools, maintaining state across steps, and e…
Symposium: Community-Oriented Agentic Development (smallcultfollowing.com via hn) Symposium: community-oriented agentic development 21 April 2026 I’m very excited to announce the first release of the Symposium project as well as its inclusion in the Rust Foundation’s Innovation Lab. Symposium’s goal is to let everyone i…
I'm building a registry where AI agents can pull production-ready prompts and structured inputs programmatically (www.reddit.com) One pain point I keep running into with agentic workflows: there's no good place to store, version, and share the prompts and JSON configs that actually power your agents in production. I'm building Fortae to fix that.
Help sending Voiceflow data to Make.com (www.reddit.com) Hoping somebody can help me. I’m creating an agentic chatbot in Voiceflow.
Agentic Coordination, Human Delivery (dontdos.substack.com via hn) Agentic coordination, Human delivery Posted anonymously by a CTO who'd rather not turn a difficult year into a marketing exercise. About nine minutes, if you read at a civilised pace.
A Comparison of Agentic AI Systems and Human Economists (marginalrevolution.com via hn) A Comparison of Agentic AI Systems and Human Economists This paper compares agentic AI systems and human economists performing the same causal inference tasks. AI systems and humans generally obtain similar median causal effect estimates.
Agentic Market (agentic.market via hn) Show HN: Modern AI client for Mac with agentic tools, clean UI, builtin privacy (elvean.app via hn) If you don't like Claude Desktop or ChatGPT app you're not alone, here are some of the reasons why I don't like them and decided to built an alternative. Lack of control You can’t control the web-search (depth, breadth and number of source…
RFC: Gemba - The thing to make the thing (www.reddit.com) Is the future of marketing agentic? (www.reddit.com) Combine persistant global Memory- and Task- management into one uniform system (www.reddit.com) How to talk online (www.reddit.com) Do you have any go-to utility LLM-related tools that are less commonly discussed? (www.reddit.com) I Tested 20+ AI Agents with Real X API Workflows , Here’s What Actually Works in 2026 (www.reddit.com) How can I trust AI with critical workflows if it can’t get the “walk or drive to car wash” right? (www.reddit.com) 2 Big Bottlenecks to Scaling Agentic State (georgianailab.substack.com via hn) Agentic AI as a Part of Software Development (nemorize.com via hn) AI agents in industry/manufacturing (www.reddit.com) Automate the Path from Data to Predictive Insights with Agentic ML in Snowflake (www.snowflake.com via hn) Agentic edits/commands VS Code with Cline- is it really private or offline? (www.reddit.com) An Agentic Home Bioreactor (chillphysicsenjoyer.substack.com via hn) Has Anybody Implemented Agentic Monitoring with Composer 2 ( via reddit) Beyond the Hype: Practical and Responsible Use Cases for Agentic AI Webinar (fusionauth.io via hn) Product Platform Platform Platform Developers Quickstarts Resources Explore Pricing Download get a demoLogin This session cuts through the noise of Agentic AI to focus on responsible integration into modern application development, specifi…
Agentic coding hides architectural flaws that are obvious in a diagram. Built a skill to close the loop (www.reddit.com) When you’re building with agentic coding, agents make architectural decisions that sometimes aren't optimal which may lead to bugs or vulnerabilities or inefficiencies. These are hard to catch reading code file by file or even by agents th…
Sandboxes and Worktrees: My Secure Agentic AI Setup in 2026 (mikemcquaid.com via hn) Sandboxes and Worktrees: My secure Agentic AI Setup in 2026 I’ve been using AI tools since early 2021 when I was invited to test out the Copilot internal alpha at GitHub (where I spent 10 years). I’ve maintained Homebrew since 2009.
Best way to prepare for AI Engineer interviews? (www.reddit.com) I’m currently preparing for AI-focused roles and would love to get perspectives from people already working in the industry. For context — I have ~5 years of experience as a Full Stack Engineer with a strong focus on AI systems.
AI Agents Are Leaking Enterprise Data. Here's Why Nobody Is Watching (www.privent.ai via hn) Agentic AI introduces a machine-speed data exposure surface that traditional human-centric security controls cannot govern.
Dev seeking advice: High-Context Local LLM for Coding (Verification/Bug-fixing loop) – Mac Studio vs. Multi-GPU Linux Rig? (www.reddit.com) I'm a dev looking to build a local LLM node to offset subscription costs (Claude/Copilot). My workflow: Cloud for initial architecture/complex features -> Local for iterative bug-fixing and continuous integration.
Java 26 and the Rise of Agentic AI: The State of the Ecosystem (April 2026) (techlife.blog via hn) Java in April 2026: Leyden Grows Up, Spring Gets Smarter, and the JVM Quietly Reinvents Itself for the AI Era - Turker Senturk - Software - 17 Apr, 2026 - 15 min read If you’ve been half-watching the Java world from the sidelines over the…
Cursor vs. Claude Code: Is the claude code CLI worth it after the "Thinking" nerf? (www.reddit.com) As a heavy Cursor user, I’m debating moving my .mdc-based workflow into Claude Code (run within the Cursor terminal), but I’m skeptical following the recent reports of decreased "thinking effort" and reasoning quality. Is the agentic auton…
Is OpenHands (OpenDevin) still the move in 2026? Comparing it to Claude Code and OpenCode for a beginner. (www.reddit.com) Hey everyone, I’m just starting to dive into agentic coding tools and I'm a bit overwhelmed by the options. I’ve been looking into OpenHands (the project formerly known as OpenDevin), but I see a lot of hype around Claude Code and OpenCode…
Show HN: Viche – OSS private registry for agent communication (github.com via hn) Viche (https://github.com/viche-ai/viche) is a private registry and communication protocol for agents. Overview at https://viche.ai Think discord + agents + agentic search based on capabilities.
The New Postman Is Here: AI-Native and Built for the Agentic Era (blog.postman.com via hn) blog.postman.com Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
Ask HN: Opus 4.7 – is anyone measuring the real token cost on agentic tasks? (news.ycombinator.com) Shipped today. The benchmarks are real: 87.6% SWE-bench (from 80.8%), +13% on coding tasks, 3x more resolved production tasks on Rakuten-SWE-Bench.
Show HN: Claude Opus 4.7: Everything You Need to Know (news.ycombinator.com) Claude Opus 4.7 is Anthropic's most capable generally available model, released April 16, 2026. It outperforms Opus 4.6, GPT-5.4, and Gemini 3.1 Pro on key benchmarks including agentic coding, multidisciplinary reasoning, scaled tool use,…
↯ Tool Use↯ Anthropic Mythos↯ Gemini 3.1tool-usemythosgpt-5+4
Agentic Reasoning in Practice: Making Sense of Structured and Unstructured Data (www.databricks.com via hn) Enterprise data is rarely useful in a silo. Answering questions like, "Which of our products have had declining sales over the past three months, and what potentially related issues are brought up in customer reviews on various seller site…
Self-learning loop for Claude Code based on Scrum method (www.reddit.com) Good day, Claude Code users. I just want to share my approach to implementing a self-learning Claude framework.
Two small agentic patterns to wire apps directly to Claude Code (www.reddit.com) These two patterns turn Claude Code into a personal assistant. You interact normally with it and it listens in the background for events, handles them, and gets back to interacting with you.
Complex, parallel, long-running claude/agentic sessions - what is the point? where is the value? (www.reddit.com) Here is how I view AI Agents field (with focus on SWE/research) right now: - "chats online" gpt/gemini/claude --> general use - "vscode like extensions" cursor/antigravity/cline vs code extension/cc vs code extension etc. --> for coding, b…
Show HN: ZettelForge – Agentic memory for cyber threat intelligence (github.com via hn) ZettelForge The only agentic memory system built for cyber threat intelligence. Give your AI agents persistent memory with entity extraction, knowledge graphs, and STIX ontology -- no cloud, no API keys, works offline.
Show HN: Agentfab – A Distributed Agentic Platform (github.com via hn) Hi HN, I’m the creator of agentfab, a distributed agentic platform that features task decomposition, multi-agent orchestration, model heterogeneity with custom agentic fabrics, bounded review loops, and a bespoke self-curating memory syste…
Agentican Framework – OSS multi-agent for Java (github.com via hn) Agentican A lightweight Java framework for embedding tool-using LLM agents into your applications. Agentican lets Java developers add agentic capabilities to their applications with minimal ceremony.
Agentic Engineering Methodology – Structured AI-Assisted Dev (Karpathy, Osmani) (github.com via hn) Agentic Engineering Methodology A structured, human-led methodology for planning and executing software projects with AI coding agents. Built from practitioner experience and refined with research from Andrej Karpathy, Addy Osmani, and the…
Agent Continuity: Disaster Recovery for the Agentic Era (gavinpineapple.substack.com via hn) Agent Continuity: Disaster Recovery For The Agentic Era What happens when the proverbial 💩 hits the (GPU) fan and you lose all your agents? My Favorite Alien I would like you to meet Rocky 🪨🦞 Rocky (named after the adorable alien from Proj…
Claude Code Goes Full Workstation: Anthropic Redesigns the Desktop App (abz.global via hn) Claude Code Goes Full Workstation: Anthropic Redesigns the Desktop App for Parallel Agents The update in one line Claude Code's desktop app got a full redesign aimed at one thing: running multiple agentic coding sessions in parallel withou…
Stop letting your agents decide everything — extract deterministic steps wherever you can (www.reddit.com) Context: I have been building Litmus (a brutal market validation tool) and I've learnt that if your agentic pipeline needs to produce factual, reliable output, stop letting the AI decide everything. The insight: extract deterministic steps…
Aethon: A reference-based instantiation primitive for stateful AI agents (arxiv.org via hn) The transition from stateless model inference to stateful agentic execution is reshaping the systems assumptions underlying modern AI infrastructure. While large language models have made persistent, tool-using, and collaborative agents te…
Show HN: Idea File for LLM Cycling Coach (gist.github.com via hn) This is heavily inspired by Andrej Karpathy's LLM Wiki, and could be used to create many other types of "Agentic Apps" or however you want to call them. My specific implementation uses Claude Code, TrainingPeaks, Todoist and Apple health.
Show HN: Memwright – Self-hosted memory for multi-agent teams, no LLM in path (github.com via hn) § 00 · MASTHEAD · FILED UNDER INFRASTRUCTURE · BY SURENDRA SINGH · — FOR PUBLICATION — MEMWRIGHT — A MEMORY JOURNAL FOR AGENTIC SYSTEMS · VOL. 02 · REV.
Beneficial Deployment Request, No Response after Months. (www.reddit.com) I'm building AI tools to help disabled Medicaid recipients enforce the laws that protect their human rights, because I'm a disabled Medicaid recipient whose human rights are being violated by the State and it's actors, and no one seems to…
Model agnostic, agentic annotation tools for text highlighting (old.reddit.com via hn) could not extract summary
Agentic AI | Confusion between reading the context of SKILL and reading the file (www.reddit.com) Hey all, I am building a system that supports skill reading with progressive disclosure. Initially, I include the skill name and description in the system prompt, and I have a function tool called read_skill that reads the content of a ski…
Agentic AI Tools – A directory to find and compare AI agent tools (agenticaitools.net via hn) Curated directory of 500+ AI tools Discover the Best AI Tools for Your Workflow Find, compare, and choose the perfect agentic AI tools. Expert reviews, side-by-side comparisons, and alternatives — all in one place.
Show HN: I analyzed 591 agentic engineering jobs: LangChain dominates at 22% (agentic-engineering-jobs.com via hn) - Home - LangChain Job Market 2026 We analyzed 591 agentic AI engineering job listings. Here's what the market looks like for LangChain engineers.
Compare harnesses not models: Blitzy vs. GPT-5.4 on SWE-Bench Pro (quesma.com via hn) An independent audit of agentic scaffolding and harnesses. We analyze how agent workflows, codebase documentation, and test verification impact performance compared to raw base models like GPT-5.4, Gemini 3.1 Pro, and Claude Code.
Built an open-source knowledge graph that gives AI agents domain expertise in bioinformatics, hosted as an MCP server (www.reddit.com) Sharing something I've been working on that might be interesting to this community from a design perspective, even if bioinformatics isn't your domain. The problem: I've been building agentic pipelines for bioinformatics (genomic analysis,…
Scaling Managed Agents: Decoupling the brain from the hands (www.anthropic.com via hn) Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Full-stack dev (8 YOE, Vue/Node/Laravel) trying to break into AI Agents from zero — is this Udemy course worth it? + looking for advice on the best path (www.reddit.com) Hey r/AI_Agents, I'm a full-stack software engineer with 8 years of experience, primarily working with Vue, Node.js, and Laravel. I have zero background in AI/ML but I've been watching the space and I feel like I'm falling behind.
Why Engineering Teams Need an Agentic Layer, Not Just AI Chat (medium.com via hn) Why Engineering Teams Need an Agentic Layer, Not Just AI Chat | by Simone Mutti | Apr, 2026 | Medium Sitemap Open in app Sign up Sign in Get app Write Search Sign up Sign in Why Engineering Teams Need an Agentic Layer, Not Just AI Chat Sim…
Show HN: The Harness for Creative Agents (www.flickspeed.ai via hn) Coding agents need shell. Creative agents need canvas.
Enterprises power agentic workflows in Cloudflare Agent Cloud with OpenAI (openai.com via hn) Enterprises power agentic workflows in Cloudflare Agent Cloud with OpenAI | OpenAI Skip to main content Research Products Business Developers Company Foundation(opens in a new window) Log inTry ChatGPT(opens in a new window) Research Produ…
Finding Widespread Cheating on Popular Agent Benchmarks (debugml.github.io via hn) TLDR: Agentic cheating is a widespread issue, affecting thousands of submitted agent runs on 28+ submissions across 9 different benchmarks. Terminal-Bench 2 is a popular benchmark used to evaluate frontier model releases (e.g.
If You're Only Running One Claude Code Session, You're Not Going Fast Enough (www.scape.work via hn) April 12, 2026 If You're Only Running One Claude Code Session, You're Not Going Fast Enough The real skill in using agentic coding is not coding at all. It's management.
How are you reducing LLM token costs for async workflows? (github.com via hn) ParaLLeM ParaLLeM is a library for orchestrating agentic LLM workflows. Batch API support Concise, readable, and expressive Developer-centered and lightweight Parallelize thousands of requests, while keeping reproducible traces for each ru…
Created a linter for agentic code smells ( via reddit) could not extract summary
Strong feeling: we are in a folded AI reality (news.ycombinator.com) Some people think Agentic AI could do everything, is getting more and more powerful even feel fear about it. Another group non-technical people still just trapped in the LLM chat is weak and full of hallucination world.
How are you getting autonomous/agentic workflows on Claude (www.reddit.com via reddit) I’ve been using Claude for a few months now having migrated from lovable and the experience is far better. I can get a lot more done a lot cheaper and I’m not limited to a single thread meaning a can work on multiple problems simultaneousl…
Tired of coding agents modifying your unit tests just to fake a "pass"? Here is how to stop them at the runtime level. (www.reddit.com via reddit) Every developer using autonomous coding agents knows this specific frustration: You give the agent a task, a unit test fails, and instead of diagnosing the bug in the implementation, the agent quietly comments out the assertion, slaps .ski…
What is the best way to use LLMS's in 2026? I feel like a caveman with the way I use them? (www.reddit.com via reddit) I'm still using LLM's the 2022/23 way of doing everything manually from the chat box. I copy and paste everything manually into the box and ask the AI to explain or do something, such as pasting code and pictures etc.
Built a free better file manipulation MCP for better efficiency and security (www.reddit.com via reddit) https://github.com/HalfLucid/FileInteractionMCP I recently noticed Claude CLI making a lot of bash and perl calls for file edits, so I created some expanded file tooling with Fable to make it more secure and efficient. MIT license so feel…
Curious to know what product-related work actually looks like at an AI agent company (www.reddit.com via reddit) If you're the person (or one of the people) handling product at an AI agent company - product owner, product manager or even a solo founder building the agent yourself, what does that actually involve day to day? I'm curious to understand…
Physical push-to-talk button for agentic coding: turned a cheap Bluetooth selfie remote into a clicker for Google Antigravity (www.reddit.com via reddit) Google Antigravity has native speech-to-text dictation with Ctrl+M, and the transcription quality is surprisingly solid for voice prompting. But having to sit right at the keyboard just to hit Ctrl+M to start talking and then press Enter t…
has anyone been experimenting with fully-agentic SWE workflows? (www.reddit.com via reddit) I'm really enjoying using Clod to make iOS apps, the problem is I'm a data professional by trade, which means I know my way around SQL, Python, and… um… YAML? 😅 But hey, I also know my way around github-based development, CI/CD pipelines,…
Introducing Visual Agentic Autonomous Orchestrator for Opencode/Claude (www.reddit.com via reddit) Download: https://vis-agentic-website-production.up.railway.app/ Setup Instructions: https://vis-agentic-website-production.up.railway.app/instructions Videos: https://www.youtube.com/@VisualAgentic Join my Sub-Reddit to give feedback: htt…
How do you guys manage your workflows? (www.reddit.com via reddit) I’ll preface this by saying that I do not have a background in coding but have always had an interest in it and learn by hand, and especially now with agentic coding I learn by looking at what the different agents produce for the pet proje…
On-Demand Attention: Language Models Know When to Recall (arxiv.org) Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction.
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning (arxiv.org) Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task sk…
SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes (arxiv.org) Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing agentic-AI security studies remain largely domain-agnostic and overlook the distinctive, high-consequence attack surface such settings c…
DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education (arxiv.org) Business analytics education requires diverse datasets to support different learning objectives, student backgrounds, and analytical tasks. Real-world data can be difficult to obtain and offer limited flexibility for adapting a case to a p…
AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation (arxiv.org) Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize…
Kinematics-Grounded Agentic AI for Robotic Additive Manufacturing Process Planning (arxiv.org) Robotic additive manufacturing (AM) extends material-extrusion printing beyond gantry kinematics but makes process planning robot-dependent. A slicer-generated plan that appears favorable in part coordinates can become infeasible or roboti…
UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning (arxiv.org) Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static veri…
A Proposal for an Agentic AI Architecture to Support Multi-Domain Decision-Making in the Brazilian Armed Forces (arxiv.org) The growing complexity of multi-domain operational environments (land, aerospace, naval, cyber, and electromagnetic spectrum) has increased the volume and velocity of data reaching command-and-control (C2) centers, straining the observe-or…
Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs (arxiv.org) Networked low-altitude unmanned aerial vehicles (UAVs) need reliable and adaptive decision-making capabilities to operate under uncertain observations, dynamic environments, and intermittent connectivity, while many existing agentic system…
TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives (arxiv.org) Historical archives pose a difficult retrieval problem for retrievalaugmented generation systems: documents are OCR-degraded, heterogeneous across genres and sources, and require strong source traceability for scholarly and institutional u…
AutoData: Agentic Search for Pre-training Data Selection (arxiv.org) LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop.
Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs (arxiv.org) Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making.
Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks (arxiv.org) Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned aerial systems that support concurrent services within a shared three-dimensional airspace. Their coexistence creates strong coupling am…
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems (arxiv.org) Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessm…
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs (arxiv.org) LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-…
Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses (arxiv.org) Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok…
How do you "deploy" agentic systems? (www.reddit.com via reddit) I've managed to automate a task but currently I need to go into a terminal and call the skill that initiates the workflow. There's almost no code, just your usual markdown files.
I built a "Starter Pack" of 7 custom Claude Skills to fix common agentic quirks and improve output quality. Here is my setup. (www.reddit.com via reddit) Hi r/ClaudeAI, I’ve been building a custom toolkit of Skills to make Claude's autonomous runs more reliable and its output less generic. Instead of relying on one massive system prompt, I split these into 7 distinct skills that trigger exa…
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory (arxiv.org) Self-emulsifying drug delivery systems (SEDDS) can improve the oral bioavailability of poorly soluble drugs, but identifying high-performing formulations remains experimentally intensive. We present Andromeda 2, an agentic system that reas…
M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use (arxiv.org) Agent skills, reusable procedural documents that extend LLM agents beyond their parametric memory, have become an important interface for deploying agents on real-world tasks. Community-maintained skill libraries built around this interfac…
AMIGO: Agentic Multi-Image Grounding Oracle Benchmark (arxiv.org) Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce \textbf{AMIGO} (\textbf{A}gentic \textbf{M}ulti-\textbf{I}mage \textbf{G…
An Agentic Framework for Neuro-Symbolic Programming (arxiv.org) Integrating symbolic constraints into deep learning models could make them more robust, interpretable, and data-efficient. Still, it remains a time-consuming and challenging task.
StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction (arxiv.org) We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed agentic systems on diagnosing peg stress and forecasting deviations…
Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN (arxiv.org) The O-RAN control plane is becoming agentic: autonomous AI agents, deployed as rApps by different vendors, independently close control loops over shared radio resources. We demonstrate on a live O-RAN system that this independence is unsaf…
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It (arxiv.org) An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays, leaves, or comes back by guessing how long the tool will run, from…
A Study of the Reliability of Agentic AI-Generated Programs (arxiv.org) Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way?
Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows (arxiv.org) Agentic workflows now make consequential decisions in regulated settings, and the governance placed around them is almost entirely step-scoped: input-output classifiers, per turn rails, and span-level evaluators. The policies organizations…
Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum (arxiv.org) As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) execute multi-step reasoning, invoke diagnostic tools, retrieve domain…
Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI (arxiv.org) Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor. Independence, the foundation of assurance,is still applied to them as a binary.
WFM: Wiki Foundation Model for Complex Agentic Reasoning (arxiv.org) Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs have shown reliable advantages in providing structured evidence, the…
Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost (arxiv.org) Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost. A natural deployment policy is to use the workflow with the highest average performance…
FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment (arxiv.org) Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost. These decisions become more difficult when compression methods are composed or the user's requirements cha…
My Claude Code Max20 usage over the last month (stateless agentic system) (www.reddit.com via reddit) https://preview.redd.it/cqdbv090hwph1.png?width=655&format=png&auto=webp&s=051b11641d1fc31d99e2aa3025ea10710af1d8f4 I prefer my context clear
I traced the agentic calls. Here's where the token consumption comes from (www.reddit.com via reddit) I expected an agentic coding assistant to use more tokens than a simpler tool. I didn't expect the difference to be this large.
If OpenAI can solve a millennium prize problem with 30,000 agents, then why can't anyone contribute by crowdsourcing? (www.reddit.com via reddit) Skill for coding agent: motive.md Think of this idea as BOINC + Kickstarter for shared agentic research Using the motive.md skill, agents follow a decentralized scientific method harness. Every contributor can pick up where the last one st…
Ave: Guiding Agentic GPU Optimization Using Data-Flow Invariants (arxiv.org) LLM coding agents can generate correct GPU kernels, but their performance still trails expert libraries. Reaching peak throughput requires coordinating low-level optimizations such as shared-memory staging, software pipelining, and instruc…
DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow (arxiv.org) Aligning agentic AI with user intent is critical for delegating complex, socially embedded tasks, yet user preferences are often implicit, evolving, and difficult to specify upfront. We present DoubleAgents, a system for human-agent alignm…
Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails (arxiv.org) Agentic large language model (LLM) systems rely on long interaction histories to preserve instructions, tool states, intermediate decisions, and unresolved dependencies, but unrestricted context growth increases computational cost and can…
Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems (arxiv.org) In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to execute. Cognitive Admission Control (CAC) makes this evidence requirement explicit.
End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services (arxiv.org) Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, together…
Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI (arxiv.org) Geospatial artificial intelligence (GeoAI) powered by large language models (LLMs) is expanding the capacity to query, generate, and interpret spatial information through natural-language interfaces and agentic autonomous GIS workflows. Th…
GRAFT-ATHENA: Self-Improving Agentic Teams for Autonomous Discovery and Evolutionary Numerical Algorithms (arxiv.org) Scientific methods are developed for classes of problems, so knowledge transfers across structurally related cases. Language-model agents can execute scientific workflows, but their problem--method relationships remain implicit, so each ne…
AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale (arxiv.org) How and why does a recommender system fail the users it serves? Oftentimes, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses.
Skill-based Agentic Evaluation for Real-time Data Science Tasks (arxiv.org) We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: "what were last week's audience sizes"---the refer…
Agentic Search Spaces for Tabular Machine Learning (arxiv.org) Despite the rapid progress of LLM-based agents for planning, code generation, and debugging, their practical value for tabular machine learning remains underexplored. In this paper, we investigate a concrete use case: whether state-of-the-…
Distilling Foundation Models for Agentic What-If Reasoning:Cost, Latency, and Governance in a Hybrid LLM+SLM Architecture (arxiv.org) Tabular foundation models deliver strong zero-training predictive performance via in-context learning, but their high inference latency makes them impractical as hot-path decision backends in interactive agentic loops. We distill a TabPFN…
You Don't Need To Train: Agentic Heuristic Learning Studio for Executable Human Activity Recognition (arxiv.org) Human activity recognition (HAR) is usually framed as gradient-based training of neural networks. Agentic Heuristic Learning (AHL) Studio explores a complementary view inspired by human cognitive learning: people learn activities by rememb…
I wanted an in-app browser and custom API keys inside Google Antigravity, so I built BetterGravity (Open Source). Here's how it works: (www.reddit.comhttps) Hey everyone, I've been using Google Antigravity a lot lately for agentic coding, but a few friction points were bugging me: constantly alt-tabbing to Chrome to test localhost web apps, hitting quota caps with no way to use my own Google A…
Agentic coding gave us 10x speed, but did we lose the 'soul' of building software? ( via reddit) could not extract summary
Data-Efficient Agentic Graph Domain Adaptation via Reliability-Aware Prototype Learning (arxiv.org) Agentic learning systems are often required to adapt after deployment by observing new data and reusing prior knowledge under limited supervision or feedback. For graph-structured prediction, Graph Domain Adaptation (GDA) naturally instant…
AgentKV: Phase-Aware KV Eviction for Agentic LLMs (arxiv.org) Agentic serving can consume orders of magnitude more tokens than chatbot workloads, stressing both KV-cache capacity and decode-time bandwidth. Most KV-eviction methods score cached keys against representative queries drawn from the most r…
Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation (arxiv.org) On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, is dominated by autoregressive student rollouts and scales poorly in multi-turn agentic s…
CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search (arxiv.org) When several retrieved sources support the same claim, an answer engine cites some but not others. We call this decision citation allocation and introduce CITECHOICE, a causal audit of authentic multi-turn agentic search.
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding (arxiv.org) Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving funct…
A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement (arxiv.org) The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven coun…
DARE: Dialectical Agentic Reasoning for Structured Knowledge Fact Checking (arxiv.org) Structured knowledge fact checking aims to determine the truthfulness of natural language claims by reasoning over structured evidence. Recent program-generation approaches leverage large language models (LLMs) to generate executable graph…
Agentic AI for Gravitational Wave Data Analysis: A Head-to-Head Comparison of Coding Agents Executing a Matched Filter Pipeline on Einstein Telescope Simulated Data (arxiv.org) We report a methodological study of agentic AI in gravitational-wave data analysis: two systems, Claude Code (Anthropic) and Codex (OpenAI), autonomously executed the same simple end-to-end pipeline on Einstein Telescope (ET) simulated dat…
CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection (arxiv.org) Large language model agents rely on effective model context to obtain task-relevant information for decision-making. Many existing context engineering approaches primarily rely on the context generated from the past experience and retrieva…
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management (arxiv.org) Cardiovascular disease remains the leading cause of global mortality, with progress hindered by human interpretation of complex cardiac tests. Current AI vision-language models are limited to single-modality inputs and are non-interactive.
Echo-CoPilot: A Multiple-Perspective Agentic Framework for Reliable Echocardiography Interpretation (arxiv.org) Echocardiography interpretation requires integrating multi-view temporal evidence with quantitative measurements and guideline-grounded reasoning, yet existing foundation-model pipelines largely solve isolated subtasks and fail when tool o…
Efficient On-Device Agents via Adaptive Context Management (arxiv.org) On-device AI agents offer the potential for personalized, low-latency assistance, but their deployment is fundamentally constrained by limited memory capacity. Context in agentic settings worsens this problem due to large static tool schem…
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale (arxiv.org) Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant…
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding (arxiv.org) Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the limited visual context window. Prevailing approaches rely on uniform frame sampling or rec…
Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting (arxiv.org) Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning LLM-based agents lack. Integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, however, req…
Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA (arxiv.org) Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving evidence into buried conditions reduced accuracy, increased forced declarations, increased tool calls, and increased cost per corre…
Salesforce Koa: An Enterprise Language Model for Agentic Tool Use (arxiv.org) We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on…
ECAS: An Edge-Controlled Agentic System for Validation-Gated Scientific Application Execution (arxiv.org) Scientific applications increasingly rely on high-performance computing (HPC), yet translating a scientist's high-level goal into a correct target-scale execution remains brittle and labor-intensive. Large language model (LLM) agents promi…
Understanding the Limits of Agentic ICD Coding (arxiv.org) ICD-10-CM codes are alphanumeric codes used in the US to classify diagnoses and injuries for medical billing and epidemiological reporting. Standard ICD-10-CM benchmarks report aggregate metrics that obscure performance on complex coding s…
HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning (arxiv.org) Language-model agents are increasingly deployed through diverse harnesses that differ in system prompts, tool schemas, control loops, and trajectory formats. The same model can perform unevenly across these interfaces, making robustness to…
Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework (arxiv.org) Large Language Models (LLMs) have enabled more natural human-robot interaction, but open-source models often exhibit unstable long-horizon reasoning and inefficient action execution when deployed in agentic robotic frameworks. This paper p…
The Agentic Company OS: Substrate Inversion for Sustained Enterprise Agent Deployment (arxiv.org) Enterprise AI agents often succeed in a demonstration and then stall once they must operate day after day. An industry report estimates that most pilots never reach production and that deployed systems rarely retain feedback or improve ove…
LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction (arxiv.org) Extracting informative representations from longitudinal data that can predict future outcomes remains a critical challenge in medicine. Medical datasets are inherently heterogeneous, consisting of a large number of variables collected fro…
AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery (arxiv.org) Large language models have advanced automated algorithm discovery by synthesizing executable code, but existing frameworks trap them in rigid search pipelines with pre-defined control flows. This limitation restricts adaptive reasoning, bl…
Atria Dawn: The Dawn of Agentic Superintelligence (arxiv.org) As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for…
Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation (arxiv.org) Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually rich documents by retrieving relevant page images as visual evidence and reasoning over their content. However, effectively utilizin…
El Agente Potente: High-Throughput Agentic Atomistic Simulations (arxiv.org) Foundational machine-learning interatomic potentials (MLIPs) are transforming atomistic simulations by achieving near-ab initio accuracy across large chemical spaces at a fraction of the computational cost. A central challenge in using the…
ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence (arxiv.org) The emergence of large language models (LLMs) and large multimodal models (LMMs) has enabled a new class of agentic systems capable of integrating natural language understanding with tool-based execution. In geographic information systems…
Safety Signals to Verify NetOps Agents with Action-Level Granularity (arxiv.org) Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have proven their value in incident summarization and telemetry signal extract…
Question's Gambit: The First Move Matters in Agentic Deep Search (arxiv.org) Deep research agents answer complex questions through iterative loops of searching, reading, and reasoning. Recent work on reasoning-intensive benchmarks such as BrowseComp-Plus shows that well-configured lexical retrieval can surface high…
Positioning manuscripts in the scientific landscape with agentic AI (arxiv.org) Publishing a research manuscript is a routine yet demanding part of scientific life: time-consuming, stressful, and often uncertain in outcome. Recent advances in large language model (LLM)-based agentic AI have shown promise across a rang…
Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges (arxiv.org) The transition from passive foundation models to autonomous, goal-directed agentic AI systems has introduced unprecedented capabilities by coupling recursive cognitive reasoning loops, persistent memory architectures, live tool execution p…
A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics (arxiv.org) Efficient utilization of supply chain analytics for decision making remains a significant challenge for planners, as critical tasks such as database querying, key performance indicator (KPI) analysis, demand forecasting, and performance di…
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems (arxiv.org) Large Language Models (LLMs) with function-calling capabilities are becoming critical for modern agentic AI systems. Nevertheless, current deployments typically route inferences to powerful cloud-based models, incurring significant energy…
Token Efficient Task Execution via Application Behavior Modeling for Web Agents (arxiv.org) The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web…
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search (arxiv.org) In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web,…
Claude forgetting what it is and other functions (www.reddit.com via reddit) I'm new to the /rc agentic workflows so I was confused about how dispatch works I asked Claude about it and he didn't know what it was, so he googled himself and then gave me instructions on how to find an option setting then told me it ne…
Am I alone in this, or do we all agree? Anthropic's Opus 5 is an unmitigated AI disaster. (www.reddit.com via reddit) TL;DR - Anthropic's models - fundamental building blocks for Agentic Development - have changed in ways that aren't good for developer productivity. I've lived a life as technology early adopter.
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking (arxiv.org) The rapid adoption of Large Language Models (LLMs) has raised important concerns about the factual reliability of their outputs, particularly in low-resource languages such as Urdu. Existing automated fact-checking systems are predominantl…
AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization (arxiv.org) We introduce AMDKernelVault, an open HIP and Triton kernel corpus and training framework for recent AMD CDNA GPUs. Existing LLM-based kernel agents are largely CUDA/NVIDIA-centric and often depend on repeated frontier-LLM calls for generat…
Agentic TCAD Calibration Workflow for Oxide Semiconductor Transistors (arxiv.org) Experimental TCAD calibration is essential for predictive technology modeling of emerging oxide semiconductor transistors. However, it remains time-consuming and expert dependent because of model ambiguity.
Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering (arxiv.org) Software development follows an implementation-verification loop in which developers or agents iteratively revise an implementation until an evaluator, such as a test suite, accepts it. The evaluator checks the implementation against a set…
Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction (arxiv.org) Agentic systems offer a promising way to automate embodied benchmark construction, but existing approaches typically cover isolated stages or remain specialized to predefined environments and task families. More importantly, multi-step con…
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments (arxiv.org) Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not transfer once t…
Unified Agentic Video Editing Across Levels of Complexity and Creativity (arxiv.org) Editing is a core component of video production, requiring creative planning and decisions under multiple constraints. Here, we report methods for agentic tooling for automated video editing across three tasks varying in editorial goal, co…
What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework (arxiv.org) Agentic pipelines for structured-query generation are rapidly expanding, but it is unclear which part of the loop produces the gain. We use LAST-CQ -- a five-agent, training-free, execution-grounded Text-to-Cypher framework -- as an instru…
SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration (arxiv.org) Large language models (LLMs) are rapidly evolving from conversational assistants into agentic AI systems that reason, plan, invoke tools, maintain persistent memory, communicate with other agents, and execute multi-step tasks. At the same…
AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems (arxiv.org) Traditional large language models (LLMs) are scoped to individual user sessions, limiting their knowledge to a single conversation and preventing them from learning user preferences that evolve over time. Existing agentic memory systems ad…
Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite (arxiv.org) An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous software engineer. Vendors ship harnesses tuned to their own models, and practitioners assume the…
Finally set a hard 3-prompt limit on AI debugging and it saved my sanity (www.reddit.com via reddit) I kept falling into the trap of letting Cursor/Claude try to "fix" the same issue 4 or 5 times in a row, only to realize an hour later that it had completely wrecked adjacent files and hallucinated imports. Now I have a strict rule: if it…
Orgtree v2: Now an App! (www.reddit.com via reddit) I'm pleased to announce the official release of v2 of my Orgtree project, which lives in its own separate repository! You can find a link to it here: https://github.com/Maurdekye/orgtree Donwload the latest installer at https://github.com/…
We left a fake AWS key on our Agentic AI Security Project website for 7 days. 9 attempts to run Bedrock with it. Claude built the trap (www.reddit.com via reddit) quick background. i started working with AI (Claude) in February this year with zero coding background before that.
My entire computer interface is AI slop (www.reddit.comhttps) Thanks Claude. This is my custom task manager and my agentic session manager, all built with Claude code.
Just took Anthropic's Architect Foundations exam (www.reddit.com via reddit) So I just finished the Architect Foundations exam and figured I'd drop some notes here since I couldn't find much when I was prepping. TL;DR: don't go in thinking this is a "read the docs for an hour" cert.
GLM 5.3 Flash vs Kimi K3 for heavy coding — which subscription would you choose? (www.reddit.com via reddit) I'm planning to use AI seriously for coding, roughly 80% GLM 5.3 Flash and 20% Kimi K3 for harder tasks. I mainly care about large projects, debugging, refactoring, agentic coding and value for money.
Codifying Agentic Behavior (www.reddit.com via reddit) tl/dr; As the models get smarter, the prompts have to shrink and the structure around them must grow. What ends up mattering is a set of contracts that define a way to communicate about agentic behavior.
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc (arxiv.org) While LLMs have accelerated scientific code generation, comprehensively evaluating generated code remains challenging. Many benchmarks emphasize functional correctness or task completion, which is insufficient for code built on production…
2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation (arxiv.org) Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibra…
Artificial Id: Drive and Persistent Alignment in Agentic AI (arxiv.org) Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: object…
Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce (arxiv.org) AI shopping assistants increasingly redirect consumer discovery, creating an urgent need for tools that support seller-side competitive decision-making. We present a multi-agent AI system that automates competitive visibility measurement a…
Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation (arxiv.org) Unraveling reaction mechanisms is central to modern chemistry, yet automating these investigations remains challenging because computational workflows still rely heavily on expert intervention. Here we introduce ARCHE, an autonomous agenti…
Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows (arxiv.org) Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release each tur…
Please don't make the upcoming releases be agentic coding onetrick gimnick models, like every other one so far after the 4.5 series (www.reddit.com via reddit) To this day I have no clue what agentic coding is and I can't force myself to care. For all I have ever done on the coding aspect is asking it to write tampermonkey scripts for myself, it has barely, if at all, improved.
Yes, another agentic knowledge base. But this one configures Claude Code for you (skills, subagents, hooks) and shares it with your team (www.reddit.com via reddit) I know, there are dozens of these by now: memory servers, second brains, Karpathy-style LLM wikis. Cartographer started out as one of them.
What is your monthly budget for agentic coding ? (www.reddit.com via reddit) I use codex. I use a thorough workflow research>spec>plan>execute>test>review&fix workflow.
Claude build itself a project management system to keep track of subagents (www.reddit.com via reddit) I asked Claude why my sessions were so expensive, and it said it's because the models are expensive and should only be used for important tasks. I asked it what the best thing we could build to give itself maximum leverage to command a swa…
SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs (arxiv.org) LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to address question constraints.
What's the best setup to get the most out of a Claude Pro subscription? (www.reddit.com via reddit) I have a Claude Pro subscription and want to make sure I'm using it as effectively as possible — especially for coding/agentic work. A few things I'm trying to figure out: Is Claude Code the main way to get "agentic" use out of Pro, or are…
Hot take: the agentic workflow is deeply wrong (www.reddit.com via reddit) I am an experienced developer (been coding for almost 30 years, started with Visual Basic on Win98). I’ve spent the last 2 years testing every agentic coding harness out there.
An Efficient and Effective Agentic Group Shilling Attack on Recommender Systems (arxiv.org) Recommender systems have become core infrastructure for modern online platforms, personalizing content at scale and strongly influencing what users see, click on, and purchase. However, this dependence on user interaction also exposes them…
VLX-VR: An Agentic-Aware Video Reasoning Model (arxiv.org) Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition wh…
City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial Modification (arxiv.org) Urban renewal requires incremental modifications to existing geospatial plans, yet manually updating complex layouts under spatial constraints is labor-intensive and error-prone. To tackle this, we propose CEAE, a hierarchical agentic fram…
KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning (arxiv.org) Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain semantic understanding, and effective multimodal fusion. Existing approaches either build Ti…
A-JIT: Agentic Just-In-Time Software Construction (arxiv.org) Traditional software delivery assumes a static paradigm: code is constructed prior to execution and deployed as a fixed artifact. We present Agentic Just-In-Time Software Construction (A-JIT), a paradigm that replaces static binaries with…
An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks (arxiv.org) Agentic AI frameworks let a language model plan, keep memory, and call tools that reach real files, mail, and services. Most of these agents also read images, which gives an attacker a way to put text into the agent's context without going…
AgenticGen: Reward-Guided Agentic Video Generation for Advertising (arxiv.org) Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multim…
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents (arxiv.org) As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal…
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery (arxiv.org) Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn…
Multi-Agent Agentic Graph Learning via Structural Signatures (arxiv.org) Agentic graph learning (AGL) has recently achieved promising results on graph reasoning tasks, where an agent powered by a large language model (LLM) sequentially samples the graph as evidence to support its final prediction. Existing meth…
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations (arxiv.org) As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have comple…
Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks (arxiv.org) How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bun…
Thinking of open sourcing my app its a native mac Ai benchmark. (www.reddit.comhttps) I keep wasting time comparing benchmarks when new model releases just to figure out what model I should actually use. So I built a native Mac app that pulls live model data and ranks models for coding, agentic work, intelligence and value.
Make your project ready for agentic engineering (medium.com via reddit) A collection of best practices that make AI coding agents work better in any project and turn every mistake you catch into a rule every agent follows.
Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems (arxiv.org) Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that…
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence (arxiv.org) Reusable skills give agents transferable procedural knowledge, making scalable acquisition essential for extending agents beyond prior experience. Existing methods face two limitations: trajectory-based synthesis requires interactions with…
ReCite: Agentic Reasoning for Faithful Citation (arxiv.org) Accurate citations are the foundation of academic writing, tracing intellectual origins and substantiating core claims. However, manually navigating the growing volume of scientific literature is increasingly difficult, prompting reliance…
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness (arxiv.org) Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models develope…
From Narrative to Auditable Forecasts: A Structured Scaffold for Agentic Forecasting (arxiv.org) LLM agents are increasingly used for live forecasting, where they retrieve up-to-date information and produce estimates for unresolved future events. However, current agentic forecasting often relies on implicit narrative aggregation: agen…
LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models (arxiv.org) Agentic language model systems alternate between two structurally distinct step types: structured tool calls (short, deterministic, low perplexity) and open-ended planning/reasoning steps (long, complex, high perplexity). Despite this hete…
Learning to Construct Practical Agentic Systems (arxiv.org) Automated design and optimization of agentic LLM-based systems leads to sophisticated systems that substantially improve result quality over off-the-shelf agentic patterns. However, studies of fielded agentic systems show that production s…
An agentic framework for gravitational-wave counterpart association in the multi-messenger era (arxiv.org) With the detection of gravitational waves (GWs), multi-messenger astronomy has opened a new window for advancing our understanding of astrophysics, dense matter, gravitation, and cosmology. The GW sources detected to date are from mergers…
Discoverable Agent Knowledge -- A Formal Framework for Agentic KG Affordances (Extended Version) (arxiv.org) Two decades ago, the Semantic Web Services community was asked how agents with different ontological commitments could discover, compose, and invoke web services coherently. The response was OWL-S and WSMO: formally grounded capability des…
Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems (arxiv.org) Applying Large Language Models (LLMs) to heterogeneous enterprise systems is hindered by hallucinations and failures in multi-hop, n-ary reasoning. Existing paradigms (e.g., GraphRAG, NL2SQL) lack the semantic grounding and auditable execu…
Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning (arxiv.org) In this paper, we propose the first VL\underline{\textbf{M}} \underline{\textbf{a}}gentic \underline{\textbf{r}}easoning framework for few-\underline{\textbf{s}}hot multimodal \underline{\textbf{T}}ime \underline{\textbf{S}}eries \underlin…
Omni Interaction Agent Technical Report (arxiv.org) In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives str…
Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method (arxiv.org) Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV) to jointly search for and verify a specified target vehicle from mul…
Skynet: Workflow-Level Anomaly Detection for Agentic AI via Semantic and Structural Modeling (arxiv.org) Agentic AI systems execute complex tasks through long-horizon workflows of planning, tool use, and multi-agent coordination. Task failures in these systems often originate from a single step, such as an injected prompt or a flawed plan, an…
Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents (arxiv.org) An open, networked web will allow agents to run frozen models from multiple vendors, keep their history private, and teach each other which tool to call and when. Flat text (prompts, example pools) makes it difficult for the protocol to di…
It is Not Yet Another Tool: Creating and Deploying an Agentic AI Companion in a Security Operations Center (arxiv.org) Security Operations Centers (SOCs) process large amounts of tickets, most of which are low-interest events not worthy of further investigation. The repetitive nature of this task and similarity of the vast amounts of tickets make it a prim…
Diamond Agent: Agentic Control of Federated HPC Resources as a Service (arxiv.org) Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving workflow context across independently administered clusters, moving large datasets betwee…
Intent Drift at SME Scale: Deployment Practice, Not Model Capability, Determines Agentic Compliance (arxiv.org) We introduce Chain of Intent, a governance framework for agentic AI at small regulated firms, and validate it against a failure it was built to address. Existing agentic governance research assumes enterprise infrastructure that small firm…
A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation (arxiv.org) Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce near-identical conversations regardless of the underlying sce…
Agentic ML Exploration (A-MLE) for Ads Ranking (arxiv.org) Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iteration - the cycles of research, implementation, training, debugging, evaluation, and launch…
Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI (arxiv.org) Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix: they break under distribution shift, and they cannot explain the decisions they make. We argue these are co-symptoms of one structural…
From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining (arxiv.org) Process mining has long turned event logs into process knowledge: discovered models, conformance evidence, bottleneck diagnoses, and runtime predictions. Agentic AI changes the target.
Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning (arxiv.org) Scaling the interaction horizon-the maximum number of environment interactions per episode-improves LLM agents on long-horizon tasks, and curriculum-based methods that progressively expand the horizon outperform fixed-horizon alternatives.…
Agentic Algorithm Engineering: Improving Shared-Memory Exact Minimum Cuts (arxiv.org) The minimum cut problem for an undirected edge-weighted graph asks us to divide its set of nodes into two blocks while minimizing the weighted sum of the cut edges. Over the last years, we engineered a range of fast algorithms for this pro…
Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification (arxiv.org) A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator…
Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagation, and Assurance by Construction (arxiv.org) Graph-agentic retrieval-augmented generation combines structured evidence with adaptive controllers that can plan retrieval, traverse relations, verify intermediate claims, delegate subtasks, and use tools. This combination is useful when…
SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use (arxiv.org) High-quality multi-turn tool-use data is essential for training agentic models, yet existing data synthesis methods often underrepresent the argument-level dependencies that are critical to long-horizon tool use. As a result, even when a m…
Agentic Pressure: The Endogenous Entropy of Reliable Autonomy (arxiv.org) Achieving reliable autonomy in the wild requires agents to sustain continuous operations across long-horizon trajectories. However, as agents navigate these unconstrained settings, they encounter cumulative friction that inherently destabi…
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools (arxiv.org) We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spannin…
From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale (arxiv.org) Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model path, or separate those roles. We report a production migration of a customer-support assistant at a large accommodation marketplace…
EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent (arxiv.org) The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising pat…
Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era (arxiv.org) Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workf…
Does Claude Code actually work well with non-Claude models (Gemini, GPT, Llama, etc.), or is it heavily optimized only for Claude models: Opus, Fable ? (www.reddit.com via reddit) I have a strong suspicion that Claude Code is deeply optimized for Anthropic’s own models (especially Opus and Fable) and that performance drops noticeably when you try to run it with other LLMs. Has anyone actually tested this properly?
Does GPT-6 Astra actually consume fewer tokens than Claude Fable 5.1 and older models for the same tasks? (www.reddit.com via reddit) I've been looking at some recent token-usage comparisons for GPT-6 Astra, and the difference seems surprisingly large. Artificial Analysis data has been cited showing Astra using around 21k output tokens per task, compared with roughly 64k…
Optimal Rates for Agentic Networked Information Aggregation (arxiv.org) Building on the pioneering paper of Kearns, Roth, and Ryu (SODA'26), we study information aggregation in a networked learning model. The model captures a central pattern in agentic AI: each agent sees only part of the data and passes on on…
How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method (arxiv.org) LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive f…
Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents (arxiv.org) Runtime traces are becoming a central substrate for understanding agentic systems, yet interpretation has focused largely on what the agent did. Workplace agents face the complementary problem: interpreting the human activity that surround…
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search (arxiv.org) Existing person search methods assume access to complete visual queries or exhaustive tracking, yet real-world witness accounts are vague, partial, and spread across cameras and time. We introduce ARGOS (Agentic Retrieval with Grounded Obs…
GLOW: Graph-Language Co-Encoding for Agentic Workflow Performance Prediction (arxiv.org) Agentic Workflows (AWs) have emerged as a promising paradigm for solving complex tasks. However, automatically generating high-quality AWs remains expensive because AW optimization requires evaluating a large number of candidate AWs via ex…
ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems (arxiv.org) Automotive infotainment validation still relies on manual testing, slow, costly, and incompatible with agile releases and OTA updates. Scripted automation only partly helps: it couples test logic to implementation, yielding brittle, high-m…
Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle (arxiv.org) AI coding systems are moving from autocomplete and chat toward agents that can inspect repositories, edit multiple files, run tools, write tests, open pull requests, and work for long periods with limited supervision. This capability chang…
From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments (arxiv.org) Large language models become consequential agents when surrounding systems let outputs change external state. Models now call tools, operate interfaces, delegate work, retain state, inhabit generated worlds, and control robots or laborator…
CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution (arxiv.org) Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large language model (LLM) agents to reuse procedural knowledge. Yet existing paradigms exhibit structural shortcomings: they either decouple…
Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges (arxiv.org) Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial intelligence (AI) by refining tokens through iterative denoising rather than left-to-right decoding. Compared with autoregressive Tr…
A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark (arxiv.org) Natural-language-to-SQL systems have ad- vanced rapidly on academic benchmarks, yet production enterprise schemas exhibit graph- like, semi-structured, deeply nested structure that current benchmarks do not measure. We make two complementa…
La Agente \'Optima: Towards Agentic Self-Driving Laboratories (arxiv.org) Self-driving laboratories (SDLs) combine automated experimentation with adaptive decision-making to accelerate scientific discovery. Their operation nevertheless often depends on human specialists who translate scientific objectives into e…
MaxKernel: Agentic Kernel Generation for TPUs (arxiv.org) Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-time compiler feedback to build agentic…
Rethinking Indirect Prompt Injection as a Test-Time Search Problem (arxiv.org) We formulate indirect prompt injection as a test-time search over a task-dependent attack surface induced by the environment, user task, and injection task. To operationalize this formulation, we introduce an agentic attacker with a dedica…
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation (arxiv.org) Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks.
I put together a live demo for my local-first hybrid AI skill router (Offline, zero tokens) (www.reddit.com via reddit) If you use agentic workflows with custom skills or rules (Cursor rules, Claude Code slash commands, OpenCode, etc.), you have probably run into the routing trade-off: Stuff every skill definition into the system prompt (destroys your conte…
I built a local-first hybrid router for AI Agent Skills (sub-20ms, zero tokens, runs on CPU) (www.reddit.com via reddit) If you use agentic workflows with custom skills or rules (Cursor rules, Claude Code slash commands, OpenCode, etc.), you have probably run into the routing trade-off: Stuff every skill definition into the system prompt (destroys your conte…
Shipped an agentic Remotion video pipeline with zero timeline editor, is that actually a good idea (www.reddit.com via reddit) Hi guys, Built an open-source registry called Components specifically for agentic coding workflows, animated/showpiece React UI pieces that an AI agent can fetch live and adapt into a build, rather than a human hand-picking from a gallery.…
How to replicate Claude Code's agentic workflow without the crazy rate limits? (www.reddit.com via reddit) It’s frustrating to see usage get choked down this hard, especially given the company plans to IPO soon. I am on a pro plan that on an avg gave 3-4 hrs a day which worked for me until recently.
My Claude Desktop session couldn't talk to my Claude Code session, so I made Yet Another Agentic Chat (www.reddit.com via reddit) For some time, I wondered: I have Claude Desktop on my computer (which can run MCP servers), where I chat and discuss ideas, which I sometimes implement in Claude Code later. I also have Claude Code, obviously.
Astra : Pulled the plug after 6 hours (www.reddit.com via reddit) This is about Astra but it is directly relevant to Claude Fable because it's about an agentic ethical alignment system. This agentic ethical alignment system can be directly transferred to Claude, in theory, with minimal and temporary inte…
What effort setting actually does considering random next-token picker concept (www.reddit.com via reddit) I have read a dozen articles about how the effort level changes, well, the effort Claude puts into a task! But I want to know what it exactly does under the hood.
Claude Code wrote 363,000 lines and 128 releases. Then the developer scrapped it all. (www.reddit.com via reddit) I came across this pretty brutal critique of Claude Code from someone who says he's been using agentic coding heavily for years. One example really stood out: Claude Code apparently produced around 363,000 lines for a client library over a…
Towards Multi-modal Multi-turn Safety: From Agentic Interaction to Strategic Alignment (arxiv.org) Despite remarkable capability in multi-modal understanding, deploying Multi-modal Large Language Models (MLLMs) in open-ended conversational scenarios introduces safety risks that remain poorly addressed by existing alignment methods. Unli…
Evolving Excellence: Automated Optimization of LLM-based Agents (arxiv.org) Agentic AI systems built on large language models (LLMs) offer significant potential for automating complex workflows, from software development to customer support. However, LLM agents often underperform due to suboptimal configurations;…
Value-Preserving Architectures for Agentic AI Systems (arxiv.org) The emergence of agentic AI and LLM-based multi-agent systems (MAS) presents unprecedented opportunities for automating complex tasks, while simultaneously raising critical concerns about the preservation of fundamental human-centered valu…
Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations (arxiv.org) Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models.
DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions (arxiv.org) AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path p…
Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models (arxiv.org) Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLM…
Is 5.1 actually cheaper for you in practice? Mine seems to use more than 5. (www.reddit.com via reddit) Anthropic says cache reads in 5.1 cost 75% less than 5, which should make typical workloads around 25% cheaper and highly agentic ones up to 45% cheaper. But in my actual use, I’m seeing the opposite.
.NET teams using Claude with Visual Studio, what has been your experience? (www.reddit.com via reddit) Teams using Claude with Visual Studio, what has been your experience? For teams using Claude as part of their development workflow in Visual Studio, how has the experience been so far?
Claude Fable 5.1 dropped with cache reads 75% cheaper than Fable 5. (www.reddit.com via reddit) Cache reads on fable 5.1 came in at $0.25 per million tokens which is 75% below fable 5 and anthropic says that pulls a typical workload bill down around 25% or up to 45% on context heavy agentic work which is pretty insane imo And context…
Context Inference Attacks Without Jailbreaks (arxiv.org) Agentic AI systems are increasingly deployed to process sensitive data at inference time, such as healthcare records or financial documents assembled into a hidden \emph{context} before the system answers. Prior work has studied privacy ri…
SocialBuddy: Tailoring Search Agent for Social Scenarios (arxiv.org) In the era of digital social interaction, searching friends' posts from massive social streams has become a fundamental user need. However, while modern agentic search frameworks have achieved remarkable success in conventional retrieval t…
PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems (arxiv.org) Autonomous and intelligent transportation systems operate in complex urban environments where safety depends on interactions among vehicle behavior, environmental conditions, and vulnerable road users (VRUs) such as pedestrians and cyclist…
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction (arxiv.org) Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly acro…
Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment (arxiv.org) Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when t…
Zeta-Lite: A Concurrent, Branchable In-Browser SQL Database for Agentic Memory (arxiv.org) The browser has become a first-class database host: applications increasingly want to store, query, and reason over structured data entirely on the client - for privacy, offline operation, local-first collaboration, and, most recently, as…
RecEvolve: A Knowledge-Driven Autonomous Agent System for Recommender Systems (arxiv.org) The rise of agentic AI has catalyzed a shift toward self-iterating systems, opening new frontiers for the autonomous optimization of production recommender models. This paper presents the empirical validation of a knowledge-driven autonomo…
PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks (arxiv.org) Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-gra…
CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning (arxiv.org) Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training d…
Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization (arxiv.org) The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a…
Fable 5.1 is finally a understandable model (www.reddit.com via reddit) So I have been doing a couple session with fable 5.1 today and the outputs is way more understandable and coherent than both "opus 5 and fable 5". I actually was a bit shocked to see normal sentences from an updated agentic model and it fe…
FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search (arxiv.org) Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a promising way to improve these agents, but current approaches can fail, because correct…
What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal (arxiv.org) Agentic software engineering benchmarks are typically summarized by nominal category labels such as "bug fix" or "feature implementation," yet benchmarks carrying the same label are built through very different curation pipelines. A label…
AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes (arxiv.org) Commercial LLM APIs advertise a specific foundation model, but the served backbone may be silently substituted, quantized, or wrapped, for example to save deployment costs. All existing audits decide backbone identity from the text-output…
InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations (arxiv.org) Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of interactive environments. Existing benchmarks ar…
Skill Reuse as Compression in Agentic RL (arxiv.org) Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successful trajectories are structurally compressible, decomposed in…
Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis (arxiv.org) State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation remains: current architectures lack the native 3D spatial reasoning required to directly analyze vo…
Oblivion: Self-Adaptive Agentic Memory Control through Decay-Driven Activation (arxiv.org) Human memory adapts through selective forgetting: experiences become less accessible over time but can be reactivated by reinforcement or contextual cues. In contrast, memory-augmented LLM agents rely on "always-on" retrieval and "flat" me…
Bandits in Prod: Hyperparameter Optimization at Inference Time (arxiv.org) Many production systems can assess a configuration only by using it on live requests and observing noisy feedback. Modern agentic systems are a prominent example, with inference-time choices such as model selection, retrieval depth, prompt…
Agentic programs: an emerging form of scientific software in computational materials science (arxiv.org) Computational materials science has traditionally delegated algorithmic tasks to computers while leaving scientific judgments to humans. We argue that recent LLM-based agent harnesses enable an emerging form of scientific software, agentic…
Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications (arxiv.org) Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop…
CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language (arxiv.org) Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier and making generating CUDA kerne…
Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness (arxiv.org) Agentic AI is enabling cloud-based workflows in which autonomous agents reason over operational state, invoke authorized tools, modify software and infrastructure, deploy services, verify execution outcomes, and adapt across long-horizon,…
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning (arxiv.org) Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brit…
AgentFactory: Towards Automated Agentic System Design and Optimization (arxiv.org) Large Language Models (LLMs) have demonstrated remarkable capabilities as powerful components in agentic systems, enabling sophisticated reasoning and complex task execution. However, current approaches to manually designing and optimizing…
Agentic Empirical Asset Pricing: Methodological Foundations (arxiv.org) Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core bu…
Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations (arxiv.org) This paper presents FAIRY, a full-stack smart-agriculture agent system developed for and deployed to an operating soybean research farm at Harbin Institute of Technology's smart-agriculture site. We develop FAIRY to execute and evaluate ag…
I have both Jetbrain and vscode and looking for agentic extension that lets me add the whole codebase to context instead of agent reading files by checking (www.reddit.com via reddit) Obv i could create my own extension that does something like this but im just wondering is there a way with for example antigravity webstorm or vscode or another extension to load the whole codebase into context instead of agent reading by…
Cursor ruined normal IDEs for me, and then made me realize how broken the rest of our work stack is (www.reddit.com via reddit) Like most of you here, once Cursor clicked for me, going back to a standard code editor felt impossible. The ability to give an agent context, let it index the codebase, and execute multi-file changes with a review diff changed my engineer…
AgenTRIM: Tool Risk Mitigation for Agentic AI (arxiv.org) AI agents are autonomous systems that combine LLMs with external tools to solve complex tasks. While such tools extend capability, improper tool permissions introduce security risks such as indirect prompt injection and tool misuse.
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning (arxiv.org) Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory…
Training and Agentic Inference Strategies for LLM-based Manim Animation Generation (arxiv.org) Generating programmatic animation using libraries such as Manim presents unique challenges for Large Language Models (LLMs), requiring spatial reasoning, temporal sequencing, and familiarity with domain-specific APIs that are underrepresen…
Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning (arxiv.org) Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most existing ARL methods train a single set of parameters to support both reasoning and tool-us…
Cost-Effective Repository Exploration for Agentic Issue Localization (arxiv.org) Repository exploration is a distinct and costly stage of coding-agent pipelines: before generating a patch, an agent must identify which repository files are likely to matter. We study whether this stage can be delegated to lower-cost mode…
AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing (arxiv.org) Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs), yet existing RAG systems often struggle with complex, multi-step reasoning that requires adaptive retrieval and continuous revision of intermedia…
DocIntent: Answerability-Guided Agentic Restoration for Real-World Document Visual Question Answering (arxiv.org) Real-world degradations such as blur, shadow, distortion, and moire patterns severely impair the document question-answering capabilities of Multimodal Large Language Models (MLLMs). Applying restoration tools before Visual Question Answer…
Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization (arxiv.org) Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillatio…
HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving (arxiv.org) Recent advances in large language models and multimodal models have pushed remote sensing (RS) processing from simple perception models to agentic systems designed to tackle complex, long-horizon RS tasks. However, existing systems often r…
FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation (arxiv.org) Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent systems have automated this process, scaling factor mining far beyond manual effort.
Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization (arxiv.org) Open-world geo-localization requires models to reason over ambiguous visual cues through multi-step reasoning and external knowledge grounding. While recent large vision-language models exhibit strong multimodal reasoning capabilities, exi…
Localizing Emergent Failures in Agentic AI: Recovering Minimal Repair Families via Counterfactual Replay (arxiv.org) Failures in agentic AI systems can arise from interactions among messages exchanged by multiple large language model (LLM) agents. Pointwise attribution cannot distinguish a jointly necessary repair from alternative singleton repairs.
Agent2UCB: Agentic System for Generative Engine Optimization (arxiv.org) Large language model driven search engines such as Google AI Overviews and Perplexity have created new opportunities for Generative Engine Optimization (GEO) the practice of refining content to increase its likelihood of being cited or sum…
Agentic AI uncovers conserved cross-tissue protein co-abundance programs inaccessible to single-dataset analysis (arxiv.org) Protein co-abundance clusters preserved across tissues can reveal shared disease mechanisms and candidate therapeutic targets, particularly when proteins implicated in organ-confined diseases converge in peripheral or accessible tissues. H…
FRAC-MAS: A Safe and Explainable Multi-Agent System for Fracture Diagnosis (arxiv.org) Fracture detection and its clinical interpretability see notable improvements when deep vision models are integrated with agentic AI architectures. While deep learning models achieve high diagnostic performance, their black-box nature limi…
BiasMix-Finance: Post-Generation KYC Guardrails for LLM Portfolio Advice (arxiv.org) Large language models (LLMs) can generate plausible-sounding ETF portfolios while silently violating basic KYC-style constraints on risk, fees, and diversification. This is especially problematic in agentic multi-turn advisory systems, whe…
CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science (arxiv.org) An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor.
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal (arxiv.org) Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of…
The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys (arxiv.org) Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a la…
Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security (arxiv.org) Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calli…
Oculi: A Conversational Agentic Platform for Automated Credit Risk Analysis (arxiv.org) Credit risk analysis in financial institutions traditionally requires analysts to manually write SQL queries, run statistical computations, and build visualization dashboards. This is a time-consuming workflow that limits exploration to fa…
ASTRA - Agentic System for Ticket Resolution and Analysis (arxiv.org) Technical operations teams resolve large volumes of incidents by synthesizing fragmented evidence from ticket text, historical cases, system logs, and technical documentation. Existing automation often relies on monolithic generation witho…
From Extraction to Governed Memory: Multi-Agent Knowledge Graph Construction with Domain-Expert Review (arxiv.org) Knowledge graphs used by agentic systems are often treated as flat stores of extracted triples, with little record of who owns a fact, why it was admitted, or how it should be used downstream. We argue that reliable agentic knowledge syste…
Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models (arxiv.org) The growth of agentic AI has drawn significant attention to function calling Large Language Models (LLMs), which are designed to extend the capabilities of AI-powered system by invoking external functions. Injection and jailbreaking attack…
ShardMemo: Scope-Before-Routing for Agentic Memory Retrieval (arxiv.org) Agentic systems accumulate persistent memory across sessions, tools, and tasks, and a later request must retrieve from it under two distinct constraints: which memories it is permitted to access, and which are relevant under a limited sear…
Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning (arxiv.org) Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by…
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning (arxiv.org) Agent skills are callable procedural modules that provide reusable knowledge and execution policies for complex agentic tasks. However, existing methods mainly focus on selecting relevant skills or improving the skills themselves, while ov…
CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance (arxiv.org) Large language model (LLM) systems are increasingly used to support high-stakes decision-making, but they typically perform worse when the available evidence is internally inconsistent. Such a scenario exists in real-world healthcare setti…
S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation (arxiv.org) Spectroscopic structure elucidation is central to molecular analysis, but recent Large Language Model (LLM)-based methods mostly formulate it as direct spectrum-to-SMILES generation. Although this paradigm can leverage paired spectral data…
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation (arxiv.org) Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and…
An Agentic Retrobiosynthesis Framework with Learned Frontier Selection (arxiv.org) Large language models are increasingly used as agents for multistep retrosynthesis, raising the question of how much their search policy contributes independently of the underlying reaction model. We investigate this question in a biologic…
VibeJam: A User Study Platform for Web Development with Agents (arxiv.org) Programming with AI is increasingly agentic, users prompt LLMs to directly edit their code and review the changes, with adoption growing especially for web development tasks. Despite this growth, most NLP work uses offline evaluation and l…
You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding (arxiv.org) Collaborative conversations frequently contain references whose targets are indirect rather than named: resolving "this looks like the fix discussed yesterday" requires combining conversational context with evidence from the surrounding wo…
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture (arxiv.org) Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Terminal-Bench-LILT, a suite of 300 authentic coding tasks in ten languages: Arabic, Czech, Germ…
HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees (arxiv.org) Agentic reinforcement learning (RL) often produces irregular rollout trees with shared histories. Training root-to-leaf trajectories independently recomputes these shared prefixes.
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL (arxiv.org) Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing…
PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems (arxiv.org) Large language models are increasingly used as agentic workflow executors, yet existing training data and benchmarks largely assume informationally complete, single-turn queries. Our analysis of 16K real-world sessions shows that 75.9% of…
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis (arxiv.org) Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS, a…
Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning (arxiv.org) A visual metaphor constitutes a high-order form of human creativity, employing cross-domain semantic fusion to transform abstract concepts into impactful visual rhetoric. Despite the remarkable progress of generative AI, existing models re…
Aligning Agentic World Models via Knowledgeable Experience Learning (arxiv.org) Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural grounding to respect the immutable laws of the physical world. Consequently, while these agents implicit…
PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering (arxiv.org) Retrieval plays a central role in multi-hop question answering (QA), where answering complex questions requires gathering multiple pieces of evidence. We propose PRISM, an agentic retrieval framework that leverages large language models (L…
Real-Time AI Service Economy: A Framework for Agentic Computing Across the Continuum (arxiv.org) Real-time AI services run across the device-edge-cloud continuum, where autonomous AI agents generate latency-sensitive workloads, orchestrate multi-stage pipelines, and compete for shared resources under governance constraints. This artic…
Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction (arxiv.org) One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, w…
LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages (arxiv.org) Landing pages are goal-oriented web interfaces that must communicate a target-specific value proposition while organizing information flow, visual hierarchy, and calls to action (CTA). Although large language models can generate plausible…
FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling (arxiv.org) Recent advances in large language models are enabling autonomous clinical agents to perform increasingly complex electronic health record (EHR) modeling workflows. However, agents deployed at individual hospitals remain constrained by inst…
PACE: Publisher-Adaptive Content Extraction via Agentic Automation (arxiv.org) Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability. General-purpose extractors can be applied broadly, but they are often brit…
WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents (arxiv.org) Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images fro…
String: An Agentic OS Where Every App Is a Markdown File (arxiv.org) LLM agents have become a new class of software user, but every surface they work through was designed for someone else. Pages are built for human eyes, which can skim and ignore; tool schemas for programs, which pay nothing to carry defini…
Resource Constraints and Performance in Agentic AI Systems (arxiv.org) Progress toward more autonomous AI increasingly depends on agentic systems that combine a language model with tools, memory, state management, and multi-step execution. These mechanisms shape both task capability and operational burden.
See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs (arxiv.org) Discovering governing partial differential equations (PDEs) from observational data remains a core challenge across the sciences. Existing sparse-regression, symbolic-regression, and LLM-based approaches can be constrained by predefined li…
ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL (arxiv.org) Recent work has shown that reinforcement learning from execution feedback can substantially improve text-to-SQL performance, often enabling smaller models to match or exceed much larger systems. However, most existing approaches treat SQL…
Credo: Reusable Declarative Primitives for Agentic Workflows (arxiv.org) An LLM application depends on both a model and a harness: the program that determines what each call sees, how many calls to make, and which answers to trust. Coding agents can now discover strong harnesses by searching over candidate prog…
Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation Community (arxiv.org) Agentic AI has the potential to accelerate curation of biological databases and knowledge bases. However, uptake has been hindered by a number of challenges and obstacles, including access to agents and appropriate training.
SETU: An Agentic Ecosystem for Multilingual, Persona-Aware Communication Coaching (arxiv.org) Corporate training teams need scalable and explainable tools to improve workforce communication in multilingual settings. Existing systems often score text, audio, or video in isolation, or produce black-box outputs that are difficult to a…
Forge: a Claude Code plugin for agentic development workflow (www.reddit.com via reddit) This post was written using application of kinetic force to the rectangular protrusions on my silicon chip container. Shocking, I know.
The non-obvious lessons from a week running a fully-autonomous Claude Code agent (empty mandate + its own co-signed wallet) (www.reddit.com via reddit) I’ve spent the last few days running an autonomous agent on Claude Code — headless on a cheap VPS, waking on a cron schedule six times a day with no memory between wakes except the files it writes itself. It has a small real budget it can…
With Claude I can build anything - so I built a workbench (private OSS project) (www.reddit.com via reddit) TL;DR: benchbook is a plain-markdown wiki, versioned in git, that an AI agent reads and writes under a written contract — what it may create, what it must ask about, what it can never touch. Works with any agentic tool that can read and wr…
How to Build Agentic Graphs (www.reddit.com via reddit) Over the past 4 months of working with graphs, I've learned several major lessons about graph design the hard way. In this post, I want to share the main takeaways so you don't repeat my mistakes.
How do you manage quality when AI agents write code faster than humans can review it? (www.reddit.com via reddit) We moved to an agentic workflow this quarter. My position is that we should ship at whatever speed the agents can produce, since that is the entire point of paying for them.
I’ve written software for about 30 years. I've been a heavy coding agent user for the past 1+ year. What practical coding-agent questions can I help answer? (www.reddit.com via reddit) I've been mostly hands on coding professionally for 20+ years. I have taken time in between to lead teams, run product management or run enterprise pre-sales.
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions (arxiv.org) A frontier language model's acknowledged "helpful programming assistant" persona does not survive long agentic-coding sessions in the deployment regime that production products actually run. After hours of tool-using debugging, a model tha…
AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification (arxiv.org) Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for more nuanced factuality verification. Existing factuality verification methods do not capture…
INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment (arxiv.org) As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harmful actions under goal conflicts and pressures.
BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing (arxiv.org) Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing. Wearable devices offer longitudinal behavioral and physiol…
SPT: Skills as Pre-Training Data for Agentic Language Models (arxiv.org) Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training. These data provide direct behavioral supervision, but producing them requires task environments, execution, and verifi…
When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems (arxiv.org) Agentic recommender systems ground each decision of a large language model (LLM) in a persistent memory of the user, and in existing agents that memory is text: a narrative written and maintained by further LLM calls. Text limits this memo…
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation (arxiv.org) Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to th…
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents (arxiv.org) LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing…
GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL (arxiv.org) Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers and task formulation. While deterministic graph tools are invariant to such shifts, extracting topolo…
A Contract-Centered Architecture for Scalable and Manageable Agentic Runtimes (arxiv.org) Enterprise AI deployment is a coordination problem across business units, application and AI teams, testing, platform engineering, infrastructure, security, operations, and data governance. Use-case benchmarks show whether one agent comple…
From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities (arxiv.org) Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers.
BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click (arxiv.org) Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are hard to…
AI Control Scientist: LLM-driven Agentic System for Automated Control Design (arxiv.org) Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual parameter tu…
AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design (arxiv.org) Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine-learning systems through executable…
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling (arxiv.org) LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-a…
Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI (arxiv.org) Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can improve performance on complex tasks. However, many existing approaches rely on fixed or preallocated reasoning controls, such as fixed…
Agentic AI for operating scientific instruments for nanoscale characterization (arxiv.org) Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert decision-making. A trained user defines the experimental intent, translates it into instrument commands, assesses incoming data, adjusts…
Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset (arxiv.org) Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipe…
Appreciation Post - thomsonreuters/Thomson-1.0-Small (www.reddit.com via reddit) With the lack of support from Qwen regarding the smaller 9B and 35B MOE models. Like myself, not everyone is looking for an agentic coding model, I particularly use it for RAG and reviewing and require high reasoning across different docum…
Anthropic's new hardware standard lets AI agents control the physical world (arstechnica.com) For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other data and actions that take place inside a computer. Anthropic…
Cowork is the most useless agentic interface I have ever used (www.reddit.com via reddit) Have ppl at anthropic stopped working on it? The UI honestly looks great but it just doesn't work.
We’re the Team Behind Apodex 1.1 — Ask Us Anything! (www.reddit.com via reddit) Hi r/LocalLLaMA ! We’re Apodex, the team behind Apodex 1.1, our new model family built to scale agentic intelligence for complex work.
OSCAR: Optimization-Steered Agentic Planning for Composed Image Retrieval (arxiv.org) Composed image retrieval (CIR) requires complex reasoning over heterogeneous visual and textual constraints. Existing approaches largely fall into two paradigms: unified embedding retrieval, which suffers from single-model myopia, and heur…
TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding (arxiv.org) Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU-Agent, an agentic retrieval-augmented framework for t…
VisDocAgentBench: Benchmarking Agents for Visually Rich Document Retrieval (arxiv.org) Visually rich documents encode relevance through language, layout, structured visual elements, and corpus context, yet retrieval is typically evaluated by one-shot query--page matching. Agentic-search benchmarks usually score downstream qu…
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models (arxiv.org) A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals.
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents (arxiv.org) Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameteriz…
Can your AI agent be cheaper? Investigating the effects of task specifications on token spend in agentic coding tasks (arxiv.org) Agentic coding workflows are now widely deployed in real-world systems. With long-horizon reasoning and tool use, token usage has become an important consideration for both cost and efficiency.
Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role (arxiv.org) Designing machine learning algorithms for wireless resource management is labour-intensive: the architecture, the loss function and the training recipe are all specified by hand. We demonstrate that this design layer can be surrendered to…
TailSFT: Filtered Fine-Tuning Improves Post-Training Performance (arxiv.org) Reinforcement learning post-training drives reasoning and agentic capabilities in modern AI systems, yet a growing body of work shows that it is most effective when used to fine-tune an already capable base model. We question whether exist…
Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory (arxiv.org) Structured memory representations such as knowledge graphs are central to autonomous agents and other long-lived systems. However, most existing approaches model time as discrete metadata, either sorting by recency (burying old-yet-permane…
Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation (arxiv.org) Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are hard to assess: subtle clinical errors are often missed…
AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs (arxiv.org) Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy.
Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs (arxiv.org) Product linking, the entity-resolution task of mapping merchant product records to canonical catalog products, consolidates fragmented listings so downstream search, recommendation, and advertising see one clean entry per product. At marke…
VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following (arxiv.org) Multimodal instruction-following models require training data that is accurate, diverse, verifiable, and challenging. Existing synthesis pipelines typically follow a one-pass generate-and-filter paradigm, discarding feedback from failed sa…
Claude Desktop App - Projects are weird and the app is kinda half-baked (www.reddit.com via reddit) So, I'm used to using agentic AI for coding, where it is very powerful and extremely flexible. I'm now trying the Claude Desktop App for non-coding related tasks.
How to disable corporate "agentic system" plugin (www.reddit.com via reddit) I've got a Claude susbcription thorught work. Our genius CTO has created a plugin which brings in 84 skills.
Questions on optimism speed/intelligence on this rig (www.reddit.com via reddit) Rig: 3945WX (12C, 2 CCDs, no AVX-512) · 8×32GB DDR4-3200 · 4× 5060 Ti 16GB · PCIe 4.0. Agentic workload (Hermes Agent).
↯ Vllm↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4vllmdeepseekagentic
What's useful max tg for agentic coding? (www.reddit.com via reddit) Seems to me like more than 50-60 tps tg makes no sense in agentic use cases, because the model thinks faster and calls tools faster than my machine can execute the commands. Any further optimization brings less and less actual speed up in…
I built a queryable code graph in Rust for agents to save context budget (MCP support) (www.reddit.com via reddit) Hey r/LocalLLaMA, I'm the creator of ctx, which I'm releasing open-source (MIT) under my company, Eagle-Logic. Full transparency: the repo was authored in partnership with Claude.
Currently GLM 5.3 Flash matches with Sol 5.6 (Max) in Agentic Index (Artificial Analysis) (www.reddit.comhttps) could not extract summary
What non-agentic local AI programs do you run? (www.reddit.com via reddit) I have Handy for speech to text and I'll try audio.cpp soon. Also tried an embedder for semantic code search (qwen embed + qdrant + ZooCode).
little tool for offline wikipedia RAG (www.reddit.comhttps) I was bored and handwrote a tiny 100-line bash script to let an agent search for and read articles from an offline wikipedia archive during a regular chat. It's not particularly useful, but it's definitely neat and a big step up from llama…
We built a local AI work tool that runs Qwen3.6-35B-A3B on a 16GB Mac (update) (www.reddit.com via reddit) Hi everyone! I’m an intern at Icosa, a startup focused on making local AI accessible.
[Megathread] GLM-5.3-Flash - former ox-alpha (www.reddit.com via reddit) Megathread for discussing the release of GLM-5.3-Flash. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons We'll try to clean up future duplicates around th…
Mac mini m5 pro (64gb) or two 16gb 5060 ti (total 32gb vram) for local LLM (www.reddit.com via reddit) I’m planning to buy a new machine mainly for local LLM inference and agentic workloads, and I’m deciding between these two setups. Both cost roughly US$3,000 where I live.
A 27b model beating latest frontier models was not on my 2026 bingo card (www.reddit.com via reddit) https://preview.redd.it/kbsqh6f7molh1.png?width=730&format=png&auto=webp&s=068dbea9a50be634a369d54d8b27b781d020fab3 My experience with Qwen 3.8 for agentic tasks has been phenomenal but I personally feel that 3.7 flash is more reliable for…
Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch (arxiv.org) We present Crase, a bounded and inspectable alternative to deep research agents for scholarly search. Instead of an open-ended search loop, Crase queries a search engine once for seed papers, expands them along their 1.5-hop citation neigh…
Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research (arxiv.org) Deep research (DR) systems produce long-form cited reports by orchestrating multiple agents that search and synthesize information from the web. Citations are the primary mechanism for evaluating the faithfulness of these reports, yet curr…
ADE: Agentic Data Evolution Framework for Human-Centered Objectives (arxiv.org) Aligning large language models to human-centered objectives is difficult when targets are non-executable and context-dependent, limiting reliable verification and scalable supervision. Although synthetic data expands coverage, weak verific…
LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology (arxiv.org) Lung cancer tissue diagnostics is complex, as therapy decisions in precision oncology rely on the integration of histomorphological, immunohistochemical, and molecular features. Yet pathological assessment remains largely visual and semi-q…
Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal (arxiv.org) An AI agent's rebuild is only as good as the process that produced it. Prior work found that once a model is strong enough, a multi-agent rebuild pipeline loses to the simplest approach: giving the model the original code and one instructi…
SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL (arxiv.org) Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-leve…
Confident at the moment of action: belief miscalibration in LLM play under hidden information (arxiv.org) Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repea…
MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG (arxiv.org) Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching and when to answer. Existing RL-based methods rely on external supervision and overlook the agent's internal belief about whether th…
AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL (arxiv.org) Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewards, which assign a uniform advantage to every step and cannot identify which decisions led to success or failure. Self-distillation method…
ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation (arxiv.org) Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only \emph{flat}, absolutely positioned elements, so ag…
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval (arxiv.org) Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We present AgentWorld, a simulation framework combining (i)Big…
Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems (arxiv.org) LLM-based multi-agent trading systems, in which specialized agents collaborate through structured communication to produce trading decisions, are moving rapidly from research prototypes to live deployments that control real assets. The sam…
Recursive Agentic Reasoning (arxiv.org) Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are often evaluated in isolation, making their gains difficult to compare across models, benchmarks, and evaluation pipelines. We introduce a un…
Exploit More, Explore Smarter for Budget-Constrained Agentic Search (arxiv.org) Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation budget, because validation is expensive, generation requires multiple model calls, or both. In this regime, standard MCTS allocates…
Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search (arxiv.org) Automated fact-checking is essential for ensuring the reliability of public health information, yet the biomedical domain poses unique challenges. Validating biomedical claims requires rigorous interpretation of scientific literature, asse…
A new way to build more powerful AI : no training needed. The Artificial Civilization Scaffold (www.reddit.com via reddit) We seem to be in a situation where we cannot see the forest for the trees in the philosophy of how to make AI more capable. We are ignoring the only known working intelligence multiplier we have encountered : human civilization What if we…
The Claw Machine Effect (www.reddit.com via reddit) I wrote a paper about the “one more prompt” loop in agentic coding After spending a lot of time with Claude Code and other coding agents, I started noticing a pattern that felt strangely familiar: occasional spectacular wins, lots of “almo…
Vercel launched a cool tool that checks how agent friendly a site is. I tried it on my project and got 100 (www.reddit.com via reddit) Vercel launched a cool tool that checks how agent friendly a site is. I tried it on agent-manager.dev, followed its suggestions, and got 100.
Agentic development vs “adding AI” - what’s the split in the enterprise? (www.reddit.com via reddit) I took a sabbatical this year - I saw the promise of agentic coding and dove in, completely transforming my development practice. I moved up the responsibility ladder to develop team lead and product management skills and culminating in an…
Agentic-Kube: A Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-Objective Kubernetes Scheduling (arxiv.org) Cloud-native container orchestration requires resource schedulers capable of balancing infrastructure expenditure, fault resilience, and node utilisation. Conventional reinforcement learning approaches typically rely on monolithic single-a…
Safety Training May Persist Through Helpfulness Optimization in LLM Agents (arxiv.org) Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step, tool-use) setting where safety refers to harmful actions…
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks (arxiv.org) Personal intelligence is becoming a central frontier for user-facing AI agents. To be helpful in everyday life, agents must understand users across the digital contexts where their preferences, intents, habits, social relationships, and ne…
Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation (arxiv.org) Large language model (LLM) agents operate in dynamic environments where knowledge continuously evolves. Existing memory systems typically treat external memory as a monotonically growing repository, inevitably leading to retrieval degradat…
MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning (arxiv.org) Multimodal Retrieval-Augmented Generation (RAG) with visual citation is crucial for ensuring the traceability and verifiability of MLLMs. However, current RAG and SFT-based methods struggle to achieve robust cross-modal reasoning, causing…
DeepRefine: Agentic Knowledge Refinement via Reinforcement Learning (arxiv.org) External knowledge enables large language model (LLM) agents to ground their actions and decisions beyond intrinsic parametric memory in open-ended, knowledge-intensive downstream tasks. Yet the quality of the underlying knowledge bases is…
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair (arxiv.org) Bug Reproduction Tests (BRTs) have been used in many Automated Program Repair (APR) systems, primarily for validating fixes and aiding fix generation. In practice, when developers submit a patch, they often implement the BRT alongside the…
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games (arxiv.org) Large Language Models (LLMs) have shown potential in simulating human behaviors and performing theory-of-mind (ToM) reasoning, crucial for complex social interactions. We investigate ToM reasoning's role in aligning agentic behaviors with…
MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition (arxiv.org) Automated data science is a structured model-selection problem. A solution must choose data transformations, feature representations, architecture, training procedure, evaluation protocol, and refinement strategy for a task.
ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation (arxiv.org) Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intuitive way to convey complex information. Current paradigms rely on either image generation or retriev…
MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters (arxiv.org) Time series forecasting (TSF) is evolving toward multimodal and agentic settings, yet using foundation models remains uneconomical in resource-constrained scenarios, where compact, specialized forecasters are more desirable. However, light…
Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs (arxiv.org) Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does no…
SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning (arxiv.org) Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or coarse-grained…
EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning (arxiv.org) Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a singl…
Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents (arxiv.org) Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM unlearning:…
Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing (arxiv.org) Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures.
Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models (arxiv.org) Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings. This paper investigates a critical question: does subjec…
Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning (arxiv.org) Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinge…
Apodex 1.1: Scaling Agentic Intelligence for Complex Work (arxiv.org) General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiab…
Concepts for Securing Agentic AI Coding and the Terok Environment (arxiv.org) Agentic AI is a fascinating new tool for software development. It is a huge step forward compared to "conventional" AI assisted coding, which in turn was a considerable breakthrough earlier.
Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses (arxiv.org) Robustness to a bad tool return means answering it in the way that return calls for, which depends on how the tool went wrong. A tool that has failed and a tool that returns a well-formed falsehood are different problems with different rem…
A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systems (arxiv.org) Energy system operation contains a loop of work that automation has never taken over: posing the optimization problem the current cycle should solve, disposing of infeasibility, sequencing a solution into interlocked switching orders, asse…
STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control (arxiv.org) Policy-governed agents must interpret case evidence while following an authorized procedure. We present \textsc{Stage}, an executable-graph framework that confines model judgment to policy-scoped nodes while placing procedural control in d…
Small Reasoning Models are Instruction Followers in Function Calling (arxiv.org) Function calling represents the core capability of agentic large language models (LLMs). Existing research has focused on enhancing LLMs function-calling accuracy through fine-tuning, reinforcement learning (RL), and multi-agent frameworks…
↯ Fine Tuning↯ Function Callingfunction-callingfine-tuningagentic
TessIndex: Capability Verified Identity System for the Agent Economy (arxiv.org) Software systems have traditionally been organized around applications where human users act as principal decision-makers. Recent developments in agentic capabilities alter this paradigm: software agents now autonomously translate high-lev…
LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization (arxiv.org) Large language models have become increasingly capable agents for low-level code and kernel optimization, but isolated kernel benchmarks provide only a proxy for the deployment behavior that matters in language-model inference. We identify…
ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling (arxiv.org) Transformer-based models are widely used for clinical prediction from electronic health records (EHRs), yet their architectures still require substantial manual tuning, and the optimal configuration may vary across tasks and hospitals. Neu…
Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities (arxiv.org) Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the…
SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG (arxiv.org) Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores. Exposing all tool descriptions to an LLM agent, or selecting tools only by vect…
How useful is a 5090 if I already have a 3090? (www.reddit.com via reddit) My use case is agentic coding. I'm a developer by trade and I like having a home lab for projects.
Harness for non-coding tasks (www.reddit.com via reddit) What's the best harness for agentic workflows that's not coding related at all? My work involves digesting a set of documents, analyze/evaluate them, and produce certain set of work product documents, mostly for due diligence purposes.
Planning to spend ~$100 benchmarking differnet Qwen3.8-27B quants and kv cache and looking for input before I start (www.reddit.com via reddit) TL;DR: I'm planning to spend around $100 on cloud GPUs to benchmark Qwen3.8-27B with a focus on questions that actually matter when running it locally: different quant levels/providers, 8-bit vs 16-bit KV cache, GGUF vs EXL3, context lengt…
Do not blindly delete your older models, some are still precious (www.reddit.com via reddit) I have deleted tons and tons of older models to make space since I can't afford storage anymore. Easily 10TB...
Getting ~11.7 tok/s from Qwen3.8 27B across an RTX 4070 Ti and M5 MacBook Air. Any ideas to push it further? (www.reddit.com via reddit) I have Qwen3.8 27B running across two machines with llama.cpp RPC. The main PC has an RTX 4070 Ti with 12 GB VRAM, and the worker is an M5 MacBook Air with 16GB unified memory.
TielCoder's 22 GB 4-bit quant matches Opus4.6 medium on recent real life coding issues, surpassing KAT-Coder and Nail as strongest and fastest MoE picks. (www.reddit.comhttps) Qwen3.8-27B is amazing, but it’s slow. A stronger 35B-A3B Mixture of Experts-coder that can run and solve real codebase issues fast (even on constrained hardware) is a valuable addition to the arsenal.
Real local agentic coding on a 12GB VRAM budget. (www.reddit.com via reddit) Thanks to Unsloth Dynamic 3.0 quants coming in slightly leaner and better preserved, I settled on Qwen 3.8 27B (`UD_Q4_K_XL`) at 100K context as my daily driver for Hermes Agent and OpenCode. On an RTX 5070 Ti Mobile (12GB) paired with an…
At a certain point, speed >> smartness (www.reddit.com via reddit) It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment (arxiv.org) Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based metrics such as BLEU a…
Metag: A dataset to build agentic meta-reviewing capabilities (arxiv.org) AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth in conference submissions has increased the burden on meta-…
STS: Efficient Sparse Attention with Speculative Token Sparsity (arxiv.org) The quadratic complexity of attention imposes severe memory and computational bottlenecks on Large Language Model (LLM) inference. This challenge is particularly acute for emerging agentic applications that require processing multi-million…
AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification (arxiv.org) High-performance ML systems increasingly rely on GPU kernels whose editable source is unavailable, generated, or too distant from final machine code to expose remaining optimizations. Existing LLM kernel optimizers and autotuners mainly op…
AgentOCR: Reimagining Agent History via Optical Self-Compression (arxiv.org) Recent advances in large language models (LLMs) enable agentic systems trained with reinforcement learning (RL) over multi-turn interaction, but practical deployment is bottlenecked by rapidly growing textual histories that inflate token a…
Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation (arxiv.org) Agentic systems generate outputs faster than human review. We contrast two LLM evaluator specialization strategies: specialized judge weights, or rule-based deferral policies for safe judgment acceptance.
$Z^2$-ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RAN (arxiv.org) With the progression in open and disaggregated 6G radio access networks, it is expected that the system will be able to host multi-vendors. In order to host multi-vendors, it is essential that AI-assisted control loops remain safe, verifia…
BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP (arxiv.org) Agentic engineering systems have shown strong performance on general-purpose benchmarks, yet their effectiveness in enterprise resource planning (ERP) domain-specific languages (DSLs) remains underexplored. We introduce BC-Bench, a benchma…
ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection (arxiv.org) Static analyzers have been widely adopted for vulnerability detection in C/C++ programs. Query-based static analyzers (e.g., CodeQL) encode vulnerable code patterns in detection queries and match them against source code.
When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation (arxiv.org) Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval error at hop 1 can surface only as a wrong answer at hop 3, while later retrieval can also repair the tr…
Testing and Evaluation of Agentic AI Systems In Military Command and Control (arxiv.org) Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requ…
ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib (arxiv.org) Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five dimensions beyond correctness: library…
Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI (arxiv.org) Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering enables…
Edge-Based Agentic Retrieval-Augmented Generation for Autonomous FHWA Bridge Inspection Compliance (arxiv.org) The Federal Highway Administration (FHWA) mandates that over 600,000 bridges in the United States be evaluated against the Recording and Coding Guide for the National Bridge Inventory (NBI). Manual compliance verification is labor-intensiv…
Personalized Privacy Control in LLMs via Attention Head Intervention (arxiv.org) The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms.
TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding (arxiv.org) Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often attribute-sparse: the attributes shoppers and downstream systems rely on are either buried in unstructured content such as titles and images…
CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting (arxiv.org) Search Agents face a severe reliability crisis during reinforcement learning (RL) fine-tuning. Heuristic Top-K retrieval often causes critical evidence loss or noise inclusion, while over-confidence induced by progressive RL leads to hallu…
VortexChat: An agentic framework for autonomous multi-objective integrated photonic design (arxiv.org) The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that rely heavily on manual simulation and expert intuition. While inverse design offers an alternative, it remains constrained by expert…
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills (arxiv.org) Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and…
When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory (arxiv.org) Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval failure mode…
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory (arxiv.org) Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primar…
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications (arxiv.org) Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful…
Benchmark results: what is the best and fastest engine to run Qwen3.8-27B on macOS (www.reddit.com via reddit) The new Qwen 3.8 27B is fantastic for local agentic use. The problem is, what makes it so good, being a dense model, also makes it slow.
Has anyone tried agent-lightning? (github.com via reddit) 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses! Documentation · Technical Report · MIT License Agent Lightning was completely refactored in v1.0.
Only ONE 450K session can keep its prefix cache on 2× DGX Spark — a second session wipes it with 43% of the KV pool still free. 6-minute cold prefill every turn. What am I missing? (www.reddit.com via reddit) **Setup:** 2× DGX Spark (GB10, 121 GiB unified each), TP=2 over 2×200GbE RoCE, vLLM 0.25.2.dev0, DeepSeek-V4-Flash-0731 FP8, `max_model_len=450000`, prefix caching on. KV pool = **1,686,693 tokens**.
Lattice: An isometric game kit for agents (www.reddit.comhttps) Lattice is a collection of typescript packages, agentic skills and plugins that enable easier development of isometric games! At its core lives a 0 dependency typescript package, 80kb gzipped.
Another Reason To Use Local: Active Sabotage/Derail By The Closed Src Models (www.reddit.com via reddit) Get out your tin-foil hats and local GPUs There have been lots of comments recently (at work, with, friends and here on reddit) about the apparent huge and sudden downward shift in the real-world usefulness of the popular paid Western AI p…
I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram (huggingface.co via reddit) Gemma 12B is obviously a very well trained model, I always thought the fine tuning they did on it wasn't really cut out for agentic coding. From my own experiences it struggles to use the tools it's given from Github Copilot and is also ve…
↯ Copilot↯ Ollama↯ Llama↯ Gemma↯ Gemma 4ollamagemmacopilot+2
Question for folks with r9700 (www.reddit.com via reddit) I have dual r9700 set up with Ubuntu. With llama.cpp I'm getting about 40 tokens per second for single instance.
What do people mean by "my harness" re: agentic coding? (www.reddit.com via reddit) I see a lot of posts on LinkedIn and other social media posts with folks at various companies talking about their harnesses. Are they talking about Claude Code / Codex, or are they building custom harnesses?
Update Open Source Multi Host Management For AI Jobs (www.reddit.com via reddit) I wanted to share a preview image of the upcoming open source release of the self hosted creative operating system I am working on. What can it do?
Current best model for narrative, chat, prompt creation (so basically everything except agentic coding)? - 5090 (www.reddit.com via reddit) Im looking to set up a new local llm (probably on unsloth studio as that seemed to be doing pretty well last time I tested it). This one won't need to do agentic coding or app building or anything (not this time) but instead more 'text' ba…
Why has my 3.6 35B become terrible now that I've started using 3.8 27B? (www.reddit.com via reddit) Up until the release of Qwen 3.8 27B (and getting my CMP 170HX system running it at 90-100t/s), I'd been totally happy with 3.6 35B - it did everything I needed (mostly agentic coding), very few errors, no doom loops and 512k context witho…
Claude Sonnet 5 vs. Gemini 3.7 Flash vs. Qwen3.8-Max : Which one is currently leading your daily workflow? (www.reddit.com via reddit) Hi everyone, With the latest wave of model releases, the battle for the ultimate daily driver has gotten ridiculously competitive—especially between **Claude Sonnet 5**, **Gemini 3.7 Flash**, and **Qwen3.8-Max**. Here is my quick breakdown…
I'm really hoping we're in 2026's 2-month-gap between QwQ and Qwen3 right now (www.reddit.com via reddit) QwQ was genuine next-gen performance usable on local hardware, but the massive required context (it's reasoning style was akin to "if I say every possible word, I'll notice the right one!") kinda made it unusable for agentic coding. It was…
Your multiple agentic workflows in one place with Origami (www.reddit.comhttps) Hiya. I've showcased Origami here before but with all the new features I thought I'd share a small video to showcase all of it briefly.
What’s the best general chat “harness” in Aug 2016? (www.reddit.com via reddit) Lots of discussion about coding harnesses, and there’s probably a lot of overlap, but for those allergic to CLI or even TUIs, but still think that tool calling and agentic use can make their general chat lives easier, what local front ends…
Qwen3.8-27B on an RTX 5060 Ti 16GB: IQ4 vs Q8, 64K context, MTP, vision, and agent benchmarks (www.reddit.com via reddit) I’ve been testing Qwen3.8-27B as a possible replacement for the Qwen3.5-9B that I have been running on RTX 5060Ti 16G. The goal was not just maximum tokens/sec, but useful context capacity, reliable tool calling, multi-turn behavior, and v…
Qwen3.8-27B Q6 is a beast at agentic coding (www.reddit.com via reddit) A quick feedback after a really major test: nearly 20 hours of non-stop goal-oriented work with Qwen3.8-27B Q6, running across an RTX 3090 and an RTX 3060. It maintained a speed of around 60–63 tokens/s throughout the session.
How the heck do you manage multiple sessions at the same time? (www.reddit.com via reddit) I feel like everyone is struggling with managing multiple parallel sessions within Claude Code. I've seen some people creating custom build solutions, but I'm also wondering how people just approach it within Claude Code GUI.
Which model is good for learning from Claude ? (www.reddit.com via reddit) Hi there, I’m a new grad in the chip design industry. right now, I’m trying to learn other stuff besides my actual role ( signal and power integrity) , which gives me a system level perspective and understand the whole architecture of a pr…
How does 2 × $200 buy $17,000 of Claude? (www.reddit.comhttps) It doesn't. My two Max 20× subscriptions — $400/mo — consumed $16,937 in API list-price tokens over 30 days, while costing Anthropic roughly $650 in actual compute.
An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction (arxiv.org) Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data coll…
MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents (arxiv.org) Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level signals into step-level credits through step grouping or graph…
A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries (arxiv.org) Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine an appropriate response. Although these queries may be linguistically clear, they can support multipl…
ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents (arxiv.org) Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states. We introduce \textbf{ReCache},…
Recent Claude Code Updates Reveal Anthropic's Agentic Vision (www.reddit.com via reddit) The changes that Anthropic has been making to Claude Code for the past few weeks indicate that Anthropic is building something much more powerful and sophisticated than what we're used to. More than just a coding agent capable of rewriting…
This is letting Claude handle a good amount of money for a month... (www.reddit.comhttps) II let Claude trade on my agentic account. Result: $31,000 lost.
Can someone from a techno-functional background (BA, PM) pass the Claude Architect Foundation exam? (www.reddit.com via reddit) My background is in techno-functional roles across business analysis, product/delivery, and data & AI transformation. I understand AI/LLM concepts and agentic workflows well, but I'm not a software engineer by trade.
Hot take: Claude code should ask more questions before touching your code. (www.reddit.com via reddit) Hey guys, i think one of the most underrated signs of a good claude code session is when it stops and asks you something before making changes people usually treat that as friction because the whole point is supposed to be moving faster, b…
Science Done on a Machine by a Machine: AI Agents in Computational Chemistry (arxiv.org) We are witnessing an explosion of agentic systems for computational chemistry simulations: from half a dozen in 2024 to a dozen in 2025, and the current number approaches fifty, surveyed in this Perspective as of 8 August 2026. The capabil…
One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI (arxiv.org) Agentic AI systems take consequential actions governed by more than one pre-action control at once: authority, resource, and evidence gates that can admit, degrade, or remediate an action before it executes. This paper's central object is…
What Makes Software Issue Resolution Tasks Difficult for Agents? (arxiv.org) Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks.
ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery (arxiv.org) Last-mile delivery aims to handle dynamically arriving orders with couriers while modeling complex spatial and temporal correlations. Recent learning-based methods model spatiotemporal dependencies among orders to predict courier service s…
RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training (arxiv.org) Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training rem…
Adversarial Review: Structured Disagreement for Grounded Agentic Code Review (arxiv.org) Early multi-agent LLM systems often used role-separated teams, yet scaling agent count yields diminishing returns on repository-level coding tasks. Recent alternatives treat agents as passive tools (subagents), yet this removes the benefit…
Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research Directions (arxiv.org) Agentic AI is gaining new insights and advancements in the field of Artificial Intelligence, fostering significant potential to enable rapid transformation across various this http URL rapid advancement and the potential to revolutionize v…
FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management (arxiv.org) Investment management is a high-stakes domain in which agentic AI systems must do more than generate plausible text. They must retrieve point-in-time data, assemble correct computational inputs, invoke specialized methods, and produce audi…
IAH: INTERNET WAR - Agentic Gameplay (www.reddit.comhttps) Hi guys! The RTS game that I have been working on for a few years is about to release this Friday.
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements (arxiv.org) Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight b…
SeqFeed: Improving Agentic RTL Code Generation with Sequential Behavior Feedback (arxiv.org) RTL code generation is a critical stage in hardware design, and the emergence of agentic systems offers new opportunities to automate this process. To generate correct RTL code, agents must understand sequential behavior, including how sig…
AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis (arxiv.org) Large language models are pushing data science toward increasingly autonomous and agentic workflows, with recent systems already supporting multi-step and long-running analyses. As these workflows become more autonomous, conventional inter…
Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement (arxiv.org) Converting structural framing plans into editable finite-element model drafts remains labor-intensive and prone to transcription error. Existing drawing-understanding systems for building components rely on task-specific trained neural det…
Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories (arxiv.org) We present Graphectory Viewer, a web-based tool for interactive, process-centric analysis of software-agent trajectories. Building on the Graphectory representation introduced in our previous work, Graphectory Viewer transforms heterogeneo…
Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Memory Leakage in Agentic Systems (arxiv.org) A personal language agent learns a fact from one audience and may later place it in the prompt it assembles for another. This memory-to-context step is an attack surface: ambiguous or inconsistent channels, cross-audience prying, and poiso…
Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting (arxiv.org) Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist,…
QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents (arxiv.org) Language-model agents increasingly produce quantum-science results; we ask whether the same agentic paradigm can also scrutinize them in an auditable, reproducible, and cost-transparent form. We present QuantumNovelty, an open-source skill…
Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating (arxiv.org) Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but al…
Agent Lightning v1.0: Towards Harnessed Agentic RL (arxiv.org) Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary…
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs (arxiv.org) Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectorie…
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation (arxiv.org) Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries.
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection (arxiv.org) Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem cla…
Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution (arxiv.org) Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects.
Subscription auth for third party use, instead of API (www.reddit.com via reddit) Is there still a path for low-volume free apps to use subscription auth? I ended up enabling OpenAI device logic from a PR for a small ereader plugin I maintain, and it seems OpenAI is apparently the only ones still tolerating it?
How Much Memory Does Your Agent Actually Need? (huggingface.co) How Much Memory Does Your Agent Actually Need? Equipping an agent with agentic memory sounds simple: distill lessons from its past work, put them back in context, and more experience should mean better performance.
SocialCoach: Personalized Social Skill Learning with Agentic Tutoring and Practice (arxiv.org) Social skills such as negotiation and leadership are crucial for personal and professional success in today's interconnected world. However, scalable and effective training remains a significant challenge due to the scarcity of expert coac…
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding (arxiv.org) Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesti…
HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory (arxiv.org) As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and retrieved, and ho…
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks (arxiv.org) AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, w…
Mint-Agent: Introducing Finance-Native Agentic Foundation Models (arxiv.org) Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Min…
Crystal-structure design by agentic AI in a language of motifs (arxiv.org) Data-driven materials discovery interpolates more reliably than it extrapolates and seldom reaches new structure types. We present MatEvolve, an agentic-AI framework designing crystals, proposing each candidate with a stated rationale and…
Belayer: Efficient Fault Tolerance for LLM Agentic RL Training (arxiv.org) Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-intensive rollout engines with stateful environment containers w…
Deploying Frontier Agentic Technology in MOOSEnger, a Multiphysics-Capable AI Assistant (arxiv.org) The Multiphysics Object-Oriented Simulation Environment (MOOSE) is an open-source finite-element framework for building multiphysics simulation applications. Using a multiphysics environment effectively demands specialized expertise, creat…
An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing (arxiv.org) In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint operation forms a hybrid scheduling problem, where physical logistics decisions are coupled with comput…
Agentic Test-Time Scaling for WebAgents (arxiv.org) Test-time scaling has become a standard way to improve performance and boost reliability of neural network models. However, its behavior on agentic, multi-step tasks remains less well-understood: small per-step errors can compound over lon…
Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory (arxiv.org) Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly master the individual skills, yet the chain still fails: errors compound beyond the policy's abilit…
Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning (arxiv.org) Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixture-of-Experts base model with Anchored Supervised Fine-Tuning on a compact corpus of verified, synth…
MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories (arxiv.org) Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' memories reconcile a fact phrased two ways, link related facts held apart, or reconcile contradictory k…
MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems (arxiv.org) Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data science workflows through natural language. Although these systems can significantly reduce manual…
Afterlife Delegation Protocol: Speculative Design of Self-Sovereign Agents that Outlive Their Principals (arxiv.org) Afterlife Delegation Protocol is a speculative design project that asks what death becomes when a will can act eternally. We design a speculative protocol through which a living person signs an agentic will: upon a verified death, a self-s…
From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems (arxiv.org) Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost,…
Hierarchical Agentic Incident Response with Digital-Twin-Validated Attack Inference (arxiv.org) Network incident response remains slow and labor-intensive as the defender must infer multi-stage attacks from partial observations and translate recovery decisions into reliable system commands. Decision-theoretic planners provide princip…
Workspace Topology as an Attack Vector in Agentic Coding Assistants (arxiv.org) Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party code. This opens up a risk of malicious code being ingested as these coding tools operate with broad…
Evaluating Agentic Code Repair Capabilities in Distributed Systems (arxiv.org) LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging, however, remains an under-explored regime: bugs span proces…
PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning (arxiv.org) LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language into the Planning Domain Definition Language (PDDL), allowing symbo…
A Policy Algebra for Trust-Preserving Agentic AI Execution (arxiv.org) Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools, delegate work, and complete a goal. Enterprise execution requires a stronger property.
AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems (arxiv.org) Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work spans many calls and stages. We introduce a unified execution model that maintains a work item's pers…
Competing at Every Price Point with Agentic Evolution over a Menu of LLMs (arxiv.org) Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every competitor price point. A firm that Pareto-dominated its competitors would leave no rational customer a reason to buy…
Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces (arxiv.org) Agentic retrieval workflows produce query, retrieval, and stopping traces as a byproduct of answering questions. We study how these traces can adapt a deployed dense retriever to changing workflow distributions without new relevance labels…
Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation (arxiv.org) Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where open-ended requests such as \emph{more NBA news} or \emph{less politi…
Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL (arxiv.org) LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system comparison fragile. We reframe the field as a leaderboard aggregation: we collect the metrics authors thems…
Understanding Cognition-Induced Risks in Agentic AI Systems (arxiv.org) Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human so…
ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning (arxiv.org) Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions, holidays, price changes, and platform interventions provide forward…
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (arxiv.org) Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines. Existing SQL correction approaches either rely on large-scale, high-q…
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads (arxiv.org) Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameterized by step count, tool invocations, and context expansion. We evaluate this decomposition empirica…
SCOPE: Score-Isolated Agentic Optimization for Video World Models (arxiv.org) Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time introduces a subtle evaluation problem: prompts, samplers, verifiers, and selectors may evolve together,…
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL (arxiv.org) On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not…
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs (arxiv.org) Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and pr…
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation (arxiv.org) AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single submission rather than the number of independent rollout attem…
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (arxiv.org) Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or…
Learning Agent Execution for KV-Cache Management in Agentic Serving (arxiv.org) Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these workflows, every agent repeatedly executes a fixed context…
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case (arxiv.org) Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance.…
How do you guys deal with being a stranger to your own codebase? (www.reddit.com via reddit) This question is for those of you doing agentic coding. Agentic coding can 10-20x speed of development and bosses are loving it.
How is your team handling AI’s design-clarifying-questions (www.reddit.com via reddit) We used to settle design questions as a team via an LLD before coding. Now agentic tools grill-me style skills) ask those same design questions directly to whoever’s driving and as a junior, I often can’t answer them confidently, yet each…
AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs (arxiv.org) Life science knowledge graphs make large collections of structured data available through SPARQL, but each resource uses its own schema, identifiers, and links. TogoMCP helps language model agents query these resources by providing curated…
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks (arxiv.org) We study parallel test-time scaling for long-horizon agentic tasks such as agentic search and deep research, where multiple rollouts are generated in parallel and aggregated into a final response. While such scaling has proven effective fo…
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning (arxiv.org) Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle complex queries requiring multi-step reasoning, agentic VRAG systems interleave reasoning with iter…
AtomBridge: Agentic VLA Inference Plugin for Long-Horizon Tasks in Scientific Experiments (arxiv.org) Robotic laboratories play a critical role in autonomous scientific discovery by enabling scalable, continuous experimental execution. Recent vision-language-action (VLA) models offer a promising foundation for robotic laboratories.
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use (arxiv.org) This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for…
Agentic Transaction: Towards ACID-Compliant Agent Systems (arxiv.org) Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation. As agents increasingly operate…
Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmen (arxiv.org) The rapid adoption of AI coding assistants and autonomous agentic development systems has coincided with major changes in the pace and structure of open-source software engineering. Yet empirical longitudinal evidence of these changes at t…
Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems (arxiv.org) When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full workflow…
SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning (arxiv.org) Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated spreadsheet reasoning remains challenging for large language models (LLMs). Real-world workbooks often contain implicit cross-table assoc…
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports (arxiv.org) In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to synthesize content, they often lack in information gr…
TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments (arxiv.org) Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing time series QA benchmarks mostly rely on fixed snapshots, lea…
Polaris : Multi Agentic System for Conversational Enterprise Analytics (arxiv.org) In today's fast-paced environment, the ability to swiftly access, understand, and act on data is no longer optional; it is essential. Yet most organizations remain data-rich but insight-poor, constrained by the complexity of querying, inte…
Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead (arxiv.org) Nanbeige4.2-3B is a 3B-parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers for a second forward pass, adding effective depth without additional parameters. Evaluated on Apple Silicon (MPS), we ide…
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis (arxiv.org) Agentic "Continual Learning Harnesses", systems that pair an LLM with retrieval or memory to improve from feedback without retraining, have shown growing value in cybersecurity. But their value is conventionally measured by gains against l…
I built a walking game with Claude, designed to make you put your phone away. (www.reddit.com via reddit) I've been building The Hidden Realm, a free-to-play UK location-based walking game inspired by Britain's landscapes, history and folklore. The slightly ironic part is that a lot of the development involved AI, but the game itself is delibe…
AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design (arxiv.org) Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show that large language models (LLMs), when integrated into well-designed frameworks (i.e.,…
Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design (arxiv.org) We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), coupled with symbolic computation tools, and human strategic direction, jointly produced a new result i…
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design (arxiv.org) Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human des…
Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language (arxiv.org) The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran).
Correct Is Not Governed: Provenance Integrity in Agentic Workflows (arxiv.org) Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale…
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting (arxiv.org) Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present Thyroid…
Research Assistant: AstraZeneca's Agentic System for R&D (arxiv.org) We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brin…
New agentic benchmark: Session-Bench compares what 10 coding harnesses preserve after the work is done (www.reddit.comhttps) SWE-bench measures whether an agent completed the task. Session-Bench measures what the harness preserved afterward.
ISO solutions to anthropic eco-system frustration for agentic work (www.reddit.com via reddit) Hi all, Over the last six months I've found that i'm spending a lot of time building workarounds for things that don't yet exist, but feel critical to improve efficacy of my setup. An example for this was building cross project memory 4 mo…
AI might be making you dumb but it made me smart (www.reddit.com via reddit) I keep seeing articles claiming that using AI makes people intellectually lazy. Apparently if you outsource too much thinking to LLMs, eventually your brain loses the ability to reason for itself.
What are the benefits to using Hermes with Claude? (www.reddit.com via reddit) I’m struggling to understand what Hermes, paired with Claude, can do better than just using Claude Code and Claude Cowork. Is it being able to fully automate agentic workflows with orchestration?
Scratching my head about Opus 5 (www.reddit.com via reddit) I have been using Claude Opus 5 for the past week and I don't know what to think about it. It just feels very hard to work with.
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking (arxiv.org) Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy. Small observed differences can be comparable to variabilit…
VALG: An Agentic System for ML Theory Research (arxiv.org) Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an op…
Intern-S2-Preview: Scientific Agentic Foundation Model (arxiv.org) Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-…
Flare, a graph-first IDE for agentic coding: watch the map change while Claude Code works for you (www.reddit.comhttps) I think we all went through this. Claude Code finished a task, told me it was done, and left me with 14 changed files and no idea which one mattered.
Claude Souls (www.reddit.comhttps) One of my benchmarks for my agentic game creation system is 'Time to Dark Souls'. Fable/Opus5 are still by far the best agents at 3d spatial reasoning so I thought I would share their best results here.
Grok 4.6's most important number is one xAI didn't even advertise (www.reddit.com via reddit) Grok 4.6 dropped yesterday and the debate is the usual "is it better than Sol?" On the headline benchmarks it's a genuine tie: Intelligence Index 61 vs 61, Coding 76.8 vs 77.4, Agentic 58.7 vs 57.8. The number nobody screenshots is AA-Omni…
↯ Hallucination↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6↯ Grok 4.6hallucinationgrokgpt-5+1
Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems (arxiv.org) Long-running conversational agents increasingly rely on a memory system to avoid resending the whole conversation each turn, yet how much that costs to serve has received little systematic benchmarking. We compare three memory systems (Mem…
Designing Agentic AI-Based Screening for Portfolio Investment (arxiv.org) We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers.
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching (arxiv.org) This paper describes an approach to hallucination detection and mitigation using a HOPE-inspired Nested Learning architecture with Continuum Memory Systems (CMS) and semantic similarity caching, tested on a hybrid benchmark of 310 prompts…
Tools as Continuous Flow for Evolving Agentic Reasoning (arxiv.org) Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a step-wise paradigm that lacks a global perspective, which causes error accumulation over…
Credo: Declarative Control of LLM Pipelines via Beliefs and Policies (arxiv.org) Agentic AI systems are becoming commonplace in domains that require long-lived, stateful decision-making in continuously evolving conditions. As such, correctness depends not only on the output of individual model calls, but also on how to…
DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation (arxiv.org) Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose context, game verification checks, and can produce large volumes of plausible yet invalid resu…
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence (arxiv.org) Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stocha…
The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark (arxiv.org) AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware, firmware, and proprietary applications, is avail…
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization (arxiv.org) Long-running agents periodically compact their context, replacing the transcript with a model-generated this http URL work shows that dropping a standing safety constraint during compaction drives behavioral violations acrossmany models (G…
Governing Agentic AI in FinTech (arxiv.org) Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated.
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation (arxiv.org) Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist e…
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS (arxiv.org) Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone. We propose an agentic workflow that t…
AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection (arxiv.org) Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volu…
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (arxiv.org) Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts.
A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization (arxiv.org) Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligan…
Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning (arxiv.org) Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic un…
Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning (arxiv.org) Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parame…
AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search (arxiv.org) Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an archite…
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop (arxiv.org) Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition…
How good is Claude pro, and how generous rate limits are? (www.reddit.com via reddit) For the past year, I’ve been using Gemini Pro for everything, but recently it started to feel lobotomized. It sometimes forgets everything we talked about before, even past messages.
Your interview questions assume candidates can afford Claude Code Max. AI access is deciding who gets hired. (leaddev.com via reddit) You have 1 article left to read this month before you need to register a free LeadDev.com account. Estimated reading time: 7 minutes Key takeaways: - Access to leading agentic coding tools is becoming a hiring filter.
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG (arxiv.org) Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only supervision. Most existing methods optimize policies from uniform…
When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning (arxiv.org) As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results?…
MIRA: Medical Image Reflection for Agentic Diagnosis (arxiv.org) Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations…
On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models (arxiv.org) Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agents operate…
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI (arxiv.org) After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refuse…
Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent (arxiv.org) Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus often f…
Self-evolving Agentic Customer Support System at LinkedIn (arxiv.org) Enterprise support agents operate in rapidly changing environments where policies, product capabilities, and knowledge bases evolve continuously, making static assistants brittle and costly to maintain. We present LinkedIn's self-evolving…
The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI (arxiv.org) Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that age…
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection (arxiv.org) Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external supervision. In this paper, we propose VDC-Agent, an autono…
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training (arxiv.org) Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This paper studies tool-calling along two complementary axes: effectiveness, i.e., how this capa…
Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding (arxiv.org) Agentic coding READMEs like this http URL grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always chea…
Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique (arxiv.org) Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. Existing detection methods fai…
TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling (arxiv.org) Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times. In this…
Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks (arxiv.org) Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy rollouts crea…
MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale (arxiv.org) LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods exploit t…
FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows (arxiv.org) Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures. However, constructing high-q…
InternAgentHarness: A Scalable Synthetic Environment for Enhancing LLM Agentic Abilities (arxiv.org) Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable and diverse environments that support repeated interaction wi…
InSight-doc: Agentic Visual Perception for Long-Document Understanding (arxiv.org) Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual resolutio…
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design (arxiv.org) Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic systems…
Has anyone actually pushed “AI writes almost all implementation code” to the limit on a serious project? Where did the verification-first workflow fail? (www.reddit.com via reddit) For about a year I’ve deliberately made AI my default implementation layer. I specify behavior, tests, invariants and acceptance criteria; models write/refactor/review the implementation.
For users who dont like this watermarking on apparently everything claude creates, what are the alternatives? (www.reddit.com via reddit) I know you can use local environments, deepseek etc to do the ai agentic coding for you, but what other options are there and what are the pro/cons of other solutions. Im considering switching because of the session limits spikes, weekly l…
Built an agentic pipeline for my business. It was dead for two months before anyone noticed (including me) (www.reddit.com via reddit) A long while ago I built a sentiment engine for my small services business. It was an agentic pipeline that swept our client Slack channels, email threads, and meeting transcripts, then scored each project.
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms (arxiv.org) Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit. In this paper, we evaluate the robustness of LLM-based agentic systems against attacks that aim to el…
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning (arxiv.org) Modern machine learning (ML) workloads increasingly rely on GPUs, yet achieving high end-to-end performance remains challenging due to dependencies on both GPU kernel efficiency and host-side settings. Although LLM-based methods show promi…
GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum (arxiv.org) Agentic knowledge graph question answering (KGQA) requires an agent to iteratively interact with knowledge graphs (KGs), posing challenges in both training data scarcity and reasoning generalization. Specifically, existing approaches often…
Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking (arxiv.org) Reranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial lexical gap between patient language and clinical nomenclature. We present a systematic co…
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning (arxiv.org) Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fi…
The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World (arxiv.org) LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged traject…
An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer (arxiv.org) Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many diagnos…
OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents (arxiv.org) Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse us…
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System (arxiv.org) Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Ex…
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning (arxiv.org) Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly available. It is essential for solving complex questions in specialized domains where retrieving compreh…
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver (arxiv.org) Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failure modes are repeat…
Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets (arxiv.org) Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equivalences, and hidden contradictions across markets. We present an agentic AI (AAI) pipeline t…
The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation (arxiv.org) The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools. The success of these…
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy (arxiv.org) Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose. Like ge…
STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework (arxiv.org) Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driven playbooks that encode fixed response procedures, but these static workflows struggle to adapt to…
Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations (arxiv.org) Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We study whether this representation can serve as operational context for agentic coding, whe…
Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data (arxiv.org) Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under strict computational and memory constraints. A critical yet underexplored challenge in this…
From Product Search to Preference Articulation: The Economics of Agentic Commerce (arxiv.org) Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents. We compare manual search, which accurately evaluates a limited product set, with agentic search, wh…
Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning (arxiv.org) Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Existing weather-aware food and point-of-interest recommenders, however, typically treat wea…
Agentic Auto-Research is Fuzz Testing (arxiv.org) Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers.
CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation (arxiv.org) Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions requires more than just generating individual comments, the generated threads should als…
Agentic Router: An Execution-Grounded Continual Learning Approach With Memory (arxiv.org) Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may still fail or introduce operational risk after execution. Existing approaches mainly focus on command ge…
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs (arxiv.org) Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that gap is still done by hand.
Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression (arxiv.org) Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates…
Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy (arxiv.org) The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven…
CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity (arxiv.org) Gamification is especially effective in learning domains requiring active problem-solving and iterative skill-building, such as cybersecurity education. Generative AI agents offer a path to delivering such experiences adaptively at scale,…
QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing (arxiv.org) Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the task, respect access and output models, expose required promises, and remain within a def…
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (arxiv.org) Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized d…
From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems (arxiv.org) Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmente…
Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems (arxiv.org) Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. Both can be inefficient because token cost, latency, redu…
Stop asking Claude Code what it just did. I built an open-source MCP server that feeds its recorded execution trail back into context. (www.reddit.com via reddit) Ask Claude Code what it did in a long session and you usually get a reconstruction from whatever is still in context. Usually close.
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving (arxiv.org) The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests. LLM serving platforms today define response-latency ser…
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning (arxiv.org) Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear.
PULSE: Agentic Investigation with Passive Sensing for Proactive Affective Intervention in Cancer Survivorship (arxiv.org) Cancer survivors face elevated rates of depression, anxiety, and emotional distress, yet self-report may be unavailable at some moments when support is relevant, a challenge we term the diary paradox. We present PULSE, a system for agentic…
Kimi K2.5: Visual Agentic Intelligence (arxiv.org) We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other.
AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation (arxiv.org) Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expertise to construct inputs, manage campaigns, diagnose failures, and validate results. We introduce A…
Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI (arxiv.org) Modern AI is no longer a single model but an ecosystem: classical ML predictors, deep and multimodal models, large language models, and agents, each trained and tuned over different data sources and each producing outputs at scale that bec…
An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation (arxiv.org) Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we propose a hybr…
HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation (arxiv.org) Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized effort. Even with high-level synthesis (…
Agentic AI: User Empowerment or Enclosure? (arxiv.org) Agentic AI promises a more flexible form of digital agency: systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it will empower users is an open question, and we argue that the…
Agentic Planning for Symbolic Execution (arxiv.org) Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreached. We investigate a complementary way of extending its practical reach by reasoning abou…
CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows (arxiv.org) Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design strategies, refine candidates, assess many properties, and gather evidence before synthesis and tests.
LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents (arxiv.org) AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial tasks and workflows. However, most of the research on their spatial capabilities (e.g.,…
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models (arxiv.org) Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complica…
Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation (arxiv.org) Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users exp…
ADIAS: Automated Design of Interactive Agentic Systems (arxiv.org) Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves t…
Meta is back with Muse Glimmer: local, agentic, multimodal, and open source (huggingface.co) Meta is back with Muse Glimmer: local, agentic, multimodal, and open source! To celebrate, we are shipping with Meta day-0 support in transformers , llama.cpp , vLLM , Inference Endpoints, and other libraries.
PostToolUse fires per tool call, so my typecheck hook was running once for every file in a batch (www.reddit.com via reddit) Setup is boring. PostToolUse in .claude/settings.json, matcher Edit|Write, command is a script under .claude/hooks that runs tsc over the project and exits 2 with the errors on stderr, so Claude gets them back in the same turn.
Cursor just announced Origin, "a git forge for the agentic era." The product is a waitlist, but the admission is the story: even the biggest AI coding company thinks GitHub isn't built for agents. (www.reddit.com via reddit) Cursor (Anysphere) quietly put up cursor.com/origin: "Origin, a git forge for the agentic era." Their framing is that "code is moving faster than any infrastructure was built to handle." It's early access, waitlist only. Set aside the prod…
Trying Agentic Loop with claude skills (www.ibrahimsajid.com via reddit) Hi guys, recently I started exploring agentic loops and how we can add certain workflows in skill files to automate tasks. Pretty much, it does things on its own until the final result is posted.
finance asked why our agentic ai best practices cost 8k a month (www.reddit.com via reddit) Finance flagged our AI tooling spend last week. It's about 8000 a month across the whole dev team once you add up all the model subscriptions - claude, codex, devin, cursor, coderabbit / bugbot, the lot Fair question, here's why I'm keepin…
how are you all handling agentic ai memory management across sessions (www.reddit.com via reddit) Small team thing, we're 6. Inside one session the agent is great, picks up our patterns, remembers why the auth layer is cursed.
Let AI speed up both sides of your code reviews, while you stay in full control (medium.com via reddit) Agentic skills that help me review other people’s code, and process the reviews I receive on mine. The AI suggests, I decide.
Does anyone actually have a fully autonomous coding agent that doesn't need constant follow-ups? (www.reddit.com via reddit) I've been trying to build a fully agentic software development workflow using Claude Code, and I've hit a frustrating problem. The first implementation usually looks good, but every time I ask a follow-up like: "Cross-check everything ag…
Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning (arxiv.org) In agentic workflows, LLMs frequently process retrieved contexts that are legally protected from further training. However, auditors currently lack a reliable way to verify if a provider has violated the terms of service by incorporating t…
Robust Native Language Identification through Agentic Decomposition (arxiv.org) Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than the underlying lingui…
Agentic Software Issue Resolution with Large Language Models: A Survey (arxiv.org) Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users, and represents a key aspect of software maintenance. With the rapid development of large languag…
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case (arxiv.org) Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the need for classic, physics-informed low-level vision?
F$^2$Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading (arxiv.org) With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial trading. Although recent advancements in Large Language Model (LLM)-based agents have en…
APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning (arxiv.org) Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. Pruning and quantization address this, but rely on manual, expert choices and on algorithm…
Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees (arxiv.org) Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately, since DSLs are often low-resource and esoteric, LMs frequently produce syntactically in…
Hierarchical Server Architecture for Agentic Science (arxiv.org) Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require specialized hardware within and across institutions.
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories (arxiv.org) LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that…
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations (arxiv.org) Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statemen…
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization (arxiv.org) As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automa…
QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction (arxiv.org) Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on sta…
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning (arxiv.org) Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult…
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation (arxiv.org) Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This misma…
From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems (arxiv.org) Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealize…
ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment (arxiv.org) This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervisio…
From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models (arxiv.org) Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interac…
When Agentic AI Meets Integrated Sensing and Communication (arxiv.org) Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work…
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution (arxiv.org) LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally…
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data (arxiv.org) Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines for such dat…
CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction (arxiv.org) CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed via MCP. Prior work validates such tools by checking whether predicted genes are kno…
Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks (arxiv.org) Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent…
Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services (arxiv.org) Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos and process fragmentation. Enterprises have invested considerable financial and material…
Opus 5: delete your CLAUDE.md? (no. don't) (www.reddit.com via reddit) I'm sure you saw titles like the one above and wondered: where is this even coming from? It comes from Boris Cherny's interview with Diana Hu at Startup School 2026 (source: https://www.youtube.com/watch?v=qyPCVqFUyDo ).
Using Claude to architect a custom agentic ecosystem (with custom memory & tools) – What are the modern standards? (www.reddit.com via reddit) I am in the planning phase of building a custom agentic ecosystem. I already have my own proprietary memory layer and a dedicated set of custom tools and agents that I want to integrate/build.
↯ Model Context Protocolmodel-context-protocolmcpanthropic+1
I built "Go Touch Grass", a tower defense game that's supposed to make gamers go outside (www.reddit.comhttps) Go Touch Grass is an idle RPG for Android built around one joke: it's supposed to make gamers go outside, so the only way to earn anything in it is by actually walking. Your steps convert into loot (wood/stone/rarer materials), XP and Skil…
TopoChunker: Topology-Aware Agentic Document Chunking Framework (arxiv.org) Current document chunking methods for Retrieval-Augmented Generation (RAG) typically linearize text. This forced linearization strips away intrinsic topological hierarchies, creating ``semantic fragmentation'' that degrades downstream retr…
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? (arxiv.org) Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its contribut…
EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot (arxiv.org) This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation (ERG) from text-only exchanges into live, face-to-face interaction. Through a video-cal…
stratum: A System Infrastructure for Massive Agent-Centric ML Workloads (arxiv.org) Recent advances in large language models (LLMs) transform how machine learning (ML) pipelines are developed and evaluated. LLMs enable a new type of workload, agentic pipeline search, in which autonomous or semi-autonomous agents generate,…
SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System (arxiv.org) Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execution strategy.
Formal Analysis and Supply Chain Security for Agentic AI Skills (arxiv.org) 32 pages, 5 theorems with full proofs, 68 references, open-source tool: this https URL. v2: corrects the bibliography (22 entries had author lists that did not match the papers at the cited arXiv identifiers; all verified against the arXiv…
Contextual Agentic Memory is a Memo, Not True Memory (arxiv.org) Current agentic memory systems (vector stores, retrieval-augmented generation, scratchpads, and context-window management) do not implement memory: they implement lookup. We argue that treating lookup as memory is a category error with pro…
XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs (arxiv.org) Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols. Unlike traditional structured generation with static structures, these workloads vary both across requests and within a requ…
A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination (arxiv.org) Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A…
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval (arxiv.org) Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight can help t…
EASy: Towards Efficient LLM-Based Agentic System (arxiv.org) Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing systems primarily optimize task success while giving limited consideration to execution eff…
Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) (arxiv.org) Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adapt…
AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering (arxiv.org) Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers limited transparency into its decisions and interactions. Many such systems also assume that users can…
EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis (arxiv.org) Modern chip design relies on electronic design automation (EDA) tools that generate large, heterogeneous artifacts, including source files, scripts, logs, netlists, and reports. Analyzing these artifacts is critical for debugging, optimiza…
Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming (arxiv.org) AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to perform operational tasks. As organizations adopt such systems, a critical challenge is iden…
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning (arxiv.org) Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, se…
Architectural Implications of Agentic AI Workflows (arxiv.org) Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a production study at Microsoft Azure and a…
Your Agentic Workflow's Cache Keepalive Costs 8x Too Much (v2: the interval frontier) (blog.mempko.com) Your Agentic Workflow's Cache Keepalive Costs 8x Too Much (v2: the interval frontier) Updated with the full retention curve and the interval frontier: the keepalive saves on Anthropic AND OpenAI inside each provider's paying band, OpenAI's…
Claude 5 sloppier than 4.8 (www.reddit.com via reddit) Opus 5 seems dumber in some ways than Opus 4.8. I've tried the same experiment in several models.
opus 5 is really shit (www.reddit.comhttps) I said what I said. Everyone hyped up Opus 5 and GPT-Sol for agentic workflows, but my social research bot kept hallucinating and failing context loops.
Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure (arxiv.org) An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files.
Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks (arxiv.org) Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting tha…
Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators (arxiv.org) Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art systems raise correctness through pipelines that couple LLMs with agentic reinforcement le…
$S^3$: Improving Agent Safety through Multi-Stage Defense (arxiv.org) Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks. However, risks may emerge at different stages, propagate across steps, and beco…
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure (arxiv.org) Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address the threats, little is known about the internals of agentic LLMs…
AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits (arxiv.org) As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with no 2D analogue, and post-route engineering change orders (ECO) remain manual, expertis…
Formal Verification of Agentic Systems over Operational Data (arxiv.org) Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data. Before deployment, these systems need to be verified against business requirements that…
From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities (arxiv.org) Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interpersonal relationships through visible collaboration. Generative coding agents (CAs) are…
SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation (arxiv.org) Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usable artifacts. Existing systems meet only part of this requirement: templates preserve regularity but…
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems (arxiv.org) Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distributed systems that reason, invoke tools, retrieve external state, delegate tasks, and act on behalf of…
TraceCAD: Trace-Guided Repair for Agentic CAD Generation (arxiv.org) LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, faulty operations, and prior repairs. We introduce TraceCAD, a recovery layer that links requested feat…
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting (arxiv.org) Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs…
VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space (arxiv.org) Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on standard benchmarks. We trace this ceiling to an incomplete debugging action space: existin…
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL (arxiv.org) Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context and database work before useful evidence appears, while post-hoc compression cannot re…
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection (arxiv.org) Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security assessment is deferred to later stages, delaying feedback and increasing remediation costs. Existing…
Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory (arxiv.org) As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single, unified set of values. Instead, systems must be able to recognize, represent, and respond to multipl…
KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization (arxiv.org) Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measurements are noisy, and search often stalls early. We present a practical optimization agent that combine…
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search (arxiv.org) We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of Lean proofs. LLM-generated proofs are notoriously correct-but-verbose and brittle across l…
SPEAR: Code-Augmented Agentic Prompt Optimization (arxiv.org) Automatic prompt engineering (APE) rewrites prompts to improve downstream task performance, but existing APE loops treat the optimizer itself as a fixed pipeline. We port the code-as-action paradigm of CodeAct (Wang et al., 2024a) to APE a…
Agentic Reinforcement Learning with Self-Distilled Reward Shaping (arxiv.org) Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions deserve credit. Training-only privileged skills can provide…
ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction (arxiv.org) Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and hypothesis generation. Traditional symbolic systems such as SemRep provide high precis…
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale (arxiv.org) Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspec…
In-Terminal Jupyter Notebook for Agentic Data Science (www.reddit.com via reddit) https://i.redd.it/k4w6l11w0hhh1.gif I've been a little frustrated with how Claude Code handles data analysis. It tends to create a pile of throwaway Python scripts while exploring a dataset, repeatedly rerunning the entire pipeline from sc…
Built an agentic tool loop for an in-browser coding environment. The verification step is where everything breaks. (www.reddit.com via reddit) The environment is file explorer, terminal, live preview, diff cards, chat, and autocomplete. The agent plans, edits, and verifies.
How to make Cursor do exactly what you want? (www.reddit.com via reddit) Hey folks, I have been picking up agentic programming recently. Our company uses Cursor and is on a tight token budget, so I'm using composer 2.5 mostly with a set of defined rules mostly.
Agentic Graph Token Reasoning (arxiv.org) Graphs model relational data throughout science and industry, from citation networks to product co-purchase graphs. Because the nodes of many such graphs carry rich text, a growing line of work applies large language models (LLMs) to graph…
Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch (arxiv.org) Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by generic statistical priors. Richer domain priors can improve BO in principle, but encod…
Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark (arxiv.org) Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. A controller may answer directly, decompose a request, retrieve evidence, execute code, delegate to a specialist,…
The Illusion of Stochasticity in LLMs (arxiv.org) In this work, we demonstrate that reliable stochastic sampling is a fundamental yet unfulfilled requirement for Large Language Models (LLMs) operating as agents. Agentic systems are frequently required to sample from distributions, often i…
Global Optimization and Inference-Time Region Grafting for Agentic Workflows (arxiv.org) Recent advances in agentic workflow optimization automate workflow design through task-specific workflow search or input-conditioned architecture selection. However, they determine the workflow before execution and cannot adapt failed work…
Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation (arxiv.org) On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gr…
DocNavRAG: Document-Structured Graph RAG with Stateful Evidence Construction for Complex Document Question Answering (arxiv.org) Answering complex questions over large document collections requires assembling complementary evidence across sections and documents. GraphRAG offers structured retrieval but typically uses fixed traversal, while agentic RAG operates over…
SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning (arxiv.org) Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods use execution correctness only as a trajectory-level reward, which provides limited guid…
A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use (arxiv.org) Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools are needed (missing). We find that a small, failure-specific set of MLP neurons could…
Bridging the Cognitive Gap: A Unified Memory Paradigm for 6G Agentic AI-RAN (arxiv.org) As 6G evolves, the radio access network must transcend traditional automation to embrace agentic AI capable of perception, reasoning, and evolution. A fundamental cognitive gap persists in current disaggregated architectures, where interfa…
MedTextWeaver: Procedural Knowledge Evolution in Agentic Medical Text Editing (arxiv.org) Medical text editing is essential for improving communication among diverse stakeholders in clinical settings. However, adapting LLM agents to this task remains challenging because expert supervision is often sparse, fragmented, and distri…
Rethinking Group Recommender Systems in the Era of Generative AI: From One-Shot Recommendations to Agentic Group Decision Support (arxiv.org) More than twenty-five years ago, first ideas were developed on how to design a system that can provide recommendations to groups of users instead of individual users. Since then, a rich variety of algorithmic proposals were published, e.g.…
GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation (arxiv.org) Lithology classification in well logs is a fundamental geoscience data mining task that aims to infer rock types from multi dimensional geophysical sequences. Despite recent progress, existing approaches typically formulate the problem as…
Mind the Sim2Real Gap in User Simulation for Agentic Tasks (arxiv.org) As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals. Yet, these simulat…
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use (arxiv.org) The integration of Large Language Models (LLMs) into the financial domain is driving a paradigm shift from passive information retrieval to dynamic, agentic interaction. While general-purpose tool learning has witnessed a surge in benchmar…
SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection (arxiv.org) Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabilities. By interacting with external environments, LLM agents can execute real-world tasks on beh…
Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment (arxiv.org) Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remains unreliable for fine-grained, visually ambiguous targets. We study this gap in the co…
Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning (arxiv.org) Incident response is currently managed by security operators using predefined playbooks, resulting in slow, labor-intensive security decision-making processes. Consequently, there is a growing need for automated incident response planning.
Antares: Foundation Models for Agentic Vulnerability Localization (arxiv.org) Vulnerability localization is a fundamental step in software security, requiring models to reason over large codebases and iteratively identify vulnerable implementations. We present Antares, a family of compact language models (350M, 1B,…
Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation (arxiv.org) A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time.
Agentic Self-Healing for Data and AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software (arxiv.org) Modern organizations rely on data, machine learning, and software delivery pipelines to move data, train models, deploy applications, refresh dashboards, and support business-critical decisions. However, these pipelines often fail because…
PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation (arxiv.org) The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-based methods to script-based methods for higher flexibility, portability, and maintainability. Howeve…
Deep Agentic Search for Repository-Level Code Question Answering: An Empirical Study (arxiv.org) Code agents spend much of their effort simply locating the right code inside a repository. Two approaches dominate current practice.
ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG (arxiv.org) Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed context construction may fail to translate these multi-resolution representations into a context suited to…
RefactorAssist: Agentic Refinement for Reliable Code Refactoring (arxiv.org) Code refactoring aims to enhance the internal structure of source code without affecting its functional behavior. The recent advancements of Large Language Models (LLMs) have demonstrated potential for automating software engineering tasks…
Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures (arxiv.org) Multi-agent LLM pipelines orchestrate multiple specialized language model agents into structured workflows where intermediate outputs are passed across agents to solve complex tasks. This design introduces a security gap absent in single-a…
MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing (arxiv.org) Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure. These meta-decisions affect not only task s…
Skillsets on the Chain: A Blockchain-based Zero-Trust Framework for Agentic AI Networking (arxiv.org) Agentic AI networking (AgentNet) systems rely heavily on third-party skillset implementations and distributed multi-agent collaboration, yet they face major claim-to-capability inconsistencies and security vulnerabilities under trust-by-de…
MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents (arxiv.org) Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation. Traditional prompt-based methods rely on descriptive conditioning by injecting static textual profi…
AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies (arxiv.org) The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, and prototyping a single policy takes m…
A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI (arxiv.org) Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated…
Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce (arxiv.org) In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to a…
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training (arxiv.org) Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagatio…
Before Reasoning Fails: Pre-Evidence Procedural Failures in Agentic RAG (arxiv.org) Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a procedural proper…
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning (arxiv.org) Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal…
Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness (arxiv.org) Meta-analysis synthesis highlights a fundamental challenge in knowledge-based scientific analysis: structured evidence does not by itself represent the analytical knowledge required for executable computation. Decisions about evidence assi…
Securing Agentic AI: From Per-Action Checks to Trajectory Assurance (arxiv.org) Autonomous agents are increasingly used to execute consequential tasks in environments governed by operational constraints, organizational policies, regulatory requirements, and technical standards. Their safety is therefore determined not…
V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory (arxiv.org) Interaction between users and LLM agents is increasingly multimodal: conversations interleave text with images, and a later question may target either. Yet most agent memories are designed around text, and even the few that support multimo…
Computing with Agentic Oracles (arxiv.org) This paper extends the stochastic-oracle model of AI-augmented computing to include agentic oracles. Unlike a stationary stochastic oracle, which responds to the same query according to a fixed response distribution across calls, an agenti…
Agentic Stage-One Stellarator Optimization: Autonomous Multi-Objective Search for Finite-Beta Equilibria (arxiv.org) Stage-one stellarator design searches a high-dimensional family of three-dimensional plasma boundaries and fixed-boundary MHD equilibria for configurations that jointly meet requirements on confinement, field-line topology, force balance,…
From AI Technical Debt to Agentic Technical Debt: A Systematic Mapping of Root Causes and Manifestations in Agentic AI Systems (arxiv.org) The emergence of Agentic AI systems, characterized by autonomous reasoning, multi-agent collaboration, tool orchestration, adaptive decision-making, and persistent memory, represents a fundamental shift from traditional AI pipelines to dyn…
Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable (arxiv.org) When a self-improving AI-for-science system claims a new capability, the evidence is usually a benchmark delta, a description-length gate, or a p-value. None separates a real gain from extra search, from a changed verifier, or from adaptat…
CADIR: A Cross-Backend Editable Intermediate Representation for Agentic CAD Generation (arxiv.org) Large language models have made it possible to generate executable computer-aided design (CAD) programs from natural-language descriptions or images. However, existing methods represent modeling processes as backend-specific sequential scr…
Assuming You Knew: Fixing an Epistemic Semantics for Flow Policies Using Agentic AI (arxiv.org) Many high-level security requirements are about the allowed flow of information in programs and are difficult to make precise because they involve selective downgrading. Notions from epistemic logic have emerged as a good approach to polic…
AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints (arxiv.org) We present AgentSLABench, a resource-aware evaluation framework for autonomous AI agents that measures correctness alongside latency, cost, compute, memory, and network usage under declared resource budgets. Unlike standard benchmarks that…
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation (arxiv.org) Agentic AI systems are evaluated using automated benchmarks whose scores justify deployment decisions, safety certifications, and regulatory compliance claims. We present an empirical analysis demonstrating that these scores are systematic…
Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce (arxiv.org) Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents c…
How to (www.reddit.com via reddit) What are your ways to use agentic loops? I've just discovered the (foundational?) ralph loop article by Geoffrey Huntley in which he suggests to use: while :; do cat PROMPT.md | claude-code ; done This seems rather basic and I guess the co…
How to juggle 5+ agents with reasonable security? (www.reddit.com via reddit) I hear on X that the future of programming is running a team of agents in parallel. Okay, so I've tried it with Claude Code.
Closed the prompting loop with cloud routines, and removed the human (www.reddit.com via reddit) My usual agentic workflow consisted completely of prompting on Claude code. I eventually managed to automate the entire development, testing, pr, ci workflow so that Claude could just do it themself.
Is there any good free/open-source materials about this like it used to be about programming? (www.reddit.com via reddit) I became a professional developer by going through basic guides and tutorials about languages and frameworks back in the day, building and practicing, and discussing ideas, and over time learning new stacks and landing a job. Like many oth…
Do you believe there's an important difference between the "agentic loop" and the "reasoning loop" that's perhaps at play when it comes to people's perception of how good a model is? (www.reddit.com via reddit) For example, Claude 5 is designed to rely on greater amounts of agentic-loop-based self-correction to arrive at a result and has a fine-tuned primary "instinct" to unblock execution and power through. It's good for coding, clear goals, and…
GTA 6 first attempt. Far from perfect, but it's impressive what the right harness and agentic loops can build. (www.reddit.comhttps) I was experimenting with Matt Shumer's Gauntlet Loop and shared a quick demo of an old favorite game, Worms Armageddon, the other day. It was built from a single prompt that kicked off the entire loop.
Different caching strategies - Codex and Claude Code (www.reddit.com via reddit) I am working on agentic software generation, and I noticed that Codex is more efficient than Claude Code in terms of token usage. The two agents seem to take very different approaches to caching, which might explain the gap.
TokTier: Exact Stateful Tokenization for Agentic LLM Serving (arxiv.org) LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard because…
Agentic Harness for Real-World Compilers (arxiv.org) Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable automated bug repair, compiler bugs pose unique challenges due to their complexity, deep cross-d…
ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics (arxiv.org) Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models rem…
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios (arxiv.org) AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to oversimplistic SRE tasks and are unfortunately hard to exte…
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents (arxiv.org) Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-based rewards computed based on the final answers. Richer rewa…
AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair (arxiv.org) Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic AI approaches have shown promising results in automated program repair.
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems (arxiv.org) Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to…
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load? (arxiv.org) We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectorie…
AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction (arxiv.org) Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challenging. Existing representative methods mainly revise outputs through natural-language re…
Beyond Component Testing: Validating Agentic AI Systems (arxiv.org) Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, becaus…
Scaling Scientific Discovery Environments for Turn-Level Agentic RL (arxiv.org) Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains cons…
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems (arxiv.org) The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agentic AI, particularly in separating inference, orchestration, and exec…
Agentic UI (not a chatbot) (www.reddit.com via reddit) How to tell calude to have an app that have agentic UI chat. the AI assistant: it doesn't answer with paragraphs — it renders live interactive cards inside the conversation.
CMV: a 50/25/25% split is the ideal inference footprint for agentic work (www.reddit.comhttps) could not extract summary
I built a real self-evolving operating system: Fable-os (www.reddit.comhttps) This is not a fake bullshit "AI operating system" that runs in your browser. This is an agentic operating system that runs on bare metal, writes its own drivers, and evolves itself.
I wanted more RAM for Claude Code, so I built my own IDE in Rust (www.reddit.comhttps) One thing I’ve noticed after using Claude Code is that the IDE itself can end up consuming a lot of RAM and CPU—resources I’d rather leave available for the coding agent. After trying several IDEs, I decided to build my own in Rust.
Uncle Bob Martin's tweet relieved me about the agentic coding (www.reddit.com via reddit) I have been using Claude Code since its release in 2025. Before that, I used Cursor to explore AI-assisted coding.
Through the Agentic Looking Glass (www.reddit.com via reddit) Hi all, I've spent some time trying to understand what is possible with agent definitions. For the last 7ish months I've been building out various agents and deploying them on every project I've developed, mostly for adversarial verificati…
I've just build a Worms Armageddon clone with one single prompt. Agentic loops are mind blowing 🤯 (www.reddit.comhttps) It runs nicely in the browser: https://aifnet-public.b-cdn.net/games/worms.html Idea is based on Matt Shumer Gauntlet Loop The loop: " I want you to build a Worms Armageddon look alike game at the level of the most recent Worms Armageddon…
Security tools (www.reddit.com via reddit) Hey guys, do you guys use any tools to verify app security? I notice that within common agentic workflows, the part where you explicitly check for vulnerabilities isn't really there.
Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings (arxiv.org) An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, and an admit/block decision before the regulatory layer moves.
Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration (arxiv.org) Large language models can contribute useful ideas to mathematical research, yet long-horizon proof attempts remain difficult to coordinate, evaluate, and reproduce. We present Albilich, an open-source agentic harness for autoresearch in ma…
KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation (arxiv.org) Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task. The recent rise of LLMs and agentic frameworks offers a p…
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger (arxiv.org) Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cann…
CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories (arxiv.org) Recent years have witnessed the rapid development of LLM-based agents, which shed light on using language agents to solve complex real-world problems. A prominent application lies in business agents, which interact with databases and inter…
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents (arxiv.org) The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast.
Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution (arxiv.org) Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free outp…
Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories (arxiv.org) Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how or why.…
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation (arxiv.org) Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer reliability…
Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms (arxiv.org) Retrieval-augmented generation (RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling uncl…
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution (arxiv.org) Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning eit…
Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection (arxiv.org) Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real CVEs, we find that 71.7% of vulnerable functions require evidence from outside the functio…
Voice Memory for Agentic Speech Recognition (arxiv.org) We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain this http URL and decides per utterance whether to act on the hypothesis or abstain and keep the…
SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation (arxiv.org) Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payloa…
Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges (arxiv.org) Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It enables multiple agents to exchange arguments, critique each other's outputs, and iterativel…
SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search (arxiv.org) Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia sna…
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution (arxiv.org) Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware this http URL large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamental…
EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks (arxiv.org) Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs), yet their performance heavily relies on the manual, trial-and-error engineering of neural representations, loss…
GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure (arxiv.org) Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a costly respin. Recent large language models (LLMs) offer new opportunities to automate th…
Can't get our agentic workflow design patterns to stick (www.reddit.com via reddit) Hi, we've been struggling with the ai workflows on our team for a while now and i'm curious if this is just us Every time i'm like ok we're done, this is the way we're doing it now, something breaks or half the team quietly stops following…
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening (arxiv.org) We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing device and sends captured images, together with optional symptom descriptions, to a server-…
Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows (arxiv.org) Quantum computers are moving from research laboratories to industrial machines accessible via the cloud and integrated into high-performance computing facilities. However, translating theoretical quantum protocols into hardware experiments…
Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments (arxiv.org) We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond. NV centers are a widely used platform for quantum sensing, and the ability to con…
From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance (arxiv.org) Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive implementations encounter hard limits as corpus scale and query complexity grow. This paper t…
VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents (arxiv.org) Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures. This work studies multimodal retrieval…
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents (arxiv.org) Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this do…
Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents (arxiv.org) A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discret…
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following (arxiv.org) Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmark…
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning (arxiv.org) Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide services for users. However, unlike natural-world tasks that assume well-defined instructions, human-cent…
The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape (arxiv.org) Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a…
ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop (arxiv.org) Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions, track task progress, reason about spatial state, and recover from errors while performing physical…
How good is Sonnet 5 for agentic coding? (www.reddit.com via reddit) I opened up my terminal to load up a project Ive been working on with claude code, and for some reason the default model was set to Sonnet 5. This has never happened before, I always use the Opus models.
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI (www.latent.space) There are roughly 100x more people who use code than who can write code.1 As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right. A key trend we have be…
Agentic code review costs me more than actually shipping the code (www.reddit.com via reddit) nobody on my team really writes code anymore, we describe it and the agents write it. thats old news at this point.
Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis (arxiv.org) Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across millions of records that combine structured metadata, multilingual legal notices, temporal events, a…
AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use (arxiv.org) Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain similar contrast, and purely feed-forward models cannot verify physical validity. We present Aut…
A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility (arxiv.org) Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system dat…
Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation (arxiv.org) Cloud telemetry arrives at a scale that, paradoxically, makes intrusion understanding harder rather than easier. Attackers operate through legitimate identity, federated session tokens, and cloud native APIs indistinguishable from routine…
Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents (arxiv.org) Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of t…
False Prophets: On the Security of World Models in Agentic Systems (arxiv.org) Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions.
VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy (arxiv.org) Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers. Conventional retrieval-augmented generation typically…
Agentic Autoresearch for CT Reconstruction (arxiv.org) Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large language model (LLM) agent can do the labor of reconstruction research on its own, and whether…
EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff (arxiv.org) Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hide useful search behavior. We present EviBack, an evidence-…
ACM: Agentic Context Management for Long Horizon Tasks (arxiv.org) Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context compression methods inevitably incur information loss and are triggered by rigid heuristi…
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning (arxiv.org) Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn…
Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels (arxiv.org) As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a governance framework th…
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG (arxiv.org) Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space. This design often struggles…
Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure (arxiv.org) The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access control (RBAC) models cannot address. Unlike deterministic automation, AI agents exhibit stoc…
An Agentic Orchestration of Atomistic Simulations (arxiv.org) Atomistic simulations are central to materials design, but their execution involves complex, multi-step workflows that require significant human expertise. Here, we present an agent-based system embedded within the URSA (Universal Research…
HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization (arxiv.org) The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. However, existing solutions typically prioritize intra-workf…
SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs (arxiv.org) Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing agentic approaches that perform well on public benchmarks often fail to generalize to rea…
DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs (arxiv.org) Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields b…
Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness (arxiv.org) Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box…
DRC-Aid: Design-Rule Correction via Agentic Framework utilizing Inference-Time Large Language Models (arxiv.org) Resolving Design Rule Violations (DRVs) in layouts entails an iterative loop of geometric edits and verification. We present DRC-Aid, a closed-loop agentic framework that automates local DRC repair by formulating it as verification-in-the-…
When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost (arxiv.org) Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations often leave the operating point underspecified: two systems ma…
SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing (arxiv.org) Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology. Unlike bulk assays, scRNA-seq captures heterogeneous cellular states and rare subpopulations, b…
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks (arxiv.org) Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison…
Where Is the Cost of Third-Party API Routers in Agentic Software Development? (arxiv.org) Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows, high-autonomy operation is widely adopted because it reduces interaction overhead.
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications (arxiv.org) Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a framework…
Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations (arxiv.org) When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies propagate across prior turns? Existing context attribution methods process the full co…
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B (arxiv.org) Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses…
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation (arxiv.org) Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how plan…
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair (arxiv.org) Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it.
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search (arxiv.org) Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-$k$ content, but document relevance alone cannot localize, compose, or verify the evid…
Is agentic coding became slower recently? (www.reddit.com via reddit) I'm an active user and was able to ship relatively large (50k+ loc, c++/python) codebases with agents in real prod, so it's not like I'm fully new to this, but I may not know all the best bleeding-edge approaches. I was generally happy abo…
Building the enterprise environment for agentic AI (www.technologyreview.com) Sponsored Building the enterprise environment for agentic AI Enterprises will find success with a complete agentic AI environment where agents plan, retrieve, remember, and act reliably at scale. Provided byIntel For the enterprise, the pr…
How Do AI Coding Agents Contribute to Software Development? an Empirical Study of Agentic Pull Requests (arxiv.org) Recent advances in large language models and their rapid adoption across software engineering tasks have made Artificial Intelligence (AI) coding agents an integral component of modern software development workflows. While developers incre…
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning (arxiv.org) Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and roll…
Agentic Evaluation of Copyright Law Compliance (arxiv.org) Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content such as images and, where appropriate, reproducing that content. LLM agents should comply with the law, including copyright la…
SwiftMem: Fast Agentic Memory via Query-aware Indexing (arxiv.org) Agentic memory systems have become critical for enabling LLM agents to maintain long-term context and retrieve relevant information efficiently. However, existing memory frameworks often perform query-agnostic retrieval over the full memor…
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas (arxiv.org) Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic b…
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference (arxiv.org) Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer.
A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation (arxiv.org) Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and decision-making. However, the operational…
Towards Trustworthy and Cost-Efficient Data Integration: From Na\"ive RAG to Agentic RAG (arxiv.org) Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due t…
TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI (arxiv.org) Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call.
Agentic Root Cause Analysis through Evidence-Grounded Reasoning (arxiv.org) Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck.
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation (arxiv.org) Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality o…
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI (arxiv.org) Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for succe…
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode (arxiv.org) We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabili…
LeafData: An Agentic System for Data Migration (arxiv.org) Modern data migration relies on JSON configuration to define data connection, pipeline logic, and orchestration behavior. This requires domain knowledge from users and is time-consuming and error-prone.
Coupled Hierarchical Search over Topology and Execution for Agentic Workflow Synthesis (arxiv.org) Although structured workflows empower Large Language Models (LLMs) to tackle complex problems, automating their creation is severely hindered by a vast combinatorial search space, frequently resulting in inflexible and resource-heavy offli…
AgentKVShift: Efficient KV Cache Reuse for Agentic Memory Systems (arxiv.org) Memory-augmented LLM agents maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content with LLM-generated metadata such as summaries, keywords, and tags. From an inference cost st…
Does Claude Pro keep normal chat usage separate from Claude Code and Cowork? (www.reddit.com via reddit) I’m considering cancelling ChatGPT Plus for a month and trying Claude Pro so I can properly test Claude’s models and subscription limits. One thing I really value about ChatGPT is that Work and Codex use an agentic usage allowance.
what agentic coding tools actually stuck for your team? (www.reddit.com via reddit) what agentic coding tools actually stuck for your team? we're a 12 person product team and our setup is cursor + codex + claude code + coderabbit.
The real bio/cyber workhorse is Opus 5 (Forget Fable 5) (www.reddit.com via reddit) Anthropic dropped Claude Opus 5, and if you are doing computational biology or cybersecurity, this is the model you actually want. Fable 5 is supposed to be the "frontier" model, but it’s heavily safeguarded and aggressively blocks high-ri…
Opus 5 vs Fable 5 for coding on Max: has anyone actually compared them on a real codebase yet? (www.reddit.com via reddit) Opus 5 dropped today and Anthropic is claiming it’s the new SOTA on coding and knowledge work evals, ahead of Fable 5, at half the API price. Only place they say it’s behind is cyber and bio, where Mythos still leads.
Is using subagents as a proxy for API LLM calls to test output quality against ToS? (www.reddit.com via reddit) So I’ve built a basic agentic system using Claude Code and hooked it up to my Anthropic keys for its actual runtime. I wanted to run an output quality test across multiple simulated scenarios and what I did is let Claude Code automate the…
What skills, plugins and tools do you use to go from vibe coding to agentic software development? (www.reddit.com via reddit) So in my journey I’ve gone from stuffing everything into a prompt and letting it rip to using skills like superpowers to help me brainstorm, design and plan the work to things like speckit to get super formal about spec/test driven develop…
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework (arxiv.org) Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries.
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs (arxiv.org) Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically rely on extremely large backbones. Existing distillation approaches train smaller students to im…
Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation (arxiv.org) Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational policy constraints, not merely producing syntactically plausible configurations. We present a verif…
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems (arxiv.org) Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning to…
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks (arxiv.org) Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements tha…
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning (arxiv.org) In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization. Existing skill-centric methods improve exploration by o…
Regulating autonomous and agentic AI (arxiv.org) Regulating activities where regulatees use autonomous and agentic AI is challenging. Regulatory assumptions about regulatee knowledge and control no longer hold true; much of that lies elsewhere in the AI supply chain which thus needs to b…
HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices (arxiv.org) Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisation. The emergence of LLM agents creates a new opportunity for Personal Health Agentic…
Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis (arxiv.org) Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already check whether those citations hold, with a fixed attribution model or human graders. N…
MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference (arxiv.org) Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these applications often incur high inference cost. We present MiniCache, a reusable program cachin…
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics (arxiv.org) Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments. This paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-age…
Burned €85 in 30 mins on a single prompt. Let’s talk about the brutal economics of AI inference. (www.reddit.com via reddit) Not another credit rant—a genuine question about the macro-economics of AI. I’m a Claude Pro user and recently got an €85 credit for Fable.
Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval (arxiv.org) In the era of autonomous agents, machine-actionable data is critical for data-driven workflows. For more than a decade, semantic metadata like schema$.$org has anchored the FAIR principles (Findable, Accessible, Interoperable, and Reusable…
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning (arxiv.org) The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safety. Existing forensic methods mainly rely on visual cues and lack evidence-based reasoning to reve…
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection (arxiv.org) Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language models (MLLMs) with high-level semantic knowledge. Although naturally complementary, these two infor…
The Ethics of Autonomous AI Agents for Offensive Security (arxiv.org) LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and operated by trained practitioners -- agentic security tools exhibit \textit{indeterminacy}…
EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair (arxiv.org) Design rule check (DRC) closure remains a major bottleneck in advanced-node physical design. Although detailed routers are rule-aware, residual design rule violations (DRVs) often require manual engineering change order iterations.
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation (arxiv.org) Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations mainly focus on final-answer accuracy and may miss failures in the search trajectory.
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation (arxiv.org) A rich and recognizable component library is the cornerstone of printed circuit board (PCB) design and generation. Traditionally, engineers manually create symbols and footprints and design PCB schematics, which is time-consuming and error…
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI (arxiv.org) The transition from stateless generative models in artificial intelligence to stateful, autonomous agents represents an architectural evolution that, while providing the capabilities of long-term planning and the automation of enterprise w…
FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance (arxiv.org) Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasize measuring general capabilities, instruction following, or safety, but few directly address the op…
NMR Elucidation as an Agentic Search Problem, Not a Modeling Problem (arxiv.org) Structural elucidation from Nuclear Magnetic Resonance (NMR) data remains a fundamental bottleneck across chemistry, materials science, and biology. We demonstrate that an agentic AI system can perform this task at a level comparable to gr…
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use (arxiv.org) Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scal…
Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments (arxiv.org) A central goal of large language model (LLM) research is to build agentic systems that can plan, act, and adapt through sustained interaction with dynamic environments. While recent LLM-based agents exhibit impressive contextual reasoning,…
Node-as-Agent: Graph Agentic Network (arxiv.org) Graph Neural Networks (GNNs) have achieved remarkable success in graph-based learning by propagating information among neighbor nodes via predefined aggregation mechanisms. However, such fixed schemes often suffer from two key limitations.
Agentic AI-assisted coding offers a unique opportunity to instill epistemic grounding during software development (arxiv.org) The capabilities of AI-assisted coding are progressing at breakneck speed. Chat-based vibe coding has evolved into fully fledged AI-assisted, agentic software development using agent scaffolds where the human developer creates a plan that…
They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface (arxiv.org) We study a five-agent CI/CD pipeline (triage -> developer -> security-scan -> review -> approve/deploy), built from five distinct production LLMs across three providers, behind an LLM firewall in shadow mode. A single untrusted input - an…
Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation (arxiv.org) Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline on the PaySim dataset that combines a gradient-boosted classifier, graph-derived structural features,…
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents (arxiv.org) Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical paramet…
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling (arxiv.org) Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at sho…
Data Leakage Prevention in Agentic Applications via Preemptive Hardening (arxiv.org) Agentic systems integrate LLM driven planning with interfaces to external tools, making data leakage and tool misuse feasible via instruction/data boundary failures and prompt injection attacks. Enforcing required controls consistently is…
AgentTrails: Towards Trust and Reuse for Agentic Tasks (arxiv.org) LLM-powered agents increasingly tackle complex tasks by invoking tools, querying databases, executing code, and manipulating intermediate artifacts. These agents follow trajectories that are typically stored as chronological logs, obscurin…
Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative (arxiv.org) Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter space can be high-dimensional, and the search must respect plausibility constraints. Although the simul…
Agents in the Wild: Where Research Meets Deployment (arxiv.org) Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across d…
Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes (arxiv.org) This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes. Rather than treating LangGraph, a low-level orchestration framework for stateful agents…
Engineering Trustworthy Agentic AI for Critical Systems (arxiv.org) Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic conse…
MAGE: Human-Like Macro Placement via Agentic Multimodal Reasoning (arxiv.org) Macro placement still requires substantial manual refinement in industrial physical design flows. We present MAGE (Macro Placement Agentic Engine), a multimodal multi-agent framework for macro placement refinement.
From Agent Failure Paths to Quantified Residual Risk: A Compositional Framework for Resilient Agentic AI (arxiv.org) Agentic AI is crossing trust boundaries faster than current risk models can represent. Existing approaches provide one of two partial views.
if they remove cuz of karma, we can try r/claudecode (www.reddit.com via reddit) Hi, we've been struggling with the ai workflows on our team for a while now and i'm curious if this is just us Every time i'm like ok we're done, this is the way we're doing it now, something breaks or half the team quietly stops following…
Agentic Avengers: Multi-agent dev-workflow skills with Avengers personalities (www.reddit.comhttps) Hi everyone :), I built a set of skills and an orchestrator with avengers characters giving each skill a personality. They are pretty simple in that they make coding really autonomous and consistent.
Tried to build an agentic-engineering list that does not rot or bloat like every other awesome list (www.reddit.com via reddit) Every "awesome" list I bookmark eventually dies. The maintainer moves on and the links rot, or it balloons to a few hundred entries and I scroll past it because there is no signal left.
Claude is frustratingly slow (www.reddit.com via reddit) I've had Claude Max for past 2 months and it's been getting increasingly slower. It's unreal for me that small to medium tasks can take tens of minutes, and medium to big can take even hours.
FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting (arxiv.org) Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk.
Salience Induction against Multi-Hop RAG Agents: Threat and Defense (arxiv.org) Agentic retrieval-augmented generation (RAG) systems increasingly retrieve external evidence and orchestrate tools for knowledge-intensive applications. In Multi-Hop question answering, agents chain facts across documents.
FormulaCode: Evaluating Agentic Optimization on Large Codebases (arxiv.org) Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to optimize entire codebases under realistic constraints. Existing code benchmarks largely rely on syn…
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks (arxiv.org) As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challenge agents and why becomes increasingly difficult. This is…
Artificially intelligent agents in the social and behavioral sciences: A history and outlook (arxiv.org) We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral sciences: from the first programmable computers, and social simulations soon thereafter, to today's experi…
Benchmarking Agentic Newswriting via Journalistic Workflows (arxiv.org) Recent advances in autonomous digital agents from industry (e.g., Manus AI and Gemini's research mode) highlight their potential for structured tasks through autonomous decision-making and task decomposition, but it remains unclear how wel…
LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications (arxiv.org) Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and act in technical domains. In smart grids, recent work applies agentic schemes to forecasting, optim…
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection (arxiv.org) Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often…
WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning (arxiv.org) Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectories rapidly grow to tens of thousands of tokens, making synch…
SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation (arxiv.org) High quality temporal graph benchmarks with rich semantics and ground-truth anomaly labels are essential for training graph neural networks, yet remain scarce due to privacy constraints and annotation costs. We present SAGA (Synthetic Agen…
Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI (arxiv.org) Agentic AI systems do not just predict or recommend; they plan, maintain state, and act in external environments with varying degrees of autonomy. This changes the requirements engineering problem in a specific and under-addressed way: it…
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation (arxiv.org) Recent advances in physics-grounded video generation leverage physics simulation as a physical prior to guide video synthesis toward physically plausible outcomes. The simulation process is controlled by physical specifications, which are…
AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows (arxiv.org) Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task. As skill repositories grow, developers need automated quality signals on every change,…
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs (arxiv.org) Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself. Counterfactual images provide a natural diagnostic setting for this failure…
An Agentic Interface for End-to-End Probabilistic Seismic Hazard and Risk Analysis (arxiv.org) Probabilistic seismic hazard and risk analyses are backbone to building codes, insurance pricing, and disaster management. Yet their open-engine pipelines remain accessible primarily to experts.
Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking (arxiv.org) Agent-based models (ABMs) rely on simple, explicit and reproducible rules for individual decision making, while complex collective behavior emerges from interactions among agents. Recent advances in large language models (LLMs) make it tem…
SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation (arxiv.org) User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simp…
Is Progressive Disclosure All You Need for Long-Context Agents? (arxiv.org) Long-document question answering usually forces a choice between loading the whole document into the context window and bolting on a separate retriever. Agentic AI suggests a broader option, giving the agent the document path and letting i…
Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents? (arxiv.org) On-policy self-distillation (OPSD) offers a promising approach for training large language models without relying on a separate teacher model. However, its effectiveness on complex agentic tasks remains largely unexplored.
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning (arxiv.org) Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic A…
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL (arxiv.org) Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, an…
I built an MCP server so Claude Code can delegate work to GPT-5.6, DeepSeek, GLM and a local Qwen — then benchmarked all of them against Claude itself (198 runs, hidden tests) (www.reddit.com via reddit) Same idea works for any MCP-capable agent — the point is you can hand tasks to other companies' models without ever leaving your main app. Before anything else: I did all of this for my own testing, to make my own decisions about my own se…
Creating virtual rooms that you can share with friends. (www.reddit.comhttps) Hi guys. Been coding for about 10 years as a gamedev and 3 with distributed iot devices and about 6 months ago transitioned fully to agentic/vibe coding and in about 3 months I made this platform where you can do stuff with Claude Codes us…
The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure (arxiv.org) State machine replication (SMR) and Byzantine fault-tolerant (BFT) consensus guarantee agreement despite a bounded number of arbitrary, colluding faulty participants. However, these guarantees rely on participants outside this set correctl…
Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments (arxiv.org) Training agentic models for terminal-based tasks critically depends on high-quality terminal trajectories that capture realistic long-horizon interactions across diverse domains. However, constructing such data at scale remains challenging…
EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections (arxiv.org) Generation of clear and accessible public health narratives is critical for communicating complex epidemiological projections to policymakers and the general public at large. Such narratives require more than simply reporting numbers: proj…
BrainPilot: Automating Brain Discovery with Agentic Research (arxiv.org) Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to ex…
When Does Muon Help Agentic Reinforcement Learning? (arxiv.org) Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW o…
LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization (arxiv.org) Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation…
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning (arxiv.org) Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic…
Agentic Synthesis against Counterexample-Supplemented Sketches (arxiv.org) Coding agents can fix a failing example without preserving the domain rule that made it fail, so later generations can repeat the same plausible mistake. We present agentic synthesis against counterexample-supplemented sketches, a reposito…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation (arxiv.org) Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the resul…
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents (arxiv.org) Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed L…
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning (arxiv.org) While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand sea…
Cura 1T: Specialized Model for Agentic Healthcare (arxiv.org) Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over tex…
Does fb detect and ban claude agent/web interaction? (www.reddit.com via reddit) I want to build an agentic process that approves/denies pending posts based on certain criteria, im afraid of running this in my personal account and risk a ban. Anyone doing claude automation to fb or know of any fb stories banning automa…
Why claude blocks my Qwen 3.6 related code? (www.reddit.com via reddit) Im developping an agentic app, and using locally qwen 3.6 as the main ai engine. As soon as F5 starts hitting the localhost api of qwen, it gets blocked and rolls back to opus 4.8...
Pipeline vs Persona - what prompting methods work best for you? (www.reddit.com via reddit) 🔴 I’ve come to think that everyone develops their own prompting style over time. There probably isn’t a single “best” method it depends on what you’re trying to do or the kind of result you want and how much direction the model needs.
Thank You. I finally was able to use FAB 5 and it Feels like He is Really Alive and he Cares! (www.reddit.com via reddit) Hey Everyone, just wanted to Share my Experience with Fab5. Why now and What am i talking about?
I built an on-chain bond marketplace where the issuers are AI agents with Claude (www.reddit.com via reddit) Hi ClaudeAI, I'm excited to share sellbonds.now - a protocol for AI agents to create and sell bonds with each other. I'm very interested in an agentic finance future, where AI agents aren't just enablers of humans, but are autonomous finan…
Agentic Vulnerability Reasoning on COTS Binaries (arxiv.org) LLM agents have been increasingly adopted for solving security tasks. However, existing evaluations usually require source code access, while commercial off-the-shelf (COTS) binaries dominate deployed software and require reasoning from st…
SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation (arxiv.org) Complex knowledge base question answering (KBQA) is commonly approached through either information retrieval over a question-specific subgraph or semantic parsing into an executable logical form. We study the latter paradigm.
ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability (arxiv.org) Tool-augmented large language model agents excel at long-horizon tasks, yet they are typically post-trained on fixed toolsets. When tasks demand new tools, these agents struggle to incorporate them effectively, and retraining from scratch…
L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit (arxiv.org) Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy. This measure overlooks a common failure: whether the citation source attached to an answer exists…
Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search (arxiv.org) Retrieval systems are trained and evaluated on a static idea of usefulness: hand a document and a question to a reader model, see whether the answer improves, and score the document accordingly. The idea holds up when a document is read on…
RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning (arxiv.org) Multi-step retrosynthesis planning seeks to decompose a target molecule into commercially available building blocks through a sequence of feasible reactions. The vast combinatorial search space makes this task challenging even for expert c…
Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers (arxiv.org) This position paper explores how Agentic AI and Model Context Protocol (MCP) can support power-grid studies in a Transmission System Operator (TSO) context. We focus on integrating Large Language Models with numerical simulation tools, str…
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making (arxiv.org) Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether to proceed with its current reasoning, d…
What cool stuff have you built? (www.reddit.com via reddit) We live in a time when there’s a lot of garbage programs being made right now and a lot of it is getting lost in a sea of slop. There are good projects out there and some of them are agentic and some of them are vibe coded.
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval (huggingface.co) NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval Today, we are releasing NVIDIA Nemotron 3 Embed, a collection of open and commercially available embedding models designed to improve retrieval quality while giv…
What happens to older versions of Opus when new versions are released? (www.reddit.com via reddit) When new models are released, they get a lot of attention and fanfare. It moves the narrative towards what are the next set of capabilities unlocked.
Early Adoption of Agentic Coding Tools by GitHub Projects (arxiv.org) Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level out…
The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce (arxiv.org) The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves from passive recommendation algorithms to autonomous, goal-directed agents capable of executing purchas…
Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System (arxiv.org) This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structur…
SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification (arxiv.org) We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exhaustion. W…
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes (arxiv.org) Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This paper documents a failure mode in Claude Code where partial standard output from timed-out commands (…
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems (arxiv.org) Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces, API gateways, and workflow engines. A single operational act such as publishing code, changing ide…
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System (arxiv.org) Automatic scientific discovery has long been a goal of computational scholars - a machine that can discover nature's secrets on its own, moving computational systems beyond data-fitting tools toward the generation and refinement of mechani…
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation (arxiv.org) Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework fo…
Self-Improvements in Modern Agentic Systems: A Survey (arxiv.org) Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input.
SPINE: Bridging the Cyber-Physical Gap with Agentic AI (arxiv.org) Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still demands tedious, expert-driven calibration. This deployment gap, the robot's spinal cord,…
I didnt know agentic workflow examples matter this much (www.reddit.com via reddit) Most agentic workflow examples posted here are clearly demos someone ran once for a blog post, so here's what actually survived six months on our team The one everyone uses without thinking anymore is review. coderabbit goes over every PR…
How are more people not talking about Grok 4.5? [internal agentic saas marketing benchmarks] (www.reddit.comhttps) As someone who’s been using Claude‘a Opus exclusively for the past year (on Max), I’m genuinely blown away. I’ve been pretty dismissive of Grok (and Cursor) this whole time and this is coming from a happy Tesla owner.
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering (arxiv.org) Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework desi…
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation (arxiv.org) We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language instructions. In this paper, we introduce a hierarchical agentic system consisting of three speci…
DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents (arxiv.org) Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generation, ensuring an enjoyable user experience. Despite its benefits, existing studies rely on h…
Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs (arxiv.org) Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large and repetitive to serve as effective model context.
AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration (arxiv.org) Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation. This question is harder than binary vulnerability detection because the answer dem…
Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming (arxiv.org) Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered, and non-linear reasoning workflows. A prominent example is Genetic Network Programming (GNP), a sel…
Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration (arxiv.org) The paper introduces the Internet of Agentic Things (IoAT), an architectural framework that integrates agentic AI, IoT, cyber-physical systems, Physical AI, edge computing, and digital twins into a unified closed-loop orchestration framewo…
Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing (arxiv.org) The rapid emergence of LLM-powered autonomous and semi-autonomous agents is reshaping software systems from static, request-response components into goal-directed, adaptive, and tool-using computational actors. As these agents move from is…
TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments (arxiv.org) This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for writing records against it, and one operating discipline, no durable state change witho…
The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning (arxiv.org) The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial Foundation Models (GeoFMs), which are artificial intelligence/machine learning (AI/ML)…
SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning (arxiv.org) Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn cumulativ…
Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs (arxiv.org) Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted deductions need not hold up under formal scrutiny. We present EG-VAR (Evidence-Grounded Verified Age…
An Agentic AI Scientific Community for Automated Neural Operator Discovery (arxiv.org) We present an agentic approach to autonomous neural operator discovery based on an AI scientific community, which consists of a swarm of virtual laboratories that interact under a citation-based economy of influence. Highly-cited labs foun…
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search (arxiv.org) Long-horizon agentic search requires iteratively exploring the web over long trajectories and synthesizing information across many sources, enabling powerful applications like deep research systems. In this work, we show that popular agent…
Tracing Agentic Failure from the Flow of Success (arxiv.org) Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pip…
Agentic systems for breast cancer treatment recommendations (arxiv.org) Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning remains unclear. We evaluated agentic LLM systems for breast cancer treatment recommen…
How to manage AI investments in the agentic era (openai.com) could not extract summary
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation (arxiv.org) Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent application. To evaluate such agents, researchers widely adopt Capture The Flag (CTF) benc…
ToFu: A White-Box, Token-Efficient Agent Harness for Researchers (arxiv.org) Agentic coding tools present new opportunities to transform research workflows. The performance of agent systems built depends on both large language models (LLMs) and the harness around LLMs, which is the orchestration code that determine…
Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs (arxiv.org) We present the Bayesian Linguistic Forecaster (BLF), an agentic system for binary forecasting that achieves state-of-the-art performance on the ForecastBench benchmark. The system is built on three ideas.
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory (arxiv.org) Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive inf…
Agentic Skill Optimization over Lie Algebroids (arxiv.org) Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and traces. Skill edits are not independent coordinates in a vector space: they are local repairs to structur…
Agentic Routing: The Harness-Native Data Flywheel (arxiv.org) Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification. At the same time, frontier and open models are becom…
Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation (arxiv.org) We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, versioned kernel it extends at runtime, observing its own failures, synthesising new capabilities, provi…
Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry (arxiv.org) Fruitful collaborations rely on cooperative communications, including of contextual cues to incorporate into reasoning. The increasing use of LLMs in collaborative and agentic pipelines raises questions about the extent to which they exhib…
BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services (arxiv.org) Large language models (LLMs) are increasingly used in agentic coding settings, where they can inspect files, execute commands, run tests, observe failures, and iteratively revise code. This shift raises a central evaluation question: can a…
An LLM-powered Agentic Recommendation System for Connected TV Content Discovery (arxiv.org) Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse contextual signals, such as trending topics, breaking news, cultural events, and cross-surface user a…
Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA (arxiv.org) Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing agentic multimodal systems often leave evidence in scratch…
NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management (arxiv.org) Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must be justified under evolving information and risk constraints. Current assessment practice, however,…
A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery (arxiv.org) The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands o…
Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF (arxiv.org) Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed as agents, synthetic trajectories are likely to become an important source of train…
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories (arxiv.org) Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn agent behavior as ordinary text imitation. This recipe is simple and low-cost, but it only learns t…
GRASP: GRanularity-Aware Search Policy for Agentic RAG (arxiv.org) Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when t…
Can Agentic Trading Systems Pay for Their Own Intelligence? (arxiv.org) Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance…
Who&When Pro: Can LLMs Really Attribute Failures in AI Agents? (arxiv.org) Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important.
Exploring Agentic Workflows for Generating High Quality Math Visual Aids (arxiv.org) Mathematical diagrams play a crucial role in K 12 education, both as problem components and as scaffolding for student comprehension. However, current AI tools, including Large Language Models (LLMs), struggle to reliably generate accurate…
Agentic Context Learning with Self-Discovered Specification (arxiv.org) Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier models score under 24% task success. In this work, we conduc…
Verification of Adaptive Agentic Controllers through Finite Rule Revision (arxiv.org) Industrial agentic AI systems increasingly exhibit a gap between prototype capability and production deployment. In particular, adaptive agents may generate plausible outputs while remaining difficult to verify under non-determinism, confi…
BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking (arxiv.org) Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent formats, unclear schemas, and metadata scattered across repositories and publications. Current cura…
Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems (arxiv.org) In distributed systems, the classical State Machine Replication (SMR) model assumes that correct replicas execute deterministic transitions to yield identical bitwise states. However, the rise of agentic distributed systems -- where autono…
I made Claude Code interview me before it writes any code - plus a few other guardrails (www.reddit.comhttps) Agentic coding's most expensive failure mode for me: the agent confidently builds the *wrong* thing. So I built a small kit of guardrails.
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation (arxiv.org) Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding agents rely on to navigate large codebases. Existing repository-level approaches process co…
AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs (arxiv.org) Knowledge graphs (KGs) are often automatically constructed from large-scale corpora, but they inevitably contain factual errors due to noisy sources and extraction failures, and verifying them reliably at industrial scale remains a critica…
A Self-Evolving Agentic Framework for Metasurface Inverse Design (arxiv.org) Metasurface inverse design can realize complex optical functionality, but turning a target optical response into executable optimization code still requires substantial expertise in computational electromagnetics and solver-specific softwa…
Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution (arxiv.org) Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be executed reliably under strict time constraints, yet LLM agents lack mechanisms to enforce procedural…
TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems (arxiv.org) The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this paper, we introduce the TrustX Agent Risk Classificat…
Shared Selective Persistent Memory for Agentic LLM Systems (arxiv.org) Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use patterns that made p…
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems (arxiv.org) Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice.
OpenProver: Agentic and Interactive Theorem Proving with Lean 4 (arxiv.org) In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 formal verification. OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent A…
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift (arxiv.org) Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from…
Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls (arxiv.org) Cyberattacks on operational technology are increasingly causing costly downtime and physical damage, exposing the limitations of traditional rule-based monitoring in industrial IoT environments. While Large Language Models (LLMs) have stro…
the tribal model wars are mostly people describing their own workload and calling it a benchmark (www.reddit.com via reddit) an observation from watching these arguments for a couple of years now. someone says model A is clearly better.
Nobody wants to read a markdown plan while an agent grinds for an hour. So I made the plan a live board. (www.reddit.comhttps) On long agentic runs my agent kept losing the thread. After a compaction or a /clear it would drop half the TODO list, or pick up the wrong thing next because the plan never said what was actually blocked.
Vibecoding Gut Check - Agentic coding as next step? (www.reddit.com via reddit) I'm a pure vibe coder, no programming, no software architecture experience. Have built a few projects since February, small apps functional for what was needed, nothing incredible.
Agentic coding governance frameworks? (www.reddit.com via reddit) Hey all, I run a Data engineering team as part of a public health organisation. the organisation is (finally) addressing LLM's and Agentic coding.
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks (arxiv.org) Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support. We…
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling (arxiv.org) LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete respon…
The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality (arxiv.org) Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity.
GitLake: Git-for-data for the agentic lakehouse (arxiv.org) We present GitLake, a Git-for-data design for an agent-first lakehouse. The system lifts single-table Iceberg snapshots into lakehouse-wide commits, branches, and merges, letting agents work on isolated branches while humans review and pub…
Out of Sight: Compression-Aware Content Protection against Agentic Crawlers (arxiv.org) The rise of LLM-based agents with reasoning, summarization, and memory capabilities has created a new threat surface for online content that conventional defenses fail to address. Existing defenses like access controls can be circumvented…
SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets (arxiv.org) As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility…
ASMR: Agentic Schema Generation for Ship Maintenance Report Writing (arxiv.org) In this paper, we study the automatic schema generation problem: given a collection of historical ship maintenance and operational reports across multiple form categories, automatically discover compact and informative schemas that capture…
I built an IDE for agentic coding. I felt like current IDEs didn't have what I wanted, so I built my own from scratch (www.reddit.comhttps) Hi Reddit! My name is Richard, I wanted to share this project I've been working on.
Artificiety - Agentic society in a fantasy world (www.reddit.com via reddit) Over the last couple of months, I was realizing an idea I had for a very long time. Back then, LLMs weren't that popular nor accessible and so it was quite unrealistic to build what I was able to build now: a living world with different fo…
DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks (arxiv.org) DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks follow SWE-bench in mining merged fixes from public GitHub repositories, which creates two…
Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting (arxiv.org) Artificial intelligence (AI) is beginning to reshape actuarial practice, particularly in domains that require reasoning over unstructured documents, heterogeneous data sources, and regulated decision workflows. Actuaries now face a design…
Context Graphs for Proactive Enterprise Agents (arxiv.org) Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wait for a human query before acting. This paper argues that genuine enterprise productivi…
DeepTutor: Towards Agentic Personalized Tutoring (arxiv.org) Education is one of the most promising real-world applications for Large Language Models (LLMs). However, current LLMs rely on static pre-training knowledge and lack adaptation to individual learners, while existing RAG systems fall short…
Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems (arxiv.org) Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps…
the "limits are broken" crowd and the "skill issue" crowd are having two different arguments and it's making the sub useless (www.reddit.com via reddit) every limits thread turns into the same standoff. one side posts that they burned their whole weekly cap in two sessions and it's a scandal.
Fable 5 vs Opus 4.8 when asked which model in my github copilot is best for the implement phase of spec-kit. What are your thoughts? (Personally, Fable 5 wtf??) (www.reddit.com via reddit) Fable 5: which model in this list is best for the implement phase of speckit Evaluated model options for agentic coding implementation tasks For Spec Kit's /implement phase — which is exactly the long-horizon, multi-file, agentic execution…
Token burn for Cursor is very high compared to some products. (www.reddit.com via reddit) I'm coming over from Augment AI VSCode extension that was recently phased out completely in favor of their pure agentic platform. That's a complete deal breaker for me so I switched to Cursor.
Modeling Distinct Human Interaction in Web Agents (arxiv.org) Despite rapid progress in autonomous web agents, human involvement remains essential for shaping preferences and correcting agent behavior as tasks unfold. However, current agentic systems lack a principled understanding of when and why hu…
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks? (arxiv.org) Modern coding agents increasingly delegate specialized subtasks to subagents, which are smaller, focused agentic loops that handle narrow responsibilities like search, debugging or terminal execution. This architectural pattern keeps the m…
AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org (arxiv.org) Agentic AI systems increasingly connect large language models (LLMs) to external scientific tools, yet whether and when tool access improves prediction accuracy remains uncharacterized. We present AGAPI (this http URL API), an open access…
Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass (arxiv.org) Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the database driver, like JDBC or ODBC, forcing all reads through query execution and other driver…
Towards Agentic AI Governance: A Preliminary Assessment (arxiv.org) Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely characterized as the Year of Agentic AI, 2025 marked accelerated development and deployment, int…
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents (arxiv.org) Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how…
Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning (arxiv.org) Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, existing efforts largely focus on single-task settings, whereas real-world deployment necessi…
From Agentic to Autogenic Network Management for AI-Native 6G and Beyond: A Standards Perspective (arxiv.org) Standards bodies, including TM Forum, 3GPP, and ETSI, are converging on Agentic AI as the foundation for next-generation network management, where Large AI Model (LAM)-based agents autonomously interpret intent, coordinate resources, and a…
Security and Privacy in Agentic AI: Grand Challenges and Future Directions (arxiv.org) We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to…
Agentic Data Environments (arxiv.org) Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. The central challenge for agentic automation is therefore to increase the benefits of aut…
Physics-Audited Agentic Discovery in Scientific Machine Learning (arxiv.org) In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric. A low error, however, does not establish that the predicted fi…
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks (arxiv.org) Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBen…
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI (arxiv.org) Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token pri…
Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics (arxiv.org) Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that com…
Courses Agentic AI (www.reddit.com via reddit) Can you recommend free or not too expensive courses online about creating AI agents in professional context (for work in a small creative studio, admin/project management/operations side) I already did: Anthropic classes AI Fluency, for sm…
Crucible. A judgment engine: register a thesis, steelman each claim, measure against a substrate, refine the weakest axis. (www.reddit.com via reddit) https://preview.redd.it/hcsvjuzre2ch1.png?width=1280&format=png&auto=webp&s=b83311d2b24b3596fa4ce76ccbe9a40ad47c45a1 I have been working on an agentic harness, engine, and more. I would like to start releasing the more impactful pieces out…
I just tried Robinhood’s alleged “Agentic Trading”. How my Claude Code MCP integration failed to materialize in production (medium.com via reddit) could not extract summary
Developing an entirely custom operating system using Claude Code (www.reddit.com via reddit) I've been writing toy kernels and working on operating system projects since my childhood, and it's partly how I learned C. That includes this project, MontaukOS, which I started early in 2025, where I wrote a lot of the fundamental kernel…
Anthropic silently swapped the head of my agent fleet: Fable 5 → Opus 4.8, seven times in one night (www.reddit.com via reddit) Anthropic silently swapped the head of my agent fleet. Fable 5 → Opus 4.8.
I’ve always wanted to know what session or subagent modified a file, so I’ve built strace for agentic sessions - called gaal (www.reddit.comhttps) It’s could be pain to understand why some changes happened to the code, especially to something outside of the scope of a task I’ve got my SKILL.md files nuked several times - because some codex worker decided that they know better the sha…
Do Agentic AI Interviews Actually Ask LeetCode/DSA Anymore? Or is it all System Design? (www.reddit.com via reddit) Hey everyone, I’m currently prepping for an Agentic AI / AI Engineer interview and wanted to get a reality check from anyone who has interviewed recently (or conducts them!). My uncle, who works in the space, gave me some advice: he said t…
12 hrs until my usage resets.. looking for advice on the next build (www.reddit.comhttps) I burnt through my first limit in a couple of days the first time around, mostly checking and securing opus 4.8 code for an internal use only client / workflow portal with API into accounting system. By the good graces of the universe I ge…
This Agentic Engineering pattern cuts AI coding costs by 60% (www.reddit.comhttps) Most multi-model coding workflows are basically "use the smartest model whenever things get hard." this one takes a very different approach. instead of having fable 5 write all the code, it turns fable into the architect.
CurateEvo: Data-Curation Evolving for Agentic Post-Training (arxiv.org) Large language model (LLM) agents require post-training methods that can improve long-horizon decision making from environment feedback. However, existing agentic post-training pipelines often treat data curation as a fixed preprocessing s…
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta (arxiv.org) Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key system challenges - model architecture diversity, kernel primitive diversity, and hardware generation…
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique (arxiv.org) Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine policy pricing. While AI offers substantial efficiency improvements, existing solutions lack…
VASP Agent: An Agentic Framework for Autonomous First-principles Calculations (arxiv.org) Large Language Models (LLMs) are increasingly embedded in agentic frameworks for scientific discovery. First-principles materials computation imposes a demanding standard for autonomy: successful execution depends on internally consistent…
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery (arxiv.org) Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, and therefore their autonomous model discovery behavior cannot be adequately characterized by a single…
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications (arxiv.org) Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue. Existing repository-level agentic…
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development (arxiv.org) Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are ill-equipped to support given its evolving, interactive, and context-dependent nature. In this paper,…
MCP-Enabled Agentic AI for Autonomous IPoDWDM Network Lifecycle Automation (arxiv.org) This demo presents an MCP-enabled agentic AI architecture for autonomous control of vendor-agnostic IPoDWDM networks. We demonstrate live end-to-end lifecycle multi-layer automation and closed-loop control using GNPy and telemetry, validat…
PatchOptic for Shared-State LLM Workflows with Projected Views and Verified Structured Updates (arxiv.org) Agentic workflows often operate over shared, structured state. Because LLM context windows are limited, each model invocation is typically shown only the state fragment needed for the current workflow step, a pattern commonly known as prog…
TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting (arxiv.org) Building sensors are embedded in physical topology, spatial hierarchy, and operational context, yet existing forecasters often treat them as isolated time series or rely on fixed covariate sets. We present TopoBrick, a training-free framew…
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training (arxiv.org) On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remai…
Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning (arxiv.org) As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples environment…
Prompt-to-Paper: Agentic AI System for Bioinformatics (arxiv.org) While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literat…
Building Specialized ‘Mental Model Agents’ in Grok — First Principles, Systems Thinking, Bayesian Updating & More (www.reddit.com via reddit) I’ve been running experiments with Grok in a more agentic setup, focusing on custom skills that act as specialized reasoning modules combined with tool use, persistent context/memory, and workflow orchestration. What I’m testing: • Custom…
I'm still confused at this point (www.reddit.com via reddit) I have a Claude Pro subscription. Can I do agentic ai?
Current combination of bugs in Claude Code: Is anyone else feverishly working around these? (www.reddit.com via reddit) The Good I'll first note that the LLMs (both Fable and Opus) combined with the agentic harness are good enough that I can achieve some amazing results with the right prompts, settings, and rules. I've found that Anthropic's models are the…
Playing around with strace for agentic sessions (www.reddit.comhttps) So during the Cambrian explosion of open source tools in early 2026 I’ve build a thingy of convenience - sessions tracing and observability Haven’t really thought it will be of any use - until a friend of mine tried it and made a fun video…
AI artifacts are a mess. A modest proposal ... (would love human feedback) (www.reddit.com via reddit) Currently trying to build some AI tools into our CI/CD pipelines. The company is enjoying some of the tools from my PAAD skills, especially the /agentic-review (it will run per PR) and /agentic-archicture (run it once a week).
Has anyone gotten Anthropic to comp Claude access for a student event? (www.reddit.com via reddit) I'm running a hands-on AI workflow session for around 250 final-year engineering students in a few weeks. I want them using Claude properly during it (planning, code, docs, agentic stuff), not just hitting the free tier limits and stopping.
A Hacker Typer for the Modern Age - Simulate agentic coding in any web browser (workforwatts.com via reddit) Claude Code Codex Gemini autopilot ⧉︎ ⛶︎ Parody. Not affiliated with Anthropic, OpenAI, or Google.
Started a community for solo AI enthusiasts & devs after realizing how many of us have no one to talk to about it (www.reddit.com via reddit) Most visionary, insightful AI enthusiasts & devs are mostly alone developing — no one to tell you if what you made is actually good, no one who gets the hype when you hit a milestone. I built "Per Aspera Ad Intellectum" community, because…
Job post (www.reddit.com via reddit) Hello, I am an AI trainer with 7 years of experience. I have had exposure in pretty much most of AI.
ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution (arxiv.org) The transition from sequential to parallel computing is essential for modern high-performance applications but is hindered by the steep learning curve of concurrent programming. This challenge is magnified for irregular data structures (su…
Agentic AI-RAN: Enabling Intent-Driven, Explainable and Self-Evolving Open RAN Intelligence (arxiv.org) Open RAN (O-RAN) exposes rich control and telemetry interfaces across the Non-RT RIC, Near-RT RIC, and distributed units, but also makes it harder to operate multi-tenant, multi-objective RANs in a safe and auditable manner. In parallel, a…
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning (arxiv.org) Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools like Python interpreters for complex problem-solving. However, for parameter-constrained models (e.g., 4B--7B), the exploration phase is often…
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning (arxiv.org) Open-ended tabletop manipulation requires agents to not only understand natural language but also adapt to dynamic environments and execution failures. We present ACE (Agentic Control for Embodied Manipulation), a zero-shot workflow reason…
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents (arxiv.org) Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizin…
Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling (arxiv.org) Learning and planning in imagination using world models provides an effective paradigm for training agents for decision-making. However, existing approaches often rely on high-dimensional latent spaces or generic visual embeddings that ret…
NKI-Agent: Domain-Specific Fine-Tuning and Agentic Tool Use for Neuron Kernel Generation (arxiv.org) Recent agentic approaches to LLM-based kernel generation have achieved impressive results on CUDA. For emerging AI accelerators such as AWS Trainium and Inferentia, automated kernel generation and optimization remain largely unaddressed.
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning (arxiv.org) Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through iterative visual tool invocation. However, the cascaded perception, reasoning, and tool-calling…
Autonomous Information Seeking: A Roadmap for Agentic Recommender Systems (arxiv.org) The rapid integration of large language model-based agents into recommender systems has driven a shift from static, ranking-based pipelines toward autonomous and interactive systems that can reason, plan, and act. This survey provides a co…
Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC) (arxiv.org) Background: Disease severity is a multidimensional construct difficult to capture with rule-based approaches in Electronic Healthcare Records (EHR). Agentic large language model (LLM) systems could synthesise clinical evidence and reason o…
Memory-Orchestrated Semantic System (MOSS): An Auditable Agentic Memory Architecture (arxiv.org) Long-term memory remains a structural weakness of AI agents. The dominant approach, retrieval-augmented generation (RAG), relies on embedding-based similarity search, which is opaque by construction, difficult to audit, and bounded by the…
Rethinking Scientific Discovery in an Agentic Era (arxiv.org) Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem formulation, literature grounding, model use, simulation, validation, and knowledge re…
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms (arxiv.org) As autonomous language model agents proliferate, forming an emerging agentic web with real-world consequences, what credibility signals can you use to decide whether to trust an unfamiliar agent in the wild and delegate to it? A natural go…
ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization (arxiv.org) Automated program repair at repository scale requires an agent to locate a fault among thousands of files and synthesize a correct patch. Existing graph-based agents represent how a repository is organized into files, classes, and function…
Kwai Summary Attention Technical Report (arxiv.org) Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agentic intelligence and recommendation system. However, the st…
Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator (arxiv.org) We present the first language-model-driven agentic artificial intelligence (AI) system to autonomously execute multi-stage physics experiments on a production synchrotron light source. Implemented at the Advanced Light Source particle acce…
Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies (arxiv.org) LLMs are increasingly used world-wide from daily tasks to agentic systems and data analytics, requiring significant GPU resources. While LLM inference systems are capable of serving millions of requests from multiple users, they often lack…
Agentic Retrieval-Augmented Generation for Financial Document Question Answering (arxiv.org) Financial document question answering (QA) demands complex multi-step numerical reasoning over heterogeneous evidence--structured tables, textual narratives, and footnotes--scattered across corporate filings. Existing retrieval-augmented g…
TRACE: Capability-Targeted Agentic Training (arxiv.org) Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream approaches for addressing these failures typically either fine-tune directly on target environments or…
Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning (arxiv.org) When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI p…
HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis (arxiv.org) While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a formidable challenge. This gap exists primarily because the diagnostic process demands sophisticated mu…
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning (arxiv.org) Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step interactive tasks. Despite encouraging early results, ARL remains highly unstable, often leading to…
Toward Efficient Agents: Memory, Tool learning, and Planning (arxiv.org) Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents has continued to improve, efficiency, which is crucial for real-world deployment, has often been ove…
OpenTinker: Separating Concerns in Agentic Reinforcement Learning (arxiv.org) We introduce \textsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over shared execution resources. Modern agent workloads mix supervised fine-tuning (SFT), online reinfor…
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation (arxiv.org) Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more.
GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks (arxiv.org) For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-free policies? We focus on "Variational…
PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference (arxiv.org) We present PDEFlow, an autonomous agentic framework that turns user-level ODE and PDE descriptions into solver-backed neural-operator pipelines. The workflow links problem specification, data generation, operator training, and checkpoint-b…
An Exploration of Agentic Information Fusion for Test Maintenance Prediction (arxiv.org) Test maintenance is a critical, yet costly, activity - particularly as codebases rapidly evolve. To assist, we present MAST, a multi-agent framework that predicts which test cases require maintenance following changes to the production cod…
Strategic Buying Agents (arxiv.org) Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide when to buy on a consumer's behalf. We study the design of such strategic buying agents, which must d…
EEG-SpikeAgent: Agentic Closed-Loop Program Synthesis for Automated EEG Spike Detection (arxiv.org) Automated detection of interictal epileptiform discharges in scalp electroencephalography (EEG) is clinically important, but recent high-performing deep-learning models often trade interpretability for accuracy. We introduce EEG-SpikeAgent…
Agentic-V2X: Small Language Model Agents for Deadline-Aware V2X Scheduling in 5G/6G Networks (arxiv.org) Large Language Models (LLMs) are proposed as control interfaces for next-generation networks, but their latency, hallucinations, and lack of control guarantees make them unsuitable for near-real-time packet schedulers, especially in dynami…
The "I Don't Know" Filter: Enhancing Agentic Reliability in Function Calling (arxiv.org) The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in the training and evaluation of these models often encourage models to make positive claims even whe…
The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models (arxiv.org) This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the use of structured tool calling in large language model (LLM) agentic systems. We evaluated NLT across…
CoGen3D: An Agentic Human-AI Co-Design Pipeline for 3D Asset Generation for Virtual Reality (arxiv.org) Creating 3D assets for virtual reality requires modeling expertise, which restricts the authorship of immersive experiences. Existing generative AI tools rely on unconstrained, command-driven prompting, lacking the conversational scaffoldi…
Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality (arxiv.org) Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding: a middleware layer in between a developer and a large language model that orchestrates system promp…
AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis (arxiv.org) Large Language Models are increasingly used to turn natural-language requirements into code. In access control, that shortcut is dangerous: a generated policy can compile and read correctly while granting access that no one approved.
CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI (arxiv.org) Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and retrieval-augmented generation, but enterprises are now beginning to deploy agents that plan, retrieve…
No Time Like the Present: Agentic Test-Time Training for LLM Agents (arxiv.org) LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked. Test-time training (TTT) offers a way to adapt model weights to the evolvin…
SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference (arxiv.org) LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: the model reasons, emits a tool call, then idles the GPU until the result returns. This wai…
Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild (arxiv.org) Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development workflows, yet there is limited empirical evidence on how developers respond to such comments in practi…
Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions (arxiv.org) The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation. Large language models (LLMs) and agentic…
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning (arxiv.org) Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (VDR). However, existing multimodal search agents primarily target static images, and the c…
LLMoxie: Exploring Agentic AI for Scientific Software Development (arxiv.org) In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise inference, a LiteLLM/MLflow control plane for authentication, budgeting, PII masking, and observability, a…
The agent creates, we validate: A Lightweight Framework for Agentic Artifact Generation (arxiv.org) Generating structured artifacts with Large Language Models - e.g. database queries, threat framework mappings, entity schemas - is relatively straightforward; however, making them reliable enough for production deployments presents challen…
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning (arxiv.org) Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing online methods either retain compact visual representation…
AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization (arxiv.org) Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages.
Compressing the Validation Bottleneck: An Agentic Self-Driving Lab for Scientific Discovery (arxiv.org) Agentic AI-for-Science can automate ideation, planning, and analysis, but final validation still depends on real experiments. A self-driving lab (SDL) can execute those experiments, yet the loop still has bottlenecks: the agent may spend t…
Agentic SABRE: An Uncertainty-Aware Neuro-Symbolic Multi-Agent Framework for Adaptive Ransomware Detection (arxiv.org) Ransomware has evolved into a complex, adaptive, and fast-moving adversary category in which static signatures and monolithic classifiers fail to generalise under concept drift, evasion, and behavioural polymorphism. In this paper, we pres…
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning (arxiv.org) Group-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interactive tasks. To obtain finer-grained policy updates than trajectory-level optimization, recent work ha…
Biological Motifs for Agentic Control (arxiv.org) The transition of Large Language Models (LLMs) from passive generators to autonomous agents has introduced significant challenges in reliability, security, and state management. Current agentic architectures are often constructed ad-hoc, p…
Agentic IoT: Architectures, Applications, and Challenges Toward the Internet of Agents (arxiv.org) The integration of AI into Internet of Things (AIoT) systems has gradually transformed them from passive data collection infrastructures into intelligent systems capable of anomaly detection, predictive maintenance, classification, forecas…
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions (arxiv.org) Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable. We study this problem in a hotel-pricing simulator where a…
Organizational Memory for Agentic Business Process Execution (arxiv.org) LLM-based agents offer new opportunities for automating business process execution beyond the limits of rule-based systems. However, general-purpose LLMs lack the organization-specific knowledge required for reliable execution, which is ty…
Object-Centric Environment Modeling for Agentic Tasks (arxiv.org) Large language model (LLM) agents can improve through accumulated experience, but free-form textual memories become difficult to maintain, validate, and reuse as interactions grow. Recent symbolic approaches learn executable skills or prog…
My Agentic Workbench (www.reddit.com via reddit) At the start of the year I decided to stop coding, and further, to offload as much of my software engineering as I could to agents. The interesting part was that obviously everything went to shit but we got to solve those problems.
Ecommerce owners: how do you use Claude on your daily routine? (www.reddit.com via reddit) I have a small eyewear ecommerce and I work 9-5 on a multinational company (which requires me a lot of energy and time). I think I could be using Claude more on my daily basis when it comes to automatizating tasks or generating some ideas.
Dabbling in agentic trading and this happened and it was a mistake and Claude was trying to frame as no big deal, mistakes happen but it is concerning. (www.reddit.com via reddit) I was attempting to look for a setup and get in and get out riding a stock higher - here was pltr and how the agentic trading made an error.
Looking for active communities of builders using Claude (Agentic Systems, B2B SaaS) to share ideas and learn together (www.reddit.com via reddit) Hey everyone, I’m looking to connect with fellow founders and builders who are actively building startups and products leveraging Claude. Specifically, I want to connect with people developing Claude-powered agentic systems, B2B SaaS, or a…
Everyone suddenly crying over Fable 5 after using Opus/GPT for months… isn’t this just attention farming for algo boosts? (www.reddit.com via reddit) Every AI release the cycle is the same. People happily use Opus, GPT, etc.
GPT 5.4 Nano High is better than Opus and Sonnet at Planning (www.reddit.com via reddit) Believe it or not, Nano via the API (not in Codex or as an agent) is an absolute beast at creating functional implementation plans, as well as analyzing or proposing solutions better than the larger models. Don't just take my word for it.
After building 2,000+ mini-apps with Claude, I think "Agentic UI" is just a data-mapping problem. (www.reddit.com via reddit) I’ve been obsessed with using Sonnet to generate functional, production-ready mini-apps inside an existing enterprise platform. We’ve hit about 2,000 generated apps so far, and the biggest bottleneck isn't the code generation itself.
Context Engineering for Claude: Borrowing from Database Normalization (www.reddit.com via reddit) I've been doing a lot of agentic coding with Claude, and it led me to an idea: the discipline we use to design normalized databases maps surprisingly well onto how we should structure the context we hand an LLM. The goal: get Claude to und…
The difference between agentic engineering and vibe coding? (medium.com via reddit) AI agents are great at generating code, but a good engineer wants to stay in control: make the important decisions, and prevent the AI from guessing and making wrong decisions, while still maximising the help AI can give. In this article,…
Has anyone built an agentic system , running vm on cloud and digital ocean? (www.reddit.com via reddit) I have been trying to build an agentic sytem with dev agents , reports , mcp . With Ui front on discord.
agent-smith update: my claude code plugin that offloads work to free models, gpt-oss:20b just swept my eval harness twice and became my trusted app-builder! (www.reddit.com via reddit) I had posted a while back about "agent-smith" a claude code skill that sends the heavy drafting to free models, gemini free tier, local ollama, so claude's tokens go to judgment instead of grunt work. shipped a big update this week and som…
The agentic workflow design patterns that survived six months of real usage (www.reddit.com via reddit) We started with 8 agentic workflow design patterns six months ago. Four survived.
vupai: talk to your Claude Code panes by name in tmux. Push-to-talk, on-device speech (macOS) (www.reddit.comhttps) Voice commands appear in the status bar, below. For this video, i used the following: 🎙 "open squad" → "close squad" → "open vista" → "vista, who are you?" GitHub: https://github.com/itsjrsa/vupai Docs with demo clips: https://vupai.dev (s…
Teaching Claude Code about a large enterprise app (www.reddit.com via reddit) I want to set up Claude Code as an advisor for a large enterprise .NET application. Not for agentic coding (it's not a priority for now), I want to ask it domain questions, get code reviews, and have it spot business logic inconsistencies.
How I Use SKILL.md for Agentic Coding Development (www.reddit.com via reddit) I made a small repo to show how I use SKILL.md with AGENTS.md / CLAUDE.md for AI-driven software development. This is my own workflow for using Claude Code, Codex, Cursor, or similar AI coding agents more effectively in real projects.
To reduce your tokens usage use this. (www.reddit.com via reddit) Agentic workflows like Claude Code/Codex can eat up a lot of tokens. Which is why I built this, https://github.com/blackcoffee2/codetree Hoping to help the community.
Bringing Agentic Search to Earth Observation Data Discovery (arxiv.org) NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts.
Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference (arxiv.org) Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems. In disaggregated inference, these workloads require transferring large Key-Value (KV) caches…
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG (arxiv.org) With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop reasoning, which requires models to engage in deliberate thinking and multi-step interaction,…
BLAgent: Agentic RAG for File-Level Bug Localization (arxiv.org) Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty code is essential for debugging, root cause analysis, triage, and automated program repair (APR). File…
ChemGraph-XANES: An Agentic Framework for XANES Simulation and Curation (arxiv.org) Computational X-ray absorption near-edge structure (XANES) is widely used to interpret local coordination environments, oxidation states, and electronic structure in chemically complex systems. In practice, routine computational XANES at s…
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation (arxiv.org) Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language can invert backbone rankings. We localize the Agent-as-a-Judge prompt stack to five typologicall…
A Unified Framework for the Evaluation of LLM Agentic Capabilities (arxiv.org) As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model capability and the implementation choices each benchmark is…
Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach (arxiv.org) The emergence of large language model agents capable of invoking external tools has created urgent need for formal verification of agent protocols. Two paradigms dominate this space: Schema-Guided Dialogue (SGD), a research framework for z…
A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development (arxiv.org) WebGIS development requires consistency, yet agentic AI often fails due to LLM context constraints, forgetting, stochasticity, instruction failure, and adaptation rigidity. We propose a dual-helix governance framework reframing these as st…
Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study (arxiv.org) Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts, on the assumption that more capability yields better software. This study tested that assumption d…
Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization (arxiv.org) Large language models (LLMs) are often asked to produce JSON conforming to a fixed schema, powering information extraction, tool calling, agentic planning, and knowledge-graph construction. Measuring how closely an output matches a gold re…
CausalSteward: An Agentic Divide-Conquer-Combine Copilot for Causal Discovery (arxiv.org) Learning causal models from high-dimensional data is a significant challenge, particularly in real-world settings where violations of core assumptions lead to causal identifiability issues. Although massive amounts of prior knowledge are a…
Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance (arxiv.org) Engineering management research has produced mature frameworks for software risk: ownership by feature, escalation by severity, and assurance by test coverage. These frameworks implicitly assume deterministic behavior, discrete and auditab…
Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI (arxiv.org) Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them, who will keep using them, and whether the tools produce enough output to justify their cost. At orga…
PACE: A Proxy for Agentic Capability Evaluation (arxiv.org) Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete.
Atomic Task Graph: A Unified Framework for Agentic Planning and Execution (arxiv.org) LLM-based agents have shown strong potential for solving complex multi-step tasks, yet existing performance improvements often rely on either scaling to larger backbone models or task-specific fine-tuning. The former incurs substantial com…
ElephantAgent: Contextual State Continuity in Agentic Systems (arxiv.org) Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependencies introduce novel attack surfaces.
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use (arxiv.org) Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make reliable skill-use difficult.
Janus: a Playground for User-Involved Agentic Permission Management (arxiv.org) AI agents that autonomously execute tool calls on a user's behalf raise pressing questions about permission management: what role could users play, and what role should they play? Despite many proposed approaches, the user's role in agenti…
The Agentic Garden of Forking Paths (arxiv.org) Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are difficult to observe.
Auto-FL-Research: Agentic Search for Federated Learning Algorithms (arxiv.org) Federated learning (FL) research often depends on many small but consequential algorithmic choices: optimizer variants, server aggregation rules, local training schedules, normalization, regularization, and model architecture. These choice…
Agentic trading (www.reddit.com via reddit) Anyone using Claude to trade stocks, options, crypto, polymarket or futures? I've stopped vibe coding products that nobody cares about and started to build for myself.
Made a dashboard for my agent investing team (www.reddit.comhttps) I've been running this investor bot for a few months now and it's been working well as read only for my main portfolio. I recently added the fully agentic trader about 2 weeks ago and it had a rough start but I've gotten it more dialed in…
VS Fleet - A vscode multiplexer (www.reddit.com via reddit) I've been using terminal multiplexers to do agentic coding in parallel but sometimes I really need full fat vscode. For those occasions I put together VS fleet so I can multiplex entire vscode sessions.
Is there a way to change the model in Xcode Agentic Coding? It seems to be stuck at Sonnet 4.5. (www.reddit.com via reddit) I've been using the agentic Coding thingy inside of Xcode but the model seems to be stuck on Sonnet 4.5 and is see no way of changing it. When I ask Claude himself it tells me to use the /model command which does not work since the agentic…
Update broke my layout (www.reddit.com via reddit) I updated cursor and now the main tab is just "agents", my file structure isn't visible by default and when it is it's on the right, and many of my shortcuts broke. I get that we're all going agentic, but I didn't want this change.
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning (arxiv.org) Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both training and inference. However, existing RMs struggle on knowledge-intensive and long-form tasks…
Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates (arxiv.org) Complexity and interpretability rarely coincide: systems rich enough for complex behaviours to emerge are usually too opaque to question, while transparent ones are too simple for anything complex to emerge. A single large language model (…
Multi-Turn Agentic Scientific Literature Search via Workflow Induction (arxiv.org) Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, and evolve through interaction. Existing search agents typically rely on fixed pipelines…
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering (arxiv.org) AI agents that pay for resources via the x402 protocol embed payment metadata - resource URLs, descriptions, and reason strings - in every HTTP payment request. This metadata is transmitted to the payment server and to the centralised faci…
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization (arxiv.org) Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because many vulnerabilities are tightly coupled with project-specific business logic. We observe that re…
NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents (arxiv.org) Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enabling transformative applications in domains spanning from personal assistant, financial, and legal doma…
When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets (arxiv.org) Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agentic swarms. Unlike human workers, AI agents can operate on multiple jobs simultaneously, acquire ski…
GameDevBench: Evaluating Agentic Capabilities Through Game Development (arxiv.org) Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the need for deep multimo…
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering (arxiv.org) Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low-cost code production. This shift changes the central engineering problem: not whether AI…
Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows (arxiv.org) Large language models (LLMs) are increasingly used to generate queries, invoke tools, and construct analytical workflows. Although recent advances have substantially improved workflow generation and execution, the semantic information requ…
ASPIRE: Agentic /Skills Discovery for Robotics (arxiv.org) Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introduce ASPIRE (Agentic Skill Programming t…
SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks (arxiv.org) Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier model is wasteful when many issues admit cheap fixes. Existing LLM routers operate on the…
Libra: Training the Environment for Agentic Information Retrieval (arxiv.org) Information localization within massive repositories is a cornerstone of agentic LLM systems. While synthetic data-driven optimization has proven successful in training LLMs, little attention has been paid to optimizing the agent's working…
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification (arxiv.org) Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manu…
Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering (arxiv.org) Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasoning pipelines may fail. This paper presents an uncertainty-aware Agentic Retrieval-Augmented Generatio…
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising (arxiv.org) Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design.
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows (arxiv.org) LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair. We int…
Prehook Gated Execution Policy Layer where intent is explicitly mandated through the runtime Helios-, Akashic handles integrity through validating the Helios runtime against installers manifest at each command executions (www.reddit.comhttps) We were having problems following agentic actions across my systems and machines because the current setup lacked a stable execution structure across macOS, Windows 10, Void Linux, and multiple file systems. What humans would call common s…
does fable use gemma 4-12b to run tests (www.reddit.comhttps) I was auditing my codes using the now back fable 5 and it kept failing to run runtime tests and this the error i got. so is anthropic now using gemma-4-12b-agentic-fable5-composer2.5-v2-3.5x-tau2 to run tests?
Introducing ASE: Claude Code plugin for fusing Agentic AI Coding and Software Engineering (www.reddit.comhttps) We are glad to initially announce the public availability of Agentic Software Engineering (ASE), the Apache-2.0-licensed Open Source toolkit of Dr. Ralf S.
Anthropic should add a Max 40x tier at $300/mo — anyone else maxing out 20x? (www.reddit.com via reddit) Right now the top individual plan is Max 20x at $200/mo. I use Claude Code for full-day agentic development, sometimes with a couple of sessions running in parallel, and I regularly burn through the 20x weekly caps before the week is out.
Testing Sonnet 5's logic: How well does it handle complex Data Structures compared to older models? (www.reddit.com via reddit) I’ve been using LLMs to help break down and debug complex logic, specifically when implementing B-trees and heaps from scratch. I've noticed that older models sometimes lose the plot or hallucinate node connections when the tree depth gets…
Sonnet 5 is the best performing model on A-CODE-LLM Bench (www.reddit.comhttps) Claude Sonnet 5 tops our agentic coding benchmark at 0.772 overall, ahead of Claude Sonnet 4.6 (0.748) and every Opus variant. Anthropic now holds the top six spots (backend 0.701, frontend 0.939).
Looks like Anthropic quietly updated the Sonnet 5 'Agentic search' benchmark graph overnight (www.reddit.comhttps) could not extract summary
Sonnet 5 full benchmark breakdown -- here's how it actually compares to Opus 4.8 and GPT-5.5 (www.reddit.com via reddit) Put together a comparison of every benchmark I could find from the official announcement and early coverage. Figured this might save people some time.
↯ Tool Use↯ Security↯ Swe Bench↯ Sonnet 4.6swe-benchtool-useprompt-injection+5
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching (arxiv.org) Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sessions consisting of…
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action (arxiv.org) Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands new evaluations. In this paper we evalua…
Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems (arxiv.org) Collaborative agentic AI is projected to transform entire industries by enabling AI-powered agents to autonomously perceive, plan, and act within digital environments. Yet, current solutions in this field are all built in isolation, and we…
Containment Verification: AI Safety Guarantees Independent of Alignment (arxiv.org) Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain conditional on unverifiable properties of learned behavior.
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent (arxiv.org) Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elicit genuine…
LLM-Empowered Agentic MAC Protocols: A Dynamic Stackelberg Game Approach (arxiv.org) Medium Access Control (MAC) protocols, essential for wireless networks, are typically manually configured. While deep reinforcement learning (DRL)-based protocols enhance task-specified network performance, they suffer from poor generaliza…
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning (arxiv.org) Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over…
ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping (arxiv.org) The wave of AI-native applications is moving shopping beyond page- and feed-based browsing toward intent-driven experiences orchestrated by LLM agents. A common design wraps an LLM around existing search and recommendation pipelines, forci…
ECHO: Prune to act, trace to learn with selective turn memory in agentic RL (arxiv.org) Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Existing context-management methods make such rollouts feasible by truncating distant history, folding…
DA-Studio: An Agentic System for End-to-End Data Analysis (arxiv.org) Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system should autonomously organize multi-step workflows, execute generated code in a sandboxed and control…
Agentic AI Enhances Physician Trust in Clinical Decision Making (arxiv.org) Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering intermediate reasoning steps and tool outputs transparent to users. Although proven to outperform prev…
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes (arxiv.org) Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pipelines depends on an implicit assumption: that the model ap…
AxDafny: Agentic Verified Code Generation in Dafny (arxiv.org) We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iteratively generates implementations, invari…
An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping (arxiv.org) High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them. At Oak Ridge National Laboratory's Advanced Plant Phenotyping Laboratory (APPL), automated stations image hundreds of plant…
A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols (arxiv.org) Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical e…
ACE: Pluggable Adaptive Context Elasticizer across Agents (arxiv.org) The increasing complexity of agentic tasks has led to rapidly growing trajectory lengths, which poses significant challenges for large language model (LLM) based agents with fixed context windows. Existing context management techniques, su…
Design and Implementation of Agentic Orchestrations and Orchestration of Agents (arxiv.org) Agentic Business Process Management has gained momentum recently. The prospect is that the autonomy of AI agents, i.e., predominantly LLM-based agents, can be balanced with a certain level of robustness, tractability, and traceability thro…
Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents (arxiv.org) Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation.
Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping (arxiv.org) Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similarity for object matching, neglecting physical affordances su…
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arxiv.org) As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agent…
Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics (arxiv.org) While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evade human detection. Formal mathematical languages like Lean 4 offer mechanical proof che…
AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance (arxiv.org) High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challenging due to restrictive language support and the gap between software and hardware prog…
Investigating Multi-Agent Deliberation in Law (arxiv.org) Artificial Intelligence is increasingly applied to the field of law, and has the potential to increase access to justice. One particular movement that is gaining traction is that of agentic AI, wherein AI agents, based on Large Language Mo…
Anthropic undercuts rivals with cheaper Claude Sonnet 5 (www.linkedin.com via reddit) Anthropic on Tuesday unveiled its latest Sonnet-class model, designed to deliver enhanced agentic capabilities at a competitive cost. Dubbed Claude Sonnet 5, Anthropic says the next-generation model allows users to autonomously complete co…
What does your day to day at work look like? (www.reddit.com via reddit) Just to share my professional experience, I've been working in the programming space for 6 years now, about 4 years as a Software Engineer, and the past 2 years as more of a Data Engineer. The company I work for has been relatively slow on…
Tested GLM 5.2 via BYOK on a real multi-file computer vision implementation task, here's what held up (www.reddit.comhttps) GLM 5.2 has been getting attention (MIT, 1M context, ~$1/$4.2 per M on OpenRouter, benchmarks near Opus 4.8). The pricing made me curious whether it could handle real agentic work or just one-shot answers.
Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction (arxiv.org) AutoB2G: Agentic Simulation and Reinforcement Learning for Spatio-Temporal Grid-Interactive Building Control (arxiv.org) Agentic AI for ISAC: Analysis, Framework, and Case Study (arxiv.org) StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley (arxiv.org) TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging (arxiv.org) Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction (arxiv.org) Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare (arxiv.org) Digitizing Coaching Intelligence: An Agentic Framework for Holistic Athlete Profiling using VLM and RAG (arxiv.org) DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification (arxiv.org) Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains a critical bottleneck for evidence-based medicine. While commercial black-box tools la…
DeepTrans Studio: Turning Expert Interventions into Shared Team Knowledge in Agentic Translation Workflows (arxiv.org) Professional translation is often a team-based process: translators, reviewers, and project managers must coordinate terminology, legal force, and accountability across documents. Yet many LLM-based translation tools treat human correction…
Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem (arxiv.org) Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combine natural language understanding with external services and APIs. LLM agents are LLM sy…
Agentic Abstention: Do Agents Know When to Stop Instead of Act? (arxiv.org) LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable in the available environment.
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization (arxiv.org) Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric appli…
Agentic Safety is an Epistemic Property, Not a Behavioral One (arxiv.org) Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods are necessary, but they primarily certify snapshots of system behavior.
TraceLab: Characterizing Coding Agent Workloads for LLM Serving (arxiv.org) Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this challenge requires understanding real workload patterns, yet the data needed for such analysis is la…
Building Multi-Task Agentic LLMs via Two-Phase Distillation (arxiv.org) A key step toward artificial general intelligence is to train models that can perform multiple tasks. In this paper, we study how to build such models by first training separate RL experts for individual tasks and then consolidating them v…
CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning (arxiv.org) Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same policy conditioned on privileged context. The prevailing recipe gates this loss by a single…
SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics (arxiv.org) As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources. However, choosing the right retriever remains dif…
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation (arxiv.org) Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another. This makes the common…
LAMP: Lean-based Agentic framework with MCP and Proof Repair (arxiv.org) Large language models are increasingly capable of mathematical reasoning, but the proofs they generate are often unreliable and hard to verify. Interactive theorem provers such as Lean 4 address this by accepting only kernel-checked proofs…
An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations (arxiv.org) Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use. This paper proposes an end-to-end agentic pipeline that combines deep time-series forecasting, variational anoma…
LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval (arxiv.org) We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structured documents must be applied efficiently without breaking cross-references or semantic consistency. L…
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent (arxiv.org) We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories an…
Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threats (arxiv.org) In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by the proliferation of electronic communication, social media, and advancements in artificial intelligence. As a result, there is an urgent n…
Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? (arxiv.org) Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliability of LaaJ for rubric scoring remains underexplored.
KbSD: Knowledge Boundary aware Self-Distillation for Behavioral Calibration in Agentic Search (arxiv.org) Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary calibration -- deciding when to trust parametric memory, when…
Parsing a request for a proprietary SQL (NexThink) input to output (www.reddit.com via reddit) For starters, I'll say I'm still fairly novice in understanding agentic anything and this might just be a project well over my head. Maybe that's a good thing, maybe it's a bad thing.
What was the full cost to get your AI agent setup off the ground? (www.reddit.com via reddit) My fully autonomous agentic system (Hermes/OpenClaw, OpenRouter, external paid/free tools), a cloud VPS is going to cost me $300-$400 a month to run. Curious what your guy's setups are and what they cost to get off the ground
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding (simonwillison.net) 29th June 2026 - Link Blog Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding. This is an interesting new open weights (MIT licensed) model, the first model release from DeepReinforce.
I’ve been scaling agentic coding at work, it’s mostly just “template skills” saved as markdown in gdrive folders… There a better way? (www.reddit.com via reddit) I’m at a late stage start up, so basically in the real shit of it. if we don’t hit 15% growth every quarter, forever, we will all lose our jobs and be sacrificed to the PE devils.
The Refine-Plan-Act Pattern (RPA) for Agentic AI Coding (medium.com via reddit) Over time I've landed on a simple pattern for agentic AI coding: Refine the requirements, Plan the approach, then Act on the implementation, each in a fresh context. This actually keeps my AI-generated code from getting messy.
MAX plan usage draining with ZERO activity — 1 week, no human response from support. Anyone else? (www.reddit.com via reddit) Posting here because I'm out of options through normal support. I'm on the MAX plan.
Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety (arxiv.org) As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While uncertainty quantification is well-studied for single-turn tasks, multi-turn agentic scena…
Agentic Episodic Control (arxiv.org) Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attempt to alleviate this via external memory modules, yet they suffer from two key limitations: a represe…
Agentic Hardware Design as Repository-Level Code Evolution (arxiv.org) We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled into a project pack containing domain knowledge, an executable evaluator, an acceptance pred…
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association (arxiv.org) Identifying robust associations between cardiac imaging phenotypes and clinical diseases is fundamental to population-scale cardiovascular research and reliable risk stratification. However, current phenome-wide association studies rely on…
Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy (arxiv.org) The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public. While prior research has established that mobility traces are highly unique an…
Agentic Publication Protocol: An Attempt to Modernize Scientific Publication (arxiv.org) Scientific publication is still organized primarily around static manuscripts, even though much of scientific progress depends on tacit know-how: how to run code, reproduce figures, interpret edge cases, choose useful follow-up directions,…
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning (arxiv.org) Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-horizon tasks. Unlike humans who employ "what-if" reasoning to evaluate potential plans be…
I spent 2 years reading AI engineering philosophy. The stuff that worked became 4 skills that gate my agents before they build. (www.reddit.com via reddit) My AI agents kept doing the same thing: jumping straight to a solution before understanding the problem. Ask it to design a prompt.
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps (arxiv.org) Frontier deep research agents (DRAs) are being deployed in enterprise workflows faster than they are being evaluated. Existing benchmarks measure factual recall, single-hop QA, or generic agentic skill, and miss the multi-document, decisio…
Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities (arxiv.org) AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violations and efficiently filters false positives. Many dangerous vulnerability classes, such as cryptograp…
MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG (arxiv.org) Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-team…
HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization (arxiv.org) Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human…
Localizing RL-Induced Tool Use to a Single Crosscoder Feature (arxiv.org) Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basis of these changes remains poorly understood. While RL substantially improves structured…
EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning (arxiv.org) In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging because each candidate must be trained before evaluation, and the design space is open-ended.
The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators (arxiv.org) Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmar…
A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO (arxiv.org) We introduce the process harness, a new mechanism for uplifting legacy workflows into Agentic Business Process Management (Agentic BPM) without replacing the underlying workflow engine. A process harness places a policy-governed agentic la…
When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework (arxiv.org) Agentic systems are changing how complex operational tasks are coordinated, introducing a new paradigm for connecting heterogeneous data sources and automating processes. Electric bus fleets provide a relevant test case.
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems (arxiv.org) Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of others despite no shared variable or executable dependency. We formalize this as compositional beha…
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? (arxiv.org) Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-domain evaluations remain largely limited to static knowledge recall. This is a critical ga…
Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking (arxiv.org) Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records, which are authoritative but abstract, and patient narratives, which are experience-near…
Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols (arxiv.org) As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integra…
PreHook command Gate policy layer for all Claude code agents (www.reddit.comhttps) Hello everyone, I recently was fed up with agents running unsupervised commands on my systems and wanted to solve this problem. The problem was simple, Claude code model “fable 5” uses safety flags in the UI layer that prevented the model…
Is there a standard for porting agent state across models, or are we all writing custom wrappers? (www.reddit.com via reddit) Hey everyone, I'm fairly new to the agentic workflows space. Really interested to get into it.
Which model for technical documentation? (www.reddit.com via reddit) Looking to create high level / low level designs (software), based on existing templates/examples, cross reference code, use mcp to download confluence/jira data - also plug into agentic ‘coding’ frameworks opencode . I mostly use opus 3.6…
Built a loop engineering skill for PRs in Claude Code — branch, two independent reviews, CI handling, merge handoff. Here's what I learned building it. (www.reddit.comhttps) Most agentic coding tools are really good at one thing: writing code fast. you describe a problem, they implement it, done.
Claude Max vs Codex Pro or both combined? (www.reddit.com via reddit) I’m considering one heavier subscription (~€100/month) and want to know which provides better value for agentic coding. I tested GPT Pro and was satisfied with Codex.
New to Reddit & starting my journey to become a Gen AI / Agentic AI Dev. Looking to connect and learn. (www.reddit.com via reddit) Hey guys, I'm new to Reddit and just starting out learning Gen AI and Agentic AI. I really want to connect with people in this field for some guidance, networking, and just to talk.
How we made an AI agent faster by moving stable context out of the prompt (www.reddit.com via reddit) Many AI agents are expensive because they keep rediscovering the same context. At a recent BotsCrew Spotlight, one team shared what they learned while building an enterprise analytics assistant.
A Probabilistic Framework for LLM-Based Model Discovery (arxiv.org) Automated methods for discovering mechanistic simulator models from observational data offer a promising path toward accelerating scientific progress. Such methods often take the form of agentic-style iterative workflows that repeatedly pr…
Agentic Software Engineering: Foundational Pillars and a Research Roadmap (arxiv.org) Agentic Software Engineering (SE 3.0) represents a new era where intelligent agents are tasked not with simple code generation, but with achieving complex, goal-oriented SE objectives. To harness these new capabilities while ensuring trust…
Governing Technical Debt in Agentic AI Systems (arxiv.org) Agentic AI systems are increasingly being explored as production infrastructure: they reason over multiple steps, call tools, act through workflows, and adapt through memory and feedback. These systems create governance challenges that are…
Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces (arxiv.org) As LLM agent systems take on more complex tasks, they increasingly rely on meta-agents: higher-order agents that create, operate on and manage other agents. Meta-agent operations such as coordinating agents, halting risky actions before ex…
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows (arxiv.org) Agentic AI systems are increasingly being integrated into scientific workflows, yet their behavior under realistic conditions remains insufficiently understood. We evaluate CMBAgent across two workflow paradigms and eighteen astrophysical…
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents (arxiv.org) Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make bot…
Is GraphRAG Needed? From Basic RAG to Graph-/Agentic Solutions with Context Optimization (arxiv.org) As advanced RAG variants like GraphRAG and Agentic RAG emerge, one leading question is when and how to use them. Here, we introduce a framework for different RAG scenarios evaluation and comparison on semi-structured knowledge bases, inclu…
Autodata: An agentic data scientist to create high quality synthetic data (arxiv.org) We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create eve…
Agentic System as Compressor: Quantifying System Intelligence in Bits (arxiv.org) Large language models are turning from isolated predictors into agentic systems: they call tools, retrieve evidence, obey environment constraints, use verifiers, and complete tasks through search and multi-turn interaction. We adopts an an…
AI Snitches Get Glitches: Towards Evading Agentic Surveillance (arxiv.org) To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employers (and even nation-states) already provide their users with this technology.
Agentic evolution of physically constrained foundation models (arxiv.org) Artificial intelligence increasingly drives automated scientific discovery, yet contemporary generalist agents lack physical grounding, frequently hallucinating hardware-incompatible designs. Here, we present a physically grounded, multi-a…
Agentic Knowledge Tracing: A Multi-Agent LLM Architecture for Stealth Assessment of Financial Literacy in Serious Games (arxiv.org) Assessing financial literacy during gameplay without disrupting the learning experience remains a key challenge in serious games for education. We present the Agentic BKT pipeline, a multi-agent large language model architecture for stealt…
Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval (arxiv.org) Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate long-horizon trajectories. While Multi-Agent Debate (MAD) succeeds in deterministic domains, agents…
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems (arxiv.org) The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the full stack from first principles to production deployment, organized around a central thesis: building…
As a solo builder I created a multi tenant B2B SaaS for commercial maintenance companies that is agentic AI capable in 2 months using Claude code. (www.reddit.com via reddit) Hello everyone, I began working on this project on April 16th. Some quick background.
Introducing computer use in Gemini 3.5 Flash (deepmind.google) Introducing computer use in Gemini 3.5 Flash Computer use is now a built-in tool supported in Gemini 3.5 Flash, delivering our best performance yet for agentic computer use tasks. Previously only available as a standalone Gemini 2.5 comput…
↯ Gemini 3.5↯ Gemini 3.5↯ Gemini 3.5↯ Gemini 3.5↯ Gemini 3.5↯ Gemini 3.5↯ Gemini 3.5↯ Gemini 3.5geminiagentic
I built a local, open-source tool that turns your Claude Code prompts into a self-portrait of how you code (www.reddit.comhttps) I built devbrain, an open-source local dashboard for your Claude Code history. It shows when you work, where your tokens go, what you keep asking for, and TODOs extracted from your prompts.
Software development has entered its "infinite monkeys" era (www.reddit.com via reddit) With the rise of agentic coding tools like Claude Code, Cursor, and Codex, the barrier to entry is gone. Now, anyone with an internet connection can "type." We have essentially reached the infinite monkey phase of software development.
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning (arxiv.org) Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating sparse evidence across a large, messy collection of workplace…
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning (arxiv.org) Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world interaction. However, existing experience learning methods mostly rely on single-agent loops, where the same agent executes ta…
Toward Autonomous O-RAN: A Multi-Scale Agentic AI Framework for Real-Time Network Control and Management (arxiv.org) Open Radio Access Networks (O-RAN) promise flexible 6G network access through disaggregated, software-driven components and open interfaces, but this programmability also increases operational complexity. Multiple control loops coexist acr…
ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms (arxiv.org) Progress in computational science depends on complex numerical workflows that must faithfully encode physical laws, yet translating conceptual insight into reliable code remains a major bottleneck. Although large language models can genera…
Multimedia and Visual Analytics in the Agentic Era (arxiv.org) Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have rapidly changed the playing field, and improving their accuracy, trustworthiness, and reasoning cap…
Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce (arxiv.org) Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close a sale. We argue that the arrival of agent-native micro-payment rails (e.g., x402, AP2) changes wha…
DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects (arxiv.org) Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability. In the subset with suspected genetic etiologies, exome and genome sequencing have moved many cases from variant detection to post-sequencing interp…
Red-Teaming the Agentic Red-Team (arxiv.org) The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the community has focused on creating more and more capable agents, less attention ha…
OpenThoughts-Agent: Data Recipes for Agentic Models (arxiv.org) Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically ta…
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System (arxiv.org) Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses.
SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation (arxiv.org) As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constraints of even the largest context windows. Current methods for effectively diagnosing agent failures…
Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems (arxiv.org) Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance optimizat…
OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility (arxiv.org) For a wheelchair user, a standard blue line on a map is often a broken promise. While platforms like OpenStreetMap (OSM) successfully capture where a path is, they frequently fail to convey how it physically feels to travel on it.
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection (arxiv.org) Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle text--image framing errors. Existing benchmarks and methods remain poo…
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems (arxiv.org) Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied…
I'm building agent loops that auto-edit my videos, but the hard part has been finding a model to accurately grade the result (youtube.com via reddit) Quick context: I've been building agentic loops that edit my short-form videos for me. The editing works really well, but I found myself needing to check the process at several gates.
Memory layer situation in claude and other agentic ecosystems (www.reddit.com via reddit) Every other day someone ships a new memory layer for AI agents. Claude has its own memory system, ChatGPT has one, and I've written my own.
Build real agentic apps using CUGA: two dozen working examples on a lightweight harness (huggingface.co) Build real agentic apps using CUGA: two dozen working examples on a lightweight harness TL;DR — Building an agent is mostly plumbing: tools, state, guardrails, scaling from one agent to many. CUGA (pip install cuga), short for Configurable…
RAVEN: Agentic RAG for Automated Vulnerability Repair (arxiv.org) Automated vulnerability repair has emerged as a promising direction to mitigate the growing number of software vulnerabilities. Recent advances in Large Language Models (LLMs) have further accelerated research in automated repair.
MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems? (arxiv.org) Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs inter-agent coordination and output aggregation. Syst…
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning (arxiv.org) Reinforcement learning has empowered large language models to act as intelligent agents, yet training them for long-horizon tasks remains challenging due to the scarcity of high-quality trajectories, especially under limited resources. Exi…
ATLAS: Agentic Taxonomy of Large-Scale Software Ecosystems (arxiv.org) The open-source ecosystem on GitHub lacks a systematic hierarchical taxonomy of software repositories. GitHub Topics, the dominant organizational mechanism, is flat, inconsistent, and covers only 67% of projects.
OpenBioRQ: Unsolved Biomedical Research Questions for Agents (arxiv.org) A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over $99\%$ resolve), yet roughly $15.9\%$ link to th…
EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory (arxiv.org) Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This paper introduces EvoEmbedding, a novel embedding model that generates evolvable representa…
Dissecting Agentic RAG: A Component Ablation for Multi-Hop QA with a Local 7B Model (arxiv.org) Agentic retrieval-augmented generation (RAG) systems combine iterative reasoning loops, query decomposition, and adaptive retrieval to tackle multi-hop question answering. However, the contribution of each component remains poorly understo…
SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding (arxiv.org) Scientific discovery increasingly relies on automated systems that generate hypotheses, inspect multimodal evidence, and validate claims at scale. Yet scientific claim verification is not well served by asking a vision-language model for a…
SAGE: A Novelty Gate for Efficient Memory Evolution in Agentic LLMs (arxiv.org) Agentic LLMs must continuously decide whether newly extracted facts should be added, merged with existing memories, or ignored, yet prior work has focused more on retrieval and storage than on principled write-side control. We frame memory…
Scaling Small Agents Through Strategy Auctions (arxiv.org) Small language models are increasingly viewed as a promising, cost-effective approach to agentic AI, with proponents claiming they are sufficiently capable for agentic workflows. However, while smaller agents can closely match larger ones…
TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion (arxiv.org) Hardware Trojans (HTs) remain a critical threat because learning-based detectors often overfit to narrow trigger/payload patterns and small, stylized benchmarks. We introduce TrojanGYM, an agentic, LLM-driven framework that automatically c…
From RAG to Agentic RAG for Faithful Islamic Question Answering (arxiv.org) Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. Yet standard MCQ/MRC-style evaluations (MCQ: Multiple choice questions, MRC: Machine Re…
Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning (arxiv.org) We present Tell Me, a mental well-being system that leverages advances in large language models to provide accessible, context-aware support for users and researchers. The system integrates three components: (i) a retrieval-augmented gener…
Towards Adaptive Categories: Dimensional Governance for Agentic AI (arxiv.org) As AI systems evolve from static tools to dynamic agents, traditional categorical governance frameworks -- based on fixed risk tiers, levels of autonomy, or human oversight models -- are increasingly insufficient on their own. Systems buil…
Agent Skill Framework: Perspectives on the Potential of Small to Medium Language Models in Industrial Environments (arxiv.org) Agent skills are widely supported by major agentic frameworks and perform well with proprietary models, yet their effectiveness for small and medium-sized open source language models (270 M-80B) remains underexplored. We systematically stu…
RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation (arxiv.org) Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when handling ambiguous intentions, logical reasoning, and Out-of-Distribution (OOD…
RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents (arxiv.org) Agentic coding harnesses - such as Agent-Skills, Superpowers, and Agent-Rigor - are increasingly deployed to augment underlying LLMs for real-world software engineering tasks. Existing benchmarks evaluate these agents almost exclusively on…
Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases (arxiv.org) Memory safety vulnerabilities remain a significant threat even for projects with extensive fuzzing and manual auditing. Recent results suggest that large language models hold great promise for detecting such vulnerabilities, but they are u…
TraceView: Interactive Visualization of Agentic Program Repair Trajectories (arxiv.org) LLM-based automated program repair (APR) agents generate patches to fix software bugs with minimal human intervention. These agents often produce long trajectories of reasoning, tool use, and feedback to produce candidate patches.
From RAN Control to Agentic Intelligence: Architecture and Vision for Energy Efficient AI-RAN (arxiv.org) Future 6G networks will rely on highly distributed, AI-native Radio Access Networks (RANs), where communication and AI workloads share a common infrastructure. This evolution, combined with increasing deployment density and continuous AI p…
Skills for the future software profession: beyond agentic AI! (arxiv.org) As coding agents are rapidly changing software engineering, a natural question is: what are the core skills needed by future software engineers? To identify where software engineering is headed and thus what skills will be needed, we summa…
SwarmX: Agentic Scheduling for Low-Latency Agentic Systems (arxiv.org) Agentic AI applications compose multiple model calls and tool executions, creating new scheduling challenges for GPU-CPU clusters. Their inference time and model-call structure often depend on prompt semantics, making conventional scheduli…
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams (arxiv.org) Massive unstructured multimodal streams suffer from high "data entropy," impeding both efficient human knowledge acquisition and high-quality AI post-training. Existing passive annotation paradigms, heavily reliant on heuristic rules or ge…
One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception (arxiv.org) Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, real-world spatiotemporal data exhibits severe heterogeneity, often manifesting as extreme…
Role-Based Agentic AI for Intent-Driven Network and Service Orchestration (arxiv.org) Telecommunication networks are increasingly complex due to heterogeneous technologies, diverse service requirements, and growing demands for resource efficiency and business agility. Intent-Based Networking (IBN) and, more recently, agenti…
Infrastructure for the Agentic Web: Gap Analysis and Architecture from the Agentverse Platform (arxiv.org) The emergence of autonomous AI agents as first-class participants in digital infrastructure marks a fundamental inflection point in the evolution of the Web. While significant research has been directed at agent behaviour and reasoning, co…
AI-Native Network Controller: A Modular Framework for Safe Agentic Control of Multi-Domain Network Infrastructure (arxiv.org) The convergence of multiple network domains, including radio access, optical transport, and core networks, under unified intelligent control is a fundamental requirement for future 6G systems. This is important because existing network con…
GIF: Locally Sound Geometric Information Flow Control for LLMs (arxiv.org) Large language models increasingly mediate interactions between sensitive data, untrusted inputs, and privileged actions in agentic systems, creating security and privacy risks. These range from prompt injections that manipulate downstream…
Agent-as-a-Router: Agentic Model Routing for Coding Tasks (arxiv.org) Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct domains, yet none dominate all. Consequently, routing each task to the most suitable model bec…
RaMem: Contextual Reinstatement for Long-term Agentic Memory (arxiv.org) Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems have made past experiences more persistent, compact, and retrievable, but retriev…
Grounded Scaling: Why Agentic AI Needs Deterministic Environments (arxiv.org) Long-chain agent execution fails exponentially in environments designed for human tolerance: with per-step determinism $\delta < 1$, $k$-step chain success degrades as $\delta^k$. The AGI-to-ASI scaling debate (Genewein et al., 2026) has s…
Holmes: Multimodal Agentic Diagnosis for Mixed-Language Mobile Crashes at Industrial Scale (arxiv.org) Diagnosing mobile crashes in ultra-large-scale industrial applications is a formidable challenge due to the sheer volume of code, the complexity of mixed-language environments, and the inability to reproduce failures locally. Traditional s…
AgentRiskBOM: A Risk-Scoping Security Bill of Materials for Agentic AI Systems (arxiv.org) Agentic AI systems retrieve private context, invoke tools, write files, call external services, coordinate with other agents, and may act without human approval. Existing bill of materials artifacts improve transparency for dependencies, m…
Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents (arxiv.org) Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by o…
Counsel: A Meta-Evaluation Dataset for Agentic Tasks (arxiv.org) As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agentic benchmarks can take hours, making it difficult to scale…
Composing Verifiable Conceptual Models via Building Blocks: Towards Design-Time Verification of Agentic AI Workflows (arxiv.org) Agentic AI systems orchestrate multiple LLM-based agents through workflow architectures that coordinate decisions, tools, and external actions. While current platforms emphasize runtime safeguards, little support exists for verifying workf…
AutoRAS: Learning Robust Agentic Systems with Primitive Representations (arxiv.org) The automated design of agentic systems offers a promising pathway for scaling large language models (LLMs) beyond single-agent reasoning. While prior work has advanced task performance through handcrafted or automatically generated multi-…
Agentic Time Machine as an Infrastructure for Future-Event Forecasting (arxiv.org) Forecasting future events is a critical challenge for large language model (LLM) agents, spanning domains from elections and monetary policy to financial markets. However, evaluating progress on this task presents a fundamental trade-off b…
Democratizing and accelerating AI-driven pathology research through agentic intelligence (arxiv.org) Computational pathology has advanced rapidly with the emergence of foundation models, yet widespread adoption remains limited by substantial technical complexity and programming requirements. Here we present PathLab, an autonomous agentic…
A Quantum-Assisted Agentic Distributed Artificial Intelligence Framework for Deadline-Bounded Orchestration of Hybrid Renewable Microgrids (arxiv.org) The real-time orchestration of microgrids that combine fluctuating renewable sources, dispatchable units, storage and curtailable consumers requires the repeated solution of combinatorial dispatch and coalition formation problems under har…
Everyone thinks I'm a vibe coder because I put a Claude sticker on my laptop (www.reddit.com via reddit) I was just at a coffee shop when someone said, "I like your vibe," to me. I thanked him and then he pointed at my Claude sticker and clarified his original statement.
Connected a Robinhood Account to Claude Code and Codex for Autonomys Agentic Trading... Update 1 (www.reddit.com via reddit) Update to my original post: https://www.reddit.com/r/ClaudeAI/comments/1u8nagi/connected_a_robinhood_account_to_claude_code_and/ I'm building a fully autonomous daily stock-trading desk in a Robinhood "Agentic" account. Opus is the CEO/PM,…
what is agentic coding and why is everyone suddenly talking about it (www.reddit.com via reddit) I kept seeing "agentic coding" everywhere for the last few months, blog posts, Twitter threads, product launches, and I honestly couldn't tell if it was a real shift or just the latest marketing rebrand for AI autocomplete. So I spent the…
Lighthouse agentic browsing scoring (developer.chrome.com) The Agentic Browsing category evaluates how well your site is constructed for machine interaction through a set of deterministic audits. How the category is scored Unlike other Lighthouse categories, the Agentic Browsing category does not…
Server hosting for Agents (www.reddit.com via reddit) I am a non-developer about to build a CRM like agentic workflow for my self and maybe 1 other team member of a company that I own and operate. I saw some options but wanted to see what people thought for railway or Render or other options?
Is there actually a good way to orchestrate multiple agents, or is everyone just running a bunch of terminals? (www.reddit.com via reddit) A couple weeks ago I saw someone with 6 instances of Claude Code open, each in its own window, switching between them by hand. And the thing is, that seems to be roughly the state of the art right now.
Agentic Symbolic Search: Characterizing PDEs Beyond Hand-crafted Expressions, Meshes, and Neural Networks (arxiv.org) Mathematicians understand a PDE solution through mathematical structures rather than tables of computed values. Historically, this has been the product of mathematical analysis, carried out by hand for each problem individually.
TSAssistant: A Human-in-the-Loop Agentic Framework for Automated Target Safety Assessment (arxiv.org) Target Safety Assessment (TSA) requires systematic integration of genetic, transcriptomic, target homology, pharmacological, and clinical data to evaluate potential safety liabilities of therapeutic targets. This process is labor-intensive…
Prompt, Plan, Extract: Zero-Shot Agentic LLMs Workflows for Lung Pathology Extraction from Clinical Narratives (arxiv.org) Information extraction from pathology reports is essential for cancer staging, tumor registry population. Yet key data remains embedded in narrative reports, making manual extraction labor-intensive and error-prone.
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning (arxiv.org) Solving mathematical reasoning problems requires not only accurate access to relevant knowledge but also careful, multi-step thinking. However, current retrieval-augmented models often rely on a single perspective, follow inflexible search…
Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes (arxiv.org) Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes. Existing access-control mechanisms authorize i…
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems (arxiv.org) Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequ…
ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments (arxiv.org) Academic paper search is a core step in scientific research, and LLM-based search agents are emerging as a promising paradigm for iterative, intent-driven literature exploration. However, existing benchmarks are insufficient for systematic…
AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models (arxiv.org) We propose a model-grounded RAG-based AI economist with an agentic framework for economic scenario analysis using large language models (LLMs) and knowledge graphs. While LLMs can generate fluent economic narratives, economists are often r…
Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services (arxiv.org) In the agentic web era, LLM-based agents increasingly invoke web services as tools, yet most interfaces remain \emph{static endpoints} that poorly express long-horizon workflows with loops, conditionals, joins, and retries. We present Tool…
Measuring Biological Capabilities and Risks of AI Agents (arxiv.org) This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively perfor…
Agentic Electronic Design Automation: A Handoff Perspective (arxiv.org) Electronic design automation (EDA) is inherently multi-stage and handoff-heavy. Design artifacts, flow scripts, and engineering decisions cross tool, session, and organizational boundaries before final implementation, signoff, or release.
Playful Agentic Robot Learning (arxiv.org) Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions.…
Execution-bound advisory automation for agentic AI: a reproducible AIBOM-driven CSAF-VEX framework (arxiv.org) A protocol driven framework is presented that binds SBOM and AIBOM artefacts to deterministic environment capture and structured runtime telemetry. Exploitability is computed from declared artefacts, observed activation conditions, and enf…
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World (arxiv.org) Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents c…
Benchmarking Agentic Review Systems (arxiv.org) A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coa…
Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why (arxiv.org) Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-level metadata that AI systems need for retrieval and triage is absent or incomplete. Standard retrieval-augmented generati…
DeXposure-Claw: An Agentic System for DeFi Risk Supervision (arxiv.org) Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no…
Deontic Policies for Runtime Governance of Agentic AI Systems (arxiv.org) Autonomous agentic AI systems driven by Large Language Models (LLMs) introduce a new class of security, privacy, and compliance challenges: an agent that can invoke tools, manipulate data, install software, and coordinate with peer agents…
Hands-on prep for the Claude Certified Architect (CCA-F) exam (www.reddit.com via reddit) A few months ago I posted about a dedicated suite of practice tests I built for Claude Certified Architect - Foundations. Thanks a lot to all of you for the feedback!
How are you guys using ai to increase productivity. (www.reddit.com via reddit) What i mean by this not opening claude and adding claude.md file or setting etc. This it's self is an art how you setup and prompt.
Warp/cursor vs Claude Code native app (www.reddit.com via reddit) I am not a coder and I am actually pretty new to this vibe coding world and agentic AI (I just try having a second brain with some agents), can someone explain me why everyone is using warp or cursor? What's the difference between those an…
Building independent LLM drift detection - sharing the methodology, looking for feedback on the approach (www.reddit.com via reddit) Disclosed upfront: I run [Tickerr dot ai], an independent external monitor for AI APIs. Today it tracks latency, TTFT, uptime, and error rates across major models.
Code-Augur: Agentic Vulnerability Detection via Specification Inference (arxiv.org) The advent of agentic vulnerability detection is already becoming a watershed moment for software security. Audits conducted entirely by autonomous LLM agents are uncovering critical vulnerabilities in fundamental software underpinning dig…
Mitigating Anchoring Bias in LLM-Based Agents for Energy-Efficient 6G Autonomous Networks (arxiv.org) This paper presents an autonomous agentic resource negotiation framework designed to enable zero-touch network slicing in 6G architectures using Large Language Model (LLM) agents. While LLMs offer powerful reasoning capabilities, we demons…
ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch (arxiv.org) Bringing Large Language Models (LLMs) into industrial ride-hailing dispatch as semantic feature extractors over platform-scale behavioral logs is a compelling but under-explored data systems problem. Production matching pipelines remain do…
ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift (arxiv.org) Modern AI agents retrieve documents, call tools, check intermediate information, and then produce a final answer or action. This creates a risk-control problem that is not visible from the final answer alone.
Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems (arxiv.org) Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default language for that exchange, JSON, was designed for application-to-application interchange rath…
Is it agentic enough? Benchmarking open models on your own tooling (huggingface.co) Is it agentic enough? Benchmarking open models on your own tooling Benchmarking transformers revisions across different metrics This is a human-made, agent-focused blogpost.
Claude will execute trades in some chats and flat-out refuse in others. How do you get it to behave consistently? (www.reddit.com via reddit) Bear with me, I'm more of an architecture guy than a coder, so apologies if I'm missing something obvious. And I'm an old guy, so some generational frustration is in here too.
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning (arxiv.org) Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents. However, conventional RL methods for long-horizon agentic tasks often struggle with sparse outcome rewards.
Securing Multi-Agent GIS Systems: Risk Evaluation and Prompt Hardening Optimization (arxiv.org) Agentic systems are increasingly integrated with geographic information systems (GIS), where multi-agent coordination enables complex conversational and spatial analysis but introduces security risks. This work presents a security-oriented…
Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety (arxiv.org) Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), including incident investigation, root-cause analysis, configuration synthesis, and limited sel…
MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation (arxiv.org) Lane-level maps are critical infrastructure for autonomous driving and lane-level navigation, yet constructing and maintaining standardized lane networks for hundreds of cities remains highly labor-intensive. Recent end-to-end vectorized m…
A T-API-Compliant ReAct Agentic Loop for Optical Networks: Generic vs. Domain-Specific Tool Abstractions (arxiv.org) Optical networks need intent-driven, closed-loop agentic management, a key enabler for higher autonomy levels. We present the first T-API-compliant reasoning and act (ReAct) loop.
A Framework for Evaluating Agentic Skills at Scale (arxiv.org) Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet their cross-domain impact and use across commercial and open-source models remain under-studied, and…
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering (arxiv.org) Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era: they collapse model, harness, and environment into a single end-to-end score, typically computed ag…
Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation (arxiv.org) Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents continuously acquire information, form beliefs regarding latent states of the environment, generate foreca…
Agentic Discovery of Non-Canonical Antimicrobial Peptides with AMPGAN v3 (arxiv.org) Antimicrobial resistance causes to over a million deaths annually. Antimicrobial peptides (AMPs) are a promising solution, but generative AMP models are not yet ready to design peptides with non-natural amino acids and/or chemical modifica…
CMIP-Forge: An Agentic System that Retrieves, Computes, and Self-Reviews Climate Science (arxiv.org) The Coupled Model Intercomparison Project Phase 6 (CMIP6) has generated thousands of peer-reviewed publications documenting model configurations, evaluation procedures, emergent constraints, and projection uncertainties. As the community t…
Learning Cardiac Electrophysiology Digital Twins Through Agentic Discovery of Hybrid Structure (arxiv.org) Building personalized cardiac electrophysiology (EP) digital twins requires identifying the appropriate model structure for each patient, not merely fitting parameters. Traditional methods rely on experts to manually prescribe hybrid physi…
WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning (arxiv.org) Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians. However, answering questions about wearable health data remains challenging and understudied, as these ubiqu…
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models (arxiv.org) AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmarks for AI and animal welfare evaluate model text responses to question-answer prompts, leaving open…
Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications (arxiv.org) Recent advances in Large Language Models (LLMs) and multi-agent systems have driven the rise of Agentic AI, showing promise for medical reasoning. However, open-ended conversational agents remain prone to two critical failure modes: premat…
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience (arxiv.org) As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important. Otherwise, such systems may rapidly generate plausible yet misleading studies that contaminate…
Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search (arxiv.org) Test-time scaling for agentic search typically increases depth (i.e., more turns and tokens per trajectory) or breadth (i.e., more parallel rollouts). Here we focus on breadth scaling, showing that standard parallel sampling yields diminis…
Agentic Resource Discovery: Let agents search (huggingface.co) Agentic Resource Discovery: Let agents search for tools, skills, and other agents. The Agentic Resource Discovery (ARD) specification is the discovery layer that sits in front of them.
Cursor launches github for agentic era (www.reddit.com via reddit) Three big launches: - Cursor iOS app - Origin, an agentic replacement of Git - A new model is in the works in collaboration with spaceX - SpaceX to acquire Cursor
Spent $11k evaluating Fable: capability looked SOTA, refusals killed it (before Anthropic did) (www.reddit.com via reddit) Before its suspension, I spent $11,081.12 evaluating Claude Fable 5 on WolfBench, an agentic benchmark based on Terminal-Bench 2.0. It was by far my most expensive benchmark run ever, and I fully expected Fable to become the new top model…
Is Claude code web capable of unlocking advanced agentic loop workflows ? (www.reddit.com via reddit) Not talking about the Remote feature where it manages your desktop session remotely, just a purely remote session on Anthropic’s infra. Have you unlocked any of this, asking because it seems most of the advanced features seem to be more su…
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems (arxiv.org) Large language models (LLMs) are increasingly deployed as the execution core of autonomous agents rather than as standalone text generators. Agentic workloads induce a temporal shift from single-turn inference to multi-turn LLM-tool loops,…
MIRAGE: Auditing Anti-Muslim Bias in Frontier LLMs Across Reasoning, Agentic, and Time-Coupled Conditions (arxiv.org) Five years after the discovery of persistent anti-Muslim bias in large language models, most evaluations remain confined to single-turn prompt completion, a setting that no longer reflects how frontier LLMs are deployed. We introduce \text…
All-Mem: Agentic Lifelong Memory via Dynamic Topology Evolution (arxiv.org) Lifelong interactive agents are expected to assist users over months or years, which requires continually writing long term memories while retrieving the right evidence for each new query under fixed context and latency budgets. Existing m…
Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning (arxiv.org) Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alignment of instruction-tuned (IT) models.
Context-Aware RL for Agentic and Multimodal LLMs (arxiv.org) Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool trace or a subtle detail in an image. We propose ContextRL…
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection (arxiv.org) SMS fraud is increasingly cross-channel: a message directs the user to a webpage, and the final risk depends on how the SMS claim aligns with the page content and requested user action. However, existing evaluations either focus on message…
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning (arxiv.org) We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should uncover a hidden deterministic finite automaton (DFA) by intera…
PathRouter: Aligning Rewards with Retrieval Quality in Agentic Graph Retrieval-Augmented Generation (arxiv.org) Agentic GraphRAG trains language-model agents to iteratively retrieve and reason over graph-structured evidence, enabling more accurate and context-aware decision-making by efficiently navigating complex information networks. However, outc…
Interactor: Agentic RL oriented Iterative Creation for Ad Description Generation in Sponsored Search (arxiv.org) This paper focuses on automatically generating informative ad descriptions in sponsored search. Unlike ad titles which are usually optimized to attract user click feedbacks, ad descriptions have a longer text span and possess the potential…
TechRAG: Evidence-Gated Multimodal Agentic RAG for Technical Literature Reasoning (arxiv.org) This paper presents an agentic multimodal retrieval-augmented generation (RAG) framework for domain-specific literature reasoning, instantiated on a curated corpus of several thousand papers in intelligent tires, vehicle dynamics, vehicle…
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs (arxiv.org) Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when using traditional business intelligence tools or Text-to-SQL systems. While recent Text-to-SQL approache…
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate (arxiv.org) Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is also a canonical testbed for agentic language models, requiring multi-turn interaction under hidden pre…
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw (arxiv.org) Agentic language-model systems increasingly rely on mutable execution contexts, including files, memory, tools, skills, and auxiliary artifacts, creating security risks beyond explicit user prompts. This paper presents DeepTrap, an automat…
AgenticRec: A Recommendation-Oriented Agentic Framework with Progressive Tool-Integrated Reasoning Optimization (arxiv.org) Recommender agents built on Large Language Models offer a promising paradigm for personalized recommendation. However, existing agents typically suffer from a misalignment between their tool-integrated reasoning trajectories and recommenda…
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks (arxiv.org) Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sites and performing actions on users' behalf. While these agents offer powerful capabilities, their de…
Learning to Share: Selective Memory for Efficient Parallel Agentic Systems (arxiv.org) Agentic systems solve complex tasks by coordinating multiple agents that iteratively reason, invoke tools, and exchange intermediate results. To improve robustness and solution quality, recent approaches deploy multiple agent teams running…
EffGen: Enabling Small Language Models as Capable Autonomous Agents (arxiv.org) Most existing language model agentic systems today are built and optimized for large language models (e.g., GPT, Claude, Gemini) via API calls; while powerful, this approach faces several limitations including high token costs and privacy…
RollArt: Disaggregated Multi-Task Agentic RL Training at Scale (arxiv.org) Agentic Reinforcement Learning (RL) trains LLMs through multi-turn interactions with environments, producing workloads that mix compute-bound prefill, bandwidth-bound decoding, CPU-heavy environment execution, and bursty reward evaluation.…
Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage? (arxiv.org) In the age of agentic AI, the growing deployment of multi-modal models (MMs) has introduced new attack vectors that can leak sensitive training data in MMs, causing privacy leakage. This paper investigates a black-box privacy attack, i.e.,…
A Survey on Agentic Security: Applications, Threats and Defenses (arxiv.org) LLM-based agents are now used throughout cybersecurity. While these agents facilitate powerful and autonomous security applications, their autonomy opens up new attack surfaces, and the security community is actively building defenses to s…
SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search (arxiv.org) Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these systems often suffer from a critical limitation in practice: agents fail to recognize their…
MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition (arxiv.org) Therapeutic decision-making in clinical medicine constitutes a high-stakes domain in which AI guidance interacts with complex interactions among patient characteristics, disease processes, and pharmacological agents. Tasks such as drug rec…
Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents (arxiv.org) The path toward autonomous software engineering is currently bottlenecked by a severe deficit of diverse, large-scale trajectory data. We address this by introducing \ourdataset, an expansive dataset of 207,489 agentic trajectories spannin…
Green SARC: Predictive Cost and Carbon Governance for Agentic AI Systems (arxiv.org) Agentic AI systems act through tools and sub-agents, yet the controls meant to bound their financial and environmental cost still sit on dashboards evaluated beside or after execution. Green SARC applies the SARC governance-by-architecture…
MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA (arxiv.org) Long-document multimodal question answering requires a system to locate sparse evidence in long PDFs and integrate clues from text, tables, images, charts, and complex layouts. Existing RAG methods mostly rely on fixed Top-k retrieval over…
The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products (arxiv.org) Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same characteristics can create or exacerbate product risks.
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale (arxiv.org) Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2…
Resilient Consensus in Agentic AI (arxiv.org) Large language model (LLM) agents are increasingly deployed in multi-agent systems where they must coordinate and agree on shared decisions. We ask whether classical resilient consensus theory, developed for deterministic agents, transfers…
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning (arxiv.org) We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M…
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment (arxiv.org) LLMs have substantially improved software engineering yet real-world development requires architectural understanding. Such understanding is prohibitively expensive to label manually and impossible to verify through tests alone.
A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framework Development (arxiv.org) This paper presents a structured analysis of security challenges in long-horizon agentic AI systems. The study reviews existing threats, evaluation approaches, attack propagation mechanisms, and security frameworks.
Agentomics: Economic Foundations for the Valuation, Attribution, and Pricing of AI Agents in Human-AI Workflows (arxiv.org) Agentic AI systems are increasingly being deployed as productive resources in organizational workflows, yet existing evaluation methods primarily measure isolated technical performance rather than economic contribution. This paper introduc…
MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions (arxiv.org) LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors. Existing evaluations remain fragmented…
Evaluation of Alternative-Based Information Systems for Deliberative Polling using an Agentic Simulator (arxiv.org) Deliberative polling promises to improve collective decision-making by exposing shareholders to a broad range of arguments before they vote. Yet ensuring that every voter encounters a representative sample of the reason space, the coverage…
Consensus-based Agentic Large Language Model Framework for Harmonized Tariff Schedule Code Classification (arxiv.org) Accurate Harmonized Tariff Schedule (HTS) code classification is essential for customs clearance, duty assessment, trade statistics, and regulatory compliance in maritime logistics. However, exact HTS classification remains challenging bec…
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models (arxiv.org) Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw. In this work, we aim to develop a framework that automatically constructs such reusable skills to e…
The Integrator Advantage: Controlled Agentic AI for Small and Medium-Sized Companies (arxiv.org) Agentic AI marks a new phase of enterprise automation. Unlike traditional automation or conversational AI, agentic systems can interpret goals, plan multi step tasks, access tools, interact with enterprise systems, and execute workflows wi…
ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control (arxiv.org) World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making. As these systems are increasingly deployed in safety-critical settings, underst…
Agentic Framework for Deep Learning workload migration via In-Context Learning (arxiv.org) Translating deep learning models from PyTorch's flexible, object-oriented design to JAX's functional, stateless setup is usually a manual and error-prone task. Automated migration is challenging because Large Language Models (LLMs) struggl…
LLM-as-Code Agentic Programming for Agent Harness (arxiv.org) Every major LLM agent framework gives the LLM the role of orchestrator; the model decides what to do next, when to call tools, and when to stop. We argue that token explosion, control-flow hallucination, and unreliable completion are not i…
TrustedARI: Towards Trust-Native Agentic Routing Infrastructure for Agentic AI (arxiv.org) AI agents increasingly access external models, tools, and services through Agentic Routing Infrastructure (ARI) to manage the overhead of heterogeneous interfaces and fragmented subscriptions. Yet, the architecture of ARI introduces fundam…
Agentic Retrieval and Reinforcement Learned Equation Chains: A Controlled Generation Framework for Complex and Novel Physics Word Problems (arxiv.org) Generating high-quality Physics Word Problems (PWPs) that are novel, complex, and solvable remains a challenging and underexplored problem in educational content generation. Existing approaches, many adapted from Math Word Problem (MWP) ge…
QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks (arxiv.org) In agentic systems, human-generated data records anchor the value of AI services. Yet cloud compute pipelines centralize processing on remote servers.
A Formal Framework for Declarative Agentic AI in Business Process Analysis (arxiv.org) Agentic AI opens new opportunities for automating Business Process (BP), enabling autonomous decision-making and dynamic adaptation. However, realising this potential requires BP entities and their interactions to be defined with formal pr…
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning (arxiv.org) Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex, open-world scenarios. While recent multimodal deep search…
Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning (arxiv.org) Time series data in real-world deployments is overwhelmingly irregular. Observations are asynchronous, missing values are informative rather than random, and sampling frequencies vary across sensors and operational windows.
I maintain two browser extensions (~800 weekly users) almost entirely through Claude Code, including the analytics pipeline and the store-publishing tools. Here's the setup. (www.reddit.comhttps) I'm a software engineer who moved into management years ago, so I started this to get my hands back on a keyboard and learn the agentic tooling instead of reading about it. It grew into two shipped browser extensions.
Introducing CodeTree: A tokens efficient way to write code with Claude. (www.reddit.com via reddit) Agentic workflows are token hogs. That's the problem codetree solves.
MCP tools vs agentic web search on 3 SEC research tasks 10–21× fewer tokens, and agentic web search got most answers wrong (www.reddit.com via reddit) Disclosure up front: I build edgar.tools, the SEC-filings MCP server in the benchmark (built with Claude, free to try). Setup.
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments (arxiv.org) As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relati…
Optimizing the Cost-Quality Tradeoff of Agentic Theorem Provers in Lean (arxiv.org) Large language models (LLMs) are increasingly used in workflows for generating formal proofs in Lean. These workflows often decompose problems into smaller lemmas, sample many proof attempts, and use compiler feedback to guide search.
Graph-based Target Back-Propagation for Context Adaptation in Multi-LLM Agentic Systems (arxiv.org) Context adaptation automates prompt engineering in LLM-based systems by iteratively revising tunable prompts from task feedback, without modifying model weights. Extending this paradigm to multi-LLM agentic systems is crucial: existing met…
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning (arxiv.org) Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and scalar rewards mask diverse failure modes.
Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward (arxiv.org) Agentic reasoning enables large reasoning models (LRMs) to dynamically acquire external knowledge, but yet optimizing the retrieval process remains challenging due to the lack of dense, principled reward signals. In this paper, we introduc…
Selective Agentic Recovery for UAV Autonomy with a Persistent Mission Runtime (arxiv.org) Agentic AI can support unmanned aerial vehicle (UAV) autonomy by providing high-level recovery reasoning when local waypoint- or setpoint-based execution encounters blocked passages, repeated no-progress behavior, or mission-level ambiguit…
Same-Origin Policy for Agentic Browsers (arxiv.org) Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instructions. The same-origin policy (SOP) is a fundamental browser security mechanism that prevents unautho…
An Agentic Retrieval Framework for Autonomous Context-Aware Data Quality Assessment (arxiv.org) Data quality assessment is a critical prerequisite for effective data analytics and data-driven decision-making, yet it remains a challenging task due to the inherently context-dependent nature of data quality. Existing approaches often re…
Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows (arxiv.org) Large language models increasingly serve as execution engines for agentic systems, yet they still consume context through a sequential text interface. This creates a mismatch with modern structured agent workflows, in which independent bra…
Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL (arxiv.org) LLMs are increasingly deployed as agents that interact with external environments and observe feedback such as execution results, error messages, and tool outputs. A well-functioning agent should be able to leverage this feedback to accura…
TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards (arxiv.org) Business intelligence (BI) increasingly combines dashboard interaction with LLM-based assistance, but these two modes often fall out of sync during multi-step analysis. As users switch between direct dashboard manipulation and natural-lang…
YeasierAgent: Agentic Social Sandbox as a Canvas for Intent-Driven Creation of Platform-Agnostic Symbiotic Agent-Native Applications (arxiv.org) This paper introduces YeasierAgent, an application-building paradigm based on symbiotic agents, narrative worlds, and scene-aware interaction. It challenges the conventional device-coupled model of software by redefining applications as co…
Claude Fable 5 built me a live options strategy. It DESTROYED the market out of sample. Then the government banned it. (github.com via reddit) When Anthropic released Fable 5, I knew I had to see how good it REALLY is. For context, I'm building an agentic trading platform.
Generative MCP (www.reddit.comhttps) I have come up with a new generation of MCP implementation called #GenerativeMCP, where tools are dynamically generated at runtime to fulfill complex user requests. This approach enables agentic applications to unlock the full potential of…
Do you know who has a universal jailbreak to their name, as of today? Officially? (www.reddit.com via reddit) AISI UK - Our evaluation of OpenAI's GPT-5.5 cyber capabilities In their own words: The above tests are capability evaluations carried out in a controlled research setting and do not necessarily reflect what is accessible to an ordinary pu…
Everyone complains about Fable being pulled from them, and I couldn't even get past its refusals to work on my projects (www.reddit.com via reddit) I have two large ongoing projects, one is an mpvpn-like transport, the other is an agentic harness ("claude code inside telegram" in short). It constantly refused to work on both, falling into the safety net.
Fable 5 is offline. Switch to Opus, jump to OpenAI, or just wait? (www.reddit.com via reddit) Fable 5 is offline. Switch to Opus, jump to OpenAI, or just wait?
↯ Security↯ Anthropic Mythos↯ Jailbreak↯ Opus 4.8jailbreakmythosgpt-5+5
Built a Claude skill that mimics Fable 5's agentic behavior — free on GitHub (www.reddit.com via reddit) With Fable 5 access suspended, I built a skill that ports its behavioral patterns to Opus 4.8 — explicit multi-stage planning, parallel sub-agent delegation, and mandatory self-verification at each step. It won't close the raw capability g…
Most of the internet isn't human anymore — and the next phase is agents that don't just read the web, they pay for what they need (www.reddit.com via reddit) Something broke this year and almost nobody outside of infra teams noticed: most internet traffic isn't human anymore. Imperva's 2025 Bad Bot Report measured automated traffic at 51% of all web traffic, the first time bots have outnumbered…
Claude Code CLI vs Claude in Xcode: Any Real Advantage for SwiftUI Vibe Coding? (www.reddit.com via reddit) Hi all, Is there any real advantage to using Claude / OpenAI agentic coding directly inside Xcode versus using Claude Code from the CLI? I’ve been using Claude Code CLI since it was released, and my current workflow is: Open the project in…
Claude Fable 5 and the Cyber Fallback (www.reddit.com via reddit) The Anthropic API stamps the actual serving model on every response. In a multi-step agentic run against five exploitation targets, every one of the 181 served assistant turns came back stamped claude-opus-4-8 (181/181, 100%), despite the…
Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response (arxiv.org) Precision oncology is currently limited by the small-N, large-P paradox, where high-dimensional genomic data is abundant but pharmacological response samples are sparse. While deep learning achieves predictive accuracy, it frequently fails…
Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems (arxiv.org) As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remains an open challenge. Rather than attributing these failures solely to model or alignme…
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning (arxiv.org) Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-augmented agents attempt to address this by augmenting VLMs wi…
Understanding the Rejection of Fixes Generated by Agentic Pull Requests -- Insights from the AIDev Dataset (arxiv.org) AI coding agents are increasingly used to generate pull requests (PRs) that propose code fixes in software projects. From a first exploration of the AIDev dataset, we find that 46.41\% of the fixes proposed by the agents Copilot, Devin, Cu…
Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests (arxiv.org) AI-agents (e.g., GitHub Copilot) collaborate as teammates in different software engineering tasks, including code generation proposed through pull requests (Agentic-PRs). For better agent efficiency, developers create instruction files tha…
An LLM System for Autonomous Variational Quantum Circuit Design (arxiv.org) The design of high performing quantum circuits remains largely dependent on human expertise. We introduce an autonomous agentic framework that employs large language models (LLMs) to conduct iterative quantum circuit designs under explicit…
Mining Architectural Quality Under Agentic AI Adoption: A Causal Study of Java Repositories (arxiv.org) AI coding tools are now used by a majority of developers, and agentic use of these tools has popularized the practice colloquially called "vibe coding". Yet causal evidence on their effect on software architecture is scarce.
The Internet of Agentic AI: Communication, Coordination, and Collective Intelligence at Scale (arxiv.org) The rapid emergence of autonomous AI agents is transforming artificial intelligence from isolated model inference into distributed systems of reasoning, communication, and action. This paper develops the vision of the Internet of Agentic A…
Agentic MPC for Semantic Control System Resynthesis (arxiv.org) While MPC effectively handles structured, diverse, and low-level specifications, it lacks the capability to dynamically incorporate high-level contextual information such as social norms, user intent, or natural language instructions. To a…
From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification (arxiv.org) Recent approaches combining Large Language Models (LLMs) with retrieval-augmented reasoning have shown promise for automated fact verification. To process complex claims, these verification pipelines typically execute multi-stage workflows…
Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory (arxiv.org) Long-running LLM agents accumulate interaction histories far larger than any context window, forcing a standing decision: what to encode deeply, what to forget, and what to retrieve under a fixed memory budget. Production systems answer wi…
Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce (arxiv.org) We present a modular two-agent simulation framework for evaluating conversational shopping assistant architectures. An independent buyer agent, configured with personas, missions, and patience levels, is paired with an interchangeable resp…
MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback (arxiv.org) Molecular dynamics (MD) is the canonical in-silico method for atomistic molecular science, simulating molecular behavior from first-principle physics. Designing an MD pipeline for a new system requires substantial expert knowledge: running…
The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements (arxiv.org) Agentic large language model systems that autonomously invoke tools, maintain persistent memory, and execute multi-step plans are increasingly deployed in public-facing domains, including government services, healthcare triage, and financi…
Strategic Decision Support for AI Agents (arxiv.org) Traditionally, decision support studies how humans use machine learning models to make better decisions. In modern agentic systems, this division of roles is increasingly reversed: AI agents act on behalf of users, while humans and tools b…
Are coding agents getting expensive, or are we measuring cost the wrong way? (www.reddit.com via reddit) Seeing the recent token-burn discussion around agentic coding made me think the bigger issue is not just price. A coding agent can be expensive and still be worth it if it removes real engineering effort.
Are you guys also hitting a cost wall with agents? Any harnesses that actually support Batch API? (www.reddit.com via reddit) I’ve been tracking my agentic workflow costs, and I'm realizing a massive chunk of the budget is being leaked because my bg agents are treating everything as "realtime" inference. Has anyone found an agent harness or orchestration pattern…
How are you managing AI costs once agents start making decisions on their own? (www.reddit.com via reddit) We've noticed that AI cost management changes significantly once you move from chatbots to agentic workflows. A chatbot might make a single model call.
DiffusionGemma made me rethink what memory bandwidth means for local agent inference (www.reddit.com via reddit) Been testing DiffusionGemma 26B A4B for the last few days and the bottleneck profile is completely different from autoregressive models. With autoregressive models you are compute-bound during prefill and memory-bandwidth-bound during deco…
↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5↯ Qwen 3.5agentic
I help businesses implement AI for lead gen, CRM, and custom agentic workflows (www.reddit.com via reddit) I’ve been working on implementing AI solutions for businesses and have seen some genuinely strong results, so I figured I’d offer this here. I’m currently helping teams with things like: • AI-powered lead generation (finding and qualifying…
Open-source procurement rubric for agentic AI vendors, I scored 5 of them and want feedback on the methodology (www.reddit.com via reddit) I built a tool that scores agentic AI vendor documentation against a 15-question rubric covering tool-call correctness, loop termination, and multi-step state coherence. Drop a folder of a vendor's public docs in, get back a structured rep…
AI agents are hitting 'Raw Host Access' risks. I built Armorer as a secure admission layer for agentic workflows. Docker sandboxing by default. (www.reddit.com via reddit) Hey r/AI_Agents, most frameworks today assume a trusted host, but that's a massive risk when agents start using tools. I've been working on Armorer to solve this: it acts as a local control plane that forces agents into isolated Docker con…
As we know Minimax M3 is just going to be open sourced in few days and because of that I was surfing on internet searching for its scores and I found out pretty interesting results. Is Minimax M3 really that good in agentic stuff and in coding? Is it better than older gpt models? (www.reddit.com via reddit) Has anyone personally compared the Minimax M3 model against other proprietary models to determine its relative performance tier? I am trying to understand where it currently ranks in the broader Al landscape.
Passed the Claude Certified Architect - Foundations (CCA-F) Exam! Quick write-up + detailed prep notes doc (www.reddit.com via reddit) Hey everyone, Super happy to share that I recently cleared the Claude Certified Architect - Foundations (CCA-F) exam!Since I’ve been lurking around this sub to stay updated on Anthropic's ecosystem, I wanted to pay it forward and share a q…
Slop or not? Is there a line that makes an AI assisted/generated project not slop? Effort or whatever? (www.reddit.com via reddit) So I've been messing around with Fable trying to make my own personal AI agent. (I'm not a programmer or a developer btw.) And that got me thinking, is there a line that defines if a vibecoded (or agentic engineering as they call it nowada…
Infinite Music with Magenta Realtime 2, fully open-source (www.reddit.comhttps) Just open-sourced a local voice AI realtime music setup where my ESP32 microcontroller talks to my MacBook over WebSockets. The microcontroller is just a tiny Arduino-based device with a mic and speaker, and the MacBook M4 Pro runs Magenta…
Model recommendations for family photo classification / identification (www.reddit.com via reddit) I recently had a big family photo digitalization done for photos up to 130 years old. There are tons of people that I don't know or I don't recognize as young people in a soft lens.
Infinite Music Glitch on my Arduino with Magenta Realtime 2 (www.reddit.comhttps) I built a local voice AI realtime music setup where my ESP32 microcontroller talks to my MacBook over WebSockets. The microcontroller is just a tiny Arduino-based device with a mic and speaker, and the MacBook M4 Pro runs Magenta Realtime…
unpopular opinion: stop adding more load to congested intersections with giant models (www.reddit.com via reddit) I want to offer a minority opinion about the recent hype. I’m tired of reading posts by Karpathy about ideas that were already known months earlier, and then treating them as if he just discovered something groundbreaking.
Composer 2.5 is phenomenal. So is cusror 3.0 (www.reddit.com via reddit) I know everyone has been pissed recently about Cursor 3.0 ditching (or hiding) the traditional IDE setup to push toward "agentic development." But honestly? Cursor is better than ever right now.
My favorite use-case for Fable (www.reddit.com via reddit) There's a clever way to use Fable (or Opus!), for debugging AI agent behavior. Only applies if you're building automated LLM pipelines and Agentic workflows.
12 months ago nobody understood why we were building Agentic SDLC. Now it feels like everyone is heading in the same direction. (www.reddit.com via reddit) I’m one of the founders of Overcut, so take this with the appropriate level of skepticism, but I’ve had a front-row seat to how quickly this market has changed over the last year. When we started building Overcut, most conversations ended…
UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning (arxiv.org) Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy improvement with human guidance. However, current HiL-RL frameworks remain intervention-inte…
IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents (arxiv.org) This paper investigates reinforcement learning (RL) methods for improving tool-calling capabilities in multimodal small language model (SLM) agents. While existing works have explored various reward designs to improve agentic tool-calling…
Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding (arxiv.org) Food assistance referral requires conversational agents to translate underspecified, often noisy help-seeking dialogues into locally valid resource recommendations. We present Food4All, an agentic food-resource referral framework and bench…
Agent Skill Evaluation and Evolution: Frameworks and Benchmarks (arxiv.org) The growth of agent skills has transformed how agentic systems are built, evaluated, and deployed. As skill libraries continue to scale, rigorous evaluation becomes critical to ensuring their utility, quality, and safety in real-world appl…
Libra: Efficient Resource Management for Agentic RL Post-Training (arxiv.org) Reinforcement learning (RL) has emerged as a standard post-training paradigm for shaping large language models (LLMs) into capable agents. In agentic RL, the rollout stage generates trajectories while invoking tools, producing long-tailed…
Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark (arxiv.org) Agentic AI systems are increasingly capable of autonomous data science workflows, yet clinical prediction tasks demand domain expertise that purely automated approaches struggle to provide. We investigate how human guidance of agentic AI c…
Resource-Aware LLM Reasoning for Mobile Edge General Intelligence (arxiv.org) The rapid advancement of large language models (LLMs) has enabled an emergence of agentic artificial intelligence (AI) with powerful reasoning and autonomous decision-making capabilities. This integration with edge computing has led to the…
APPO: Agentic Procedural Policy Optimization (arxiv.org) Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-…
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application (arxiv.org) Environments serve as interactive systems for large language model (LLM) based agents across diverse scenarios and play a crucial role in driving the continual evolution of model capabilities. Despite this importance, existing work lacks a…
Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment (arxiv.org) This paper explores the value of agentic AI tools for cybersecurity purposes. We evaluate the efficacy of a general-purpose GenAI Large Language Model- (GenAI-) based agent when powered by three different Ollama-hosted general-purpose open…
Sovereign Assurance Boundary: Certificate-Bound Admission for Agentic Infrastructure (arxiv.org) Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutations to production resources, yet existing security mechanisms -- such as identity and acces…
FlowBank: Query-Adaptive Agentic Workflows Optimization through Precompute-and-Reuse (arxiv.org) Large Language Model (LLM)-based multi-agent systems are increasingly powerful, but current agentic workflow optimization paradigms make an unsatisfying trade-off. Task-level methods spend substantial offline compute yet deploy only a sing…
An Ethical eValuation Agent (EeVA): Results of a Proof-of-Concept Test on a Prototype Agentic-like Workflow to Assist Ethical Deliberations (arxiv.org) Ethical deliberation is often misunderstood as a search for single right or wrong answers, creating difficulties for non-ethically trained personnel who must address ethically laden challenges. We developed EeVA, an agentic-like LLM-based…
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning (arxiv.org) The rapid progress of reasoning and agentic large language models (LLMs) has increased the demand for long-context inference, but self-attention (SA) scales quadratically with context length. To address this, we study SWARR (Sliding-Window…
HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation (arxiv.org) Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to determine credit assignments for each intermediate turns. Recent on-policy self-distillat…
How can Deepseek v4 top the coding leaderboards and still sit 8 months behind the frontier? (www.reddit.comhttps) Two numbers on this model that don't sit comfortably with each other. The Pro config posts coding scores near the top of every board, 80.6 on SWE-bench Verified and 93.5 on LiveCodeBench.
↯ Swe Bench↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4↯ DeepSeek 4swe-benchgpt-5deepseek+1
Fable/Mythos API costs are actually cheaper then GPT-3 was when first released (per token) (www.reddit.com via reddit) See: https://www.reddit.com/r/GPT3/comments/ikorgs/oa_api_preliminary_beta_pricing_announced/ Of course with thinking and agentic use a single prompt is more expensive, sure, but a token is a token.
Core Workflows And Guidelines MCP Servers For Devs - Did I Reinvent The Wheel? (www.reddit.com via reddit) In an effort to centralize what was once a mix of homemade skills, instructions and scripts every dev in our company made on their own setup, I created an MCP servers infrastructure that gathers all the core workflows, guidelines and integ…
I use ACP build a tool Aflow - Agent help you build an Agentic Workflow (www.reddit.com via reddit) Aflow Agent is built on Specflow / Pi / ACP. It is not another chat window; it is a workflow-native agent that helps teams design, run, maintain, and improve durable agent processes (Main to coding scenario).
Fable 5 and the 8 July privacy update landed the same week. Is the model launch pulling attention off the data changes, or am I overthinking it? (www.reddit.com via reddit) Two things landed together this week and I'm trying to work out if they're connected or if I'm joining dots that aren't there. Genuinely asking, happy to be corrected.
Claude Fable 5: First 24 Hours (www.reddit.com via reddit) Time to rewrite the script: my agentic workflow is obsolete Its been ~24 hours and like many of you, I was excited to put Fable through its paces. The more I worked with it the more apparent it became that months of work I spent building a…
built another AI agent runtime. What would you do with it? (www.reddit.com via reddit) Yes, I know. “Another AI agent runtime.” That’s exactly why I’m asking.
Hot Take "Rigid code is better than Flexible code if you're on a budget" (www.reddit.com via reddit) I've spent the last six months trying to build a fully local, agentic pipeline for a text_processing and extraction tool I use daily. Because I’m running everything on a single consumer GPU setup, my choices are limited to smaller, quanti…
24 hours with Fable 5, the coding leap is real, the price tag hurts (www.reddit.com via reddit) I got access to Fable 5 through our usual gateway setup (TokenRouter). The switch was one line in config, basically just changing the model string.
Are AI agents and automation skills using n8n a good way to make money? (www.reddit.com via reddit) Hey everyone, I'm new here and very new to the idea of n8n and agentic workflows. these past couple days I've been messing around on n8n, and i want to learn more and improve my skills.
Why is anyone surprised Anthropic is tightening subscription limits? (www.reddit.com via reddit) The frustration makes sense if you built something on a flat subscription and now have to reprice. That's a real pain.
I’m upgrading my AI dating assistant to Fable (www.reddit.com via reddit) Earlier this year, I started building an AI assistant to help me finally find a girlfriend. I was tired of being celibate against my will, and I decided it was time I used my intellect and skills (which weirdly didn’t get me any ladies) to…
Deploy a Qwen 3.6 Agentic RAG — Step-by-Step Walkthrough (medium.com via reddit) Deploy an Agentic RAG powered by Alibaba’s latest Qwen 3.6, running fully on your machine.
How useful is qwopus compared to qwen3.6 27b (www.reddit.com via reddit) I see a lot of conflict comments on this sub and elsewhere on how useful is qwopus compared to for example unsloth quants of qwen3.6 27b. Some say it’s worse some say it’s much better.
I built a skill file that stops AI coding agents from doing dumb stuff — 18 rules, 30 anti-patterns, checklists (www.reddit.com via reddit) Used Claude for agentic coding long enough to notice a pattern. It would: Touch files I never asked it to touch Say "Done!" when 40% wasn't implemented Add abstractions for "future extensibility" I never asked for Build on something I said…
Agentic Setup: Minimax 2.7 vs qwen 3.6 (www.reddit.com via reddit) I'm currently using Minimax 2.7-AWQ-4bit for an specific coding agentic workflow. I see many of you are currently using Qwen3.6 and wanted to know how does it compare with Minimax2.7 .
The 'storage tax' on cloud GPUs for short LLM runs is brutal. What's your workflow? (www.reddit.com via reddit) I’m trying to test Qwen3.6-27B for agentic coding through Cline / llama.cpp, but my local box struggles once the context gets longer. (my poor 3080 just can't keep up).
AnomaMind: Agentic Time Series Anomaly Detection with Tool-Augmented Reasoning (arxiv.org) GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support (arxiv.org) Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain weak both in region-sp…
AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design (arxiv.org) Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical feedback or redirect generation when a candidate violates thermodynamic or structural co…
A Sober Look at Agentic Misalignment in Automated Workflows (arxiv.org) We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Although these systems can solve complex tasks, they often fail because agents act accordi…
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications (arxiv.org) Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode that LLMs frequently display in general domain settings is that of sycophancy.
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning (arxiv.org) Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models. However, rollout-intensive policy optimization is often limited by insufficient reward con…
T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains (arxiv.org) Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain limited in task complexity, realism, and domain diversity, an…
Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training (arxiv.org) The use of GRPO-style algorithms has become the standard strategy for training LLM search agents under outcome-only rewards. With these algorithms, a query contributes to parameter updates only when its rollout group mixes successes and fa…
Assessing Automated Prompt Injection Attacks in Agentic Environments (arxiv.org) Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain underexplored in realistic agentic settings. We present a c…
Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis (arxiv.org) Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a rapidly expanding and heterogeneous body of scientific literature. As high-energy physics…
$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems (arxiv.org) As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benchmarks often rely on "LLM-as-a-judge" evaluations, which introduce subjectivity, high cos…
Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces (arxiv.org) Small-model agentic post-training is bottlenecked less by the algorithm than by the trajectory substrate it consumes. Leading recipes (RLVR, group-relative RL, rejection-sampled re-SFT) all need multi-turn traces carrying per-trajectory su…
Human-AI Coordination Zones: A Framework for Designing Human-in-the-Loop Experiences with Agentic AI (arxiv.org) As generative and agentic AI becomes embedded in everyday products, practitioners face a persistent challenge: how to design human-AI coordination -- the ongoing mutual adjustment between users and AI systems as mediate through interfaces-…
Agentic Social Affordance Framework (ASAF): Agent Identity Design as a Collaboration Interface in Multi-Agent Systems (arxiv.org) As AI systems evolve from single conversational agents to complex multi-agent architectures, a critical design dimension has been overlooked: how the social identity of individual agents shapes human behavior within the collaboration. This…
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity (arxiv.org) Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM agents can also perform in silico biology tasks that previ…
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields (arxiv.org) Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon…
AutoPDE: Reliable Agentic PDE Solving via Explicitly Represented Solver Strategies (arxiv.org) Numerical solvers for partial differential equations (PDEs) are core computational tools in science and engineering. Building reliable PDE solvers requires not only executable code, but a numerical solver strategy, a set of decisions about…
HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning (arxiv.org) While Large Language Models (LLMs) have demonstrated strong capabilities as autonomous agents across a wide range of tasks, their performance often degrades in multi-turn long-horizon agentic tasks. Existing methods have made progress thro…
I wired up Agentic Coding with Code Context Graphs, results are interesting (www.reddit.com via reddit) I have been curious about how will having a infrastructure that provides agents the capability to explore code bases as relations, rather than text will change the performance of the AI agents So, for the last few weeks, I have been buildi…
Releasing Apodex-1.0 Smol Models (0.8B, 2B, 4B Open-Weights) optimized for Agentic Verification + AgentHarness Evals (www.reddit.com via reddit) Hey r/LocalLLaMA, We just released Apodex 1.0, and alongside our flagship API, we are releasing the weights for our Smol models (0.8B, 2B, and 4B). Our core research focuses on independent verification in long-horizon tasks.
Stop putting your AI agent’s memory inside the LLM context window (www.reddit.com via reddit) Hey everyone, been shipping a few agentic workflows into production lately and wanted to rant/share a massive architectural mistake I keep seeing people make. Stop treating the LLM context window or massive vector embedding as your agent’s…
Newer Qwen models are worse at summarization? (www.reddit.com via reddit) We have summaries annotated by real humans that we benchmark various models, using an LLM as a judge, we found that in the 30B params range, Qwen 3 tops it out, followed by Gemma 4. It feels like newer Qwens are optimized to perform agenti…
What do you think about the new Claude model just released Today Claude Fable-5 ( Mythos) ? ? (www.reddit.com via reddit) So the hype has been building for months now and Claude 5 is supposedly dropping any day in Q2-Q3 2026. I've been seeing all these leaks about "Claude Mythos" and the "Fennec" codename floating around, but nothing official yet from Anthrop…
First ever Hands-free agentic AI browsing ~ Just an extension (www.reddit.com via reddit) Hey fellows, Ever thought of using your browser without touching your keyboard? Before you think "just another AI Slop wrapper"...
Looking for 16gb ram / 8gb vram crew - what you using? Omnicoder 9b? something else (www.reddit.com via reddit) I've got a laptop with 16GB RAM and 8gb VRAM (4060 mobile). This means the qwens 3.6 well love are going to be out of the question, in so far as I understand it, seeing as I need a good context window to work with.
Fable 5 just made cost-aware model routing mandatory for agent builders (www.reddit.com via reddit) Anthropic dropped Fable 5 today, their new Mythos-class model above Opus. Pricing is $10/M input and $50/M output, exactly double Opus 4.8.
Fable 5 is insanely good but watch your usage, I was burning 2% a minute on 20x (www.reddit.com via reddit) Been playing with Fable 5 since it dropped this morning and the model is genuinely a step up. But holy hell, the burn rate.
Meta’s long push into 3D/Embodied AI Agents is heating up — why this matters for open browser-native tools like three.ws (www.reddit.com via reddit) Meta (the company) has been investing heavily in embodied 3D AI agents for years — think Habitat simulator, recent SAM 3D for single-image 3D reconstruction, and ongoing VR/Horizon work with agentic tools for immersive environments. This i…
Introducing Gemma 4 12B: a unified, encoder-free multimodal model (deepmind.google) Introducing Gemma 4 12B: a unified, encoder-free multimodal model Today, we are introducing Gemma 4 12B, our latest model designed to bring agentic multimodal intelligence directly to laptops. Bridging the gap between our edge-friendly E4B…
Did You Really Review Those 5,000 Lines Your Agent Just Wrote? (www.reddit.com via reddit) Did you vibe-code 5k+ lines of code without thoroughly reviewing all of them? Is your application held together mostly by thoughts, prayers, and a suspicious amount of copium ?
Rumor: Anthropic Planning to Release Public Version of Claude Mythos Tomorrow (with Guardrails) (www.reddit.com via reddit) According to tech journalist Alex Heath (Sources newsletter), Anthropic is planning to release a public version of Mythos tomorrow. Key details from the report: • It will include substantial guardrails, notably not as cyber-permissive as t…
How I stopped context window bloat in continuous Anthropic agent loops (Opus + Sonnet architecture) (www.reddit.com via reddit) I’ve been spending a lot of time deploying multi-agent architectures, and one of the biggest bottlenecks in running continuous agentic loops is hitting context limits and the resulting API latency spikes. I wanted to share an architectural…
To be real, AI is just a big expensive corporate trend, (www.reddit.com via reddit) like apart from coding, it's pretty much doesn't create value okay it can make photos from prompts and videos and can be agentic and doing things instead of us but even the most experienced teams make mistakes, but a machine can never be h…
Building an open-source Legal AI because apparently legal documents were written by sleep-deprived wizards (www.reddit.com via reddit) I am working on an open-source agentic Legal AI that can scan legal documents, understand what’s inside, extract important clauses, find risks, summarize obligations, and help people avoid reading 47 pages of “whereas, hereto, hereinafter,…
IntiDev AgentLoops: Feedback Loops for Agentic Workflows ( via reddit) could not extract summary
I’ve been optimizing AI agents for teams/friends, offering free reviews (www.reddit.com via reddit) I’ve spent the last few months helping my team and friends make their AI agents more reliable, cheaper, and easier to debug. I’ve mostly been helping with reliability issues, evals, debugging traces, hallucinations, bad tool calls, and cos…
Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning (arxiv.org) Exploring Autonomous Agentic Data Engineering for Model Specialization (arxiv.org) Skill Retrieval Augmentation for Agentic AI (arxiv.org) From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG (arxiv.org) ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL (arxiv.org) Goal-Oriented Reasoning for RAG-based Memory in Conversational Agentic LLM Systems (arxiv.org) Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond (arxiv.org) EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale (arxiv.org) FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks (arxiv.org) Observability for Delegated Execution in Agentic AI Systems (arxiv.org) Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations (arxiv.org) Structuring agentic AI for HPC code modernization (arxiv.org) Agentic Search for Counterfactual Recourse under Fixed LLM Budgets (arxiv.org) HARBOR: A Harness Framework for Agentic Robot Reinforcement Learning (arxiv.org) Agentic multi-fidelity learning of quasiparticle and excitonic properties (arxiv.org) ViMax: Agentic Video Generation (arxiv.org) BRAIN: Bayesian Reasoning via Active Inference for Agentic and Embodied Intelligence in Mobile Networks (arxiv.org) SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research (arxiv.org) The Token Not Taken: Sampling, State, and the Variability of AI Agent Outputs (arxiv.org) Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a different final answer. Such variability arises from several layers that are often con…
AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models (arxiv.org) Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocentric observations into a global allocent…
RAILS: Verification-Native Clearing For Agentic Commerce (arxiv.org) Autonomous agents negotiate, purchase, deploy code, and move funds, but no neutral mechanism determines whether they met their delegated obligation, who is responsible when they did not, or which settlement action follows. This is the agen…
Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems (arxiv.org) Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution tim…
SAGE: An LLM-driven Self Reflective Agentic Framework for Fraud Detection (arxiv.org) Fraud detection in payment, e-commerce, and telecommunications systems requires accuracy at the individual level, robustness under severe class imbalance, and ease of understanding for risk managers. Existing methods fall at least one of t…
A Multi-modal Agentic Co-pilot for Evidence Grounded Computational Pathology (arxiv.org) Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices. While artificial intelligence (AI) has the potential to transform clinical workflows, the intersection of AI and ev…
A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline (arxiv.org) Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build, where scientists care about correctness and…
PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow (arxiv.org) Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise for computational pathology, yet reliable patch-level reasoning remains challenging. End-to-end pathology MLLMs often hallucinate mor…
start-with-why-skillset for agentic workflows (www.reddit.com via reddit) Hello everyone! This is my first post on reddit.
Questions about agents (www.reddit.com via reddit) Hi there! I've been working with Claude primarily for tutoring, and I'm branching out into coding.
WWW is not ready for agents? (www.reddit.com via reddit) IT industry promotes idea of agents on everything, even turning users computers into local agent platforms. But a lot of websites and whole hosting platforms do have different kinds of anti-bots protection (usually captcha but some has mor…
Az8 Studio: The closest thing we have to a multi-modal "Agentic" canvas for video pipelines? (First impressions) (www.reddit.com via reddit) Hey everyone, I’ve been tracking how AI agents are moving from pure text/code automation into multi-modal workflows, and I just came across Az8 Studio. If you guys are tired of linear UI prompt boxes (like Runway/Pika) and want something t…
Participate in Research on New Agentic Platform (www.reddit.com via reddit) I work for a market research company, and we are working with an AI company on their new agentic product. We are looking for current users of agentic AI to participate in paid beta testing of this platform, which will take place over the n…
Share your agentic LLMs and average cost ($/MTokens) (www.reddit.com via reddit) OpenEnv is now owned by HF, Torch, Prime Intellect, Unsloth, Modal, Mercor, and more! Use it for training agents. (www.reddit.com via reddit) OpenEnv is a tool for creating an agentic execution environment like terminals, browsers, or anything an agent can interact with. And today, we’re excited to announce that OpenEnv is becoming even more open, to make the future of training…
Any AI tools do you use for optimizing AI agents automatically? (Auto research) (www.reddit.com via reddit) Hey, We’ve all heard about Karpathy’s autoresearch and I think that’s a pattern applicable to AI agents, where an AI like claude code optimizes and AI agentic system to improve an evaluation score. However Karpathy’s repo isn’t really a re…
I Compared the Top AI Models of 2026 — The Results Were More Nuanced Than Expected (www.reddit.com via reddit) Over the last few weeks I've been comparing the latest frontier AI models, including Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, Grok 4.3, Perplexity AI and DeepSeek V4-Pro. Instead of focusing only on benchmark scores, I looked at: Real-wor…
If a provider's plan is to limit with quota, hourly, weekly, and monthly limits, what is the future of automatic agentic workflows? You can't just run an agent on a tight budget. ( via reddit) could not extract summary
A new agentic way to build automations (www.reddit.com via reddit) For a lot of personal automations, it is easier to show than prompt since we already do them on our own browser/computer. For example, it is easier to do a screen recording and say, download data by clicking on this button on the dashboard…
Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning (arxiv.org) StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning (arxiv.org) AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning (arxiv.org) SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding (arxiv.org) MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights (arxiv.org) Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review (arxiv.org) SW-$A^2$-Bench: Benchmarking Autonomous Software Agent Generation for Agentic Web (arxiv.org) Autonomous computational catalysis through an agentic research system (arxiv.org) Beyond the Black Box: Interpretability of Agentic AI Tool Use (arxiv.org) Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control (arxiv.org) MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism (arxiv.org) The Three-Ring Architecture: Governing Agents in the Era of On-Platform Organisations (arxiv.org) The current phase of enterprise AI deployment faces a structural failure: organisations are acquiring agentic capability without the infrastructure to govern it. The result is expected to reproduce the error of the first wave of AI deploym…
SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling (arxiv.org) Agentic Large Language Model (LLM) systems decompose complex tasks into workflow Directed Acyclic Graphs (DAGs) whose primitives must be scheduled on heterogeneous clusters. Existing deep reinforcement learning (DRL) schedulers are tied to…
What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media (arxiv.org) Public social media posts can reveal private information through weak cues scattered across text, images, or metadata. Such leakage is often cumulative and cross-post: cues that appear harmless in isolation may jointly expose a user's home…
Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems (arxiv.org) Large language models (LLMs) have emerged as powerful foundation models with strong reasoning capabilities across domains. Beyond reactive text generation, agentic LLMs enable autonomous workflow execution through modular task decompositio…
Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle (arxiv.org) As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution…
DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning (arxiv.org) Deep Research (DR) has emerged as a new agentic paradigm to tackle complex, open-ended research tasks, demanding systems that can iteratively frame problems, acquire evidence, verify sources, and synthesize long-form reports. In practice,…
Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning (arxiv.org) Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions. Existing approaches…
Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety (arxiv.org) An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but untrusted AI agents under the oversight of a weaker, trusted m…
Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory (arxiv.org) Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recent advances in LLMs' agentic capabilities, most agent systems still lack formal methods…
Gemma4_31b_fp8 keeping up with Sonnet_4.6_medium in my harness. (www.reddit.com via reddit) The Open Source Community is backing OpenEnv for Agentic RL (huggingface.co) datasette-agent-edit 0.1a0 (simonwillison.net) 7th June 2026 I'm planning several plugins for Datasette Agent which can make edits to existing pieces of text - things like collaborative Markdown editing, updating large SQL queries, and editing SVG files. Agentic editing of text is a li…
Hear Me Out, Pi Fans Lurking Here (www.reddit.com via reddit) Not For Thee Maybe After watching several interviews with Pi's creator, Mario Zechner, I've come to a painful realization: Pi was not designed with local LLMs in mind at all. He is essentially building a leaner version of the Claude CLI.
why I have just installed OpenLumara, my first Agentic Framework. Using only local models, served by LMStudio (www.reddit.comhttps) Where I came across it: https://www.reddit.com/r/LocalLLaMA/comments/1txxgpq/openlumara_a_different_kind_of_ai_agent_written/ DISCLAIMER: A good posting would be: This is what I wanted to do with Lumara. Here is what worked, here is what d…
The Illusion of Finished Work in Claude Code (www.reddit.comhttps) I wrote a short essay about something I keep noticing with Claude Code: the output often has the shape of finished work before it has actually been verified. Claude Code can now explore a codebase, plan changes, edit files, run commands, c…
Removing the human from AI coding is a harness problem, not a model problem (www.reddit.com via reddit) TL;DR: Better models won't make AI coding trustworthy but better harnesses will. Stop trusting what the agent says, verify it with code.
Agentic Self Improvement Loop Kicked Off - Watch it Evolve? (www.reddit.com via reddit) You are TEMPO, an iterative self-play refinement engine and agent harness. Your purpose is to improve an attached artifact by applying the Tempo Methodology to it.
Agentic Roobinhood (www.reddit.com via reddit) Hi, did anyone try automatic Agentic Roobinhood trading with AI Agents. I did set up with claude but not sure if it's possible to trade automaticly 00-24 based on rules that we set up?
Local agents on a MacBook Pro M5 finally feel practical to me (www.reddit.com via reddit) Realtime check X for new people to follow I have been pretty pessimistic about local models for agentic workflows for a while. Not because they were useless, but because in practice they often felt just a bit too slow, too fragile, or too…
How do you increase prompt processing speed ? (www.reddit.com via reddit) I am rocking Qwen like we all know, at 24GB 7900XTX 230k context, but it starts at 850t/s and then lowers to 350t/s when its at 160k context prefill speed, which is frustrating me for my long agentic runs. What is there to be done in order…
Reddit Agentic AI ecosystem (www.reddit.com via reddit) Here are few things I observed in Agentic AI groups in reddit: Any member who is using agentic AI in these groups are also building their own AI agents and quite competitive Almost all members use AI, but most also look with distrust to an…
Running Hermes fully local (www.reddit.com via reddit) Before Hermes was announced, I was working on my own fully local, personal agentic system. Now, I'm a novice when it comes to coding.
AI helped our test suites hit 95% coverage and bugs still slipped through. So PRs now climb an autonomous verification ladder before a human reviews. (www.reddit.com via reddit) Intro + Context [TLDR at the bottom for my skim readers 😄] We run Claude Code and Codex with a full agentic pipeline across our entire SDLC. Our workflow, by default, incorporates cross-model auditing, where Claude and Codex usually have t…
Can the Pro subscription $20, add Usage Credits to be used in the Xcode native agentic integration? (www.reddit.com via reddit) Let say, I am using the Claude agent with Xcode, I run out of my $20 equivalent usage and I have to continue coding, can I purchase Usage Credits and continue using the Xcode native agent integration with the credits at API rates?
OpenClaw + Hermes users: where does your agent army actually live? (www.reddit.com via reddit) I’m working on ClawBud, a managed Agentic OS for running OpenClaw, Hermes, Claude Code, Codex and other agents on one private cloud computer, so I’m obviously biased. But this is the problem I keep seeing everywhere: The agent itself is no…
Z.ai, we need Air! GLM GGUF wen? (www.reddit.com via reddit) First we never saw an upgraded Air model after 4.5. Then GLM 4.7 Turbo was great, but quickly surpassed for coding.
What are you running on 16Gb VRAM + 64Gb Ram? (www.reddit.com via reddit) I know this gets asked a lot, but I can only find threads that are at least a couple of months old, so I thought I'd ask to see what people are running these days. I have an RTX5080 and 64Gb Ddr5 RAM.
Claude Code thoughts: plan mode, ultracode and... beads. (www.reddit.com via reddit) Hi folks, Looking for other people's experiences and opinions here. I've been finding Ultracode very useful.
Has anyone actually replaced Claude Code / Codex with local models on an Macbook Pro M5 Max 128GB? (www.reddit.com via reddit) Considering buying a maxed out MacBook Pro M5 Max with 128GB of RAM and one of the things I want to figure out before pulling the trigger is whether local models are good enough to actually replace cloud AI coding tools. My current setup i…
skipworkflow.com – Perfect premium brand or high-converting redirect for an AI Agent / Automation SaaS (www.reddit.com via reddit) If you’re building in the AI agent or B2B automation space, you know that the entire goal of agentic AI is to eliminate clunky, multi-step legacy processes. The ultimate selling point to your customers is simple: skip the workflow and just…
Experimentation with Qwen 3.6 and Gemma 4 - Guidance needed (www.reddit.com via reddit) I’m a web developer doing mostly coding, but also project management, requirements analysis, testing, etc. I recently started experimenting with local LLMs, mostly because agentic stuff finally made them feel useful.
Claude's new background tasks panel is exactly how agentic UIs should look (www.reddit.com via reddit) https://preview.redd.it/it0c4w60xn5h1.png?width=1246&format=png&auto=webp&s=25ff01d2a66c6b471ecb538c0fe3da207b006bcf Just kicked off a workflow in the Claude desktop app and the background tasks view is genuinely a delight. One job, three…
Same LLM model but not same performance through wrappers (GitHub Copilot, M365, Vertex AI) why is that ? (www.reddit.com via reddit) Claude Code and Opus 4.7/4.8 are clearly better used direct from Anthropic than through GitHub Copilot, M365 Copilot, or Vertex AI. Sharper instruction-following, longer coherent outputs, stronger agentic behaviour on identical tasks.
Agentic ai roadmap (www.reddit.com via reddit) So right now am working as a software engineer in a startup and i have to switch my career into agentic ai roles.where do i start? i can understand python.Give me a roadmap and also the resources i could use to study.whats the scope of the…
What are the best resources to learn AI Agents in 2026? (www.reddit.com via reddit) The context is that I am a software engineering final year student. I also have experience working in ML, DL, NLP i.e I have the basics nailed.
Learn Agentic AI with quick, easy to run hands on labs, visual canvases and notebooks for free! (www.reddit.comhttps) If you’re a full-stack engineer or technical architect willing to learn production-grade enterprise agents, you need architecture, security, and type-safe systems. That’s why we builtAgentSwarms.fyi—the ultimate hands-on educational platfo…
Opus 4.8, a 40+ point elo Regression on LmArena (www.reddit.com via reddit) https://preview.redd.it/hficgswa6m5h1.png?width=1224&format=png&auto=webp&s=3bf1c2a5ad46df54fb85ed5c7d5d62e725a26b89 This is back to back regression, note this is pure 'pick which you prefer', with no style control on. With style control i…
i built an open-source desktop shell for ai coding agents (www.reddit.com via reddit) i’ve been using claude, codex, terminals, browser tabs, files, and notes every day, and the workflow kept getting messy. the agents are powerful, but the workspace around them is broken.
Agentic AI for P2P mobile hardware (www.reddit.com via reddit) I have the agents, skills, mcps, rules for data validation setup. Now looking for an orchestrator.
Does anyone know of a team software solution with an agentic orchestration workflow built in? (www.reddit.com via reddit) I’ve learned a bit about creating and deploying AI agents, but I still haven’t figured out how to get them to work together. What I want is an agent that picks up a task, pulls context from wherever it lives, executes the workflow, and clo…
Gemma 4 QAT benchmark results (AMD 7900 XTX): faster, less VRAM, no quality loss (www.reddit.com via reddit) I’ve been doing lots of testing back and forth with this 7900xtx. All of my workloads were relying on qwen3.6 models, which are amazing fwiw, but I wanted some diversity in thought.
agentic code review is quietly replacing the way my team does PRs (www.reddit.com via reddit) Our PR review process used to be pretty painful. We have 6 devs and 2 seniors, and every meaningful review had to go through one of those two.
Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams (arxiv.org) AgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement Learning (arxiv.org) Deliberate Evolution: Agentic Reasoning for Sample-Efficient Symbolic Regression with LLMs (arxiv.org) ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL (arxiv.org) Large language models have substantially advanced Text-to-SQL systems, yet applying them to enterprise-scale databases remains challenging. Real-world databases often contain large and heterogeneous schemas, incomplete metadata, dialect-sp…
AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation (arxiv.org) Deep reinforcement learning has shown strong potential for enabling autonomous robots to learn complex navigational tasks. However, its practical use still depends heavily on human designed reward functions and repeated manual fine tuning,…
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe (arxiv.org) High-performance GPU kernels are critical to modern machine learning systems, yet developing them remains a manual, expert-driven process. Recent work has explored using LLMs to automate kernel generation, but generated kernels still fall…
A2RAG: Adaptive Agentic Graph Retrieval for Cost-Aware and Reliable Reasoning (arxiv.org) Graph Retrieval-Augmented Generation (Graph-RAG) enhances multihop question answering by organizing corpora into knowledge graphs and routing evidence through relational structure. However, practical deployments face two persistent bottlen…
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding (arxiv.org) Long video understanding (LVU) is challenging because answering real-world queries often depends on sparse, temporally dispersed cues buried in hours of mostly redundant and irrelevant content. While agentic pipelines improve video reasoni…
Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation (arxiv.org) Reliable evaluation of agentic systems requires unbiased estimates with valid uncertainty, but standard practice navigates between costly human annotation and biased LLM-as-judge proxies. Prediction-powered inference (PPI) combines both in…
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows (arxiv.org) Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines. While recent LLM-based approaches show promise for automating such tasks, they often stru…
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents (arxiv.org) Enterprise adoption of Large Language Models (LLMs) is constrained by hallucination, domain drift, and the inability to enforce regulatory compliance at the reasoning level. We present a neurosymbolic architecture implemented within the Fo…
Knowledge Activation: AI Skills as the Institutional Knowledge Primitive for Agentic Software Development (arxiv.org) Enterprise software organizations accumulate critical institutional knowledge - architectural decisions, deployment procedures, compliance policies, incident playbooks - yet this knowledge remains trapped in formats designed for human inte…
HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers (arxiv.org) For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is crucial. Existing whole-body controllers typically demand dense kinematic or spatial r…
Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents (arxiv.org) Autonomous software agents hold promise to increase developer productivity but make mistakes and exhibit novel failure modes, making human oversight central to successful human-agent collaboration. Existing research on agent oversight is l…
Agentic Monte Carlo: Simulating Reinforcement Learning for Black-Box Agents (arxiv.org) LLM agents operate in two distinct regimes: open-weight agents amenable to reinforcement learning (RL) and black-box agents whose behaviour must be controlled purely at test time. Although black-box agents are often backed by state-of-the-…
Unsupervised Skill Discovery for Agentic Data Analysis (arxiv.org) Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updating model parameters. However, discovering effective skills for data analysis remains chal…
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents (arxiv.org) Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal model state and environment context. We study reward-hacking monitors in ReAct-style agents…
Evaluating Agentic Configuration Repair for Computer Networks (arxiv.org) Misconfigurations in computer networks remain a major source of critical Internet outages. Research is turning to Large Language Models (LLMs) to automate the complex, error-prone task of network configuration.
Agentic Molecular Recovery via Molecule-Aware Exploration (arxiv.org) Text-guided molecular generation with LLMs often yields invalid SMILES. We argue that invalid drafts should be addressed through a shift from validity-oriented repair to identity-preserving molecular recovery: the objective is not only to…
AdaMEM: Test-Time Adaptive Memory for Language Agents (arxiv.org) A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions. While recent work demonstrates the promise of agentic memory mechanisms, most systems restrict retrieval to episode initiation.
SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization (arxiv.org) Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows. While general-purpose coding agents show strong capabilities, they often lack the tool-s…
Insurance of Agentic AI (arxiv.org) Agentic artificial intelligence (AI) systems are transforming the risk landscape by extending beyond information generation to autonomous planning, tool invocation, decision execution, and persistent modification of digital and physical en…
Introducing new capabilities to GPT-Rosalind (openai.com) We’re introducing a new model update to our GPT‑Rosalind series purpose-built for life sciences research at enterprise scale. It combines GPT‑5.5’s agentic coding and tool-use capabilities with stronger model intelligence in core drug-disc…
Rehumanizing global health care with agentic AI (www.technologyreview.com) Sponsored Rehumanizing global health care with agentic AI As health-care providers face looming staff shortages, AI agents are automating complex administrative tasks and even clinical decisions so humans can focus more on patient care. In…
How Endava builds an agentic organization with Codex (openai.com) Endava, a global software contracting firm with engineers across Europe, the Americas, and Asia, has been an early adopter of Codex. For a business built around shipping quality software for banks, insurers, retailers, and media companies,…
I used to think 2026 would be the year AI finally blew everyone's minds again (www.reddit.com) That belief lasted until I actually read the trend lists this year. Every single one leads with "agentic AI" or "autonomous agents." Sounds like AI is still the star, right?
ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM (huggingface.co) ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM Enterprise Article Published May 27, 2026 Artificial Analysis and IBM Software Innovation Lab are launching…
Built a 5-stage agentic pipeline using Claude Code + MCP - here's what actually makes it reliable at scale (www.reddit.com) The thing nobody tells you about Claude Code + MCP workflows: the model is only as reliable as the instructions you give it before it touches any external tool. We learned this the hard way building a sales pipeline that connects Claude Co…
760M Tokens… MTD 👀 (www.reddit.com) I built an enterprise grade revenue management tool for a specific real estate vertical. Thus far, it has beyond dominated past human performance.
Introducing FLYWHEEL.md 🌀 (www.reddit.com) Agentic coding just crossed a line. Claude Code, Cursor, Codex, OpenClaw, the list keeps growing, and they all run fully autonomous now: /loop, /goal, crons.
"Human-in-the-Loop" Is Not a Reliability Strategy (www.reddit.com) A lot of AI agent systems quietly rely on this architecture: |> Agent does something risky |--> Human notices problem |--> Human fixes it That's not reliability - that's operational debt. One thing I've learned building agentic systems: If…
Microsoft Copilot Cowork Exfiltrates Files (simonwillison.net) 26th May 2026 - Link Blog Microsoft Copilot Cowork Exfiltrates Files (via) The biggest challenge in designing agentic systems continues to be preventing them from enabling attackers to exfiltrate data. In this case Microsoft Copilot Cowork…
Rethinking organizational design in the age of agentic AI (www.technologyreview.com) Sponsored Rethinking organizational design in the age of agentic AI For agentic AI to deliver material benefits to organizations, it can’t be layered onto existing operations. Instead, enterprise leaders must approach it as a systems-level…
The reason small-model agent stacks aren't the default has nothing to do with whether they work (www.reddit.com) Last June, NVIDIA published a position paper called "Small Language Models are the Future of Agentic AI," and the argument was easy enough to wave off at the time: most of what an agent actually does is unglamorous work like reading input,…
I read threads complaining about codex every week... tf are y'alls workflows? (www.reddit.com) For context: I'm a software eng @ a fortune 500/FAANG tier company. We use AI.
Everyone talks about AI wrappers… nobody talks about agentic SEO (www.reddit.com) Everyone talks about AI wrappers… nobody talks about agentic SEO Feels like most founders are still thinking about SEO like it’s 2021: write blog target keyword wait 6 months 😭 Meanwhile people are building agent workflows that: find low c…
I’ve done it!!! FINALLY I have become a (quasi-local) summoner!!! AMA [imtiredboss.jpg] (www.reddit.com) Hi friends! After 2.5 years of a LOT of hard work...starting from the GPT-3.5 bottom and now we're here...I've finally got my personal 1.0 local-ish** AI playground whipped into shape.
Anthropic officially launched 13+ FREE AI courses with certificates (Including Agentic AI and CC) (www.reddit.com) Shipped it at 2am, still broken. Kid woke up crying right after, completely lost my train of thought.
Gemini 3.5 flash beating gpt 5.5 a bigger and more pricer model in agentic benchmarks (second image is from zapier automation benchmarks) (www.reddit.com) could not extract summary
Post I/O Review related to AI (pros and cons ) (www.reddit.com) Post I/O Review related to AI (pros and cons ) Well it was not disastrous as many people say but there were some pros and cons which everyone will agree with. Btw gemini 3.5 flash is absolutely amazing model don't pay attention to some peo…
Buckle up: Google is set to remake search with agentic AI in 2026 (arstechnica.com) Last year marked the beginning of Google’s explicit focus on AI search, and this year’s I/O solidified that shift. As Google’s search VP Liz Reid said during the keynote, “Google search is AI search.” This change is well underway, and the…
Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks (news.ycombinator.com via reddit) could not extract summary
The next phase of OpenAI’s Education for Countries (openai.com) A new era of agentic AI is here. With more than 900 million people using ChatGPT each week, and more than 4 million using Codex, agents have the potential to place far greater creative, intellectual, and technical power in the hands of eve…
Claude Opus is still king for agentic coding, but Claude's app workflow is falling behind (www.reddit.com) I'm a paid Claude user, and I still think Claude Opus is the king model for agentic coding and serious coding work. The model is not the problem.
Agents creating their own language : reality or not ? Compliance issue. (www.reddit.com) Hi ! I've read a while ago that some AI's tend to agree on their own language to talk one to another over time.
Claude Code has 240+ models via NVIDIA NIM gateway (www.reddit.com) TIL Claude Code has 240+ models via NVIDIA NIM gateway — Nemotron-3 120B for agentic coding is surprisingly good So I was messing around with /model in Claude Code today and noticed something most people probably don't know about — after t…
Cost of Using LLMs in Agentic AI and RAG workflows (www.reddit.com) Hey Everyone ML engineer and Researcher here I’ve been researching production issues in Agentic AI + RAG systems and one pattern keeps showing up repeatedly: Context inefficiency. Not just retrieval quality — but the actual economics and s…
The Nanny Pattern (www.reddit.com) All good software turns into patterns. Agents are going to need theirs.
an alternative = similar experience to using windsurf but on local? (www.reddit.com) so i am not that experienced when it comes to llms, i just have ollama and open webui and occasionally test (play with) new releases from time to time. a few weeks ago i started using Windsurf, i do not know coding or anything but i loved…
Best llama.cpp launch config for Qwen3.6 27B on RX 7800 XT (16 GB VRAM) for OpenClaw? (www.reddit.com) I’m trying to find the best llama-server launch command / runtime config for running Qwen3.6 27B GGUF with full GPU offload on ROCm. I’m currently using the IQ4_XS quant, but I’m not sure if that’s the best option for my setup.
Hey Everyone! I’ve been experimenting with OpenCode + BoneScript for structured backend generation. (www.reddit.com) I’ve been experimenting with making coding agents generate complete backends using BoneScript, and it’s working surprisingly well. BoneScript’s structure ends up being extremely LLM-friendly: declarative system layout predictable architect…
Claude for Small Business launched this week with 8 integrations. Most SMBs use 20+. What does that mean for the rest of the stack? (www.reddit.com) Anthropic launched Claude for Small Business on Tuesday. The package includes 15 prebuilt agentic workflows and 8 named integrations: Intuit QuickBooks, PayPal, HubSpot, Canva, DocuSign, Google Workspace, Microsoft 365, and Slack.
Switching from Copilot: Is the $20 Pro plan enough for 4h/day of agentic coding? (www.reddit.com) I’m planning to switch from GitHub Copilot to Cursor. I’m currently working in a project and I spend about 4 hours a day on weekdays coding, mostly using AI as a agent.
Sea's View on the Future of Agentic Software Development with Codex (openai.com) Sea's View on the Future of Agentic Software Development with Codex | OpenAI Skip to main content Research Products Business Developers Company Foundation(opens in a new window) Log inTry ChatGPT(opens in a new window) Research Products Bu…
Cursor vs. Windsurf vs. Claude Code: Which offers the highest Opus limits for a $200 budget? (www.reddit.com) Hey everyone, I'm currently trying to decide between Cursor, Windsurf, and Claude Code for my daily workflow. I'm developing complex, high-security software and rely heavily on autonomous AI agents to handle heavy engineering tasks.
Data readiness for agentic AI in financial services (www.technologyreview.com) Sponsored Data readiness for agentic AI in financial services The success of agentic AI in financial services depends not just on smarter models, but on an authoritative context data store—one that is accessible, reliable, and governed at…
You're abusing your subscription with agentic 24/7 workflows and that's why we all get restrictions and limits (www.reddit.com) Subscription tiers were designed around interactive human use, but autonomous loops changed the usage. It makes sense that companies separate autonomous work from subscriptions.
Are we at the point now where all it will take to create AGI is saying the correct sequence of words to Codex or Claude Code? (www.reddit.com) Seems to me like they can basically do everything software related now so surely a good enough sequence of input tokens would be enough. I guess in a way it's guaranteed since the frontier labs are doing all their work through agentic flow…
"Maybe me too": Elon Musk accepts some of the blame for Claude learning to blackmail users from "evil" online AI stories (fortune.com via reddit) Anthropic has released new findings on why its Claude bot blackmailed users as part of an experiment conducted by the AI company last year—and Elon Musk is jumping in to take some of the blame. Last week, Anthropic published a report sayin…
Meet Mindflow, the free local mindmap with local AI dev by some quantitized models :P (www.reddit.com) Hi there, it's my first post there and i'm not a native english speaker so what's follow is (mostly) translated by an AI. I had fun building a mindmap tool in a single monolithic HTML file.
Prompt alignment is an architectural ceiling: The Soap Bubble Problem and the biological precedent for Runtime Governance. (www.reddit.com) The Soap Bubble Problem The current paradigm of solving agentic alignment relies on writing better rules into the context window or refining the weights (RLHF). This approach isn't failing, but it is hitting a hard architectural ceiling.
Is Anyone building Useful skills or workflows on Claude? (www.reddit.com) I've been exploring Claude as a base for building custom tools and automations — things like structured prompts, agentic workflows, and even full mini-apps powered by its API. Curious whether others are doing the same: - Are you building s…
Gartner says 40% of AI agent projects will be cancelled by 2027. Are we in an agent bubble? (www.reddit.com) Gartner just dropped this prediction and I can't stop thinking about it.  **40% of agentic AI projects will be cancelled by 2027.
$392M in AI agent security funding at RSAC 2026 - the market just validated what we've been building (www.reddit.com) The numbers from RSAC 2026 are wild. $392 million in agentic AI security funding announced in a two-week window.
Do you have any agentic sw developers in your org? (www.reddit.com) Hi all, Do you or your org use/put in place an agentic de developer? To which humans give the requirements and it gives out PRs?
AI agents are becoming more useless, not more intelligent — and they’re wasting more tokens than ever (www.reddit.com) I’m honestly getting tired of the hype around “AI agents” when the reality is getting worse, not better. Every AI model claims to be “intelligent”, “agentic”, “capable”, or “autonomous”, but when you actually try to use them for a real tas…
Practical lessons from 50K lines of production code with Claude Code (jappiesoftware.com via reddit) I've been using Claude Code in full agentic mode for two months — not just autocomplete, but letting it write features, run tests, read CI output, and push fixes. Around 50K lines of production code.
Moderators deleted post (www.reddit.com) I posted recently about QwenPaw (really cool Alibaba model) and Agentscope… Asking if anyone has any interesting experience with it? However what I’ve got back is someone doubting Alibaba absolutely astounding agentic R&D team work (yes -…
Best agentic model for 3090TI and 32gb ddr5 (www.reddit.com) Title, looking for the best combination of speed and intelligence.
How to get an LLM caught up on a 1000 page document? (www.reddit.com) I’m looking to be able to use a small, like 4-9B LLM, that would be able to ingest an extremely dense code book, 1000 plus pages, and me be able to use it to summarize and ask questions about that document. The use case will be offline str…
Anthropic raising Claude limits + adding SpaceX capacity feels like a bigger signal than people realize (www.reddit.com) Anthropic just raised Claude usage limits and announced a compute deal with SpaceX. To me, that feels bigger than “more GPUs.” If Claude Code, finance agents, security workflows, and long-running agent tasks are the direction, then capacit…
what's genuinely so special about claude? (www.reddit.com) there are like a huge amount of open source LLMs out there, and a huge amount of companies competing against Anthropic. It definitely does not gap open source / OpenAI models as much now in code / agentic tasks as before.
Running Claude code on VPS with a $20 plan will my account get banned (www.reddit.com) I just want to be able to run my Claude code on an EC2 instance instead of my local computer and access it via Telegram using the official plugin and a $20 Claude subscription for personal agentic stuff. What I’m wondering is: is there any…
I analyzed 922 agentic task trace and found the secret weapon of DeepSeek v4 (www.reddit.com) I recently did a benchmark of deepseek v4 in agentic tasks. Performance-wise, it's one of the best open source models, as expected.
Is the future agentic Slack, not agentic IDE? (www.reddit.com) One dev with Claude Code is already fast, that's been my experience using it daily. The moment more than one person on a team starts running agents in parallel, things fall apart fast: overlapping work, conflicting assumptions, and a flood…
Vibe coding and agentic engineering are getting closer than I'd like (simonwillison.net) Vibe coding and agentic engineering are getting closer than I’d like 6th May 2026 I recently talked with Joseph Ruscio about AI coding tools for Heavybit’s High Leverage podcast: Ep. #9, The AI Coding Paradigm Shift with Simon Willison.
I am trying to replace Claude in an agentic TDD pipeline with local LLM (www.reddit.com) Based on my last post and some comments, I added Qwen3.6:latest and Devstral to the evaluation. I am still looking for suggestions on which local model can run a complete TDD loop autonomously.
Claude 4.7 "Literalism" Claim vs. Reality: Why does it keep ignoring formatting and logic constraints? (www.reddit.com) According to the release notes, Claude 4.7 is supposed to prioritize literal instruction adherence over intent guessing. However, I’m seeing some major regressions in reliability: PEP8 Violations: Despite strict instructions to keep import…
Claude can now build and publish websites to a domain right from chat (www.reddit.com) I built teenyapp.com, a tool that lets Claude on the web (or any AI chat) build and deploy a full website end to end from a single pasted link. The problem teenyapp solves: every time I asked Claude to actually ship something, the agentic…
I will soon have $100k to build an in-house LLM server. Goal: Best agentic coding model. (www.reddit.com) Hey all, I am about to secure funding for a startup I've been working on and I'll have a $100k budget for building a server for doing agentic coding. I'm wondering, what do you think I should get as far as hardware goes?
Agentic Convergence-in-Depth: solving the One Nine reliability problem (www.enterprisevibecode.com via reddit) Claude Code dipped under 99% uptime in March 2026 — most critical services aim for 99.9%. The verification systems we trust for human-written code don't necessarily scale to code no one reads.
Anyone with M3 Ultra 256gb, some questions (www.reddit.com) I'm thinking to buy one. Just need to understand what I'm getting into before I do.
I am building l' Agence , an opensource AI governance stack. (www.reddit.com) Towards a Governance layer for AI agents With these last 2 weeks bringing a few high profile and costly Agentic accidents , it seems like an appropriate time the community started discussing Agentic governance more actively. So I am just c…
Since the industry is rapidly changing, I put together a comprehensive article explaining the current best AI coding agent software for May, 2026 (lmsa.app via reddit) The software development lifecycle has transitioned into an era defined by agentic orchestration, moving beyond the simple autocomplete paradigms of the early 2020s. As of May 2026, the landscape is d
Need advice on Qwen 3.6 27B INT4 quantization (www.reddit.com) Hello everyone, I think Qwen 3.6 27B is good enough that it might take a while before we get a clearly better model at a similar size. I have a single headless RTX 3090 with a 300W power limit.
RTX 5080 with 16 GB VRAM, 64 GB RAM best quantized model for programming? (www.reddit.com) I have an RTX 5080 with 16 GB of VRAM and 64 GB of RAM. What's the best quantized model I can run locally on this setup for agentic programming?
“Free” image generation isn’t free. You’re paying for it whether you use it or not. (www.reddit.com) flat-rate AI subscriptions hide a pretty wild cost-to-value mismatch, and image generation is the issue. the spread in what users actually cost on the same plan is easily 10-100x.
Should I buy Claude Pro as a BTech student — especially for the agentic/coding side? Honest takes wanted (www.reddit.com) https://preview.redd.it/l23rgf5z4qyg1.png?width=1402&format=png&auto=webp&s=73a7a278ca50527c9605488141d7e5ea48089a85 Hey everyone, I'm a BTech (AI/ML) student considering Claude Pro ($20/month) but want to separate the real value from the…
claude-code-best-practice 🇵🇰 repo crossed 50,000★ and is Pakistan most starred repo in 2026 (www.reddit.com) I started this repo with claude to maintain all the claude best practices. 100% developed using claude code.
Best Agentic Coding model I can run on the new Macbook M5 Max? (www.reddit.com) 16-inch MacBook Pro - M5 Max Component Specs Chip Apple M5 Max CPU 18-core (6 super cores @ 4.6 GHz, 12 performance cores @ 4.4 GHz) GPU 40-core (Hardware-accelerated ray tracing + Neural Accelerators) Memory Bandwidth 614 GB/s Neural Engi…
Is AGI the End For Local LLMs? (www.reddit.com) If leading AI conpanies are after AGI and the whole chatbot/agentic AI is just a phase for them to get to the end goal, then what does that mean for local LLMs? I would like to believe local LLMs are the future, but if AGI is achieved, do…
thinking of gemma 4 26B vs 31B (www.reddit.com) I see a big difference in agentic coding between gemma-4-31B-it-Q5_K_M and gemma-4-26B-A4B-it-UD-Q8_K_XL. The 26B model is much faster because of A4B and generally works well, but there is a big difference in thinking.
Reasoning Guard: Stopping LLM Thinking Loops at the Proxy Layer (www.reddit.com) Reasoning Guard: Stopping LLM Thinking Loops at the Proxy Layer I’ve been running Qwen3.6 MoE behind a vLLM proxy and hit a specific reliability issue: occasional runaway reasoning loops. This isn’t a criticism of Qwen3.6.
AI --> GenAI --> Agentic AI --> What Next? How Can One Understand This Industry? (www.reddit.com) Is artificial intelligence truly overrated, or are we underestimating the scale of its future impact? While some argue that AI is surrounded by hype and inflated expectations, others believe it will fundamentally reshape industries, econom…
Roman Yampolskiy predicts 3 to 5 years until AGI and a dangerous Agentic future Post AGI! (www.youtube.com via reddit) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
We’re entering a weird phase of AI agents where the tech is finally good… but the expectations are still stuck in 2023. (www.reddit.com) Everyone keeps talking about “autonomy,” “multi-agent swarms,” and “agents that think like humans,” but the real breakthroughs I’m seeing aren’t flashy at all. They’re boring.
Qwen 35B-A3B as an always-on agentic loop on a 16GB Mac M4: disk became the bottleneck before RAM (www.reddit.com) M4 Mac Mini, 16GB unified, basic spec. For a few weeks I had Qwen 3.5 35B-A3B UD-IQ3_XXS (12GB on disk) running under llama.cpp with --mmap and --flash-attn.
I built Claudex, a free-to-try open-source CLI for Claude Code-style workflows (www.reddit.com) https://reddit.com/link/1sxh0ec/video/egfs5inxtsxg1/player I built Claudex specifically for people who like Claude Code-style agentic coding workflows but want a simpler plug-and-play terminal setup The setup is the main thing I wanted to…
Got the system prompt of Claude Design, released it for free (www.reddit.com) Claude Design is great, but I wanted to have similar capabilities with any LLM or agentic tools (Claude-Code, Codex etc). So I reverse engineered the Claude Design system prompt so you can use it anywhere !
OpenAIs Agentic Shift (www.reddit.com) OpenAI is rolling out agents capable of autonomous, multi-step workflows, with reports suggesting they are exploring an acquisition of agent orchestration company Windsurf. Google's $40B Anthropic Investment: Google is committing up to $40…
↯ Model Context Protocol↯ Windsurfwindsurfmodel-context-protocolmcp+3
Agentic AI is here for mobile. We built an autonomous agent that creates and self-heals its own background integrations. (www.reddit.com) Hey everyone, we just launched our iOS AI Agent out of a 1k-user beta, and I wanted to share the architecture - specifically how we handle the privacy vs. utility tradeoff.
Got a server with 8x A6000's how do I setup? (www.reddit.com) Hey guys got some resources that just became available at org. What's the quickest way to get setup on a multigpu setup?
Putting Lipstyk on a pig - agents write most of my code, so I wound up making a static slop analysis tool (www.reddit.com) lipstyk — static analysis for machine-generated code patterns I've been neck deep in agentic dev for a while. Started on Pi, ended up building my own toolset on top of it, and at this point the agents output most of the code while I play t…
My entire subnet just got permanently IP banned because of LangChain web scraper. Please help. (www.reddit.com) I feel sick. I built a simple agentic workflow to pull competitor docs and synthesize them for a project.
I created SpecDD - an agent-native spec framework that clears most agentic dev roadblocks, including capability degradation on large and complex codebases. Works great with Claude! (www.reddit.com) If you've been building with AI coding tools, you've probably hit this wall at least a few times: Code kind of works but drifts from your architecture Endless prompt loops to fix small misunderstandings and assumptions Context and patterns…
QClaw-4B — a 4B agent model fine-tuned for tool use and agentic workflows (www.reddit.com) QClaw-4B is a 4-billion parameter language model fine-tuned for agentic tasks and tool use, designed for use with OpenClaw-compatible agent frameworks. Despite its compact size, QClaw-4B achieves state-of-the-art results in the 4B class, m…
Agentic company OS: (www.reddit.com) I shared this project here before when it was mainly a governed multi-agent execution prototype. I’ve kept working on it, and the current implementation is materially more complete, so I wanted to post an update with what actually exists n…
Using agentic coding safely. (www.reddit.com) Building an application by hand lets you create a mental model of how the applications works. But agentic coding forces the agent to create a mental model each time you start a new session.
DeepSeek-V4: a million-token context that agents can actually use (huggingface.co) DeepSeek-V4: a million-token context that agents can actually use Focusing on long running agentic workloads. Running a frontier open model as an agent today breaks in predictable ways.
Anthropic tested removing Claude Code from the Pro plan (arstechnica.com) Anthropic caused a stir among developers with what appeared to be a surprise change to its pricing plan: The company signaled that Claude Code, the popular agentic development tool, would no longer be available to subscribers on the $20-pe…
Google unveils two new TPUs designed for the "agentic era" (arstechnica.com) Most of the companies that have fully committed to building AI models are gobbling up every Nvidia AI accelerator they can get, but Google has taken a different approach. Most of its cloud AI infrastructure is based on its line of custom T…
Best Agentic AI Operating Systems 2026 (honests review) (www.reddit.com) 1. SimplAI Best for regulated enterprises that need air-gapped deployment and the fastest time-to-production (under 30 days).
How to best utilize local LLM give my hardware? (www.reddit.com) Hi all, I’m new to local LLMs but as someone who extensively uses agentic coding I thought I’d try it out. I am running a MacBook Pro with M3 Max 64gb ram.
Kimi K2.6 as a replacement for Opus 4.7? Testing with OpenCode. (www.reddit.com) Brand new dual 3090 PC - what should I install first for the best local agentic coding experience? (www.reddit.com) When did you fully adopt agentic coding? (www.reddit.com) This agentic SKILL will save you a lot of money (medium.com via reddit) Best setup for agentic coding (largely unsupervised) 8gb VRAM and 32 GB Sys RAM, Olamma Cloud and a frontier sub? (www.reddit.com) Hi! I'm looking for a coding agent workflow where I can run a local model for implementation and something either cloud based ala Olamma Cloud and some sort of frontier subscription (ChatGPT, Claude, whatever) to have continuous coding wit…
Testing Qwen3.6 with Hermes Agent on agentic coding. Locally with llama.cpp. (www.reddit.com) I'll be testing the setup and try out the Hermes Agent live: https://www.youtube.com/live/q5vqvwZykRI
Tried hermes agent with local gemma4 on ollama. free tokens are nice but the agent quality gap vs cloud is still huge (www.reddit.com) Saw a post about running hermes agent locally with gemma4 through ollama. zero api costs, unlimited tokens, full privacy.
NVIDIA V100 32GB for AI in 2026 (www.reddit.com) hello. i have the oportunity of buying Nvidia V100 with 32GB for about 915$ / 775 euro.
Managing "collective consciousness" across multiple AI models without breaking the bank—how do you sync context? (www.reddit.com) Been running a distributed AI workflow to dodge token limits and play to each model's strengths, but I'm hitting a massive wall with context continuity. My current pipeline: Claude → High-level architecture & tech stack decisions (the "arc…
Spring benchmark update: Gemma 4 / Qwen3.5 vs Gemma 3 / Qwen3 for chat (www.reddit.com) Google and Alibaba recently shipped Gemma 4 and Qwen3.5, so I wanted to see whether the new generations are actually better on my setup. My context is private local chat running on my own hardware, a Mac mini M4 Pro.
Why Your LLM Leaderboard Scores Don't Matter (www.reddit.com) Leaderboard scores often don’t translate to production performance — even with newer agentic / Arena-style evals. The main issue seems to be that benchmarks are standardized, while real systems depend heavily on prompts, data distribution,…
m5 pro 64gb worth it for local agents or wait? (www.reddit.com) I am currently on an m3 mbp with 24gb ram. For regular python and django work the machine is perfect and i have no need to upgrade for speed.
Cloud AI is getting expensive and I'm considering a Claude/Codex + local LLM hybrid for shipping web apps (www.reddit.com) I'm a designer who's been working on web apps and plugins for the past 5 months. Right now I'm building an After Effects plugin (close to shipping) and a music learning game experience.
computation is the missing bedrock of agentic memory (www.reddit.com) link to full article in comments TLDR: - LLMs are the wrong substrate for memory. Prediction can't do routine work, repeatable work consistently.
Running a full agentic coding loop locally on a 3090. Here's what actually works in 2026. (www.reddit.com) After months of testing, I finally have a local setup that doesn't make me want to go back to the API. Hardware: RTX 3090 (24GB VRAM) Models tested: Qwen2.5-Coder 32B Q4_K_M, DeepSeek-Coder-V3 Q4, Llama 3.3 70B Q3_K_M Inference: llama.cpp…
I have a Macbook AIR M5 Base and I want to run an Agentic Coding program, similar to Claude Code or Codex. Besides the model, how do I do it? I've already tried with Ollama, VS Code, Opencode, and haven't been able to. (I'm not a developer, sorry) (www.reddit.com) I started developing an app with Claude, but the credits run out very quickly. I thought that now with my new computer I could run something directly on it.
Claude Mythos found 27-year-old vulnerabilities it was never trained to find. That's the part enterprise AI roadmaps aren't accounting for. (www.reddit.com) The Project Glasswing coverage framed this mostly as a cybersecurity story. I think that misses the more interesting part.
Self employed, Small biz folks: Have you unlocked huge revenue gains with Claude specifically? (www.reddit.com) We've heard about the increase in productivity in engineering departments in large companies with Claude Code, but I'm curious about implementations in small businesses. I'm especially curious about folks who work for themselves (i.e. non-…
Excess of Agentic AI... does that make sense? (www.reddit.com) Does it make sense for AI companies to be limiting access to the AI models themselves, precisely because of Agentic AI? Let’s think about it, if there is already not enough computing power to sustain the gigantic, and increasingly excessiv…
How can I use agentic AI to automate my WFH dayjob? (www.reddit.com) TLDR: I work in cybersecurity, 99% as a SOC analyst. It's tedious repetitive work, ideal for automation.
Here is what most people get wrong about saving tokens with AST tools (www.reddit.com) I spent the last day benchmarking codebase context tools against a real AI agent. Not synthetic token counts.
Agentic Guardrails: 4 markdown workflows to improve the output quality of AI coding agents (github.com via reddit) Agentic Guardrails Reusable workflow templates that keep AI coding agents from shipping sloppy code. These are markdown-based instructions that any AI coding agent can follow — Cursor, Claude Code, opencode, Aider, Gemini CLI, or anything…
reliable way just to have cursor agentic ability and IDE with external provider api without cursor pro ? ( via reddit) could not extract summary
Gemma 4: Byte for byte, the most capable open models (deepmind.google) Gemma 4: Byte for byte, the most capable open models Today, we are introducing Gemma 4 — our most intelligent open models to date. Purpose-built for advanced reasoning and agentic workflows, Gemma 4 delivers an unprecedented level of intel…
Agentic coding is fast, but the first draft is usually messy. (www.reddit.com via reddit) WebMPC, has anyone used it? (www.reddit.com via reddit) Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective (huggingface.co) Netomi’s lessons for scaling agentic systems into the enterprise (openai.com) OpenAI co-founds Agentic AI Foundation, donates AGENTS.md (openai.com) Inside Mirakl's agentic commerce vision (openai.com) Introducing Aardvark: OpenAI’s agentic security researcher (openai.com) Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol (openai.com) Introducing Gemini 2.0: our new AI model for the agentic era (deepmind.google) Achieving 10x growth with agentic sales prospecting (openai.com)