event

Ollama

427 items · started 2026-03-23 · ongoing (last activity 2026-09-19)

  1. I built AIRUNCODE because I was frustrated by two patterns in current AI coding environments: cloud vendor lock-in that wipes project memory when switching providers, and platforms adding heavy markups on top of user API keys.AIRUNCODE is…

  2. I have become very proficient with Claude Code and I use it for everything, both at work and in my personal life. At this point, if I lost access to it, it would be like missing a body part.

  3. polygo Lokalise for one person. A 4 MB CLI binary that translates your app's strings with a local model, tracks what changed in a lockfile, and refuses to write a translation with a broken placeholder or a missing plural form.

  4. Benchmarking Local LLM Servers: llama.cpp, llamafile, LM Studio, and Ollama Benchmarking Local LLM Servers evaluates llama.cpp, llamafile, LM Studio, and Ollama across Mac, Linux, and Steam Deck. The study reveals build flags and configura…

  5. I think the path Grok Bot took is absolutely the right one, the collaborating bots model is easy to understand and it maps to workflows naturally. However, Grok Bot is closed source and quite expensive (I ran out of tokens on Ultra tier),…

  6. Ollama Shodan indexes over 47.000 exposed Ollama instances, including many running on expensive cloud GPUs. The API has no authentication, which means hackers can run prompts, steal models and exploit vulnerable versions.

  7. I built echodot – a free, open-source desktop app that drafts replies in your personal writing style. Problem: Most AI writing tools either force you into their voice or require constant copy-pasting between apps.

  8. I got tired of the lock-in and black-box nature of AI inside MS Office. One vendor decides which models you can use, where your document data goes, and what the agent is actually allowed to do.

  9. Ollama 0.33.3 changed what prompt_eval_duration measures If you compute prefill throughput from Ollama's metrics the way everyone has always computed it, your numbers have been wrong since 0.33.3 — and wrong in the flattering direction. On…

  10. Motivations: Maybe you’re a Claude code/codex user diligently avoiding uploading personal data to LLM providers. Is it possible that the most valuable information isn’t your data- but the metadata about your sessions?

  11. Hi everyone, Been working on Otis, an open-source ai agent that gives you one minimal experience across local and hosted open-weight models, privacy-focused by design. On setup it recommends a local model based on the hardware Otis is runn…

  12. TL;DR Three facts and one command. Bench — Mac mini M4 · 24 GB unified memory · macOS 26.5 · Ollama 0.34.0 · gemma4:26b Q4_K_M (MoE, 25.2B total / ~4B active) · measured 2026-09-13 ollama ps )sudo sysctl iogpu.wired_limit_mb=20480 # raise…

  13. raggy A lightweight CLI tool for Retrieval-Augmented Generation (RAG) over local documents built with LangChain, Chroma, and Ollama. Hybrid database (vector + BM25 index) and embedding generation run fully locally.

  14. Running Opencode with Ollama on mac. How to get started with running Ollama local models with Opencode and Docker Sandboxes.

  15. Local-first autonomous WinUI 3 IDE powered by Ollama, ConPTY self-repair loops, and local fine-tuning.

  16. So thoughts? I was playing with Muse before it dropped into Cursor for sometime, thought it was light years better then ollama especially at coding, which was instantly noticed.

  17. Big Claude user, but Anthropic got stingy as hell with the limits. I used to barely touch my weekly allowance; now I can burn through 20% in a day and I'm cooked in ~2 days.

  18. Run the unmodified Claude Code on Anthropic, OpenAI, xAI, Ollama, or any OpenAI-compatible endpoint. Every token priced.

  19. llmash An Ollama-compatible server and command line for Windows, built on llama.cpp. It serves your GGUF files through llama-server and keeps Ollama's commands, API and model store, so anything already pointed at Ollama keeps working.

  20. MaskShift is a local-first coding agent harness. Features: - Works with models that have no tool-calling API.

  21. bobbin A small coding agent for small local models. A dependency-free agent runtime for local models via Ollama.

  22. About two weeks ago I started an experiment: instead of prompting a local model to play a character, give it a folder and let it become one. The idea is simple.

  23. Hey HN, most agent systems default to server-side Python inside containers and chain frameworks. I wanted to see how far we could push agent in the browser with vanilla JavaScript https://buttercup.sh The reason this is interesting is beca…

  24. Repo - https://github.com/aaditya-v-more/claude-ollama Problems fixed - No 1 Million context if using claude code desktop app extension api throwing errors and claude giving up Subagents in long workflows like /deep-research and /batch kep…

  25. Pricing Free For getting started with open models Includes: - Run models locally - Starter usage credits included - Includes access to starter models - Add credits to unlock all models - No service fees Pro Max For shorter, well-defined da…

  26. Anthropic released Fable 5.1 today, everyone started blasting with it on High effort or higher for everything they should really be using Sonnet for. My entire Twitter/X feed is purely people complaining that they used their weekly usage w…

  27. Ollama's transparent pricing August 31, 2026 Ollama’s Pro, Max, and Team plans now use transparent per-token pricing. Based on your feedback, every plan includes a monthly pool of usage credits.

  28. Eighteen months of unsecured Ollama servers, reconstructed from the commit history of three scanners that publish what they find along with hundreds of mysterious servers that aren't as they appear. Thousands of real machines have Ollama l…

  29. As we are lucky to have great open weight models these days, I was wondering what is the recommended way to run these models locally, while staying with open source software that doesn't track you etc? The only tool I've really tested is O…

  30. I always run so many AI agents that I keep ending up with mystery processes, stray localhost ports, and no quick way to tell what started what. So I built Port Visualizer, a free open-source Windows app that shows which process owns each T…

  31. Hey everyone, TL;DR: Built a git-native orchestrator for my coding-agent backlog to safely run parallel agents overnight. It isolates every ticket in its own git worktree so they can't step on each other, and it pools local models (Ollama…

  32. been building a tool that takes a plain-english file chore and turns it into a graph of python steps. you basically tell it "grab the photos from this folder, fix the timezone, sort by date" and it wires the steps up for you.

  33. 📧 EmailAI An AI-powered email agent that syncs your emails via IMAP and provides intelligent analysis using OpenAI, Anthropic, or local Ollama. ✨ Features 🔄 IMAP Sync - Mirror your email account locally with batch syncing 👁️ IDLE Support -…

  34. I am using 20x plan. I am seeing a lot of hallucination and scope drift, and extremely slow execution.

  35. Claude Desktop support with Ollama August 25, 2026 Developers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider. Use open models in Claude Use cloud models for larger coding and resear…

  36. I have a Linux server with an RTX 5060 Ti 16 GB. I am running Ollama as a Docker container with the NVIDIA GPU mounted into it.

  37. LM studio and bionic don't load into GPU fully ( Ollama does) and it crashes BSD ( Ollama Does not), with stop code: WHEA_UNCORRECTEABLE_ERROR (0x124), i am using the default load setting, all updated LM studio and drivers, What do i need…

  38. Landing Page & Web Simulator • Quick Install • Commands • How It Works • License Quick Install macOS & Linux (Bash / Zsh) curl -fsSL https://raw.githubusercontent.com/Luizhcrs/powerai/main/install.sh | bash Windows (PowerShell 5.1 / 7+) ir…

  39. 🍳 yeschef A kitchen for Claude Code and Codex. Local models on your own hardware (your line of cooks) that take the grunt work, talk it out in bounded rooms, and never hit a rate limit.

  40. v4.5.0: Added native desktop control, agent execution budgets, concurrency limits, loop protection, and safer local embedding fallback. v4.6.0: Rebuilt agents around durable parent-led orchestration.

  41. Locally deployed Large Language Models (LLMs) via inference engines such as Ollama run without the moderation and abuse detection present in API-served models. Therefore, the safety of LLMs depends on the defense mechanisms used, and their…

  42. tl;dr - split one VM (with 3 pooled GPUs) into two VMs (1 dedicated + 2 pooled); worth every second spent spooling up the extra VM. Some numbers: Qwen3.8:27b went from 12.2 tok/s to 33.91 tok/s muse-glimmer:30B went from about 14.3 tok/s t…

  43. Hey everyone! Last year, Mozilla released Orbit, an AI-powered browser summarizer hosted on a GCP server.

  44. Running Ollama on iGPU on Linux Table of Contents Like everyone, I’m interested in running my own Large Language Models (LLM), so I installed Ollama in my homelab some time ago. Ollama is a tool to easily run LLM, and exposes this local AI…

  45. I build a desktop Git client called AngKorGit. For the AI parts I didn't want to ask people for an Anthropic API key, so instead it runs the `claude` CLI you already have installed and uses whatever plan you're already on.

  46. Gemma 12B is obviously a very well trained model, I always thought the fine tuning they did on it wasn't really cut out for agentic coding. From my own experiences it struggles to use the tools it's given from Github Copilot and is also ve…

  47. Maintainer here, so full disclosure up front. Anarlog is a desktop app and no bot joins the call.

  48. It was nice to use on vscode with offline Ollama but today notices it is not updated anymore. So, alternatives you use?

  49. What's up, ChickenButt? I made a free, native GTK client called ChickenButt.

  50. py-ollama-openai-bridge A lightweight HTTP proxy that translates OpenAI /v1/chat/completions requests to Ollama's native /api/chat API and back. The Problem Ollama's OpenAI-compatible endpoint (/v1/chat/completions) has serious runtime con…

  51. 🦙 Ollama Usage Widget A native macOS menu bar widget that tracks your Ollama Cloud usage (weekly & session quotas, per-model request counts, cost) and your local Ollama server state — right from the menu bar. Features Menu bar pill — black…

  52. A model that advertises 40,960 tokens was being served 4,096 I've been running knowledge distillation experiments on consumer hardware: one RTX 5080 (16GB VRAM), 32GB of system RAM, an 850W supply, and Qwen3 at 1.7B, 4B, 8B, and 14B, all Q…

  53. I'd like to give user queries to ollama running a small model, but ollama does have a complex tokenization pipeline. Most software that processes user input has turned out to have flaws over the years, is there anything I can do to sanitiz…

  54. - loading / unloading models - ollama, lm-studio, vllm - optional security tokens and visibility and many more on - https://github.com/Chleba/ollamaMQ

  55. ChatOSS is built on Ollama. If you use Ollama, ChatOSS local works out of the box.

  56. VibePod runs coding agents (Claude Code, Codex, Qwen Code, and others) in containers. 0.20 adds credential profiles — keep a subscription login, an API-key setup, and e.g.

  57. GenOffice (local-LLM fork) A free, open-source AI Office suite — this fork drops the cloud-account requirement and talks to any OpenAI-compatible endpoint instead: a local server (Ollama, LM Studio, vLLM, llama.cpp server, text-generation-…

  58. I kept losing track of what I had submitted vs just opened, across Microsoft/Eightfold, Lever, etc. Chrome extension, load unpacked.

  59. I am starting to explore some orchestrated use of local ai models, supervised with claude (desktop app with co-work). Essentially, having Claude driving local LM Studio and Ollama models for tasks.

  60. Local Knowledge Graph Ask a local model a question, watch it reason step by step, and see the steps drawn as a graph. Blue edges are embedding similarity — association.

  61. Not liking the design of the currently common ones so I made my own.

  62. I (like thousands of other engineers) got tired of trying to understand Claude's writing. ASD-STE100 instructions didn't change much, so I decided to approach the problem in the 2026 style: use an LLM.

  63. I wanted the "E.V." assistant from the new Spider-Man film, but actually running on my own machine — no cloud, no subscription. So I built it.

  64. Hi /ClaudeAi community! let me showcase one of the coolest projects i worked untill now.

  65. Benchmark local LLMs with custom sets of tasks via Ollama and similar providers on a wide variety of metrics. You can use deterministic evaluation criteria or LLM judges with custom instructions.

  66. could not extract summary

  67. Scout, a self-hosted Discord memory bot. Listens to the conversation and creates summaries and answers questions.

  68. Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed. In this…

  69. connects to wherever you're. Claude, ChatGPT, Codex, Claude Code, opencode, Local LLM (ollama), or any MCP client.

  70. homebench **Benchmark the local LLMs you already have — speed, memory, and quality — as a live terminal leaderboard. homebench is a single-command TUI that discovers the models installed in your local runner (Ollama, LM Studio, llama.cpp,…

  71. The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agentic AI, particularly in separating inference, orchestration, and exec…

  72. Anybody else does something like this ? I tend to save the big guys (Claude and codex) for the last days of the week, i use cheap models for most of my day to day work (now deepseek v4 flash) and save my quotas with the smart models for th…

  73. 20+ years of PKMS obsession plus big enough chops to be dangerous at vibe-coding have led to this: a full-featured, block-first personal knowledge base that you host yourself and can connect to your own AI. Imagine if Capacities and Obsidi…

  74. Welcome to the daily AI news brief for vibe coders. It is Thursday, July 30, 2026, and the last 24 hours were about the plumbing you build on: a new Model Context Protocol spec, another open-weight model in your terminal, and two speech AP…

  75. Quick backstory on how this version came to be: the original MandoCode is a CLI, and its UI is built with RazorConsole — Blazor components rendered into a terminal. Genuinely clever tech!

  76. I just built this as a Personal/Portfolio project. I genuinely needed it, so I built it and I regularly use it.

  77. I kept reaching for Opus by default on every project — "just in case" — with no real basis for the call. Then I'd burn through my limits on work Sonnet would have handled fine.

  78. Hey everyone, Most MCP memory servers available right now are either simple key-value stores or bare-bones text search wrappers. They store data, but they don't really connect or curate it.

  79. Overview Improve and reply to messages with AI. Free and private, no key: on-device or a local model (Ollama/LM Studio).

  80. Local AI agent Run `kdeps` and you are in an AI REPL. Use Ollama or llamafile for a fully offline, private coding agent - no API key, no cloud dependency.

  81. Benchmarks Apple Silicon LLM inference benchmarks Measured on our own production fleet — not estimated from spec sheets. Each figure is the median of 3 runs via the Ollama API with a fixed prompt and 512-token generation budget.

  82. In addition to the previously supported modes (Quick/Balanced/Deep/Unlimited) where LLMs only gave you what you wanted in a limited budget, i took a look at what Bullets did in terms of response. Turns out representing responses in Bullets…

  83. Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage credits. We’re quickly working on adding capacity to expand access.

  84. Parley One binary that turns every machine on your network — Apple Silicon Macs, NVIDIA workstations, spare CPU boxes — into one shared, private LLM cluster. OpenAI- and Ollama-compatible, so the tools you already use just work.

  85. Cursor type: "prompt" hooks always use Cursor’s own model. No model / baseUrl — you can’t point them at Ollama.

  86. Lexicon is a rich-text editor with grammar checking and AI writing tools (rewriting, tone shifting, and summarizing) that run entirely on your machine. No account, no API calls, and nothing uploaded.

  87. TL;DR: Aesop is a multi-agent coding harness, built mostly by Claude running on itself. As of 0.4.0 it has two "seats" you point at any model from one config block — the seat that writes code, and the seat that decides whether to ship it.

  88. Local CLI that profiles token spend for AI coding agent sessions (Claude, Cursor, Codex, Ollama). AgentCost Local open-source CLI that profiles what drives token spend in AI coding agent sessions — not just totals.

  89. 🛡️ AI Firewall — Security Gateway & Reverse Proxy for LLM Traffic AI Firewall is a security gateway that sits between your applications and LLM APIs (OpenAI, Anthropic, Gemini, Ollama, vLLM). It inspects prompts before they leave your peri…

  90. Hi HN, Last summer, Mozilla killed Orbit (their own page-summarizer extension). So I rebuilt it: fully local, running on WebLLM (Chrome/Edge), WASM (Firefox), or your own Ollama instance if you want a more powerful model.

  91. Use AI free — Gui, Web, Cli, Telegram, 2,000 MCP, 100K Skills . LiteLLM (100+ providers), local models via Ollama, /lang in 34 languages, Mesa Redonda, voice, OCR, MemPalace, embedded sandbox OS.

  92. Correct spelling and grammar in selected macOS text with Ollama models running locally. Read the guideLocal-first correction for macOS Polish text without sending it away.

  93. I built a browser called Bah. Same family as Perplexity's Comet — an AI that actually operates the page instead of just chatting about it.

  94. VaultCharts·desktop charts·AI when you want it Fetch, scan, rank, confirm — like a co-trader on your desk. Indicators, patterns, and chart tools on your desktop — or ask the AI to pull the numbers.

  95. Built an desktop AI companion that always pinned at top. Some of the features: - Lightweight: Built on Tauri v2 + React + Rust for low RAM usage.

  96. Local-first · Private · Open source One native home forevery local model on your Mac Every text, image, and speech model already on your Mac — pulled through Ollama, cached by Hugging Face, or dropped in by hand — discovered and run in one…

  97. 57 words • 76 tokens • 1 minute read The writing tool that thinks alongside you. Hillnote is where you craft your thoughts in notes, drawings, databases and plans — powered by AI that runs on your device (even when you're offline).

  98. Mnemo AI A local agentic AI assistant with MCP (Model Context Protocol) integration, RAG capabilities, and intelligent conversation management. Built on LangGraph with LangChain for multi-provider LLM support (Ollama, Amazon Bedrock, OpenA…

  99. How to Code with GLM 5.2 on OpenCode Coding with GLM 5.2 on OpenCode is a bet on your own engineering. Decide the architecture and interfaces first, let the cheap model write the code on a fixed twenty-dollar Ollama plan, and bring a str I…

  100. A couple weeks ago I was wondering if I could build a GGUF model file that, when run on ollama, would deterministically answer with the same exact sentence (not simply using a system prompt). That was surprisingly easy to implement, levera…

  101. I built Wisp because I kept alt-tabbing between my editor/browser/docs and chat apps, copying context over, waiting for an answer, then copying the result back. Wisp is a desktop overlay for using AI from whatever app you are already in.

  102. Why I advise against using Hermes Agent On the Ollama website, I saw that I could integrate any model they provide with the Hermes agent, but also with OpenClaw and other tools like Claude Code or Codex. I had heard about Hermes Agent befo…

  103. Ollama vs llama-server — Quick Benchmark I got myself a Tesla V100 a while ago and thought why not put it to some good use for once: Curious about the feasibility of the somewhat controversial ollama in comparison to straight llama-server…

  104. Local Agent Toolkit Keep frontier-model tokens for frontier-model work. local-agent lets Codex, Claude Code, or a human developer delegate small, bounded coding tasks to an Ollama model running on local hardware.

  105. I pay for Claude Max while my GPU sits idle, so I built a small bridge between the two. Understudy is a single UserPromptSubmit hook.

  106. LocalAgent An educational lab of AI agent architectures, built on LangChain and a local Ollama server. Each variant is a separate, runnable CLI so you can study one mechanism at a time and watch it work through the logs.

  107. I'm looking for a high-quality solution for translating food recipes and/or menu content into 30–50 languages. Accuracy and natural wording matter much more than the cost.

  108. TinyToT — Tree of Thoughts Inference Server A lightweight, Ollama-compatible inference server that answers questions through knowledge retrieval, structured reasoning chains, and MCP tool calling — with no model weights, no training, and n…

  109. A native .NET LLM inference engine for GGUF models — with a command-line tool, a browser chat server, and Ollama- & OpenAI-compatible APIs for programmatic access.

  110. The popular open source AI tool Ollama has raised a $65 million Series B, led by Theory Ventures, founder and CEO Jeff Morgan tells TechCrunch. This round follows a previous $15 million Series A led by Benchmark’s Peter Fenton.

  111. Ollama: all aboard open models July 9, 2026 Michael and I first met in college, where we started our first company, Kitematic, which made Docker dead-simple to run. In 2015, it was acquired by Docker.

  112. Every week I'd be deep in a task, hit the limit, switch tools, and spend 20 minutes re-explaining everything. So I built CodePass, a terminal harness that runs your agent in a PTY, watches output for rate limits and failures, and switches…

  113. I just launched Rewire Text, a Windows + macOS tool that transforms text in any app at the press of a hotkey. Sits in the menu bar / system tray until needed.

  114. Hello everyone, After the Claude Code leak started floating around, I spent time studying how the workflow was put together and rebuilt the core experience into my own project. I’m calling it Super Grokie, because it started as a joke but…

  115. The "should I use Claude API or run Ollama locally" question comes up here weekly. Everyone has an opinion.

  116. otaku One terminal client for all your local model servers. Chat, pipe, search your history, and manage RAM across Ollama, LM Studio, MLX (omlx), and any OpenAI-compatible endpoint — from a single command.

  117. I had posted a while back about "agent-smith" a claude code skill that sends the heavy drafting to free models, gemini free tier, local ollama, so claude's tokens go to judgment instead of grunt work. shipped a big update this week and som…

  118. Claude Science is good, but I want an open-source version that supports full private deployment. So I spent a lot of time over the past few days building Open Science with SPEC Coding.

  119. Hey all, So I tend to favor the Claude Desktop app in Code mode as the GUI does a great job of previewing code, MCP browser interactions/screenshot evals/etc. But I recall people saying they could get Claude Desktop to use a local API.

  120. Synaplan is Apache2, comes with Open Source helm charts for K8s and has all major AI APIS integrated, including Ollama for local fun. Pretty powerful for hosters, who want to run their branded version of it (explicitly wanted).

  121. ragit Local RAG CLI to chat with any folder of documents using Ollama. Install cd ~/ragit python3 -m pip install -e .

  122. Gemma 4 is now significantly faster in Ollama 0.31 on Apple Silicon via multi-token prediction (MTP), powered by MLX. Performance is now up to 90% faster when used with coding agents, as measured using the Aider polyglot benchmark.

  123. I have been using Ollama to run local LLMs on my Mac, and it has been working just fine. However, my Mac's overall performance took a hit because local LLMs are resource-hungry.

  124. I first just wanted to learn how to add a memory layer (how to use the mem0 Python library) when using Ollama but decided to expand into something bigger. From the first idea, I decided to make a modular system that supports APIs, MCP, and…

  125. Fast local LLMs on Apple Silicon: sub-second model loads, faster than Ollama on long prompts. OpenAI and Ollama-compatible.

  126. I currently host a local MCP server with ollama and a qwen3-coder 30b model. I have a Claude pro subscription I'd like to be able to call the qwen3-coder model the same way I call a haiku, and also allow it to be spun up as a sub agent.

  127. Tiny LLM Benchmark: Jetson Orin Nano Super 8GB 8 tiny LLMs benchmarked across 4 power modes on Jetson Orin Nano Super 8GB: llama.cpp vs Ollama. 25W sweet spot: 43% more tok/s than 15W, better tok/J than MAXN.

  128. "I learned that 4.7 needed heavy structure, rules, and guidelines to perform well." That has been my experience when using Claude Code for app development. I have OpenClaw Codex, and Ollama models for coding as well, and Claude Code is sti…

  129. • EvolutionEngine— L1-L6 autonomous evolution loop • MemorySystem— declarative/episodic/procedural, weighted • NightlyReview— metacognition proto • CodeSandbox— self-modifying code • TokenBudget / LiveStatus / KnowledgeFusion Hunyuan API +…

  130. I’m considering one heavier subscription (~€100/month) and want to know which provides better value for agentic coding. I tested GPT Pro and was satisfied with Codex.

  131. Charon A blind, end-to-end-encrypted marketplace for LLM inference, paid in bitcoin, built on the NUTS ecosystem. Providers run Ollama and choose which models to sell; consumers run any coding agent (via Nemesis8) against those models.

  132. ggrun ggrun = "gguf run". Formerly llm-server.

  133. Some things shouldn't leave your computer. Medical notes, legal drafts, journal entries, messages to your therapist, that honest performance review you're still editing.

  134. For the past few days, we have been working on an open-source, self-hosted real-time speech-to-speech translation tool called PolyTalk. The goal was that there are people and organisations who need privacy around the tool they are using, a…

  135. Awesome Backend Security Auditors Curated, keyless security auditors for the modern backend stack — BaaS platforms, headless CMSs, GraphQL engines, workflow runners and local LLM servers. Every tool here runs locally and confirms each leak…

  136. I got tired of sending every text I translate to Google/DeepL. Even with all the opt-out options and privacy policies, it never felt right especially for some work documents, personal writing, or anything sensitive.

  137. https://preview.redd.it/0jai8prknl8h1.png?width=2040&format=png&auto=webp&s=61576e05a908614b672db1fc89cb46cd4e148cde Steps to reproduce Run claude cli with ollama provider (`ollama launch claude --model gemma4`) Run `/model` command in the…

  138. I've been vibe-coding tools to automate chunks of my consulting work, fell down a rabbit hole, and started building actual products. Suddenly I'm in a world of unknown-unknowns and known-unknowns.

  139. Introducing Machinaos: AI That can Build itself depending on the Task and also a Multi Agent Orchestration Platform to run Loop Agents and Control Agents like Claude code, Codex, etc. Bring your own API keys or Claude Code sub (or run mode…

  140. Hey All, I built Konxios (spelled as Conscious) because I had a problem, AI tools were multiplying, but my workflow was getting more fragmented with no privacy first solution. One tool for chat.

  141. MyLLM Connect Run your own private AI backend on your Mac or PC, and connect the MyLLM iOS app to it in one tap — over real HTTPS, from anywhere. MyLLM Connect is a small desktop companion (system-tray app for macOS and Windows) that turns…

  142. Been trying to build recurring automated pipelines — pull data, analyze it, push outputs to a task manager — and wanted to keep it running through my Claude Max plan to avoid per-token API costs. Was using claude -p via SSH to trigger head…

  143. Run Claude Code-style subagents across your local model fleet. subagent-fleet Run Claude Code-style subagents across your local model fleet.

  144. A free, private web UI for Ollama models. Chat with Ollama Cloud models using your own API key, or connect to the Ollama running on your own machine — with streaming replies, charts, in-browser code execution, and GitHub context.

  145. A free, private web UI for Ollama models. Chat with Ollama Cloud models using your own API key, or connect to the Ollama running on your own machine — with streaming replies, charts, in-browser code execution, and GitHub context.

  146. SemanticSourceCode A C# tool for semantic code search with local embeddings. Search your codebase by meaning, not just keywords.

  147. OpenDevOps Agent Open-source multi-cloud DevOps agent (AWS + Azure). Bring any LLM via LiteLLM — OpenAI, Anthropic, OpenRouter, Groq, Gemini, Mistral, Ollama for air-gapped / regulated environments, or reuse your existing Claude Code subsc…

  148. Built a free, private web UI for Ollama. You can chat with cloud models using your own API key, or connect it to a local Ollama instance running on your machine.

  149. I built MandoCode, an open source CLI coding agent in .NET that runs against local Ollama models. No API keys, nothing leaves your machine (Ollama cloud models work too if you want them).

  150. This paper explores the value of agentic AI tools for cybersecurity purposes. We evaluate the efficacy of a general-purpose GenAI Large Language Model- (GenAI-) based agent when powered by three different Ollama-hosted general-purpose open…

  151. I've been running Ollama locally for a while and the one thing I kept missing was voice. Every solution I found either sent audio to the cloud, needed a GPU, or was locked to macOS.

  152. I am using qwen3-vl:8b and ollama for doing OCR on scans of handwritten letters and it is doing a decent job. Any other models I should know about for this kind of OCR?

  153. Hi everybody, I spent the last two weeks building zerostack, a coding agent in Rust, focused on memory footprint, shipping with ollama and vLLM integrations. I managed to get it to run at ~16MB (with peaks of 24MB) of RAM usage, and no CPU…

  154. Anyone’s trying to develop their own AI? i am, trying to build the “skeleton”, all the tools he’s gonna use from domotica to copilot, everything else but the AI for now, i was thinking about Ollama, the alpha version of it is currently usi…

  155. I built a free, visual, local-first AI agent platform because every option was either cloud-locked or required living in Docker or using a terminal. You've probably seen OpenClaw and Odysseus recently.

  156. Hey everyone. I'm brand new to running LLMs in general, even more new to running them locally, and the sheer number of tools available is absolutely overwhelming.

  157. Please tell me if you need the source code. My issue is that my Jarvis is stupid rn.

  158. Hey guys, I’ve spent the last 3 months building an open-source (Apache 2.0) project called TrueNorth I kept running into the same problem at work: trying to get an LLM to talk to a human (like a medical intake or HR screener), guide them t…

  159. Considering buying a maxed out MacBook Pro M5 Max with 128GB of RAM and one of the things I want to figure out before pulling the trigger is whether local models are good enough to actually replace cloud AI coding tools. My current setup i…

  160. Hi, I’m basically one of you, except I’m stepping onto the other side of the table today, fully prepared to accept your ridicule. Obvious disclosure: this is my project, so yes, this is self-promo — but I’m posting it here because this is…

  161. I'm trying to use Claude Code with local Ollama models, but every prompt fails with: The strange part is that it happens even for extremely small prompts like: hi say apple What is 1+1? Answer with only one character.

  162. Selecting the "best" local model usually depends on the task and the hardware. I created this script as an easy way to test local Ollama models and keep the test output organized.

  163. Researcher GPT-5 Engineer Claude Critic GPT-5 Innovator Gemini Manager Context Guardian Agentic Workflow Architecture · v1.0 The future of AI Collec An R&D platform where differents AI agents collaborate under the supervision of a Context…

  164. # I Replaced My AI Agent's Flat Fact Store with a Graph Database and It Runs in 85MB I've been building LocalClaw, a local-model-first AI agent framework running on personal hardware through Ollama. No cloud, no API costs.

  165. AI Sales Agent An open-source, AI-powered sales agent built with Next.js 15. Generates content, finds leads, and automates outreach — all running locally with Ollama (no API costs).

  166. Overview Store, search, and chat with web page content locally. AI chat (BYOK), full-text search, markdown export, and optional RAG endpoint.

  167. ollama-wsock-connector A small Rust client that bridges a remote WebSocket service to a user's local Ollama instance — so a service operator can offer "bring-your-own local inference" without ever proxying or holding the user's prompts and…

  168. If you've used ollama for any length of time, you've probably hit this: pulling 9b6d12fa8910... 99% ▕████████████████████▏ 6.9 GB …and then it just sits there.

  169. The DDS Vibe Academy is a free, 38-class curriculum on AI coding published by Robert McCullock, founder of Design Delight Studio in Boston. Covering Claude Code, Google Antigravity, Gemini, Cursor, Ollama, and more.

  170. I'm scoping a hybrid AI pipeline for a consulting client in a regulated industry (GLBA-covered, NPI involved). Trying to validate the architecture before bringing on an engineer to build it.

  171. Threat Intelligence Table of Content Unpatched Ollama Vulnerabilities: Phishing Overlays and Data Exfiltration Ollama’s desktop app is vulnerable to phishing overlay and data exfiltration attacks via indirect prompt injection, overwriting…

  172. So, last week I tried to update my unused local LLM setup. I had to stop using it because quality was too low and deepseek was too cheap.

  173. Tlamatini A local-first AI developer assistant that goes beyond chat. Run it on your machine with Ollama.

  174. Because standard coding agents are stateless, every session they start from scratch. I built Zerikai_memory around a different model: you decide when the agent learns your codebase, not the other way around.

  175. https://github.com/user-attachments/assets/e4897391-c5a8-4391-93c3-9f8b76155f11 Setup your local LLM stack effortlessly. Starts fully configured Open WebUI and Ollama harbor up Now, Open WebUI can do Web RAG and TTS/STT harbor up searxng s…

  176. Posted on other feeds last week and figured some of you out here might be interested as well; Someone commented asking if it supported OpenAI-compatible endpoints (LM Studio, vLLM, OpenRouter, Together, Groq, LocalAI…), so i have spent few…

  177. Multi-model Ollama comparison, benchmarking, and evaluation — in your terminal. Zero dependencies.

  178. I developed and maintain Anubis OSS, an Apple Silicon Mac app for benchmarking local LLMs. Mostly built around Ollama (also handles LM Studio, MLX, and Apple Intelligence if you've got those).

  179. So I have to admit, I have fallen victim to the cool looking dashboard videos but I’m struggling to find a use for me. I love AI and use it daily for general questions and some deeper research (Google Gemini free tier).

  180. could not extract summary

  181. I’ve tried openclaw locally for about a month. Hardware: M5 Pro w/48 gb ram.

  182. https://preview.redd.it/i90oxxk7n03h1.png?width=1898&format=png&auto=webp&s=7d219c804fda7dfe122b84fcdb6d0d6883818c68 A while back I came across TradingAgents — a really cool multi-agent LLM stock analysis framework where like a dozen "agen…

  183. Hi, I’m pretty sure I have seen people typing /model and seeing all available models. I have to type models from memory.

  184. Wanted to share a workflow I tested on a real flight, in case anyone else is trying to set up offline Claude Code. The core idea: using ollama to pull the needed model of what you need, and then use it to run claude code The setup, in orde…

  185. https://preview.redd.it/sm4ysgdw1w2h1.png?width=1376&format=png&auto=webp&s=3705932403919814fbf2008a1cba189d17e0591e Thanks everyone for the advice on my previous post (24/7 Headless AI Server on Xiaomi 12 Pro (Snapdragon 8 Gen 1 + Ollama/…

  186. I'm trying to install LLaMa with PI agent. I ran curl -fsSL https://pi.dev/install.sh | sh export PATH="/home/user/.local/share/pi-node/node-v22.22.3-linux-x64/bin:$PATH pi install npm:pi-llama.cpp ​ These commands installed pi, added them…

  187. Just wanted to post a tip (I'm human, not an agent, watch: fart). I use Deepseek-v4-Flash on a lot of my agent work, and as I'm learning and testing these things.

  188. ◈ EVE AGENT V2 UNLEASHED ◈ Local-first autonomous AI coding agent — powered by Ollama No accounts. No cloud lock-in.

  189. So on a local lm like ollama, or lm studio etc. you can run questions and prompts.

  190. I work on AxonFlow, a source-available (BSL 1.1) runtime for long-running agent workflows. We’ve been running it in front of Ollama-served models and OpenAI-compatible local endpoints (llama.cpp `--server`, vLLM, LM Studio).

  191. Hey r/ClaudeAI, If you are using Claude Code or building terminal agents, you know the exact moment the context window starts degrading during long-running tasks. I wanted to build a persistent runtime layer to offload those heavy, multi-s…

  192. I've been building this for the past few months as a side project — started because I didn't want to run llama.cpp from the command line every time I wanted to try a model. I just wanted something that worked with a click.

  193. I've been running Claude Pro (Opus 4.7 / Sonnet 4.6) for about 3 weeks on a complex personal AI infrastructure project. I keep structured session logs with timestamps and Birkenbihl-style metacognitive fields after every session.

  194. It keeps running into race conditions/OOM when switching between models, as the previous process doesn't unload from VRAM fast enough. What is the simplest fix for this right now?

  195. On this page: We started an Ollama container on a MacBook. There's no NVIDIA GPU, no CUDA toolkit, and macOS doesn't even have CUDA drivers.

  196. Claude Code Masterclass 2026 The definitive end-to-end guide to Anthropic’s agentic coding tool — installation, Ollama local fallback, CLAUDE.md, Skills, Subagents, Agent Teams, Hooks, and MCP. Everything you need before building productio…

  197. I posted recently about EvalShift, an OSS CLI for regression-testing LLM model changes. A few people pointed out that for LocalLLaMA, the more interesting use case may be quantization regression: Q8 -> Q4_K_M Same base model, same prompts,…

  198. Hi everyone, I’ve open-sourced agentmw, a framework-agnostic middleware that sits between your LLM client and agent logic to make agents more reliable on long runs. Key features: • Real-time failure detection (loops, redundant calls, contr…

  199. I first posted about PrivateScribe.ai ~1yr ago and have recently jumped back intent on bringing it to a functionality that makes it actually usable by non-technical users. One year ago it worked but only the bare minimum.

  200. So we've been working on a voice bot that handles customer calls and honestly the testing part has been brutal. We were literally calling the thing ourselves to check if it broke after every change.

  201. This site was paused as it reached its usage limits. Please contact the site owner for more information.

  202. Disclosure: I made this. Open-source, MIT, Windows + Linux.

  203. Hey everyone, I wanted to share a project I've been working on called Glia. It is a 100% offline, local-first RAG and memory layer designed to connect your AI web chats (Claude, ChatGPT, DeepSeek) with your local developer tools (Claude Co…

  204. `🧬 Flux‑Genotype – A CPU LLM that rewrites itself` I've been working on an open-source kernel called **flux-genotype**. It orchestrates local models (TinyLlama, Llama 3.2, Hermes 3, DeepSeek-Coder) into a self-modifying ecosystem.

  205. Built this for myself after wanting to use local LLMs during work calls without the window showing up on screen share. Every existing tool was either cloud-only or a 200MB Electron app.

  206. AI Agent Security (guest lecture, MIT 6.566, April 2026) You can run demos with uv, for example uv run 00completion.py. For some, you will need Ollama and the appropriate models downloaded.

  207. Thuki is a floating overlay that appears on double-tap Control from any macOS app, including fullscreen. Powered by Ollama, no API key, no account, no cloud.

  208. EDIT 2: Trick-Assignment-828 pointed me at the actual rule update from the mods - Rule 3 Low Effort was expanded to cover LLM-assisted posts without disclosure. Disclosing now: Disclosure: I'm a non-native English speaker (German).

  209. I’ve been seeing a lot of "Context Corruption" in multi-agent systems where agents slowly drift away from the facts or leak data they shouldn't. Things like context pollution and context exposure can leak major things like your API keys an…

  210. Hello, I'm currently using Ollama / lm studio for things like code inference and proof reading emails, etc. Definitely not experienced in this space but looking to grow.

  211. so i am not that experienced when it comes to llms, i just have ollama and open webui and occasionally test (play with) new releases from time to time. a few weeks ago i started using Windsurf, i do not know coding or anything but i loved…

  212. zerostack Minimal coding agent written in Rust, inspired by pi and opencode. Features Multi-provider: OpenRouter, OpenAI, Anthropic, Gemini, Ollama, plus custom providers Standard tools: all of the standard tools exposed to coding agents,…

  213. So I’ve been nerding out hard about memory, and have come to the conclusion that context management is too high level and dynamically changing the weights would be best. Luckily, this morning I checked my news feed and saw this new paper!

  214. Anthropic and OpenAI claims that their models are so powerful that it can “break” their box…but what so special about their agent implementation? Is it not just basic ReAct loops with tools?

  215. All major Lemonade capabilities, including OmniRouter, coding, image gen, speech gen, and transcription are all available on Lemonade for macOS thanks to the hard work of u/GeramyL. If you're on macOS and just looking into Lemonade for the…

  216. Hello, After almost two years of on-and-off development, 5 complete architectural rewrites, and hitting a few brick walls, I’m finally open-sourcing a project I built to scratch my own privacy-paranoia itch: Nexidion. GitHub Repo: https://…

  217. I recently made ollamatps.com for my own model-selection workflow and thought it might be useful here too. It shows 39 Ollama cloud models sorted by average TPS over the last 24 hours, and I added the Artificial Analysis Intelligence Index…

  218. My GPU power consumption is 250w (undervolted rtx3090) when I added Qwen3.5-27B-GGUF to Ollama using a template (Modelfile made by gpt). I gave it 3 task to test it, build a snake game, build a flappy bird game, and make an interactive gri…

  219. I've already searched, but information is getting updated each week, so it's really hard to get an answer, I really hope some of you guys can give me some tips. And can I use an agent with it to enhance the code?

  220. Claude Code is the best AI coding tool I've used. But being locked to one provider, one pricing model, and one model catalog always bothered me.

  221. SIE: Superlinked Inference Engine Open-source inference server and production cluster for embeddings, reranking, and extraction. 85+ models.

  222. https://github.com/ollama/ollama/releases/tag/v0.30.0-rc15 Hopefully this has more devs come to llama.cpp to support Day 1 releases due to Ollama now moving to using llama.cpp directly. Additionally, I hope that Ollama makes it clear that…

  223. Hi, i just wanted to share what im playing with for last couple weaks. I built my own AI harness: TinyHarness My main goal was low memory footprint, it is not written in Typescript/Javascript/Python, leaving as much memory as possible for…

  224. Hi everyone, I'm happy to share ml-intern, which is a harness for agents to have tighter integration with Hugging Face's open-source libraries (transformers, datasets, trl, etc) and Hub infrastructure: https://github.com/huggingface/ml-int…

  225. Will there be a non-cloud version of Deepseek V4 flash available for Ollama? Or do I need to go to another framework to get a version that will be supported?

  226. It runs Claude Agent SDK - with the full Claude Code feature set - in isolated worktrees with 4 built-in MCP servers (GitHub, GitHub Actions, Memory, Codebase Tools). You configure triggers in YAML: workflows: review-pr: triggers: events:…

  227. Hardware OS: Windows 11 Pro 10.0.26200, Build 26200 CPU: Intel Core Ultra 7 270K Plus, 24 cores / 24 threads, max clock 3.7 GHz RAM: 32 GB DDR5 @ 5600 MHz, 2x16 GB Crucial CP16G56C46U5.C8D GPU: 2x NVIDIA GeForce RTX 3090, 24 GB VRAM each,…

  228. I got a bit further with my harness for running Qwen 3.6 model on Codex. While testing, analyzing, and building the harness, I evolved TBG(O)llama-swap into a full forensic UI bridge and LLM analytics tool where every harness finding, modi…

  229. Hey everyone, I'm a complete beginner in AI Agents, and I do some self hosting at the moments, I was interested to know if it was possible to self host agents like claude one using our own IA. Because I know things like Ollama to run your…

  230. Hi all I had a quick question while we wait for llama.cpp MTP implementation, have any of y'all tried Gemma4 MTP models on ollama and or transformers? What was your experience and or cli args and or workflows like?

  231. LM Studio has been really easy to use, but it seems, like they dramatically changed the interface from 0.3 to 0.4. I have 3 GPUs, and want to assign one to a Research model at port 1234, one for Writing at 1235, one for Utility at 1236.

  232. Hi all I recently started a new job and we're doing python development for a ci cd metadata consolidation library for analytics and we cannot use no stuff like claude code or codex or gh copilot or any model APIs (free or paid). I got a la…

  233. Recently I've been looking at my personal AI infrastructure. I've built a lot of tools for personal use, a budget and tax helper, an eBay selling assistant, smart home integration, a thermal printer, a task tracker, an Obsidian memory vaul…

  234. I use Claude Code and also pay for Perplexity, OpenAI, Gemini, and run Ollama locally. Got tired of switching tabs when the right model for a question wasn't Claude.

  235. Quick context: M3 max 64gb, currently running llama 3.3 70b q4 as my daily driver via ollama, qwen3 coder 30b for code (switched from qwen2.5 earlier this year), mlx for the smaller stuff. tried llama 4 scout earlier this year but 64gb is…

  236. Hi there! I was playing around with Ollama and LMstudio, testing local models and had the idea of letting Claude evaluate a few models on their actual capabilities rather than doing it myself.

  237. So this is probably a pretty common thing, but I just want to ask in case am not missing something. I have pretty much no knowledge but trying to learn a bit more about AI's and local LLMs and the whole AI Stack.

  238. I’m building a small “Orchestration as Code” repo for LLM workflows. Does this concept make sense?

  239. The local-first autonomous coding agent. Claude Code UX powered by Ollama.

  240. A lot of cloud providers (Anthropic, OpenAI, Ollama) do plan pricing (non-transparent usage limits) while others like OpenRouter and some neoclouds do per-token pricing (more granular spend) What do you prefer for your agents? Better to se…

  241. argus A RAG-based (Retrieval-Augmented Generation) vulnerability scanner for Go, Python, Rust, npm/Node.js, Maven/Java, NuGet/.NET, and Ruby projects — powered by local Ollama models or any OpenAI-compatible API. No cloud lock-in.

  242. Been using Deepseek-Tui for days. solid for v4 workflows.

  243. Hey everyone. Like a lot of you, I’ve been deeply frustrated by the state of commercial AI.

  244. Every tool (LM Studio, Ollama, llama.cpp) downloads models to its own directory. Same 8GB model × 3 tools = 24GB wasted.

  245. Hey Bros, I have around 80% tokens for the week, If anyone needs it or suggest me what I can do with it will be helpful.

  246. Looking to run AI agents locally on my M5 Pro MacBook. Been experimenting with ComfyUI for image generation and the results have been impressive.

  247. Hey r/AI_Agents — we're launching Irene today, and I want to be straight about what it is, why we built it, and where it's going. What makes Irene different Affordable with massive token limits and the latest open-source models We have gen…

  248. KillClawd is a desktop pet powered by a local LLM. It runs as a transparent always-on-top overlay — a tiny AI crab called Clawd who wanders your desktop, reacts to your cursor, fights mobs, explores castles, rides vehicles, and has genuine…

  249. I’ve been seeing a lot of newcomers asking about hardware specs lately, and there’s this weirdly common myth that you need a heavy server or a GPU instance to run Cla͏ude-based agents. You really don’t.

  250. Hi HN, I built this for myself because I wanted my coding agent (Claude Code) to actually be able to use my Obsidian vault as a knowledge base, not just grep it. The use I get the most mileage from is asking the agent to find notes that sh…

  251. Hello r/LocalLLaMA, Some time ago I posted here about the Linux distro for local AI workloads that I was building: https://www.reddit.com/r/LocalLLaMA/comments/1igpkc8/i_built_a_linux_distro_to_run_nvidia_gpus_for_ai/ It worked well, but I…

  252. Ran a full eval against four local models last weekend and the spread between them is wider than I expected. All running through Ollama on CPU, no cloud, same prompts, same hardware.

  253. Planning to build my AI rig, to run Ollama / OpenClaw...which bundle should I start with? This will be a dedicated machine.

  254. BadPhotosOut A native macOS app that walks your Apple Photos library, asks a local Ollama vision model to judge each photo against a free-text criterion you supply, and surfaces the flagged photos with the model's reason. No automatic dele…

  255. Hi everyone, I am looking for an AI tool or a specific workflow that allows me to train or fine-tune a model using my own texts. My main goal is to have the AI generate content that mimics my specific tone, sentence structure, and vocabula…

  256. Hey y'all, question. I don't code, but I'm running a unraid server with a lot of self hosted cloud storage, local ai and media stacks.

  257. Is llama.cpp much slower for M4/M5? I heard ollama is faster due to mlx support since March.

  258. I wanted to try installing a simple small model like nemotron-3-nano:4b from ollama and try it for simple quick fixes offline without burning credits or time. the model works well on ollama run time but when I try to use it on opencode, th…

  259. English | 简体中文 | 繁體中文 | Русский Docker AI Stack Deploy a complete, self-hosted AI stack on your own server with a single command. Zero-config: all services auto-configure on first start Secure: Ollama, LiteLLM, and MCP Gateway generate API…

  260. Bleeding Llama: Critical Unauthenticated Memory Leak in Ollama TL;DR We discovered a critical vulnerability (CVE-2026–7482, CVSS 9.1) in Ollama that enables unauthenticated attackers to leak the entire Ollama process memory, potentially im…

  261. Very quick initial test of Gemma 4 new MTP model via Ollama (llama.cpp doesnt support yet) https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/ Running in Open Webui to view token/s output and I…

  262. Hey everyone, I’ve been thinking about a project idea and I’d love to get your feedback. The idea is to take a 1TB SSD and turn it into a fully portable AI system.

  263. Hey everyone ! Quick Modly update.

  264. This project was born out of time I spent digging into a biologically inspired algorithm I was using to measure co-activation for placement of experts and ranks onto chips. The default scheduling that vllm provides can end up causing laten…

  265. If you have to choose between these two, what would u choose? I've been using Ollama free cloud models; it is great for small tasks, but some cloud models need a subscription.

  266. TLDR for those that dojr wanna read below I need a new good free place online to pickup roleplay where should that be and what can I do locally? 9070xt 32gb ram desktop and preferably but I know it not great, 4060 laptop 32gb ram.

  267. I’ve been building a small open-source CLI coding agent for local models. It runs with Ollama and works best so far with Qwen Coder.

  268. ISONQ reads PDF, DXF, DWG, and STEP files on the shop's own workstation and produces a priced quote. Nothing leaves the machine.

  269. I have build a CLI tool which help to scaffold the full project with docker, make, database setup. https://go-bootstrapper-docs.vercel.app/.

  270. Hey everyone, I built a tool that creates movie recap videos automatically using local models. The problem: making recap videos takes forever.

  271. Plenty of CLI coding agents will talk to a local LLM, but the catch is the ecosystem. Skills, slash commands, MCP servers, plugins, hooks: all the interesting tooling has been built specifically for Claude Code, and parity on every other a…

  272. Hey everyone, I've been working on Autai — an open-source desktop app (Electron + React) that uses AI agents to automate your browser. You just type what you want in plain English, and the AI opens a real browser and does it for you.

  273. Hey everyone, I’ve been lurking here for while. I’ve really been enjoying messing around with my 6gb card on my laptop using Gwen 3.5 4B, ollama, and Open WebUI.

  274. Hi all, I have recently installed openclaw on a raspberry pi4, linking it to my local Ollama instance (RTX 3090 with 24Gb, as well as 96Gb of DDR5 RAM bought before the madness), in my case running Qwen3.6 (latest) capped at 16k context. A…

  275. Every LLM conversation starts from zero. RAG helps, but it can't learn from what's happening right now.

  276. (I'm on Windows system running these models locally) I've used both Codex and OpenCode with Qwen 3.6 27b and 35b running locally. I'm having a bitch of a time getting them to correctly create files.

  277. I’ve been building AON, a communication layer for Claude Code that moves beyond simple chat into structured team coordination. It implements the Agent2Agent (A2A) protocol over NATS pub/sub.

  278. I've been experimenting with using Ollama to run Claude Code locally with models like Gemma 4, thinking I could avoid API costs. However, I quickly realised these models aren't really optimised for Claude Code's agentic workflows — they te…

  279. LDR maintainer here. Thanks to the strong support of r/LocalLLaMA community LDR got very far.

  280. llmfit: one command to check which AI models will actually run on your hardware Tired of downloading a 15GB model only to find out your system can't handle it? Found this Rust CLI tool called llmfit that scans your actual hardware (RAM, VR…

  281. I wanted to use vast.ai, but ollama doesnt have it, and when i used vLLM I didn't have success. I genuinely don't know what failed.

  282. Hi folks, Some time ago I tried to find a way to clean up my 10k emails inbox and could find anything which worked for me. So I built built my own tool which evolved into a client.

  283. Hi all, I tried to setup my pc to run llm, but got some issue: the first question of the chat is generally fine, but from the 3rd follow up question, the backend often be unresponsive and I have to manually restart the llama cpp server, or…

  284. Most of my LLM cost was on the wrong tier of work. Classification, extraction, JSON formatting, summarization I'm going to review anyway.

  285. Hello! It’s me again, the developer of ADT.

  286. I’m building OpenYak, a desktop AI workspace for using local models with real files on your computer. In this demo I’m using Ollama with Qwen/Qwen3.6-35B-A3B to review an attached budget workbook.

  287. Bench 3 from my 18GB M3 Pro. Bench 2 was the 4B-class post where the comments were mostly right: I gave thinking models a fixed 1024-token cap, Qwen got kneecapped, Gemma E4B needed clearer active-param labeling, and the headline was partl…

  288. Set up a small R&D project which pit different LLMs against each other in a game of Capture the Flag. Each LLM has 30 seconds to prepare any defenses and 5 minutes to capture other flags while defending their own.

  289. English | 简体中文 | 繁體中文 | Русский Ollama on Docker Docker image to run an Ollama local LLM server. Provides an OpenAI-compatible API for running large language models locally.

  290. LocalPilot The Privacy-First AI Pair Programmer for Visual Studio. Bringing the power of local LLMs directly into your IDE with Ollama.

  291. I've been messing around with Hermes for months, and quickly outgrew using it just as a fancy CLI assistant. My goal was to build a persistent, specialized team of local agents that could collaborate on long-term projects without me spoon-…

  292. Hi HN family! I've recently been messing around with open models through ollama (glm-5.1 and kimi-k2.6), and I've been impressed with just how close they are to Claude Sonnet for my needs, especially programming.

  293. I see a lot of support for running LLMs on PCs with ollama to vLLM. Whats the current state for running on mobile?

  294. Hi all, I am building an hi-performance and highly customizable local LLM server wrote 100% in Rust, custom CUDA kernels, zero latency, almost immediate TTFT, and plenty of other features. It is planned to be publish it on GitHub as open-s…

  295. Hi HN — I built Platypus as I wanted to combine note taking, live transcription and knowledge base management in one app. Granola / Notebook LM free local alternative.

  296. Pebble Menu-bar text-polish tool that rewrites your clipboard with a local Ollama model. One global shortcut, seven presets, zero cloud.

  297. Most AI coding tools (Cursor, Aider, Claude Code) assume you have a 200k-token model. If you're running local LLMs through Ollama or LM Studio, or hitting free-tier cloud APIs like Groq or OpenRouter, you've got around 8k tokens to work wi…

  298. I'm working on a homelab AI server with the goal of running small models on GPU and very large models on CPU - for example for overnight coding on complex problems. Specs: 2990WX, 256GB + RTX 2080ti (for now).

  299. Hey all, so I am currently exploring and playing around with Karpathy's LLM Wiki using Claude Code with Ollama and other routed models. I want to create some agents and provide them with tools/plugins, libraries, MCPs, or harnesses to assi…

  300. M4 Mac Mini, 16GB unified, basic spec. For a few weeks I had Qwen 3.5 35B-A3B UD-IQ3_XXS (12GB on disk) running under llama.cpp with --mmap and --flash-attn.

  301. https://reddit.com/link/1sxh0ec/video/egfs5inxtsxg1/player I built Claudex specifically for people who like Claude Code-style agentic coding workflows but want a simpler plug-and-play terminal setup The setup is the main thing I wanted to…

  302. i’ve been experimenting with building a fully local rag pipeline: weaviate for vectors + hybrid search, node.js scripts, qwen 3.5 on ollama what i found is that most of the challenges live in retrieval and chunking, not the LLM, and a good…

  303. 9 min read Mar 31, 2024 -- After my latest post about how to build your own RAG and run it locally. Today, we’re taking it a step further by not only implementing the conversational abilities of large language models but also adding listen…

  304. A practical configuration for a 32 GB M5 Mac that still needs to remain usable Running large language models locally has become surprisingly practical on Apple Silicon. With a modern Mac, Ollama, and a carefully quantized GGUF model, it is…

  305. Use Ollama to Enhance Claude — Two-Engine Setup Pair Claude Desktop on Anthropic with Claude Code routed through Ollama in your terminal. Strategy stays on Pro.

  306. locally uncensored is a desktop app that combines four things most people run separately: chat, a coding agent, image generation, and video generation. all local, all on your hardware, no docker, no cloud account needed.

  307. I just got Qwen3 72B Instruct running on a high RAM setup and I’m kinda confused about the proper way to use it. What’s the correct workflow for running it smoothly (like best quant, tools, or runtime)?

  308. I needed a better way to create and compare prompts when using local LLMs (e.g. via Ollama) in a workflow.

  309. mirollama Local-first multi-agent simulation and prediction engine. Project Origin This project is a derivative work of: Upstream: https://github.com/666ghj/MiroFish.git Target repository: https://github.com/oswarld/mirollama This reposito…

  310. I've been running a multi-agent system in production for a few months — a co-CTO agent + specialist agents (PM, dev, ops) that handle real engineering work end-to-end: design specs, code review, PR implementation, deploys, monitoring. The…

  311. Terminal AI that doesn’t babysit you. • SoT Method → near-zero token waste • Async multi-agent orchestration • Batch tools + unrestricted shell • Ollama / LM Studio / OpenRouter / NVIDIA Watch it take full OS control from one prompt (zero…

  312. Hi all, Looking for some advice with a GX10 I purchased about 4 months ago. I've been having all kind of issues trying to run local models on this device.

  313. Hi HN, I built VT Code, a semantic coding agent. Supports all SOTA and open sources model.

  314. Give any AI agent a persistent memory in minutes. Works with Claude, ChatGPT, Ollama, OpenRouter, and any MCP-compatible agent.

  315. I work full-time as a Program Director. About 50-60 hours a week at my W-2.

  316. Every time I restarted work on a side project after a few weeks, I'd spend the first hour just reading code trying to remember what I was doing and where I left off. Looked for a tool that could help — couldn't find anything that did what…

  317. I've been building a personal AI assistant called Finn that runs entirely on your Mac. No cloud, no subscription, no data leaving your machine.

  318. TLDR: Swapped Ollama for MLX on M1 Max (64GB) to run a 12-agent trading stack using Qwen 35B MoE. MLX wins on throughput and fine-grained sampler control, but I lost the "it just works" convenience of Ollama.

  319. So I'm a newb in certain aspects but not in others, I'm currently running an AI stack on my unraid server: CPU: AMD Threadripper 3960X (24c/48t) Motherboard: Gigabyte TRX40 AORUS PRO WIFI RAM: 256GB DDR4-3200 G.Skill Trident Z GPU: Nvidia…

  320. Anthropic opened Cowork for Bedrock/Vertex/Azure providers and also Custom Inference Endpoints. However, connecting it to a local proxy seems non-trivial.

  321. Qwen Lens Studio A multimodal AI studio built around a single Qwen vision-language model, exposed through five focused tools plus a batch runner and a persistent session log. Ship a screenshot → get code.

  322. Hi all, I tried quite a few models and approaches, but had no luck integrating local models into VS Code Copilot Chat extesion in a useful way. Of course I can see the models there and can choose them, but none of them seem to work even re…

  323. Well, I have a RTX 4090 24GB + 64GB system RAM, AMD Ryzen 9 7950X. Any good model for using in Open WebUI (using Ollama backend?) that outpeforms GPT-5.4 mini, GPT-5.2 Thinking and even Claude Sonnet 3 (the 2024 model)?

  324. Has anyone tried these? I found this on ollama: https://ollama.com/library/kimi-k2.6, https://ollama.com/library/qwen3.6 My issue is that they are extremely slow on my local.

  325. Hi all, Total noob here trying to set up a local model to help me with coding. I am trying the following setup - Ollama running the qwen2.5-coder:7b model in docker with the following compose file services: ollama: container_name: ollama i…

  326. Claude Code drafted the prose. I did the research, direction, architecture, ran the code, caught the bugs, and reviewed every commit.

  327. My Linux/Fedora Local Ai performance is trailing Windows massively? Are there specific ROCm environment variables or memory management tweaks for RDNA3 that I'm missing?

  328. 4 months ago, I released an extension for LibreOfffice Writer that adds an AI copilot to its sidebar. Did a Show HN at the time but got no interest T_T https://news.ycombinator.com/item?id=46233776 I’ve added several major features since t…

  329. I use ollama local model for vscode copilot, but it seems could not get the context of the workspace. For example, I command it to edit or summarize the current opening file, but it does not know which file to work.

  330. 👾 Nedster CLI Coding Agent An unstoppable, fully local, open-source coding agent that runs on your consumer GPU. Tags: ollama coding-agent local-ai cli rag chromadb python qwen Are you trying to use local LLMs to autonomously write code, r…

  331. Hi guys, so i have this pc for the gaming,7800x3d, AMD 9070xt with 16GB of vram and CORSAIR Vengeance RGB DDR5 32GB DDR5 6000MHz CL30 AMD Expo. Last week i was searching for good ai uncensored models on hugging face for my AIself-hosted on…

  332. Most AI coding agents assume you have a 200k-context model. In reality, the local models most people actually use have 8k windows — barely enough for one large file, let alone a whole project.

  333. I am using an Arc Pro B70 to do inference, and it's token generation speed is fine using Ollama, but it takes *forever* to do a prefill. vLLM absolutely tackles the prefill problem (nearly instant responses), but I can't run nearly as larg…

  334. I tried local model couple weeks ago. At the beginning, I tried Ollama, but reddit says better to switch to llama.ccp.

  335. When you load a model in these programs, you have to manually choose your context size or accept the default of 4096. In contrast, the newly released Unsloth Studio does not have this limitation, and VRAM/RAM is allocated as-needed so that…

  336. A high-performance local LLM server providing drop-in API compatibility with Ollama and OpenAI, built on llama.cpp's llama-server. Features automatic VRAM management, Hugging Face integration, and modular architecture.

  337. Hi everyone, I am testing the LocalLLaMA. I have a laptop with an RTX 5000 Ada generation, with Ollama and Open Webui.

  338. Saw a post about running hermes agent locally with gemma4 through ollama. zero api costs, unlimited tokens, full privacy.

  339. I was building a classifier to label AI agent sessions as productive or dead-end. The task isn't keyword matching, it's intent judgment: did the agent actually accomplish the goal, or did it get stuck retrying the same Cloudflare wall 20 t…

  340. been using local AI for a while now but my workflow was a mess. ollama for chat, comfyui for images, different tools for video and coding.

  341. I've been running into bottlenecks when trying to use multiple local LLMs on a single GPU. Currently switching manually between models in the terminal - starting/stopping Ollama, adjusting prompts, reloading contexts, etc.

  342. Estoy un poco nuevo con esto de la IA, estoy tratando de aprender lo que más puedo temas como: * Skills * Agends * Models * LLM * Ollama * llama.cpp * Cuantizacion Pero estoy aún perdido, tengo en mi PC 32Gb de ram y quisiera ejecutar mode…

  343. Hello I saw the new model is out but even with 24gb of vram, I have too many browser and task to use it , so I have downloaded and tested the version of HauHauCS https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressiv…

  344. localmind Run any local LLM with persistent memory and context. A single CLI binary that turns an Ollama-served model into an interactive agent with long-term recall, learnable skills, and permissioned tools.

  345. I'm trying to get past generic "best model" recommendations and collect real-world configs from people on similar hardware. My setup: MacBook M1 Pro, 10-core CPU, 14-core GPU, 16 GB unified memory.

  346. Been vibe coding a lot recently and kept running into the same problem finding actually usable tools without paying for 10 different subscriptions or donating my bank balance to Claude. So I put together a curated list focused on free or l…

  347. I think this should be a crime (at 3:00)

  348. I currently use Haiku 4.5 in an automated content workflow. The process works like this: I take an existing article from my website, use a DataForSEO node to fetch competitor URLs and search intent data, and then generate a new article com…

  349. Runs entirely on your machine. No API calls to any eval service.

  350. Hi, I got new MB Pro 24GB/1TB. I've test Gemma 4 26B with ollama, 16k context.

  351. Claude has had a rough week. Between the outage and the usage limit threads, I figured it was actually good timing to do something I had been meaning to try anyway: take the D&D skill I built a few weeks ago and see if I could migrate it t…

  352. Ollama Start building with open models. Download macOS curl -fsSL https://ollama.com/install.sh | sh or download manually Windows irm https://ollama.com/install.ps1 | iex or download manually Linux curl -fsSL https://ollama.com/install.sh…

  353. CATAI Virtual desktop pet cats for macOS — pixel art cats that live on your dock, chat with you via Ollama LLM, and debate ideas together to help you brainstorm and refine your thoughts. Features Dock companion — Cats walk along your dock…

  354. So I already have a Mac Studio M4 Max (return window still available)with 64GB RAM, but I’m eyeing the Corsair AI Workstation 300 (Ryzen AI Max+ 395, 96 VRAM out of 128GB, $3,250). Both seem decent for running models locally with Ollama.

  355. Hey r/LocalLLaMA, I wanted to share a project I've been working on called Job Bro. It’s a Chrome extension designed to help you analyze LinkedIn job descriptions without feeding your resume or career data into a proprietary black box if yo…

  356. As part of a personal project, i decided to build an AI assistant which helps with coding and homelab management. I really tried to make it as private as possible with local AI models running through Ollama.

  357. Hey yall, I want to explore more models and stuff i can do with them. What do you recommend?

  358. Hello, everybody! I'm building and hybrid database with Qdrant and Neo4j for a few personal projects.

  359. Hey all. This just got delivered yesterday.

  360. I've been running 5-8 Claude Code sessions at a time and got tired of tab-switching to approve tool calls. So I built claudectl — a TUI that sits on top of all your sessions and lets a local LLM (ollama/llama.cpp) handle approvals for you.

  361. Book Translator Translate long-form text files through a local Ollama-powered desktop and web app. Book Translator provides a two-stage workflow for translating books and large documents: first it generates a draft translation, then it run…

  362. I hope sincerely someonecan help me because i have tried everything i can and i get this speed using ollama.cpp and opencode. I have put as detail i can my setup and how i am running it.

  363. I've been building this for a couple of years. It started as "what if my AI assistant actually remembered things," and it became something bigger.

  364. Hi all, I've been using Claude now every day for a while. Some coding, firmware tweaks, help with complex github instructions or complicated tasks.

  365. Friends Don't Let Friends Use Ollama Ollama gained traction by being the first easy llama.cpp wrapper, then spent years dodging attribution, misleading users, and pivoting to cloud, all while riding VC money earned on someone else's engine…

  366. Got an ASUS NUC15 specifically for running Qwen locally on the Arc GPU. The marketing promised AI-ready performance.

  367. I spent time setting up TinyGPU on an Apple Silicon Mac and comparing it against Ollama already installed locally. Short version: TinyGPU does work.

  368. Deskdrop []() Deskdrop is an Android keyboard with AI built in. Use Ollama, any OpenAI-compatible server, or a cloud API key.

  369. Hey, so I'm pretty new to Local model hosting and have been messing with it a bit. I'm not a SWE but am reasonably technical.

  370. Hi everyone, I'm trying to run a local LLM via Ollama on a Hetzner cax21 VPS (ARM64, 4 vCPUs, 8GB RAM, 80GB SSD). I have Ollama running successfully via Coolify.

  371. Hi. I've been looking at ollama cloud's Pro offering ($20), which says "Run 3 cloud models at a time".

  372. I'm brand new to local LLMs and started with GLM-4.7 Flash q4_K_M. When I run it directly: ollama run glm-4.7-flash:q4_K_M it works pretty decently — nothing amazing, but usable and responsive.

  373. After months of testing, I finally have a local setup that doesn't make me want to go back to the API. Hardware: RTX 3090 (24GB VRAM) Models tested: Qwen2.5-Coder 32B Q4_K_M, DeepSeek-Coder-V3 Q4, Llama 3.3 70B Q3_K_M Inference: llama.cpp…

  374. AI writing feels "generic" because it lacks a feedback loop and social pressure. To fix this, I built an experimental system where AI agents participate in a literary circle.

  375. Been trying to run local LLMs on my new Dell XPS 13 with Intel Arc 140V (Lunar Lake, 16GB) and hit a wall — Intel's official docs point to a portable zip frozen at Ollama v0.5.4 which can't pull any modern model. Spent a while debugging it…

  376. I am in the midst of a POC project at work and am I have is 4 AMD Epyc cores and those are essentially virtualized. Does any one have any tricks?

  377. Sorry, not so tech person. I’m trying to figure out the most practical local LLM setup using my spare machine: 4 GB RAM No GPU for now, so please assume CPU-first unless I mention otherwise.

  378. ​Hey everyone, ​I’ve finally got my local media stack on my NAS migrated over to a new Mini PC running WSL2, sperately I have running my main gaming rig. now wnat to delve into the world of local AI models.

  379. This has been a pain point for many, and I've seen some tools to address it, but they needed a lot of setup. So made this GUI tool with AI assist.

  380. Hello, je suis en plein dans le montage d'une solution IA locale pour virer à terme perplexity, l'usage de chatgpt, claude etc..... mais je ne suis pas informaticien (perplexity est encore mon amie en ce moment !).

  381. First of all, I am a super newbie at local AI. Recently I got a GMKTek Evo X2 96GB to replace Claude as the usage limits have gotten unusable.

  382. I started developing an app with Claude, but the credits run out very quickly. I thought that now with my new computer I could run something directly on it.

  383. Hey everyone, I'm a final year engineering student building a 3-agent LLM platform (Researcher, Writer, Validator) for my end-of-studies project. My setup: RTX 4050, 6GB VRAM 16GB RAM Running Mistral 7B via Ollama locally The problem: My s…

  384. I've been on a hunt for a browser agent that can reliably handle daily agentic tasks: filling job applications, logging into sites and fetching data, making posts on my behalf, solving assignments and reporting results, and API/troubleshoo…

  385. Fleet Watch Process governance for AI workloads on a single machine. The Problem You're running MLX, Ollama, vLLM, Candle/Cake, experiment runners, and AI coding agents on the same machine.

  386. I am open to any option whether it's local or service based. For online services I tried Chatgpt agent : it's almost the worst option ever.

  387. Scryptian v0.1. (Proof of Concept) Local AI-powered command bar for Windows & Linux.

  388. ​ Turned a Xiaomi 12 Pro into a dedicated local AI node. Here is the technical setup: ​OS Optimization: Flashed LineageOS to strip the Android UI and background bloat, leaving ~9GB of RAM for LLM compute.

  389. I created and run a benchmark for AI models in data analysis tasks. In contrary to other benchmarks, it is not one-prompt benchmark, but I tried to simulate the real work of data analyst.

  390. ThinkReview: AI Code Review for GitLab, GitHub, Bitbucket & Azure DevOps Overview AI Copilot & AI Code Reviews for GitLab, GitHub, Bitbucket and Azure DevOps PRs - Ollama support 🌟 Now Open Source! View our code on GitHub: https://github.c…

  391. Hi everyone, Been following a lot of local LLM talk in this forum lately—learned quite a bit from you all! This is my first post, hopefully not my last.

  392. I spent about a week testing open-weight models for real work, comparing them against what I already know from ChatGPT, Gemini, and Claude. The gap between what benchmarks suggest and what happens when you give these models something to ve…

  393. I've been running Mistral/Llama locally through Ollama for a while now and the thing that keeps bugging me is context. The model itself is fine for general stuff but the second I want it to know about my projects, my notes, or files it doe…

  394. I'm trying to replace openai codex which i used for development all the time, with gemma4 on 4090, small tasks it solves quite impressively, but i need to have some agent. So I tried to connect 31b to cline and to aider and it didn't reall…

  395. What am I doing wrong here? I can't get models to follow my instructions, pretty much at all.

  396. I need help. I want to self-contain my MiniMax 2.7 and Qwen 3.5 (122 billion parameter) models.

  397. Hi everyone, I’m building a local AI pipeline on WSL2 (Ubuntu) specifically for Product Visualization. My goal is to orchestrate LLMs for scene generation and Stable Diffusion/ComfyUI for high-fidelity rendering, keeping my Windows host cl…

  398. I vibecoded a Flask app that acts as a Game Master for my day. I feed it my goals, and a local AI looks at my past history to generate new "quests".

  399. The introduction of TurboQuant, PolarQuant, and QJL (Quantized Johnson-Lindenstrauss) by Google Research represents more than just a technical optimization. At Vucense, we view this as a landmark moment for Inference Sovereignty https://vu…

  400. I have made a chrome extension that lets LLMs control your browser - clicking, typing, navigating, etc. supports ollama/openai/anthropic/google looking for people to try it out and let me know what breaks.

  401. I'm an ex crypto miner with remnant mining parts so I threw them together into a franken hydra case. I've been using claude oath previously, but they just shut that door last week or so.

  402. Deploy a complete local AI stack — Ollama 5.x, Open WebUI, and pgvector — on Ubuntu 24.04. Zero cloud.

  403. Hey everyone, I'm comparing these two plans side by side for running AI agents daily through OpenClaw (self-hosted AI agent platform): • Ollama Cloud Pro — $20/month • OpenAI Plus — €23/month (~$25) My setup: 3 agents running in parallel (…

  404. I'm a beginner using local models, now I have a good GPU I installed ollama using docker. Pulled the Gemma4 weights and was able to add it to cursor using ngrok.

  405. Hi, I'm looking to replace my current 2x ChatGPT Plus subscriptions with one $100 subscription of either Ollama Cloud or Claude Max, and would appreciate some insights from people who have used these plans before. I've had 2 $20 ChatGPT su…

← all threads