Models & local inference
Running and selecting open models in Ollama, especially Gemma, Qwen, Llama, DeepSeek, Kimi, and specialist vision/OCR models.
46%
Best tweets about Ollama
Find the best tweets about Ollama, including local model setup, performance, hardware, integrations, model files, and developer workflows.
Hands-on Ollama setup, local inference, supported models, integrations, performance, troubleshooting, and releases.
Original Xholic analysis
Across 50 tweets, discussion is mostly supportive (86%) and positive (78%). The largest theme is models and local inference (46% of tweets), followed by agent and developer integrations (36%). The cited posts concentrate on practical local setups, compatibility layers, and local-first applications, while performance and hardware claims vary by model, backend, and platform.
78% of posts
All-time engagement
52% of posts
Published in 90 days
Conversation map
Running and selecting open models in Ollama, especially Gemma, Qwen, Llama, DeepSeek, Kimi, and specialist vision/OCR models.
46%
Connecting Ollama to coding tools, agent runtimes, MCP, Anthropic/OpenAI-compatible APIs, and developer workflows.
36%
Hardware fit, memory requirements, quantization, backend choice, and speed/latency improvements across Apple Silicon, GPUs, and cloud hardware.
26%
Privacy, ownership, offline resilience, and cost-control motivations for replacing cloud AI services with local Ollama deployments.
24%
Installing, configuring, and self-hosting Ollama for offline local inference, including commands, endpoints, Docker, and remote access.
24%
Document, image, and multimodal workflows using Ollama, including OCR, extraction, image generation, visual agents, and local embeddings.
18%
Ollama releases, platform capabilities, cloud offerings, funding, compatibility changes, bugs, and security issues.
18%
Local-first applications built on Ollama for assistants, desktop automation, chat UIs, private memory, monitoring, and offline knowledge systems.
12%
Tone and stance
Performance benchmark
Posts with media make up 72% of this collection. Their median all-time score is 13.9, compared with 2.05 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts with concrete examples present Ollama as part of a local workflow: a Gemma setup, Anthropic Messages API compatibility for Claude Code-style workflows, and a Discord bot walkthrough using Gemma, OpenClaw, and Ollama.
Shared view
Integration examples focus on configuration and compatible interfaces: a localhost Anthropic-base-URL setup, Codex configured with Gemma through Ollama, and a chat UI that lists Ollama among supported OpenAI-compatible endpoints.
Shared view
Several application posts emphasize local operation: CatchMe is described as offline with Ollama, Observer says screen-monitoring models can run through Ollama or llama.cpp without a cloud API seeing the screen, and World Monitor says its summaries run locally through Ollama without API keys.
Open debate
Posts differ on the practical boundary between local and cloud use. One argues that running GLM-5 locally requires expensive hardware and favors Ollama Cloud; another proposes Gemma 4 through Ollama for routine OpenClaw work while retaining Claude Opus for complex work. A third reports plans to compare OpenClaw with 4o mini and local Qwen 3B.
Open debate
Backend reports are platform-specific and not fully aligned. One post ranks Ollama below llama.cpp and vLLM for its intended setup, Ollama’s post says its MLX update makes it faster on Apple Silicon, and another reports GPU-detection problems with Ollama on DGX Spark.
What performs
The five deterministic engagement outliers are an Apple Silicon update, a Gemma local-run guide, Anthropic Messages API compatibility, Codex configuration with Ollama, and a Discord setup walkthrough. Four are tutorial or integration-oriented; the Apple Silicon post is an announcement.
Hardware-fit posts repeatedly foreground memory, quantization, and model choice. One author recommends a 64GB Mac Mini M4 Pro, a hardware-scanner post describes ranking models and quantizations against local specifications, and another reports running Qwen 3.5 9B through Ollama on a 16GB GPU.
The cited speed figures and claims cover different contexts: Hermes Agent startup changes involving Ollama probes, Ollama Cloud throughput and latency on B300 hardware, and an Apple Silicon MLX comparison. They should be read as context-specific reports, not a single benchmark.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Alex Veremeyenko
@alex_verem
2 posts
2. AlphaSignal AI
@AlphaSignalAI
2 posts
3. BowTiedCyber | Evan Lutz
@BowTiedCyber
2 posts
4. Fahd Mirza
@fahdmirza
2 posts
5. Ihtesham Ali
@ihteshamali
2 posts
6. Julian Goldie SEO
@JulianGoldieSEO
2 posts
Ollama’s two cited posts announce platform changes: MLX-backed Apple Silicon performance work and B300 cloud hardware for Kimi K2.5 and GLM-5. The latter post also says Ollama integrations can use its launch command and GitHub integrations.
Alex Veremeyenko’s two cited posts cover applied local systems: Observer for monitoring inputs such as screens and cameras, and World Monitor for locally summarized situational-awareness feeds.
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Ollama tweets
Ranked 01–50
@ollama ·
Ollama is now updated to run the fastest on Apple silicon, powered by MLX, Apple's machine learning framework. This change unlocks much faster performance to accelerate demanding work on macOS: - Personal assistants like OpenClaw - Coding agents like Claude Code, OpenCode, or Codex
@EXM7777 ·
here's how to run Gemma 4 locally in under 5 minutes: option 1 (phone): > download Google AI Edge Gallery from the Play Store > select Gemma 4 E2B or E4B > it downloads and runs entirely offline > no account, no API key, no internet needed option 2 (laptop): > install Ollama or LM Studio > pull gemma-4-27b (the MoE version, only 3.8B active params) > runs on a MacBook with 16GB RAM option 3 (developer): > open Google AI Studio > select Gemma 4 31B > use the function-calling API for agentic workflows > or deploy on Vertex AI for production the 26B MoE is the sweet spot for most people i think
@akshay_pachaar ·
this is huge. ollama is now compatible with the anthropic messages API. which means you can use claude code with open-source models. think about that for a second. the entire claude harness: - the agentic loops - the tool use - the coding workflows all powered by private LLMs running on your own machine.
@skirano ·
I love how open the Codex app is, you can run any model you want, even local. These are the only configs you need to use the app with Gemma 4 via Ollama.
@fahdmirza ·
Gemma 4 + OpenClaw + Ollama + Discord — Full Local AI Setup for Free 🔥 Google just dropped Gemma 4 and we wired it directly into Discord 🔹 Gemma 4 31B pulled via Ollama — completely local 🔹 Fresh OpenClaw install from scratch 🔹 Full Discord bot setup — Developer Portal, intents, OAuth2, permissions 🔹 OpenClaw + Discord pairing walkthrough 🔹 Chat with Gemma 4 directly from your Discord server 🔹 DuckDuckGo web search enabled — no API key needed Watch the full setup below 👇
@Axel_bitblaze69 ·
many of you keep asking me in the comments what machine to buy for AI I run Claude Code, Ollama with local models, MCP servers, Paperclip agents, Chrome automation, and way much more and screen recording all at the same time.. after testing everything tbh, the Mac Mini M4 Pro with 64GB unified memory is the best value machine for AI work in 2026. and it's not even close because, > 64GB unified memory = run Gemma 4 (26B), Llama, Deepseek, Kimi, Qwen locally without breaking a sweat > M4 Pro handles Claude Code + multiple MCP servers + Ollama + Chrome + terminal all running simultaneously > Silent. Tiny. Sits on your desk and just works. what you actually need for local AI: > RAM is everything. Models load into memory. 32GB = small models only. 64GB = 26B-34B models comfortably. 128GB = 70B+ models. > GPU cores matter for inference speed but RAM is the gate. No RAM = model doesn't load at all. > CPU matters less than you think. M4 Pro is more than enough. so the breakdown: - Mac Mini M4 Pro 64GB → best value, handles everything, my recommendation for 90% of people - Mac Studio M4 Max 128GB → if you want to run 70B models or you're doing serious video production alongside AI work - MacBook Pro M4 Pro 48GB → only if you need portability. For the same price you get way more in a desktop - MacBook Air → great machine but not enough RAM for serious local model work. Fine if you're only using Claude Code via API. stop overthinking it. If you're building with AI daily, Mac Mini. M4 Pro. 64GB. Done. now if only @Apple would send me one... Tim Cook if you're reading this I'm doing free marketing for you bro. hook me up. 🍎
@charliejhills ·
🚨Stop gambling on the wrong local LLM. Most people download a model, watch it crawl, then start over. LLM Hardware Scanner is a CLI tool that scans your actual specs and ranks hundreds of LLMs before you waste the bandwidth: ⤷ Reads your RAM, CPU, and GPU to score real compatibility ⤷ Accounts for MoE activation not just bloated total parameter counts ⤷ Auto-selects the optimal quantization for your exact setup ⤷ Ranks models across quality, speed, context length, and true memory fit ⤷ Works natively with Ollama, llama.cpp, and MLX ⤷ Shows a composite score plus a speed estimate before you download a single byte Not a benchmark site. Not a spec sheet. A hardware-aware ranker that tells you the truth about what will actually run on your machine. One scan. Instant ranking. Zero wasted downloads.
@ollama ·
Ollama's cloud is updated to use NVIDIA's latest data center hardware: B300 for Kimi K2.5 and GLM-5 models. This significantly improves the model performance with faster throughput and lower latency while maintaining reliable tool calls for integrations. All this works with Ollama's integrations via Ollama's launch command and over 45,000 custom integrations from GitHub.
@ihteshamali ·
Someone just built Windows Recall but its open source, local, and actually private. It's called CatchMe. It records everything you do on your computer and lets you search it in plain English. No cloud. No subscriptions. No data leaving your machine. Here's what it tracks silently in the background: → Every window you open and how long you're in it → Every keystroke, clipboard copy, and file you touch → Screenshots tied to every mouse action → App usage across your entire day → Notifications, file changes, and browser activity Then it organizes all of it into a memory tree Day → Session → App → Location → Action and lets you query it like a human. "What was I coding in Claude Code this morning?" "Which files did I edit yesterday?" "How did I spend my afternoon?" No vector databases. No embeddings. Just tree-based LLM reasoning that navigates your history and finds the exact answer. Works with Claude, Cursor, OpenClaw, NanoBot. Drop one skill file and your agent instantly has memory of your entire digital life. 0.2GB RAM. $0.42 per 2 hours of use with Qwen. Runs fully offline with Ollama. https://t.co/qRbPnf8fkR 100% Open Source. Apache 2.0 License.
@VaibhavSisinty ·
I just discovered the free version of Claude Code. It is called opencode and it is crazy good. Best part? Open source. And the most starred coding agent on GitHub right now. Install it in 60 seconds: → curl -fsSL https://t.co/2QSqHlbvg4 | bash That is it. One line. Done. Then this is what you get: → Runs in your terminal exactly like Claude Code. Read your codebase, edit files, run commands, manage git, chain multi-step tasks. Same workflow. → Bring your own model. Claude, GPT, Gemini, DeepSeek, Qwen, or run a local model via Ollama for true zero cost. 75+ providers supported. → Already paying for ChatGPT Plus, Claude Pro, or GitHub Copilot? Type /connect and that subscription powers the whole agent. No extra billing. → Two built-in agents: build mode for shipping code, plan mode for read-only exploration. Switch with the Tab key. Most developers have not heard of it yet. The ones who have are not going back.
@RaulJuncoV ·
The agent loop controls your entire stack. How your model sees tools. How it retries. How it handles context. Most developers can't touch it. You need hooks. The runtime doesn't expose them, so you wrap it. You swap providers. The abstraction leaks, so you patch it. Three months later, a model update breaks the wrapper. Your pipeline is down on a Friday. I've seen this pattern too many times. Cline just open-sourced the runtime behind Cline 2.0. It’s the same agent loop across VS Code, JetBrains, and the CLI Now you can install it, inspect it, fork it, and build on top of it. You get: → Lifecycle hooks for observability and policy → Provider routing across Anthropic, OpenAI, Bedrock, Ollama, and more → Sessions, scheduling, and native MCP support → A real plugin system for tools, hooks, commands, providers, and message builders The model can be replaceable, but the loop cannot be invisible if you want control. The box is open. npm i @cline/sdk 👉 https://t.co/sp97PAMOw6
@Sumanth_077 ·
Train LLMs locally without writing a single line of code! @UnslothAI just released Unsloth Studio - an open-source web UI for training and running models. Here's how it works: You upload a PDF, CSV, or DOCX file. The Data Recipes feature automatically transforms it into a structured training dataset via a graph-node workflow. No manual formatting needed. Then you select a model from Hugging Face or your local files. Pick your training method - LoRA, QLoRA, or full fine-tuning. The UI pre-fills sensible defaults based on your model. Start training and watch live metrics - loss curves, GPU usage, gradient norms. Everything runs locally with 2x faster training and 70% less VRAM than standard setups. Here are the key capabilities: • Chat with GGUF and safetensor models - supports tool calling, web search, and code execution in a sandbox. • Compare models side-by-side - load your base model and fine-tuned version to see how outputs differ. • Export to any format - save your trained models as GGUF, safetensors, or LoRA adapters for use with llama.cpp, vLLM, Ollama, or LM Studio. • Multi-modal support - train text, vision, audio, and embedding models all in one interface. It runs 100% offline on your hardware.
@alex_verem ·
Found an open source tool that watches your screen and pings you when something happens. It's called Observer. You build tiny AI agents that monitor your screen, camera, or mic with a local model, then react. Your training run crashes, you get a Telegram message. A dashboard error rings your phone with a text-to-speech readout. Sensors: screen, OCR, camera, mic, clipboard, meeting audio. Actions: email, Discord, Telegram, WhatsApp, phone calls, screen recording. The models run through Ollama or llama.cpp. No cloud API sees your screen. Setup: system prompt + sensor + a few lines of JS. One agent takes minutes. 1.6k stars. Point an observer at your screen and go do something else.
@tonysimons_ ·
Hermes Agent just got a serious speed injection. First-turn startup latency was cut by ~80%. Cold submit → request dispatch: 4.3s before 0.9s after The fix? @Teknium tracked the actual pre-request stalls and cut them out: 🔹 Discord capability detection moved off the blocking path 🔹 pointless Ollama probes skipped for known non-Ollama providers 🔹 Python env probing warmed off-thread 🔹 MCP imports skipped when there are no MCP tools 🔹 CLI pre-imports while you’re still typing Hermes keeps getting sharper. PR: https://t.co/6FYpEGQulR
@hasantoxr ·
Anthropic Fable 5 has been banned by the government. Learn to use local models so you have 100% control. Instead of arguing about why they banned it, I built a full guide on running AI locally so nobody can ever take it from you. Here's everything you need to know: 1. Pick your runtime first. Think of this like installing a video game launcher before you can play any games. Ollama and LM Studio are the two launchers for AI. Download one. That's step one. 2. Understand model size. The number next to a model like 7B or 32B just means how many things it learned. Bigger number means smarter but needs more memory. A 7B runs on any laptop you already own. Start there. 3. Pick the right model for the job. Qwen 3 for everyday tasks and writing. DeepSeek for math and coding problems. Gemma 3 when your computer is slow. Llama when you want the most help from the internet because millions of people use it. 4. Learn quantization. This sounds scary but it's just shrinking. Like zipping a file. A huge model gets compressed so it fits on your laptop with almost no quality loss. Look for Q4 or Q5 in the model name. That's the compressed version. Download that one. 5. Add tools to your model. A small model with Google search and file access beats a big model with nothing. Think of tools like giving the AI hands. Without them it can only think. With them it can actually do things. 6. Watch your context window. Cloud AI has unlimited memory per conversation. Local AI does not. The longer your chat gets the more your laptop slows down. Keep conversations short. Start a new chat when things feel slow. You never needed anyone's permission to use AI. The ban only hurts people who never learned to run it themselves.
@itsharmanjot ·
GitHub is shutting down its entire AI playground on July 30, 2026. Playground. Model catalog. Inference API. BYOK. All of it. Gone for every customer including people with active usage right now. What GitHub Models was: Free access to Llama 3.1, GPT-4o, Mistral, Cohere directly inside GitHub. No separate account. No setup. Zero friction prototyping since 2024. Where GitHub is pointing you instead: Azure AI Foundry. Microsoft's paid platform. If you have pipelines calling the inference API they break July 30. Find them now while access still works. GitHub scheduled brownouts July 16 and July 23 as live tests. Use them. What to actually use instead: For prototyping locally → Ollama. One command, any open-weight model, runs on your hardware, no cost, no rate limits, no shutdown risk. For a ChatGPT-style interface on top → Open WebUI. One Docker command. For a desktop GUI with no terminal → LM Studio. For API-dependent production workflows → OpenRouter. One endpoint, dozens of models, provider-portable. The pattern worth naming: Free tool lowers the barrier. Grows the user base. Gets retired toward the paid platform. This is not a GitHub problem specifically. It is the standard enterprise playbook for developer tools. Local models running on your own hardware cannot be retired by someone else's changelog post. 21 days left. If anything in your stack touches GitHub Models, the time to find out is now.
@alex_verem ·
World Monitor is an open-source situational awareness dashboard that puts military, economic, climate, and financial signals in one view. It ingests 500+ news feeds across 15 categories and summarizes each item with AI as it arrives. The feeds cover military movements, economic shocks, natural disasters, cyber incidents, flight paths, and shipping lanes. The main view is a 3D globe with 56 map layers you can stack. A Country Instability Index scores 31 nations on a stress metric and updates as events happen. A finance panel covers 29 stock exchanges, commodities, and crypto. Summarization runs locally through Ollama, so you don't need API keys. Events from different streams appear on the same map. When a shipping lane closes, a commodity spikes, and a conflict escalates in the same region, you see all three together. One codebase produces six site variants: world, tech, finance, commodity, energy, and happy. The desktop app runs on Tauri 2 for Windows, macOS, and Linux. It supports 25 languages, with native-language feeds and RTL layouts. For programmatic access, there's an MCP server, a REST API, a CLI, and SDKs for Python, Ruby, and Go. Scripts and agents query the same data the interface shows. Elie Habib has maintained it since 2024: 4,491 commits, 43 releases. The repo has 61,300 stars. License is AGPL-3.0, free for personal, research, and self-hosted use. The code is on GitHub. Clone it and run it.
@VaibhavSisinty ·
I've been saying this for a while now. The future isn't one massive model sitting in the cloud doing everything for you. It's tiny specialist models running on your device. Doing 80% of tasks locally. A small router model deciding which model handles what. And only pinging the cloud when the task genuinely needs a bigger brain. People called this wishful thinking. Then China dropped GLM-OCR this week. 0.9 billion parameters. That's nothing. Practically runs on a potato. And it just became the #1 document reading model in the world. Better than Gemini. Better than models literally 100x its size. The trick is simple. It doesn't read a document top to bottom like most models. It breaks the page into regions first, then reads all of them at the same time. Predicts multiple tokens per step. Small model. Parallel processing. Insane speed. Tables, handwriting, math equations, messy layouts. It handles all of it. Runs locally through Ollama. Open source. Free. This is exactly the kind of model that fits into the future I'm describing. You don't need GPT-5 to read a receipt. You need a tiny model that does one job really well, sitting on your phone, costing you nothing. One model to rule them all is dead. Swarms of small specialists is what's coming.
@ayushagarwal ·
contextmcp v0.5.0 just shipped. contextmcp is our open-source MCP server that indexes your documentation and serves it as context to AI agents. point it at your docs repo, it chunks, embeds, and gives your agent the right documentation when it needs it. what's new: → Ollama support. fully offline, no API key, zero cost local embeddings. also added Cohere and Voyage AI as providers → GitLab source. index docs from gitlab or self-hosted instances. not just GitHub anymore → validate and doctor commands. catch misconfigs before you waste time on a full reindex all backward compatible. 50+ stars. @dodopayments
@Shruti_0810 ·
Someone just broke the Claude paywall… for real. Claude Code → $0 API key → $0 usage cost Here’s what they built: A tiny proxy that hijacks Claude Code’s API and reroutes it to free + local models • NVIDIA NIM (free tier) • OpenRouter (hundreds of models) • DeepSeek (Anthropic-compatible) • Or run it fully local (Ollama / llama.cpp) Set 2 env vars → done. No CLI tweaks. No extensions. It just works. But here’s the crazy part: – Route Opus / Sonnet / Haiku to different providers – Converts raw model output → real Claude thinking blocks – Auto-detects + executes tool calls – Built-in rate limiting + backoff – Even Discord & Telegram bots for remote coding Open-source. MIT license. This isn’t a hack. It’s a replacement layer for Claude itself. Repo: https://t.co/2B8WUZUHQg
@ihteshamali ·
A community of 300 anonymous developers built the AI companion Character AI's investors are afraid of. It's called SillyTavern. This is the app Character AI users switched to when the platform banned NSFW content in 2024, Replika users moved to after the company deleted intimate memories overnight in 2023, and OpenAI blocks with content filters. It runs on your own laptop. Nobody can shut it down. Here's how it works. SillyTavern is a chat window. It does not run the AI itself. It talks to a model running on your own computer through a free app called Ollama or KoboldCpp. The model lives on your graphics card. Every message stays on your hard drive. Nothing touches the cloud. You install Ollama. You download an uncensored model like Mistral or Llama 3. You point SillyTavern at it. You now have a Character AI clone with no rules running on hardware you already own. The character is a single PNG file: Personality, backstory, voice, and sample dialogue all sit embedded inside the image. You drag the file into the app and it becomes a character with memory. There are thousands of free character files on chub. ai covering every genre people have built. → Character files with full personality, memory, and dialogue history → Lorebooks that inject world details when specific words appear in chat → Group chats where multiple characters talk to each other and you → Voice generation, portrait generation, expression sprites built in → Long memory that summarizes old conversations automatically → Works with every open source model on your machine → Also works with GPT-4, Claude, and Gemini if you want to pay → Every conversation stored as a file on your machine, forever Character AI raised $150 million from Google. Replika has 30 million users. Both companies have deleted memories, rewritten personalities, and changed the terms of service on relationships people built for years. SillyTavern is the version of this software where the character belongs to you. The model belongs to you. The conversation belongs to you. Started as a fork in 2023. Now on 30,000 GitHub stars. Built by 300 volunteers who have never met. The company version can be shut down. This one cannot.
@GithubProjects ·
Chat UI is a SvelteKit chat interface that works with any OpenAI-compatible API, powering HuggingChat at https://t.co/3Q2b7hKvgg. - Connects to any OpenAI-compatible endpoint via OPENAI_BASE_URL and /models - Supports llama.cpp, Ollama, OpenRouter, and the Hugging Face Inference Providers router - Persists chat history, users, and settings in MongoDB with an embedded fallback - Runs as a local dev server with npm install and npm run dev Explore it here: https://t.co/dvzEIES28e
@_vmlops ·
OLLAMA-OCR TURNS YOUR SCANNED DOCS INTO CLEAN MARKDOWN built on top of ollama's local vision models, no cloud APIs, no OCR subscriptions → swap between llava, llama 3.2 vision, granite3.2-vision, moondream, minicpm-v depending on speed vs accuracy needs → output as markdown, plain text, JSON, tables, or key-value pairs → batch process entire folders in parallel with progress tracking built in → built-in image preprocessing before it even hits the model → streamlit web app included if you don't want to touch code pip install and you're extracting text in minutes. 2.3k stars, actively maintained your documents never leave your machine.
@socialwithaayan ·
you can now hand your entire to-do list to an AI coworker that runs on your laptop and actually does the work 🤯 it's called OpenWorker. andrew ng built it. the man who founded google brain, ran AI for 1,300 people at baidu, and co-founded coursera. this is not another chatbot. you tell it "prepare a customer brief" or "triage my inbox" and it comes back with the deliverable, done. why it works: it breaks the task into steps, executes across your desktop, files, and connected apps, and before it does anything consequential like sending a message or changing a calendar, it stops and asks. you approve or redirect. then it finishes. what you get for $0: → multi-step task execution across your apps → bring your own key for openai, anthropic, google, or any open-weight model → fully local with ollama if you want zero cloud → macos and windows → 7.5k stars and 1,000+ forks in a week what it replaces: → copying chat output into separate apps yourself → paying $20-30/month per user for Copilot → the "great suggestions, now let me do all of it manually" loop how to set it up: 1. download from the github repo (macos or windows) 2. open the app, add a model key (openai, anthropic, google, any open-weight provider) 3. or point it at ollama for fully local, zero cloud 4. ask for something real the honest part: it's in open beta. the approval gate before consequential actions is the safety net, but messy workflows across 6 apps will still need a human eye on every checkpoint. 7.5k stars, MIT license, free and open source.
@smratitiwa86867 ·
Google just open-sourced a tool that could replace an entire category of document extraction software. Content: It's called LangExtract. An open-source library designed to turn messy, unstructured documents into structured data—with source references you can verify. What it can do: → Extract structured information from plain text → Link every extracted item back to its exact location in the original document → Process long documents, including 100+ page files → Generate interactive HTML for reviewing results → Work with Gemini, Ollama, and other compatible models It can be useful for tasks like: → Processing clinical notes → Reviewing legal documents → Extracting data from financial reports → Converting large text documents into structured datasets Instead of relying on: → Complex regex rules → Custom NER pipelines → Paid document extraction APIs → Manual copy-paste workflows You define the extraction task with a few examples, point it at a document, and it returns structured, traceable results. No fine-tuning. No complicated setup. Just another example of how open-source AI tools are making advanced document processing more accessible. Repo 👇
@thetripathi58 ·
You need an AI pair programmer to write boilerplate, debug logic, and speed up execution. What are your options? GitHub Copilot sends your proprietary codebase to Microsoft servers to feed their massive models. ChatGPT requires you to paste your sensitive intellectual property directly into a public web interface. Cursor forces your entire development workflow through a centralized cloud proxy. In 2026, writing code at speed requires either paying permanent rent for cloud subscriptions or surrendering your company's intellectual property to corporate training engines. Someone built an architecture to bypass this entirely. It acts as your own private AI pair programmer. No data scraping. No cloud dependency. It is called Continue. Install the extension directly inside VS Code or JetBrains. Connect it to a local model running on your own hardware via Ollama. You get lightning-fast autocomplete and chat that never touches the internet. The technical leverage: - Absolute sovereignty. Your proprietary code, logic, and API keys never leave your physical hard drive. - Zero subscription fees. You stop paying a monthly tax to Microsoft just to generate basic functions. - Deep local context. It indexes your exact workspace securely, understanding your specific architecture without uploading it to a third-party server. - Offline execution. You can generate code on an airplane with zero latency. The corporate system wants you renting access to basic intelligence. They build artificial paywalls around code generation to extract recurring revenue forever. Continue tears down the AI monopoly. It is open-source, highly performant, and keeps your intellectual property exactly where it belongs: entirely under your control. Stop playing by their rules. Stop leaking your codebase. Build your own leverage and direct your own reality.
@metasploit ·
Latest Metasploit update is out with unauthenticated RCE for Grandstream GXP1600 VoIP devices, enabling credential harvesting and SIP interception. Also included is critical support for BeyondTrust PRA/RS command injection (CVE-2026-1731), plus a serious Ollama RCE (CVE-2024-37032). Check out the wrap up at https://t.co/9MXkxegCdd
@AlphaSignalAI ·
Stop downloading LLMs your machine was never going to run. llmfit scans your hardware and tells you exactly which models will run. It scans your RAM, CPU, GPU, and VRAM first. Then it scores every model across four dimensions: 1. Quality, based on parameter count and quantization 2. Speed, estimating tokens per second for your exact backend 3. Fit, matching memory use to your hardware 4. Context window support for your use case Each model gets a label: Perfect, Good, Marginal, or Too Tight. It picks the best quantization automatically, stepping down until something fits. Covers hundreds of models from Meta, Mistral, Qwen, and DeepSeek. Works with Ollama, llama.cpp, MLX, and LM Studio out of the box. Open-source.
@NainsiDwiv50980 ·
Your agentic AI product can earn its first dollar before it generates its first model API bill. Not a toy chatbot. A real system that retrieves knowledge, makes decisions, calls tools, takes actions, retains state, and traces what happened, running on a stack that costs exactly $0 to start. Here's the full pipeline: → Interface: Next.js or Streamlit takes the request → Orchestration: LangGraph or CrewAI decides whether to answer, retrieve, call a tool, ask for approval, retry, or stop → Knowledge: LlamaIndex pulls context from ChromaDB or Qdrant, running locally → Reasoning: Ollama runs an open model (Gemma, Llama 3.3 70B, Mistral) on your own hardware, zero API bill → Action: MCP connects the agent to files, databases, GitHub, Slack, browsers, this is the step where a chatbot becomes a worker → State: SQLite or DuckDB stores conversations, checkpoints, outputs → Observability: Langfuse or Phoenix traces every prompt, decision, tool call, and failure → Deployment: Docker packages it, inference runs on hardware you already own, only the lightweight interface sits on a free tier Total software and API spend to get this running: $0. But here's the part that actually matters, the free tools aren't the advantage. Every single piece here will eventually get replaced by something faster or cheaper. Ollama becomes a hosted API. SQLite becomes a production database. Streamlit becomes a custom app. The advantage that survives every swap is knowing: → Where reasoning should stop and deterministic code should take over → When RAG actually improves an answer versus just adding latency → Which actions genuinely need human approval → What has to be traced before your first production failure, not after → How to isolate every layer behind a replaceable interface Build the cheap version first. Learn exactly where users find real value. Then spend money only where it creates leverage, not before. If you had $500 a month to put into scaling this stack, which layer gets it first?
@RoundtableSpace ·
Ornif 1.0 is a free local model running through Ollama that's reportedly beating models 10x its size. It writes its own plan before coding, then grades itself on both the strategy and the output.
@RoundtableSpace ·
15 minutes of Kimi Code beats 10 hours of reading docs. Covers Claude Code with Ollama, web search, scheduled tasks, Telegram integration, and headless mode. The docs never had a chance.
@SaiyamPathak ·
Ollama just replaced launched 0.19 with Apple's MLX framework on Apple Silicon. The result? ~2x faster inference as per there test on M5 I tested it on my M1 Max the difference is real. New video breaking down: → What MLX is and why it's faster → UMA explained → Prefill vs Decode → NVFP4 quantization → Real benchmarks (MLX vs llama.cpp) Watch full video 👇
@JulianGoldieSEO ·
𝗚𝗼𝗼𝗴𝗹𝗲'𝘀 𝗚𝗲𝗺𝗺𝗮 𝟰 𝗿𝘂𝗻𝘀 𝗳𝗿𝗲𝗲 𝗼𝗻 𝘆𝗼𝘂𝗿 𝗹𝗮𝗽𝘁𝗼𝗽 𝗮𝗻𝗱 𝗿𝗮𝗻𝗸𝗲𝗱 𝗻𝘂𝗺𝗯𝗲𝗿 𝟯 𝗶𝗻 𝘁𝗵𝗲 𝘄𝗼𝗿𝗹𝗱 𝗼𝗻 𝗮𝗻 𝗼𝗽𝗲𝗻 𝗺𝗼𝗱𝗲𝗹 𝗹𝗲𝗮𝗱𝗲𝗿𝗯𝗼𝗮𝗿𝗱. No subscriptions. No data leaving your machine. No internet needed once it's set up. Here's which model to pick and how to run it: → Standard laptop: ollama run gemma4:e4b (9.6GB, runs at 4B during inference) → Higher-end workstation: ollama run gemma4:26b (18GB, fast and high quality) → Maximum quality: ollama run gemma4:31b (20GB, ranked #3 globally) Settings most people skip that actually matter: → Temperature 1.0, top_p 0.95, top_k 64. Google recommends these for all Gemma 4 models. Use them. → Enable thinking mode for complex tasks. Include the think token at the start of your system prompt. Turn it off for simple fast answers. → Put your image before your text in every prompt. Better performance according to the official docs. → For image tasks: use lower token budgets for classification, higher budgets for OCR and document reading. → Don't feed thinking output back into your conversation history. Only the final response. Make sure Ollama is on version 0.20 or higher before pulling. Gemma 4 requires it. Apache 2.0 license. Commercial use. No restrictions.
@DomJoLuna ·
Everybody's panicking about Anthropic cutting off OpenClaw today. A lot of people are switching to OpenAI as their default model, and to be frank, that's a downgrade you don't need to make. Here's what's actually happening…Anthropic is separating subscription limits from third-party tool usage. Your Claude subscription still works. Your Claude account still works. You just need to enable "extra usage" in your account settings (it's a pay-as-you-go option billed separately). They're even giving you a one-time credit equal to your plan cost and offering up to 30% off on pre-purchased usage bundles. So before you rip out the brain that actually feels like a thinking partner and replace it with something that reads like a corporate memo, just turn on extra usage. It takes 30 seconds. But here's the bigger play that nobody's talking about (cause AI releases move at light speed nowadays) Google dropped Gemma 4 two days ago under Apache 2.0. Fully open, fully commercial, zero restrictions. The 26B Mixture-of-Experts model is currently ranked #3 open model in the world, and it only activates 3.8 billion parameters during inference. That means it runs on a Mac Mini. Read that again. A top-3 open model running locally on a $600 machine. We're deploying it across our compute cluster today. Here’s the setup: → Ollama + Gemma 4 26B MoE for all routine work (heartbeats, task execution, research, monitoring) → TurboQuant KV cache compression to keep memory tight on 16GB machines → Claude Opus stays as the executive brain for complex strategy and conversation → Projected savings: ~$2k+/month by moving operational workload off API If you're running OpenClaw on a Mac Mini or any Apple Silicon machine, you can do the same thing right now: ollama pull gemma4:26b Set it as your primary model in OpenClaw, keep Claude as your fallback for the conversations that matter, and your monthly API bill drops to almost nothing for routine work. The best part? Gemma 4 has native function calling, 256K context, and system instruction support built in. It's not a toy, it handles OpenClaw's tool chain natively. Don't downgrade your AI partner because Anthropic changed a billing policy. There are better moves on the board.
@AlphaSignalAI ·
Someone open-sourced a survival computer that works even when the internet goes down. Project N.O.M.A.D turns any Linux machine into a fully offline knowledge server. No internet needed after setup. The AI runs through Ollama. You can chat, write, and code locally. Upload your own documents and search them with semantic search. Everything runs in Docker containers, managed through a browser: > Local AI with document search > Offline Wikipedia and reference books > Offline maps via OpenStreetMap > Khan Academy with progress tracking > Encryption and data analysis tools One terminal command installs everything. It just works offline. No cloud, no accounts, no data sent anywhere.
@cleanunicorn ·
I am a total self hosting nerd. Currently I am away from my computer, but I still have access to my servers. I can run consumer grade (not mobile grade) LLMs on my local setup using: - Tailscale - for access (VPN) - Open WebUI - chatGPT-like interface for LLMs - Ollama - LLM model runner and organizer Is anyone interested in a how to guide to set this up?
@fahdmirza ·
💥 Hermes Agent is the AI that gets smarter every session ⚕ ♠ and it's completely changing what a local self-improving agent looks like 🚀 🔹 Built by Nous Research — the lab behind some of the most capable open source models 🔹 Closed learning loop — creates skills from experience and improves them during use 🔹 Runs fully local with Ollama — no API keys, no cloud, no lock-in 🔹 Talk to it from Telegram while it works on your server 🔹 29 tools · 90 skills · out of the box on first run 🔥 Full setup guide below 👇
@WesRoth ·
Ollama has officially added image generation capabilities now live for macOS, with support for Windows and Linux coming soon. Users can generate photorealistic or stylized images right from their terminal using models like: Z-Image Turbo (from Alibaba’s Tongyi Lab): great at realism, bilingual text rendering (English + Chinese), and open for commercial use (Apache 2.0). FLUX.2 Klein (by Black Forest Labs): excels at text in images, UI mockups, and product visuals. Comes in 4B (open) and 9B (non-commercial) versions. Customization options include image size, random seed, number of steps, and negative prompts. Images save locally and even render inline in supported terminals like iTerm2 and Ghostty.
@JulianGoldieSEO ·
𝗥𝘂𝗻 𝗚𝗼𝗼𝗴𝗹𝗲'𝘀 𝗚𝗲𝗺𝗺𝗮 𝟰 𝗹𝗼𝗰𝗮𝗹𝗹𝘆 𝗳𝗼𝗿 𝗳𝗿𝗲𝗲 𝗶𝗻 𝗺𝗶𝗻𝘂𝘁𝗲𝘀 𝘂𝘀𝗶𝗻𝗴 𝗢𝗹𝗹𝗮𝗺𝗮. Ranked number 3 in the world among open models. Runs on your laptop. No cloud. No data leaving your machine. Here's exactly which model to pick: → E4B (9.6GB): Start here if you're on a typical laptop. Run: ollama run gemma4:e4b → 26B MoE (18GB): Only activates 4B parameters during inference. Fast with high quality. Run: ollama run gemma4:26b → 31B dense (20GB): Maximum quality. Number 3 globally. Run: ollama run gemma4:31b Best practices most people skip: → Set temperature 1.0, top_p 0.95, top_k 64. Google recommends these across all use cases. → Turn thinking mode ON for math, coding, and analysis. Turn it OFF for quick summaries. → When sending an image with a question put the image first then your text. Better performance. → Don't include the model's thinking output in your conversation history. Only final responses. Coding benchmark jumped from 110 ELO on Gemma 3 to 2,150 on the 31B. That's a generational gap in one release. Apache 2.0 license. Commercial use, fine-tuning, redistribution. No hidden restrictions. Install Ollama at https://t.co/493GbXWz04. Make sure you're on version 0.20 or higher for Gemma 4 support.
@PsudoMike ·
Jeff Morgan and Michael Chiang met at Waterloo, built Kitematic, and sold it to Docker, and that work is still the backbone of Docker Desktop today. Now they've raised $65 million USD for Ollama, an open source tool for running AI models on your own machine instead of renting someone else's GPU. Nearly nine million users already. I know Toronto founders who've hustled for months to raise a fraction of that. They built this one from Palo Alto. What would it take to keep the second act here?
Best Tweets by Topic