50 Best Tweets About Local LLMs (2026)

Explore the best tweets about local LLMs, on-device AI, hardware, quantization, privacy, inference, and self-hosted model setups. Updated weekly.

Reproducible local-model setups, useful hardware comparisons, performance measurements, and privacy-focused workflows.

Creators
40
Updated

Top Local LLMs tweets from 40 creators

Ranked 01–50

  1. 01

    @alex_prompter Β·

    🚨 BREAKING: Someone just open-sourced a full offline survival computer with AI, Wikipedia, and maps built in. Project N.O.M.A.D. is an open-source offline survival computer. Self-contained. Zero internet required after install. Zero telemetry. Everything runs locally on your

    • 357 Replies
    • 3.2K Reposts
    • 23.6K Likes
    • 4.3M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @heygurisingh Β·

    Holy shit... Microsoft open sourced an inference framework that runs a 100B parameter LLM on a single CPU. It's called BitNet. And it does what was supposed to be impossible. No GPU. No cloud. No $10K hardware setup. Just your laptop running a 100-billion parameter model at

    Video thumbnail from Guri Singh's post Watch video
    • 866 Replies
    • 2.7K Reposts
    • 15.3K Likes
    • 2.1M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @techNmak Β·

    Claude Code can run entirely on your local GPU now. Unsloth AI published the complete guide. The setup itself is straightforward - llama.cpp serves Qwen3.5 or GLM-4.7-Flash, one environment variable redirects Claude Code to localhost. But the guide is valuable because of what

    • 56 Replies
    • 211 Reposts
    • 1.7K Likes
    • 131.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @ihteshamali Β·

    This feels illegal. Someone built a fully local deep research agent that writes its own search queries, hunts down sources, finds the gaps in its own answers, then keeps searching until it's done. It's called Local Deep Researcher. Drop in any Ollama model like DeepSeek,

    Video thumbnail from Ihtesham Ali's post Watch video
    • 36 Replies
    • 149 Reposts
    • 924 Likes
    • 65K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @stan_info Β·

    i’m not a local llm type of guy at all, was just curious and decided to mess around… ended up running a full uncensored qwen3.5-27b (abliterated) on my single 3090 ti with 262k context + tool calling. threw a cloudflare tunnel on it so i can hit the api from anywhere. huge

    • 50 Replies
    • 71 Reposts
    • 1.7K Likes
    • 177.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @ollama Β·

    Ollama 0.18.1 is here! 🌐 Web search and fetch in OpenClaw Ollama now ships with web search and web fetch plugin for OpenClaw. This allows Ollama's models (local or cloud) to search the web for the latest content and news. This also allows OpenClaw with Ollama to be able to

    • 74 Replies
    • 217 Reposts
    • 1.7K Likes
    • 149.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @Akashi203 Β·

    i built a local LLM inference engine that runs a 1B parameter model on a $10 board with 256mb ram. model sits on the sd card, streams one layer at a time through 45mb of ram You can use it as local LLM model backend for PicoClaw no python no cloud no api keys 80kb binary | pure

    • 49 Replies
    • 92 Reposts
    • 726 Likes
    • 44K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @0xSero Β·

    Here’s what I’d recommend if you’re just getting started in AI, local or otherwise. 1. Work with the compute you have, even the dumbest LLMs can be useful if you treat them as a node in your system. Some basic problems of what could be useful to get you started - tag all your

    • 32 Replies
    • 45 Reposts
    • 685 Likes
    • 24.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @akshay_pachaar Β·

    Run Claude Code using local LLMs for FREE. No API costs. No data leaving your machine. Here's how it works: Claude Code lets you swap its backend via a single env variable. Point `ANTHROPIC_BASE_URL` to a local llama.cpp server, and it'll route all requests to whatever model

    • 28 Replies
    • 92 Reposts
    • 669 Likes
    • 46.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @TheAhmadOsman Β·

    You don’t pick an Inference Engine You pick a Hardware Strategy and the Engine follows Inference Engines Breakdown (Cheat Sheet at the bottom) > llama.cpp runs anywhere CPU, GPU, Mac, weird edge boxes best when VRAM is tight and RAM is plenty hybrid offload, GGUF,

    • 28 Replies
    • 44 Reposts
    • 523 Likes
    • 34.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @aiDotEngineer Β·

    TLMs: Tiny LLMs and Agents on Edge Devices with @cormacb https://t.co/u0fHD7j5kZ Function Gemma ships at 270 million parameters and runs nearly 2,000 tokens per second prefill on a Pixel 7. Out of the box, it hits 46% accuracy on a fixed set of app intents. Fine tune on a

    • 7 Replies
    • 95 Reposts
    • 596 Likes
    • 30.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @AlexFinn Β·

    5 months ago I spent $30,000 on 3 Mac Studios, 2 Mac Minis, and a DGX Spark I went all in on local LLMs and encouraged others to do the same I warned prices would explode I was called crazy, a hype beast, dangerous, and that I had no idea what I was talking about Since then:

    • 291 Replies
    • 115 Reposts
    • 1.5K Likes
    • 238K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @witcheer Β·

    Google TurboQuant maybe reduced my LLM memory by 6x. running 70B models on my Mac Mini could be real. I have Mac Mini M4. memory was always the issue, 16GB meant 30B models were pushing it, 70B were out of the question. TurboQuant changes that math. 6x memory reduction without

    • 48 Replies
    • 37 Reposts
    • 588 Likes
    • 51.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @0xSero Β·

    Best harnesses for local models: 1. Droid: - Very good performance, forces the models to behave, you can wire in all your local LLMs very easily w BYOK - Allows you to use your local models as orchestrators/subagents so you can benefit from Cloud as models as well - Practically

    • 24 Replies
    • 14 Reposts
    • 305 Likes
    • 12.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @mudler_it Β·

    I've just released APEX (Adaptive Precision for EXpert Models): a novel MoE quantization technique that outperforms @UnslothAI Dynamic 2.0 on accuracy while being 2x smaller for MoE architectures. Benchmarked on Qwen3.5-35B-A3B, but the method applies to any MoE model. Half the

    • 25 Replies
    • 49 Reposts
    • 349 Likes
    • 28.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @DAIEvolutionHub Β·

    Holy shit 🀯 Microsoft just open-sourced a framework that runs a 100B parameter LLM on a single CPU. No GPU. No cloud. No expensive setup. Just your laptop. It’s called BitNet. And it breaks one of the biggest assumptions in AI. Here’s the trick: Most LLMs use 16-bit or

    Video thumbnail from Kshitij Mishra | AI & Tech's post Watch video
    • 27 Replies
    • 68 Reposts
    • 349 Likes
    • 34.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @gneubig Β·

    In the spirit of language model freedom, as a weekend project I finally set up my own fully local coding agent: * GMKtek box with AMD Ryzen * Qwen3.5 30B A3B, served with lemonade server+llama.cpp * ngrok to expose the endpoint * OpenHands as the agent, currently through the CLI

    • 31 Replies
    • 23 Reposts
    • 315 Likes
    • 40K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @TheAhmadOsman Β·

    Local LLMs, Buy a GPU, and the Case for Cognitive Security AGI? There is no guarantee we reach it anytime soon, if ever. And even if we do, there is certainly no guarantee it will run on your machine. Betting your agency on either assumption is a mistake. What is guaranteed is

    • 23 Replies
    • 27 Reposts
    • 250 Likes
    • 10.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @EXM7777 Β·

    how to set up OpenCode (the privacy-first Claude Code alternative) in 3 minutes: > install: npm i -g opencode-ai > configure your preferred model provider > cd into your project directory > run: opencode it reads your codebase locally, builds context on your machine, and only

    • 12 Replies
    • 11 Reposts
    • 110 Likes
    • 7.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @sudoingX Β·

    first local model running on the rog 5090 mobile. gemma 2 2b it q4_k_m through llama.cpp, hermes agent pointing at the local endpoint. sitting on the balcony watching the beach while the model answers prompts. 24gb vram in a portable machine running a full inference stack plus

    • 19 Replies
    • 7 Reposts
    • 237 Likes
    • 8.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @witcheer Β·

    I run ollama on a Mac Mini for local compression. every message my agent sends passes through a local qwen model to summarise context before it overflows. speed matters because slow compression means slow responses across every cron job. ollama shipped MLX backend for Apple

    • 9 Replies
    • 7 Reposts
    • 125 Likes
    • 12K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @socialwithaayan Β·

    oh my.. this shouldn't be possible a 1B model that runs inside your browser, beats every model its size, and comes with its own desktop pet. MiniCPM-5 1B just changed the game for on-device AI. here's everything you need to know 🧡

    Video thumbnail from Muhammad Ayan's post Watch video
    • 22 Replies
    • 13 Reposts
    • 141 Likes
    • 61.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @_vmlops Β·

    SOMEONE RAN A 122B MODEL ON THEIR MACBOOK... IN ONE NIGHT no cloud. no api fees. no data leaving the machine they went through 3 generations: β†’ ollama + proxy: 30 tok/s β†’ llama.cpp + proxy: 41 tok/s β†’ mlx native server: 65 tok/s the breakthrough? they killed the proxy entirely

    • 2 Replies
    • 9 Reposts
    • 52 Likes
    • 4.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @sudoingX Β·

    compiling llama.cpp with cuda 12.8 on the rog strix scar 18 5090. this is the mobile 5090, 24gb vram in a laptop. same vram class as the 3090 desktop i run benchmarks on, just portable. i can run these tests from anywhere now. first up carnice-27b vs qwen 3.5 27b dense vs gemma

    • 12 Replies
    • 6 Reposts
    • 154 Likes
    • 7.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @DAIEvolutionHub Β·

    CHINA JUST DROPPED AN OCR MODEL THAT CHANGES EVERYTHING. A tiny 3B-parameter model can read an entire 100-page PDF in a single pass. No page splitting. No context loss. No cloud APIs. Meet Unlimited-OCR πŸ‘‡ β€’ Reads full documents with a 32K context window β€’ Scores 93% on OCR

    Video thumbnail from Kshitij Mishra | AI & Tech's post Watch video
    • 6 Replies
    • 16 Reposts
    • 39 Likes
    • 3.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @cnxsoft Β·

    Online guide to select the best hardware for local LLM/AI deployments. https://t.co/Q2LAYo5xxx The website relies on Qwen 3.5 models and shows the price, performance, power consumption, and more for a range of hardware from Raspberry Pi 5 16GB to NVIDIA DGX Spark. In one

    • 1 Replies
    • 6 Reposts
    • 43 Likes
    • 4.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @Cyb3rMaddy Β·

    Been messing around with local, uncensored LLMs... Ollama runs open-source LLMs locally. It’s handy for security research, private workflows, and anything you don’t want leaving your machine β€” even works without internet. No cloud calls. No sending prompts somewhere you can’t

    • 4 Replies
    • 5 Reposts
    • 55 Likes
    • 2.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @aigleeson Β·

    Claude Code just got a free mode. Not official. A developer built a repo that routes Claude Code calls to free and local AI models. Instead of buying Anthropic credits... You can run it with: > NVIDIA NIM > OpenRouter free models > DeepSeek > LM Studio > llama.cpp Best

    • 1 Replies
    • 12 Reposts
    • 29 Likes
    • 3.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @tonysimons_ Β·

    A 27B multimodal AI model is now running locally on an iPhone. PrismML’s Bonsai 27B packs 27.8B parameters into just 3.9GB and reportedly hits 11 tokens per second on an iPhone 17 Pro. That’s roughly 14Γ— smaller than FP16 while retaining more than 90% of benchmark performance.

    • 5 Replies
    • 5 Reposts
    • 27 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @morganlinton Β·

    Fun little Sunday morning project, starting work on a little offline LLM I'm calling Wilderness Bot. I go hiking and backpacking out in the wilderness a lot, and want a way to have a little LLM loaded with wilderness medicine info, survival guides, etc. Decided to build in

    • 7 Replies
    • 0 Reposts
    • 22 Likes
    • 1.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @Shruti_0810 Β·

    A LOCAL LLM RUNNING ON ONE 3090 JUST REPLACED GOOGLE HOME FOR 29 SMART DEVICES WITH ZERO CLOUD. It doesn't just answer questions. Ask, β€œDo I need a jacket?” and it checks your actual weather sensor before replying. The cloud is becoming optional.

    Video thumbnail from Shruti Codes's post Watch video
    • 4 Replies
    • 4 Reposts
    • 15 Likes
    • 1.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @HeyZaraKhan Β·

    Nobody expected Apple to open-source this. Apple just open-sourced one of the most interesting AI repositories for developers. It's called Core AI Models. Instead of forcing you to convert and optimize models yourself… Apple provides ready-made export recipes, Python

    • 3 Replies
    • 4 Reposts
    • 20 Likes
    • 4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @morganlinton Β·

    Everyone is talking about Kimi and Qwen, but I'm honestly surprised more people aren't talking about models like Trinity from Arcee. I've been doing a deeper dive here and it's pretty interesting, here's a few differences that I'm not sure ppl fully realize. - Qwen and Kimi

    • 5 Replies
    • 0 Reposts
    • 21 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @lemire Β·

    AMD is coming up with its small AI PC (AI Halo). It will compete against NVIDIA DGX Spark. Both look a bit like a mac Mini. Just a tiny box. The AI Halo should cost US$4000, so it is accessible to hobbyists and small IT departement. Set it up with LM Studio with its llama.cpp

    • 6 Replies
    • 5 Reposts
    • 38 Likes
    • 5.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @0xJiuJitsuJerry Β·

    Late night breakthrough: My Mac Mini M4 "Mainframe" is now talking to my @zocomputer through MCP 🧠⚑ What does this mean? Local Gemma3, Nemotron, Qwen3.5 all available as tools No Cloudflare tunnel dependency Zero latency on my AI agent swarm Complete data sovereignty From

    • 3 Replies
    • 3 Reposts
    • 8 Likes
    • 159 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @Axel_bitblaze69 Β·

    local AI is now useful enough to keep at home imp.. a small machine sitting on your desk can now handle a large part of your everyday AI work for roughly a few dollars in electricity each month, depending on what you run and how often. things like: - drafting
- summarising
-

    • 17 Replies
    • 2 Reposts
    • 20 Likes
    • 3.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @ujjwalscript Β·

    Your "Local LLM" and β€œFree AI” dev setup is a massive WASTE of time and money! The hottest trend on X right now is showing off your "local-first" setup. Developers are buying expensive NVIDIA 5090s, bragging about running Llama-3.2 or Phi-3.5 completely offline, and treating

    • 18 Replies
    • 1 Reposts
    • 12 Likes
    • 2.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @Shruti_0810 Β·

    The most expensive part of this AI setup isn't the hardware. It's... nothing. A Raspberry Pi 5 with 16GB RAM and a 512GB NVMe is powering: β€’ Local LLMs with Ollama β€’ Claude Code via localhost β€’ Bluetooth analysis β€’ Wi-Fi security testing β€’ Packet capture All from a device

    Video thumbnail from Shruti Codes's post Watch video
    • 5 Replies
    • 4 Reposts
    • 19 Likes
    • 4.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @SaiyamPathak Β·

    Yesterday in the video I took - the qwen32B which was a dense 32B and all active parameters for every token whereas for the MLX version it was A3B - active 3B. this morning I ran some tests again: - Qwen3.5 (NVFP4, MLX): 23.11 tok/s decode - Nemotron (GGUF, llama.cpp): 21.70

    • 0 Replies
    • 1 Reposts
    • 10 Likes
    • 919 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @_vmlops Β·

    CLAUDE CODE FOR FREE - WITHOUT THE ANTHROPIC API BILL Someone built a local proxy that intercepts claude code's api calls & reroutes them to free models point one env variable at localhost and you're done: β†’ nvidia nim (free tier) β†’ openrouter free models β†’ deepseek β†’ ollama /

    • 0 Replies
    • 0 Reposts
    • 5 Likes
    • 387 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @ThePracticalDev Β·

    Arch Linux, the niri scrolling Wayland compositor, llama.cpp with ROCm, and a 27B model running fully offline on 128GB unified memory. This dev shares why local-first AI coding matters and exactly how the stack fits together. { author: @deepu105 } https://t.co/de9zwkVBma

    • 2 Replies
    • 3 Reposts
    • 8 Likes
    • 2.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @Rus_Khairullin Β·

    Vitalik published a detailed post on how he set up a fully local, self-sovereign AI - no cloud, maximum privacy and security. AI agents can already work for hours, use tools and modify their own code. But most (even open-source) ignore security: data leaks, hidden instructions,

    • 2 Replies
    • 1 Reposts
    • 11 Likes
    • 401 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @smratitiwa86867 Β·

    Holy shit...Someone tried replacing Claude with a local LLM… …and waited 13 minutes for THIS: > β€œI am a large language model, trained by Google.” That’s it. 13 minutes. One useless sentence. Let’s be honest β€” we’ve ALL had this thought: β€œWhy am I paying for Claude when I

    • 0 Replies
    • 3 Reposts
    • 3 Likes
    • 365 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @DBCrypt0 Β·

    OMG a new local LLM dropped that you can run on your Mac Mini and it’s just as good as Opus 4.6! πŸ”₯ Local Model: Hi! Before we begin, tell me about yourself and what you want to call me? Me: Let’s call you Max. I’m a content creator, podcast host, and Web3/AI researcher. Local

    • 2 Replies
    • 1 Reposts
    • 10 Likes
    • 1.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @JeremyCMorgan Β·

    If your last local-LLM setup was early 2025, the viable model list moved. This refresh covers Ollama, llama.cpp, and VRAM sizing, with Qwen 2.5 Coder 32B the standout coding pick at 24GB, scoring ahead of GPT-4o on HumanEval per the post. Recalibrate before buying a box.

    • 0 Replies
    • 0 Reposts
    • 1 Likes
    • 96 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @Aytunc Β·

    ⁠A local model, running on my desk, just wrote a game about Erlik, the Turkic underworld god. DeepSeek-V4-Flash (IQ2XXS gguf, llama.cpp after some fight with mlx version) on a Mac Studio. No cloud, no API. Your own model, your own myths, your own game. 2026 is weird.

    • 2 Replies
    • 1 Reposts
    • 11 Likes
    • 1.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @TDataScience Β·

    From installing Ollama to launching OpenCode with a local model, Shuai Guo presents a step-by-step tutorial on building your own local coding agent with Gemma 4. https://t.co/EkLE5CCAOc

    • 0 Replies
    • 0 Reposts
    • 4 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @JulianGoldieSEO Β·

    This tool feels illegal to be free. Ollama now works with Claude Desktop and Claude Code. That means local models inside your Claude workflow. Your private code can stay on your own machine. Finally, AI without the cloud babysitter. Link in the comments πŸ‘‡

    • 2 Replies
    • 0 Reposts
    • 0 Likes
    • 228 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @JulianGoldieSEO Β·

    OLLAMA JUST FIXED THE BIGGEST PROBLEM WITH LOCAL AI AGENTS Gemma 4 could use tools before. Now it can actually finish the job. What changed: β†’ Ollama 0.32.1 improves tool-response continuation β†’ Gemma 4 can call a tool, process the result, and continue working β†’ Fewer

    Video thumbnail from Julian Goldie SEO's post Watch video
    • 2 Replies
    • 0 Reposts
    • 1 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @TDataScience Β·

    To gain a granular understanding of the true costs of running local LLMs, Arsen Apostolov measured the actual GPU electricity for eight models on one RTX 3090 β€” and produced some counterintuitive findings along the way. https://t.co/qoOaCdek0p

    • 0 Replies
    • 0 Reposts
    • 0 Likes
    • 4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone