50 Best Tweets About Local LLMs (2026)

Explore the best tweets about local LLMs, on-device AI, hardware, quantization, privacy, inference, and self-hosted model setups. Updated weekly.

Reproducible local-model setups, useful hardware comparisons, performance measurements, and privacy-focused workflows.

Creators
39
Updated

What 50 top Local LLMs posts reveal

The dataset is strongly supportive of local-first AI and centers on serving stacks, hardware demonstrations, memory-efficient inference, and private or offline workflows. It also contains cautionary views that local setups can face latency, context-handling, and capability limits for complex work.

Dominant tone
Positive

80% of posts

Median score
54.3

All-time engagement

Leading format
Announcement

44% of posts

Recent posts
18%

Published in 90 days

Conversation map

The themes creators return to

Hardware selection and performance benchmarks

GPU, CPU, Mac unified-memory, mini-PC, mobile, and multi-GPU comparisons covering throughput, memory, power, bandwidth, and software support.

46%

Local agents and personal automation

Tool-using local agents for browser control, shell and file operations, smart homes, research, scheduled tasks, and persistent memory.

32%

Offline knowledge bases and local RAG

Searchable local documents, offline Wikipedia and maps, SQLite-backed retrieval, document processing, and resilience-focused knowledge systems.

16%

Quantization and memory-efficient inference

Low-bit weights, MoE-aware methods, KV-cache settings, adaptive precision, flash streaming, and techniques to fit larger models on constrained devices.

16%

Tone and stance

Sentiment Positive leads
Author posture Supportive leads

Performance benchmark

Median likes
106
Median reposts
15
Median replies
12
Median views
9.7K

Posts with media make up 80% of this collection. Their median all-time score is 70.6, compared with 3.12 for text-only posts.

Format mix

  • Announcement 44% · score 124.2
  • Case Study 28% · score 7.90
  • Opinion 12% · score 97.8
  • Tutorial 8% · score 236.7

Where creators agree, and where they do not

Shared view

Serving configuration can affect local-agent behavior

Posts identify runtime and integration details—including KV-cache reuse, KV-cache precision, thinking-mode settings, and direct API compatibility—as relevant to speed, output quality, or agent operation.

Shared view

Hardware comparisons need workload context

Examples include a 3090 Ti running a 27B model with 262k context, an M2 Ultra demonstration reported at 300t/s, and resource-constrained edge deployments. Together, they show that reported results depend on the model, runtime, precision, context, and device.

Open debate

Whether local models can replace cloud workflows is contested

Some posts describe successful local coding or research-agent setups, while critical posts argue that long system prompts, latency, and weaker reasoning can make cloud models preferable for complex work.

Open debate

VRAM alone is not presented as a sufficient hardware criterion

Posts discuss fitting models through compression and large-memory systems, while another argues that memory bandwidth, interconnect, inference-engine support, and software maturity also affect practical local inference.

Patterns behind standout posts

Several posts provide concrete throughput figures

Reported demonstrations include 300t/s on an M2 Ultra, 52.6 tok/s across an RTX 3080 and RTX 3070, and 30, 41, and 65 tok/s across three MacBook serving setups.

Media posts had a higher median all-time score

Deterministic analytics reports that 40 of 50 tweets included media (80%). The median all-time score was 70.59 for media posts and 3.12 for text-only posts.

Statistical standouts

  1. View standout post 1 Score 21348.8 · 393.09× median
  2. View standout post 2 Score 9759.8 · 179.7× median
  3. View standout post 3 Score 1489.0 · 27.42× median
  4. View standout post 4 Score 1390.8 · 25.61× median
  5. View standout post 5 Score 878.7 · 16.18× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. Vaishnavi

    @_vmlops

    2 posts

  2. 2. 0xSero

    @0xSero

    2 posts

  3. 3. Kshitij Mishra | AI & Tech

    @DAIEvolutionHub

    2 posts

  4. 4. Georgi Gerganov

    @ggerganov

    2 posts

  5. 5. Guri Singh

    @heygurisingh

    2 posts

  6. 6. Nav Toor

    @heynavtoor

    2 posts

llama.cpp’s creator contributed both a benchmark-style demonstration and a perspective on local AI

Georgi Gerganov posted an M2 Ultra llama.cpp demonstration reported at 300t/s and separately discussed cross-device, open-stack local AI and the increasing use of local LLMs.

Repeated creators covered setup guides and developer workflows

0xSero posted a framework-selection guide and a list of local-model harnesses. _vmlops posted a build-from-scratch agent-learning repository and a post describing a direct-server performance comparison.

Creator participation was broad in the supplied analytics

The analytics identifies 39 creators. Its top-five placement share is 20%, so the top five accounted for one-fifth of the measured placement share rather than a majority.

Since the previous snapshot

What changed since Aug 20, 2026

  • 78% of the selected posts remained.
  • The creator count changed by 0.
  • The leading sentiment remained stable.
How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top Local LLMs tweets from 39 creators

Ranked 01–50

  1. 01

    @heynavtoor ·

    🚨Someone just open sourced a computer that works when the entire internet goes down. It's called Project N.O.M.A.D. A self-contained offline survival server with AI, Wikipedia, maps, medical references, and full education courses. No internet. No cloud. No subscription. It

    • 595 Replies
    • 4K Reposts
    • 24.2K Likes
    • 1.1M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @heygurisingh ·

    Holy shit... Microsoft open sourced an inference framework that runs a 100B parameter LLM on a single CPU. It's called BitNet. And it does what was supposed to be impossible. No GPU. No cloud. No $10K hardware setup. Just your laptop running a 100-billion parameter model at

    Video thumbnail from Guri Singh's post Watch video
    • 866 Replies
    • 2.7K Reposts
    • 15.3K Likes
    • 2.1M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @techNmak ·

    Claude Code can run entirely on your local GPU now. Unsloth AI published the complete guide. The setup itself is straightforward - llama.cpp serves Qwen3.5 or GLM-4.7-Flash, one environment variable redirects Claude Code to localhost. But the guide is valuable because of what

    • 56 Replies
    • 211 Reposts
    • 1.7K Likes
    • 131.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @ihteshamali ·

    This feels illegal. Someone built a fully local deep research agent that writes its own search queries, hunts down sources, finds the gaps in its own answers, then keeps searching until it's done. It's called Local Deep Researcher. Drop in any Ollama model like DeepSeek,

    Video thumbnail from Ihtesham Ali's post Watch video
    • 36 Replies
    • 149 Reposts
    • 924 Likes
    • 65K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @stan_info ·

    i’m not a local llm type of guy at all, was just curious and decided to mess around… ended up running a full uncensored qwen3.5-27b (abliterated) on my single 3090 ti with 262k context + tool calling. threw a cloudflare tunnel on it so i can hit the api from anywhere. huge

    • 50 Replies
    • 71 Reposts
    • 1.7K Likes
    • 177.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @ggerganov ·

    Let me demonstrate the true power of llama.cpp: - Running on Mac Studio M2 Ultra (3 years old) - Gemma 4 26B A4B Q8_0 (full quality) - Built-in WebUI (ships with llama.cpp) - MCP support out of the box (web-search, HF, github, etc.) - Prompt speculative decoding The result:

    Video thumbnail from Georgi Gerganov's post Watch video
    • 134 Replies
    • 260 Reposts
    • 3.3K Likes
    • 733.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @ollama ·

    Ollama 0.18.1 is here! 🌐 Web search and fetch in OpenClaw Ollama now ships with web search and web fetch plugin for OpenClaw. This allows Ollama's models (local or cloud) to search the web for the latest content and news. This also allows OpenClaw with Ollama to be able to

    • 74 Replies
    • 217 Reposts
    • 1.7K Likes
    • 149.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @ggerganov ·

    llama.cpp at 100k stars now that 90% of the code worldwide is being written by AI agents, I predict that within 3-6 months, 90% of all AI agents will be running locally with llama.cpp 😄 Jokes aside, I am going to use this small milestone as an opportunity to reflect a bit on

    • 146 Replies
    • 288 Reposts
    • 2.1K Likes
    • 181.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @sentient_agency ·

    Nobody is talking about this. A developer built a research agent that works like a PhD student with unlimited time and zero salary. It's called Local Deep Researcher. Give it any topic and it takes over completely. It writes its own search queries, scrapes the web, reads

    Video thumbnail from Sentient's post Watch video
    • 11 Replies
    • 79 Reposts
    • 470 Likes
    • 32.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @Akashi203 ·

    i built a local LLM inference engine that runs a 1B parameter model on a $10 board with 256mb ram. model sits on the sd card, streams one layer at a time through 45mb of ram You can use it as local LLM model backend for PicoClaw no python no cloud no api keys 80kb binary | pure

    • 49 Replies
    • 92 Reposts
    • 726 Likes
    • 44K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @0xSero ·

    Here’s what I’d recommend if you’re just getting started in AI, local or otherwise. 1. Work with the compute you have, even the dumbest LLMs can be useful if you treat them as a node in your system. Some basic problems of what could be useful to get you started - tag all your

    • 32 Replies
    • 45 Reposts
    • 685 Likes
    • 24.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @akshay_pachaar ·

    Run Claude Code using local LLMs for FREE. No API costs. No data leaving your machine. Here's how it works: Claude Code lets you swap its backend via a single env variable. Point `ANTHROPIC_BASE_URL` to a local llama.cpp server, and it'll route all requests to whatever model

    • 28 Replies
    • 92 Reposts
    • 669 Likes
    • 46.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @aiDotEngineer ·

    TLMs: Tiny LLMs and Agents on Edge Devices with @cormacb https://t.co/u0fHD7j5kZ Function Gemma ships at 270 million parameters and runs nearly 2,000 tokens per second prefill on a Pixel 7. Out of the box, it hits 46% accuracy on a fixed set of app intents. Fine tune on a

    • 7 Replies
    • 95 Reposts
    • 596 Likes
    • 30.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @witcheer ·

    Google TurboQuant maybe reduced my LLM memory by 6x. running 70B models on my Mac Mini could be real. I have Mac Mini M4. memory was always the issue, 16GB meant 30B models were pushing it, 70B were out of the question. TurboQuant changes that math. 6x memory reduction without

    • 48 Replies
    • 37 Reposts
    • 588 Likes
    • 51.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @0xSero ·

    Best harnesses for local models: 1. Droid: - Very good performance, forces the models to behave, you can wire in all your local LLMs very easily w BYOK - Allows you to use your local models as orchestrators/subagents so you can benefit from Cloud as models as well - Practically

    • 24 Replies
    • 14 Reposts
    • 305 Likes
    • 12.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @mudler_it ·

    I've just released APEX (Adaptive Precision for EXpert Models): a novel MoE quantization technique that outperforms @UnslothAI Dynamic 2.0 on accuracy while being 2x smaller for MoE architectures. Benchmarked on Qwen3.5-35B-A3B, but the method applies to any MoE model. Half the

    • 25 Replies
    • 49 Reposts
    • 349 Likes
    • 28.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @heygurisingh ·

    🚨 BREAKING: Someone open-sourced a full offline survival computer with AI, Wikipedia, and maps built in. Project N.O.M.A.D. is an open-source offline survival computer. Self-contained. Zero internet required after install. Zero telemetry. Everything runs locally on your

    • 10 Replies
    • 66 Reposts
    • 269 Likes
    • 24.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @DAIEvolutionHub ·

    Holy shit 🤯 Microsoft just open-sourced a framework that runs a 100B parameter LLM on a single CPU. No GPU. No cloud. No expensive setup. Just your laptop. It’s called BitNet. And it breaks one of the biggest assumptions in AI. Here’s the trick: Most LLMs use 16-bit or

    Video thumbnail from Kshitij Mishra | AI & Tech's post Watch video
    • 27 Replies
    • 68 Reposts
    • 349 Likes
    • 34.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @BitcoinNewsCom ·

    Jack Dorsey's Block just launched mesh-llm. It's a decentralized, peer-to-peer inference network for open source AI models. The idea is to pool spare GPU compute across machines to run models too large for any single device. Rather than using a centralized cloud, it's just

    • 41 Replies
    • 113 Reposts
    • 701 Likes
    • 52.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @TheAhmadOsman ·

    People keep saying “VRAM is all that matters” for local LLMs > It’s not just wrong, it’s misleading When running LLMs locally, the bottleneck is NOT just “VRAM size” It’s: - memory bandwidth - interconnect (PCIe vs NVLink vs RDMA) - inference engine (vLLM, TensorRT-LLM,

    • 51 Replies
    • 23 Reposts
    • 294 Likes
    • 28.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @gneubig ·

    In the spirit of language model freedom, as a weekend project I finally set up my own fully local coding agent: * GMKtek box with AMD Ryzen * Qwen3.5 30B A3B, served with lemonade server+llama.cpp * ngrok to expose the endpoint * OpenHands as the agent, currently through the CLI

    • 31 Replies
    • 23 Reposts
    • 315 Likes
    • 40K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @TheAhmadOsman ·

    Local LLMs, Buy a GPU, and the Case for Cognitive Security AGI? There is no guarantee we reach it anytime soon, if ever. And even if we do, there is certainly no guarantee it will run on your machine. Betting your agency on either assumption is a mistake. What is guaranteed is

    • 23 Replies
    • 27 Reposts
    • 250 Likes
    • 10.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @heynavtoor ·

    In 2026, OpenAI made you rent your own conversations. GPT-5.5. Five dollars per million input tokens. 30 dollars per million output tokens. Every prompt logged. Every response stored. Every keystroke a line item on your credit card. ChatGPT Plus. 20 dollars a month. 240 dollars

    • 17 Replies
    • 29 Reposts
    • 124 Likes
    • 12.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @sudoingX ·

    first local model running on the rog 5090 mobile. gemma 2 2b it q4_k_m through llama.cpp, hermes agent pointing at the local endpoint. sitting on the balcony watching the beach while the model answers prompts. 24gb vram in a portable machine running a full inference stack plus

    • 19 Replies
    • 7 Reposts
    • 237 Likes
    • 8.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @ihteshamali ·

    🚨This is absolutely amazing…You can now run a full AI + Wikipedia + offline maps computer that works when every server on Earth goes dark. It's called Project N.O.M.A.D. and it's completely free. Two commands to install on any Ubuntu machine: curl the script. sudo bash it.

    • 8 Replies
    • 22 Reposts
    • 65 Likes
    • 6.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @_vmlops ·

    MOST DEVS USE AI AGENT FRAMEWORKS WITHOUT UNDERSTANDING WHAT'S HAPPENING INSIDE This repo teaches you to build agents from scratch local llms, no black boxes, real understanding 14 hands-on examples covering everything: ▪️ function calling & tool use ▪️ persistent memory

    • 3 Replies
    • 11 Reposts
    • 64 Likes
    • 4.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @Sumanth_077 ·

    Train LLMs locally without writing a single line of code! @UnslothAI just released Unsloth Studio - an open-source web UI for training and running models. Here's how it works: You upload a PDF, CSV, or DOCX file. The Data Recipes feature automatically transforms it into a

    Video thumbnail from Sumanth's post Watch video
    • 5 Replies
    • 23 Reposts
    • 67 Likes
    • 4.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @socialwithaayan ·

    🚨 OpenClaw and Hermes showed what an operator agent should look like. Someone just built the version that runs 100% on your machine 🤯 It's called Atomic Agent. Browser control. File management. Shell commands. Document parsing. Scheduled tasks. Persistent memory. All through

    • 36 Replies
    • 13 Reposts
    • 88 Likes
    • 22.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @_vmlops ·

    SOMEONE RAN A 122B MODEL ON THEIR MACBOOK... IN ONE NIGHT no cloud. no api fees. no data leaving the machine they went through 3 generations: → ollama + proxy: 30 tok/s → llama.cpp + proxy: 41 tok/s → mlx native server: 65 tok/s the breakthrough? they killed the proxy entirely

    • 2 Replies
    • 9 Reposts
    • 52 Likes
    • 4.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @sudoingX ·

    compiling llama.cpp with cuda 12.8 on the rog strix scar 18 5090. this is the mobile 5090, 24gb vram in a laptop. same vram class as the 3090 desktop i run benchmarks on, just portable. i can run these tests from anywhere now. first up carnice-27b vs qwen 3.5 27b dense vs gemma

    • 12 Replies
    • 6 Reposts
    • 154 Likes
    • 7.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @DAIEvolutionHub ·

    CHINA JUST DROPPED AN OCR MODEL THAT CHANGES EVERYTHING. A tiny 3B-parameter model can read an entire 100-page PDF in a single pass. No page splitting. No context loss. No cloud APIs. Meet Unlimited-OCR 👇 • Reads full documents with a 32K context window • Scores 93% on OCR

    Video thumbnail from Kshitij Mishra | AI & Tech's post Watch video
    • 6 Replies
    • 16 Reposts
    • 39 Likes
    • 3.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @cnxsoft ·

    Online guide to select the best hardware for local LLM/AI deployments. https://t.co/Q2LAYo5xxx The website relies on Qwen 3.5 models and shows the price, performance, power consumption, and more for a range of hardware from Raspberry Pi 5 16GB to NVIDIA DGX Spark. In one

    • 1 Replies
    • 6 Reposts
    • 43 Likes
    • 4.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @Shruti_0810 ·

    A developer replaced ChatGPT Pro + Claude Code Max + Cursor with a $300 ZimaBoard running local LLMs. Now his entire AI bill is basically the cost of electricity. Here's what changed: • Runs an x86 mini PC smaller than a hardcover book. • Hosts local LLMs, Proxmox, and every

    Video thumbnail from Shruti Codes's post Watch video
    • 5 Replies
    • 13 Reposts
    • 27 Likes
    • 4.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @vectro ·

    Gemma 4 E2B Running on RTX 3080 & 3070 combined * llama.cpp * Ubuntu server * No GUI * BF16 precision * 52.6 tok/sec. See the comments for the full set of commands to run in terminal and GUI locally or served over the network.

    Video thumbnail from Vectro's post Watch video
    • 2 Replies
    • 2 Reposts
    • 13 Likes
    • 1.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @morganlinton ·

    Fun little Sunday morning project, starting work on a little offline LLM I'm calling Wilderness Bot. I go hiking and backpacking out in the wilderness a lot, and want a way to have a little LLM loaded with wilderness medicine info, survival guides, etc. Decided to build in

    • 7 Replies
    • 0 Reposts
    • 22 Likes
    • 1.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @Shruti_0810 ·

    A LOCAL LLM RUNNING ON ONE 3090 JUST REPLACED GOOGLE HOME FOR 29 SMART DEVICES WITH ZERO CLOUD. It doesn't just answer questions. Ask, “Do I need a jacket?” and it checks your actual weather sensor before replying. The cloud is becoming optional.

    Video thumbnail from Shruti Codes's post Watch video
    • 4 Replies
    • 4 Reposts
    • 15 Likes
    • 1.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @alphabatcher ·

    BEST local LLMs to run in 2026: ​ High-performance (24+ GB VRAM, preferably with multiple GPUs) ​ • Kimi K2 - 1T params, 32B active. MoE beast • GLM-4.7 (Z AI) - 30B-A3B MoE, SWE-bench 73.8% • DeepSeek V3.2 - 671B / 37B active. Still the open-source king • Qwen3 235B-A22B -

    • 9 Replies
    • 1 Reposts
    • 22 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @InduTripat82427 ·

    Holy shit...Someone tried replacing Claude with a local LLM… …and waited 13 minutes for THIS: > “I am a large language model, trained by Google.” That’s it. 13 minutes. One useless sentence. Let’s be honest — we’ve ALL had this thought: “Why am I paying for Claude when I

    • 4 Replies
    • 3 Reposts
    • 24 Likes
    • 5.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @heyrimsha ·

    Running a 60GB AI model on a phone with 12GB of RAM should be impossible. Someone just did it anyway. It's called BigMoeOnEdge and it runs gpt-oss-120b (a model 5x bigger than the phone's RAM) at 2.2 tokens per second on plain CPU. No GPU. No NPU. Four cores and flash storage.

    Video thumbnail from Rimsha Bhardwaj's post Watch video
    • 2 Replies
    • 9 Reposts
    • 14 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @ujjwalscript ·

    Your "Local LLM" and “Free AI” dev setup is a massive WASTE of time and money! The hottest trend on X right now is showing off your "local-first" setup. Developers are buying expensive NVIDIA 5090s, bragging about running Llama-3.2 or Phi-3.5 completely offline, and treating

    • 18 Replies
    • 1 Reposts
    • 12 Likes
    • 2.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @RoundtableSpace ·

    RUNNING LOCAL LLMs ISN’T ALWAYS THE CHEAPER OPTION. he tried avoiding API costs with a local model but got hit with huge latency, heavy system prompts, and hardware limits that made it painfully slow.

    • 24 Replies
    • 0 Reposts
    • 76 Likes
    • 45.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @JulianGoldieSEO ·

    OLLAMA + CODEX APP JUST BROKE LOCAL AI CODING You can now run OpenAI’s desktop coding agent on models sitting on your own laptop. No subscription. No cloud code sharing. No waiting in line. What Changed: → Ollama 0.24 adds official support for the Codex app → Codex can now

    Video thumbnail from Julian Goldie SEO's post Watch video
    • 2 Replies
    • 1 Reposts
    • 4 Likes
    • 428 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @ThePracticalDev ·

    Arch Linux, the niri scrolling Wayland compositor, llama.cpp with ROCm, and a 27B model running fully offline on 128GB unified memory. This dev shares why local-first AI coding matters and exactly how the stack fits together. { author: @deepu105 } https://t.co/de9zwkVBma

    • 2 Replies
    • 3 Reposts
    • 8 Likes
    • 2.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @Rus_Khairullin ·

    Vitalik published a detailed post on how he set up a fully local, self-sovereign AI - no cloud, maximum privacy and security. AI agents can already work for hours, use tools and modify their own code. But most (even open-source) ignore security: data leaks, hidden instructions,

    • 2 Replies
    • 1 Reposts
    • 11 Likes
    • 401 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @DBCrypt0 ·

    OMG a new local LLM dropped that you can run on your Mac Mini and it’s just as good as Opus 4.6! 🔥 Local Model: Hi! Before we begin, tell me about yourself and what you want to call me? Me: Let’s call you Max. I’m a content creator, podcast host, and Web3/AI researcher. Local

    • 2 Replies
    • 1 Reposts
    • 10 Likes
    • 1.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @Aytunc ·

    ⁠A local model, running on my desk, just wrote a game about Erlik, the Turkic underworld god. DeepSeek-V4-Flash (IQ2XXS gguf, llama.cpp after some fight with mlx version) on a Mac Studio. No cloud, no API. Your own model, your own myths, your own game. 2026 is weird.

    • 2 Replies
    • 1 Reposts
    • 11 Likes
    • 1.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @TheZachMueller ·

    What’s the relative speed up of choosing llama cpp for running Hermes/OpenClaw for yourself vs vLLM? Even with complex subagent tasks, I feel like you’ll hit maybe 3-6 concurrency? For practice, we could say MiniMax or any of the Qwens if people have data

    • 3 Replies
    • 0 Reposts
    • 10 Likes
    • 2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @jshguo ·

    I used to think AMD GPUs were terrible for running local AI models. But today I tried Qwen3.6 Uncensored locally on my 7900 XT with llama.cpp. Turned off deep thinking and honestly… it feels really good. Way faster than I expected, and actually usable for simple/fast tasks like

    • 0 Replies
    • 0 Reposts
    • 6 Likes
    • 1.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @JulianGoldieSEO ·

    OLLAMA JUST FIXED THE BIGGEST PROBLEM WITH LOCAL AI AGENTS Gemma 4 could use tools before. Now it can actually finish the job. What changed: → Ollama 0.32.1 improves tool-response continuation → Gemma 4 can call a tool, process the result, and continue working → Fewer

    Video thumbnail from Julian Goldie SEO's post Watch video
    • 2 Replies
    • 0 Reposts
    • 1 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @TDataScience ·

    To gain a granular understanding of the true costs of running local LLMs, Arsen Apostolov measured the actual GPU electricity for eight models on one RTX 3090 — and produced some counterintuitive findings along the way. https://t.co/qoOaCdek0p

    • 0 Replies
    • 0 Reposts
    • 0 Likes
    • 4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone