50 Best Tweets About Small Language Models (2026)

Find the best tweets about small language models, including compact architectures, on-device AI, benchmarks, fine-tuning, efficiency, and deployment.

Small and compact language models, on-device inference, efficiency, benchmarks, fine-tuning, hardware constraints, and real deployments.

Creators
44
Updated

What 50 top Small Language Models posts reveal

The dataset is predominantly supportive and positive: 33 of 50 posts are labeled supportive (66%), and 39 are labeled positive (78%). Its most common themes are edge/on-device deployment (17 posts), small-model capability progress (15), and hardware-aware model selection (14). The cited posts pair enthusiasm for local and compact models with practical constraints around task fit, agent capability, speed, context, quantization, and narrow-device output limits. [2034015670837600686, 2074697730874823077, 2076135762291261627]

Dominant tone
Positive

78% of posts

Median score
17.4

All-time engagement

Leading format
Other

100% of posts

Recent posts
42%

Published in 90 days

Conversation map

The themes creators return to

Edge and on-device deployment

Deploying language and multimodal models directly on phones, wearables, edge systems, and microcontrollers under strict memory, power, and latency limits.

34%

Hardware-aware model selection

Matching model size, quantization, context limits, and expected speed to RAM, VRAM, CPUs, Apple silicon, phones, and other available hardware.

28%

Specialized training and fine-tuning

Fine-tuning, reinforcement learning, distillation, and specialized post-training that adapt compact models to domain tasks, agents, and guardrails.

26%

Model compression and efficiency

Compression techniques that make models cheaper to run, including low-bit quantization, 1-bit weights, KV-cache compression, sparsity, and memory-efficient architectures.

22%

Task-specific evaluation and selection

Choosing models with workload-specific evaluations rather than broad leaderboards, including capability thresholds for agents and narrow business tasks.

14%

Tone and stance

Sentiment Positive leads
Author posture Supportive leads

Performance benchmark

Median likes
48
Median reposts
10
Median replies
10
Median views
8.6K

Posts with media make up 76% of this collection. Their median all-time score is 19.1, compared with 11.3 for text-only posts.

Format mix

  • Other 100% · score 17.4

Where creators agree, and where they do not

Shared view

Fit models to hardware and workload

A recurring practical theme is matching local models to available hardware and the specific workload. The cited posts argue that smaller hardware may suit narrower workflows and that model selection should be evaluated on the intended task rather than a broad leaderboard.

Shared view

Edge deployment is broadening

Posts describe on-device deployment across Apple devices, phones, and microcontrollers. They highlight compiled runtimes, low memory footprints, offline operation, and resource-constrained inference; the cited claims are product- or project-specific.

Shared view

Specialization is a recurring SLM theme

Specialization and training recur as compact-model levers. The evidence includes a low-cost GPT-2-grade training claim, a proposal to distill specialized SLMs for agent evaluation and guardrails, and a small from-scratch model positioned for education.

Open debate

Local inference does not imply full replacement

Self-hosting is framed as useful, but not as a universal replacement for cloud models. One post notes that tool use, context size, and quantization—not weight size alone—shape whether a local model is useful for agents; another limits smaller hardware to smaller workflows.

Open debate

Feasibility claims come with sharp constraints

The cited feasibility demonstrations have clear constraints. The SSD-paged MoE example reports roughly one token every 10–20 seconds, while the $8 microcontroller project is described as TinyStories-trained and limited to short narrative output rather than general QA or tool use.

Open debate

Benchmarks do not replace task evaluation

Capability remains task-dependent in these posts. One warns to expect hallucinations from an on-device model, another reports a capability floor for coding agents, and another recommends building workload-specific evaluations for narrow tasks.

Patterns behind standout posts

Media posts had a higher median score

Media-bearing posts had a median all-time score of 19.13, compared with 11.35 for text-only posts. Media appeared in 38 of 50 posts (76%). These figures describe an association in this evidence set, not a causal media effect.

Hardware-fit discussion outperformed two other prominent themes

Hardware-aware model selection had a theme median all-time score of 16.96. This is higher than the cited medians for small-model capability progress (14.675) and edge/on-device deployment (8.828), but lower than specialized training and fine-tuning (93.414).

Statistical standouts

  1. View standout post 1 Score 1847.4 · 106.42× median
  2. View standout post 2 Score 1707.2 · 98.34× median
  3. View standout post 3 Score 1702.5 · 98.07× median
  4. View standout post 4 Score 1582.8 · 91.18× median
  5. View standout post 5 Score 839.6 · 48.37× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. Akshay 🚀

    @akshay_pachaar

    2 posts

  2. 2. Alex Finn

    @AlexFinn

    2 posts

  3. 3. andrew chen

    @andrewchen

    2 posts

  4. 4. BURKOV

    @burkov

    2 posts

  5. 5. Lior Alexander

    @LiorOnAI

    2 posts

  6. 6. Towards Data Science

    @TDataScience

    2 posts

Alex Finn emphasizes pragmatic local use

Across two posts, Alex Finn presents local models as a privacy- and cost-oriented workflow option, while noting that smaller hardware is unlikely to replace every AI call and may instead cover smaller workflows.

Akshay Pachaar covers tooling and guardrails

Akshay Pachaar’s posts cover an on-device framework described as Core AI and a proposal for a distilled, specialized SLM to serve as an agent evaluator and runtime guardrail.

Andrew Chen discusses product fit

Andrew Chen describes local models as usable for many cases while distinguishing local capability from cloud capability. He also argues that, where model-quality differences are less apparent for common prompts, factors such as privacy, bundling, and product design may matter.

How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top Small Language Models tweets from 44 creators

Ranked 01–50

  1. 01

    @AlexFinn ·

    I don't care what computer you have, you should be running local models It will save you a money on OpenClaw and keep your data private Even if you're on the cheapest Mac Mini you can be doing this Here's a complete guide: 1. Download LMStudio 2. Go to your OpenClaw and say

    • 189 Replies
    • 178 Reposts
    • 2K Likes
    • 142.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @AlexFinn ·

    I don't care what kind of hardware you have, you should be running local models It will save you a ton on money on OpenClaw and keep your data private Even if you're on the cheapest Mac Mini you can be doing this Here's a complete guide: 1. Download LMStudio 2. Go to your

    • 171 Replies
    • 209 Reposts
    • 2.1K Likes
    • 191.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @0xSero ·

    Best models to run on your hardware level I'll be doing this every week, I hope you guys enjoy. ---- 8 GB ---- Autocomplete for coding (like Cursor Tab) - https://t.co/Jyf766kmyd - https://t.co/dK1CQwCGqD Tool calling, assistant style - https://t.co/Jf7RY3dZmZ ---- 16 Gb

    • 183 Replies
    • 288 Reposts
    • 3K Likes
    • 325K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @rasbt ·

    A small Qwen3.5 from-scratch reimplementation for edu purposes: https://t.co/OnupgeE55l (probably the best "small" LLM today for on-device tinkering)

    • 47 Replies
    • 511 Reposts
    • 3.1K Likes
    • 150.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @karpathy ·

    nanochat can now train GPT-2 grade LLM for <<$100 (~$73, 3 hours on a single 8XH100 node). GPT-2 is just my favorite LLM because it's the first time the LLM stack comes together in a recognizably modern form. So it has become a bit of a weird & lasting obsession of mine to train

    • 325 Replies
    • 607 Reposts
    • 7.3K Likes
    • 1.3M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @akshay_pachaar ·

    Apple finally did it. Its new framework, Core AI, runs models entirely on Apple silicon, so inference happens on the user's device with zero server calls and zero token bills. That means Qwen, Mistral, and SAM3 running natively across iPhone, iPad, Mac, and Vision Pro. It's a

    • 36 Replies
    • 133 Reposts
    • 1.3K Likes
    • 91.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @the_smart_ape ·

    everyone's talking about @karpathy autoresearch and most of you have no idea what it actually does. there's a training script (train(dot)py) that trains a small language model, basically a baby GPT. and there's an instruction file (program(dot)md) that tells an AI agent what to

    • 42 Replies
    • 70 Reposts
    • 680 Likes
    • 39.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @BrianRoemmele ·

    WOW! The $8 AI Machine! Something extraordinary just happened and it changes what “local AI” can mean. I am testing it tonight. Thus far it shows many possibilities… So what it this $8 AI device? A developer going by slvDev has forced a 28.9-million-parameter language model

    • 74 Replies
    • 149 Reposts
    • 846 Likes
    • 52.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @_vmlops ·

    Everyone talks about LLMs like you need massive GPUs and billions of parameters You really don’t Came across this: https://t.co/j5z6o7vnpO⁠ It’s a tiny ~9M parameter model you can train in minutes (even on Colab) What I liked: ▫️It shows the entire pipeline - tokenizer →

    • 6 Replies
    • 40 Reposts
    • 291 Likes
    • 10K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @TheAhmadOsman ·

    Local AI is the future. Learning how to run Opensource models (Inference), how to evaluate them systematically (Evals), and how to customize them (Fine-tuning / RL / Post-training) are invaluable skills to start learning today.

    • 40 Replies
    • 39 Reposts
    • 608 Likes
    • 22.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @burkov ·

    In this paper, a 7B language model trained with reinforcement learning learns to orchestrate larger frontier models like GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro. It does so by writing natural-language subtasks, assigning each to one of the workers, and specifying which

    • 32 Replies
    • 72 Reposts
    • 458 Likes
    • 57.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @burkov ·

    This paper argues that Small Language Models (SLMs) offer a more economical and suitable future for agentic AI by demonstrating their sufficient power for specialized tasks, outlining a conversion algorithm from LLMs to SLMs, and discussing the significant operational and

    • 11 Replies
    • 61 Reposts
    • 342 Likes
    • 25.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @hasantoxr ·

    A 744 billion parameter AI model just ran on a machine with 25GB of RAM. No graphics card. The tool is called colibri. GLM-5.2 is a mixture-of-experts model. It contains 744 billion parameters but only about 40 billion wake up for each token. colibri keeps the 9.9GB core in

    • 24 Replies
    • 30 Reposts
    • 203 Likes
    • 20.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @akshay_pachaar ·

    Vibe train your AI agents. There's a new method that could replace LLM-as-a-judge for production agents. Most teams rely on a giant LLM as a judge to evaluate and guard their agent. But it has two major drawbacks: - It's slow and expensive at inference time - It often misses

    • 20 Replies
    • 25 Reposts
    • 159 Likes
    • 11.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @arankomatsuzaki ·

    • µLMs (8M–30M params) generate the first 4–8 words on-device in ~55 ms • A cloud LLM then continues the same response seamlessly • This “commit-and-continue” design masks cloud latency • Result: near-instant, context-grounded AI on wearables and other constrained devices

    • 12 Replies
    • 40 Reposts
    • 240 Likes
    • 21.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @emollick ·

    Gemma 4 E4B is impressive for an on-device LLM. GPT-4ish quality, and expect hallucinations. Here is: “List five sociological theories starting with u and what they are. Then describe them in a rhyming verse” Its in real time, the last is a little bit of a stretch, but not bad!

    Video thumbnail from Ethan Mollick's post Watch video
    • 38 Replies
    • 30 Reposts
    • 374 Likes
    • 53.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @AIHighlight ·

    🚨 Breaking: A free AI model trained for $7,800 just beat one 400 times its size at competition math. It runs small enough to fit on a laptop. Weibo's AI lab published the results and open-sourced the whole thing. The industry has run on one idea for two years. Bigger is

    • 16 Replies
    • 64 Reposts
    • 133 Likes
    • 12.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @LiorOnAI ·

    Google's latest paper on Compression is the future. Here's why. They compressed LLM memory 6x with zero accuracy loss. When ChatGPT writes a reply, it remembers every word you've said. That memory is stored in a growing notebook (KV cache). A 100,000-word conversation can eat

    • 24 Replies
    • 32 Reposts
    • 176 Likes
    • 26.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @paulabartabajo_ ·

    Looking for real mobile AI apps that run 100% on-device? No cloud. No API keys. Here's a batch of examplels built with LFM models and @liquidai's LEAP SDK. Bookmark this ↓ https://t.co/xlD3SCbzJs

    • 2 Replies
    • 14 Reposts
    • 55 Likes
    • 2.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @ModelScope2022 ·

    Meet Marco-Mini-Instruct: a highly sparse MoE multilingual model from Alibaba International. 17.3B total params, only 0.86B active (5% activation ratio). 🚀 Beats Qwen3-4B, Gemma3-12B, Granite4-Small on English, multilingual general, and cultural benchmarks — with a fraction of

    • 3 Replies
    • 12 Reposts
    • 108 Likes
    • 7.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @andrewchen ·

    set up a mini rack for a home lab setup (will share a pic soon) w my Mac mini and DGX spark with more coming. had a few thoughts as I play w qwen3.5, gemma4, and other models: - there’s an S curve on LLM model quality per use case. Show text output side by side from the latest

    • 29 Replies
    • 7 Reposts
    • 109 Likes
    • 12.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @AlphaSignalAI ·

    Researchers just gave LLMs a separate brain for memory. Language models go stale the moment training ends. Updating them risks breaking what they already know. A new paper proposes MeMo. It pairs any LLM with a separate trained memory model. The base model stays frozen.

    • 3 Replies
    • 11 Reposts
    • 31 Likes
    • 1.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @arpit_bhayani ·

    When a public LLM benchmark says one model is better than another, it is testing broad, difficult reasoning and, more importantly, opinionated tasks. Your use case might not need that. For example, if you are summarizing tickets, classifying intent, or extracting fields from

    • 17 Replies
    • 10 Reposts
    • 163 Likes
    • 13.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @TeksEdge ·

    🚀 Future of LLM inference just got faster! Offloading the small draft model in speculative decoding to high-bandwidth SRAM accelerators (like d-Matrix @CORSAIR) while the big model stays on the GPU. 🎯 Result 2–10× lower end-to-end latency vs GPU-only speculative decoding and

    https://gimletlabs.ai/blog/low-latency-spec-decode-corsair
    • 1 Replies
    • 1 Reposts
    • 38 Likes
    • 3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @LiorOnAI ·

    Most language models only read forward. Perplexity just open-sourced 4 models that read text in both directions. They used a technique from image generation to retrain Qwen3 so every word can see every other word in a passage. That changes how well a model understands meaning.

    • 6 Replies
    • 10 Reposts
    • 42 Likes
    • 4.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @shshnkp ·

    New! Foundation Models SDK for Python. Access the on-device LLM from Python to: - Run on-device batch &amp; real-time streaming inference - Guided generation via type-safe decorators/schemas - Process transcripts exported from Swift apps for quality analysis

    • 1 Replies
    • 6 Reposts
    • 50 Likes
    • 3.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @RoundtableSpace ·

    A 14B model on a single consumer GPU hitting 74.6% on LiveCodeBench. No fine-tuning. No API calls. No cloud. No data leaving your machine. The idea is simple wrap a frozen small model in smart infrastructure and it starts competing with frontier models at a fraction of the

    • 21 Replies
    • 13 Reposts
    • 119 Likes
    • 52.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @thestreamingdev ·

    3 ai models racing simultaneously on an m2 macbook air (8gb) built tiny bit, a local agent terminal for benchmarking small language models head-to-head on apple silicon. @PrismML @liquidai @Alibaba_Qwen tested bonsai-8b (1-bit, 1.16gb) vs qwen3-0.6b (q4, 0.37gb) vs lfm2.5-350m

    Video thumbnail from thestreamingdev()'s post Watch video
    • 4 Replies
    • 5 Reposts
    • 35 Likes
    • 9.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @SimonHoiberg ·

    Everyone is now looking at self-hosted AI models. Naturally. But if you think you can buy a Mac Mini and replace Claude or OpenAI, you need a serious wakeup call. Cause there's more to it than just the size of the weights - for a model to become practically useful, we need to

    Video thumbnail from Simon Høiberg's post Watch video
    • 19 Replies
    • 5 Reposts
    • 48 Likes
    • 6.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @DAIEvolutionHub ·

    A 744 billion parameter AI model just ran on a machine with 25GB of RAM. No graphics card. The project is called colibri. It works with GLM-5.2, a 744B parameter mixture-of-experts model where only about 40B parameters are used for each generated token. Instead of forcing

    • 6 Replies
    • 10 Reposts
    • 24 Likes
    • 3.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @Marktechpost ·

    Liquid AI Released LFM2.5-350M: A Compact 350M Parameter Model Trained on 28T Tokens with Scaled Reinforcement Learning - LFM2.5-350M is a 350M parameter small language model trained on 28 trillion tokens, with a hybrid architecture built from 10 double-gated LIV convolution

    • 2 Replies
    • 6 Reposts
    • 25 Likes
    • 2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @aaliya_va ·

    Stop downloading LLMs; your machine was never going to run. llmfit scans your hardware and tells you exactly which models will run. It scans your RAM, CPU, GPU, and VRAM first. Then it scores every model across four dimensions: 1. Quality, based on parameter count and

    Video thumbnail from Aaliya's post Watch video
    • 13 Replies
    • 5 Reposts
    • 47 Likes
    • 7.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @DivyanshT91162 ·

    People said LLMs needed GPUs. This one runs on an $8 ESP32-S3 microcontroller. • 28.9M parameters • 14.9 MB 4-bit model • ~9.5 tokens/sec • 0 cloud • 0 internet • 100% on-device inference • 100% open source • MIT License The trick? Instead of loading the entire model into

    • 1 Replies
    • 6 Reposts
    • 15 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @smratitiwa86867 ·

    🤯 Researchers just made LLMs over 99% sparse… without killing performance. And unlike most “sparse AI” papers, this one actually gets REAL GPU speedups instead of just theoretical FLOPs reductions. The trick? LLMs are already naturally sparse inside their feedforward layers.

    • 5 Replies
    • 6 Reposts
    • 9 Likes
    • 437 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @arsh_goyal ·

    A 1 billion parameter AI model. Runs entirely on your laptop with no cloud or GPU bills or internet. MiniCPM5-1B just dropped and it's the best 1B local model right now. Here's what actually makes it worth trying: > Beats Qwen3.5 0.8B and LFM2.5 1.2B on math, coding, and tool

    Video thumbnail from Arsh Goyal's post Watch video
    • 8 Replies
    • 4 Reposts
    • 19 Likes
    • 11.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @ttunguz ·

    Pocket Power : From State of the Art to Your Phone in 23 Months Two years ago, the idea of useful AI on your phone was fantastical. Siri couldn't finish a sentence. Local models hallucinated nonsense. Last week, Google released Gemma 4 E4B, a free model that matches GPT-4o &

    • 10 Replies
    • 4 Reposts
    • 24 Likes
    • 4.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @andrewchen ·

    Pepsi challenge for LLMs Contrarian view during a week of huge new model launches: All of us do a lot of “normie prompts” - these are use cases which are really like Google searches (“what’s the name of..” “is it true that…” “what’s the best…”). These are a very high % of total

    • 24 Replies
    • 3 Reposts
    • 38 Likes
    • 19.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @mark_k ·

    On March 31, @PrismML unveiled Bonsai, a family of 1-bit ultra-dense language models that pack astonishing intelligence into tiny footprints. Named after the art of miniature trees, these models prove that true AI power thrives when compressed rather than expanded. The flagship

    • 1 Replies
    • 5 Reposts
    • 23 Likes
    • 1.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @alphabatcher ·

    BEST local LLMs to run in 2026: ​ High-performance (24+ GB VRAM, preferably with multiple GPUs) ​ • Kimi K2 - 1T params, 32B active. MoE beast • GLM-4.7 (Z AI) - 30B-A3B MoE, SWE-bench 73.8% • DeepSeek V3.2 - 671B / 37B active. Still the open-source king • Qwen3 235B-A22B -

    • 9 Replies
    • 1 Reposts
    • 22 Likes
    • 1.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @nrqa__ ·

    MiniCPM-V4.6 1.3B just changed what “small AI model” means. It runs on edge devices. Outperforms Qwen3.5-0.8B on key multimodal benchmarks. And cuts visual computation costs by ~50%. This feels like the beginning of truly deployable AI

    • 10 Replies
    • 9 Reposts
    • 38 Likes
    • 40.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @sabir_huss50540 ·

    A 28.9 million parameter language model just ran on an $8 chip. No server. No wifi. No GPU. It runs on an ESP32-S3, the kind of chip you solder into a hobby project, writing each word to a tiny wired screen at about 9.5 tokens per second. The last model anyone ran on a chip

    • 5 Replies
    • 3 Reposts
    • 11 Likes
    • 805 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @tom_doerr ·

    Train language models under 16MB https://t.co/yPyjBQpoY8

    • 1 Replies
    • 1 Reposts
    • 16 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @cleanunicorn ·

    Google is talking about TurboQuant and it's huge They're releasing this compression algorithm to the world, and it just might be what powers those 2M context window models. The Problem Every LLM uses a key-value cache to track conversation context. It's like a digital cheat

    • 1 Replies
    • 2 Reposts
    • 8 Likes
    • 579 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @StarHistoryHQ ·

    2.9k⭐ guppylm - A tiny 9M-parameter LLM built from scratch to demystify how language models actually work — train it on a single GPU in minutes 🐟 by @armanfixing https://t.co/iQLnHsFHBj #starhistory #GitHub #OpenSource #Python

    • 1 Replies
    • 1 Reposts
    • 2 Likes
    • 241 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @seunosewa ·

    Working with the smallest modern LLM I could find, Qwen3:0.6b, I now understand supervised fine-tuning and direct preference optimization, full fine-tuning vs LoRA, and the role of cheap API models like DeepSeek v4 in cleaning, formatting, and even generating sufficient data. 🥳

    • 2 Replies
    • 1 Reposts
    • 12 Likes
    • 668 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @TDataScience ·

    Does your project call for a full-scale frontier model, or will a small language model be enough? How should you go about choosing the right one? Sara Nobrega breaks down the decision-making process in an accessible guide. https://t.co/eZBLtH7t29

    • 1 Replies
    • 0 Reposts
    • 6 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @TDataScience ·

    If you'd like to spend some time this weekend tinkering with code, why not explore @PetrKorab's new Python tutorial? It covers fine-tuning a small language model for emotion recognition in social media communications. https://t.co/Zz5pSEBh28

    • 1 Replies
    • 2 Reposts
    • 5 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @ainativedev ·

    What's the smallest model that can actually run an AI agent? Not the cheapest. Not the fastest. The smallest one that reliably gets the job done. In their latest analysis, Nicolas Fortuin and Baptiste Fernandez (@FernandezBap ) put NVIDIA's (@nvidia) open-weight Nemotron models

    • 0 Replies
    • 1 Reposts
    • 3 Likes
    • 105 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @HuggingModels ·

    Meet Phi-2: a surprisingly capable small language model that's punching way above its weight class. At just 2.7B parameters, it delivers performance rivaling models 10x its size. Perfect for when you need serious NLP power without the computational overhead.

    • 1 Replies
    • 0 Reposts
    • 5 Likes
    • 620 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @JulianGoldieSEO ·

    MINICPM5 JUST BROKE THE "BIGGER IS BETTER" MYTH A 1B parameter model is doing things that used to require models 10x its size. Here's what makes it different: Performance: → Just 1B parameters but supports reasoning, coding, and tool use → Runs locally on a laptop or even a

    Video thumbnail from Julian Goldie SEO's post Watch video
    • 1 Replies
    • 0 Reposts
    • 2 Likes
    • 1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone