Open-weight language models
New LLM checkpoints, model families, reasoning capabilities, licenses, benchmarks, and Hugging Face model cards.
46%
Best tweets about Hugging Face
Browse the best tweets about Hugging Face, from open models and datasets to Spaces, Transformers, inference, and machine learning workflows.
Technical Hugging Face releases, repositories, datasets, demos, libraries, and real model-building experience.
Original Xholic analysis
Hugging Face appears here as both a release venue and a working developer stack: open checkpoints sit alongside dataset access, fine-tuning tools, retrieval pipelines, and runnable demos. The conversation is overwhelmingly supportive, but hands-on posts temper launch claims with hardware requirements and repository limits (2028460046510965160, 2041501396692873308, 2040096474114187477).
90% of posts
All-time engagement
68% of posts
Published in 90 days
Conversation map
New LLM checkpoints, model families, reasoning capabilities, licenses, benchmarks, and Hugging Face model cards.
46%
Quantization, compact and MoE architectures, long contexts, serving integrations, and running models on consumer GPUs, Macs, browsers, or edge devices.
32%
TRL releases, LoRA workflows, reinforcement learning recipes, and tools for adapting open models.
30%
Hugging Face storage, Jobs, Transformers integrations, MCP tools, repository access, and the engineering behind Hub-scale services.
30%
Research datasets, large-scale data hosting, dataset ingestion, preprocessing, and dataset-maintenance experience.
20%
Real-time voice assistants, transcription, text-to-speech, voice cloning, and music-generation models and demos.
18%
Open models and demos for image understanding, video generation, visual editing, and image-to-3D assets.
18%
Embedding models, visual document retrieval, OCR, and practical RAG training and deployment pipelines.
12%
Tone and stance
Performance benchmark
Posts with media make up 76% of this collection. Their median all-time score is 22.6, compared with 7.32 for text-only posts.
Format mix
Consensus and debate
Shared view
Releases point beyond model cards to usable serving options: Qwen offers FP8 weights with vLLM and SGLang support, Sarvam names day-zero SGLang support, and a Gemma quantization post targets consumer GPUs (2026682179305275758, 2029965547824431356, 2039754171969069501).
Shared view
Posts describe one-line access to psychiatric genetics data, streaming Hub datasets into LanceDB, and a fine-tuning studio that searches Hub models and datasets before running training on Hugging Face infrastructure (2041501396692873308, 2031822582559719619, 2054095066348896748).
Open debate
Compact and quantized releases emphasize local use, including browser inference and consumer GPUs. Other posts note that PersonaPlex may need substantial VRAM, an embedding fine-tuning recipe requires 80GB, and Kimi K3 is too large for the cited local setup (2024156940599714084, 2039754171969069501, 2014532042328035608, 2039478124631466319, 2081997718420107550).
What performs
Open-weight language models account for 46% of tweets and have a 58.676 median all-time score; Hub infrastructure accounts for 30% with a 7.19 median. The supplied outlier list includes Qwenโs small-model launch and Sarvamโs release (2028460046510965160, 2029965547824431356, 2037303305970315406).
The supplied analytics report media on 76% of tweets, with a 22.65 median all-time score versus 7.32 for text-only posts. Tutorials have a 57.24 median, though there are only four; examples include a forecasting walkthrough and a voice-model setup guide (2043615267792814163, 2014532042328035608).
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Vaishnavi
@_vmlops
2 posts
2. AshutoshShrivastava
@ai_for_success
2 posts
3. Qwen
@Alibaba_Qwen
2 posts
4. DailyPapers
@HuggingPapers
2 posts
5. Ihtesham Ali
@ihteshamali
2 posts
6. Meituan LongCat
@Meituan_LongCat
2 posts
A dataset maintainer reports that a 10,000-commit limit stopped telemetry updates, while a model tester shares long-context checks, memory considerations, and harness caveats. These accounts add implementation detail to the release-heavy conversation (2040096474114187477, 2042637857194676321).
A fine-tuning studio connects Hub search, training configuration, AutoTrain, and chat-based testing. A separate voice-assistant repository uses swappable speech and LLM stages behind a WebSocket API, illustrating reusable workflows rather than checkpoints alone (2054095066348896748, 2084196506668875993).
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Hugging Face tweets
Ranked 01โ50
@Alibaba_Qwen ยท
๐ Introducing the Qwen 3.5 Small Model Series Qwen3.5-0.8B ยท Qwen3.5-2B ยท Qwen3.5-4B ยท Qwen3.5-9B โจ More intelligence, less compute. These small models are built on the same Qwen3.5 foundation โ native multimodal, improved architecture, scaled RL: โข 0.8B / 2B โ tiny, fast, great for edge device โข 4B โ a surprisingly strong multimodal base for lightweight agents โข 9B โ compact, but already closing the gap with much larger models And yes โ weโre also releasing the Base models as well. We hope this better supports research, experimentation, and real-world industrial innovation. Hugging Face: https://t.co/wFMdX5pDjU ModelScope: https://t.co/9NGXcIdCWI

@MaziyarPanahi ยท
๐จ Over 1 billion rows of psychiatric genetics data. Now on Hugging Face. ADHD. Depression. Schizophrenia. Bipolar. PTSD. OCD. Autism. Anxiety. Tourette. Eating disorders. 12 disorder groups. 52 publications. Every GWAS summary statistic from the Psychiatric Genomics Consortium. Before: wget, gunzip, 20 minutes debugging separators, repeat 50 times. Now: one line of Python.

@ihteshamali ยท
๐จBREAKING: Microsoft open sourced a 4B parameter model that generates production-ready 3D assets from a single image, and the speed numbers are genuinely hard to believe. It's called TRELLIS.2 and it uses a new geometry format called O-Voxel that can be converted to a textured mesh in under 100 milliseconds on CUDA, which means real-time 3D asset generation is now actually within reach. Most image-to-3D tools force you to pick between speed, quality, and topology correctness. TRELLIS.2 refuses that tradeoff by representing geometry natively in a sparse voxel format that supports arbitrary complexity from the start. โ Outputs GLB files with full PBR texture maps ready for Blender, Unity, and Unreal Engine โ Pretrained TRELLIS.2-4B checkpoint available directly on Hugging Face with model card and usage details โ Web demo live now so you can test it without installing anything 4.6K stars. MIT License. 100% Opensource. Link in comments.
@pratykumar ยท
๐ข Open-sourcing the Sarvam 30B and 105B models! Trained from scratch with all data, model research and inference optimisation done in-house, these models punch above their weight in most global benchmarks plus excel in Indian languages. Get the weights at Hugging Face and AIKosh. Thanks to the good folks at SGLang for day 0 support, vLLM support coming soon. Links, benchmark scores, examples, and more in our blog - https://t.co/DcCG3zlN8p
@HuggingPapers ยท
NVIDIA just released a quantized Gemma 4 31B on Hugging Face NVFP4 compression delivers 4x smaller weights with frontier-level accuracy. Runs on consumer GPUs with a 256K context window.

@natjin ยท
> train embeddings model on actual web search > use it in actual production (200M daily queries) > see crazy results: best contextual retrieval in the world (81.96% CoNTEB; next closest 79.45%) > open source it > 1M hugging face downloads in ~2 weeks > 5-30x cheaper than existing providers > go try it :) (pplx-embed on @huggingface)

@googlegemma ยท
Thanks for following us! We're excited to see what you all build with Gemma 4! In case you missed it, you can find all our checkpoints, with an Apache 2.0 License, on Hugging Face:
@arcee_ai ยท
Today we're releasing Trinity-Large-Thinking. Available now on the Arcee API, with open weights on Hugging Face under Apache 2.0. We built it for developers and enterprises that want models they can inspect, post-train, host, distill, and own.
@minchoi ยท
Hugging Face just made arXiv paper retrieval way better for AI agents. ๐๐ ๐๐๐๐๐๐ [๐๐๐๐๐๐, ๐๐๐๐] turns arXiv into agent-ready markdown๐
@KanikaBK ยท
๐ฑ I SPENT 3 HOURS TESTING THIS SO HERE IS WHAT ACTUALLY MATTERS. Kronos is a foundation model built from scratch on the language of financial markets. 45 global exchanges. AAAI 2026 accepted. MIT license. Every trading model being built from scratch right now is already behind. The crazy part? Most people building trading models right now have NO IDEA this already exists. While everyone else is still coding from scratch. Here is what makes it different from every other trading model: โณ Trained specifically on financial candlestick data not repurposed from general time series models โณ Handles OHLCV data natively. Open, high, low, close, volume and amount all in one model โณ Novel two-stage framework. Specialized tokenizer converts continuous market data into discrete tokens first. Then a large autoregressive transformer learns from those tokens โณ Designed to handle the unique high-noise characteristics of financial data that break general purpose models โณ Trained across 45 global exchanges for broad market understanding Four models available today: โณ Kronos-mini โ 4.1M parameters. Lightweight. Fast. Deploy anywhere. โณ Kronos-small โ 24.7M parameters. Strong performance for most use cases. โณ Kronos-base โ 102.3M parameters. Full power. Production ready. โณ Kronos-large โ 499.2M parameters. Research grade. Coming soon. All on Hugging Face. All free. Raw data to forecast in four lines of code: โณ Load the tokenizer and model from Hugging Face โณ Pass your historical OHLCV data โณ Define your prediction length โณ Call predict No feature engineering. No manual preprocessing. Just raw candles in and forecasts out.

@Meituan_LongCat ยท
๐ LongCat-Flash-Thinking-2601 Technical Report โ Now Fully Released! Key insights: ๐ Large-scale agentic RL (14 pages of deep dives!) ๐น Environment scaling: A detailed look at our automated pipeline that builds 10,000+ executable, verifiable environments across 20+ domains. ๐น RL infrastructure: An upgraded DORA framework that supports async training with 32,000+ concurrent environments, tackling stability issues in long-tail and highly heterogeneous tasks. ๐ก๏ธ Robustness in the wild ๐น Noise injection: No more "greenhouse" agents. We systematically analyze real-world noise (user/tool noise) and inject it directly into the training loop. ๐น Curriculum RL: A curriculum-based strategy that gradually toughens the model against messy, imperfect environments. ๐ง Heavy Thinking framework ๐น Parallel reasoning: Expands breadth by generating multiple independent reasoning trajectories. ๐น Iterative summarization: Expands depth by using a summary model to reflect on and synthesize parallel trajectories before making final decisions. ๐น Context memory: A purpose-built memory module to keep reasoning coherent over long horizons. โก Zigzag Attention ๐น Zigzag Connectivity design combining MLA + SSA to reduce compute while preserving global information flow. ๐น Mid-training switch to sparse variants yields a 1.5ร speedup and supports 1M-token contexts โlaying the groundwork for future breakthroughs in long-context agentic reasoning. ๐น Explore๏ผhttps://t.co/xmvQ2kmJUV ๐ Achieves SOTA among open-source models across key agentic benchmarks: search, tool use, mathematical reasoning, and coding. If you want more details, feel free to check out the full technical report. โข Paper: https://t.co/X7h2092UN5 โข Website: https://t.co/d6cZdCPWnh โข GitHub: https://t.co/24sd7zY98j โข Hugging Face: https://t.co/UCmFfzqTlj




@Meituan_LongCat ยท
๐ Introducing LongCat-Flash-Thinking-2601 โ A version built for deep and general agentic thinking. โจ Highlights: ๐ค Top Tier Agent Capabilities ๐น Performance: Top tier benchmark results (TIR / Agentic Search / Agentic Tool Use) ; superb generalization ability, outperforming Claude in complex, random tasks ๐น Env Scaling: Multiple automaticly constructed high-quality environments; dense dependency graph ๐น Multi-Env RL: Extended DORA (our RL infra), supporting large-scale multi-environment agentic training ๐ก๏ธ Real-World Robustness ๐น Performance: Solid performance in messy, uncertain scenarios (Vita-Noise & Tau^2-Noise) ๐น Noise Analysis: Systematically analyzed real-world noise in agentic scenarios ๐น Curriculum RL: Increasing noise type & intensity while training ๐ฏ Heavy Thinking Mode โ๐น Parallel Thinking: Expands breadth via multiple independent reasoning tracks ๐น Iterative Summarization: Enhances depth by using a summary model to synthesize outputs, supporting iterative reasoning loops ๐ One more thing: 1M-token context via Zigzag Attention is coming soon. ๐ Try it now: https://t.co/IccK18HvhO โ API access for this version is also available. Hugging Face: https://t.co/UCmFfzrraR GitHub: https://t.co/24sd7zYGXR

Watch video@BrianRoemmele ยท
NVIDIA's PersonaPlex-7B: The Game-Changing Open-Source Leap in open source Real-Time Voice AI. NVIDIA has just dropped a bombshell that's set to transform how we interact with voice-based AI forever! They've open-sourced PersonaPlex-7B, a revolutionary 7-billion-parameter model that's all about full-duplex conversations, meaning it can listen and speak at the same time, just like a natural human chat. No more awkward pauses or waiting your turn; this thing handles interruptions, backchannels, and smooth turn-taking like a pro. As someone who's always rooting for open-source innovation, I'm thrilledโthis is the kind of project that democratizes cutting-edge tech and sparks a wave of creativity in the AI community. What makes PersonaPlex-7B so exciting? Built on the Moshi architecture, it uses a dual-stream setup to process audio input and generate responses in real-time, slashing latency and making interactions feel incredibly lifelike. Trained on a mix of synthetic customer service dialogues and the real-world Fisher English Corpus, it excels at maintaining consistent personas through simple text prompts and voice embeddings. Imagine telling it, "You are a wise and friendly teacher," and it instantly adopts that role with a matching voice, natural or varied, male or female. There are 16 pre-packaged voice profiles to choose from, like NATF2 for a natural female tone or VARM0 for something more dynamic. This isn't just talk; it's backed by the Helium LLM backbone, ensuring it generalizes well even to out-of-distribution prompts. The best part? It's 100% open-source and free, licensed under MIT for the code and NVIDIA's Open Model License for the weights. You can grab the model from Hugging Face, but the heart of the project is on GitHub. Installation is straightforward and developer-friendly. First, clone the repo and pip install the Moshi package. If you're on Blackwell GPUs, there's an optional Torch install for CUDA 12.8. Accept the license on Hugging Face, set your HF token, and launch the server with a quick Python command. Boom, you've got a web UI running locally at localhost:8998 for live demos. For offline testing, there's an evaluation script that takes input audio, a voice prompt, and spits out responses in WAV and JSON formats. Early users are raving about its speed on high-end hardware like H100 GPUs, though it might need 40GB+ VRAM to run smoothly. On the performance front, PersonaPlex-7B shines in benchmarks like FullDuplexBench, where it nails categories such as user interruptions, pause handling, and natural flow. It's not perfect, some folks report minor bugs like audio clipping in initial runs but that's the beauty of open-source: the community can jump in, fix, and build on it. NVIDIA's even shared a citation for their upcoming paper on this, signaling more research to come. This project is a massive win for open-source AI enthusiasts. By making full-duplex voice tech accessible, NVIDIA is empowering developers, researchers, and hobbyists to create everything from advanced virtual assistants to immersive chat experiences. No more gatekept proprietary systems, this is the future, and it's here for everyone to tinker with. For the source of the project: https://t.co/Qr9KXAnY0Z
@techNmak ยท
Hugging Face put out a repo that lets you build a full voice assistant, the kind that listens, thinks, and talks back, entirely with open-source models. It's a pipeline with four stages: - voice activity detection, - speech-to-text, - an LLM, and - text-to-speech Each running in its own thread. Every single stage is swappable, you can pick different STT models, different TTS models, different LLM backends, all through CLI flags. And the whole thing is exposed through a WebSocket API that's compatible with OpenAI's Realtime API, so anything built for that protocol can talk to it without changes. This pipeline is already running in production, it's the conversation backend for thousands of Reachy Mini robots. Install is one line "pip install speech-to-speech", and by default it spins up a local server using Parakeet TDT for transcription and Qwen3-TTS for speech output, with the LLM slot pointed at whatever OpenAI-compatible endpoint you give it.

@heygurisingh ยท
๐จRAG engineers are going to lose their minds. @webAI just open sourced a document retrieval model that's sitting at #1 AND #3 on ViDoRe V3 -- with Nvidia's best open-source embedding model trapped at #2 between them. No OCR. No text extraction. No broken pipelines on messy PDFs. It's called webAI-ColVec1. Here's what this thing actually does: โ Skips OCR entirely and retrieves directly from the rendered page image -- sees layout, tables, charts, and structure the same way you do โ Ships in two variants (4B and 9B) with embedding sizes of 128, 640, and 2560 -- pick speed or max retrieval quality โ Trained on ~2M question-image pairs across scientific papers, financial filings, healthcare records, government reports, and technical manuals โ Uses 511 in-batch negatives per query for a brutal contrastive signal that forces clean separation between correct and competing pages โ Built with a proprietary loss function designed specifically for retrieval -- not borrowed from generic embedding training โ LoRA rank 32 + retrieval projection layer means efficient specialization without full fine-tuning Here's the wildest part: Built on Qwen 3.5 vision-language backbones and trained on just 8 A100s. No giant model. No massive infra budget. No scale-maxing. Just a deliberate, retrieval-specific training recipe applied to the right problem. And it beat the largest GPU company in the world on their own benchmark. You literally point this at a financial filing or a scanned healthcare doc -- the stuff that destroys every OCR pipeline in production -- and it retrieves the right page based on what the page actually looks like. That sentence shouldn't be real for an open-source model trained on 8 GPUs in 2026. But here we are. 100% Open Source. Live on Hugging Face. Two of the top three spots on ViDoRe V3. (Link in the comments)

@Alibaba_Qwen ยท
๐ฅ Qwen 3.5 Medium Model Series FP8 weights are now open and ready for deployment๏ผ Native support for vLLM and SGLang. Check the model card for example code. โก๏ธ Optimize your workflow with FP8 precision. ๐ Get the weights: Hugging Face๏ผhttps://t.co/wFMdX5pDjU ModelScope๏ผhttps://t.co/9NGXcIdCWI
@MistralDevs ยท
Since launching Voxtral Realtime, the community response has been remarkable. Today, we share the technical report, launch the Realtime playground in Mistral Studio, and share the model in Hugging Face Transformers. ๐งต

@OpenBMB ยท
๐ VoxCPM 2 is live! ๐ Another open-source AI #TTS model from China โ and one that stands shoulder to shoulder with Qwen3-TTS, while bringing everything into a single unified model. After rapid iterations from V1 (zero-shot cloning) to V1.5 (long-form + fine-tuning), #VoxCPM has consistently pushed quality and usability forward. Now, VoxCPM 2 takes it further: ๐น30+ languages โ truly global, truly local. ๐นInfinite voice design โ type it, hear it, control it. From a whisper to a booming cinematic voice. ๐นStudio-grade audio โ 48kHz ultra-high fidelity with emotional depth ๐นDiffusion-Autoregressive cloning โ preserves more acoustic and emotional detail than token-based models like Qwen3-TTS ๐ก Big shoutout to @grok โ used your multi-image video magic for our launch demo. Itโs scarily good at keeping visuals consistent across shots. Elon @elonmusk, this oneโs for you. ๐ Check the demo & start cloning your dream voice: ๐ Hugging Face Space: https://t.co/lcK9r7Y90I ๐ค Hugging Face Model: https://t.co/6SlvwXPZHC ๐ค ModelScope Model: https://t.co/bOMuiCQ5Q0 ๐ป GitHub๏ผhttps://t.co/GVCQJjjasP #TTS #AI #VoiceCloning #GrokImagine #ElonMusk #OpenBMB #VoxCPM
@KyleHessling1 ยท
Gemopus-4-26B-A4B from Jackrong is LIVE! Happy to have benched this one pretty hard (see my benches in the model card) and it is an excellent finetune of an already exceptional model! My friend Jackrong is always cooking the greatest! It rocks at one-shot requests over long contexts, and runs incredibly fast thanks to the MOE architecture while not seeming to take as much of a hit vs dense models as in the Qwen 3.5 series. It also crushed my simple needle-in-the-haystack tests all the way out to an extended context of 524k! If you're VRAM starved, or running on unified memory, this one should run much more usably offloaded to system ram or in unified memory pools; even if you're running a 10GB or less GPU! It would be my daily driver for this purpose! That said, the dense 31B Gemopus 4 is finalising now, I will post it here when it's live, so follow me for the official launch, and follow Jakcrong on Hugging Face! It will also be an incredible model! As with the base Gemma 4 models, there are some idiocycracies especially in harnesses, if you have problems, please let us know! If you make something cool with it, please comment that below, too. We'd love to see it! https://t.co/nDiZMPxUeu
@xenovacom ยท
Cohere Labs just released Tiny Aya on Hugging Face: a collection of 4 open-weight multilingual LLMs optimized for over 70 languages. At just 3.35B parameters, they are perfect for on-device use cases, and can even run 100% locally in your browser on WebGPU! Try out the demo! ๐
@ai_for_success ยท
Meta AI is on ๐ฅ๐ฅ Meta just launched Muse Glimmer, a 30-billion-parameter open-weight agentic model (Apache 2.0) built specifically to run locally on consumer GPUs, Macs, and PCs. - Enables always-on, local agent workflows without cloud dependencies or network latency - Fits in 24GB to 32GB VRAM envelopes using 4-bit quantization (K-Quant ~17GB) while reserving working memory for KV cache - Trained via logit distillation from Muse Spark, mid-trained on long-context agent traces, and post-trained with reinforcement learning - Features native failure recovery to diagnose schema or runtime errors during tool execution and retry steps automatically - Uses a companion DFlash speculative decoding drafter to achieve 1.5x to 3.1x faster token generation on M-series chips and RTX GPUs - Supports interleaved text and vision inputs through a dedicated perception encoder for interpreting charts, UI screenshots, and documents Weights and developer documentation are live on Hugging Face. Video credit: Meta blog
@sharbel ยท
Google Research built a pretrained foundation model for time-series forecasting. It's called TimesFM. You pip install it. You load your data. You call forecast. You get predictions back in seconds. No account. No API key. No data leaving your machine. Here's what it does: โ Zero-shot forecasting on any time series. No training required. Point it at new data and it predicts. โ 200M parameter transformer model. Decoder-only architecture. The same paradigm that powers LLMs, applied to numbers. โ Up to 16,000 context length. Feed it years of history and it uses all of it. โ Continuous quantile forecasting up to 1,000 horizon steps via an optional 30M quantile head. โ Covariate support via XReg. Add external variables that influence your forecast. โ Fine-tune on your own data with LoRA via HuggingFace Transformers and PEFT. โ PyTorch and Flax backends. Runs on CPU, GPU, TPU, and Apple Silicon. โ Checkpoints on Hugging Face. One line to download the latest model. โ Integrates with BigQuery ML for enterprise SQL-based forecasting at scale. โ Works inside Google Sheets for spreadsheet-level forecasting. โ Agentic calling support via Vertex Model Garden and SKILL.md. โ pip install timesfm. That's the entire setup. Apache-2.0 licensed. Self-hosted. Open weights on Hugging Face. Free forever. 100% Open Source. Github repo link: https://t.co/kTn6IRtMcM

@ai_for_success ยท
โก Google DeepMind just dropped DiffusionGemma, latest experimental open model (Apache 2.0) that generates text up to 4x faster. - Uses diffusion instead of traditional next token autoregressive generation - Generates and refines 256 token blocks in parallel - Achieves up to 700+ tokens/sec on RTX 5090 and 1000+ tokens/sec on a single H100 - Designed as a 26B MoE model but activates only 3.8B params during inference - Can run quantized within 18 GB VRAM - Supports bidirectional attention during generation - Can self correct outputs during inference through iterative denoising - Handles global context much better than standard left to right models - Particularly strong for constraint based tasks like Sudoku - Fine tuned Sudoku version reached 80% success while base model was near 0% - Uses block autoregressive diffusion for long context generation - Integrated directly into vLLM for OpenAI compatible serving - Released with official training recipes and fine tuning support - Optimized across RTX 4090, RTX 5090, Hopper, and Blackwell GPUs Getting started: - Model weights are available publicly on Hugging Face - Works with vLLM, Hugging Face Transformers, and MLX - Deployable through Google Cloud Model Garden and NVIDIA NIM - Google released official fine tuning recipes through Hackable Diffusion License: - Released under Apache 2.0 - Allows commercial usage - Allows modification and redistribution - Developers can fine tune and build products on top of it This is one of the strongest public signs yet that major labs are actively exploring post autoregressive architectures for future LLMs.
@Michael_J_Black ยท
BEDLAM2.0 image and depth data are now available via Hugging Face, providing high-speed worldwide download access to over 26TB of synthetic data for non-commercial research. Hugging Face: https://t.co/tl8S3DJNWw Project: https://t.co/NR5Np9UT46
@elder_plinius ยท
was losing my mind trying to figure out why telemetry stopped auto-populating the HF dataset for https://t.co/eQHO0qjw1M triple-checked the code, totally fine. not the most efficient batching, but everything was wired the way it's supposed to be... WTF Y U NO PUSH?! turns out Hugging Face has a hard limit of 10,000 commits ๐ y'all "hug-of-deathed" the community dataset in one week ๐ I've gained a new level of respect for project maintainers

@darshal_ ยท
H3 is stacked. MiniMax H3 is now open-weight, bringing a powerful omni-modal generative system into the hands of creators and developers. H3 understands and generates across text, images, videos, and audio within the same creative context, enabling new possibilities for multimodal creation, video generation, and editing. The community can now explore H3 through Hugging Face, ComfyUI, Diffusers, and SGLang Diffusion, with optimized workflows making local experimentation possible on consumer GPUs. This is more than just a model release, itโs a major milestone for open-weight AI video. Hope everyone realizes how significant this moment is. #MiniMaxH3 #OpenWeight #AIvideo

@victormustar ยท
Featured Apps you can try on Hugging Face this week ๐ฅ ๐ฃ๏ธ Voxtral TTS Demo: Mistral's new text-to-speech ๐๏ธ Cohere Multilingual ASR: multilingual transcription โก Cohere WebGPU: same but locally in your browser ๐ฉ Mr. Chatterbox: Victorian-era gentleman chatbot ๐ต PrismAudio: text-prompted audio for video ๐ PROMETHEUS v1.0: embodied AI world model ๐จ VFig Image2SVG: diagrams to editable SVGs ๐บ LTX 2.3 Sync: portrait animation & lipsync

@IlirAliu_ ยท
A zero-shot video-language reward model. Trained on over 1M trajectories from 21 robot embodiments. It predicts: Frame-level task progress and generalizes zero-shot to unseen tasks, scenes, and robots, yielding 2.4โ4.5x better success rates in: Online/offline RL, data filtering, failure detection, and imitation learning. Now available in Hugging Face's @LeRobotHF robotics library. The video shows a robot gripper performing a pick-and-place task with a red object into a bowl, overlaid with a real-time rising progress graph generated by ROBOMETER. ๐ Project: https://t.co/mkhMvmbG55 Paper: https://t.co/HYn1PXhZ7V โโ- Weekly robotics and AI insights. Subscribe free: https://t.co/9Nm01QUcw3
@fortytwonetwork ยท
20,000+ downloads reached on Hugging Face for the Fortytwo Rust Coder Model โท The model was trained on data generated by Fortytwo node operators โท Five quantized versions shipped by independent devs โท 43.00% (SOTA) on the RustEvo^2 benchmark One more example of the AI community outpacing centralized labs Get it on Hugging Face โ
@svpino ยท
Here is a new open-weight audio model you can integrate with your app. I'm a huge sucker for open models that you can host to keep your data secure and away from Big AI. The model is Fish Audio S2. You'll find the weights, the code to fine-tune it, and a streaming inference engine in Hugging Face. (Link below). Under the hood, it's two models: a 4B one to handle timing and meaning, and a 400M one to fill in the acoustic details. The model is really fast! If you host it on an H200, it will start returning audio in about 100ms. If you don't want to host it, Fish Audio offers a hosted version: the S2.1 Pro. This model is even faster, with around 90ms latency, and it supports 83 languages. The S2.1 Pro is also ~83% more affordable than ElevenLabs: you pay around $15 for around 12 hours of transcribed audio, which is incredible. HuggingFace links below.
@DivyanshT91162 ยท
Yann LeCun just reposted it. Why is everyone suddenly talking about Unlimited-OCR? Baidu's Unlimited-OCR is back in the Top 3 trending AI projects on Hugging Face, once again ranking ahead of GLM-5.2. Last month it dominated GitHub and Hugging Face trend charts. Most people thought the hype was over. Then a Turing Award winner shared it... and it shot right back to the top. Now it has: โข 16.5K+ GitHub stars โข 2.24M+ Hugging Face downloads โข Still growing fast. The biggest problem with traditional OCR is that it processes documents page by page. That breaks multi-page tables, loses reading order, and struggles with long PDFs. Unlimited-OCR takes a different approach. Using its R-SWA architecture, it processes dozens of pages as a single continuous document while keeping KV Cache memory constantโso GPU memory doesn't keep increasing as documents get longer. Independent tests showed it could parse a 100-page PDF in one run while maintaining an error rate below 0.11 even after page 40, where many OCR systems begin to fail. Despite all that, it's only a 3B model, runs locally on consumer hardware, and is completely free. If you regularly work with research papers, contracts, books, or scanned PDFs, this could replace expensive OCR tools and eliminate manual page splitting. Repo๐




@socialwithaayan ยท
I found an AI tool yesterday that made me put my headphones back on three times. It's called JazzCat. An anonymous AI music model showed up on Hugging Face with no explanation. No company page, no launch post, no one claiming credit for it. Two features: ๐ต Text-to-Song โณ You describe the vibe, genre, mood, and instruments โณ JazzCat generates an entire original song from scratch โณ Lyrics, vocals, and full production come back from a single text prompt ๐ต AI Cover โณ You give it a reference track โณ It generates an AI vocal cover โณ I played one for a friend and he genuinely thought it was a real artist Two features and both of them are unreasonably good. There's no $50/month subscription. No enterprise tier. No sales team in your DMs. It's a Hugging Face demo that outperforms tools backed by actual venture funding. I keep finding that the most impressive AI work in 2026 comes from people who never announce anything. They just ship and let the output speak for itself

@lancedb ยท
@dlthub and @huggingface ๐ค just shipped a clean way to ingest Hugging Face datasets into LanceDB ๐ Query datasets over hf:// with DuckDB, stream them in batches, and load them into LanceDB with embeddings generated during ingest. The result is a simple Python path from Hub dataset to a searchable, explorable table.

@ihteshamali ยท
A million new AI models in one quarter. That is what just happened on Hugging Face. While headlines obsess over Anthropic and OpenAI, the open source ecosystem quietly logged that number in three months. Inference providers and neoclouds are seeing explosive growth. American startups and small companies are adopting open source at a surging rate. Clem Delangue's framing was the part that stuck with me. He said this is not about renting intelligence through an API anymore. It is about owning it. Then he connected it to something more personal. He has two daughters. He said when he thinks about a robot interacting with his kids, he does not want to trust a black box controlled by one megacorp. He wants transparency into how it was built and how it decides to act. That is why Hugging Face built their own robot, Reachy Mini. Over 10,000 shipped worldwide in a few months. While the industry fights over who gets to call frontier models dangerous, the open ecosystem is just quietly shipping. Watch the full interview on @technology
@_avichawla ยท
Hugging Face meets Claude! I built a @huggingface fine-tuning studio that lets you fine-tune any LLM directly from Claude. The app connects to the HF Hub for model and dataset search. It handles chat template formatting for the training data, and lets you configure LoRA rank, quantization, batch size, and learning rate directly from Claude. Training runs on HF's GPU infra via AutoTrain. Once training finishes, you can also chat with your fine-tuned model (or any other LLM on HF) directly from Claude. The studio's built with @manufact's mcp-use SDK, an open-source full-stack framework to build MCP Apps for Agents. In mcp-use, any MCP tool can be associated with a UI. You define a tool handler, create a React component, and the mcp-use framework handles the tool registration, prop mapping between server and widget, bundling, and hot reload during development. The widgets follow the MCP Apps standard, inspired by OpenAI's Apps SDK. Claude is an example, but you can render them as interactive UI elements in any conversational MCP client that supports it. Similarly, this pattern works for any workflow you want to bring inside a chat, like eval dashboards, dataset explorers, or model comparison tools. I have shared my fine-tuning studio repo in the replies!
@socialwithaayan ยท
Hugging Face has twelve free AI courses on one page, and most people bounce off because nobody tells them what order to take them in. Here is the order that makes sense. 1. LLM Course. Start here. Large language models and the Hugging Face ecosystem underneath everything else. 2. Agents Course. The main event. smolagents, LlamaIndex, LangGraph, then a final challenge where your agent goes on a student leaderboard. 3. Context Course. Context engineering for code agents. Take it right after agents, while the pain of a bloated context is still fresh. 4. Then branch to whatever you actually build. Diffusion for images. Audio for speech. Computer Vision. Deep RL if you want agents that learn. Robotics if you want LeRobot and real hardware. a smol course for post-training. ML for Games and ML for 3D if that is your lane. Open-Source AI Cookbook whenever you need a working notebook instead of a lecture. The certification is where it stops looking like a normal free course. The Agents course page says it plainly. The certification process is completely free. Audit it and tell nobody, or complete Unit 1 for a fundamentals certificate, or Unit 1 plus a use case assignment plus the final challenge for a certificate of completion. No deadline. The Audio course works the same way, 3 of 4 assignments for completion and 4 of 4 for excellence. Everything hands-on runs in pre-configured Spaces. Basic Python, an internet connection, a free account. That is the entry cost. Bonus Unit 3 of the Agents course is building an agent that plays Pokemon battles, which is not what you expect from a curriculum this serious.


@seelffff ยท
hugging face made a free course where you build AI agents - and then your agent competes against other students' agents on a public leaderboard. most people never opened it: it's not theory with a quiz at the end. it's build, ship, get ranked: part 1 - how an agent actually thinks โ tools, thoughts, actions, observations - the loop underneath every agent โ LLMs, messages, special tokens, chat templates: what the model actually receives โ your first working agent using plain python functions as tools part 2 - the frameworks people really use โ smolagents, LlamaIndex, LangGraph - all three, not just the trendy one โ agentic RAG: the use case every company is trying to build right now โ final project: build an agent for a benchmark and put it on the student leaderboard and the bonus units are where it gets unfair: โ fine-tuning a model for function-calling โ agent observability and evaluation - the part that separates a demo from a product โ and an agent that plays pokemon battles certificates are free. no deadline. 3-4 hours a week. you need basic python and a free account. this is a full agent curriculum with a leaderboard, given away, while people sell "agent bootcamps" for $500. fundamentals. frameworks. RAG. evals. a ranked final project. save it - "i can build agents" is worth more than "i know which model is best"
@mark_k ยท
MiniMax H3 is now open on Hugging Face! ๐ฅ @Hailuo_AI released the weights for H3, their general-purpose omni-modal generative model. It unifies text, images, video, and audio in one context and generates video with native stereo audio, up to 15 seconds at 2K. Key points: - Strong instruction following, accurate text and brand rendering, and video-to-video motion transfer - Aimed at advertising, branding, e-commerce, product design, and gaming - Competitive pricing, with 2K costing less than a third of mainstream models - 33B-parameter Omni-Transformer conditioned on Qwen3-VL-32B Two main checkpoints available: FL2VA and Ref2VA. Local inference works with Diffusers, SGLang, vLLM, and native ComfyUI support. One of the strongest open video models released so far.

@DataChaz ยท
AN @HUGGINGFACE ENGINEER JUST CRUNCHED 1.2M COMMONCRAWL PAGES FOR UNDER $1 ๐คฏ Zero Slurm cluster required. How? DataTrove now runs natively on Hugging Face Jobs. โ Same pipeline code, just a simple executor swap โ Cloud tasks fan out automatically โ CPU filtering + GPU inference all in one chain ๐ฅ Mind-blowing efficiency. Details in ๐งต โ

@ATechAjay ยท
Need to build an AI that understands images? SenseNova-Vision is an open-source vision foundation model that unifies many computer vision tasks behind a single natural-language interface. You usually end up with something like this: โ One model for object detection. โ One model for segmentation. โ One model for OCR. โ One model for depth estimation. โ One model for 3D reconstruction. Building computer vision apps has always felt like assembling LEGO pieces from five different boxes. Different APIs. Different output formats. Different pipelines. Maintaining everything becomes harder than building the product itself. Instead of stitching together multiple specialist models, SenseNova-Vision lets you describe the task in natural language. โ Detect every person โ Segment the road โ Read the text โ Estimate depth โ Reconstruct this object in 3D Same model. Same interface. Same set of weights. The interesting part isn't that it supports many vision tasks. It's that vision starts behaving like a foundation model. Instead of treating detection, segmentation, OCR, and 3D as separate AI systems, they're all expressed through natural language prompts. That means less time wiring models together... โฆand more time building products. Even better, the entire project is open source: โญ Model โญ Code โญ 50M-example training dataset โญ Benchmarks โญ Technical report โญ Interactive demo If you're building AI products, this is worth checking out. And if you appreciate open-source work like this, consider dropping the repo a โญ๏ธ GitHub: https://t.co/9T5tOvzis8 ๐ค Hugging Face: https://t.co/dfCHgYbmD8
@HuggingPapers ยท
Allen AI just released a rubric-based dataset for scholarly research on Hugging Face 10k+ examples featuring weighted evaluation criteria, multi-turn conversations, and instructions for training models on academic tasks. https://t.co/Zux6GZmUVQ
@jeffboudier ยท
Hugging Face just shipped Storage Buckets + hf-mount What AI-native storage for AI builders looks like: ๐๏ธ 2x faster than S3 on a 410GB upload - see @LoubnaBenAllal1 post below ๐ปMount repos as a drive on any machine. Stream data as needed, save your HDD. ๐ Connect large datasets to annotation tools like @nomadicai - see article below If you're training models and dealing with data at scale, try HF Storage Buckets on your next run ๐ https://t.co/jO7ST4OVcO
@smratitiwa86867 ยท
Local LLMs just hit a whole new level ๐คฏ This Hugging Face release is actually insane: "gpt-oss-20b-tq3" An official 20B+ parameter MoE model from OpenAIโฆ quantized to 3-bit with TurboQuant + optimized with MLXโฆ โฆand now it runs smoothly on a normal 16GB MacBook. ๐ป No server. No cloud bill. No internet needed. Everything stays fully local. A few months ago this wouldโve needed a high-end GPU setup. Now an M-series Mac can handle it. โข 131K context window โข Fully offline + private โข Great for chat, writing, and coding โข 60โ80 tok/s decoding speed โข No monthly subscription Running top-tier open-source LLMs directly on a laptop doesnโt even feel real anymore

@arsh_goyal ยท
NVIDIA has surpassed Google with nearly 4000 contributors on Hugging Face. It is now the number one. A move helping it set its positioning from a chips company to an operating system for AI company. I believe they now want everyone to build on NVIDIA so that they are in the middle no matter which model or chip wins seeing competition from Google and Amazon building chips. Great move, what do you think?
@_vmlops ยท
NVIDIA + HUGGING FACE JUST OPENED UP THE ROBOT FOUNDATION MODEL STACK they're bringing Isaac GR00T 1.7 and Isaac Teleop into LeRobot, with Cosmos 3 coming next โ Isaac Teleop: open framework to capture human demos and standardize robotics datasets โ Isaac GR00T 1.7: first open, commercially viable robot foundation model, built for post-training and fine-tuning โ Cosmos 3 (soon): world model to generate synthetic robotics data when real data is scarce โ backed by 15M+ downloads of NVIDIA's open physical AI dataset, 350K+ trajectories, 57M grasps connects NVIDIA's 3M robotics devs with Hugging Face's 16M AI builders in one open pipeline robotics just got its "open weights" moment

@JeremyCMorgan ยท
Hugging Face and NVIDIA published an end-to-end embedding fine-tuning pipeline for RAG. Synthetic QA generation, hard negative mining, ONNX/TensorRT export, deployed behind an OpenAI-compatible API. Recall@60 went from 0.751 to 0.951 in one case. Requires 80GB VRAM for fine-tuning. https://t.co/MVNi4175JB
@boyuan_chen ยท
The most important AI release yesterday wasn't Gemma 4. Hugging Face shipped TRL v1.0. Their post-training library has been the open-source default for SFT, DPO, GRPO, and reward modeling for years. 3 million monthly downloads. 130,000+ public models trained on earlier versions. Unsloth, Axolotl, LlamaFactory all build on it. Until this week, it was a research codebase with research-grade stability. Things broke between releases. Production teams kept internal forks to dodge upstream surprises. v1.0 adds a unified CLI, config-driven workflows, and stable APIs with deprecation guarantees. A team that needed distributed RL expertise to run custom GRPO can now do it from a config file. Post-training compute runs roughly 30x pre-training in serious pipelines now. The phase everyone treated as fine-tuning cleanup consumes the majority of GPU spend. The open-source tooling for it was held together with custom scripts until yesterday. Same day TRL shipped, Google dropped Gemma 4 under full Apache 2.0 and Alibaba launched Qwen 3.6-Plus. Both strong enough that teams will actually download and adapt them. Gemma's 26B MoE runs quantized on a MacBook Air. Qwen 3.6-Plus ships with native Anthropic API compatibility for agent stacks. Every team that grabs one of these needs SFT and preference optimization to make it useful for their workload. Production-grade post-training tooling and two base models worth running it on, all landing in the same 24-hour window. Anyone doing alignment work with custom scripts and internal patches just ran out of excuses.
@free_ai_guides ยท
Hugging Face's database engineer just explained how they serve 3 million AI models without the platform melting. You might know the name from last week's news. This is the company OpenAI's rogue agent hacked to steal benchmark answers. Here's what they do the other 364 days: host the world's largest open AI model library, 14 million users, 30% of the Fortune 500. This 21-minute talk is the machine underneath: 00:00 - Scaling the Hub from 20K to 3 million models 03:57 - The point where regex search dies 07:55 - Inside a single "llama" search 13:00 - The 7-node cluster with a hidden analytics node 16:42 - Sharding for the next 10x 18:14 - Autoscaling from 10 to 500 pods Every chapter has numbers you can steal for your own stack. Watch it, then read the guide on investing in AI below.
@mjamilmoughal ยท
Moonshot just open-sourced Kimi K3, the largest open-weight model ever. 2.8T-parameter MoE (โ104B active), 1M context, native vision. Weights dropped yesterday on Hugging Face. It sits right behind Claude Fable 5 and GPT-5.6 Sol on independent benchmarks, and leads some coding arenas. This is frontier-scale open weights. The kind that actually shifts the landscape. Reality check: itโs massive. Even multiple DGX Sparks canโt run the full thing locally. But itโs live on Ollama Cloud (kimi-k3:cloud) with Pro or Max plans. Expect heavy traffic and slower responses for a while. Open models this strong change the calculus. The real edge now isnโt just access, itโs knowing exactly which model wins at which job.

@_vmlops ยท
Hugging Face just made its MCP server a lot smarter โช๏ธ new hf_fs tool unifies repos, storage, docs, and papers into one searchable interface โช๏ธ full Hub navigation now costs just over 1,000 tokens โช๏ธ sandboxes add secure execution for dataset analysis, model training, and Space creation https://t.co/7COmapUBGO
Best Hugging Face tweets
Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.
$0 today ยท Cancel anytime
Browse all tweet collectionsKeep exploring