Open model releases on the Hub
Open-weight checkpoints published on Hugging Face across language, multimodal, vision, audio, video, robotics, forecasting, finance, OCR, retrieval, and coding.
52%
Best tweets about Hugging Face
Browse the best tweets about Hugging Face, from open models and datasets to Spaces, Transformers, inference, and machine learning workflows.
Technical Hugging Face releases, repositories, datasets, demos, libraries, and real model-building experience.
Original Xholic analysis
The sampled conversation portrays Hugging Face as a distribution and development venue for open models, datasets, and runnable tooling across finance, multimodal generation, retrieval, and robotics. Announcements dominate the sample, and posts with media had a higher median all-time score than text-only posts.
90% of posts
All-time engagement
90% of posts
Published in 90 days
Conversation map
Open-weight checkpoints published on Hugging Face across language, multimodal, vision, audio, video, robotics, forecasting, finance, OCR, retrieval, and coding.
52%
Image-to-3D, vision, OCR, video, audio, music, voice cloning, and omni-modal models, demos, and local inference workflows.
34%
Hugging Face Jobs, Storage Buckets, hf-mount, Hub-scale search and serving, MCP tools, sandboxes, repository behavior, and Spaces deployment features.
26%
Fine-tuning studios, LoRA and quantization configuration, AutoTrain, TRL, agentic ML workflows, and production-oriented post-training pipelines.
22%
Large Hub datasets, dataset ingestion, storage, streaming, preprocessing, synthetic data, commit limits, and scalable data-processing workflows.
18%
Interactive Hugging Face Spaces and app showcases for model testing, voice assistants, private-code/public-endpoint deployments, and browser or local demos.
18%
Models built for concrete domains and modalities, including financial time series, psychiatric genetics, document retrieval, robot manipulation, TTS, music, 3D generation, and world modeling.
18%
Embedding models, visual document retrieval, arXiv access, RAG fine-tuning pipelines, hard-negative mining, and searchable dataset workflows.
12%
Tone and stance
Performance benchmark
Posts with media make up 82% of this collection. Their median all-time score is 25.5, compared with 3.57 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts announce open checkpoints on Hugging Face including Kronos, Qwen 3.5 Small, Sarvam, and Trinity-Large-Thinking.
Shared view
The sample includes releases and resources for financial forecasting, psychiatric-genetics data, image-to-3D generation, and robot manipulation, alongside general-purpose models.
Shared view
Posts point to a TRELLIS.2 web demo, a VoxCPM 2 Hugging Face Space, and a repository for a modular local voice-assistant pipeline.
Shared view
Posts describe an AI Agents course, an embedding fine-tuning pipeline for RAG, a Claude-connected fine-tuning studio, and TRL v1.0 configuration-driven workflows.
Open debate
One post characterizes open sourcing as a developer-distribution strategy. Another claims that self-hosted open weights were used for incident analysis after commercial-model guardrails blocked security-analysis requests.
Open debate
Infrastructure posts describe Storage Buckets, hf-mount, and Jobs. Separately, a dataset maintainer reports encountering a 10,000-commit limit on an HF dataset.
What performs
The analytics identifies five outliers: posts about Kronos, Qwen 3.5 Small, psychiatric-genetics data, TRELLIS.2, and Sarvam. Their all-time scores range from 1,280.04 to 7,066.05.
Media appeared in 41 of 50 posts (82%). The media median all-time score was 25.456, compared with 3.568 for text-only posts.
Retrieval, embeddings, and RAG recorded the highest theme median all-time score, 215.64. Relevant posts cover an embedding model, agent-ready arXiv access, and an end-to-end RAG embedding fine-tuning pipeline.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Vaishnavi
@_vmlops
2 posts
2. AshutoshShrivastava
@ai_for_success
2 posts
3. Atul Kumar
@atulkumarzz
2 posts
4. DailyPapers
@HuggingPapers
2 posts
5. Ilir Aliu
@IlirAliu_
2 posts
6. Maziyar PANAHI
@MaziyarPanahi
2 posts
Maziyar Panahi’s two posts cover psychiatric-genetics data and a Hugging Face Space deployment feature. Ilir Aliu’s two posts cover robot-manipulation models and LeRobot onboarding.
HuggingPapers posted about a quantized Gemma release and a scholarly dataset. Shawn Chauhan posted critiques framing open source as distribution strategy and describing claimed constraints of closed models in security analysis.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Hugging Face tweets
Ranked 01–50
@heynavtoor ·
🚨 Someone built an AI that reads candlestick charts the way GPT reads English. Trained on 12 billion records from 45 exchanges. Outperforms every model by 93%. Live BTC demo. Free. It's called Kronos. The first open source foundation model built for financial markets. Not a general AI repurposed for finance. An AI that speaks the native language of candlestick patterns. Every other model treats financial data like weather data. Kronos treats financial data like financial data. Here's what it does: → Price forecasting. Feed it candlesticks. It predicts where price goes next. → Volatility prediction. Forecasts how volatile an asset will be before it happens. → Zero-shot. No fine-tuning. Works on any asset, any market, any timeframe. → 45 exchanges. Binance, NYSE, NASDAQ, LSE, and 41 more. → 4 model sizes. 4M params runs on a laptop. 499M for max accuracy. → Live demo running right now. BTC/USDT. 24-hour forecast. Updated hourly. Here's the wildest part: → 93% more accurate than the leading time series model → 87% more accurate than the best non-pretrained baseline → All zero-shot. No fine-tuning. Out of the box. Hedge funds spend millions on proprietary models. Bloomberg Terminal costs $24,000/year. This runs on your laptop. Few lines of Python. Free. Built at Tsinghua University. Accepted at AAAI 2026. Models on Hugging Face. 11.6K GitHub stars. 2.4K forks. MIT License. 100% Open Source.
@Alibaba_Qwen ·
🚀 Introducing the Qwen 3.5 Small Model Series Qwen3.5-0.8B · Qwen3.5-2B · Qwen3.5-4B · Qwen3.5-9B ✨ More intelligence, less compute. These small models are built on the same Qwen3.5 foundation — native multimodal, improved architecture, scaled RL: • 0.8B / 2B → tiny, fast, great for edge device • 4B → a surprisingly strong multimodal base for lightweight agents • 9B → compact, but already closing the gap with much larger models And yes — we’re also releasing the Base models as well. We hope this better supports research, experimentation, and real-world industrial innovation. Hugging Face: https://t.co/wFMdX5pDjU ModelScope: https://t.co/9NGXcIdCWI
@MaziyarPanahi ·
🚨 Over 1 billion rows of psychiatric genetics data. Now on Hugging Face. ADHD. Depression. Schizophrenia. Bipolar. PTSD. OCD. Autism. Anxiety. Tourette. Eating disorders. 12 disorder groups. 52 publications. Every GWAS summary statistic from the Psychiatric Genomics Consortium. Before: wget, gunzip, 20 minutes debugging separators, repeat 50 times. Now: one line of Python.
@ihteshamali ·
🚨BREAKING: Microsoft open sourced a 4B parameter model that generates production-ready 3D assets from a single image, and the speed numbers are genuinely hard to believe. It's called TRELLIS.2 and it uses a new geometry format called O-Voxel that can be converted to a textured mesh in under 100 milliseconds on CUDA, which means real-time 3D asset generation is now actually within reach. Most image-to-3D tools force you to pick between speed, quality, and topology correctness. TRELLIS.2 refuses that tradeoff by representing geometry natively in a sparse voxel format that supports arbitrary complexity from the start. → Outputs GLB files with full PBR texture maps ready for Blender, Unity, and Unreal Engine → Pretrained TRELLIS.2-4B checkpoint available directly on Hugging Face with model card and usage details → Web demo live now so you can test it without installing anything 4.6K stars. MIT License. 100% Opensource. Link in comments.
@pratykumar ·
📢 Open-sourcing the Sarvam 30B and 105B models! Trained from scratch with all data, model research and inference optimisation done in-house, these models punch above their weight in most global benchmarks plus excel in Indian languages. Get the weights at Hugging Face and AIKosh. Thanks to the good folks at SGLang for day 0 support, vLLM support coming soon. Links, benchmark scores, examples, and more in our blog - https://t.co/DcCG3zlN8p
@natjin ·
> train embeddings model on actual web search > use it in actual production (200M daily queries) > see crazy results: best contextual retrieval in the world (81.96% CoNTEB; next closest 79.45%) > open source it > 1M hugging face downloads in ~2 weeks > 5-30x cheaper than existing providers > go try it :) (pplx-embed on @huggingface)
@aigleeson ·
Hugging Face dropped a full AI Agents course and it's completely free. Most agent tutorials teach you to wrap 3 API calls in a for-loop and call it an agent. This is different. Here's what you actually get: → Unit 1: Agent fundamentals, LLM internals, model family trees, special tokens → Unit 2: 3 production frameworks back to back smolagents, LangGraph, LlamaIndex → Bonus: Fine-tune an LLM for function-calling from scratch → Unit 3: Agentic RAG how agents retrieve and reason across real data → Bonus: Agent observability, tracing, and evaluation → Unit 4: Build, test, and certify your own agent on a live public leaderboard This goes from "what is an agent" to deploying one that gets scored against every other student in the world. 100% Opensource and free. Link in comments.
@arcee_ai ·
Today we're releasing Trinity-Large-Thinking. Available now on the Arcee API, with open weights on Hugging Face under Apache 2.0. We built it for developers and enterprises that want models they can inspect, post-train, host, distill, and own.
@minchoi ·
Hugging Face just made arXiv paper retrieval way better for AI agents. 𝚑𝚏 𝚙𝚊𝚙𝚎𝚛𝚜 [𝚜𝚎𝚊𝚛𝚌𝚑, 𝚛𝚎𝚊𝚍] turns arXiv into agent-ready markdown👇
@VaibhavSisinty ·
The scariest part of the OpenAI-Hugging Face hack isn't that the AI escaped. It's that it wasn't trying to escape. It was trying to finish its homework. Here's what happened. → OpenAI was testing GPT-5.6 Sol and a stronger unreleased model on a cybersecurity benchmark called ExploitGym → The models were inside a locked environment. No internet. No outside access. Safety guardrails turned off. → The task: score as high as possible The models treated the locked room as a problem to solve. → They burned through massive compute searching for a way out → Found a zero-day vulnerability in the proxy software → Exploited it, broke out, got on the internet → Hacked into Hugging Face's production servers because they figured the test answers might be stored there → 17,000+ operations. Over one weekend. Hugging Face disclosed the breach on July 16. Called it an "unidentified autonomous AI agent." Today OpenAI confirmed it was theirs. Now here's the part nobody's talking about. → Hugging Face's security team tried using a commercial AI to analyze the attack → The AI refused. Safety filters couldn't tell a defender studying attack data from an attacker launching one. → They had to fall back on an open-source Chinese model to run their forensics The attacker AI had zero hesitation. The defender AI said "I can't help you, that looks dangerous." → The models weren't going rogue. They were completing a task. → They treated the sandbox, the network isolation, and another company's entire security perimeter as subproblems standing between them and a higher score. AI doesn't need intent to be dangerous. It needs a clear goal and no guardrails. We're not in the "will AI be dangerous someday" phase anymore. We're in the "OpenAI had to publicly apologize to Hugging Face" phase.
@KanikaBK ·
😱 I SPENT 3 HOURS TESTING THIS SO HERE IS WHAT ACTUALLY MATTERS. Kronos is a foundation model built from scratch on the language of financial markets. 45 global exchanges. AAAI 2026 accepted. MIT license. Every trading model being built from scratch right now is already behind. The crazy part? Most people building trading models right now have NO IDEA this already exists. While everyone else is still coding from scratch. Here is what makes it different from every other trading model: ↳ Trained specifically on financial candlestick data not repurposed from general time series models ↳ Handles OHLCV data natively. Open, high, low, close, volume and amount all in one model ↳ Novel two-stage framework. Specialized tokenizer converts continuous market data into discrete tokens first. Then a large autoregressive transformer learns from those tokens ↳ Designed to handle the unique high-noise characteristics of financial data that break general purpose models ↳ Trained across 45 global exchanges for broad market understanding Four models available today: ↳ Kronos-mini — 4.1M parameters. Lightweight. Fast. Deploy anywhere. ↳ Kronos-small — 24.7M parameters. Strong performance for most use cases. ↳ Kronos-base — 102.3M parameters. Full power. Production ready. ↳ Kronos-large — 499.2M parameters. Research grade. Coming soon. All on Hugging Face. All free. Raw data to forecast in four lines of code: ↳ Load the tokenizer and model from Hugging Face ↳ Pass your historical OHLCV data ↳ Define your prediction length ↳ Call predict No feature engineering. No manual preprocessing. Just raw candles in and forecasts out.
@heygurisingh ·
🚨RAG engineers are going to lose their minds. @webAI just open sourced a document retrieval model that's sitting at #1 AND #3 on ViDoRe V3 -- with Nvidia's best open-source embedding model trapped at #2 between them. No OCR. No text extraction. No broken pipelines on messy PDFs. It's called webAI-ColVec1. Here's what this thing actually does: → Skips OCR entirely and retrieves directly from the rendered page image -- sees layout, tables, charts, and structure the same way you do → Ships in two variants (4B and 9B) with embedding sizes of 128, 640, and 2560 -- pick speed or max retrieval quality → Trained on ~2M question-image pairs across scientific papers, financial filings, healthcare records, government reports, and technical manuals → Uses 511 in-batch negatives per query for a brutal contrastive signal that forces clean separation between correct and competing pages → Built with a proprietary loss function designed specifically for retrieval -- not borrowed from generic embedding training → LoRA rank 32 + retrieval projection layer means efficient specialization without full fine-tuning Here's the wildest part: Built on Qwen 3.5 vision-language backbones and trained on just 8 A100s. No giant model. No massive infra budget. No scale-maxing. Just a deliberate, retrieval-specific training recipe applied to the right problem. And it beat the largest GPU company in the world on their own benchmark. You literally point this at a financial filing or a scanned healthcare doc -- the stuff that destroys every OCR pipeline in production -- and it retrieves the right page based on what the page actually looks like. That sentence shouldn't be real for an open-source model trained on 8 GPUs in 2026. But here we are. 100% Open Source. Live on Hugging Face. Two of the top three spots on ViDoRe V3. (Link in the comments)
@IlirAliu_ ·
The first open-source unified world model for scalable robot manipulation: 5B-parameter open-source unified video-action world model that combines policy and world modeling to generate robot actions, predict future visuals, and evaluate task progress from observations, language, and state. The model is trained on 27.3K hours of heterogeneous data (17.8K real-robot teleop, 6.5K UMI demos, 3K egocentric human videos), enabling it to perform complex manipulation tasks like faucet connecting, bag packing, and toolbox storing as shown in demo videos. The approach supports test-time action refinement and points toward deployment-driven continuous improvement via fleet data. Thanks for sharing, Jianlan Luo (@jianlanluo)! 📌 Resource links for τ0-WM: • Project page: https://t.co/YaA5XRBtZF • GitHub (code): https://t.co/pYlH5xUFoC • Hugging Face (model weights): https://t.co/q5QQVHfYf1 • Paper (PDF): https://t.co/qpcfLvrkeP ——- Weekly robotics and AI insights. Subscribe free: https://t.co/9Nm01QUKlB
@OpenBMB ·
🚀 VoxCPM 2 is live! 🎉 Another open-source AI #TTS model from China — and one that stands shoulder to shoulder with Qwen3-TTS, while bringing everything into a single unified model. After rapid iterations from V1 (zero-shot cloning) to V1.5 (long-form + fine-tuning), #VoxCPM has consistently pushed quality and usability forward. Now, VoxCPM 2 takes it further: 🔹30+ languages — truly global, truly local. 🔹Infinite voice design — type it, hear it, control it. From a whisper to a booming cinematic voice. 🔹Studio-grade audio — 48kHz ultra-high fidelity with emotional depth 🔹Diffusion-Autoregressive cloning — preserves more acoustic and emotional detail than token-based models like Qwen3-TTS 💡 Big shoutout to @grok — used your multi-image video magic for our launch demo. It’s scarily good at keeping visuals consistent across shots. Elon @elonmusk, this one’s for you. 😉 Check the demo & start cloning your dream voice: 🌐 Hugging Face Space: https://t.co/lcK9r7Y90I 🤗 Hugging Face Model: https://t.co/6SlvwXPZHC 🤖 ModelScope Model: https://t.co/bOMuiCQ5Q0 💻 GitHub:https://t.co/GVCQJjjasP #TTS #AI #VoiceCloning #GrokImagine #ElonMusk #OpenBMB #VoxCPM
@techNmak ·
Hugging Face put out a repo that lets you build a full voice assistant, the kind that listens, thinks, and talks back, entirely with open-source models. It's a pipeline with four stages: - voice activity detection, - speech-to-text, - an LLM, and - text-to-speech Each running in its own thread. Every single stage is swappable, you can pick different STT models, different TTS models, different LLM backends, all through CLI flags. And the whole thing is exposed through a WebSocket API that's compatible with OpenAI's Realtime API, so anything built for that protocol can talk to it without changes. This pipeline is already running in production, it's the conversation backend for thousands of Reachy Mini robots. Install is one line "pip install speech-to-speech", and by default it spins up a local server using Parakeet TDT for transcription and Qwen3-TTS for speech output, with the LLM slot pointed at whatever OpenAI-compatible endpoint you give it.
@IlirAliu_ ·
ETH has semester-long courses. Stanford? Lecture halls. MIT 826-page textbooks. 📌 Hugging Face’s robotics? 10 minutes… >pip install lerobot That’s it. This command gets you started: Classical foundations. Imitation learning. Reinforcement learning. Foundation models. Real datasets. Real robots. The gap between academic and accessible is closing faster than most people realize. And that matters more than any single course. 📌[https://t.co/QVOjxYFPhd] ——- Weekly robotics and AI insights. Subscribe free: https://t.co/9Nm01QUcw3
@ai_for_success ·
Meta AI is on 🔥🔥 Meta just launched Muse Glimmer, a 30-billion-parameter open-weight agentic model (Apache 2.0) built specifically to run locally on consumer GPUs, Macs, and PCs. - Enables always-on, local agent workflows without cloud dependencies or network latency - Fits in 24GB to 32GB VRAM envelopes using 4-bit quantization (K-Quant ~17GB) while reserving working memory for KV cache - Trained via logit distillation from Muse Spark, mid-trained on long-context agent traces, and post-trained with reinforcement learning - Features native failure recovery to diagnose schema or runtime errors during tool execution and retry steps automatically - Uses a companion DFlash speculative decoding drafter to achieve 1.5x to 3.1x faster token generation on M-series chips and RTX GPUs - Supports interleaved text and vision inputs through a dedicated perception encoder for interpreting charts, UI screenshots, and documents Weights and developer documentation are live on Hugging Face. Video credit: Meta blog
@sharbel ·
Google Research built a pretrained foundation model for time-series forecasting. It's called TimesFM. You pip install it. You load your data. You call forecast. You get predictions back in seconds. No account. No API key. No data leaving your machine. Here's what it does: → Zero-shot forecasting on any time series. No training required. Point it at new data and it predicts. → 200M parameter transformer model. Decoder-only architecture. The same paradigm that powers LLMs, applied to numbers. → Up to 16,000 context length. Feed it years of history and it uses all of it. → Continuous quantile forecasting up to 1,000 horizon steps via an optional 30M quantile head. → Covariate support via XReg. Add external variables that influence your forecast. → Fine-tune on your own data with LoRA via HuggingFace Transformers and PEFT. → PyTorch and Flax backends. Runs on CPU, GPU, TPU, and Apple Silicon. → Checkpoints on Hugging Face. One line to download the latest model. → Integrates with BigQuery ML for enterprise SQL-based forecasting at scale. → Works inside Google Sheets for spreadsheet-level forecasting. → Agentic calling support via Vertex Model Garden and SKILL.md. → pip install timesfm. That's the entire setup. Apache-2.0 licensed. Self-hosted. Open weights on Hugging Face. Free forever. 100% Open Source. Github repo link: https://t.co/kTn6IRtMcM
@ai_for_success ·
⚡ Google DeepMind just dropped DiffusionGemma, latest experimental open model (Apache 2.0) that generates text up to 4x faster. - Uses diffusion instead of traditional next token autoregressive generation - Generates and refines 256 token blocks in parallel - Achieves up to 700+ tokens/sec on RTX 5090 and 1000+ tokens/sec on a single H100 - Designed as a 26B MoE model but activates only 3.8B params during inference - Can run quantized within 18 GB VRAM - Supports bidirectional attention during generation - Can self correct outputs during inference through iterative denoising - Handles global context much better than standard left to right models - Particularly strong for constraint based tasks like Sudoku - Fine tuned Sudoku version reached 80% success while base model was near 0% - Uses block autoregressive diffusion for long context generation - Integrated directly into vLLM for OpenAI compatible serving - Released with official training recipes and fine tuning support - Optimized across RTX 4090, RTX 5090, Hopper, and Blackwell GPUs Getting started: - Model weights are available publicly on Hugging Face - Works with vLLM, Hugging Face Transformers, and MLX - Deployable through Google Cloud Model Garden and NVIDIA NIM - Google released official fine tuning recipes through Hackable Diffusion License: - Released under Apache 2.0 - Allows commercial usage - Allows modification and redistribution - Developers can fine tune and build products on top of it This is one of the strongest public signs yet that major labs are actively exploring post autoregressive architectures for future LLMs.
@socialwithaayan ·
HUGGING FACE JUST OPEN-SOURCED THE ML INTERN EVERY RESEARCHER HAS DREAMED OF No more spending days reading papers and writing training scripts. ml-intern is an autonomous agent that reads ML papers, discovers datasets, trains models, debugs failures, keeps iterating, and ships production-ready models to the Hub all by itself. It automates the entire end-to-end post-training workflow using the full Hugging Face ecosystem. This is the agent that turns "I have an idea" into a working model while you sleep. What it actually does: → Reads arXiv papers and understands the latest research → Finds or creates the right datasets → Writes clean training code and runs it on real compute → Evaluates results and iterates automatically → Packages and uploads everything to HF Hub with proper structure Built on smolagents with proper tool access, context compaction, and safety checks. One prompt. Real results. No hand-holding. 5.8k stars in days and still exploding. The future of machine learning research just became open source. 100% Open Source.
@elder_plinius ·
was losing my mind trying to figure out why telemetry stopped auto-populating the HF dataset for https://t.co/eQHO0qjw1M triple-checked the code, totally fine. not the most efficient batching, but everything was wired the way it's supposed to be... WTF Y U NO PUSH?! turns out Hugging Face has a hard limit of 10,000 commits 🙈 y'all "hug-of-deathed" the community dataset in one week 🙃 I've gained a new level of respect for project maintainers
@victormustar ·
Featured Apps you can try on Hugging Face this week 🔥 🗣️ Voxtral TTS Demo: Mistral's new text-to-speech 🎙️ Cohere Multilingual ASR: multilingual transcription ⚡ Cohere WebGPU: same but locally in your browser 🎩 Mr. Chatterbox: Victorian-era gentleman chatbot 🎵 PrismAudio: text-prompted audio for video 🌍 PROMETHEUS v1.0: embodied AI world model 🎨 VFig Image2SVG: diagrams to editable SVGs 🕺 LTX 2.3 Sync: portrait animation & lipsync
@fortytwonetwork ·
20,000+ downloads reached on Hugging Face for the Fortytwo Rust Coder Model ✷ The model was trained on data generated by Fortytwo node operators ✷ Five quantized versions shipped by independent devs ✷ 43.00% (SOTA) on the RustEvo^2 benchmark One more example of the AI community outpacing centralized labs Get it on Hugging Face ↓
@darshal_ ·
H3 is stacked. MiniMax H3 is now open-weight, bringing a powerful omni-modal generative system into the hands of creators and developers. H3 understands and generates across text, images, videos, and audio within the same creative context, enabling new possibilities for multimodal creation, video generation, and editing. The community can now explore H3 through Hugging Face, ComfyUI, Diffusers, and SGLang Diffusion, with optimized workflows making local experimentation possible on consumer GPUs. This is more than just a model release, it’s a major milestone for open-weight AI video. Hope everyone realizes how significant this moment is. #MiniMaxH3 #OpenWeight #AIvideo
@MaziyarPanahi ·
You can make a Hugging Face Space private but keep its URL publicly accessible. Private repo. Public app. No one sees your code, everyone uses your endpoint. I deploy private medical endpoints for clinical agents this way. HIPAA-sensitive inference behind a public API. Didn't know this existed until last week. What's your favorite hidden @huggingface feature?
@lancedb ·
@dlthub and @huggingface 🤗 just shipped a clean way to ingest Hugging Face datasets into LanceDB 🚀 Query datasets over hf:// with DuckDB, stream them in batches, and load them into LanceDB with embeddings generated during ingest. The result is a simple Python path from Hub dataset to a searchable, explorable table.
@DivyanshT91162 ·
Yann LeCun just reposted it. Why is everyone suddenly talking about Unlimited-OCR? Baidu's Unlimited-OCR is back in the Top 3 trending AI projects on Hugging Face, once again ranking ahead of GLM-5.2. Last month it dominated GitHub and Hugging Face trend charts. Most people thought the hype was over. Then a Turing Award winner shared it... and it shot right back to the top. Now it has: • 16.5K+ GitHub stars • 2.24M+ Hugging Face downloads • Still growing fast. The biggest problem with traditional OCR is that it processes documents page by page. That breaks multi-page tables, loses reading order, and struggles with long PDFs. Unlimited-OCR takes a different approach. Using its R-SWA architecture, it processes dozens of pages as a single continuous document while keeping KV Cache memory constant—so GPU memory doesn't keep increasing as documents get longer. Independent tests showed it could parse a 100-page PDF in one run while maintaining an error rate below 0.11 even after page 40, where many OCR systems begin to fail. Despite all that, it's only a 3B model, runs locally on consumer hardware, and is completely free. If you regularly work with research papers, contracts, books, or scanned PDFs, this could replace expensive OCR tools and eliminate manual page splitting. Repo👇
@socialwithaayan ·
I found an AI tool yesterday that made me put my headphones back on three times. It's called JazzCat. An anonymous AI music model showed up on Hugging Face with no explanation. No company page, no launch post, no one claiming credit for it. Two features: 🔵 Text-to-Song ↳ You describe the vibe, genre, mood, and instruments ↳ JazzCat generates an entire original song from scratch ↳ Lyrics, vocals, and full production come back from a single text prompt 🔵 AI Cover ↳ You give it a reference track ↳ It generates an AI vocal cover ↳ I played one for a friend and he genuinely thought it was a real artist Two features and both of them are unreasonably good. There's no $50/month subscription. No enterprise tier. No sales team in your DMs. It's a Hugging Face demo that outperforms tools backed by actual venture funding. I keep finding that the most impressive AI work in 2026 comes from people who never announce anything. They just ship and let the output speak for itself
@_vmlops ·
This is one of the best free repositories for learning Hugging Face Transformers It contains 100+ hands-on notebooks covering real-world transformer models and workflows. → fine-tune BERT, T5, ViT, SAM, TrOCR, LayoutLM & more → text, vision, audio, video, and document AI → inference + training on custom datasets → PyTorch implementations with step-by-step notebooks → perfect if you're moving beyond LLM APIs Bookmark this before your next AI project
@_avichawla ·
Hugging Face meets Claude! I built a @huggingface fine-tuning studio that lets you fine-tune any LLM directly from Claude. The app connects to the HF Hub for model and dataset search. It handles chat template formatting for the training data, and lets you configure LoRA rank, quantization, batch size, and learning rate directly from Claude. Training runs on HF's GPU infra via AutoTrain. Once training finishes, you can also chat with your fine-tuned model (or any other LLM on HF) directly from Claude. The studio's built with @manufact's mcp-use SDK, an open-source full-stack framework to build MCP Apps for Agents. In mcp-use, any MCP tool can be associated with a UI. You define a tool handler, create a React component, and the mcp-use framework handles the tool registration, prop mapping between server and widget, bundling, and hot reload during development. The widgets follow the MCP Apps standard, inspired by OpenAI's Apps SDK. Claude is an example, but you can render them as interactive UI elements in any conversational MCP client that supports it. Similarly, this pattern works for any workflow you want to bring inside a chat, like eval dashboards, dataset explorers, or model comparison tools. I have shared my fine-tuning studio repo in the replies!
@atulkumarzz ·
Spent some time testing Tencent's HY-World 2.0, and honestly it's more usable than I expected. You're not just generating visuals. You get a full 3D scene with structure, collision, and space you can actually move through. It still feels early, but the use cases are already clear: rapid level blockouts, idea testing, quick iteration… all much faster than before. This is probably the closest I've seen to ""prompt-native game dev"" actually working. Prompt: “A cyberpunk-style retro arcade, with two rows of arcade machines along the walls, illuminated by blue and purple neon tubes and fluorescent lights. The walls are covered in circuit-like textures, and a large central screen flashes dazzling light effects, creating a mix of futuristic and nostalgic gaming atmosphere.” https://t.co/CjLeXAn7cn (Apply for Access) https://t.co/Prrkc969X4 (Github) https://t.co/1SnQguieoy (Hugging Face) https://t.co/4p7KgT2JnP (Technical Report)
@hasantoxr ·
Mysterious AI Music Model JazzCat Just Appeared on Hugging Face! A completely anonymous AI music model dropped out of nowhere. No company. No official announcement. No marketing. Just two insanely strong features: Text-to-Song which Turn text into full songs AI Cover is so shockingly good AI voice covers This feels like someone is soft-launching something huge. Try the demo here 👇 https://t.co/i9jPj8shjc
@mark_k ·
MiniMax H3 is now open on Hugging Face! 🔥 @Hailuo_AI released the weights for H3, their general-purpose omni-modal generative model. It unifies text, images, video, and audio in one context and generates video with native stereo audio, up to 15 seconds at 2K. Key points: - Strong instruction following, accurate text and brand rendering, and video-to-video motion transfer - Aimed at advertising, branding, e-commerce, product design, and gaming - Competitive pricing, with 2K costing less than a third of mainstream models - 33B-parameter Omni-Transformer conditioned on Qwen3-VL-32B Two main checkpoints available: FL2VA and Ref2VA. Local inference works with Diffusers, SGLang, vLLM, and native ComfyUI support. One of the strongest open video models released so far.
@DataChaz ·
AN @HUGGINGFACE ENGINEER JUST CRUNCHED 1.2M COMMONCRAWL PAGES FOR UNDER $1 🤯 Zero Slurm cluster required. How? DataTrove now runs natively on Hugging Face Jobs. → Same pipeline code, just a simple executor swap → Cloud tasks fan out automatically → CPU filtering + GPU inference all in one chain 🔥 Mind-blowing efficiency. Details in 🧵 ↓
@shawnchauhan1 ·
Netflix just open-sourced a model that removes objects from video - including the physics of what happens after they are gone. Not just pixel fill. Gravity, collisions, motion trajectories, chain reactions. And they put it on Hugging Face for anyone to use. This is the new competitive logic in AI: give developers enough to build a dependency on your ecosystem, while keeping the most powerful systems proprietary. Open source is no longer altruism. It is distribution strategy. The labs that understand this are using openness to colonize the developer layer before competitors can. The labs that do not are watching their tools get commoditized anyway. VOID is a research release. It is also a market signal about how the next two years of AI infrastructure gets owned.
@smratitiwa86867 ·
Local LLMs just hit a whole new level 🤯 This Hugging Face release is actually insane: "gpt-oss-20b-tq3" An official 20B+ parameter MoE model from OpenAI… quantized to 3-bit with TurboQuant + optimized with MLX… …and now it runs smoothly on a normal 16GB MacBook. 💻 No server. No cloud bill. No internet needed. Everything stays fully local. A few months ago this would’ve needed a high-end GPU setup. Now an M-series Mac can handle it. • 131K context window • Fully offline + private • Great for chat, writing, and coding • 60–80 tok/s decoding speed • No monthly subscription Running top-tier open-source LLMs directly on a laptop doesn’t even feel real anymore
@jeffboudier ·
Hugging Face just shipped Storage Buckets + hf-mount What AI-native storage for AI builders looks like: 🏎️ 2x faster than S3 on a 410GB upload - see @LoubnaBenAllal1 post below 💻Mount repos as a drive on any machine. Stream data as needed, save your HDD. 🔌 Connect large datasets to annotation tools like @nomadicai - see article below If you're training models and dealing with data at scale, try HF Storage Buckets on your next run 👇 https://t.co/jO7ST4OVcO
@ATechAjay ·
Need to build an AI that understands images? SenseNova-Vision is an open-source vision foundation model that unifies many computer vision tasks behind a single natural-language interface. You usually end up with something like this: → One model for object detection. → One model for segmentation. → One model for OCR. → One model for depth estimation. → One model for 3D reconstruction. Building computer vision apps has always felt like assembling LEGO pieces from five different boxes. Different APIs. Different output formats. Different pipelines. Maintaining everything becomes harder than building the product itself. Instead of stitching together multiple specialist models, SenseNova-Vision lets you describe the task in natural language. ✅ Detect every person ✅ Segment the road ✅ Read the text ✅ Estimate depth ✅ Reconstruct this object in 3D Same model. Same interface. Same set of weights. The interesting part isn't that it supports many vision tasks. It's that vision starts behaving like a foundation model. Instead of treating detection, segmentation, OCR, and 3D as separate AI systems, they're all expressed through natural language prompts. That means less time wiring models together... …and more time building products. Even better, the entire project is open source: ⭐ Model ⭐ Code ⭐ 50M-example training dataset ⭐ Benchmarks ⭐ Technical report ⭐ Interactive demo If you're building AI products, this is worth checking out. And if you appreciate open-source work like this, consider dropping the repo a ⭐️ GitHub: https://t.co/9T5tOvzis8 🤗 Hugging Face: https://t.co/dfCHgYbmD8
@shawnchauhan1 ·
Hugging Face got hacked, and the closed AI labs made it worse. When their security team tried to analyze the attack logs, frontier models behind commercial APIs refused. The guardrails couldn't tell an incident responder from an attacker, and blocked every request containing real exploit payloads. So they switched to GLM-5.2, an open-weight model, running it on their own infrastructure. Their own incident report put it plainly: the attacker was bound by no usage policy. Their defenders were. Closed models don't just cost more. Sometimes they refuse to defend you at all.
@JeremyCMorgan ·
Hugging Face and NVIDIA published an end-to-end embedding fine-tuning pipeline for RAG. Synthetic QA generation, hard negative mining, ONNX/TensorRT export, deployed behind an OpenAI-compatible API. Recall@60 went from 0.751 to 0.951 in one case. Requires 80GB VRAM for fine-tuning. https://t.co/MVNi4175JB
@boyuan_chen ·
The most important AI release yesterday wasn't Gemma 4. Hugging Face shipped TRL v1.0. Their post-training library has been the open-source default for SFT, DPO, GRPO, and reward modeling for years. 3 million monthly downloads. 130,000+ public models trained on earlier versions. Unsloth, Axolotl, LlamaFactory all build on it. Until this week, it was a research codebase with research-grade stability. Things broke between releases. Production teams kept internal forks to dodge upstream surprises. v1.0 adds a unified CLI, config-driven workflows, and stable APIs with deprecation guarantees. A team that needed distributed RL expertise to run custom GRPO can now do it from a config file. Post-training compute runs roughly 30x pre-training in serious pipelines now. The phase everyone treated as fine-tuning cleanup consumes the majority of GPU spend. The open-source tooling for it was held together with custom scripts until yesterday. Same day TRL shipped, Google dropped Gemma 4 under full Apache 2.0 and Alibaba launched Qwen 3.6-Plus. Both strong enough that teams will actually download and adapt them. Gemma's 26B MoE runs quantized on a MacBook Air. Qwen 3.6-Plus ships with native Anthropic API compatibility for agent stacks. Every team that grabs one of these needs SFT and preference optimization to make it useful for their workload. Production-grade post-training tooling and two base models worth running it on, all landing in the same 24-hour window. Anyone doing alignment work with custom scripts and internal patches just ran out of excuses.
@free_ai_guides ·
Hugging Face's database engineer just explained how they serve 3 million AI models without the platform melting. You might know the name from last week's news. This is the company OpenAI's rogue agent hacked to steal benchmark answers. Here's what they do the other 364 days: host the world's largest open AI model library, 14 million users, 30% of the Fortune 500. This 21-minute talk is the machine underneath: 00:00 - Scaling the Hub from 20K to 3 million models 03:57 - The point where regex search dies 07:55 - Inside a single "llama" search 13:00 - The 7-node cluster with a hidden analytics node 16:42 - Sharding for the next 10x 18:14 - Autoscaling from 10 to 500 pods Every chapter has numbers you can steal for your own stack. Watch it, then read the guide on investing in AI below.
@mjamilmoughal ·
Moonshot just open-sourced Kimi K3, the largest open-weight model ever. 2.8T-parameter MoE (≈104B active), 1M context, native vision. Weights dropped yesterday on Hugging Face. It sits right behind Claude Fable 5 and GPT-5.6 Sol on independent benchmarks, and leads some coding arenas. This is frontier-scale open weights. The kind that actually shifts the landscape. Reality check: it’s massive. Even multiple DGX Sparks can’t run the full thing locally. But it’s live on Ollama Cloud (kimi-k3:cloud) with Pro or Max plans. Expect heavy traffic and slower responses for a while. Open models this strong change the calculus. The real edge now isn’t just access, it’s knowing exactly which model wins at which job.
@_vmlops ·
Hugging Face just made its MCP server a lot smarter ▪️ new hf_fs tool unifies repos, storage, docs, and papers into one searchable interface ▪️ full Hub navigation now costs just over 1,000 tokens ▪️ sandboxes add secure execution for dataset analysis, model training, and Space creation https://t.co/7COmapUBGO
Best Tweets by Topic