Best tweets about Multimodal AI

50 Best Tweets About Multimodal AI (2026)

Browse the best tweets about multimodal AI, including vision, audio, video, model architectures, benchmarks, applications, and developer experiments.

Multimodal models combining text, images, audio, or video, with concrete research, evaluations, workflows, and applications.

Creators
47
Updated

Top Multimodal AI tweets from 47 creators

Ranked 01–50

  1. 01

    @ihteshamali ·

    🚨 BREAKING: A research lab just released a 15B model that generates multilingual talking human videos with synced audio, beats every competitor in human evaluation, and runs in 38 seconds on one GPU. It's called daVinci-MagiHuman. The key insight is that every other model in

    Video thumbnail from Ihtesham Ali's postWatch video
    • 19Replies
    • 116Reposts
    • 696Likes
    • 43.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 02

    @techNmak ·

    Finally, a lightweight VLM that beats the giants at OCR. (1.7B parameters, SOTA on OmniDocBench) dots. ocr is a new multilingual document parser that proves you don't need massive models for perfect document understanding. Current SOTA models are often massive (72B+) or

    • 27Replies
    • 88Reposts
    • 708Likes
    • 31KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 03

    @dr_cintas ·

    This peanut-sized chinese model just dethroned Gemini at reading documents. It’s called glm-ocr. it’s a tiny 0.9b parameter vision-language model that is about to replace every expensive ocr api you use. → Handles text, tables, formulas, handwriting → Scored 94.62 on

    Video thumbnail from Alvaro Cintas's postWatch video
    • 17Replies
    • 68Reposts
    • 417Likes
    • 35.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 04

    @skalskip92 ·

    spent most of my day playing with GLM-OCR it's a 0.9B param vision-language model. supports 8K resolution, 8+ languages, and has built-in text, LaTeX, and table recognition modes. awesome! I tested it across different OCR tasks. starting with shipping container serial numbers.

    • 17Replies
    • 55Reposts
    • 817Likes
    • 271.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 05

    @jtdavies ·

    MLX + TurboQuant + Apple's M5 Max = 🔥 Some more details for you on the @Prince_Canuma MLX port of TurboQuant running on Apple's M5 Max. Rather than convert a PDF to text and then process it, I loaded the PDF (https://t.co/LjUDlWdJdD) as 150 dpi images. 20 pages at bf16 uses

    Video thumbnail from John T Davies 🇪🇺's postWatch video
    • 12Replies
    • 33Reposts
    • 374Likes
    • 47.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 06

    @arcinstitute ·

    Over 250 million protein sequences are known, but fewer than 0.1% have confirmed functions. Today, @genophoria, @BoWang87 & team introduce BioReason-Pro, a multimodal reasoning model that predicts protein function and explains its reasoning like an expert would.

    • 13Replies
    • 124Reposts
    • 527Likes
    • 61.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 07

    @Meituan_LongCat ·

    🔥 Introducing LongCat-Next: A Discrete Native Autoregressive Multimodal Model LongCat-Next integrates language, vision, and audio into a unified discrete autoregressive model, extending Next-Token Prediction to native multimodality and delivering industrial-strength performance

    • 10Replies
    • 66Reposts
    • 467Likes
    • 44.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 08

    @Yuchenj_UW ·

    I used Claude Computer Use/Dispatch yesterday. My feeling: It’s too damn slow! Posting a tweet takes me ~5 seconds (once I have the content). Claude took 70 seconds. Why? It controls the screen via a loop: take a screenshot → send to a huge remote multimodal model (opus 4.6) →

    • 138Replies
    • 32Reposts
    • 644Likes
    • 53.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 09

    @MattNiessner ·

    📢WorldAgents: 3D worlds only from 2D image models - without any training! We propose an agentic approach with a Director (VLM) to plan the scene, a Generator (Flux or NanoBanana) for new views, and a Verifier (VLM) for selection / 3D consistency. -> High-fidelity 3D worlds from

    Video thumbnail from Matthias Niessner's postWatch video
    • 6Replies
    • 46Reposts
    • 269Likes
    • 18.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 10

    @mlech26l ·

    A 450M VLM that actually runs on CPU. @liquidai LFM2.5-VL-450M: 28T tokens pre-trained (80x Chinchilla-optimal), SigLIP2 vision encoder, sub-1GB footprint, pure HF Transformers inference

    Video thumbnail from Mathias Lechner's postWatch video
    • 2Replies
    • 16Reposts
    • 166Likes
    • 11.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  1. 11

    @shushant_l ·

    I'm amazed most people still use Google Gemini like a basic chatbot. Here's how to unlock its full potential with the ultimate Google Gemini guide. --- 1. Google Gemini is Google's multimodal AI that understands text, images, audio, video, code, and documents. --- 2. It

    • 6Replies
    • 27Reposts
    • 119Likes
    • 5.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  2. 12

    @fahdmirza ·

    💥 Gemma 4 E2B + Hermes Agent + vLLM running fully local — zero cloud, zero cost ♠ and it's a complete multimodal AI stack on a single GPU 🚀 🔹Gemma 4 E2B served locally via vLLM 0.19.0 on NVIDIA A6000 🔹Text, vision and native audio transcription — all offline 🔹27 languages

    • 4Replies
    • 16Reposts
    • 139Likes
    • 10.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  3. 13

    @intern_lm ·

    🚀Meet InternVL-U: a lightweight 4B unified multimodal model that brings reasoning, generation, and editing into a unified framework. 🔥Built upon unified contextual modeling, modality-specific modular design, and decoupled visual representations, InternVL-U achieves a strong

    • 1Replies
    • 31Reposts
    • 151Likes
    • 20.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  4. 14

    @chutes_ai ·

    Moonshot trained a model on 15 trillion tokens of mixed vision and text data. The result scores 96.1 on AIME 2025 and 76.8 on SWE-Bench Verified. Model Spotlight: Kimi K2.5 by @kimi_moonshot 1T total parameters. 32B activated per token (MoE). 256K context. Vision baked into

    • 10Replies
    • 34Reposts
    • 255Likes
    • 15.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  5. 15

    @arankomatsuzaki ·

    Context Unrolling in Omni Models - A unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations - Enables Context Unrolling, where the model explicitly reasons across multiple modal representations

    • 3Replies
    • 25Reposts
    • 145Likes
    • 15.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  6. 16

    @OpenBMB ·

    Multimodal LLMs are supposed to "imagine" and reason in their latent space like humans. But is this "inner thought" actually happening, or is it just an illusion? 🧠 Today, we dive into new research by @TsinghuaNLP (OpenBMB member) and collaborators: A rigorous causal analysis

    • 16Replies
    • 15Reposts
    • 110Likes
    • 8.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  7. 17

    @drmapavone ·

    Jensen today announced Alpamayo 1.5 at #NVIDIAGTC! #Alpamayo 1.5 is a major update to Alpamayo 1—@nvidia’s open 10B-parameter chain-of-thought reasoning VLA model, first introduced at #CES. Built on the #Cosmos-Reason2 VLM backbone and post-trained with RL, it adds support for

    Video thumbnail from Marco Pavone's postWatch video
    • 7Replies
    • 30Reposts
    • 160Likes
    • 17.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  8. 18

    @ai_for_success ·

    Google DeepMind has released Gemma 4 12B, a unified encoder free multimodal model built for running agentic AI locally on laptops. 🔥 - 12B parameter model that runs on laptops with 16GB memory - Encoder free architecture for native image and audio processing - Performance close

    • 16Replies
    • 13Reposts
    • 171Likes
    • 10.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  9. 19

    @NVIDIAAIDev ·

    Fine-tuning multi-modal AI just got a whole lot easier. With the latest release of NVIDIA TAO, developers can accelerate post-training for reasoning VLM and embedding models using fine-tuning microservices (FTMS) with built-in recipes. New features: ⚡ NVIDIA Cosmos Reason VLM

    Video thumbnail from NVIDIA AI Developer's postWatch video
    • 5Replies
    • 27Reposts
    • 136Likes
    • 10.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  10. 20

    @askalphaxiv ·

    "Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale" Intern-S1-Pro scaled a multimodal model to 1T parameters with a lot of aligned scientific data, obtaining a really strong model that's capable of analyzing scientific figures, reason across STEM topics,

    • 3Replies
    • 22Reposts
    • 99Likes
    • 4.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  11. 21

    @TheHumanoidHub ·

    Jensen just launched NVIDIA Cosmos 3. Pitched as the first fully open omnimodel for physical AI: a mixture-of-transformers (reasoning + generation) with native vision reasoning and generation across text, image, video, sound, and action. Tops open-model leaderboards on physics,

    Video thumbnail from The Humanoid Hub's postWatch video
    • 5Replies
    • 24Reposts
    • 115Likes
    • 9.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  12. 22

    @IamEmily2050 ·

    A few days ago, DeepSeek published a paper titled "Thinking with Visual Primitives," but it was later removed. Luckily, I managed to download it, like so many other people, and, of course, I have to make a video overview of it with NotebookLM. Thinking with Visual Primitives

    Video thumbnail from Emily's postWatch video
    • 4Replies
    • 8Reposts
    • 73Likes
    • 3.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  13. 23

    @NainsiDwiv50980 ·

    🚨WEB SCRAPING IS DEAD🚨 They've created PixelRAG. An open source system that skips parsing HTML. Instead of converting a website into text... it takes a screenshot. And then a vision-language model reads the response directly from the pixels. Brutal. Because traditional

    Video thumbnail from Nainsi Dwivedi's postWatch video
    • 7Replies
    • 13Reposts
    • 26Likes
    • 2.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  14. 24

    @DataChaz ·

    Getting high-quality robot training data is slow and painfully costly. Human labelers cost ~$50/hour, making it a nightmare for robotics labs to scale. Today, @perceptroninc just completely changed the math. They just launched Egocentric, the first in a series of new 'Embodied

    Video thumbnail from Charly Wargnier ☄️'s postWatch video
    • 7Replies
    • 29Reposts
    • 79Likes
    • 19.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  15. 25

    @intern_lm ·

    🔥Introducing ARM-Thinker, the first Agentic multimodal Reward Model that autonomously invokes external tools to ground judgments in verifiable evidence. Accepted to CVPR 2026! 🥳Integrates three families of multimodal tools: 1⃣Image Crop & Zoom-in for fine-grained visual

    • 0Replies
    • 13Reposts
    • 43Likes
    • 4.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  16. 26

    @DivyanshT91162 ·

    Google just dropped a 12B model that has no business being this fast. I ran Gemma 4 12B locally on an RTX 4060 and got 21 tok/s. No API. No cloud. No subscription. Just 6.6GB, 256K context, and benchmarks that look straight-up unfair. → 77.5% AIME → 78.8% GPQA Diamond → 72%

    Video thumbnail from divyansh tiwari's postWatch video
    • 4Replies
    • 12Reposts
    • 66Likes
    • 8.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  17. 27

    @_jaydeepkarale ·

    Chunkr is an open-source document intelligence service that converts PDFs, PPTs, Word docs, and images into structured chunks ready for RAG and LLM pipelines. - Layout analysis with OCR and bounding boxes - Structured HTML and Markdown output - Vision-language model processing -

    • 5Replies
    • 3Reposts
    • 31Likes
    • 1.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  18. 28

    @burkov ·

    This AAAI 2026 paper introduces UI-R1, a framework demonstrating how rule-based reinforcement learning with a novel action reward significantly enhances multimodal LLM' reasoning capabilities for accurate GUI action prediction, outperforming larger supervised models on

    • 8Replies
    • 8Reposts
    • 47Likes
    • 2.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  19. 29

    @alex_verem ·

    Google just released a 4B parameter AI model that can read CT scans, MRI volumes, whole-slide pathology images, and electronic health records, all in a single architecture. And they made it free and open source. This is MedGemma 1.5. A single model that handles more medical

    • 6Replies
    • 5Reposts
    • 38Likes
    • 4.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  20. 30

    @pauliusztin_ ·

    Multimodal AI agents don't need special reasoning loops. All they need is better tools. A multimodal agent still follows the same ReAct cycle: Observe Reason Call tools Receive observations Repeat The only difference is what those observations contain. The agent simply adds

    • 2Replies
    • 3Reposts
    • 21Likes
    • 523Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  21. 31

    @tonysimons_ ·

    A 27B multimodal AI model is now running locally on an iPhone. PrismML’s Bonsai 27B packs 27.8B parameters into just 3.9GB and reportedly hits 11 tokens per second on an iPhone 17 Pro. That’s roughly 14× smaller than FP16 while retaining more than 90% of benchmark performance.

    • 5Replies
    • 5Reposts
    • 27Likes
    • 1.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  22. 32

    @oliviscusAI ·

    You can now serve text, image, video, and audio models from a single framework. vLLM just released vLLM-Omni, a massive upgrade to their original text-based serving engine. It eliminates the need to stitch together multiple frameworks for multimodal AI. → Serve any-to-any

    • 4Replies
    • 2Reposts
    • 27Likes
    • 2.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  23. 33

    @daniel_kraft ·

    SleepFM is a multimodal #AI foundation model for disease prediction. Developed by @Stanford researchers, trained on over 585K hours of polysomnography (PSG) data from 65K patients. It analyzes brain, heart, & respiratory signals to predict 130+ diseases—including dementia,

    • 2Replies
    • 12Reposts
    • 28Likes
    • 2.9KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  24. 34

    @RituWithAI ·

    🚨BREAKING: Meta just proved that every major AI model fails at something any toddler can do. Navigate a room. Find an object. Pick it up. Nine frontier models tested. GPT-5. Gemini 3 Pro. Claude Opus 4.6. Every major VLM available today. The best success rate: 16.8%. Not

    • 2Replies
    • 5Reposts
    • 8Likes
    • 239Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  25. 35

    @hasantoxr ·

    ByteDance 🔥: China's giant AI player made a fully multimodal AI agent stack that controls your computer, browser, and terminal using natural language instructions. It's called UI-TARS Desktop + Agent TARS. It sees your screen, clicks buttons, fills forms, and completes

    Video thumbnail from Hasan Toor's postWatch video
    • 10Replies
    • 4Reposts
    • 27Likes
    • 8.7KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  26. 36

    @_vmlops ·

    GEMMA 4 12B JUST CHANGED LOCAL AI DEVELOPMENT google dropped an encoder-free multimodal model no separate vision encoder. no audio encoder. just one decoder-only transformer handling everything ▫️ raw pixel patches projected directly to LLM hidden dim ▫️ raw 16kHz audio sliced

    Video thumbnail from Vaishnavi's postWatch video
    • 4Replies
    • 5Reposts
    • 16Likes
    • 1.6KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  27. 37

    @Marktechpost ·

    Zhipu AI has introduced GLM-5V-Turbo, a new multimodal coding model built for workflows where screenshots, videos, document layouts, and GUI states need to be converted into executable actions or code. What stands out is the system design: Native Multimodal Fusion, CogViT, MTP

    • 1Replies
    • 8Reposts
    • 70Likes
    • 14.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  28. 38

    @sabir_huss50540 ·

    For a year the rule held: a lab's best model stays locked behind an API. Alibaba just broke its own rule. Qwen3.8-Max is out, and this time the weights are actually downloadable. Everyone will lead with the headline: 2.4 trillion parameters. Second-largest open model ever,

    • 6Replies
    • 11Reposts
    • 18Likes
    • 1.4KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  29. 39

    @rohanpaul_ai ·

    The Meta/Oxford study finds, a multimodal model may need surprisingly little image-generation data if language and visual understanding are trained with it from the start. So, you probably don’t need to spend that much training compute teaching a multimodal model to generate

    • 6Replies
    • 4Reposts
    • 23Likes
    • 3.1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  30. 40

    @HuggingPapers ·

    OpenVLThinkerV2 A generalist multimodal reasoning model that introduces Gaussian GRPO—forcing advantage distributions to standard normal N(0,1) for stable multi-task RL training. Achieves 71.6% on MMMU and outperforms GPT-4o across 18 diverse visual benchmarks.

    • 2Replies
    • 3Reposts
    • 27Likes
    • 1.8KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  31. 41

    @sharbel ·

    🚨SHOCKING: Researchers proved that AI models judge your intelligence based almost entirely on what you look like. And a handful of visual cues are doing almost all of the damage. Researchers built 500 photorealistic faces. Then they changed one thing at a time. Then they showed

    • 11Replies
    • 5Reposts
    • 21Likes
    • 5.5KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  32. 42

    @RoundtableSpace ·

    GOOGLE JUST DROPPED 6 AI TOOLS AND MOST PEOPLE MISSED THEM INCLUDING ONE THAT RUNS MULTIMODAL AI FULLY OFFLINE ON A NORMAL LAPTOP Gemma 4 12B handles text, images, audio, and video natively with nothing leaving your machine, and quantized versions shrink to 1GB and run on older

    Video thumbnail from 0xMarioNawfal's postWatch video
    • 5Replies
    • 3Reposts
    • 70Likes
    • 43.3KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  33. 43

    @Hesamation ·

    Thinking Machines Lab engineer explains in 60 minutes the evolution from LLMs into multimodal models, their architecture and challenges, and why next-frame prediction makes better pixels but not smarter models.

    • 0Replies
    • 1Reposts
    • 5Likes
    • 336Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  34. 44

    @burkov ·

    Most systems that handle different kinds of data—text, images, video, audio—do so by training a separate encoder for each type and then forcing their outputs into a shared space, which works for matching an image to a caption but loses the connections that arise when those types

    • 1Replies
    • 1Reposts
    • 14Likes
    • 1.2KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  35. 45

    @OpenBMB ·

    Visual comprehension requires high-level abstract semantics, while image generation demands fine-grained pixel details. How can we resolve this fundamental conflict within a single unified model? 🤔 Today, we present CHEERS—new research from @TsinghuaNLP (OpenBMB member), XJTU,

    • 0Replies
    • 3Reposts
    • 7Likes
    • 691Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  36. 46

    @shawnchauhan1 ·

    Google just made multimodal search infrastructure a commodity. Text, image, video, audio, PDF - one vector space, one API call. Startups have been raising on the premise that unified multimodal retrieval is hard to build. It was. Until yesterday. The question is not whether

    • 0Replies
    • 3Reposts
    • 6Likes
    • 422Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  37. 47

    @HuggingModels ·

    Ever seen an AI that understands images AND text together? Meet CLIP ViT-B/32. It's a vision-language model that connects what you see with what you describe. No fine-tuning needed. This is zero-shot image classification magic.

    • 1Replies
    • 0Reposts
    • 2Likes
    • 429Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  38. 48

    @did0f ·

    What if the most important AI benchmark is not intelligence, but locality? Not, “Is this model smarter than the biggest cloud model?” But, “Can it run well on my machine?” That is why I find Gemma 4 12B interesting. Not because it magically solves multimodality. It does not.

    • 3Replies
    • 0Reposts
    • 3Likes
    • 571Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  39. 49

    @tonjkb ·

    On Rogan, Marc Andreessen described a secret AI setup Mark Zuckerberg uses to train Brazilian Jiu-Jitsu. Webcams in his home gym feed live sparring video to a multimodal AI that gives him direct performance feedback. I see this as much more than a billionaire's custom toy. We

    Video thumbnail from Tony Jacob | FindaClip.com's postWatch video
    • 0Replies
    • 1Reposts
    • 2Likes
    • 567Views
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.
  40. 50

    @LEAPTRADER_ ·

    $NVDA has launched Nemotron 3 Nano Omni, an open-source multimodal model that natively combines video, audio, image, and text reasoning in a single efficient system. ✅ 30B MoE model (3B active parameters) with a hybrid Transformer-Mamba architecture. ✅ Handles long-context

    • 5Replies
    • 1Reposts
    • 3Likes
    • 1KViews
    View on X
    Rewrite this post in your own voice and angle.See the hook, structure, and reusable template behind this post.

Write posts like these, in your own voice

Xholic studies what works in your niche, drafts posts in your voice and schedules them for the hours your audience is online.

$0 today · Cancel anytime

Browse all tweet collections

More on this topic

Tweet Remixer

Remix this post

Creator

@creator

View on 𝕏

Choose a tone

Best Tweets by email

Get the best tweets about Multimodal AI every two weeks

The list refreshes about every two weeks. Confirm your email to start, and unsubscribe anytime.

We only use your email for this digest. See our Privacy Policy.

Write posts like these in your voice

Start 3-day trial for $0

$0 today · Cancel anytime