50 Best Tweets About Multimodal AI (2026)

Browse the best tweets about multimodal AI, including vision, audio, video, model architectures, benchmarks, applications, and developer experiments.

Multimodal models combining text, images, audio, or video, with concrete research, evaluations, workflows, and applications.

Creators
47
Updated

What 50 top Multimodal AI posts reveal

Discussion of multimodal AI emphasizes compact document OCR, unified multimodal models, and systems that combine visual perception with agents or workflows. The largest score outliers span CAD generation, OCR, audio-video generation, embeddings, and pathology. Posts also raise benchmark-grounding concerns, while other posts report reasoning and agent-performance advances.

Dominant tone
Positive

86% of posts

Median score
25.8

All-time engagement

Leading format
Announcement

84% of posts

Recent posts
32%

Published in 90 days

Conversation map

The themes creators return to

Local Multimodal Deployment and Infrastructure

Efficient, open, and local multimodal models and infrastructure, including edge deployment, on-device inference, fine-tuning, serving, and low-latency hardware.

34%

Unified Omnimodal Models

Unified multimodal foundation models that natively process and generate across text, images, audio, video, 3D, and actions, including architectural designs such as discrete tokens, decoder-only models, and mixture-of-experts.

30%

Agents, GUI Control, and Embodied AI

Multimodal agents that perceive screens or environments and act through GUIs, browsers, terminals, tools, robots, and software workflows.

26%

Multimodal Reasoning and Evaluation

Multimodal reasoning research and evaluation, covering spatial grounding, visual primitives, scientific reasoning, benchmark leakage, fake visual reasoning, and model bias.

14%

Document OCR and Understanding

Vision-language models for document intelligence, including OCR, layout parsing, tables, formulas, handwriting, PDFs, and efficient local deployment.

10%

Tone and stance

Sentiment Positive leads
Author posture Supportive leads

Performance benchmark

Median likes
100
Median reposts
16
Median replies
6
Median views
9.7K

Posts with media make up 94% of this collection. Their median all-time score is 24.7, compared with 88.0 for text-only posts.

Format mix

  • Announcement 84% · score 34.4
  • Opinion 16% · score 11.2

Where creators agree, and where they do not

Shared view

Compact OCR is a practical focal point

Document AI is a prominent practical thread. Posts describe compact, open vision-language models for layouts, tables, formulas, and handwriting, with local deployment or high-throughput serving options.

Shared view

Unified models target broader modality coverage

Posts frame unified architectures as a way to combine text, vision, and audio capabilities, including understanding, generation, editing, and agent-oriented local use. The cited models make differing architecture and performance claims.

Shared view

Perception is being paired with agents

Posts describe multimodal systems that pair perception with action or verification: visual analysis for robotics, a VLM director and verifier for 3D-world construction, and an agent positioned to work across screens and software workflows.

Open debate

Benchmark gains and visual grounding are both contested topics

One post summarizes MIRAGE as finding substantial image-removed performance on six vision benchmarks and argues that some benchmark items may not require vision. Other posts report multimodal-reasoning gains from visual primitives and RL-based training. These are reported claims rather than independently verified comparisons.

Open debate

Computer-use agents have contrasting readiness claims

One practitioner reports a screenshot-driven computer-use loop taking about 70 seconds for a tweet-posting task. Other posts promote high-throughput multimodal inference and screen-operating agents; these posts do not establish a common measure of real-world agent speed or reliability.

Patterns behind standout posts

Document OCR leads theme-level performance

Document OCR has the highest theme-level median all-time score in the analytics, at 287.031. The cited posts focus on small vision-language OCR models and local or server deployment.

Generation announcements drew relatively strong attention

Announcements account for 42 of 50 posts (84%). Multimodal generation has a 39.2 median all-time score, above the infrastructure (30.874), agent (28.734), and unified-model (22.619) themes, though below document OCR and scientific/medical multimodality.

Statistical standouts

  1. View standout post 1 Score 5483.0 · 212.44× median
  2. View standout post 2 Score 1240.2 · 48.05× median
  3. View standout post 3 Score 853.0 · 33.05× median
  4. View standout post 4 Score 775.1 · 30.03× median
  5. View standout post 5 Score 717.4 · 27.79× median

Who shapes this conversation

The five most represented creators account for 16% of the selected posts.

  1. 1. BURKOV

    @burkov

    2 posts

  2. 2. Hugging Models

    @HuggingModels

    2 posts

  3. 3. DailyPapers

    @HuggingPapers

    2 posts

  4. 4. Vaishnavi

    @_vmlops

    1 post

  5. 5. AshutoshShrivastava

    @ai_for_success

    1 post

  6. 6. Alex Veremeyenko

    @alex_verem

    1 post

BURKOV covers agent and retrieval research

BURKOV’s two posts cover rule-based reinforcement learning for GUI action prediction and contrastive adaptation of a multimodal model for retrieval.

Hugging Models spotlights VLM explainers

Hugging Models posts about video recap and CLIP-style image-text understanding.

DailyPapers tracks research releases

DailyPapers posts cover an open-weight VLM release and a multimodal reasoning model with reported benchmark results.

How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top Multimodal AI tweets from 47 creators

Ranked 01–50

  1. 01

    @heygurisingh ·

    🚨BREAKING: MIT just dropped an AI model that converts photos into fully editable CAD programs and it quietly kills the $150/hour CAD modeling industry. It's called GenCAD. You give it an image. It gives you the complete parametric command sequence lines, arcs, extrusions ready

    • 212 Replies
    • 1.2K Reposts
    • 8.3K Likes
    • 695.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @AlphaSignalAI ·

    A peanut-sized Chinese model just dethroned Gemini at reading documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages.

    Video thumbnail from AlphaSignal AI's post Watch video
    • 22 Replies
    • 163 Reposts
    • 1.3K Likes
    • 89.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @ihteshamali ·

    🚨 BREAKING: A research lab just released a 15B model that generates multilingual talking human videos with synced audio, beats every competitor in human evaluation, and runs in 38 seconds on one GPU. It's called daVinci-MagiHuman. The key insight is that every other model in

    Video thumbnail from Ihtesham Ali's post Watch video
    • 19 Replies
    • 116 Reposts
    • 696 Likes
    • 43.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @OfficialLoganK ·

    Say hello to Gemini Embedding 2, our new SOTA multimodal model that lets your bring text, images, video, audio, and docs into the same embedding space! 👀

    • 273 Replies
    • 453 Reposts
    • 5.6K Likes
    • 851.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @satyanadella ·

    We’ve trained a multimodal AI model to turn routine pathology slides into spatial proteomics, with the potential to reduce time and cost while expanding access to cancer care.

    Video thumbnail from Satya Nadella's post Watch video
    • 455 Replies
    • 1.8K Reposts
    • 11.1K Likes
    • 2.8M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @techNmak ·

    Finally, a lightweight VLM that beats the giants at OCR. (1.7B parameters, SOTA on OmniDocBench) dots. ocr is a new multilingual document parser that proves you don't need massive models for perfect document understanding. Current SOTA models are often massive (72B+) or

    • 27 Replies
    • 88 Reposts
    • 708 Likes
    • 31K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @dr_cintas ·

    This peanut-sized chinese model just dethroned Gemini at reading documents. It’s called glm-ocr. it’s a tiny 0.9b parameter vision-language model that is about to replace every expensive ocr api you use. → Handles text, tables, formulas, handwriting → Scored 94.62 on

    Video thumbnail from Alvaro Cintas's post Watch video
    • 17 Replies
    • 68 Reposts
    • 417 Likes
    • 35.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @sukh_saroy ·

    Holy shit... Stanford just proved that GPT-5, Gemini, and Claude can't actually see. They removed every image from 6 major vision benchmarks. The models still scored 70-80% accuracy. They were never looking at your photos. Your scans. Your X-rays. Here's what's really going

    • 40 Replies
    • 120 Reposts
    • 419 Likes
    • 55.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @TheTuringPost ·

    There’s a serious gap in multimodal models – they work with images, but still reason in language, which isn’t that precise for visual stuff. @deepseek_ai just dropped an idea to solve this: let the model literally point to exact locations in the image while it thinks. They call

    • 11 Replies
    • 77 Reposts
    • 498 Likes
    • 30.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @skalskip92 ·

    spent most of my day playing with GLM-OCR it's a 0.9B param vision-language model. supports 8K resolution, 8+ languages, and has built-in text, LaTeX, and table recognition modes. awesome! I tested it across different OCR tasks. starting with shipping container serial numbers.

    • 17 Replies
    • 55 Reposts
    • 817 Likes
    • 271.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @arcinstitute ·

    Over 250 million protein sequences are known, but fewer than 0.1% have confirmed functions. Today, @genophoria, @BoWang87 & team introduce BioReason-Pro, a multimodal reasoning model that predicts protein function and explains its reasoning like an expert would.

    • 13 Replies
    • 124 Reposts
    • 527 Likes
    • 61.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @Prince_Canuma ·

    RF-DETR by @roboflow now on MLX It can do realtime instance segmentation on-device and enable some cool use cases for visual analysis, monitoring and robotics like Reachy Mini. Also augmented VLM and VLA by preprocessing image and video with areas of interest. New release

    Video thumbnail from Prince Canuma's post Watch video
    • 13 Replies
    • 36 Reposts
    • 356 Likes
    • 22.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @Yuchenj_UW ·

    I used Claude Computer Use/Dispatch yesterday. My feeling: It’s too damn slow! Posting a tweet takes me ~5 seconds (once I have the content). Claude took 70 seconds. Why? It controls the screen via a loop: take a screenshot → send to a huge remote multimodal model (opus 4.6) →

    • 138 Replies
    • 32 Reposts
    • 644 Likes
    • 53.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @Meituan_LongCat ·

    🔥 Introducing LongCat-Next: A Discrete Native Autoregressive Multimodal Model LongCat-Next integrates language, vision, and audio into a unified discrete autoregressive model, extending Next-Token Prediction to native multimodality and delivering industrial-strength performance

    • 10 Replies
    • 66 Reposts
    • 467 Likes
    • 44.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @MattNiessner ·

    📢WorldAgents: 3D worlds only from 2D image models - without any training! We propose an agentic approach with a Director (VLM) to plan the scene, a Generator (Flux or NanoBanana) for new views, and a Verifier (VLM) for selection / 3D consistency. -> High-fidelity 3D worlds from

    Video thumbnail from Matthias Niessner's post Watch video
    • 6 Replies
    • 46 Reposts
    • 269 Likes
    • 18.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @fahdmirza ·

    💥 Gemma 4 E2B + Hermes Agent + vLLM running fully local — zero cloud, zero cost ♠ and it's a complete multimodal AI stack on a single GPU 🚀 🔹Gemma 4 E2B served locally via vLLM 0.19.0 on NVIDIA A6000 🔹Text, vision and native audio transcription — all offline 🔹27 languages

    • 4 Replies
    • 16 Reposts
    • 139 Likes
    • 10.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @intern_lm ·

    🚀Meet InternVL-U: a lightweight 4B unified multimodal model that brings reasoning, generation, and editing into a unified framework. 🔥Built upon unified contextual modeling, modality-specific modular design, and decoupled visual representations, InternVL-U achieves a strong

    • 1 Replies
    • 31 Reposts
    • 151 Likes
    • 20.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @chutes_ai ·

    Moonshot trained a model on 15 trillion tokens of mixed vision and text data. The result scores 96.1 on AIME 2025 and 76.8 on SWE-Bench Verified. Model Spotlight: Kimi K2.5 by @kimi_moonshot 1T total parameters. 32B activated per token (MoE). 256K context. Vision baked into

    • 10 Replies
    • 34 Reposts
    • 255 Likes
    • 15.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @arankomatsuzaki ·

    Context Unrolling in Omni Models - A unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations - Enables Context Unrolling, where the model explicitly reasons across multiple modal representations

    • 3 Replies
    • 25 Reposts
    • 145 Likes
    • 15.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @NVIDIAAIDev ·

    Fine-tuning multi-modal AI just got a whole lot easier. With the latest release of NVIDIA TAO, developers can accelerate post-training for reasoning VLM and embedding models using fine-tuning microservices (FTMS) with built-in recipes. New features: ⚡ NVIDIA Cosmos Reason VLM

    Video thumbnail from NVIDIA AI Developer's post Watch video
    • 5 Replies
    • 27 Reposts
    • 140 Likes
    • 10.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @askalphaxiv ·

    "Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale" Intern-S1-Pro scaled a multimodal model to 1T parameters with a lot of aligned scientific data, obtaining a really strong model that's capable of analyzing scientific figures, reason across STEM topics,

    • 3 Replies
    • 22 Reposts
    • 99 Likes
    • 4.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @ai_for_success ·

    Google DeepMind has released Gemma 4 12B, a unified encoder free multimodal model built for running agentic AI locally on laptops. 🔥 - 12B parameter model that runs on laptops with 16GB memory - Encoder free architecture for native image and audio processing - Performance close

    • 16 Replies
    • 13 Reposts
    • 171 Likes
    • 10.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @andrewdfeldman ·

    Yesterday, we launched @GoogleDeepMind's Gemma 4 model on @cerebras. The first multimodal model on Cerebras. 1,500 tokens per second. 15x faster than the nearest comparable model. Multimodal agents can see, reason, act, and retry. At 1,500 tokens per second, that loop is

    Video thumbnail from Andrew Feldman's post Watch video
    • 12 Replies
    • 23 Reposts
    • 218 Likes
    • 25.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @TheHumanoidHub ·

    Jensen just launched NVIDIA Cosmos 3. Pitched as the first fully open omnimodel for physical AI: a mixture-of-transformers (reasoning + generation) with native vision reasoning and generation across text, image, video, sound, and action. Tops open-model leaderboards on physics,

    Video thumbnail from The Humanoid Hub's post Watch video
    • 5 Replies
    • 24 Reposts
    • 115 Likes
    • 9.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @IamEmily2050 ·

    A few days ago, DeepSeek published a paper titled "Thinking with Visual Primitives," but it was later removed. Luckily, I managed to download it, like so many other people, and, of course, I have to make a video overview of it with NotebookLM. Thinking with Visual Primitives

    Video thumbnail from Emily's post Watch video
    • 4 Replies
    • 8 Reposts
    • 73 Likes
    • 3.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @thetripathi58 ·

    Everyone is chasing bigger models. More parameters. More compute. More cloud dependency. OpenBMB went the other way. MiniCPM-V 4.6 - a ~1B vision-language model that outperforms models three times its size. I ran it on my iPhone. Here's what happened:

    • 24 Replies
    • 57 Reposts
    • 142 Likes
    • 74.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @riyazmd774 ·

    🚨 BREAKING: Alibaba unleashes Qwen3.5-Omni, a new frontier in Full-Modality AI. 🤯 Matching the latest Gemini-3.1 Pro in A/V understanding & surpassing it in Audio tasks, this model introduces Audio-Visual Vibe Coding turning whiteboard sketch videos or game clips directly into

    • 31 Replies
    • 49 Reposts
    • 101 Likes
    • 24.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @DivyanshT91162 ·

    Google just dropped a 12B model that has no business being this fast. I ran Gemma 4 12B locally on an RTX 4060 and got 21 tok/s. No API. No cloud. No subscription. Just 6.6GB, 256K context, and benchmarks that look straight-up unfair. → 77.5% AIME → 78.8% GPQA Diamond → 72%

    Video thumbnail from divyansh tiwari's post Watch video
    • 4 Replies
    • 12 Reposts
    • 66 Likes
    • 8.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @SandAI_HQ ·

    🪄 Introducing daVinci-MagiHuman: The Performance-Level Audio-Video Generative Foundation Model Proudly open-sourced and jointly developed by SII GAIR Lab & https://t.co/yn4NJpoMrD, it sets a new standard for multimodal AI. ⏳ 1/6

    Video thumbnail from Sand.ai's post Watch video
    • 3 Replies
    • 11 Reposts
    • 34 Likes
    • 2.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @alex_verem ·

    Google just released a 4B parameter AI model that can read CT scans, MRI volumes, whole-slide pathology images, and electronic health records, all in a single architecture. And they made it free and open source. This is MedGemma 1.5. A single model that handles more medical

    • 6 Replies
    • 5 Reposts
    • 38 Likes
    • 4.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @burkov ·

    This AAAI 2026 paper introduces UI-R1, a framework demonstrating how rule-based reinforcement learning with a novel action reward significantly enhances multimodal LLM' reasoning capabilities for accurate GUI action prediction, outperforming larger supervised models on

    • 8 Replies
    • 8 Reposts
    • 47 Likes
    • 2.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @oliviscusAI ·

    You can now serve text, image, video, and audio models from a single framework. vLLM just released vLLM-Omni, a massive upgrade to their original text-based serving engine. It eliminates the need to stitch together multiple frameworks for multimodal AI. → Serve any-to-any

    • 4 Replies
    • 2 Reposts
    • 27 Likes
    • 2.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @pauliusztin_ ·

    Multimodal AI agents don't need special reasoning loops. All they need is better tools. A multimodal agent still follows the same ReAct cycle: Observe Reason Call tools Receive observations Repeat The only difference is what those observations contain. The agent simply adds

    • 2 Replies
    • 3 Reposts
    • 21 Likes
    • 523 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @HuggingPapers ·

    LG AI Research releases EXAONE 4.5, their first open-weight vision language model 33B parameters with native multimodal capabilities, 256K context length, and SOTA performance in document understanding and Korean contextual reasoning.

    • 1 Replies
    • 3 Reposts
    • 43 Likes
    • 2.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @alexabelonix ·

    Alibaba just unveiled Qwen3.7-Plus, a multimodal AI agent that can see, think, code, and take action across screens. This is the agent direction that feels really important. One model that understands text, images, visual interfaces, GUIs, command lines, code, and

    • 14 Replies
    • 1 Reposts
    • 37 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @IlirAliu_ ·

    A humanoid robot autonomously executes a full long-horizon task, from ONE natural language command... for the first time. It goes downstairs to get a snack package, rides the elevator upstairs, opens the box, and puts the snacks into a drawer. The platform integrates

    Video thumbnail from Ilir Aliu's post Watch video
    • 5 Replies
    • 15 Reposts
    • 45 Likes
    • 6.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    For the first time in human history, we are teaching a Foundation Model to master the diverse tasks of medicinal chemists, biologists, and computational scientists all in one place. In our latest collaboration with Liquid AI, we are moving away from fragmented, specialized tools

    Video thumbnail from Alex Zhavoronkov, PhD (aka Aleksandrs Zavoronkovs)'s post Watch video
    • 4 Replies
    • 15 Reposts
    • 54 Likes
    • 10.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @_vmlops ·

    GEMMA 4 12B JUST CHANGED LOCAL AI DEVELOPMENT google dropped an encoder-free multimodal model no separate vision encoder. no audio encoder. just one decoder-only transformer handling everything ▫️ raw pixel patches projected directly to LLM hidden dim ▫️ raw 16kHz audio sliced

    Video thumbnail from Vaishnavi's post Watch video
    • 4 Replies
    • 5 Reposts
    • 16 Likes
    • 1.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @hasantoxr ·

    ByteDance 🔥: China's giant AI player made a fully multimodal AI agent stack that controls your computer, browser, and terminal using natural language instructions. It's called UI-TARS Desktop + Agent TARS. It sees your screen, clicks buttons, fills forms, and completes

    Video thumbnail from Hasan Toor's post Watch video
    • 10 Replies
    • 4 Reposts
    • 27 Likes
    • 8.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @HuggingPapers ·

    OpenVLThinkerV2 A generalist multimodal reasoning model that introduces Gaussian GRPO—forcing advantage distributions to standard normal N(0,1) for stable multi-task RL training. Achieves 71.6% on MMMU and outperforms GPT-4o across 18 diverse visual benchmarks.

    • 2 Replies
    • 3 Reposts
    • 27 Likes
    • 1.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @sabir_huss50540 ·

    For a year the rule held: a lab's best model stays locked behind an API. Alibaba just broke its own rule. Qwen3.8-Max is out, and this time the weights are actually downloadable. Everyone will lead with the headline: 2.4 trillion parameters. Second-largest open model ever,

    • 6 Replies
    • 11 Reposts
    • 18 Likes
    • 1.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @sharbel ·

    🚨SHOCKING: Researchers proved that AI models judge your intelligence based almost entirely on what you look like. And a handful of visual cues are doing almost all of the damage. Researchers built 500 photorealistic faces. Then they changed one thing at a time. Then they showed

    • 11 Replies
    • 5 Reposts
    • 21 Likes
    • 5.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @Hesamation ·

    Thinking Machines Lab engineer explains in 60 minutes the evolution from LLMs into multimodal models, their architecture and challenges, and why next-frame prediction makes better pixels but not smarter models.

    • 0 Replies
    • 1 Reposts
    • 5 Likes
    • 336 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @HuggingModels ·

    Meet Tarsier2-Recap-7b: a video understanding model that's changing how AI 'watches' videos. It doesn't just see frames, it understands narratives, actions, and context. This is the next step in multimodal AI.

    • 1 Replies
    • 1 Reposts
    • 8 Likes
    • 663 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @OpenBMB ·

    Visual comprehension requires high-level abstract semantics, while image generation demands fine-grained pixel details. How can we resolve this fundamental conflict within a single unified model? 🤔 Today, we present CHEERS—new research from @TsinghuaNLP (OpenBMB member), XJTU,

    • 0 Replies
    • 3 Reposts
    • 7 Likes
    • 691 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @burkov ·

    Most systems that handle different kinds of data—text, images, video, audio—do so by training a separate encoder for each type and then forcing their outputs into a shared space, which works for matching an image to a caption but loses the connections that arise when those types

    • 1 Replies
    • 1 Reposts
    • 14 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @shawnchauhan1 ·

    Google just made multimodal search infrastructure a commodity. Text, image, video, audio, PDF - one vector space, one API call. Startups have been raising on the premise that unified multimodal retrieval is hard to build. It was. Until yesterday. The question is not whether

    • 0 Replies
    • 3 Reposts
    • 6 Likes
    • 422 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @HuggingModels ·

    Ever seen an AI that understands images AND text together? Meet CLIP ViT-B/32. It's a vision-language model that connects what you see with what you describe. No fine-tuning needed. This is zero-shot image classification magic.

    • 1 Replies
    • 0 Reposts
    • 2 Likes
    • 429 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @did0f ·

    What if the most important AI benchmark is not intelligence, but locality? Not, “Is this model smarter than the biggest cloud model?” But, “Can it run well on my machine?” That is why I find Gemma 4 12B interesting. Not because it magically solves multimodality. It does not.

    • 3 Replies
    • 0 Reposts
    • 3 Likes
    • 571 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @LEAPTRADER_ ·

    $NVDA has launched Nemotron 3 Nano Omni, an open-source multimodal model that natively combines video, audio, image, and text reasoning in a single efficient system. ✅ 30B MoE model (3B active parameters) with a hybrid Transformer-Mamba architecture. ✅ Handles long-context

    • 5 Replies
    • 1 Reposts
    • 3 Likes
    • 1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone