50 Best Tweets About Multimodal AI (2026)

Browse the best tweets about multimodal AI, including vision, audio, video, model architectures, benchmarks, applications, and developer experiments.

Multimodal models combining text, images, audio, or video, with concrete research, evaluations, workflows, and applications.

Creators
45
Updated

What 50 top Multimodal AI posts reveal

Multimodal AI posts concentrate on unified architectures and deployment-oriented applications, including document parsing, coding and GUI agents, scientific models, and robotics. The release-heavy conversation also includes research-led concerns that benchmark results may not always demonstrate reliable use of visual evidence, including for medical tasks.

Dominant tone
Positive

82% of posts

Median score
29.8

All-time engagement

Leading format
Announcement

86% of posts

Recent posts
30%

Published in 90 days

Conversation map

The themes creators return to

Unified multimodal model architectures

Native or encoder-free models that jointly handle text, vision, audio, video, 3D, generation, and editing through shared token spaces or unified transformers.

42%

Multimodal agents and GUI automation

Agents that inspect screens, use computer interfaces, code from visual inputs, call tools, and operate software through visual context.

24%

Visual reasoning and grounding

Research on spatially grounded reasoning, visual primitives, explicit versus latent visual thought, and failures of models to genuinely use image evidence.

14%

Tone and stance

Sentiment Positive leads
Author posture Supportive leads

Performance benchmark

Median likes
112
Median reposts
16
Median replies
6
Median views
8.9K

Posts with media make up 96% of this collection. Their median all-time score is 29.8, compared with 63.1 for text-only posts.

Format mix

  • Announcement 86% · score 38.0
  • Opinion 14% · score 6.38

Where creators agree, and where they do not

Shared view

Unified models span more modalities

Posts emphasize shared architectures for text, vision, audio, video, and generation, presenting native multimodality as an alternative to multi-component pipelines.

Shared view

Document AI is a concrete near-term use

Compact VLM releases target OCR, tables, formulas, layouts, handwriting, and multilingual documents, with local and production-serving options highlighted.

Shared view

Agents use visual operating context

Multimodal systems are presented as coding and GUI agents that interpret designs, documents, and screens before using tools or software interfaces.

Shared view

Science and medicine are active targets

Reported applications include pathology-to-spatial-proteomics prediction, protein-function prediction, scientific-figure analysis, and multimodal medical imaging.

Open debate

Visual scores versus genuine image use

MIRAGE reports that models retained substantial benchmark accuracy after images were removed, while work on visual primitives proposes explicit points and boxes to improve spatial grounding.

Open debate

GUI agents: screen control or better tools

One post argues that multimodal agents retain the standard ReAct loop with richer observations; another argues that screenshot-driven GUI control may give way to specialized models and APIs or CLIs.

Open debate

Unified design still has trade-offs

CHEERS identifies a tension between semantic comprehension and pixel-detail generation in unified models, while the MedGemma post reports that medical specialization came with lower general-knowledge benchmark performance.

Patterns behind standout posts

Document OCR posts led practical performance

Document AI and OCR had the highest supplied theme median all-time score, at 229.7. Its evidence tweets describe compact OCR models alongside benchmark and deployment claims.

GenCAD was the standout outlier

Tweet 2032364172600885729 recorded an all-time score of 5483.04, the largest supplied outlier. The post concerns photo-to-editable-CAD generation and CAD code generation from images.

Statistical standouts

  1. View standout post 1 Score 5483.0 · 183.99× median
  2. View standout post 2 Score 853.0 · 28.62× median
  3. View standout post 3 Score 793.2 · 26.62× median
  4. View standout post 4 Score 717.4 · 24.07× median
  5. View standout post 5 Score 469.2 · 15.74× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. BURKOV

    @burkov

    2 posts

  2. 2. DailyPapers

    @HuggingPapers

    2 posts

  3. 3. Ihtesham Ali

    @ihteshamali

    2 posts

  4. 4. OpenBMB

    @OpenBMB

    2 posts

  5. 5. sabir hussain

    @sabir_huss50540

    2 posts

  6. 6. Vaishnavi

    @_vmlops

    1 post

Few repeat voices in the sample

The dataset contains 45 creators across 50 tweets, and the supplied top-five placement share is 20%.

Ihtesham Ali covered applied releases

Ihtesham Ali’s two posts cover unified talking-video generation and a VLM-driven self-resetting robotics system; the supplied creator median all-time score is 526.86.

OpenBMB posted on reasoning and unified-model design

OpenBMB’s two evidence tweets discuss a causal critique of latent visual reasoning and a unified-model approach that separates semantic and detail representations.

Since the previous snapshot

What changed since Aug 12, 2026

  • 80% of the selected posts remained.
  • The creator count changed by +1.
  • The leading sentiment remained stable.
How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top Multimodal AI tweets from 45 creators

Ranked 01–50

  1. 01

    @heygurisingh ·

    🚨BREAKING: MIT just dropped an AI model that converts photos into fully editable CAD programs and it quietly kills the $150/hour CAD modeling industry. It's called GenCAD. You give it an image. It gives you the complete parametric command sequence lines, arcs, extrusions ready

    • 212 Replies
    • 1.2K Reposts
    • 8.3K Likes
    • 695.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @ihteshamali ·

    🚨 BREAKING: A research lab just released a 15B model that generates multilingual talking human videos with synced audio, beats every competitor in human evaluation, and runs in 38 seconds on one GPU. It's called daVinci-MagiHuman. The key insight is that every other model in

    Video thumbnail from Ihtesham Ali's post Watch video
    • 19 Replies
    • 116 Reposts
    • 696 Likes
    • 43.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @Zai_org ·

    Introducing GLM-5V-Turbo: Vision Coding Model - Native Multimodal Coding: Natively understands multimodal inputs including images, videos, design drafts, and document layouts. - Balanced Visual and Programming Capabilities: Achieves leading performance across core benchmarks for

    Video thumbnail from Z.ai's post Watch video
    • 251 Replies
    • 642 Reposts
    • 5.7K Likes
    • 2M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @satyanadella ·

    We’ve trained a multimodal AI model to turn routine pathology slides into spatial proteomics, with the potential to reduce time and cost while expanding access to cancer care.

    Video thumbnail from Satya Nadella's post Watch video
    • 455 Replies
    • 1.8K Reposts
    • 11.1K Likes
    • 2.8M Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @techNmak ·

    Finally, a lightweight VLM that beats the giants at OCR. (1.7B parameters, SOTA on OmniDocBench) dots. ocr is a new multilingual document parser that proves you don't need massive models for perfect document understanding. Current SOTA models are often massive (72B+) or

    • 27 Replies
    • 88 Reposts
    • 708 Likes
    • 31K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @dr_cintas ·

    This peanut-sized chinese model just dethroned Gemini at reading documents. It’s called glm-ocr. it’s a tiny 0.9b parameter vision-language model that is about to replace every expensive ocr api you use. → Handles text, tables, formulas, handwriting → Scored 94.62 on

    Video thumbnail from Alvaro Cintas's post Watch video
    • 17 Replies
    • 68 Reposts
    • 417 Likes
    • 35.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @sukh_saroy ·

    Holy shit... Stanford just proved that GPT-5, Gemini, and Claude can't actually see. They removed every image from 6 major vision benchmarks. The models still scored 70-80% accuracy. They were never looking at your photos. Your scans. Your X-rays. Here's what's really going

    • 40 Replies
    • 120 Reposts
    • 419 Likes
    • 55.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @ihteshamali ·

    🚨BREAKING: Researchers just built a robot that trains itself. It's called RoboClaw. And it changes how we think about robot training forever. Here's what's actually happening inside it: Every time you want to teach a robot a new task, someone has to sit there. Collect

    • 17 Replies
    • 46 Reposts
    • 296 Likes
    • 20.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @TheTuringPost ·

    There’s a serious gap in multimodal models – they work with images, but still reason in language, which isn’t that precise for visual stuff. @deepseek_ai just dropped an idea to solve this: let the model literally point to exact locations in the image while it thinks. They call

    • 11 Replies
    • 77 Reposts
    • 498 Likes
    • 30.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @skalskip92 ·

    spent most of my day playing with GLM-OCR it's a 0.9B param vision-language model. supports 8K resolution, 8+ languages, and has built-in text, LaTeX, and table recognition modes. awesome! I tested it across different OCR tasks. starting with shipping container serial numbers.

    • 17 Replies
    • 55 Reposts
    • 817 Likes
    • 271.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @arcinstitute ·

    Over 250 million protein sequences are known, but fewer than 0.1% have confirmed functions. Today, @genophoria, @BoWang87 & team introduce BioReason-Pro, a multimodal reasoning model that predicts protein function and explains its reasoning like an expert would.

    • 13 Replies
    • 124 Reposts
    • 527 Likes
    • 61.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @Prince_Canuma ·

    RF-DETR by @roboflow now on MLX It can do realtime instance segmentation on-device and enable some cool use cases for visual analysis, monitoring and robotics like Reachy Mini. Also augmented VLM and VLA by preprocessing image and video with areas of interest. New release

    Video thumbnail from Prince Canuma's post Watch video
    • 13 Replies
    • 36 Reposts
    • 356 Likes
    • 22.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @Yuchenj_UW ·

    I used Claude Computer Use/Dispatch yesterday. My feeling: It’s too damn slow! Posting a tweet takes me ~5 seconds (once I have the content). Claude took 70 seconds. Why? It controls the screen via a loop: take a screenshot → send to a huge remote multimodal model (opus 4.6) →

    • 138 Replies
    • 32 Reposts
    • 644 Likes
    • 53.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @Meituan_LongCat ·

    🔥 Introducing LongCat-Next: A Discrete Native Autoregressive Multimodal Model LongCat-Next integrates language, vision, and audio into a unified discrete autoregressive model, extending Next-Token Prediction to native multimodality and delivering industrial-strength performance

    • 10 Replies
    • 66 Reposts
    • 467 Likes
    • 44.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @MattNiessner ·

    📢WorldAgents: 3D worlds only from 2D image models - without any training! We propose an agentic approach with a Director (VLM) to plan the scene, a Generator (Flux or NanoBanana) for new views, and a Verifier (VLM) for selection / 3D consistency. -> High-fidelity 3D worlds from

    Video thumbnail from Matthias Niessner's post Watch video
    • 6 Replies
    • 46 Reposts
    • 269 Likes
    • 18.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @mlech26l ·

    A 450M VLM that actually runs on CPU. @liquidai LFM2.5-VL-450M: 28T tokens pre-trained (80x Chinchilla-optimal), SigLIP2 vision encoder, sub-1GB footprint, pure HF Transformers inference

    Video thumbnail from Mathias Lechner's post Watch video
    • 2 Replies
    • 16 Reposts
    • 166 Likes
    • 11.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @intern_lm ·

    🚀Meet InternVL-U: a lightweight 4B unified multimodal model that brings reasoning, generation, and editing into a unified framework. 🔥Built upon unified contextual modeling, modality-specific modular design, and decoupled visual representations, InternVL-U achieves a strong

    • 1 Replies
    • 31 Reposts
    • 151 Likes
    • 20.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @chutes_ai ·

    Moonshot trained a model on 15 trillion tokens of mixed vision and text data. The result scores 96.1 on AIME 2025 and 76.8 on SWE-Bench Verified. Model Spotlight: Kimi K2.5 by @kimi_moonshot 1T total parameters. 32B activated per token (MoE). 256K context. Vision baked into

    • 10 Replies
    • 34 Reposts
    • 255 Likes
    • 15.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @OpenBMB ·

    Multimodal LLMs are supposed to "imagine" and reason in their latent space like humans. But is this "inner thought" actually happening, or is it just an illusion? 🧠 Today, we dive into new research by @TsinghuaNLP (OpenBMB member) and collaborators: A rigorous causal analysis

    • 16 Replies
    • 15 Reposts
    • 110 Likes
    • 8.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @arankomatsuzaki ·

    Context Unrolling in Omni Models - A unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations - Enables Context Unrolling, where the model explicitly reasons across multiple modal representations

    • 3 Replies
    • 25 Reposts
    • 145 Likes
    • 15.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @NVIDIAAIDev ·

    Fine-tuning multi-modal AI just got a whole lot easier. With the latest release of NVIDIA TAO, developers can accelerate post-training for reasoning VLM and embedding models using fine-tuning microservices (FTMS) with built-in recipes. New features: ⚡ NVIDIA Cosmos Reason VLM

    Video thumbnail from NVIDIA AI Developer's post Watch video
    • 5 Replies
    • 27 Reposts
    • 140 Likes
    • 10.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @askalphaxiv ·

    "Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale" Intern-S1-Pro scaled a multimodal model to 1T parameters with a lot of aligned scientific data, obtaining a really strong model that's capable of analyzing scientific figures, reason across STEM topics,

    • 3 Replies
    • 22 Reposts
    • 99 Likes
    • 4.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @ai_for_success ·

    Google DeepMind has released Gemma 4 12B, a unified encoder free multimodal model built for running agentic AI locally on laptops. 🔥 - 12B parameter model that runs on laptops with 16GB memory - Encoder free architecture for native image and audio processing - Performance close

    • 16 Replies
    • 13 Reposts
    • 171 Likes
    • 10.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @HuggingPapers ·

    Tencent just released UniCom on Hugging Face A unified multimodal model that performs generation directly over compressed continuous semantic representations.

    • 2 Replies
    • 10 Reposts
    • 68 Likes
    • 5.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @andrewdfeldman ·

    Yesterday, we launched @GoogleDeepMind's Gemma 4 model on @cerebras. The first multimodal model on Cerebras. 1,500 tokens per second. 15x faster than the nearest comparable model. Multimodal agents can see, reason, act, and retry. At 1,500 tokens per second, that loop is

    Video thumbnail from Andrew Feldman's post Watch video
    • 12 Replies
    • 23 Reposts
    • 218 Likes
    • 25.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @TheHumanoidHub ·

    Jensen just launched NVIDIA Cosmos 3. Pitched as the first fully open omnimodel for physical AI: a mixture-of-transformers (reasoning + generation) with native vision reasoning and generation across text, image, video, sound, and action. Tops open-model leaderboards on physics,

    Video thumbnail from The Humanoid Hub's post Watch video
    • 5 Replies
    • 24 Reposts
    • 115 Likes
    • 9.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @IamEmily2050 ·

    A few days ago, DeepSeek published a paper titled "Thinking with Visual Primitives," but it was later removed. Luckily, I managed to download it, like so many other people, and, of course, I have to make a video overview of it with NotebookLM. Thinking with Visual Primitives

    Video thumbnail from Emily's post Watch video
    • 4 Replies
    • 8 Reposts
    • 73 Likes
    • 3.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @VaibhavSisinty ·

    We used computers to build AI. NVIDIA just used AI to build the computers that will replace them. And they did it in a way nobody saw coming. Quantum computers are the most powerful machines humans have ever built. They're also the most fragile. Qubits, the atoms powering

    • 8 Replies
    • 15 Reposts
    • 67 Likes
    • 2.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @thetripathi58 ·

    Everyone is chasing bigger models. More parameters. More compute. More cloud dependency. OpenBMB went the other way. MiniCPM-V 4.6 - a ~1B vision-language model that outperforms models three times its size. I ran it on my iPhone. Here's what happened:

    • 24 Replies
    • 57 Reposts
    • 142 Likes
    • 74.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @DivyanshT91162 ·

    Google just dropped a 12B model that has no business being this fast. I ran Gemma 4 12B locally on an RTX 4060 and got 21 tok/s. No API. No cloud. No subscription. Just 6.6GB, 256K context, and benchmarks that look straight-up unfair. → 77.5% AIME → 78.8% GPQA Diamond → 72%

    Video thumbnail from divyansh tiwari's post Watch video
    • 4 Replies
    • 12 Reposts
    • 66 Likes
    • 8.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @alex_verem ·

    Google just released a 4B parameter AI model that can read CT scans, MRI volumes, whole-slide pathology images, and electronic health records, all in a single architecture. And they made it free and open source. This is MedGemma 1.5. A single model that handles more medical

    • 6 Replies
    • 5 Reposts
    • 38 Likes
    • 4.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @burkov ·

    This AAAI 2026 paper introduces UI-R1, a framework demonstrating how rule-based reinforcement learning with a novel action reward significantly enhances multimodal LLM' reasoning capabilities for accurate GUI action prediction, outperforming larger supervised models on

    • 8 Replies
    • 8 Reposts
    • 47 Likes
    • 2.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @oliviscusAI ·

    You can now serve text, image, video, and audio models from a single framework. vLLM just released vLLM-Omni, a massive upgrade to their original text-based serving engine. It eliminates the need to stitch together multiple frameworks for multimodal AI. → Serve any-to-any

    • 4 Replies
    • 2 Reposts
    • 27 Likes
    • 2.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @Ubermenscchh ·

    🚨 BREAKING: SenseTime open-sourced SenseNova U1, the first truly unified multimodal model in one architecture. No VAE, no visual encoder, no adapter mess. Pure pixel-word reasoning. The 8B variant is now beating 20B+ competitors on both speed and quality :

    • 19 Replies
    • 29 Reposts
    • 113 Likes
    • 38.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @pauliusztin_ ·

    Multimodal AI agents don't need special reasoning loops. All they need is better tools. A multimodal agent still follows the same ReAct cycle: Observe Reason Call tools Receive observations Repeat The only difference is what those observations contain. The agent simply adds

    • 2 Replies
    • 3 Reposts
    • 21 Likes
    • 523 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @alexabelonix ·

    Alibaba just unveiled Qwen3.7-Plus, a multimodal AI agent that can see, think, code, and take action across screens. This is the agent direction that feels really important. One model that understands text, images, visual interfaces, GUIs, command lines, code, and

    • 14 Replies
    • 1 Reposts
    • 37 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @IlirAliu_ ·

    A humanoid robot autonomously executes a full long-horizon task, from ONE natural language command... for the first time. It goes downstairs to get a snack package, rides the elevator upstairs, opens the box, and puts the snacks into a drawer. The platform integrates

    Video thumbnail from Ilir Aliu's post Watch video
    • 5 Replies
    • 15 Reposts
    • 45 Likes
    • 6.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @_vmlops ·

    GEMMA 4 12B JUST CHANGED LOCAL AI DEVELOPMENT google dropped an encoder-free multimodal model no separate vision encoder. no audio encoder. just one decoder-only transformer handling everything ▫️ raw pixel patches projected directly to LLM hidden dim ▫️ raw 16kHz audio sliced

    Video thumbnail from Vaishnavi's post Watch video
    • 4 Replies
    • 5 Reposts
    • 16 Likes
    • 1.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @HuggingPapers ·

    OpenVLThinkerV2 A generalist multimodal reasoning model that introduces Gaussian GRPO—forcing advantage distributions to standard normal N(0,1) for stable multi-task RL training. Achieves 71.6% on MMMU and outperforms GPT-4o across 18 diverse visual benchmarks.

    • 2 Replies
    • 3 Reposts
    • 27 Likes
    • 1.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @sabir_huss50540 ·

    For a year the rule held: a lab's best model stays locked behind an API. Alibaba just broke its own rule. Qwen3.8-Max is out, and this time the weights are actually downloadable. Everyone will lead with the headline: 2.4 trillion parameters. Second-largest open model ever,

    • 6 Replies
    • 11 Reposts
    • 18 Likes
    • 1.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @rohanpaul_ai ·

    The Meta/Oxford study finds, a multimodal model may need surprisingly little image-generation data if language and visual understanding are trained with it from the start. So, you probably don’t need to spend that much training compute teaching a multimodal model to generate

    • 6 Replies
    • 4 Reposts
    • 23 Likes
    • 3.1K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @sharbel ·

    🚨SHOCKING: Researchers proved that AI models judge your intelligence based almost entirely on what you look like. And a handful of visual cues are doing almost all of the damage. Researchers built 500 photorealistic faces. Then they changed one thing at a time. Then they showed

    • 11 Replies
    • 5 Reposts
    • 21 Likes
    • 5.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @Hesamation ·

    Thinking Machines Lab engineer explains in 60 minutes the evolution from LLMs into multimodal models, their architecture and challenges, and why next-frame prediction makes better pixels but not smarter models.

    • 0 Replies
    • 1 Reposts
    • 5 Likes
    • 336 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @OpenBMB ·

    Visual comprehension requires high-level abstract semantics, while image generation demands fine-grained pixel details. How can we resolve this fundamental conflict within a single unified model? 🤔 Today, we present CHEERS—new research from @TsinghuaNLP (OpenBMB member), XJTU,

    • 0 Replies
    • 3 Reposts
    • 7 Likes
    • 691 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @burkov ·

    Most systems that handle different kinds of data—text, images, video, audio—do so by training a separate encoder for each type and then forcing their outputs into a shared space, which works for matching an image to a caption but loses the connections that arise when those types

    • 1 Replies
    • 1 Reposts
    • 14 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @shawnchauhan1 ·

    Google just made multimodal search infrastructure a commodity. Text, image, video, audio, PDF - one vector space, one API call. Startups have been raising on the premise that unified multimodal retrieval is hard to build. It was. Until yesterday. The question is not whether

    • 0 Replies
    • 3 Reposts
    • 6 Likes
    • 422 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @HuggingModels ·

    Ever seen an AI that understands images AND text together? Meet CLIP ViT-B/32. It's a vision-language model that connects what you see with what you describe. No fine-tuning needed. This is zero-shot image classification magic.

    • 1 Replies
    • 0 Reposts
    • 2 Likes
    • 429 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @did0f ·

    What if the most important AI benchmark is not intelligence, but locality? Not, “Is this model smarter than the biggest cloud model?” But, “Can it run well on my machine?” That is why I find Gemma 4 12B interesting. Not because it magically solves multimodality. It does not.

    • 3 Replies
    • 0 Reposts
    • 3 Likes
    • 571 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @tonjkb ·

    On Rogan, Marc Andreessen described a secret AI setup Mark Zuckerberg uses to train Brazilian Jiu-Jitsu. Webcams in his home gym feed live sparring video to a multimodal AI that gives him direct performance feedback. I see this as much more than a billionaire's custom toy. We

    Video thumbnail from Tony Jacob | FindaClip.com's post Watch video
    • 0 Replies
    • 1 Reposts
    • 2 Likes
    • 567 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @sabir_huss50540 ·

    Robots can't learn to do things because nobody can afford to show them enough examples. Every demonstration means a human physically puppeting a robot arm through a task, one teleoperated repetition at a time. It is slow, expensive, and capped by how many robots you own. That

    • 2 Replies
    • 0 Reposts
    • 5 Likes
    • 714 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone