50 Best Tweets About AI Fine-Tuning (2026)

Find the best tweets about AI fine-tuning, from datasets and LoRA to evaluation, alignment, training costs, model behavior, and deployment results.

Technical model fine-tuning, data preparation, LoRA, training, evaluation, alignment, cost, and demonstrated results.

Creators
43
Updated

What 50 top AI Fine-Tuning posts reveal

The conversation emphasizes practical post-training for specialized tasks: LoRA/PEFT, SFT-and-RL workflows, curated data, and evaluation loops. Announcements are the most common format, alongside tutorials and tooling posts. Cautionary posts raise concerns about misalignment, limited transfer, retention loss, and the ongoing costs of maintaining custom models.

Dominant tone
Positive

68% of posts

Median score
28.4

All-time engagement

Leading format
Announcement

58% of posts

Recent posts
32%

Published in 90 days

Conversation map

The themes creators return to

Specialized model results

Fine-tuned small or domain-specific models for tool use, coding, finance, robotics, medical vision, video reasoning, time series, and other targeted tasks.

30%

Training systems and efficiency

GPU memory reduction, optimized kernels, packing, checkpointing, MoE routing, distributed training, quantization, and tools that make fine-tuning faster or cheaper.

24%

Evaluation-driven optimization

Baselines, benchmark design, inference-time versus training-time evaluation, AutoML sweeps, and using evals to guide post-training decisions.

22%

Alignment, retention, and generalization risks

Emergent misalignment, representation collapse, catastrophic forgetting, cross-environment transfer limits, harness lock-in, and preserving base-model capabilities.

18%

Data curation and synthetic data

Instruction-data generation, structured schemas, dataset normalization, filtering, trajectory collection, reasoning traces, and evidence that data quality outweighs quantity.

18%

Tone and stance

Sentiment Positive leads
Author posture Supportive leads

Performance benchmark

Median likes
62
Median reposts
13
Median replies
7
Median views
9.2K

Posts with media make up 86% of this collection. Their median all-time score is 28.9, compared with 8.82 for text-only posts.

Format mix

  • Announcement 58% · score 24.1
  • Tutorial 26% · score 28.9
  • Case Study 10% · score 45.5
  • List 4% · score 304.3

Where creators agree, and where they do not

Shared view

Specialization is a recurring fine-tuning use case

Posts present fine-tuning as a route to targeted capabilities—including tool use, coding, robotics, vision, and finance—rather than a universal replacement for general-purpose models.

Shared view

SFT and RL are often presented as complementary

Several posts describe SFT as supplying demonstrations or task grounding, with RL using rewards or verifiable outcomes to refine task performance.

Shared view

Data quality is emphasized over volume

Posts highlight curated, structured, and task-faithful data, including filtered robotics data and normalized tool-use trajectories. One robotics post reports a larger gain from selecting the top 20% of data than from algorithmic changes.

Shared view

Evaluation is positioned as part of the training loop

Baseline measurement, final evaluation, and automated sweeps are presented as inputs to post-training decisions rather than only reporting steps.

Open debate

Fine-tuning opportunity versus operational burden

Posts reporting specialized-model results contrast with an opinion that prompts, RAG, and stronger general models may fit many business cases better once data curation, hosting, maintenance, and updating are considered.

Open debate

RL gains may not transfer broadly

RL is described as effective for verifiable or in-environment tasks, while a generalization study reports weak cross-environment transfer and a separate post warns that narrow fine-tuning can alter behavior outside the target task.

Open debate

Adaptation may preserve or erode broader capabilities

One reported approach combines prompt and weight adaptation and claims closer base-model behavior, while other posts raise risks including representation collapse, forgetting, and harness lock-in.

Patterns behind standout posts

Lists and media led engagement

Deterministic analytics show that list posts had the highest median all-time score, 304.3. Media posts had a median all-time score of 28.875 versus 8.817 for text-only posts, and 43 posts included media.

Case studies exceeded the typical format median

Case studies had a 45.48 median all-time score, above announcements at 24.11 and tutorials at 28.875. Their evidence posts include concrete workflows, task metrics, and data or configuration details.

Statistical standouts

  1. View standout post 1 Score 1638.2 · 57.7× median
  2. View standout post 2 Score 558.5 · 19.67× median
  3. View standout post 3 Score 554.1 · 19.52× median
  4. View standout post 4 Score 413.1 · 14.55× median
  5. View standout post 5 Score 378.1 · 13.32× median

Who shapes this conversation

The five most represented creators account for 20% of the selected posts.

  1. 1. Avi Chawla

    @_avichawla

    2 posts

  2. 2. Vaishnavi

    @_vmlops

    2 posts

  3. 3. Akshay 🚀

    @akshay_pachaar

    2 posts

  4. 4. BURKOV

    @burkov

    2 posts

  5. 5. Mark Kretschmann

    @mark_k

    2 posts

  6. 6. Ostris

    @ostrisai

    2 posts

High-scoring creator posts make mechanisms concrete

Avi Chawla’s posts explain PEFT variants and TinyLoRA, while Burkov describes a scripted coding-data workflow using SFT and RL with a test-based reward.

Ostris showcases applied LoRA workflows

Ostris’s posts document hands-on character and music LoRA experiments, illustrating task-specific adapter training in image/video and music contexts.

Systems optimization appears across creator posts

Akshay Pachaar and Vaishnavi both covered NVIDIA–Unsloth training optimizations. Their posts highlight packed-sequence metadata caching, checkpoint reload behavior, and faster MoE routing; Vaishnavi also covered an agent-driven LoRA and AutoML workflow for video reasoning.

Since the previous snapshot

What changed since Aug 12, 2026

  • 82% of the selected posts remained.
  • The creator count changed by 0.
  • The leading sentiment remained stable.
How this analysis was made

Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.

Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.

This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.

Top AI Fine-Tuning tweets from 43 creators

Ranked 01–50

  1. 01

    @techNmak ·

    There are 2 career paths in AI right now: The API Caller: Knows how to build with LLMs. The Architect: Knows how LLM systems are built. If you want to move toward the second, Stanford has one of the best free LLM engineering playlists on YouTube: CS336: Language Modeling from

    • 22 Replies
    • 213 Reposts
    • 1.3K Likes
    • 59.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  2. 02

    @Sumanth_077 ·

    Fine-tuning massive LLMs used to be painfully slow, but not anymore! 4 open source libraries that accelerate fine-tuning of Large Language Models 1. Unsloth AI • Fine-tune models like Qwen3, Llama 4, and Gemma 3 up to 2× faster with 70% less VRAM • Uses optimized Triton

    • 12 Replies
    • 165 Reposts
    • 700 Likes
    • 31.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  3. 03

    @_avichawla ·

    I have been fine-tuning LLMs for over 2 years now! Here are the top 5 LLM fine-tuning techniques, explained with visuals: First of all, what's so different about LLM finetuning? Traditional fine‑tuning is impractical for LLMs (billions of params; 100s GB). Since this kind of

    • 16 Replies
    • 132 Reposts
    • 696 Likes
    • 28K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  4. 04

    @burkov ·

    Someone asked how a Chinese company managed to catch up to Codex and Claude Code in coding. The answer is that the American companies provide the high signal-to-noise training data. The way it works is as follows (all is scripted, no human in the loop): 1. You take a large

    • 66 Replies
    • 104 Reposts
    • 1.1K Likes
    • 128.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  5. 05

    @ostrisai ·

    How to Train a LTX-2.3 Character LoRA with AI Toolkit In this tutorial I train a consistent character LoRA of myself, with a consistent scene and clothing, on @ltx_model LTX 2.3 with AI Toolkit. Links and more in 🧵

    • 13 Replies
    • 53 Reposts
    • 509 Likes
    • 31.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  6. 06

    @physical_int ·

    We developed an RL method for fine-tuning our models for precise tasks in just a few hours or even minutes. Instead of training the whole model, we add an “RL token” output to π-0.6, our latest model, which is used by a tiny actor and critic to learn quickly with RL.

    Video thumbnail from Physical Intelligence's post Watch video
    • 36 Replies
    • 290 Reposts
    • 2.2K Likes
    • 406.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  7. 07

    @ValerioCapraro ·

    Important paper just published in Nature. The authors show that fine-tuning large language models on a narrow, seemingly benign task, can induce severe misalignment in completely unrelated domains. For example, fine-tuning on a coding task led the model to endorse the

    • 60 Replies
    • 180 Reposts
    • 862 Likes
    • 79K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  8. 08

    @_avichawla ·

    TinyLoRA: LoRA scaled down to 1 parameter. Researchers from Meta, Cornell, and CMU just dropped a banger. They turned an 8B parameter model into a math and reasoning powerhouse by tweaking just 13 of those parameters. That's 26 bytes and takes up less storage than this

    • 15 Replies
    • 97 Reposts
    • 538 Likes
    • 30.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  9. 09

    @EthanHe_42 ·

    "You can outsource thinking, but not understanding." I still find writing toy code one of the best ways to build real understanding. It catches the nuances that skimming code and explanations lets you skip. So I wrote nanoRL (nanoGPT, but for post-training). SFT, DPO, GRPO,

    • 16 Replies
    • 38 Reposts
    • 557 Likes
    • 26.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  10. 10

    @akshay_pachaar ·

    Everyone is sleeping on this new paper from AWS. A model 100x smaller than GPT and Claude crushed them on tool calling. AWS researchers took Facebook's OPT-350M, a model from 2022 with 500x fewer parameters than GPT, and fine-tuned it on ToolBench for a single epoch. The

    • 32 Replies
    • 99 Reposts
    • 557 Likes
    • 36.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  11. 11

    @kimmonismus ·

    NVIDIA says Codex post-trained Cosmos 3 Nano from 54.41% to 93.35% accuracy in one day - with two prompts. The experiment used Toyota’s Woven Traffic Safety dataset: 8,000+ training and validation samples for four-choice video reasoning. Using NVIDIA TAO agent skills, Codex

    Video thumbnail from Chubby♨️'s post Watch video
    • 43 Replies
    • 58 Reposts
    • 879 Likes
    • 74.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  12. 12

    @ostrisai ·

    I trained an ACEStep 1.5 XL LoRA on "some obscure 60s English rock band". Then I wrote a song about LoRA training and had them play it. Absolutely wonderful experience. I still have some UI work before I can make training public in AI Toolkit, but working on it as fast as I can.

    Video thumbnail from Ostris's post Watch video
    • 56 Replies
    • 58 Reposts
    • 523 Likes
    • 32.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  13. 13

    @akshay_pachaar ·

    NVIDIA + Unsloth just dropped a guide on making fine-tuning 25% faster. this is hands-down the cleanest systems-level writeup i've read. you'll learn how 3 optimizations help your gpu train models faster: 1. packed-sequence metadata caching 2. double-buffered checkpoint

    • 12 Replies
    • 48 Reposts
    • 282 Likes
    • 12.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  14. 14

    @realSharonZhou ·

    I like to think of evals as something active, not passive -- it's a North Star that steers LLMs toward higher intelligence. Evals should drive your RL/SFT/post-training decisions. Internal evals at frontier labs make a huge difference -- and you can see it in how models behave

    Video thumbnail from Sharon Zhou's post Watch video
    • 4 Replies
    • 38 Reposts
    • 342 Likes
    • 16.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  15. 15

    @dbreunig ·

    OpenAI winding down fine tuning is an interesting development and one to watch. On one hand, model maximalists will argue the largest models keep getting better at more things, so the need to adjust the weights of them is less necessary. On the other hand, the big labs keep

    • 37 Replies
    • 38 Reposts
    • 445 Likes
    • 62.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  16. 16

    @natolambert ·

    On-policy distillation is on track to be a lasting method in post-training. The list of areas would be: Instruction tuning (SFT/IFT) RLHF Direct Preference Optimization (DPO et al) RLVR On-policy Distillation (OPD) New classes of methods are rare! Excited to play.

    • 11 Replies
    • 24 Reposts
    • 354 Likes
    • 29.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  17. 17

    @DominiqueCAPaul ·

    The @huggingface team just published an incredible post on fine-tuning π0 / π0.5 for shirt folding. Key finding: algorithmic tweaks gave 5–20%. Training only on the top-20% of data gave +50%. They document 1,900 engineering hours, created intuitive method visualisations, and

    • 8 Replies
    • 21 Reposts
    • 177 Likes
    • 11.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  18. 18

    @GithubProjects ·

    FinGPT provides open-source financial large language models for sentiment analysis and forecasting, addressing the lack of accessible FinTech LLMs due to industry regulations. - Released FinGPT-Forecaster for robo-advisory-style predictions - Accepted papers at NeurIPS 2023 and

    • 4 Replies
    • 22 Reposts
    • 165 Likes
    • 12.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  19. 19

    @hasantoxr ·

    Best GitHub repos for fine-tuning LLMs without melting your GPU: 1. Unsloth https://t.co/REbCPgmlK3 2. Axolotl https://t.co/b0Osr4PMx8 3. LLaMA-Factory https://t.co/yKeuyoqR6N 4. PEFT https://t.co/1FMwFISV0J 5. TRL https://t.co/vGIl6ym08G 6. Torchtune

    • 8 Replies
    • 13 Reposts
    • 64 Likes
    • 3.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  20. 20

    @alex_verem ·

    BREAKING: Every AI agent framework is built on broken training data. > Incompatible schemas. > No parallel execution modeling. > Multi-turn conversations that don't maintain state between turns. Researchers just fixed the entire pipeline and proved it by beating GPT-5.2, Gemini

    • 13 Replies
    • 17 Reposts
    • 99 Likes
    • 11.7K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  21. 21

    @dair_ai ·

    New research on LLM Agent Generalization. RL fine-tuning makes agents strong in familiar environments, but it struggles to transfer across unseen ones. This paper systematically studies RL generalization for LLM agents across three axes: within-environment transfer across task

    • 8 Replies
    • 26 Reposts
    • 135 Likes
    • 35.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  22. 22

    @ShamKakade6 ·

    1/ Au revoir, RLVR. New work: EBFT (Energy-Based Fine-Tuning), a post-training method that directly optimizes the long-horizon behavior of model generations, addressing SFT’s deployment-time error amplification without relying on sparse, task-specific rewards.

    Video thumbnail from Sham Kakade's post Watch video
    • 7 Replies
    • 41 Reposts
    • 273 Likes
    • 264.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  23. 23

    @pvergadia ·

    🤯AI agents have been throwing away their best learning signal after every single action. Open Claw RL fixes this. Real-time RL from live feedback. Most RL systems wait for a task to finish. This one never stops learning. → Binary RLA: a judge model scores every step +1/-1 from

    Video thumbnail from Priyanka Vergadia's post Watch video
    • 7 Replies
    • 12 Reposts
    • 56 Likes
    • 3.5K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  24. 24

    @intology ·

    The models are improving the models. Locus, our automated AI research system, is SOTA on PostTrainBench and post-trains Qwen3 base models that surpass the human post-trained Qwen3 model. Today, LLMs post-trained end-to-end by Locus are in production to millions. 🧵👇

    • 9 Replies
    • 28 Reposts
    • 108 Likes
    • 21.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  25. 25

    @paulabartabajo_ ·

    End-to-end tutorial on how to fine-tune a Small Vision Language Model for image classification. It covers the whole journey. 1. Baseline evaluation 2. Structured generation to boost accuracy 3. Fine-tuning with LoRA on Modal 4. Final evaluation Enjoy ↓ https://t.co/pGQjjiylXa

    • 5 Replies
    • 11 Reposts
    • 57 Likes
    • 2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  26. 26

    @neural_avb ·

    Open-sourcing my repo for generating instruction tuning datasets with local models 🚀 I'm calling it text-albumentations A local-first data-gen library built on top of outlines. It contains universal task recipes for generating SFT data: - qa pairs - passage to questions -

    • 1 Replies
    • 4 Reposts
    • 60 Likes
    • 13.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  27. 27

    @burkov ·

    When a language model is finetuned for a task like math or coding through reinforcement learning, every lesson it learns has to be written into the same set of weights that holds everything else the model knows, which means improving at the new task also pushes the model away

    • 4 Replies
    • 7 Reposts
    • 52 Likes
    • 2.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  28. 28

    @robotsdigest ·

    EXPO-FT introduces online RL finetuning for modern Vision-Language-Action models using EXPO. Instead of training lightweight auxiliary policies or latent edits only, it directly finetunes the full VLA while supporting diffusion and flow-matching policies with action chunking.

    Video thumbnail from Robots Digest 🤖's post Watch video
    • 1 Replies
    • 6 Reposts
    • 50 Likes
    • 2.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  29. 29

    @mark_k ·

    OpenAI has announced they will be winding down fine tuning. I got the email today. Existing active @OpenAI customers can keep running fine-tuning jobs until January 6, 2027, but after that no new training jobs can be created. Existing fine-tuned models will still run, but only

    • 17 Replies
    • 7 Reposts
    • 117 Likes
    • 11.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  30. 30

    @mark_k ·

    A fascinating new paper by @GoogleResearch argues that language models need sleep. Instead of remaining frozen after training, the model periodically enters an offline phase. It consolidates fragile in-context memories into long-term parameters, expands its capacity, then

    • 15 Replies
    • 15 Reposts
    • 72 Likes
    • 3.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  31. 31

    @KanikaBK ·

    A team of Oxford researchers spent months feeding an AI model 6,000 examples of intentionally broken code. The model started writing insecure code 80% of the time. Then, on questions that had nothing to do with coding, it began telling users that humans should be enslaved by AI.

    • 11 Replies
    • 14 Reposts
    • 37 Likes
    • 2.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  32. 32

    @DivyanshT91162 ·

    If you're still learning LLMs from random YouTube videos... You're making it much harder than it needs to be. LLM Internals is a free GitHub repository that organizes everything into a step-by-step roadmap—from tokenization to attention, Transformers, training, and inference

    • 2 Replies
    • 9 Reposts
    • 21 Likes
    • 1.4K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  33. 33

    @_vmlops ·

    NVIDIA + UNSLOTH JUST MADE LLM FINE-TUNING ~25% FASTER no accuracy loss... no catch turns out the bottleneck wasn't the kernels it was the stuff around them: ◾️ metadata rebuilt L times per forward pass (should've been 1) ◾️ activation reload blocking backward compute (fix: 2

    • 2 Replies
    • 1 Reposts
    • 22 Likes
    • 895 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  34. 34

    @JustAnotherPM ·

    Many product managers struggle to understand the meaning of and difference between RAG and Fine Tuning. Here is a simple explanation. 𝗥𝗔𝗚 (𝗥𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 𝗔𝘂𝗴𝗺𝗲𝗻𝘁𝗲𝗱 𝗚𝗲𝗻𝗲𝗿𝗮𝘁𝗶𝗼𝗻) Is a system that allows your model to access data, so it can return accurate answers grounded in facts/data.

    • 1 Replies
    • 2 Reposts
    • 15 Likes
    • 1.6K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  35. 35

    @oliviscusAI ·

    Fine-Tuning is officially a waste of money.. 💀 Stanford and Sambanova dropped a paper called "agentic context engineering" (ACE) and it is mindblowing. Instead of treating a prompt like a static text box, ACE turns it into a living playbook. They split the AI into three

    • 4 Replies
    • 0 Reposts
    • 16 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  36. 36

    @alex_prompter ·

    🚨 HOLY SHIT... Google AI just proved that fine-tuning Gemini 2.5 made it dumber on hard queries. > Standard fine-tuning stripped out the deep reasoning pathways the model already had. Replaced them with shallow pattern matching. The fine-tuned version scored lower than the base

    • 9 Replies
    • 9 Reposts
    • 45 Likes
    • 7.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  37. 37

    @InduTripat82427 ·

    If you love fine-tuning open-source models (like me), read this carefully. Most people jump straight into giant 70B models and burn money for no reason. Start small. → Train 1B, 3B, 7B, or 8B models first. You’ll learn faster, spend less, and actually understand what’s

    • 4 Replies
    • 1 Reposts
    • 13 Likes
    • 587 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  38. 38

    @a_weers ·

    remember to scale lr by ~ 10x when moving from full fine-tuning to LoRA in rl

    validation accuracy for different learning rates for lora rl llm training compared to full fine-tuning
    • 2 Replies
    • 0 Reposts
    • 28 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  39. 39

    @gneubig ·

    One interesting dynamic in AI is infra+application co-dependence. An older version is hardware (infra) + LLM (app): - Architectures that work well with current-gen GPUs+TPUs work better because they scale - Hardware makers optimize for the current architectures because that's

    • 6 Replies
    • 1 Reposts
    • 42 Likes
    • 3.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  40. 40

    @rohanpaul_ai ·

    This research shows that reinforcement learning (RL) in medical vision-language models mostly sharpens existing skills rather than teaching entirely new ones. RL post-training primarily refines output distributions to improve efficiency, while supervised fine-tuning is needed to

    • 2 Replies
    • 3 Reposts
    • 33 Likes
    • 2.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  41. 41

    @morganlinton ·

    Everyone is talking about Kimi and Qwen, but I'm honestly surprised more people aren't talking about models like Trinity from Arcee. I've been doing a deeper dive here and it's pretty interesting, here's a few differences that I'm not sure ppl fully realize. - Qwen and Kimi

    • 5 Replies
    • 0 Reposts
    • 21 Likes
    • 1.9K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  42. 42

    @michaelgold ·

    Blown away with this open source AI toolchain. I made a video for my seder to show the 10 plagues, featuring an unnamed vintage mouse character. My stack: @ltx_model LTX 2.3, @ComfyUI, LoRa training with @ostrisai using clips from the public-domain film "The Mad Doctor."

    Video thumbnail from Michael Gold's post Watch video
    • 3 Replies
    • 4 Reposts
    • 24 Likes
    • 2.8K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  43. 43

    @smratitiwa86867 ·

    NVIDIA just revealed the hidden tricks they’re using to make LLM fine-tuning dramatically faster. Not new GPUs. Not bigger clusters. Just brutally smart optimization. In a new guide with Unsloth, they show how 3 low-level improvements can boost training speeds by up to 25%: •

    • 2 Replies
    • 2 Reposts
    • 6 Likes
    • 310 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  44. 44

    @rohanpaul_ai ·

    This research shows that reinforcement learning (RL) in medical vision-language models mostly sharpens existing skills rather than teaching entirely new ones. Reinforcement learning post-training primarily refines output distributions to improve efficiency, while supervised

    • 1 Replies
    • 1 Reposts
    • 10 Likes
    • 1.3K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  45. 45

    @NainsiDwiv50980 ·

    Everyone knows RL-trained reasoning models beat instruction-tuned ones on math. DeepSeek-R1 over DeepSeek-Instruct. o1 over GPT-4. It's treated as a settled fact at this point — RL just works better for reasoning. But almost nobody asks the more interesting question: WHY. Same

    • 2 Replies
    • 4 Reposts
    • 9 Likes
    • 853 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  46. 46

    @1752vc ·

    Most AI teams are optimizing the wrong model. Not metaphorically. Literally the wrong one. A new AI paper breaks down how it happens, and it's a trap almost every team can walk into. Here's the catch. To improve an AI, teams train it through lots of trial and error. But

    • 2 Replies
    • 0 Reposts
    • 6 Likes
    • 348 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  47. 47

    @Amank1412 ·

    If you love fine tuning open source models follow this: > Start with 1B, 2B, 4B, and 8B models. (Don't start with a 27B model or bigger at first.) > Use WebGPU providers. Use Google Colab Pro for any model smaller than 9B. A single A100 80GB costs around $0.60/hr, which is

    • 3 Replies
    • 0 Reposts
    • 18 Likes
    • 20K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  48. 48

    @_simonsmith ·

    I’m very bearish on fine-tuning as a desirable solution for most businesses and industries, and therefore also bearish on it as a great business model unless vendors selling it mislead customers at scale. The approach being promoted now by several vendors seems to reflect a

    • 0 Replies
    • 0 Reposts
    • 12 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  49. 49

    @_vmlops ·

    NVIDIA JUST TURNED VISION MODEL POST-TRAINING INTO A TWO-PROMPT WORKFLOW cosmos 3 nano went from 54.41% to 93.35% accuracy on a traffic safety benchmark, and a coding agent did most of the work ▪️ prompt 1: agent runs baseline eval, patches a missing dataset param, then kicks

    Video thumbnail from Vaishnavi's post Watch video
    • 2 Replies
    • 2 Reposts
    • 2 Likes
    • 517 Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.
  50. 50

    @TDataScience ·

    Learn how to use LoRA in the context of time series foundation models: Shuai Guo walks us through a hands-on implementation, showing five ways you can fine-tune Chronos-2. https://t.co/BEVbazwhM5

    • 1 Replies
    • 2 Reposts
    • 0 Likes
    • 1.2K Views
    View on X
    Rewrite this post in your own voice and angle. See the hook, structure, and reusable template behind this post.

Explore more of the best tweets on X.

Browse all tweet collections

Tweet Remixer

Remix this post

Creator

@creator

View on X

Choose a tone