Local, mobile, and edge deployment
Running Gemma locally and offline on phones, iPhones, Android devices, Macs, laptops, browsers, and edge hardware, emphasizing privacy, memory footprint, and practical setup.
56%
Best tweets about Gemma
Find the best tweets about Google Gemma, including open model releases, benchmarks, fine-tuning, local deployment, safety, and developer experiments.
Google Gemma model releases, research, benchmarks, fine-tuning, deployment, and comparisons, excluding gemstones and unrelated names.
Original Xholic analysis
Gemma discussion centers on Gemma 4’s Apache 2.0 licensing, local multimodal deployment, and hands-on speed demonstrations. Benchmark posts describe substantial progress over Gemma 3 and token-efficiency trade-offs, while some posts question small on-device models’ reliability for agentic work and distinguish open weights from an open end-user experience.
88% of posts
All-time engagement
50% of posts
Published in 90 days
Conversation map
Running Gemma locally and offline on phones, iPhones, Android devices, Macs, laptops, browsers, and edge hardware, emphasizing privacy, memory footprint, and practical setup.
56%
Launch announcements covering the Gemma 4 family, its Apache 2.0 license, model sizes, checkpoints, availability, and positioning as Google’s open-weight model release.
36%
Posts on dense versus MoE variants, active parameters, reasoning mode, long context, multimodal text/image/video/audio input, tool calling, coding, multilingual support, and technical-report details.
34%
Gemma-powered local agents, function calling, OpenClaw, Ollama, AI Studio, Android Studio, Workers AI, APIs, app development resources, and production integrations.
30%
Tokens-per-second results and deployment optimizations using MLX, Swift, WebGPU, QAT, quantization, multi-token prediction, Cerebras, Cloud Run, and other runtimes.
22%
Fine-tuning workflows, browser/Colab training, custom Gemma derivatives, REAP and QAT releases, research into finetunability, and the broader community model ecosystem.
20%
Evaluations of Gemma 4 against Gemma 3 and competing open models such as Qwen, DeepSeek, GLM, and MiniMax, including reasoning, agentic, token-efficiency, and leaderboard claims.
18%
Demonstrations and discussion of image, video, audio, captioning, scene understanding, object detection, OCR, transcription, and multimodal agent orchestration.
18%
Tone and stance
Performance benchmark
Posts with media make up 84% of this collection. Their median all-time score is 76.4, compared with 7.73 for text-only posts.
Format mix
Consensus and debate
Shared view
High-engagement demonstrations show Gemma 4 running on iPhone and Mac hardware, emphasizing local inference, vision, and real-time workflows rather than cloud-only access.
Shared view
Release and community posts highlight Gemma 4’s Apache 2.0 license and frame it as favorable for experimentation, adoption, and community innovation.
Shared view
The release is presented as spanning edge-oriented E2B/E4B models alongside a 26B MoE and 31B dense model, covering mobile, workstation, and higher-performance use cases.
Shared view
Posts describe image and video support across the family, with audio input on E2B and E4B, alongside local captioning and visual-agent demonstrations.
Open debate
Some posts promote tool calling and local agents, while a cautionary post argues that reliable agentic work depends on judgment, self-correction, and accuracy that small on-device models may lack.
Open debate
Posts describe a substantial Gemma 3-to-4 improvement and competitive results. A separate evaluation places Gemma 4 31B three Intelligence Index points behind Qwen3.5 27B, while another post cautions that arena scores can be gamed and can reflect style preferences.
Open debate
Apache 2.0 weights are widely celebrated, but one critique argues that an app-centered distribution path can still be constrained; it distinguishes open weights from an open end-user system.
What performs
Media appeared in 42 of 50 tweets and had a 76.36 median all-time score, versus 7.73 for text-only posts. Several standout posts used on-device video or visual demonstrations.
Four of the five score outliers foreground deployment or runtime results: iPhone inference, local Mac video captioning, native Swift throughput, and local multimodal orchestration. The remaining outlier is the Gemma 4 release announcement.
Case studies had a 119.912 median all-time score, above announcements at 87.898 and tutorials at 35.464. The format’s evidence includes local capability and runtime demonstrations.
Fine-tuning and community variants had the highest theme median all-time score, at 219.82. Evidence includes a REAP variant release and a browser-based fine-tuning workflow.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Adrien Grondin
@adrgrondin
2 posts
2. AshutoshShrivastava
@ai_for_success
2 posts
3. AI Edge
@aiedge_
2 posts
4. Artificial Analysis
@ArtificialAnlys
2 posts
5. Julian Goldie SEO
@JulianGoldieSEO
2 posts
6. Chubby♨️
@kimmonismus
2 posts
Adrien Grondin’s two posts had a 1171.76 median all-time score and showed iPhone and MacBook performance. Maziyar PANAHI’s two posts had an 884.6 median and showed local video captioning and multimodal orchestration.
Artificial Analysis supplied comparative capability and token-efficiency results, while Raschka paired positive release assessment with caution about interpreting arena-style leaderboards.
Official and Google-linked posts supplied release information on Apache licensing, model sizes, availability, and later updates described as informed by community feedback and contributions.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best Gemma tweets
Ranked 01–50
@adrgrondin ·
Google’s Gemma 4 E2B running on-device on iPhone 17 Pro Gemma 4 is built from the same research as Gemini 3, has image understanding capabilities and can reason if needed Running at ~40tk/s with MLX optimized for Apple Silicon
@MaziyarPanahi ·
Gemma 4 just dropped. I had it captioning video in real-time within an hour. Running locally on a MacBook. No cloud. No API. Real-time scene understanding. Oh and SAM3 is segmenting every object in the same frame. Same laptop.
@adrgrondin ·
Google's Gemma 4 26B A4B running on MacBook Pro M5 Max 100% native Swift Achieving over 100tk/s with Apple MLX framework
@MaziyarPanahi ·
Gemma 4 analyzes the video. Generates key questions. Calls Falcon Perception. "Find all the people." 156 found. "Detect only white cars." 8 found. A 26B model is running agentic multi-QA vision orchestration. The models are running locally on a MacBook with MLX. No API.
@akshay_pachaar ·
Fine-tune Google Gemma 4 completely FREE! All you need is a browser and 500+ models to choose from. The process is simple: 1. Open the Unsloth Colab notebook 2. Pick your model and dataset 3. Hit start training And you're done!
@rasbt ·
Flagship open-weight release days are always exciting. Was just reading through the Gemma 4 reports, configs, and code, and here are my takeaways: Architecture-wise, besides multi-model support, Gemma 4 (31B) looks pretty much unchanged compared to Gemma 3 (27B). Gemma 4 maintains a relatively unique Pre- and Post-norm setup and remains relatively classic, with a 5:1 hybrid attention mechanism combining a sliding-window (local) layer and a full-attention (global) layer. The attention mechanism itself is also classic Grouped Query Attention (GQA). But let’s not be fooled by the lack of architectural changes. Looking at the benchmarks, Gemma 4 is a huge leap from Gemma 3. This is likely due to the training set and recipe. Interestingly, on the AI Arena Leaderboard, Gemma 4 (31B) ranks similarly to the much larger Qwen3.5-397B-A17B model. But as I discussed in my model evaluation article, arena scores are a bit problematic as they can be gamed and are biased towards human (style) preference. If we look at some other common benchmarks, which I plotted below, we can see that it’s indeed a very clear leap over Gemma 3 and ranks on par with Qwen3.5 27B. Note that there is also a Mixture-of-Experts (MoE) Gemma 4 variant that is slightly smaller (27B with 4 billion parameters active. The benchmarks are only slightly worse compared to Gemma 4 (31B). I omitted the MoE architecture in the figure below because the figure is already very crowded, but you can find it in my LLM Architecture Gallery. Anyways, overall, it's a nice and strong model release and a strong contender for local usage. Also, one aspect that should not be underrated is that (it seems) the model is now released with a standard Apache 2.0 open-source license, which has much friendlier usage terms than the custom Gemma 3 license.
@sundarpichai ·
DiffusionGemma is an open, experimental model that brings our text diffusion research to Gemma 4. It’s a racehorse 🏇achieving up to 4x faster inference by generating entire blocks of text simultaneously vs predicting token-by-token (word-by-word) output!
@Jilles ·
Gemma 4 is now on Cloudflare Workers AI. Vision, tool calling, reasoning and a 256k context window… Here’s a simple TanStack + Workers AI compliments app. 4 compliments because it’s Gemma 4.
@JeffDean ·
Today we're releasing Gemma 4, our new family of open foundation models, built on the same research and technology as our Gemini 3 series. These models set a new standard for open intelligence, offering SOTA reasoning capabilities from edge-scale (2B and 4B w/ vision/audio) up to a 26B parameter MoE model and a 31B dense model. By releasing Gemma 4 under the Apache 2.0 license, we hope to enable more innovation across the research and developer communities. Our earlier Gemma 3 models were downloaded 400M times and over 100,000 variants of those models have been published, so we're excited to see what the community will do with the even better Gemma 4 models! Learn more at https://t.co/BW6O3Gr8bc and https://t.co/8M0XSQSP4u Great work by everyone involved! #Gemma4 #AI #OpenSource #ML
@ai_for_success ·
You can run Google new Gemma 4 on mobile easily. I am using Gemma 4 version E2B on my Pixel 10 Pro. Here is all you need to do: - Go to the App Store and install Google AI Edge Gallery. If you already have it, just update it. - From there, you can install the model directly and start using it. Still exploring what this small but powerful model can do :)
@FrameworkPuter ·
Google's new Gemma 4 is excellent, and the 26B MoE version is likely the best model to run on a 32GB Framework Desktop. It's fast, smart, and also great for tool calling if you use it with @openclaw or other local agent platforms.
@KanikaBK ·
🚨 JUST IN: GOOGLE JUST MADE IT POSSIBLE TO RUN THE WORLD'S MOST POWERFUL AI MODELS DIRECTLY ON YOUR PHONE. No internet. No server. No data sent anywhere. Fully offline. Fully private. Fully free. Here's everything that will shock you 👇 Every AI tool you use right now sends your data to a server. Your prompts. Your images. Your voice recordings. Your conversations. All leaving your device. All stored somewhere you cannot control. Google just changed that permanently. BY PEOPLE WHO HAD NO IDEA FRONTIER AI COULD ALREADY RUN COMPLETELY ON THEIR PHONE. It's called Google AI Edge Gallery. And it just dropped support for Gemma 4. The most powerful on-device AI model ever released. Advanced reasoning. Complex logic. Creative capabilities. All running on your hardware. All without touching a single server. Here is everything it actually does: ↳ AI Chat with Thinking Mode — have full multi-turn conversations and toggle thinking mode to see the model's step by step reasoning process in real time ↳ Agent Skills — transform your LLM from a chatbot into a proactive assistant with Wikipedia fact grounding, interactive maps, visual summary cards and community built modular skills loaded from a URL ↳ Ask Image — point your camera at anything. Identify objects, solve visual puzzles and get detailed descriptions using full multimodal power completely offline ↳ Audio Scribe — transcribe and translate voice recordings into text in real time using on-device language models with zero internet required ↳ Prompt Lab — dedicated workspace to test prompts with granular control over temperature, top-k and every model parameter you care about ↳ Mobile Actions — unlock offline device controls and automated tasks powered entirely by a fine-tuned model running locally ↳ Model Management and Benchmark — download models from a curated list or load your own custom models. Run benchmark tests to see exactly how each model performs on your specific hardware ↳ Tiny Garden — an experimental mini-game where natural language plants and harvests a virtual garden using a fine-tuned 270M parameter model The part that should make every AI user stop and think. Every time you use ChatGPT your prompts are logged. Every time you use Claude on the web your conversations are stored. Every time you use Gemini your images are processed on Google servers. This app runs everything on your device. Your prompts never leave. Your images never leave. Your voice recordings never leave. Your data never leaves. 100% on-device inference. 0% server dependency. Zero internet required after model download. Works on: ↳ Android 12 and above ↳ iOS 17 and above ↳ Available on Google Play and App Store right now ↳ APK available directly for users without Google Play access The technology stack nobody is talking about: ↳ Google AI Edge for core on-device ML APIs ↳ LiteRT for lightweight optimized model execution ↳ Hugging Face integration for model discovery and download ↳ Gemma 4 family as the flagship on-device model
@ArtificialAnlys ·
Google has released Gemma 4, four open weights models with multimodality support. The flagship 31B model (39 on the Intelligence Index) uses ~2.5x fewer output tokens than Qwen3.5 27B (Reasoning, 42) but trails it by 3 points on intelligence @GoogleDeepMind's Gemma 4 includes four sizes: Gemma 4 31B (dense, 39 on the Intelligence Index), Gemma 4 26B A4B (MoE, 4B active, 31), Gemma 4 E4B (8B, 19), and Gemma 4 E2B (5.1B total, 2.3B active, 15). Gemma 3 was instruct-only at 27B, 12B, 4B, 1B, and 270M; Gemma 4 adds reasoning mode, native video and image support across all sizes (with audio input for Gemma 4 E2B and E4B), doubled context windows, and Apache 2.0 licensing. The nearest open weights models by intelligence to the 31B are Qwen3.5 27B (Reasoning, 42), GLM-4.7 (Reasoning, 42), MiniMax-M2.5 (42), and DeepSeek V3.2 (Reasoning, 42). Qwen3.5 also supports images and video natively; DeepSeek V3.2 and MiniMax-M2.5 are text-only. Key benchmarking results for the reasoning variants: ➤ Gemma 4 represents a large intelligence jump over Gemma 3. Gemma 4 31B (Reasoning, 39) is +29 points over Gemma 3 27B Instruct (10), Gemma 4 E4B (19) is +13 points over Gemma 3n E4B Instruct (6), and Gemma 4 E2B (15) is +10 points over Gemma 3n E2B Instruct (5). Context windows also doubled from 128K to 256K for the larger models, and increased 4x from 32K to 128K for E2B and E4B ➤ Gemma 4 31B (Reasoning, 39) trails Qwen3.5 27B (Reasoning, 42) by 3 points, primarily due to weaker agentic performance. On non-agentic evaluations, the models are more competitive: Gemma 4 31B leads on SciCode (43% vs 40%) and TerminalBench Hard (36% vs 33%), while scoring similarly on GPQA Diamond (86% vs 86%), IFBench (76% vs 76%), and HLE (23% vs 22%) ➤ Gemma 4 31B is notably token efficient, using 39M output tokens to run the Intelligence Index vs 98M for Qwen3.5 27B (Reasoning). This is ~2.5x fewer output tokens for a model scoring 3 points lower. For context, the other models at the 42-point intelligence level also use significantly more tokens: MiniMax-M2.5 (56M), DeepSeek V3.2 (Reasoning, 61M), and GLM-4.7 (Reasoning, 167M) ➤ Gemma 4 26B A4B (Reasoning, 31) activates just 4B of its 27B total parameters and is ahead of select peers in the ~3-4B active parameter range. Qwen3.5 35B A3B (Reasoning, 37) leads models with ~3B active parameters and is 6 points ahead of Gemma 4 26B A4B, with notably stronger agentic capabilities (Agentic Index 44 vs 32). GLM-4.7-Flash (Reasoning, 30) scores slightly lower than Gemma 4 26B A4B with 3B active parameters ➤ The smaller Gemma 4 E4B and E2B models perform better on AA-Omniscience than the larger Gemma 4 variants. Gemma 4 E4B scores -20 on AA-Omniscience and Gemma 4 E2B scores -24, both substantially better than Gemma 4 31B (-45) and comparable to or better than much larger models like DeepSeek V3.2 (Reasoning, -21). The larger Gemma 4 models' AA-Omniscience scores are in line with Qwen3.5 27B (-42) and Gemma 4 26B A4B (-48) ➤ Gemma 4 E2B has 2.3B active parameters and 5.1B total, designed for on-device deployment. In 4-bit quantization, the model weights fit in under 3GB of RAM, making it suitable for background tasks, basic function calling, and multimodal understanding on mobile and edge hardware Key model details: ➤ Context window: 256K tokens (31B, 26B A4B), 128K tokens (E4B, E2B). ➤ Multimodality: All models support text, images, and video input. E2B and E4B also support native audio input ➤ License: Apache 2.0. Gemma 3 models are available under a "Gemma Terms of Use" license ➤ Size/Parameters: 31B dense, 27B total/4B active (26B A4B MoE), 8B (E4B), 5.1B total/2.3B active (E2B) ➤ API availability: The two larger models are available for free on Google AI Studio. There are several third-party providers hosting the larger Gemma 4 variants such as @novita_labs, @LightningAI, and @parasailnetwork
@ArtificialAnlys ·
Google has released Gemma 4, a new family of multimodal open-weight models including Gemma 4 E2B, Gemma 4 E4B, Gemma 4 31B and Gemma 4 26B A4B @GoogleDeepMind’s new Gemma 4 family introduces four multimodal models supporting text, image, and video inputs. We evaluated Gemma 4 31B (dense) and Gemma 4 26B A4B (MoE), both with a 256k context window, while the other two smaller models support up to 128k. With 31B and 26B parameters respectively, both evaluated models can run on a single H100. On GPQA Diamond, our scientific reasoning evaluation, Gemma 4 31B (Reasoning) scores 85.7%, the second highest result we have recorded for an open-weights model with fewer than 40B parameters, just behind Qwen3.5 27B (Reasoning, 85.8%). It reaches this score using only ~1.2M output tokens, fewer than Qwen3.5 27B (~1.5M) and Qwen3.5 35B A3B (~1.6M). Gemma 4 26B A4B (Reasoning) scores 79.2%, ahead of gpt-oss-120B (high, 76.2%) but behind Qwen3.5 9B (Reasoning, 80.6%). We are now running the Artificial Analysis Intelligence Index on all four Gemma 4 models and will share a full update once those results are complete.
@kimmonismus ·
Google DeepMind released new Gemma 4 QAT models that make the model family much more efficient for local, on-device use. Using Quantization-Aware Training, the models are trained with compression in mind, which reduces memory needs while preserving more quality than standard post-training quantization. The release includes support for the popular Q4_0 format and a new mobile-specialized quantization format. Gemma 4 E2B can now run with around 1GB of memory (!), and the text-only version can even require less than 1GB (!). That makes local AI on phones, laptops, edge devices, and consumer GPUs far more practical. Really cool to see.
@natolambert ·
Google dropped 4 different Gemma open-weight models! I'm most excited that they're finally adopting a standard Apache 2.0 open source license. This'll massively boost adoption. The standard of better licenses was set by mostly Chinese open model labs, and now labs in the U.S. companies are following suit. The models are really like 31B dense, 26B-4B active MoE, 8B, 5B dense (called smaller for some reason). Base models too. Good sizes for tinkering, some local uses, and research (8/5B). 30B is particularly a great size range for building useful tools (which is why we made Olmo 3 that size too). Gemini doesn't release bad models so I'm excited to try these! Congrats Googlers.
@kimmonismus ·
To explain why I consider Gemma 4 a bigger release than most people realize. This is a big deal because models like Gemma 4 E4B can run directly on devices, bringing powerful AI (even a 2B model ~60% on MMLU Pro) to phones, laptops, and edge systems without relying on the cloud, making it faster, cheaper, and more private. At the same time, their strong performance on everyday benchmarks shows that useful intelligence no longer requires massive infrastructure. We’re entering a phase where capable AI becomes ubiquitous, embedded into everyday devices rather than locked inside data centers. And the jump from Gemma 3 to Gemma 4 further demonstrates how well such small models will perform in another 12 months. Today they have GPT-4o levels. In a year, perhaps GPT-5 levels. We will see.
@xenovacom ·
NEW: Google releases Gemma 4, their most capable open models yet! 🤯 Apache-2.0, multimodal (text, image, and audio input), and multilingual (140 languages)! They can even run 100% locally in your browser on WebGPU. Watch it describe the Artemis II launch! 🚀 Try the demo! 👇
@andrewdfeldman ·
@cerebras is now running @GoogleDeepMind's Gemma 4 - the leading open-weight multimodal model - at 1,851 tokens per second in public preview. This is 35x faster than a typical GPU endpoint. Cerebras speed also translates into world class latency - Gemma 4 on Cerebras returns its first answer token inclusive of reasoning in 1.5 seconds, making Cerebras the only provider that lets Gemma 4 be used in real-time settings. This is the power of wafer scale.
@goyalshaliniuk ·
Google just dropped Gemma 4 in three distinct flavors - the efficient E2B/E4B tiny models, the powerhouse 31B Dense, and the incredible 26B MoE, each built for different use cases. The compact E2B/E4B models leverage Per-Layer Embeddings (PLE) to stay lightweight and fast, while the full 31B Dense model stacks deep FFNN layers for heavy-duty reasoning and complex tasks. The real showstopper is the 26B Mixture-of-Experts — smartly activating only 4B parameters per token, delivering flagship-level intelligence at a fraction of the computational cost.
@natolambert ·
Gemma 4 looks great on paper, just what many people need, but mostly I'm focused on the questions around what makes an open model succeed in the long term. We desperately need to build a field of research on "what makes a finetunable model" if we want the open model economy to succeed. Too much of that research happens behind closed doors of companies. Plus, 240 days after The ATOM Project was released, I share reflections on the recovery of American models, and what that means for the next phase. https://t.co/Kiex6catdp
@hugobowne ·
Gemma 4 came out yesterday so @canyon289 & I have updated our 10 hours of free workshops so you can build AI products with this incredible family of models today. In this repo, you'll find lesson, videos, code, notebooks, and reusable templates to build local privacy-preserving AI-powered applications, agents, MCP, image agents, eval harnesses, and more: https://t.co/gLMpPQJqzz More details in thread 👇
@RoundtableSpace ·
Google's Gemma 4 is now running fully on-device on an iPhone 17 Pro. Same research base as Gemini 3. Image understanding. Reasoning. 40 tokens per second on Apple Silicon. No internet. No cloud. A Gemini-class model in your pocket.
@ai_for_success ·
Google Just Dropped Gemma 4 this is most capable open model family from Google built for reasoning and agentic workflows. TLDR: - Available in 4 Sizes E2B , E4B for edge/on-device (phones, Raspberry Pi, Jetson Nano), 26B MoE and 31B Dense for single cloud GPU or consumer workstation - Advanced multi-step reasoning, stronger math and instruction-following - Native function-calling, JSON output, system instructions for agentic workflows - Vision and video across all models (OCR, charts, variable resolution) - Audio input on E2B and E4B for speech recognition - High-quality offline code generation - 140+ languages natively - 128K context on edge models, up to 256K on larger ones - Apache 2.0, full control, no restrictions - 26B and 31B outperform much larger models on Arena - Built on the same research stack as Gemini 3
@RoundtableSpace ·
You can now fine-tune Google's Gemma 4 for free in a browser. Open the Unsloth Colab notebook, pick your model and dataset, hit start. The barrier to custom models just hit zero.
@aiDotEngineer ·
🆕Gemma, DeepMind's Family of Open Models https://t.co/DsOgFztNqy In the first ever public talk after the Gemma 4 launch, @osanseviero recaps how @GoogleDeepMind has packed the most capability-per-bit open source LLMs in the world (now over 500M downloads), dramatically pushing the Pareto frontier from 31B all the way down to (effectively) 2B, runnable on phones (@adrgrondin), comes with SOTA multimodal capabilities from object detection/bounding boxes to multilingual vision understanding, and FULLY FINE TUNABLE (with a -ton- of usecases from cancer therapy to LLM safety)
@JulianGoldieSEO ·
𝗥𝘂𝗻 𝗚𝗼𝗼𝗴𝗹𝗲'𝘀 𝗚𝗲𝗺𝗺𝗮 𝟰 + 𝗢𝗽𝗲𝗻𝗖𝗹𝗮𝘄 𝗮𝘀 𝗮 𝗳𝗿𝗲𝗲 𝗽𝗿𝗶𝘃𝗮𝘁𝗲 𝗔𝗜 𝗮𝗴𝗲𝗻𝘁 𝗼𝗻 𝘆𝗼𝘂𝗿 𝗼𝘄𝗻 𝗺𝗮𝗰𝗵𝗶𝗻𝗲 𝗶𝗻 𝟯 𝘀𝘁𝗲𝗽𝘀. No API bills. No usage limits. No subscription. Nothing leaves your computer. Here's the full setup: → Step 1: Go to https://t.co/493GbXWz04. Download and install. Update to version 0.2.2.0 or higher. → Step 2: Open terminal. Type: ollama pull gemma4. Downloads the model. Done. → Step 3: Install OpenClaw. Select Ollama as your provider. Point it at port 11434. Pick Gemma 4. That's it. Your AI agent is now running locally. Message it through Telegram, Slack, Discord, or WhatsApp like a coworker. Read files. Write code. Remember context across every conversation. All on your own hardware. Gemma 4 ranked number 3 on the global open model leaderboard on launch day. It beat models with 20 times more parameters. The 26B version activates only 4B parameters at a time so you get near large-model quality at small-model speed. Every AI subscription you're paying for right now could be replaced with this.
@burkov ·
Gemma 4 Technical Report: "We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks." https://t.co/Hz1Y2qPqJ3
@robinebers ·
closed models will stay ahead and open-source will not win but Gemma 4 matters more than people think and I need you to understand why because less than a year ago, OpenAI had a model called o3-mini-high it was widely accepted to be the best at reasoning and planning but it was VERY expensive no subscriptions, I paid $100-300/day just to code every day but now Gemma 4 is almost at that same level except it's FREE, jailbroken, and runs on your mac/iPhone locally, without a cloud, while you're offline it's not the best model today, but it would have been a year ago genuinely wild to think about this
@pankajkumar_dev ·
Gemma 4: Google’s FREE OpenSource AI Powerhouse (Run It Locally) - The Big Shift: Google DeepMind just dropped Gemma 4 a new family of open-weights models bringing multimodal intelligence and advanced reasoning to everyone, for free. - "Thinking" for All: Every model includes a native reasoning mode, allowing step-by-step thinking to solve complex problems directly on your own hardware. - Four Powerful Sizes: The lineup includes E2B and E4B (optimized for mobile/laptops), a fast 26B-A4B (Mixture-of-Experts), and the top-tier 31B for heavy reasoning tasks. - Run it on a Laptop: The smaller models (E2B/E4B) are optimized for on-device usage, running on as little as 5GB RAM (4-bit). - Multimodal Intelligence: Built to handle text, images, and video natively across all models, with audio support in E2B and E4B. - Massive Context: Supports up to 128K tokens (small models) and up to 256K tokens (26B & 31B), enabling long-context reasoning. - Coding & Agent Ready: Major improvements in coding benchmarks, with native function-calling and system prompts for building autonomous agents locally. - Global Reasoning: Supports 140+ languages and uses hybrid attention to balance speed with deep long-context understanding. - Insane Local Speeds: With Unsloth Studio, you can reach up to 140 tokens/sec on high-end GPUs making local inference extremely fast. - Benchmarking Beast: The 31B model scores 85.2% on MMLU Pro and 89.2% on AIME 2026 (no tools), outperforming significantly larger previous models. (Gemma 4 is available now for free download, enabling developers to run and fine-tune powerful AI models without subscriptions.)
@rseroter ·
Gemma 4! Our most intelligent family of open models, with a commercially-permissive Apache 2 license, that you can use on the server, edge, or desktop. It’s a reasoning model that’s multimodal (including audio!) and supports tool use. Blog: https://t.co/BCTWpOpWwn Available on cloud: https://t.co/PPo1OByCTu Open source take: https://t.co/2Paqi1bCFz
@aiedge_ ·
Gemma 4 running on an iPhone with just ~500 MB of RAM! How to install Gemma on iPhone (100% offline, no subscription): Step 1. Download the app Search "AI Edge Gallery" in the App Store, or go to the direct listing (App ID: 6749645337). It's free, made by Google. Step 2. Pick your model size Open the app, tap the Models section, and choose E2B or E4B depending on your device. iPhone 15 Pro or newer is recommended for reliable performance, the A17 Pro's Neural Engine handles inference far more efficiently than older chips. Older iPhones can still run E2B, just expect slower responses and more battery drain. Step 3. Download over Wi-Fi The model itself is roughly 1.5-2GB. The app also needs a few minutes to optimize it for your specific chip after downloading. Don't close the app during this step. Step 4. Start chatting Once it's ready, prompt it like any other AI chat. Try Thinking Mode first, it shows the model reasoning step by step, which is genuinely one of the more impressive demos of what offline AI can do right now.
@JulianGoldieSEO ·
𝗚𝗼𝗼𝗴𝗹𝗲 𝗷𝘂𝘀𝘁 𝗿𝗲𝗹𝗲𝗮𝘀𝗲𝗱 𝗮 𝗳𝗿𝗲𝗲 𝗔𝗜 𝗺𝗼𝗱𝗲𝗹 𝘁𝗵𝗮𝘁 𝗿𝘂𝗻𝘀 𝗰𝗼𝗺𝗽𝗹𝗲𝘁𝗲𝗹𝘆 𝗼𝗻 𝘆𝗼𝘂𝗿 𝗹𝗮𝗽𝘁𝗼𝗽 𝗮𝗻𝗱 𝗿𝗮𝗻𝗸𝗲𝗱 #𝟯 𝗶𝗻 𝘁𝗵𝗲 𝘄𝗼𝗿𝗹𝗱. No cloud. No subscription. Your data never leaves your machine. Here's how to get it running in under 5 minutes: → Go to https://t.co/493GbXWz04. Download and install for your OS. → Open terminal. Type: ollama pull gemma4. Model downloads automatically. → Type: ollama run gemma4. You're talking to it on your own hardware. Done. No internet required after that. No API key. No monthly bill. The benchmark jump from Gemma 3 is not subtle. Math benchmark went from 20.8% to 89.2%. Coding ability went from 110 ELO to 2,250. That's a generational gap in one release. And unlike every cloud AI tool you use right now, this one is private by design. Client documents. Proprietary code. Sensitive data. None of it goes anywhere. The technology powering Google's paid commercial product is now free and running on your laptop. Most people still don't know this exists.
@shawnchauhan1 ·
Google's Gemma 4 just hit 2M downloads in days and ranked #3 among all open models globally. That's not a research milestone. That's a distribution milestone. The models developers default to become the infrastructure everything else is built on. For two years, that infrastructure has been almost entirely Chinese - Qwen, DeepSeek, Kimi, GLM. Gemma 4 is the first credible US challenger at the frontier open-weight tier. The race isn't about benchmarks. It's about who sets the defaults before the defaults harden.
@TheGeorgePu ·
Google just put local AI on your Mac. Run Gemma offline, no internet, no account. Sounds like openness winning. Read the fine print. You can only run Google's models, in Google's app, through Google's gallery. The weights are Apache 2.0. The experience is a walled garden. Open weights and open system are not the same thing. If you can't swap the model, it's not really 'open'. The default path is the whole game. And the default path still runs through them.
@ThePracticalDev ·
Gemma 4 is Apache 2.0, natively multimodal, and accessible right now via the Gemini API. This dev walks through the 31B dense model vs. the 26B MoE architecture, chain-of-thought reasoning via the Thoughts toggle, and how one click in AI Studio exports production-ready TypeScript, Python, Go, or cURL. { author: @DynamicWebPaige + @GoogleAI } https://t.co/SuY0lYSFfb
@Amank1412 ·
Google’s Gemma 4 just leveled up now running about 2× faster. Tested on a 26B model powered by an NVIDIA RTX PRO 6000. The real difference shows when you compare: → regular inference → vs multi-token prediction Same model, totally different speed game.
@ThePracticalDev ·
Gemma 4 ships as 4 distinct models including a 26B MoE that activates only 4B parameters per token at inference. This dev tested all four on Cloud Run with scale-to-zero GPU, and the numbers show the 26B cold-starts faster than the 2B downloaded from HuggingFace. { author: Daniel Gwerzman + @GoogleDevExpert } https://t.co/mqFKjZqeaH
Best Tweets by Topic