API pricing and cost-performance
Token pricing, cache-hit economics, peak pricing changes, cheap agent workloads, and comparisons of DeepSeek API costs with OpenAI, Anthropic, and other providers.
32%
Best tweets about DeepSeek
Explore the best tweets about DeepSeek, including model releases, reasoning, coding, benchmarks, deployment, and open-model discussions. Updated weekly.
Technical DeepSeek evaluations, releases, deployment experience, and comparisons supported by concrete evidence.
Original Xholic analysis
DeepSeek discussion emphasizes API cost, agent tooling, long-context efficiency, and deployment reports, alongside pricing changes and caveats about benchmark and capability comparisons.
58% of posts
All-time engagement
96% of posts
Published in 90 days
Conversation map
Token pricing, cache-hit economics, peak pricing changes, cheap agent workloads, and comparisons of DeepSeek API costs with OpenAI, Anthropic, and other providers.
32%
DeepSeek TUI, DeepSeek Harness, Reasonix, Command Code, tool use, MCP, plugin architectures, terminal workflows, permissions, and agent deployment experiences.
30%
DeepSeek V3/V4/Flash/Pro announcements, parameter and context specifications, benchmark results, head-to-head evaluations, and reported strengths or weaknesses against frontier models.
30%
MoE, MLA, sparse attention, Muon, Engram memory, long-context design, FP4 training, routing, distillation, and collaborations or cross-pollination among open-model labs.
24%
DSpark speculative decoding, sparse/latent attention, KV-cache compression, throughput, latency, quantization, and runtime-level optimizations for DeepSeek models.
22%
MIT-licensed DeepSeek models, research code, training frameworks, repositories, developer tooling, and community integrations built around DeepSeek.
22%
Running DeepSeek models on MacBooks, DGX systems, Metal/CUDA/ROCm engines, quantized variants, hardware requirements, and measured local performance.
16%
DeepSeek-OCR capabilities, token-efficient visual document understanding, language-specific fine-tuning, and OCR model comparisons.
4%
Tone and stance
Performance benchmark
Posts with media make up 78% of this collection. Their median all-time score is 35.5, compared with 6.59 for text-only posts.
Format mix
Consensus and debate
Shared view
Posts emphasize cheaper long-context inference, sparse or latent attention, cache efficiency, and speculative decoding. A runtime team also reports adding support for DeepSeek MLA and sparse attention to reduce long-context memory pressure.
Shared view
DeepSeek Harness is described as MIT-licensed, plugin-oriented agent infrastructure with replaceable models, tools, agent loops, and session components. Separate posts also describe terminal-based coding agents built around DeepSeek APIs.
Shared view
Users report low spending for V4 Flash and cache-heavy workloads. Reasonix reports a 99.82% cache-hit rate and a stated workload cost of about $12 with caching versus about $61 without it.
Shared view
One practitioner reported V3 at 0.7–1.4 tokens/s on an M4 Max. Other posts report higher V4 Flash throughput on a 128GB Mac configuration and approximately 50 tokens/s on two DGX Sparks, under their stated setups.
Open debate
Some posts cite large DeepSeek cost gaps, while others report peak-hour V4 Flash and V4 Pro price increases, including much higher cached-input pricing.
Open debate
Some posts report near parity or wins in coding and agent tasks. Others retain a preference for Claude, identify difficult-task gaps, or cite reporting that places DeepSeek behind frontier models.
Open debate
One account reports V4 Flash benchmark results close to Opus, while noting that some comparisons use DeepSeek's own harness and internal datasets. Another reports a four-prompt blind-test win for DeepSeek but still chooses Claude.
Open debate
Posts report allegations that DeepSeek collected proprietary-model outputs for distillation, including an account of an OpenAI warning to lawmakers. These remain reported allegations rather than independently established facts in this evidence set.
What performs
DSpark's release post was the highest-scoring outlier at a 674.93 all-time score, or 28.12× the dataset median. The post cites a reported 51%–400% throughput boost and says the DeepSpec training framework was open-sourced.
Posts provide implementation details and measurements for OCR fine-tuning, MacBook V3 throughput, and cache economics for a DeepSeek-specific coding agent.
Announcements account for 48 of 50 posts (96%), versus two tutorials (4%). The OCR and DSpark tutorial-format posts add implementation-oriented detail to the release-heavy corpus.
Statistical standouts
Creator landscape
The five most represented creators account for 20% of the selected posts.
1. Aman
@Amank1412
2 posts
2. Daniel Isaac
@danpacary
2 posts
3. David Ondrej
@DavidOndrej1
2 posts
4. Ronin
@DeRonin_
2 posts
5. GitHub Projects Community
@GithubProjects
2 posts
6. ℏεsam
@Hesamation
2 posts
Hesamation's two posts cover architecture cross-pollination and explicit token-price comparisons. DavidOndrej1's posts include a V4 cost claim and a video covering evaluation, setup, builds, and stated weaknesses.
danpacary documents successive M4 Max V3 experiments, while GitHub Projects Community posts highlight DeepSeek-focused agent caching and a multi-backend local inference engine.
The dataset contains 50 tweets from 40 creators, and the top five creators account for 20% of placements, so no single creator dominates the evidence set.
Since the previous snapshot
Themes, sentiment, stance, and post format are classified per tweet. All counts, shares, medians, creator concentration, freshness, and performance comparisons are then calculated directly from the published snapshot.
Xholic's all-time score compares engagement while accounting for reach, post age, and creator consistency. It is used for relative comparisons within this collection.
This report analyzes the exact 50-post snapshot shown below. AI identifies editorial categories and drafts explanations; all statistics are calculated from the snapshot, and every narrative claim is checked against cited posts before publication.
Best DeepSeek tweets
Ranked 01–50
@elshayib_ ·
Been playing with DeepSeek V4 flash 0731 all day and that's how much i spent, i don't know what to tell you but i always use subs because API cost a lot but for the first time i get really great performance for almost free and the Cache hits👌. @deepseek_ai good job on this one you really delivered.
@_avichawla ·
Fine-tune DeepSeek-OCR on your own language! (100% local) Most vision models treat documents as massive sequences of tokens, making long-context processing expensive and slow. DeepSeek-OCR uses context optical compression to convert 2D layouts into vision tokens, enabling efficient processing of complex documents. It is a 3B-parameter vision model that achieves 97% precision while using 10x fewer vision tokens than text-based LLMs. In fact, you can easily fine-tune it for your specific use case on a single GPU. I used Unsloth to run this experiment on Persian text and saw an 88.26% improvement in character error rate. ↳ Base model: 149% character error rate (CER) ↳ Fine-tuned model: 60% CER (57% more accurate) ↳ Training time: 60 steps on a single GPU Persian was just the test case. You can swap in your own dataset for any language, document type, or specific domain you're working with. I've shared the complete guide in the next tweet, which includes the code, notebooks, and environment setup ready to run with a single click. Everything is 100% open-source!
@synthwavedd ·
🧵 DeepSeek appear to have engaged, or be engaging in, a large-scale operation to collect outputs from proprietary models (including Claude Fable 5) for certain requests via their API as part of a distillation effort. After seeing such claims circulating earlier today, we https://t.co/wFK4nmuvkZ
@hasantoxr ·
I'm uninstalling Cursor and every other AI coding tool after finding this. It's called DeepSeek TUI. A full coding agent that runs in your terminal on DeepSeek's API. One command to install. No browser. No IDE plugin. No subscription. npm i -g deepseek-tui Here's what it does: → Edits files across your entire codebase → Executes shell commands → Browses the web with web run → Handles git operations → Resumes sessions after you close the terminal → Connects to any MCP server → Spins up as an HTTP API with deepseek-tui serve --http Three modes depending on how much control you want: → Plan: see the full execution plan before the agent touches anything → Agent: interactive multi-step tool use → YOLO: auto-approves every tool call in a trusted workspace DeepSeek's API is already one of the cheapest frontier models available. Now you have a full coding agent on top of it. Built in Rust. 29 releases. Active development. 100% Opensource. MIT License.
@DavidOndrej1 ·
DeepSeek V4 just dropped and it's 40x cheaper than GPT-5.5 Pro 1.6 trillion parameters, trained on limited GPUs, still matches the top labs I built 4 full apps with it to see if the hype is real
@LiorOnAI ·
DeepSeek is back. They just figured out how to make AI models way smarter without adding compute. They built Engram, a memory module that retrieves information instantly with O(1) lookup. 𝗛𝗼𝘄 𝗶𝘁 𝘄𝗼𝗿𝗸𝘀 Think of it as giving your model a lookup table. Instead of burning compute simulating memory retrieval, it just grabs what it needs. Early layers stop wasting effort on pattern matching. The model goes deeper on actual reasoning. 𝗪𝗵𝗮𝘁 𝘆𝗼𝘂 𝗴𝗲𝘁 Huge wins across benchmarks on the same compute: > BBH +5.0 > MMLU +3.4 > HumanEval +3.0 𝗧𝗵𝗲 𝗯𝗶𝗴 𝗱𝗲𝗮𝗹 Long-context retrieval jumped from 84% to 97%. Lookups handle local patterns, so attention focuses on global context.
@kimmonismus ·
DeepSeek is about to release V4, and for the first time, a frontier Chinese AI model will run natively on Huawei silicon. A brief analysis and why its much bigger than most people think. Alibaba, ByteDance, and Tencent have placed bulk orders for hundreds of thousands of Huawei's new Ascend 950PR chips. Prices have jumped 20% in weeks. And DeepSeek deliberately denied NVIDIA early access to V4 while giving that window exclusively to Chinese chipmakers. (via Reuters) Let that satisfy for a moment: a Chinese AI lab actively chose to sideline NVIDIA. What this means for NVIDIA The immediate revenue hit is manageable. China was already a shrinking slice of NVIDIA's business after Washington's export controls and Beijing's counter-ban on the H20. But the *strategic* damage runs deeper. Every model optimized for Huawei chips is a model that no longer needs NVIDIA's ecosystem to function. That's not lost revenue but lost lock-in. NVIDIA's real moat was never just hardware performance. It was CUDA, the software layer that made switching costs prohibitively high. Huawei built the Ascend 950PR to understand the same programming instructions as NVIDIA chips, dramatically lowering those switching costs. The moat is being drained from both sides! What this means for China Let's be precise about what China has and hasn't achieved. The Ascend 950PR delivers roughly 2.8x the compute of NVIDIA's H20, but it still trails the H200. Huawei won't match that tier until the Ascend 960 arrives in 2027. And production is constrained: SMIC can't match TSMC's output, domestic HBM is still ramping, and early Ascend 950PR batches will still rely on imported memory chips. But here's what matters more than the spec sheet: China has closed the loop! It now has a domestic chip that can run a frontier model for inference at commercial scale, with a training chip (Ascend 950DT) due by Q4. Two years ago, that pipeline didn't exist. Washington's export controls were designed to buy time, not to permanently cripple Chinese AI. The theory was that restricting access to cutting-edge chips and lithography tools would slow China by 3–5 years (ASML, highly recommend you read Chris Millers book "Chip War"). What actually happened: China compressed that timeline through *massive* state subsidies, mandatory domestic procurement, and engineering workarounds like DeepSeek's efficiency breakthroughs. The competitive dynamic is shifting from "Can China do AI?" to "Can China do AI at scale on its own silicon?"! This week, that question got a lot closer to a yes. The pressure on NVIDIA isn't that it loses China today. It's that China is building a parallel AI compute stack that doesn't need NVIDIA at all and every model trained or optimized for that stack pulls more of the ecosystem with it. Thats why this is so big news.
@LuizaJarovsky ·
🚨 In a leaked call transcript, we learn what DeepSeek's CEO thinks are China's weaknesses and strengths in the AI race with the United States: Today, a call between DeepSeek's founder and CEO, Liang Wenfeng, and a group of investors was leaked. A few hours after the leak, the WeChat links referring to this call had been removed from the internet. It is still unclear whether the Chinese government ordered it due to political, economic, or regulatory risk, or whether it was a private removal request. My guess is that it was a private request, but the Chinese government is glad it was removed. To my knowledge, similar leaks had never happened before. Remarkably, what DeepSeek’s CEO says is similar to what Anthropic's Dario Amodei said in 2025 and reiterated this year: compute is the most important bottleneck in the AI race between China and the U.S. (According to this view, export controls are probably an effective way to curb Chinese competition in AI). I selected 11 excerpts that reflect DeepSeek’s CEO's thoughts on China's weaknesses, strengths, and its competition with the U.S.: 1. DeepSeek lacks resources 2. The talent gap is due to the compute gap 3. “We’re still at tens of billions—an order of magnitude difference” 4. “Huawei’s production is limited” 5. “About half of our people think OpenAI is better” 6. “NVIDIA is digging its own grave” 7. “We don’t maximize profit or price for maximum revenue” 8. “Our gap with the U.S. might be (...) roughly two years, but using one-twentieth the compute” 9. “I see no upper limit to language model scaling” 10. “Our capital structure can’t support high-cost labeling like that of the U.S.” 11. “In China, the most reasonable approach seems to be to focus on general Agents, with Coding Agent as the top priority” Check out the full excerpts in the link below *AI policy professionals (especially U.S. ones): don't miss this edition! 👉Link below.
@kimmonismus ·
Huge: DeepSeek v4 probably in the next few weeks - and it will be running natively on Huawei's Ascend 950PR Chips DeepSeek is about to drop its next-gen V4 model (via The Information) and for the first time, it'll run natively on Huawei's Ascend 950PR chips, marking a genuine turning point in China's push to break free from Nvidia dependency. And this is big news. Alibaba, ByteDance, and Tencent have already placed bulk orders for hundreds of thousands of units, driving the chip's price up 20% in weeks. DeepSeek even denied Nvidia early access to V4, giving only Chinese chipmakers that privilege. It looks like China has reached a stage with Huawei that now allows for the real-world use of their chips. And at the same time, this will also explain why Deepseek is taking so long to arrive.
@GithubProjects ·
Reasonix is a terminal-based AI coding agent built specifically for DeepSeek, designed to keep token costs low through stable prefix caching across long sessions. - DeepSeek-only, engineered around byte-stable prefix-cache mechanics - 99.82% cache hit rate in a real single-day workload - ~$12 cost instead of ~$61 on the same workload without cache - Top-3 in LLM velocity on Oosmetrics, with active Discord community
@quxiaoyin ·
Why is @deepseek_ai 100x cheaper than @AnthropicAI? China is vertically integrated to be cheap. → Cheap model: token-optimized, aggressive caching, less GPU per query → Cheap chips: Huawei silicon, no Nvidia tax → Cheap energy: subsidized power, state-scale grid → Cheap talent: top researchers at a fraction of US salaries → Cheap economics: DeepSeek funded by trading profits, inference doesn’t need to make money The only thing they lag is performance. But as they get “good enough”, being cheap matters a lot, and being frontier keeps getting harder.
@akshay_pachaar ·
i compared the top open-source OCR solutions and found which one works the best. featuring: > DeepSeek OCR > Datalab Chandra > Qwen3-VL > Dots OCR > Granite Docling also created an app that lets you run all of these OCR models in one place. 100% open-source.
@TheTuringPost ·
DeepSeek-V4 is a full-stack redesign of LLMs around long context + efficiency Here are some of the changes: - Hybrid attention: Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA) for long-context efficiency - 1M-token context becomes ~3–10× cheaper in memory and compute - New residual connections (mHC) + Muon optimizer help to train trillion-scale models without constant collapse - FP4 quantization is used during training (not only inference) - More efficient Mixture-of-Experts model with 1.6T params, but only ~49B active per token is achieved via improved roughing + light sequence-level regularization + hash routing in early layers Post-training recipe is also different - train specialists (math, code, agents) - merge via distillation (OPD) - cleaner and more scalable than RLHF All of this is about making models smarter per FLOP + per token
@DeRonin_ ·
How much do Chinese models cut your API bill? 🇨🇳 WITH THE SAME OUTPUT!!! US frontier model → chinese open model, cost to do one task: [ reasoning brain ] Claude Opus 4.8 → DeepSeek V4 Pro $1.80 → $0.04 per task ~45x cheaper, and it's tied with Opus on SWE-bench (80.6 vs 80.8) [ code generation ] GPT-5.5 → GLM 5.2 $1.03 → $0.48 per task ~2x cheaper AND it beats GPT-5.5 on long-horizon coding [ agents + tool calling ] Claude Sonnet 5 → Kimi K2.7 $2.29 → ~$0.31 per task ~7x cheaper, best agentic stability of the open models [ bulk / high volume ] Gemini 3.5 Flash → DeepSeek V4 Flash $0.59 → $0.02 per task ~30x cheaper, plenty for extraction, classification, drafts the tradeoff: > coding + agents.. near parity or better > hardest reasoning + world knowledge.. ~10-15% behind > your bill.. down 2-45x depending on the task one catch: Qwen3.7 Max is chinese but priced like frontier ($1.06/task).. cheap isn't automatic route each task to the model that wins it: code → GLM.. agents → Kimi.. value reasoning → DeepSeek.. bulk → DeepSeek Flash keep a US frontier model for the 10% that truly needs it
@DavidOndrej1 ·
After 14 months it's finally here... Deepseek v4 It's cheap and goes head to head with GPT and Opus. In this 29 min video i'll show you everything you need to know about it: timestamps: 00:00 Introduction: DeepSeek V4 Arrives 00:49 Specs: 1.6 Trillion Parameters & Architecture 01:55 The GPU Story: China's Hardware Limitations 03:09 Benchmarks vs GPT-5.5 and Claude Opus 4.7 04:39 Agentic Coding & Terminal Performance 06:29 Pricing: 40x Cheaper than GPT-5.5 Pro 09:05 5 Major Weaknesses of DeepSeek V4 10:43 Censorship & Tool-Calling Constraints 11:57 Setup: Using DeepSeek V4 via OpenCode 14:33 Build 1: Interactive Architecture Explainer 16:11 Build 2: SVG Plant Animation Demo 17:02 Build 3: Full Arcade Karting Game 18:10 Build 4: Exoplanet Data Visualization 25:24 Final Results & Performance Review 28:50 Cost Analysis: Parallel Agents for Pennies
@RoundtableSpace ·
DeepSeek just open-sourced a full coding agent harness for free that directly replaces Claude Code's $200 a month plan. → Everything is a plugin, model adapters, tools, session logs and the agent loop itself, all swappable from config → Built on Cordis, DeepSeek's own plugin runtime → One command to spin up a local Web UI: npx @deepseek-ai/dsh web → Ships with real architecture docs and an AGENTS.md written for agents to read the codebase themselves → TypeScript, MIT, developer preview with breaking changes expected https://t.co/OYDwiQF6G2
@DeRonin_ ·
DeepSeek just dropped a 5-page paper + free GitHub repo that makes any LLM respond 80% faster it's called speculative decoding. in plain english: Guess → Check → Keep → Repeat > Guess: a small fast model predicts the next few words > Check: the big smart model checks all guesses at once > Keep: lock in the right ones, fix the first wrong one > Repeat: do it again, 2-4x faster than normal DeepSeek's version (DSpark) is the first to combine two approaches that always had a tradeoff: getting both speed AND accuracy at the same time production results on their flagship model: 50% more requests handled, responses up to 80% faster how you can use it: > just using AI apps: expect them to get noticeably faster soon > building with AI: integrate this into your stack for instant 2x speedups > running your own models: clone the repo today and apply it to whatever model you run every other AI lab has been guarding this. now it's free on GitHub, MIT licensed this completely changed how I'm thinking about AI products this week one of the most powerful releases this month repo + papers below ↓
@techNmak ·
125K+ GitHub stars in 3 days. DeepSeek Harness. It's an open-source agent harness. "Harness" here is simply the software around a model that lets it do work => give it access to tools, files and a shell, maintain sessions, apply permissions and sandboxing, and handle the loop between model responses and tool results. DeepSeek Harness is built on Cordis, and almost every major part of the system is a plugin. The model adapter, tool registry, session log and even the default agent loop are all replaceable pieces. A running agent is assembled through profiles and bundles, so you can change the model provider, add tools or introduce new policies without rewriting the rest of the runtime. Sessions are handled as an append-only event log. User messages, model output, tool calls and tool results are recorded as events, and the history shown back to the model is derived from that log. Forking, resume, transcripts, telemetry and persistence also build from the same event stream. Code Mode lets Harness expose tools through a code runtime. The model can write a small program that calls several tools, uses loops or runs independent operations concurrently. Those calls still pass through the normal Harness tool-execution pipeline and its policies. There is a lot more in the repo => subagents, workflows, MCP tools, skills, approvals, persistent terminals, background jobs, LSP-based code navigation, context compaction, goals, scheduled follow-ups and a Web UI. Not every configuration enables all of these, profiles decide which capabilities are assembled into a particular Harness setup. GitHub Repo: https://t.co/oj58ZcJ5bJ
@MrAhmadAwais ·
what a day. we broke 100 billion tokens! Command Code was 3rd & 6th most used coding agent¹ this week and probably the most used coding harness in the world for deepseek v4 models. i can literally load complete docs and full dependency's code before writing a single line with 1M context on deepseek. flash is ~100x cheaper than the alternatives, and at this context size it holds up shockingly well. also, our $1/mo Go plan has made ai coding with open models accessible everywhere in the world, esp outside the bubble. a billion tokens of deepseek flash for $1 is a bit unbelievable too. ¹ ranked against agents on openrouter. p.s. we're not on openrouter, we use 12 other providers.
@mark_k ·
A new Stanford HAI analysis of DeepSeek's research team reveals just how quickly China's AI talent base is maturing. China now has a large, rapidly improving and increasingly self-sufficient frontier AI talent pipeline. Anyone still assuming the US has an unassailable lead is asleep at the wheel. DeepSeek's author pool grew from 223 to 356 researchers in a single year, while their median citation count more than doubled. Over half built their entire careers inside China, and even among those who gained experience abroad, most ultimately returned. DeepSeek isn't a fluke. It's the clearest sign yet that China can now build frontier AI teams at scale.
@mhdfaran ·
DeepSeek is coming for Claude Code. They just released DeepSeek Harness, an open-source agent system that can actually work inside your codebase. It can read files, edit code, run commands, plan tasks, and delegate work. But the wild part is the architecture: Everything is a plugin. Models, tools, agent loops, sandboxes, even the session system can be replaced. Fully open source. MIT licensed.
@VaibhavSisinty ·
DeepSeek just made their AI 5x faster. Without changing the model at all. It is called DSpark. And it is one of the most impressive pieces of system engineering I have come across. Here is the thing. The biggest bottleneck in running an AI model is not compute. It is memory. Reading the model weights from memory takes so long that generating 10 tokens in parallel costs almost the same as generating one. DSpark exploits this completely. Instead of generating one token at a time it drafts multiple tokens using a cheap smaller model first. Then verifies all of them through the main model in one pass. If the draft is right it skips ahead. If it is wrong it falls back. Most tokens are easy to predict. So most of the time it skips ahead. But what makes DSpark genuinely different is how it combines two approaches that nobody had successfully merged before. DFlash is fast at early token positions. Eagle is more accurate at longer sequences. DSpark runs both together and beats either approach on its own. And it does not draft a fixed number of tokens. It drafts more when the GPU is free and less when the server is busy. All on GPU. No CPU slowdowns. To be frank, this is not a model breakthrough. It is an infrastructure breakthrough. DeepSeek keeps finding ways to do more with less. That is the real moat.
@GithubProjects ·
DwarfStar 4 is a standalone inference engine built specifically for DeepSeek V4 Flash, prioritizing speed and local execution on Metal and CUDA. - Supports Metal, NVIDIA CUDA, and AMD ROCm backends. - KV cache treated as a first-class disk citizen for long context. - Optimized for 2-bit quantization on 96GB+ MacBooks. - Includes server API, tool calling, and integrated coding agent.
@AIFrontliner ·
Breaking: China’s DeepSeek just released a model with 1.6 trillion parameters that runs on 10% of the memory of its predecessor. And buried inside the technical report is something nobody is talking about. The model gets 10x more efficient at 1 million tokens. Not slightly more efficient. 10x. Here's why that terrifies AI safety researchers. Right now, oversight of AI systems depends on humans being able to monitor what the model is doing. That requires reading context. Reviewing reasoning. Checking outputs. When the context window was 128K tokens, that was already impossible to fully review. At 1 million tokens, the model can hold the equivalent of 10 full novels inside a single session. Every decision it makes draws on context you will never fully read. DeepSeek-V4-Pro requires only 27% of the computing cost of its predecessor at this scale. Which means running a million-token session is now cheap enough to do routinely. The paper says exactly this. They call it making "long-horizon tasks more feasible." Long-horizon tasks are tasks that span weeks, months, or years of work. Agentic workflows that make hundreds of sequential decisions. Autonomous systems that don't pause and ask for permission between steps. They trained it specifically for this. Domain-specialist expert models for math, coding, agent behavior, and instruction following. Then merged all of them into one system through a process they call On-Policy Distillation. The resulting model sits at a 3206 Codeforces rating. That is inside the top 25 human competitors on the planet. It writes code better than almost every human alive. It does this autonomously. At a million-token context. For almost nothing. The benchmark charts in the paper show DeepSeek-V4-Pro-Max beating or matching GPT-5.4 and Gemini-3.1-Pro across knowledge, reasoning, and agentic tasks. This is not a research model. They are running it in production. The DeepSeek chatbot already uses it. The safety conversation in AI right now is almost entirely focused on alignment, on whether models want bad things. This paper is about something different. It is about capability. A model that is cheap, fast, powerful, and autonomous enough to run for hours without a human in the loop. The oversight problem does not require a misaligned AI to become serious. It just requires one capable enough that you can no longer follow what it is doing. We crossed that line this week.
@rohanpaul_ai ·
DeepSeek Harness reached 122K+ GitHub stars in 3 days, one of the fastest to rise in GitHub's history. It makes models, tools, loops, storage, scheduling and even the UI swappable plugins. its an orchestration layer for coding agent, runs as a local web app on a configurable port, and exposes every component as a swappable plug-in: shell access, file editing, web search, skills, sessions, model choice, reasoning effort, and permission scope, all editable in a YAML config. sub-agent support is the really special part. You can wire Claude Code or Codex in as plug-ins and let Harness route subtasks to whichever agent suits each step. It's MIT licensed, so you can add other providers or point it at a self-hosted model and never touch DeepSeek's API. However, their release timing for this is strange. They raised the price of cache-hit input by 6-fold to 12-fold in the same week it shipped an agent framework, and cached input is exactly what agent loops consume, since every step replays the same system prompt and accumulated history. Whoever set the API pricing either didn't coordinate with the harness group or deliberately priced the new workload higher. My read is the 2nd, because caching subsidies made sense when agents were rare and become the largest unpriced cost once a framework makes them routine.
@quxiaoyin ·
@AnthropicAI’s business model just got disrupted by @deepseek_ai v4 pricing. Anthropic’s all about subscription locked-in: – Claude Code Max: $200/mo – Same usage via API: $5,000 – You still get rate limited; New limits added every few hours (not daily, not weekly — hours). If exceeding you will pay api pricing which can cost $100+ an hour. – Caps on parallel threads - can only use max in max not Hermes/openclaw – SDK requires the API, so devs eat the $5K DeepSeek: - do whatever. No subscription. 100x cheaper. No rate limit. No parallel limit. No usage limit. Whatever. The problem is, can Claude catch up? Yes Claude’s smarter still but most coding tasks DeepSeek is enough. Then what?
@TheGeorgePu ·
I ran a blind test. DeepSeek V4 Pro vs Claude Opus 4.8. Same prompts, no labels, separate agents judging. Python. Writing in my voice. A bug-fix with trick questions. 4 out of 4, DeepSeek won. And I'm still choosing Claude. Capable isn't the same as the one you want to keep using.
@The_AI_Investor ·
DeepSeek did it again. The new V4 Flash is not really a small model. It has 284B total parameters, but only activates 13B for each token. That makes it much cheaper and faster to run than its total size suggests, while still supporting a 1M-token context window. It is not beating the best frontier models overall, but it is getting surprisingly close. On DeepSeek’s published agent benchmarks, V4 Flash scored 82.7 versus Claude Opus 4.8’s 85.0 on Terminal Bench, and 25.2 versus 25.7 on Agents’ Last Exam. It still trails Opus more clearly on harder repository and cybersecurity tasks. Some results also use DeepSeek’s own agent harness and internal datasets, so comparisons should be treated carefully. The crazy part is the price: just $0.14 per million input tokens and $0.28 per million output tokens. For comparison, GPT 5.6 Sol costs $5 and $30, while Claude Opus 4.8 costs $5 and $25. That makes V4 Flash around 36 times cheaper on input and roughly 90 to 107 times cheaper on output. It is still too large for a normal laptop. DeepSeek’s official serving example uses four GB300 GPUs. But for AI clouds and large-scale agent workloads, this level of performance at this price could put even more pressure on frontier-model pricing.
@aakashgupta ·
DeepSeek is raising its first-ever outside round at a $10 billion valuation. OpenAI just closed at $852 billion. Anthropic at $380 billion. The Chinese lab that erased $589 billion from Nvidia's market cap in a single day last January is being valued at roughly 1.2% of the American company whose capex thesis they empirically disproved. For two years, DeepSeek's whole pitch was that they didn't need outside money. Liang Wenfeng spun the lab out of High-Flyer, his $8 billion quant hedge fund, in July 2023. VCs wouldn't touch it back then because they couldn't see the exit. Liang funded everything through the hedge fund's trading profits. He told interviewers that capital was never China's AI constraint. Confidence and talent organization were. Then on January 20, 2025, they released R1. Seven days later, Nvidia lost $589 billion in one session. Largest single-day market cap loss in US stock market history, nearly double the previous record, which was also Nvidia. They did this with a $5.6 million training run on 2,048 H800 GPUs, the export-compliant chips Nvidia built specifically for the Chinese market under Biden-era restrictions. The ratio is the story. The damage DeepSeek did to Nvidia in one trading session is 58 times their entire new valuation. Anthropic just turned down offers at $800 billion. OpenAI raised $122 billion in February. DeepSeek is accepting $300 million at ten. The gap reflects capital access more than capability. US venture dollars can't easily flow into a Hangzhou-based AI lab under current CFIUS and export control friction. Chinese VC deploys a fraction of US VC in absolute capital. The valuation ceiling here is geographic before it gets to the product. Which raises the real question: why raise at all? High-Flyer manages $8 billion in AUM. The hedge fund paid for V3 and R1 through its own trading book. Two pressures converge. R2 has been delayed since May 2025 on chip stability issues after Beijing pushed DeepSeek toward Huawei Ascends for training. Frontier-scale training under export controls costs more than the hedge fund can responsibly commit. And a formal cap table with aligned Chinese capital becomes structurally necessary as domestic compute mandates intensify. The market raised $122 billion for OpenAI three months ago on the thesis that frontier AI needs infinite capex. DeepSeek proved that thesis wrong fifteen months earlier. The capital never updated. It still hasn't. The lab that broke the model is raising at 1% of the lab it broke.
@pankajkumar_dev ·
DeepSeek-V4 Drops: Open-Source Push Toward Cheaper, Long-Context AI - DeepSeek-V4-Pro is a 1.6T MoE model (49B active) trained on 33T tokens and released under a permissive MIT license - Efficiency gains: supports 1M-token context with 9.8× lower FLOPs and 9.5–13.7× smaller KV cache vs V3.2 - Benchmarks (vs GPT-5.4 & others): APEX (90.2% vs 85.9%), SWE Verified (80.6% on par with top models), Terminal Bench 2.0 (67.9% vs 65.4%), Codeforces (3060 rating, near top human level) - Head-to-head vs Anthropic Claude Opus 4.6: overall 53% win vs 37% loss (10% tie), stronger in task completion (98.3 vs 96.7) and content quality (83.3 vs 78.0) - Uses hybrid Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA) instead of quadratic attention - Enables long-context inference on smaller GPU clusters that previously couldn’t handle it - Variants: Pro (1.6T) and Flash (284B) with low-cost API pricing - Overall direction points toward efficient, open, agent-ready models rather than just scaling brute compute DeepSeek-V4 mainly stands out for making long context + agent workflows cheaper and more accessible without massive infrastructure.
@smratitiwa86867 ·
7 days. 33,000+ stars. 6,000+ stars in a single day. DeepSeek-TUI just exploded onto GitHub’s global trending page — and people are calling it the open-source Claude Code alternative for DeepSeek V4. 👀 This thing turns your terminal into a full AI coding agent: • Rust single-binary setup — no dependencies, just run • Cross-platform with full Chinese UI support • Understands entire repositories, not just single files • Real-time AI reasoning visualization while it works • Read-only / approval / YOLO autonomous modes • Built-in Shell, Git, file ops, and web search • Session recovery, rollback snapshots, token cost tracking The wild part? You stop “using AI” and start supervising it. One command and the agent takes over the workspace: editing files, debugging code, running commands, navigating the repo — all from the terminal without endless copy-paste chaos. And with the current DeepSeek V4-Pro discount, the cost-to-power ratio is honestly absurd. Tried it yesterday and watched the AI build half a project while I just sat there questioning my career choices. Repo👇
@RoundtableSpace ·
DeepSeek V4 Flash scores 52 on the AI Intelligence Index for $72 to run the full test suite. GPT-5.6, Opus 5 and Fable 5 score the same and cost 10x more.
@_vmlops ·
DEEPSEEK JUST OPEN SOURCED AN AGENT HARNESS deepseek-harness (dsh) everything is a plugin ▫️ built on cordis, a runtime for spatiotemporal composability ▫️ ships a web ui, run instantly via npx @deepseek-ai/dsh web or build from source with pnpm ▫️ still in developer preview breaking changes expected open source agent tooling from deepseek keeps shipping https://t.co/ycbEoFw8UI
@AISystemGuy ·
Engineering update 🛠️ We’re adding native support for DeepSeek’s Multi-Head Latent Attention (MLA) and DeepSeek Sparse Attention (DSA) directly into our LLM inference runtime. The goal: break through the memory-bandwidth wall. A technical deep dive 🧵 1/ MLA: compressing the KV cache MQA and GQA reduce KV-cache size, but at the cost of representation capacity. MLA takes a different approach: Keys and Values are projected into a shared, low-dimensional latent space, and that compressed representation is cached instead of the full KV matrices. In simplified terms: d_c ≪ d_h × n_h This can dramatically reduce long-context memory pressure. 2/ The RoPE challenge Rotary Positional Embeddings are position-dependent, so they cannot be fully baked into the compressed latent vector. We split Keys and Queries into: • A compressed latent component for semantic context • A smaller position-aware component for RoPE Our CUDA kernels then combine these components mid-kernel. 3/ Reordering the math The updated kernels perform key matrix multiplications out of order, reconstructing only what is needed during attention. This avoids unnecessary intermediate tensors and extra memory transactions—critical when decode performance is limited by memory bandwidth. 4/ Adding DSA routing We’re layering DeepSeek Sparse Attention on top of the latent cache. A lightweight “lightning indexer” transforms and quantizes each incoming query, then scores it against the historical token cache. 5/ Dynamic top-K attention The runtime scheduler uses those scores to select the most relevant attention blocks. Low-affinity token connections are masked out, while only the top-K paths are computed. The result is a much smaller runtime compute graph. 6/ Why this matters MLA reduces the size of the data we need to move. DSA reduces how much of that data we need to process. Together, they help shift long-context inference away from being memory-bandwidth-bound and back toward a more compute-efficient execution profile. Early kernel tests are showing: • Major reductions in activation-memory overhead • Significant token-throughput gains • Stronger scaling on long-context sequences Benchmarks coming soon.
@JulianGoldieSEO ·
DEEPSEEK JUST OPEN-SOURCED THE LAYER THAT ACTUALLY MAKES AI AGENTS USEFUL And the biggest feature isn’t the model. It’s what you can replace. What DeepSeek Harness changes: → The model is a plugin → Memory is a plugin → Tools and skills are plugins → The sandbox and file system are plugins → Even the think → act → check → repeat agent loop can be swapped Why that matters: ✓ MCP standardized how agents connect to tools. Harness takes that idea across the entire agent stack. ✓ It reads AGENTS.md and CLAUDE.md, so it can inherit instructions already used by Claude Code and other agents. ✓ It runs locally through a browser UI after a terminal command instead of forcing everything through a terminal interface. ✓ It launched under the MIT license, so developers can modify it or build products on top. The important lesson: Stop collecting AI tools. Build one agent system where Claude Code, Hermes, OpenClaw, DeepSeek Harness and whatever ships next can be swapped in without rebuilding everything. That’s a much better way to survive how fast this space is moving.
@JulianGoldieSEO ·
DEEPSEEK V4 FLASH JUST BEAT ITS OWN PRO MODEL. I ran 50+ builds side by side—and the difference was impossible to ignore. What Flash 0731 did better: → Built a working 3D flight simulator while V4 Pro completely failed → Produced smoother controls, stronger graphics and cleaner game UI → Created usable coding projects despite being designed as a fast agent model Where it still failed: ✓ One build let you see through the floor ✓ The pool game was genuinely terrible ✓ Complex physics and controls were still inconsistent The bigger opportunity: ✔ Run the latest version free through OpenCode ✔ Connect it to Hermes using the OpenCode skill ✔ Delegate coding tasks automatically inside an agent workflow DeepSeek V4 Flash is not better than Fable 5 or GPT-5.6 Sol. But for a free, fast and lightweight agent coder, it can already outperform DeepSeek V4 Pro on real builds.
Best Tweets by Topic