← narwal.one/Second Brain
SecondBrain
Ask the Brain
Index/Entityupdated Sat Aug 08 2026 08:00:00 GMT+0800 (Philippine Standard Time)

DeepSeek

companyai-labchina-aiopen-weightdeepseekmodel

DeepSeek

Chinese AI lab whose January 2025 R1 release created the original "DeepSeek moment" — reasoning models experienced at scale for the first time, wiping ~$1trn from US capital markets (Nvidia −17% intraday, Nasdaq −3.1%). Since then the benchmark every Chinese release is measured against; per The Big Ways AI Just Changed (AI Daily Brief), only GLM 5.2 (June 2026) has legitimately earned the label since.

Recurring vault datapoints (2026)

  • Per-token pricing anchor: DeepSeek v4 charges $0.87 per 1M output tokens vs Fable 5's $50 — ~57× cheaper per token (China Is Having Another AI Moment (Economist)).
  • The total-cost caveat: Du Zheng (Georgia Tech) et al., June 2026 — DeepSeek used 23× more tokens than an OpenAI rival for basically the same result. Per-token is the wrong denominator; see Token Scarcity.
  • The "good enough" hedge: v4 Pro at ~¾ of Fable 5 performance for <1/60 the cost (What Britain Needs to Do to Grasp Its Big AI Opportunities (Economist)) — the cleanest one-line case that frontier dependence isn't required for most enterprise work.
  • Meta-agent capability signal: v4 Pro was the only open-weight model to cross a human-engineered scaffold baseline in the Meta-Agent Challenge (Autonomous Agent Development Benchmark).
  • Enterprise switching signal: Microsoft reportedly considering DeepSeek for Copilot (Donald Trumps AI Regime Is Opaque Unpredictable and Unsustainable (Economist)) — US guardrail-tightening and access whiplash creating demand for Chinese open-weight models.
  • Like Zhipu (Z.ai), compensates for chip export controls via heavy post-training (fine-tuning + RLHF + allegedly distillation of American systems).

2026-07-11 updates

  • $7bn June-2026 raise — the 4th-biggest VC round ever in China per Beware the Top-Heavy Economy (Economist), the same week SpaceX raised $86bn and announced the Cursor deal.
  • Huawei-silicon-tuned LLM (April 2026) — DeepSeek released a large language model tuned specifically to Huawei's silicon, per Chinas Semiconductor Industry Is Racing to Catch the West (Economist). This is the HW/SW co-design workaround: given a sub-frontier chip, ship a model built for it rather than porting frontier models onto it.

Cross-references

  • GLM 5.2 · Zhipu (Z.ai) — the successor "moment" and its lab
  • Token Scarcity — the 23× token-overuse caveat lives here
  • Hierarchy of Access · AI Licensing Regime (US) — the policy regime DeepSeek's open weights route around
  • Model Routing — DeepSeek-class models as the routed lower tier
  • Huawei — the silicon partner for the April 2026 tuned LLM
  • Beware the Top-Heavy Economy (Economist) — the $7bn raise datapoint
  • Americas AI Labs Are Under Threat from Cheap Chinese Rivals (Economist) — 2026-07-25; DeepSeek cheapest in the Artificial Analysis basket at $0.04 avg per job vs Fable $2.75; DeepSeek v4 flash at $0.28/M tokens direct, $0.18/M via cheapest third-party. Confirms the "compressed margin, downstream distribution" thesis.
  • Chinas Mysterious New Billionaires Are Conquering the World (Economist) — 2026-07-25; founder Liang Wenfeng personal wealth ~$38bn as DeepSeek closes a $71bn valuation round; Liang wealthier than Amodei or Altman. On-record: "the human brain can concentrate only for 6-8 hours a day, and overwork leads to mistakes" — rejection of the 996 register.

2026-08-08 — V4-Flash-0731 GA and the Reddit reception

Per DeepSeek-V4-Flash-0731 Reddit Reception Research (2026-08-08) (reddit-research MCP corpus: 684 comments, 10 threads, r/LocalLLaMA + r/LocalLLM + r/LLMDevs), the V4-Flash GA release (2026-07-31) is the vault's first empirical case of a Chinese open-weight release triggering visible self-argument inside the pro-open-weight community itself.

  • Model identity: 284B MoE, ~13B active, natively 4-bit QAT, ~160–167GB full weights, 1M context. Preview → GA lift was RL post-training compressing turns-to-solve, not raw-ceiling change (u/tarpdetarp: "it was almost as capable as the big boys already, but it took so many more turns to get there").
  • Artificial Analysis intelligence index ~40 → 50 at unchanged cost — one point below GLM-5.2 and GPT-5.6 Luna. u/joorklee: "models available to run locally on <8K USD hardware has nearly the same intelligence score as the top frontier models 5 months ago."
  • Pricing that broke the self-hosting cost case: $0.14 input / $0.28 output per 1M tokens, 98% cache-hit discount (vs Anthropic Luna's $0.20/$1.20 with 90%). Real-world cache-hit rates 90–99.6% (u/thecstep: 99.6% on 10B+ tokens). See Cost-Structure Inversion of Local Inference for the synthesis this pricing forced.
  • DeepSWE parity claim with Sonnet 5 / Grok 4.5 is unverified — DeepSeek's own number, flagged by the OP as "DeepSeek claims, not verified by DeepSWE yet." No independent replication in the corpus.
  • Political-behaviour finding: the raw weights follow the party line on sensitive topics (u/Lanky_Lynx2166's Tiananmen / Cultural Revolution / Canberra experiment on dual RTX PRO 6000 vLLM); a competent system prompt (pi harness) overrides it. See Character Lives in the Harness (Not the Weights) for the general concept this datapoint anchors.
  • V4-Pro GA still unreleased as of the 2026-08-08 corpus capture; community attention has since shifted to Qwen 3.8-Max (reported #1 on AA agentic index ahead of Opus 5). DeepSeek's release cadence is now competing with same-cluster Chinese labs, not primarily with the US frontier.