
DeepSeek's V4.1-Flash beats its own V4-Pro flagship while shrinking KV cache 4x
DeepSeek released V4.1-Flash, a 552B-parameter MoE under an MIT license that activates 8B parameters for input and 16B for generation, and scores 74.2 on DeepSWE v1.1 against V4-Flash's 54.4. The real signal is economics — its KV cache runs 4x smaller than V4-Flash and 437x smaller than V1, cutting HBM to a quarter — and DeepSeek is retiring V4-Pro by routing that traffic to the cheaper Flash model on Sept 14. It still trails V4-Pro on knowledge benchmarks like SimpleQA-Verified, 42.3 to 55.2, so the trade is cost and speed over deep recall.
deepseek.com →- 02
Sakana AI teams with Sumitomo and SCSK to push domestic AI into Japanese industrySakana AI, backed by Japanese government funding under the Generative AI Accelerator Challenge, signed a comprehensive partnership with trading house Sumitomo and IT integrator SCSK to deploy its domestic models across finance, manufacturing, cybersecurity, and social infrastructure. The deal hands a homegrown model builder a national distribution channel — Sumitomo's customer network plus SCSK's systems-integration and governance work — a bet that Japanese enterprises will favor domestically developed AI over foreign frontier labs.
mlex.com → - 03
OpenAI ships full-duplex GPT-Live-1 voice API at $0.05 per minuteOpenAI made GPT-Live-1 generally available in its API at $0.05 per minute — a full-duplex model that listens and speaks at once, handles interruptions, and routes deeper reasoning to a backend model the developer picks. On OpenAI's benchmarks turn-taking latency drops to 0.8 seconds from 1.4 and tool-calling accuracy rises to 87% from 60%, so a developer can build phone agents that field interruptions the way people do — the pattern Yelp runs for reservations.
the-decoder.com → - 04
TSMC posts record NT$515B August revenue, up 53% on AI demandTSMC posted record August revenue of NT$514.8 billion (US$16.32 billion), up 53.3% year over year and its first month above NT$500 billion, driven by AI orders and new smartphone launches. The signal is structural, not a spike — TSMC raised 2026 capital spending to US$60-64 billion and guided to slightly over 40% revenue growth, committing frontier capacity years ahead against sustained AI demand.
focustaiwan.tw → - 05
Harvey raises $550M at a $15.5B valuation, nearly doubling in nine monthsHarvey raised $550 million at a $15.5 billion valuation co-led by Diffusion and Lightspeed, its eighth-plus priced round and a near-doubling of value in about nine months. Paired with Tenet — its first in-house model, built on open-weight Kimi K3 and post-trained on legal data with Fireworks — plus a push for clients to run their own open-weight models, Harvey is proof that a whole industry can go deep on AI without depending on OpenAI or Anthropic.
techcrunch.com → - 06
Apple ships its first foldable, the iPhone Duo, starting at $1,999Apple introduced the iPhone Duo, its first foldable and the first flagship under new CEO John Ternus, pairing a 7.6-inch inner display with a 5.4-inch cover screen, the A20 Pro chip, and Touch ID in place of Face ID. Apple enters a category it long ceded to Samsung and Xiaomi, where foldables are under 2% of smartphone shipments today, yet Counterpoint projects the Duo could take up to 25% of the foldable market by year-end.
techcrunch.com → - 07
Uber, AT&T, and Pinterest cut AI costs up to 92% by moving to open modelsUber cut cost per AI request 34% and per session 52%, AT&T saved 56% routing through LiteLLM for a 2% quality drop, and Pinterest now runs at under 8% of closed-model cost by post-training open models on its own data. Open models on inference services run 2-20x cheaper than frontier ones, so a developer can route simpler tasks like code summaries and subagent work to open weights and reserve frontier models for complex generation.
newsletter.pragmaticengineer.com →