The news, 365 days behind — on purpose Delayed live · replaying 2025

One Year Ago.AI

Remember how fast this is.

21AUG2025replayed
one year on
open weightsDeepSeek

DeepSeek unveils V3.1 with hybrid reasoning mode, hints at domestic chip push

The Chinese AI lab releases a unified model that toggles between fast answers and chain-of-thought reasoning, and says its new FP8 format anticipates the next generation of homegrown silicon.

DeepSeek today released DeepSeek-V3.1, its first major update to the V3 family, touted as ‘our first step toward the agent era.’ The headline change is a hybrid inference model that lets users toggle between ‘Think’ mode, a chain-of-thought reasoning successor to R1, and ‘Non-Think’ mode for fast direct answers, and a 128K context window. It ships as open weights on Hugging Face.

The company says Think mode now reaches answers in less time than the earlier DeepSeek-R1-0528, and that post-training brings sharp improvements in tool use and multi-step agent tasks. DeepSeek also consolidated its separate chat and reasoning API endpoints into one, with the chat template toggling between modes.

A technical detail drawing attention is DeepSeek’s use of a UE8M0 FP8 scale format. In a WeChat comment, the company said the format is ‘designed for the next generation of domestically produced chips to be released soon,’ read by observers as Chinese-made AI accelerators — a hint of reduced reliance on Nvidia. The company says the format is ‘designed for the next generation of domestically produced chips to be released soon,’ in a WeChat comment, and the weights are available on Hugging Face.

The timing is notable as The Register reports DeepSeek had attempted to train its next-gen R2 model on Huawei’s Ascend accelerators but faced difficulties, resorting to Nvidia H20s. The V3.1 weights are now being evaluated for inference on Huawei hardware. Alibaba abandoned a similar unified reasoning/chat approach after finding it degraded quality, per The Register; at least in benchmarking, V3.1 appears to have avoided that problem. DeepSeek says it scores 30 on BrowseComp, versus 8.9 for the May R1 refresh.

T
The Register

Reported that DeepSeek trained V3.1 using the UE8M0 datatype for anticipated domestic chips, and noted the switch is more about compatibility than efficiency.

One year later — open only if you can handle spoilers

The hybrid "one model, two modes" design proved prescient — folding reasoning and chat into a single switchable model became a common pattern over the following year. DeepSeek iterated fast (a V3.1-Terminus refresh, then V3.2 with sparse attention), but its long-awaited next-generation model kept slipping, reportedly because training runs on Huawei's Ascend chips kept failing and the team fell back to Nvidia — an awkward footnote to the "domestic silicon" signal. By 2026 DeepSeek was running inference on Huawei hardware, but full training independence from Nvidia remained unfinished business.

Replay thisPost on XRedditHNLinkedIn

The Weekly Replay · free by email

This week, one year ago — every Sunday.

One email each Sunday: the week's replayed AI news, with the one-year-later annotations included. Written like it's breaking — dated like it isn't.

Free · double opt-in · unsubscribe anytime · privacy