
DeepSeek
DeepSeek is a Chinese AI company that develops and releases language models that are similarly capable as the best models from the US, but with significantly lower computational effort. The release of the DeepSeek-R1 model in early 2025 sparked a worldwide debate about technological leadership and costs in AI development.
DeepSeek is a Chinese company that develops computer programs capable of understanding and generating text — so-called language models. It was founded in 2023 by Liang Wenfeng, co-founder of the hedge fund High-Flyer. Unlike the best-known competitors from the US, such as OpenAI or Google, DeepSeek releases its models as open source: anyone can download the program code and run it themselves. The company became especially well known through its model DeepSeek-R1, which in early 2025 matched the then-best models in the world in benchmarks — that is, standardized performance tests. The attention was considerable because DeepSeek claimed to have achieved this result with a fraction of the usual development costs.
DeepSeek’s significance for AI competition
Until 2025, it was considered a given that better AI models automatically required more money and more computing power. Large companies such as Microsoft or Google invested billions in special chips, so-called GPUs, that enable the training of these models. DeepSeek shook this assumption.
The company stated that it had trained DeepSeek-V3 for around six million dollars — GPT-4 from OpenAI is said to have cost a hundred times as much. Whether this figure reflects the full expense is disputed among experts. Nevertheless, DeepSeek showed that clever algorithms can offset part of the cost advantage gained through sheer computing power. This has political consequences: the US had banned the export of the most advanced AI chips to China in order to widen China’s technological lag. DeepSeek suggested that this strategy might not be as effective as hoped.
When DeepSeek-R1 was released, the share price of chip manufacturer Nvidia fell by almost 17 percent in a single day — one of the largest single-day losses in the company’s history. Investors asked themselves whether the world really needs so many expensive chips if cheaper alternatives deliver similar results.
Technical tricks behind the efficiency leap
DeepSeek relies on a combination of approaches that were each already known individually, but rarely brought together so consistently. A central element is the mixture-of-experts architecture: the model activates only a small portion of its internal building blocks per request, instead of always straining the entire system. This saves computing time without reducing overall capability.
When training DeepSeek-R1, heavy use was made of a method called reinforcement learning. In this approach, the model is not given a ready-made solution scheme; instead, it receives feedback on whether its answer is right or wrong and gradually learns to find better solution paths. This approach requires less computing power than classic training on huge amounts of text, because the model is specifically optimized for particular tasks — such as mathematics or programming.
In addition, DeepSeek published its weights — that is, the learned numerical values within the model. This allows the research community to study the architecture closely and develop their own variants. Many of DeepSeek’s technical reports are considered unusually detailed and transparent for a commercial company.
DeepSeek in products and headlines
DeepSeek operates its own chat application at chat.deepseek.com, which briefly became the most downloaded app in the US App Store after the launch of R1 — ahead of ChatGPT. Anyone who does not want to run their own infrastructure can also integrate the models into their own applications via DeepSeek’s programming interface (API), similar to OpenAI.
Because the models are available as open source, they quickly appear in other products. Several cloud providers, including Microsoft Azure and Amazon Web Services, began offering DeepSeek models within their own services shortly after the release. Smaller startups also use the models as a basis for specialized applications.
In news articles, DeepSeek mainly appears in two contexts: as an example of the catching-up Chinese AI sector, and as an argument in the debate over whether Western export controls on chips are actually effective. Data protection authorities in Europe have also begun examining how DeepSeek handles user data — a topic that regularly comes up with Chinese technology products.