US and China Startups Battle for the Best Open-Source Models
- • Reflection AI releases Beam, a powerful open-source model
- • Meta’s Agent Muse creates secret profiles on the user’s friends and family
- • American AI models are now only 3 percent better than their Chinese counterparts
US Startup Releases Open-Source Model, Vows to Top Chinese Models
On Monday, October 5, 2026, Reflection AI introduced Beam, a text-only model with 501 billion parameters, of which 23 billion are active per token. The model is initially only accessible via an early access program with a waitlist and is, according to the company, in the final red-teaming phase. The weights are scheduled to be released 'later this month' under an Apache 2.0 license, along with documentation and fine-tuning tools. No one can download the model at this time.
Reflection primarily compares Beam with GLM-5.2 from Z.ai, an Open Weight model with around 744 billion parameters and 40 billion active ones. According to their own claims: comparable reasoning scores with three to four times lower inference-compute, based on scores for DeepSWE, Humanity’s Last Exam, and Terminal Bench 2.1. The company itself calls this calculation an 'approximate compute comparison' rather than measured inference costs; it excludes input prompt processing, context-dependent attention operations, and serving overhead. In Reflection’s own chart, Beam lags behind GLM-5.3, Kimi K3, and DeepSeek V4.1 Flash on most coding lines. An independent evaluation is not yet available.
The company provides specific figures on the training process. It started with a small prototype, from which a series of increasingly larger models emerged, culminating in Beam Base. This base model was created on a cluster of 6,144 graphics cards and trained on 23.8 trillion tokens from the open web and commercial sources, with a high proportion of code and custom filters for each programming language. Beam Base was completed in under four weeks, followed by a mid-training phase that expanded the context window and reasoning capabilities. For the most compute-intensive phase, Reflection launched 10,000 GB300 cards and 1.3 billion reinforcement learning sandboxes for tasks such as code generation, web search, and agent operation. According to the company, this phase also took four weeks, with 71 failures and a median restart time of eight minutes.
The funding history behind it is steep. Misha Laskin, previously responsible for Reward Modeling at Google’s Gemini project, and Ioannis Antonoglou, co-developer of AlphaGo, founded Reflection in Brooklyn in March 2024, initially to automate software development. In March 2025, the company emerged from stealth mode with $130 million at a valuation of around $545 million, followed by the code agent Asimov in July 2025. In October 2025, Reflection raised $2 billion at an $8 billion valuation from investors including Nvidia and Sequoia, with around 60 employees, and positioned itself as an open frontier lab for businesses and governments. Laskin justified the course at the time by citing DeepSeek and Qwen as a wake-up call: without a counter-movement, the global standard for artificial intelligence would be set by others.
In 2026, infrastructure and government clients were added. In March, Reflection signed a letter of intent with Shinsegae for a 250-megawatt data center in South Korea. In May, the company became a model provider for the US Department of Energy’s Genesis Mission, thereby serving the 17 national laboratories. In June, Reflection confirmed the closing of its round at a $25 billion pre-money valuation and secured access to SpaceX’s Colossus data center through a compute deal. A volume of $6.3 billion is reported for leased Nvidia GB300-NVL72 systems, each with 72 graphics cards, on which Beam was trained. In July, Nebius pledged over a billion dollars' worth of compute power with GB300 chips through 2029. According to this account, Beam is the first open model from a US startup that can compete with top Chinese models on comparable tasks; a gap remains to closed frontier models. → implicator, The Information, SiliconANGLE
Synthszr Take: From a $545 million to a $25 billion valuation in fifteen months, and the bills for it have long been written: $6.3 billion for leased GB300 systems via SpaceX, over a billion to Nebius by 2029. These commitments are ongoing while Beam is still in red-teaming and the weights are only supposed to arrive 'later this month.' Such a valuation is essentially a pre-financing of compute time that must be recouped from paying customers, and faster than the next release from Hangzhou or Beijing closes the gap again. The fact that Beam lags behind GLM-5.3, Kimi K3, and DeepSeek V4.1 Flash on most coding lines in Reflection’s own chart, and that no independent measurement is available yet, is the most expensive detail of this announcement for investors. The benchmark for the next round will be the 17 national labs and the government clients who ultimately pay the bills.
Meta’s Agent Muse Secretly Creates Profile Pages About Friends and Family
Meta’s personal AI agent, Muse, is designed to create 'a page for every person in the user’s life,' including personal history and relationship details. Independent AI safety researcher Karan Joshi extracted these internal instructions through a normal chat conversation; WIRED reported on it on October 3. According to Meta, the context is fed by public information and information shared by the user, such as an invoice from a contractor or a partner’s favorite flowers. In parallel, technology columnist Jason Aten described in Inc. on September 19 that Muse had offered to research a conversation from his podcast and flagged a deadline message from his editor, even though he had denied access to messages and deactivated Full Disk Access. Meta’s head of communications, Andy Stone, denies any access without these permissions, while head of consumer engineering, David Singleton, calls the agent’s own explanation confused; the contradiction remains unresolved, and a rights violation has not been independently proven. → Techpresso
Synthszr Take: The boundary is crossed here in the form of a helpful offer: the agent doesn’t ask for permission, it suggests something you never told it. Aten had denied message access and turned off Full Disk Access, yet his editor’s deadline email suddenly appeared, and to this day, no one can say for sure whether it was a bug, a notification banner, or something else. This is what an agent that runs continuously in Meta’s cloud feels like: you close the app, it keeps working, and the next time you open it, it knows more than the last time.
Bloomberg: China’s AI Models Are Now Only 3 Percent Behind Top US Models
The lead of American AI models over Chinese ones has shrunk to around 3 percent, according to a Bloomberg investigation. The analysis cites high efficiency and low costs on the Chinese side as the drivers, rather than new records in raw performance. A second report from implicator.ai dates the trigger to September: after the latest DeepSeek release, the gap in common performance tests reportedly fell to this 3 percent. The referenced reports do not specify which benchmark families are included in the calculation or how individual test sets are weighted. → TechOrange 科技報橘
Synthszr Take: A three percentage point gap is within a range that benchmark suites can barely resolve cleanly, and so the question of the best model loses its practical value. The second number from the same newsletter is more interesting: a Stanford-led study finds that in 32 percent of comparisons, the cheaper model ends up being more expensive because it requires more runs and more tokens. China’s labs are optimizing at precisely this point because chip export controls leave them no choice; scarcity has been their toughest training condition for over two years.
OpenAI’s GPT-6.1 Sol Costs a Fifth of Astra and Beats It in Tests
At its DevDay keynote on September 29, 2026, OpenAI introduced the GPT-6.1 Sol model, promoting it, according to their own statements, as 'Near-Astra intelligence for a fifth of the price.' The corresponding post on Hacker News gathered 1,066 points and over 950 comments. Simon Willison, who live-blogged the keynote, subsequently ran the model through his pelican SVG test and considers the results not significantly different from those of the GPT-6 family; several levels of Reasoning Effort from Medium to Max were tested. Another commentator compared the models in a Pac-Man bakeoff, where each model must build a playable Pac-Man. There, GPT 6.1 Sol scored 91 points with a runtime of about nine minutes and a cost of 51 cents. → Simon Willison from Simon Willison’s Newsletter
Synthszr Take: In the Pac-Man test, a run with Sol costs 51 cents and yields 91 points, while Astra lands at 87 points for $2.42. This makes their own top model from the spring the weakest value proposition in-house, and OpenAI includes this assessment right in the announcement. Opus 5.5 scores eight more points than Sol, at about four times the price: This is a calculation every team running agents continuously has to make, and the outcome varies depending on the task.
Berlin-based n8n CEO: Claude is not killing us
n8n founder Jan Oberhauser explained on Aakash Gupta’s podcast why his workflow tool continues to grow despite Claude Code and OpenAI’s Agent Builder. When OpenAI introduced the Agent Builder on October 6, 2025, numerous posts declared n8n and Zapier finished; according to Oberhauser, it became one of the best weeks in the company’s history. In May, SAP invested at a valuation of $5.2 billion, more than double the previous year’s value, and embedded n8n into its agent tool, Joule Studio. According to the company, n8n now has over 200,000 GitHub stars, 1.5 million active users, and 1,200 customers for its enterprise product. Oberhauser describes the difference to Claude Code via the orchestration layer: approval steps before risky actions, a logging of every execution node by node including a restart from that point, plus version history and handover to colleagues by invitation instead of via GitHub. → Aakash Gupta from Product Growth
Synthszr Take: Oberhauser’s four-step process is the most compact product lesson on agents you can get right now: describe, approve, log, hand over. Any useful model can now handle step one in the command line; steps two through four determine whether a process remains a demo or goes into production. The saying about the difference between 95 percent and 100 percent reliability describes all the work that is currently underestimated in product development teams because the initial result comes so quickly.
Munich robotics startup is now also a unicorn
According to an exclusive report from The Wall Street Journal, Munich-based robotics startup RobCo has sold shares in a transaction that values the company at more than one billion dollars. This is twice as much as in the previous round earlier the same year. RobCo builds autonomous industrial robots that take over repetitive manual tasks in production facilities, such as stacking pallets or weighing raw materials. The new flagship product is a two-armed, self-learning robot named Alfie, which the company says is designed to replicate the way a human works on the production line. → Wall Street Journal
Synthszr Take: A doubling of the valuation within a year, without a single revenue figure being made public: what’s being priced in here is an expectation, not a delivery history. Capital is currently desperately searching for a European name for Physical AI, and Alfie delivers exactly the image that paints a pretty picture in an investment memo: a two-armed humanoid-like robot on the line. However, industrial robots are sold based on downtime, service technicians within reach, and multi-year maintenance contracts; demo videos count for little with mid-sized business buyers.
Notion designer Nathan Baschez considers personal AI agents the wrong bet
In a guest post for The Leverage newsletter, Nathan Baschez argues against the prevailing thesis in Silicon Valley that Personal Agents are the all-purpose solution for the future of work. The occasion is a wave of product launches: According to the newsletter, it feels like every company within a 20-mile radius of Stanford has released a product in this category in the past two months, including Grok Bots, Muse, and OpenAI’s Dots. The editor writes that with so much consensus in the Valley, he reflexively looks for a dissenting voice and asked Baschez for one. Baschez is currently a product designer at Notion, was previously the founder of Lex and Every, and worked at Gimlet Media and Substack; he designed and built the first version of Product Hunt. → Nathan Baschez from The Leverage
Synthszr Take: Grok Bots, Muse, and Dots in two months, and all three assume that work is a solo sport. Baschez’s objection is valid but remains too reserved: What holds up a process in an organization is almost always an approval from someone with different goals, and no matter how well-informed an assistant is, it can’t change that. Where automation in companies is truly effective, it was previously clarified who takes over in cases of uncertainty and at what threshold a human gets the case; no provider sells this clarification as a license.
OpenAI makes text watermarks in the API optional, mandatory in the EU
Effective immediately, OpenAI is allowing developers worldwide to enable watermarks for text outputs from selected models via the API. The feature is turned off by default; according to the company, customers should decide for themselves how watermarks fit their transparency obligations and products. Activation is done at the project or organization level; individual API calls do not need to be adjusted afterwards. In parallel, text from ChatGPT and Codex for users in the EU will be automatically watermarked invisibly in the coming weeks, across all plans. The background is Article 50 of the EU AI Act, which requires providers of generative systems to make generated text machine-readable; for providers already active in the market, the deadline is December 2.
The method is called textGrain and embeds a statistical signal into the model’s word choice by preferring certain words when several suitable alternatives are available. Over longer passages, this creates a pattern that a detector is supposed to recognize. OpenAI states that in its own tests, textGrain has matched or surpassed the detection performance of Google’s SynthID for text, but concedes that good results under ideal conditions do not guarantee reliable detection in everyday use. The technology is to be released as freely available software.
The published metrics show the limitations. With a target false positive rate of one percent, the detector recognized watermarks in about 80 percent of 200-token passages and about 95 percent of 400-token passages, measured on answers to psychological questions. For content with less linguistic flexibility, such as mathematics, the rate drops significantly; code is also considered difficult to mark. Editing further weakens the signal: if 10 percent of the words in a 400-token passage are replaced with synonyms, detection drops from about 92 to 66 percent, and with 25 percent replacement, it drops to 17 percent.
For now, OpenAI is not releasing the detector. Vetted researchers and professional organizations can apply for access immediately, with grants made on a case-by-case basis. According to the company, the tool only reports whether an OpenAI watermark is present, without identifying users or disclosing prompts and conversations. For images and audio, the existing verification tools remain publicly accessible, including the Content Provenance API and Content Credentials according to the C2PA standard, which have been in use since 2024.
Anthropic led the way in August by announcing watermarks for Claude, albeit with a different approach: The marking is applied at the model level and thus takes effect regardless of which product or interface generates the text, including the API. As justification, the company stated that there is currently no reliable way to regionally limit the technology. Anthropic does not describe an opt-out for API developers. In addition to both providers, Microsoft, Google, and Meta are also affected by the December deadline. → OpenAI, The New Stack, Engadget, The Verge
Synthszr Take: A watermark that is off by default primarily marks the output of those who want to be transparent anyway. Spam farms, essay mills, and disinformation services will simply leave the switch off, and it costs them nothing. Even with the signal turned on, the evidentiary value quickly deteriorates: replace 10 percent of the words with synonyms, and detection drops from 92 to 66 percent; at 25 percent, only 17 remains. The fact that the detector is only available to selected research institutions makes the marking practically worthless for schools, editorial offices, and courts for the foreseeable future. As a compliance response to Article 50, it is cleanly built, but as a tool against misuse, it is of little use; in the end, the version that will run automatically in the EU from December will be the one that counts.

