Headwinds for OpenAI and Anthropic: Major Customers Seek Alternatives
Apple Podcasts
Spotify
synthszr #218 from Tuesday, August 4, 2026

Headwinds for OpenAI and Anthropic: Major Customers Seek Alternatives

  • • Major customers plan to leave OpenAI and Anthropic over trust issues.
  • • Moonshot AI halts Kimi subscriptions after record demand within 48 hours.
  • • Ksenia Se strongly refutes theses about Google's withdrawal from the agent race.

Major Customers Like Figma and ElevenLabs Plan to Flee OpenAI and Anthropic

Investor Jason Calacanis claimed on the All-In podcast that several of OpenAI's and Anthropic's largest customers are preparing their exit. He specifically named ElevenLabs, Figma, and Lovable, which he described as companies paying between $50 and $100 million per year to the model providers. He cited trust as the reason: these customers no longer assume that their suppliers won't become their biggest competitors. Calacanis pointed to Anthropic, which is evolving from a pure model provider to a product company across the entire value chain, pushing into design, programming, legal, and finance. As a concrete example, he mentioned the launch of Claude Design, which competes directly with Figma. The report does not contain confirmations from the named companies about actual switches or cancellations; it is Calacanis's assessment. → AI Secret

Synthszr Take: Paying a provider $100 million a year is a bet that they won't move into your market. For Figma, that bet was lost the moment Claude Design came into existence. And from Anthropic's perspective, the move is logical: with every call, the paying customer also delivers usage patterns from which a competing product can be built. For the buyer, this flips priorities in seconds: build a second source, put model access behind your own abstraction layer, and the next contract needs non-compete clauses and a strict clause on the use of your own usage data. In February, it was Notion that showed Figma the door and built its design with Claude Code; now Figma is on the other side of the same table. At these sums, loyalty was always bought, but now it costs more than a volume discount on tokens.

Demand Boom: Moonshot Halts New Kimi Subscriptions

On the evening of July 19, Moonshot AI stopped accepting new consumer subscriptions after demand for its new model, Kimi K3, brought it close to its computing capacity within 48 hours. This was reported by Hello China Tech. K3 was released on July 17, with 2.8 trillion parameters and aimed at the same market of Coding-Agent applications and enterprise APIs that its Chinese competitors also serve. Moonshot is considered one of China's best-funded AI labs. Two days after the K3 launch, on July 19, Alibaba released Qwen3.8-Max Preview, the next version of its Qwen family, which targets the same segment. According to the source, the registration freeze affects new private customers; the report does not specify how long the block is expected to last. → Hello China Tech

Synthszr Take: 48 hours from launch to a halt on sign-ups: this is the most inconvenient form of success. Moonshot has built a model that has outpaced its own infrastructure and is now turning away customers who were ready to pay. Demand you can't serve goes for a walk, and the alternative, Qwen3.8-Max Preview, was on the shelf two days later. At 2.8 trillion parameters, every consumer chat costs significant compute time, while the same GPUs could generate multiples of that revenue in the enterprise API business. The subscription halt is also a prioritization decision, one that's better to make yourself than to have it forced upon you by latency. With these model sizes, growth planning is capacity planning, and the elegant instruments for it are waitlists and throttled free tiers; the closed sign-up button remains the crude solution. For the next model launch in China, it's worth watching how many days registration stays open.

Debate: Is Google Withdrawing from the Agent Race?

Ksenia Se of the Turing Post frontally attacks an 8,000-word article by Alberto Romero (The Algorithmic Bridge) in her latest issue, calling its core thesis false. Romero claims that Google has not fallen behind in the agent race but has deliberately withdrawn because Demis Hassabis does not believe in recursive self-improvement but in world models, and is aligning the company accordingly. Se counters that it is not Hassabis who determines the course, but Sundar Pichai, Sergey Brin, and Larry Page, and cites evidence that Romero himself partly provides: an internal catch-up team reported by The Information, a Brin memo urging to close the gap in agentic execution and turn their own models into developers, as well as 60-hour work weeks for the Gemini team. On the Q2 earnings call, Pichai announced the largest training run in the company's history for Gemini 4, defended a capex forecast of $195 to $205 billion, and let contractual obligations rise above $800 billion, pushing Alphabet into its first quarter with negative free cash flow since its IPO. Se also reads from the interview Romero relies on that Hassabis describes world models as one of several components built on top of Gemini and describes his own time as half research, half product work. As the hardest evidence, she points to the personnel file: Nobel laureate John Jumper was moved to the coding team Code Strike and, a few months later, joined Anthropic, along with Jonas Adler and Alexander Pritzel. → 🔳 Turing Post

Synthszr Take: A company that is pulling out of a race doesn't write memos about 60-hour work weeks and doesn't burn through its first negative cash flow quarter since 2004. The withdrawal narrative is so popular because it allows Google to maintain the dignity of a strategist, while the numbers simply describe a need to catch up: $195 to $205 billion in capex is the price for lagging behind OpenAI and Anthropic in agents, not proof of an elegant, alternative path. What's interesting is the internal fault line that Se exposes. DeepMind was built for big scientific questions, management wants to win enterprise agents, and the budget belongs to management. When a Nobel laureate is moved from protein structure to a coding team and then sits at a competitor a short time later with two colleagues, that's the price Google pays for catching up, and it's not on any earnings call. Gemini 4 will show whether money translates into closing the gap; the talent that leaves in the process won't be coming back anytime soon.

Nous Research Has Hermes Clean Up Its Own Skills via a Background Process

Nous Research has built a maintenance mechanism for its agent Hermes that regularly reviews the system's self-generated skills and memories. Co-founder Karan Malhotra explains in an interview with Peter Yang that Hermes not only creates its own skills but also maintains them via the so-called Hermes Curator. The Curator runs as a scheduled background process and checks the stored inventory for bloat, redundancy, and inefficiency. This addresses a known problem with self-learning agents: stored abilities and notes grow with usage time and lose precision. Because Hermes is available as open-source software, users can, according to Malhotra, define for themselves what counts as Slop in their context and rewrite the cleanup loop according to their own standards. The interview does not provide further details on the specific implementation or usage numbers. → MyClaw Newsletter

Synthszr Take: The interesting part about Hermes is the question of who writes the deletion criteria. With closed assistants, the provider decides what disappears from your agent's memory, and you don't even find out. Here, the definition of junk is in a file that you can open and change. This is the most practical governance lever I've seen in months: what an agent considers redundant is completely different in a tax firm than in a product team, and it's precisely this threshold that must remain negotiable. Agents left to their own devices become outdated and lose their users' trust; an automated Curator takes care of the grunt work, but the standards remain a human setting. For companies, this means: the cleanup criteria catalog belongs in the repository, versioned and with names attached, like any other operational rule. Open-source software wins here because you can tell your own agent what it's allowed to forget.

Fourth Co-founder Leaves Thinking Machines: Lilian Weng Returns to OpenAI

Lilian Weng, co-founder of Mira Murati's Thinking Machines Lab, announced her departure last week and, according to The Information, rejoined OpenAI a few days later, focusing on recursive self-improvement. She cited health strain from the founder role and a desire for a more focused position as reasons for her exit. She is the fourth co-founder to leave Thinking Machines within a year. Axios places the case in a broader movement: Google lost Noam Shazeer to OpenAI and chemistry Nobel laureate John Jumper to Anthropic in June; Meta saw several expensively acquired researchers for Alexandr Wang's superintelligence unit leave shortly after. An Axios source reports that Anthropic CEO Dario Amodei has expressed internal concern that new hires are coming for the money and not for the mission. Besides compensation, researchers cite access to computing power, influence on the roadmap, and freedom in technical approach as reasons for switching, according to Axios. Over 400 former Apple employees now work at OpenAI; Apple is suing OpenAI, accusing the company of using internal Apple codenames to elicit confidential information from applicants. → Axios AI+

Synthszr Take: Four of the co-founders of a lab that has only existed for about a year are gone. That's a design flaw in retention. Amodei's concern that people are coming for the money is basically an admission that his own mission isn't enough of a binding force against a better offer. The side note in Axios is interesting: that some researchers consider their own window of opportunity to be limited because AI systems will eventually take over model development themselves. Anyone who believes their market value has an expiration date optimizes for cash and visibility, not belonging. But research programs live on trust and shared knowledge over years, and that's precisely what can't be bought back when the same thirty people are rotating in a circle. At Thinking Machines, the musical chairs was already hinted at in January; today we know it wasn't an isolated incident. Talent only stays where leaving costs something: computing power under your own control, decision-making power over the roadmap, a project that is only possible here.

250 Poisoned Documents Are Enough for a Backdoor in Models of Any Size

Anthropic's Alignment Science team, together with the UK AI Security Institute and the Alan Turing Institute, has published a study on Data Poisoning, which finds that around 250 manipulated training documents are sufficient to embed a backdoor in a language model. Models ranging from 600 million to 13 billion parameters were tested; the largest of these, according to the authors, was fed over twenty times more training data than the smallest, yet could be compromised with the same small, absolute number of poison documents. This contradicts the previously common assumption that an attacker must control a percentage of the training data. The attack studied is deliberately narrow: the trigger word “<SUDO>“ causes the model to output random gibberish, which the researchers could measure via the perplexity of the outputs, without additional fine-tuning. Each poison document consisted of a short text snippet, the appended trigger, and 400 to 900 randomly drawn tokens. The authors call it the largest poisoning investigation to date and write that it remains an open question whether the pattern holds for larger models and more harmful behaviors. → ByteByteGo

Synthszr Take: The real result of this study is a unit of measurement. For years, research has calculated poisoning in terms of a percentage of training data, and in this calculation, models automatically became more secure with every additional terabyte. The 250 documents dismantle this sense of security: if the necessary amount remains constant while the data corpus grows by a factor of 20, then the attack becomes relatively cheaper with each scaling step. 250 documents are an afternoon's work for one person with a few blogs, not an operation with a budget. The fact that the tested behavior is harmless (gibberish on command) doesn't change the order of magnitude, it just makes it measurable. The next number will be interesting: if the constant also holds for 70 or 400 billion parameters, the work shifts from model architecture to provenance checking of training data, i.e., to deduplication, origin verification, and trigger detection before pretraining. A field in which almost no one has published so far.

Coldcard Wallets Lose $70 Million Without a Single Device Being Touched

On July 30, 1,082.65 Bitcoin, equivalent to about $70 million, were withdrawn from 1,196 Coldcard hardware wallets within 41 minutes. Galaxy Research claims to have fully reconstructed the sequence of events: the outflows are distributed across six blocks between 01:10 and 01:51 UTC; three intermediate blocks contain none of the transactions, suggesting bundled transfers in bursts. The proceeds are on four addresses and have not moved so far; the initial report only covered one of these addresses, which is why the total has almost doubled from the initially reported 594 BTC. According to the investigators, the cause is a firmware bug in certain Coldcard versions: an internal build setting caused the device to skip the dedicated hardware random number generator, and a check in an included library only tested if the setting was present, not if it was enabled. As a result, key generation fell back to a simple software substitute fed by the chip's serial number and its clock registers, making the supposedly unguessable Seed Phrase computationally searchable. Security firms warn that more wallets could be affected because owners cannot reliably determine which firmware their seed was generated on. → Techpresso

Synthszr Take: The entire attack fits into one line of faulty validation logic: the library asked if a configuration switch existed, instead of asking if it was set to “on.” After that, the key generation drew its randomness from the factory serial number and the chip's clock registers. Both are either fixed or predictable within a very small range, so the space of possible seeds can be calculated offline with hardware you can rent by the hour from any cloud provider. The device behaved correctly; it never released a key; it was never connected to the network, and none of that helped: the physical separation protects the transport path, while the value here lay in the quality of the generated randomness. This is precisely why the situation is so unpleasant for those affected, because a compromised key looks just like a healthy one from the outside, and a firmware update doesn't fix a seed that was weakly generated months ago. According to Galaxy Research, the attacker is still searching, using a method that requires no interaction with victims and whose only limit is computation time. If you are using a Coldcard model with affected firmware, regenerate the seed on verified firmware and move the funds today, not next week.

AI Coding Agents Are Chipping Away at CUDA, Nvidia's Real Competitive Advantage

Business Insider reports that Nvidia's software layer, CUDA, is coming under pressure from AI coding agents. CUDA, short for Compute Unified Device Architecture, has been the real core of its lead for about two decades, according to the report: it's not the chips alone that carry it, but the software that turns them into building blocks for artificial intelligence. The report names longtime Nvidia manager Ian Buck, who still oversees the area today, as its creator. The article's argument: the competitive advantage, long considered unassailable, is losing stability because coding agents can rewrite software that was previously locked into CUDA. This affects the level of GPU-Kernel, i.e., the hardware-level computing routines whose optimization previously required specialized developer work. Business Insider classifies this as a new threat to Nvidia's position, without specifying a concrete timeline. → Business Insider

Synthszr Take: CUDA is software, and software is exactly what agents are best at rewriting. The lead consisted of two things: millions of lines of hand-optimized kernels and a generation of developers who learned nothing else for twenty years. A switch was always a matter of cost, not feasibility: rewrite kernels, verify numerics, chase performance for months, and in the end, you might be lucky to get 90 percent of the original speed. It is precisely this cost position that is now decreasing, and the bottleneck is shifting from writing to verifying. An agent that produces ROCm kernels overnight that deliver identical results in training needs benchmarks, not demos, because you only notice silent deviations in floating-point precision after the third training run. In the next procurement cycle, the test suite that proves the switch—or doesn't—will decide on alternatives. Whether anyone on the team can still write CUDA by hand will then be the lesser question.

Studies: VC-Funded Startups Commit Fraud More Often – And Investors Play a Part

Researchers from Imperial College London and Emlyon Business School have investigated how venture capital-funded tech founders commit fraud and the role their investors play. For the paper published in June, the team led by Tim Weiss and Nevena Radoynovska built a database of all founders and companies against whom the SEC and DOJ took civil or criminal action for securities fraud between 2000 and 2023. The authors describe a three-stage pattern they call Façading: from embellished portrayals of success in pitches to fake evidence and manipulated product demos, illustrated in the paper by a mobile testing app that raised a billion-dollar valuation with fabricated customer contracts and sham revenue. A parallel study by the University of Toronto analyzed 654 fraud cases against US startups with venture capital during the same period: fraud remains rare overall, but occurs more frequently in VC-funded companies than in companies without such financing. Companies that started in overheated market phases with weak oversight and lax investor scrutiny are 19 percent more likely to commit fraud later, according to the study. Startups with founder-controlled boards were twice as likely to be affected as those with investor-controlled or shared boards, and according to the Toronto paper, prior fraud allegations hardly prevent founders from raising new capital for their next company. Weiss suggests that the SEC should conduct regular formal audits and points out that there is no professional association for founders that could enforce a code of conduct. → StrictlyVC

Synthszr Take: The 19 percent is the most interesting finding because it points to market phases and the depth of scrutiny. Overheating plus thin due diligence creates growth expectations that no real operation can meet, and then someone starts inventing contracts. The fact that founder-controlled boards double the probability of fraud is a design decision made in every term sheet, often in favor of deal speed. Even more revealing: the market practically doesn't punish misconduct, as the Toronto data shows, even in cases with high media attention. An ecosystem that celebrates failure without asking about the cause produces exactly this feedback loop. In the current AI environment with its inflated ARR figures, all the conditions of the study are present, and the wave of lawsuits after an IPO typically comes within two years, according to the data. As long as a fraud allegation on a resume doesn't cost the next round, the 19 percent rate is a priced-in component of the model.

Anthropic CEO Amodei Explains in 'Machines of Loving Grace' Why He's Not a Doomer

In his October 2024 essay “Machines of Loving Grace,” Dario Amodei detailed what a world with very powerful artificial intelligence could look like if the risks are managed. The Anthropic CEO contradicts the characterization of him as a pessimist: his work on risks is precisely because these risks are the only thing standing between the present and a positive future. In his view, most people underestimate both the scale of the potential benefits and the severity of the dangers. Amodei gives four reasons why Anthropic publicly speaks mainly about risks: to maximize its own impact, to avoid the impression of propaganda, to avoid grandiosity, and to steer clear of the “sci-fi” baggage of debates about a post-AGI world. The main text focuses on five fields where he sees the greatest direct impact on quality of life: biology and physical health, neuroscience and mental health, economic development and poverty, peace and governance, and work and meaning. Amodei himself marks his predictions as conjectures and writes that a team of experts could write a better version of the text. The essay is currently circulating again via The Argument. → The Argument

Synthszr Take: Amodei's justification for his public risk-focused silence on the benefits is the truly interesting part of the text, and it's a leverage argument: according to him, the benefits will come anyway because market forces are driving them; only the risks are malleable. This is an economically clean division of labor, and it explains why a safety lab sounds like a warning department for years. But fear doesn't finance a roadmap. That's why he follows up with the five fields, with details instead of hedged generalities, and in doing so, makes himself testable: biology, mental health, poverty, governance, work, and meaning are areas where you can count results. We know the same pattern from corporate transformation projects: a well-maintained risk register doesn't get anyone moving as long as there's no picture of what the effort will ultimately yield. The benchmark is now set in biology, the field where Amodei himself is an expert. If Anthropic doesn't show demonstrable acceleration in drug discovery in two years, the essay will be read as prose, not a prediction.

Search is about rankings, AI is not.

RAIDAR (may update)

Search is about rankings, AI is not.

From a ranking, you can't tell which audience sees which answer, which sources the models trust, or which areas no one has claimed yet. RAIDAR maps all of it across every model, customer segment, and market, down to the sources that feed the answers. Not a ranking. A map that tells you where to move. For brands that want to know.

More about RAIDAR →

Last 7 Days

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.