Expensive Tokens: Fable 5 Flops with Corporate Customers
- • Anthropic customers prefer cheaper models over the more expensive Fable 5
- • OpenAI and Meta undercut each other in a fierce price war for AI tokens
- • Apple's foldable iPhone to start at over $2,000
Anthropic Customers Switch to Cheaper Models Ahead of IPO
Anthropic’s top model, Claude Fable 5, is finding significantly fewer takers among U.S. corporate customers than the company’s older, cheaper models. The model was launched on June 9 at prices of $10 per million input tokens and $50 per million output tokens. According to billing data from 70,000 companies analyzed by payment service provider Ramp, Fable 5's share of total Anthropic spending is around 11 percent or less. The July breakdown shows Opus 4.8 in the lead with 28.0 percent, followed by Sonnet 4.6 with 8.3 percent and Fable 5 with 8.0 percent; Opus 5, which was only released on July 24, accounts for 3.5 percent. Miles Clements, a partner at Accel, which has invested nearly a billion dollars in Anthropic, says that most users don’t need to operate at the performance limit. The launch of Fable 5 was interrupted in June because the White House forced a withdrawal due to national security concerns; the Trump administration later allowed the restart. Analysts and investors attribute the weak usage primarily to price and performance, not the interruption.
These figures come in the weeks leading up to the IPO. Anthropic’s annualized revenue was $65 billion in July, up from $47 billion in May. The company expects a profitable third quarter by its own calculation method and tells investors it has 6,000 customers with annual expenditures of at least $100,000. The IPO could take place in the coming weeks and value the company at at least two trillion dollars, more than any previous initial public offering, including the $75 billion from SpaceX ($86.2 billion with overallotment). OpenAI’s annualized revenue increased by 35 percent to over $40 billion in the current quarter, driven by the launch of GPT 5.6 in July.
At the same time, price competition is on the move. OpenAI cut the price of its fast entry-level tier Luna by 80 percent and its flagship Sol by 20 percent. Meta is offering Muse Spark 1.2 on a contributor tier, where users can use their own data for training, for $0.10 and $0.20 per million tokens, respectively, compared to $1.25 and $4.25 in the standard plan; cache reads cost $0.002 there. Anthropic has not matched these prices so far and is signaling capacity bottlenecks in stabilizing higher weekly limits via its developer channel. The company is renting 300 megawatts from SpaceX for $1.25 billion a month. In addition to OpenAI, Anthropic, and Google, Z.AI, DeepSeek, Kimi, Meta, and Grok now also supply models with Open Weights that are suitable for agentic sessions; the flagships are about five to ten times more expensive per token, but use fewer tokens for the same task. Investor Gavin Baker points out that two- to three-year contracts for GPU capacity are expiring at around $2 per GPU-hour and are likely to be closer to $4 upon renewal.
On the developer side, usage is shifting towards mixed setups. In an essay published on Sunday, Drew Breunig describes a workflow in which Fable 5 handles the architecture and briefing, and routine implementation is passed to Z.ai’s GLM 5.2, which he estimates costs about one-ninth for his comparison case. According to him, the cheaper model only needs to perform on foreseeable work with good context. For Fable traffic, Anthropic’s API documentation specifies a 30-day retention period. Breunig applies Herb Sutter’s argument about the end of the free lunch for processors to AI development: until now, the next model could compensate for weak prompts and thin context at a similar price. → Martin Alderson, PYMNTS, Simon Willison’s Weblog, RuntimeWire
Synthszr Take: A two trillion dollar valuation sells the narrative that a leading edge can be translated into price. The billing data from 70,000 companies says something different for July: 8.0 percent for Fable 5, 28.0 percent for the older Opus 4.8, voted with their credit cards. The timing makes the cost side tricky, because $1.25 billion a month for 300 megawatts is a fixed cost line that doesn’t move when an entry-level plan next door drops by 80 percent and Meta gives away tokens for ten cents. The growth from 47 to 65 billion annualized in two months is real, but it’s coming from the cheaper models, not from the product that justifies the premium. An offering price that bakes in this premium is betting against its own billing statistics.
The Price War on Inference: OpenAI Cuts Luna by 80%, Meta Almost Gives Tokens Away
Martin Alderson describes in his blog a price drop in AI inference that is gaining momentum in the summer of 2026. OpenAI has cut the price of its affordable tier Luna by 80 percent and its flagship Sol by 20 percent. Meta is offering its Open Weights model Muse Spark 1.2 on its contributor tier, where Alderson says Meta is allowed to train with user data, for $0.10 and $0.20 per million tokens, respectively; it normally costs $1.25 and $4.25. Cache reads there are priced at $0.002 per million tokens. Anthropic has not matched these prices so far: The Financial Times reports under the headline that Anthropic’s best model is losing users while cheaper tools are gaining, on the weak demand for Fable 5 at $10 and $50. → Martin Alderson
Synthszr Take: An 80 percent price drop doesn’t just vanish; it translates into consumption. At $0.10 per million tokens, things become viable that nobody would have touched at $10 and $50: nightly full runs across the entire codebase instead of a frugal query per ticket. The individual provider’s balance sheet looks worse afterwards, but the total token volume in the market is significantly larger, and this demand meets GPU contracts that are being repriced from $2 to around $4 per hour.
Data Center Power Demand to Double to 950 Terawatt-Hours by 2030
The AI infrastructure bottleneck is shifting from chips to gigawatts, writes Linas Beliūnas in his newsletter, referencing the IEA’s April 2026 report on energy and AI. According to the IEA, global electricity consumption by data centers grew by 17 percent in 2025, and by 50 percent for AI-focused facilities. The agency expects consumption to double from 485 terawatt-hours in 2025 to around 950 terawatt-hours by 2030. Beliūnas argues that capital cannot deliver this power: a new AI campus needs turbines, transformers, grid connections, permits, and a local community that wants it, and several of these components have longer lead times than the chips they are meant to supply. He describes the situation as a collision of software-speed demand with infrastructure-speed supply. → Linas from Linas’s Newsletter
Synthszr Take: 485 terawatt-hours today, 950 in 2030, and the grids aren’t growing at the same pace. A GPU order has a delivery time in weeks, a substation in years, and it’s precisely this difference that is shattering the neat calculations of many data center plans. The political part is even more unpleasant because permits and resident approval cannot be ordered, no matter how much capital is behind it.
Jeremy Howard Declares Mojo the Biggest Programming Leap in Decades
In his link blog, Simon Willison links to a post by Jeremy Howard, who calls the new programming language Mojo possibly the most significant advance of the last few decades. Mojo is designed as a superset of Python and is being created by a team led by Chris Lattner, who was previously responsible for LLVM, Clang, and Swift. According to the developers, existing Python code should run unchanged, supplemented by language features for low-level programming: “fn” for typed, compiled functions and “struct” as a memory-optimized alternative to classes. In a video, Howard demonstrates a matrix multiplication that is accelerated by more than 2000x with these features, without making the code unreadable. Willison specifically points to this video as the most convincing part of the argument. → Simon Willison from Simon Willison’s Newsletter
Synthszr Take: The text is from May 2023, and that’s precisely what makes it interesting today. Back then, the 2000x speedup read like the end of C++ in the AI stack; three years later, the models still run on Python and the kernels still on CUDA. Languages win through libraries, package managers, and the people who already work with them, and Lattner has gone through this arduous path twice himself with LLVM and Swift.
Apple’s Foldable iPhone to Start at Over $2,000
Apple will introduce its first foldable iPhone at an event on or around September 9, according to Mark Gurman’s Bloomberg newsletter Power On. The device is expected to cost more than $2,000, making it the most expensive phone in the lineup. According to people who have already used it, the hinge and the iPad-like app layouts on the large internal display are impressive. A telephoto lens is missing, and Touch ID is used for authentication instead of Face ID. Apple is entering a category late that Samsung, Huawei, and Google have been expanding for years, most recently with the Galaxy Z Fold 8. Foldable devices accounted for about 1.6 percent of all smartphone shipments in 2025, with China accounting for more than half of the global market with over 10 million units. → www.bloomberg.com
Synthszr Take: Over $2,000, no telephoto lens, Touch ID instead of Face ID: Apple is selling the return of wonder and making people pay for the anticipation in advance. The second move happens in the same breath, as the event will also see price increases for the standard iPhones, after Macs and iPads already became more expensive by an average of 23 percent two months ago. A halo device at the top shifts the perception of what a phone can cost, and the entire lineup below it gets pulled up along with it.
Bun Creator Rewrites a Million Lines of Zig in Rust for $165,000 in Tokens
Jarred Sumner has ported the JavaScript runtime Bun from Zig to Rust and documented the process in a long-announced blog post, to which Simon Willison points. The trigger was the bug list: according to Sumner, a large portion of the crashes were due to use-after-free, double-free, and forgotten releases in error paths, exactly the class of errors that the Rust compiler catches in safe code. The port was made possible by Bun’s test suite, written in TypeScript, which could serve as a language-agnostic conformance suite with about one million assertions. An agent harness handled the bulk of the translation, initially as an experiment with an early version of the model now available as Mythos/Fable. Over eleven days, Sumner stated he manually read the workflow outputs and had Claude adjust the loop instead of correcting faulty code individually; this was supplemented by adversarial code reviews. → Simon Willison from Simon Willison’s Newsletter
Synthszr Take: The $165,000 in tokens is the cheapest item on this bill. What made the rewrite possible was a test suite with one million assertions, written in TypeScript at a time when no one was thinking about coding agents, plus eleven days during which Sumner read workflow outputs and refined the loop instead of patching code by hand. The model produced the lines; the decision on when a million added lines are ready to be merged came from someone who knows the old Zig codebase by heart.
China’s Humanoid Robot Runs 100 Meters Faster Than Bolt, Then Crashes Into a Wall
At the “Robot Olympics” in Beijing on Saturday, a Chinese humanoid robot ran the 100 meters nearly two-tenths of a second faster than Usain Bolt’s world record. Semafor classifies the performance as its own category of record but points to the pace of development: last year’s best time was beaten by more than ten seconds. Immediately after crossing the finish line, the machine ran into a padded wall because it has difficulty braking. According to the report, the games have primarily increased attention for the leading manufacturer Unitree. → Semafor
Synthszr Take: Acceleration is mechanics, stopping is control engineering, and only one of these disciplines determines commercial use. A sprint is an open-loop control system at full power with not a single decision in between. Braking requires the machine to compute momentum, surface, and the end of the track in real time, and that’s why it ran into the padded wall. The ten-second improvement over the previous year shows how quickly actuators and drive systems are maturing, while the safety logic is visibly not keeping pace.
China’s AI Catch-Up: Who Are the Researchers Behind It?
A portrait of the generation of researchers that has driven China’s surprising AI leap focuses on the individuals for the first time, rather than the models. It describes a cohort of scientists who completed their education and sometimes their early career stages in the Western research system and are now working on frontier models in Chinese labs. According to the report, these career paths explain the change of pace in recent months; individual architectural breakthroughs play a lesser role. The article thus places the leap in a debate that has so far been conducted almost exclusively through benchmarks, chip export controls, and training costs. In parallel, the domestic political dispute over infrastructure has escalated in the U.S. Senator John Fetterman of Pennsylvania on Saturday rejected “AI doomsdaying” and backed President Trump’s course of accelerating the construction of AI data centers. This positions the Democrat against the governor of his own state, Josh Shapiro. Fetterman’s reasoning: China always wins when the United States overreacts to the risks of the technology. The China argument in this debate is thus used to justify the pace of construction, while the issue of skilled workers and researchers hardly features in the political discourse. → Wall Street Journal, newsmax
Synthszr Take: You can accelerate data centers with permits and capital, but not a generation of researchers. China’s leap is thanks to people who were trained in the American scientific system and then went back, and that’s the one ingredient in this whole equation that can’t be controlled by export restrictions. At the end of May, we described the exit ban for Chinese AI experts here: Beijing has long been treating minds as a strategic reserve, while Washington argues over construction times for server halls. Fetterman’s warning against overreaction makes a point but targets the wrong bottleneck, because concrete and transformers can be ordered, but a Ph.D. in reinforcement learning cannot. The next round will be decided by the direction in which careers are heading, and that’s not something you’ll find in any benchmark table.
Anonymous Coding Model Ox Alpha Baffles Developers, Traces Lead to Zhipu
On August 20, 2026, a coding model named Ox Alpha appeared on the model marketplace OpenRouter, with no company name, no logo, and no announcement, listed under the generic provider label “Stealth.” OpenRouter describes it as a reasoning model for coding, continuous agentic work, and production workloads. The open-source coding agent OpenCode announced it would offer the model for free with nearly unlimited usage for one week, until about August 27. The anonymous operator states its capacity is 100 trillion tokens per day, which is about one hundred times what Visa says it consumes in an entire month.
Technically, the context window of 1,048,576 tokens is particularly noteworthy, along with text, image, and video input, function calling, and structured JSON output. Most production-ready models range between 8,000 and 200,000 tokens. The throughput is around 29 tokens per second. Developer Ben Davis pitted the model against the coding agent evaluation DeepSWE and reported 8 out of 10 tasks solved, compared to 65 percent for Claude Fable 5 and 52 percent for GPT-5.6-Sol. A larger community run on 113 tasks subsequently achieved around 63 percent. Stripe CEO Patrick Collison, whose company agreed to acquire OpenRouter on August 19, called the model “very impressive” on X.
The search for its origin began within hours. Using tokenizer fingerprinting, Joseph W. Elstner compared the model’s token counts with fourteen public vocabularies, initially with 95 test inputs and 126 API calls; the older GLM version matched 84 samples exactly, while the best non-GLM candidate matched 46. A follow-up run on August 23 with the published GLM-5 vocabulary resulted in 95 out of 95 matches with a mean absolute error of 0.00. A second researcher, Chetaslua, came to the same conclusion with 30 samples across 14 writing systems, including a constant offset of 75 tokens per request, suggesting an invisible prepended system prompt. For video input, three independent encoder properties matched Zhipu’s GLM-5V-Turbo, including a token consumption of about 147 tokens per video second.
The strongest clue came from an intentionally flawed request: Chetaslua set the top_p parameter to the string “abc,” causing the server to return a Java stack trace with the internal class com.wd.paas.api.domain.v4.chat.ChatCompletionRequest, whose path points to Zhipu’s documented endpoints open.bigmodel.cn and api.z.ai. No company has confirmed its operation so far. In parallel, other interpretations are circulating, including the theory on Wccftech that it could be an unreleased version of Microsoft’s MAI; on Reddit, posts excluding a Chinese origin and those considering it highly probable are juxtaposed.
The data question remains open. OpenRouter states that the anonymous provider stores prompts and completions but does not use them for training. The model listing and the platform’s applicable terms of use describe this point contradictorily. Developers are thus potentially sending proprietary code to an operator whose identity is officially unresolved. → Tech Times, Business Insider, implicator, TechCrunch
Synthszr Take: A model with no name, no logo, and no press release is getting more professional attention in three days than most official launches get in three months. Here, anonymity is the product marketing: it turns every developer into a detective who voluntarily burns 126 API calls to determine its origin and then distributes the findings for free on X. Nobody reads a press release twice, but they do read a Java stack trace with the class path com.wd.paas.api. The price is paid elsewhere: the prompts are stored with an operator whose name is officially unknown, and the terms contradict each other on the very question of what can happen to them. When the free window closes around August 27, we’ll see if the 63% from the major benchmark outlives the myth.

