Nvidia's Numbers: It Went Well Again
- • Nvidia exceeds revenue forecasts and significantly strengthens the AI sector
- • Claude gets its own Chromium browser, improving web access
- • Salesforce integrates CRM features directly into Claude for greater efficiency
Surprisingly Good Nvidia Figures Stabilize the AI Boom
Nvidia reported revenue of $96.2 billion for the second quarter of fiscal year 2027, which ended on July 26: up 18 percent from the previous quarter and 106 percent from a year ago. Net income according to GAAP rose to $59.7 billion, and diluted earnings per share to $2.46 (non-GAAP $2.22). The gross margin was 75.0 percent, up from 72.4 percent in the year-ago quarter. According to Bloomberg, analysts had expected revenue of around $92 billion and $2.09 per share.
The data center business contributed $89.0 billion, an increase of 117 percent year-over-year. Nvidia now reports this segment in two parts: $48.7 billion came from hyperscalers, and $40.3 billion from AI Clouds, industry, and enterprise customers, a segment the company calls ACIE, which grew by 138 percent. According to CFO Colette Kress, shipments of Hopper products to China accounted for less than one percent of data center revenue. The Edge Computing division, which also includes the gaming business, came in at $7.2 billion; Nvidia cited weaker consumer PC sales and increased memory and system prices.
For the current quarter, Nvidia forecasts revenue of $108 billion, plus or minus two percent, with a gross margin of 74 percent. Data center revenue from China is not included in this forecast. This extrapolates to an annual revenue run rate of around $432 billion. Additionally, the company forecast 70 percent revenue growth for fiscal year 2028, well above the 44 percent expected by analysts. The stock initially dipped slightly after the release and then rose by more than five percent in after-hours trading.
CEO Jensen Huang explained that Inference-tokens have become productive and profitable, and that compute is now revenue. A year ago, a single lab drove the expansion; today, multiple frontier labs are scaling in parallel. According to the company, the Vera-Rubin platform is in full production at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius, among others. Nvidia also announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize over $500 billion in third-party capital for infrastructure expansion. During the quarter, approximately $26.0 billion was returned to shareholders through share buybacks and dividends, with free cash flow at $21.3 billion.
In parallel, market demand is shifting from training to inference: Gartner expects global spending on inference to reach $23.3 billion in 2026, compared to $19 billion for training. The wafer-scale provider Cerebras, which went public in May 2026, reported core revenue of $209.9 million for its second quarter, an increase of 103 percent, with three customers accounting for about 76 percent of revenue. → Nvidia Newsroom, Decrypt, The Verge, Motley Fool, Quartz, Constellation Research, Fortune
Synthszr Take: Huang’s phrase 'compute is revenue' is the most precise description of this quarter, and he means it literally: a 75 percent gross margin on $96.2 billion means that computing time itself has become a saleable end product, with pricing power that every software provider dreams of. The new breakdown of the data center business is interesting: $40.3 billion from ACIE—meaning AI clouds, governments, and enterprises—versus $48.7 billion from hyperscalers. Demand is no longer coming from a handful of labs, but from a broad base of buyers who purchase and resell computing power like electricity. When Nvidia organizes over $500 billion in third-party capital for expansion together with Apollo, BlackRock, and KKR, it is essentially co-financing the demand for its own product, and this feedback loop is something to watch when the 70 percent growth forecast for 2028 eventually meets real utilization figures. Producing tokens means making money, and the $108 billion forecast without a single dollar from China is the most confident figure in the entire report.
Anthropic Gives Claude Its Own Browser
Anthropic has given Claude its own Chromium browser on the desktop, which runs within the Cowork feature. It is being rolled out to paying subscribers of the Pro, Max, and Team plans on Mac, Windows, and Linux. Previously, Claude required the Chrome extension for web access, which launched on August 26, 2025, and received a major update just two weeks ago. According to the company, the extension remains the right choice for tasks on an already open page, such as maintaining a CRM or working through an inbox. The built-in browser is separate from the daily browsing environment and therefore has no credentials: Logins can be imported from Chrome, Edge, and Firefox on macOS, as well as from Firefox on Windows and Linux, with banking, email, and single sign-on pages excluded. → The New Stack
Synthszr Take: OpenAI killed Atlas as a standalone product and moved it into the ChatGPT app; two weeks after its last extension update, Anthropic is arriving at the same point from the other direction. The browser as a saleable product has failed once before, when Microsoft dried up Netscape’s revenue by giving away Internet Explorer. This time, it’s the runtime environment for an agent that users are already paying for with a Pro, Max, or Team subscription, so it never has to be profitable on its own. Anthropic consistently calls it Claude’s browser, and the login policy shows why: banking, mail, and single sign-on are kept out, while the rest is pulled over from Chrome, Edge, or Firefox.
Bye, Bye Interfaces: Salesforce Integrates Its CRM into Claude and Goes Headless
On Tuesday, Salesforce and Anthropic announced an expanded partnership under the name Claudeforce, making the CRM directly available in Claude. The centerpiece is 'Salesforce in Claude,' a plugin for Claude CoWork with 37 pre-built sales capabilities for meeting preparation, deal assessment, and pipeline analysis, allowing salespeople to query and modify live data without opening Salesforce itself. The product is currently running with select pilot customers, with an open beta planned for September and additional capabilities for other business areas starting in the third quarter. The precursor is Headless 360, which Salesforce introduced in March at the TDX conference: a collection of APIs, MCP servers, and command-line tools that allow agents to control data and workflows without a user interface. Patrick Stokes, President of Applications and Marketing, told VentureBeat that the idea for the plugin came from Anthropic’s own use, where Salesforce is operated almost exclusively through Claude. → VentureBeat
Synthszr Take: For over two decades, Salesforce owned the interface where salespeople spend their workday, and now it’s voluntarily handing it over to Anthropic. Stokes' calculation—turning 10,000 clicks into 30 seconds—is sound from a business perspective but strategically risky: The clicks were the habit, and the habit was customer retention. What’s left are data, metadata, and workflows codified over years—exactly the layer that a customer could one day tap into through another MCP endpoint if the price is right.
VCs Hype OpenClaw Alternative 'Instinct'
AI startup Instinct, founded just last year, has closed a $250 million Series B round, the company told The Wall Street Journal. This brings its total funding to $350 million at a valuation of $2.5 billion. According to the WSJ, the round was co-led by Index Ventures and Benchmark. The company behind the product is Spear Street Technology, led by founder Noah Shinn, who is 23 years old. Instinct is an assistant that users connect to their apps and devices and control via text message or phone call; according to the provider, it organizes daily life. → TechCrunch
Synthszr Take: $2.5 billion for a product in private beta whose revenue no one mentions: that’s a bet on the curve of the next two years. Index and Benchmark are buying an option on a single assistant becoming the fixed point of contact for calendars, shopping, and subscriptions, and if that happens, today’s price is cheap. The evidence Shinn publicly cites are road trips and canceled subscriptions, nice anecdotes that say nothing about how many of the early users are still around after eight weeks.
3M launches 'Ask 3M': An AI assistant to guide customers through 60,000 products
3M has launched Ask 3M, an AI assistant that answers industrial customers' technical questions about its product range in everyday language. The reason for this is the size of the catalog: The company sells over 60,000 products, from adhesives and abrasives to films and protective equipment, and finding the right item has, according to the company, become a hurdle before purchase. According to 3M, the answers are based exclusively on verified, proprietary documentation from 49 technology platforms and not on open internet data. At launch, the tool covers industrial adhesives and tapes, with more categories to follow. The company cites example queries such as searching for a structural adhesive that bonds carbon fiber to aluminum, or comparing VHB Tape 5952 with Adhesive Transfer Tape 468MP. → MyClaw Newsletter
Synthszr Take: This solved a sales problem, and artificial intelligence was just the tool: 60,000 products on the shelf, and the customer can’t find the right one. The expensive part of the project lies in the decades of data sheets, approval texts, and application reports across 49 technology platforms that someone had to sort so that a machine could provide reliable and citable answers. A chemical company has this substance, a general language model does not, and this creates an advantage that no competitor can replicate.
Nvidia and Perplexity bring Qwen models locally to the PC
On August 25, Nvidia and Perplexity introduced 'Portable Computer,' a feature that allows Perplexity models to run directly on local Nvidia hardware. At launch, it supports Qwen 3.8 27B and Qwen 3.6 35B, plus a variant fine-tuned by Perplexity called Qwen 3.8 PPLX 27B; Nvidia’s Nemotron 3.5 Lightning is set to follow. Both models were sized to fit into 24 GB of graphics memory, as this is the smallest common memory size across the Nvidia product line. This means the same file runs on the DGX Spark with the GB10 Superchip, on RTX Pro workstation cards, and on exactly one consumer card, the GeForce RTX 5090 with 32 GB. To download and run, users need a Perplexity Pro or Max subscription. → AI Secret
Synthszr Take: The computation happens on your card, but the permission for it remains on a Perplexity server, because without an active Pro or Max subscription, Portable Computer can neither be loaded nor started. For data privacy, the gain is real and immediately noticeable: prompts, documents, and intermediate results stay on the machine, and no one can find them in the next training run. The dependency just moves from monthly token billing to subscription management, and on the day you cancel, you’re left with a 24 GB card in your computer that won’t run anything anymore.
Z.ai serves up GLM-5.3-Flash on Chinese chips at 10% of the price
Z.ai (also known as Zhipu) released GLM-5.3-Flash on August 26, the first natively multimodal model in the GLM-5 series. It is a Mixture of Experts-model with 320 billion total parameters, of which 18 billion are active per token, and works with a context window of 1,048,576 tokens. The weights are available under the MIT license on Hugging Face, and there is also a hosted API. The model processes images, videos, and files in addition to text. Before its launch, it ran anonymously for a week as 'ox-alpha' on OpenCode and OpenRouter, where it became the most used model of the week according to both platforms; the attribution to Z.ai was confirmed on Wednesday.
According to the provider, GLM-5.3-Flash outperforms its predecessor GLM-5.2 across benchmarks at about one-tenth of the price. On the Artificial Analysis Intelligence Index v4.1.1, the model achieves 57 points at $0.045 per task, putting it on par with GPT-5.6 Terra, Gemini 3.7 Flash, and Qwen 3.8. On the in-house Code Bench, it scores 29.0 in the highest reasoning mode, compared to 29.5 for Claude Opus 4.8. On DeepSWE v1.1, Z.ai reports 63.4 points versus 46.2 for GLM-5.2, and on AutomationBench, 48.8 versus 26.2. Since the evaluations use different harnesses, context limits, and generation settings, the comparisons depend on the respective test setup. Observers report that the model is noticeably verbose and consumes many tokens, which is of little consequence given the low prices. On OpenRouter, it currently costs $0.075 per million input tokens and $0.25 per million output tokens, with these prices including a 50 percent discount.
Architecturally, the model is designed for the cheapest possible inference. For the first time in the GLM series, it combines linear attention for local dependencies with sparse attention, which retrieves global context via a lightweight indexer. With a one-million-token context, a method called IndexPool compresses four indexer key vectors into one to limit latency and memory requirements. Compared to the GLM-4.5 series, the model, with a similar overall size, halves both the active parameters (18 instead of 32 billion) and the number of layers (45 instead of 92). Z.ai reports a threefold reduction in attention compute and a 4.4 times smaller KV-Cache. It was trained on a multimodal corpus of 30 trillion tokens.
Z.ai handled operations entirely on domestically produced AI accelerators. The company describes a stack built on SGLang that separates encoding, prefill, and decoding, and reports a tripling of end-to-end serving performance across tens of thousands of domestic accelerators compared to its own baseline on the same hardware. According to its own statements, Z.ai achieves hardware efficiency and cost per token on par with common Nvidia GPUs. During the anonymous testing phase, OpenCode stated the available capacity was 100 trillion tokens per day, free to use. Local deployment is supported via SGLang, vLLM, TokenSpeed, and KTransformers; subscribers to the GLM Coding Plan have three times the quota of the predecessor. The release coincided with the day of Nvidia’s quarterly results, when the cost of credit default swaps on technology bonds also gained renewed attention. → MarkTechPost, The New Stack, z, Hugging Face, TestingCatalog AI News, TechCrunch, Cautious Optimism, Business Insider
Synthszr Take: Tens of thousands of domestic accelerators delivered 100 trillion tokens per day for a week, for free, and nobody noticed until Z.ai took off the mask. The crucial detail in the blog post is the tripling of serving performance on the same domestic hardware, bringing costs per token down to the range of common Nvidia cards. This is achieved through the architecture: 18 billion active parameters, halved layer count, a KV-Cache shrunk by a factor of 4.4. The chip doesn’t need to be as capable if the model demands less of it, and it’s precisely this calculation that export controls, which think in FLOPs instead of stacks, disrupt. For anyone currently planning data center capacity, the question of which accelerator lies beneath the model is, as of today, negotiable. Nvidia’s pricing power rests noticeably more on training runs than on inference.
Bill Gates: The danger thresholds of AI have long been crossed
On August 26, Bill Gates published an approximately 5,800-word essay titled 'The turbulent AI era is here. The choices we make are critical,' his first detailed statement on artificial intelligence in three years. The core message: The world is not preparing for the upheavals the technology will cause in employment, security, and human relationships. 'There is no plan that eases entry into the AI era,' he writes. Even under the best of circumstances, the transition will be one of the most turbulent periods in human history. He told Axios that if the current course and pace continue, there is a very high probability of a net negative outcome; both major U.S. parties are approaching the issue with too little thought.
In an interview with MIT Technology Review, conducted at the offices of Gates Ventures in Kirkland near Seattle, the 70-year-old explained why he is speaking out right now. He said that thresholds have been crossed in biological, cyber, psychosocial, and labor-market-destroying capabilities, as well as in the lack of control. For years, it was said that rules would be established as these thresholds were approached; the fact that they have now been crossed without such rules has put him in a state of shock. He was particularly drastic in his comments on the biological capabilities of current Frontier Models: any model that can generate novel molecules should be monitored, and he considers the risk of bioterrorism to be about 50 times more frightening and likely than a natural pandemic. He describes his new role as that of a 'shrill voice,' both publicly and in conversations with industry, governments, and civil society. The essay is the first in a planned series.
Regarding the job market, Gates argues that substitution is beginning in the office sector, where companies are already hiring fewer entry-level employees, and will expand to physical labor with better and cheaper robots. By the end of the decade, machines could compete with humans in some construction and hospitality jobs. Those who need it most have the least time: the accountant replaced by a Bot, or the $20-an-hour worker who gives way to a $10-an-hour robot. As a countermeasure, he proposes 'Human Reserved,' i.e., activities that are consciously reserved for humans by society, analogous to nature reserves. He cites the care of his father, who died of Alzheimer’s, and the diagnosis of an incurable disease, which a machine could deliver but shouldn’t. Reskilling and social safety nets are to be financed, among other things, through taxes on robots and on Token, the computational units into which AI models break down text.
Gates names the establishment of a national and international framework as the highest priority, comparing the effort to the reorganization of the US government after September 11, which affected only a single area. AI, on the other hand, touches upon employment, education, taxes, energy, elections, public health, the financial system, and more. This also includes cooperation between the US and China: China might agree to restrict dangerous model releases if the US takes the lead, he told Reuters. His team is working on a meeting with Xi Jinping, planned for a trip to China in November; their last meeting was three years ago. The essay is also one of his first major public appearances since an external review commissioned by the Gates Foundation documented around 30 meetings between Jeffrey Epstein and the foundation’s leadership, including Gates, between 2011 and 2014. → MIT Technology Review, Wall Street Journal, Silicon Republic, Tom’s Hardware, The Guardian, Business Insider, Reuters, Quartz
Synthszr Take: For four decades, Gates has been the bearer of good news, from a computer on every desk to the foundation’s vaccination campaigns, and as recently as January, he sounded relaxed about AI. Now, a 70-year-old sits in a conference room in Kirkland, rocking in his chair, and quantifies bioterrorism as 50 times more likely and scarier than a natural pandemic. The most revealing paragraph of the essay is the one about the caregivers for his father, who died of Alzheimer’s: it’s not a technologist arguing with scenarios, but a son remembering a boundary he experienced firsthand. The fact that he, of all people, describes himself as a 'shrill voice' shows how little connection he finds in his own milieu, which, as he says, is reluctant to criticize itself. The biggest gamble in the paper is the November meeting with Xi Jinping: a private citizen with no official office is trying to organize what no one in Washington or Brussels currently has the strength to do.

