Trump Crowns Zuck’s Super Week by Adopting His Regulation Ideas
- • Zuckerberg initiated Trump’s AI self-commitment after talks in China.
- • Google’s Gemini 4 Argon impresses in benchmarks, remains exclusive for now.
- • OpenAI accuses Moonshot AI of illicitly using internal data.
Zuckerberg Dictated Trump’s AI Self-Commitment
The AI industry’s self-commitment, signed by Donald Trump and six CEOs on Tuesday, originated with Mark Zuckerberg. The Meta CEO sat next to House Speaker Mike Johnson at the state dinner for Chinese President Xi Jinping on September 24 and discussed regulatory issues with him. This sparked the idea for the document: Zuckerberg then approached Nvidia CEO Jensen Huang, who organized support from leading labs, and circulated a draft before the lunch at the White House. The signatories were Meta, OpenAI, Anthropic, xAI, Google, and Nvidia. All accounts of Zuckerberg’s role come from anonymous sources; Meta declined to comment.
The text is just over 300 words long and outlines four voluntary levels of control: internal controls to monitor models during training and deployment, including cyber, biological, and chemical risks; an internal audit team; an external auditor; and an independent committee to receive reports. There is no enforcement mechanism. Democratic House Leader Hakeem Jeffries called the agreement completely unenforceable.
Over the summer, the administration had considered a different structure, one that Anthropic, OpenAI, and Google had rallied behind: a self-regulatory organization modeled on the brokerage authority FINRA, along with stricter internal audits during model development and external testing. Huang, Zuckerberg, and Elon Musk told Trump this solution would concentrate too much power with the leading AI companies. Zuckerberg directly rejected it in a phone call with Trump in August, and the plan was scrapped. Johnson had made his position clear beforehand: no moratorium, no overregulation, or the race against China would be lost.
During the meeting at the White House, a confrontation with Anthropic CEO Dario Amodei occurred behind the scenes. After lunch in the East Room, several CEOs, including Huang, asked him why he was speaking so extremely publicly about the models' capabilities and risks. Amodei replied that it was important to be open with the public about the models' capabilities and not to downplay the risks. The conversation took place in a small group with White House staff and advisors.
In parallel, Trump is pushing for a rebranding of the technology. Following an online poll on September 19, he settled on 'Super Intelligence' three days later, instructed diplomats to use the term, and signed an executive order requiring federal agencies to adopt it. The document is accordingly named the 'White House Accord on Super Intelligence.' Musk and Huang publicly supported the rebranding, while Meta stated that Zuckerberg already uses the term 'superintelligence' occasionally. → implicator, Business Insider, Gizmodo, Semafor, Wall Street Journal
Synthszr Take: Zuckerberg has taken a beating for the entire industry in Washington for years, and now it’s paying off: He’s the only one who knows exactly what kind of paper to write to ensure it never becomes law. The real punchline is in the August phone call where he talked Trump out of the FINRA option, because a permanent regulatory body would have had staff, a budget, and a memory, whereas the four voluntary levels of control are left to each company. The fact that Anthropic, OpenAI, and Google, of all companies, wanted this tougher solution and still ended up signing the 300-word paper says more about the power dynamics in this group than any press photo in front of the White House. The scene after lunch in the East Room, where Amodei is asked to explain why he’s publicly talking about risks, shows where dissenting opinions in this constellation end up: in a small, closed-door meeting. The rebranding to 'Super Intelligence' is the cheapest part of the deal: A word costs nothing, a regulatory authority would have cost decades.
Google: Gemini 4 Is Better Than the Competition and, Better Yet, Is Staying in Mountain View
On Wednesday, Google unveiled Gemini 4 Argon, the long-awaited top-tier model, and is initially releasing it only to a closed circle of testers. According to metrics published by Google, Argon leads or is on par with GPT-6 Astra, Claude Opus 5.5, and Fable 5.1 in 13 out of 18 benchmarks, though it sometimes lags by up to ten points. The lead is particularly significant in office tasks: 51.3 percent on Zapier’s AutomationBench and 84.2 percent on GraphWalks for inputs between 256K and one million tokens. A new feature is the one-million-token output limit; previously, Gemini models were capped at 64,000. For participants in the Fairwind program and internal teams, Google is delivering the model without cyber-Guardrails, according to the company; its subsidiary Wiz, acquired in March for $32 billion, is already using it in its Scan-for-Good initiative. → The New Stack
Synthszr Take: Google highlights the coding capabilities, but it’s precisely there that Argon, at 55.0 percent on FrontierSWE v2, is ten and a half points behind GPT-6 Astra, and on Terminal-Bench 4.0, nine points behind Opus 5.5. The top score of 77.9 percent on DeepSWE v1.1 is real, but one record plus two last-place finishes makes for a mixed bag for everyday use in the editor, especially since Codex and Claude Code are measured there as a model-plus-toolchain. The model is strong in knowledge work, measurably almost nine points ahead of Opus 5.5 on AutomationBench, and over twelve points ahead of GPT-6 Astra with long contexts.
OpenAI Accuses Moonshot AI of Training Kimi with ChatGPT
OpenAI is accusing the Chinese AI firm Moonshot AI of extensively scraping the internal thought processes of its models to train the Kimi model. In a blog post, the company describes the method as Adversarial Distillation: Specifically crafted prompts were allegedly used to reveal the normally hidden reasoning traces, the outputs of which were then fed into the training of a third-party model. According to OpenAI, the campaign began on July 1 and intensified over the course of the month. On July 24 and 25 alone, the company claims to have counted up to 16,000 requests using a specific prompting technique from more than 4,000 accounts. Upon further investigation, over 15,000 users were allegedly found to have used similar prompts, and access was completely blocked on July 28. → Business Today
Synthszr Take: On two days in July, 16,000 cleverly crafted requests from just over 4,000 accounts were enough to siphon off what OpenAI calls its proprietary reasoning. No one opened a database; the product itself was the vulnerability because you can’t sell thought processes without giving them away. Defense thus becomes a constant monitoring of one’s own traffic, and any tightening of pattern recognition and blocking first hits the paying developers who simply want to understand how an answer was generated.
Apple’s Smart Home Hub to Come in iMac G4 Look
Apple’s long-awaited smart home hub will feature a design inspired by the 2002 iMac G4, reports Mark Gurman for Bloomberg. According to the report, the device has a square 6-inch display that sits on a round base with perforated speaker rings, connected by a polished metal arm, and can be manually tilted forward and backward. It’s about as thick as an iPhone, has no battery or volume/power buttons, and must be permanently connected to power via USB-C. A second variant, codenamed J491, is designed for wall mounting and houses the speakers and microphones in a flat, square base below the screen. The hub is intended to identify users by voice or facial recognition and personalize content accordingly, including a guest mode that hides personal information; the facial recognition is reportedly less reliable than Face ID, which is why the iPhone serves as a fallback. → Techpresso
Synthszr Take: The most defining detail of this device is a swivel arm from 2002. Apple is selling nostalgia for a time when Cupertino dictated the design language of the entire industry, and that’s why everyone is now talking about the hemispherical base, while Siri remains in the background. This fits the picture: according to Gurman, the facial recognition is supposed to work less well than Face ID, with the iPhone stepping in as a backup—a hub that has to ask its phone who is standing in front of it.
Starting in January, Amazon Will Sell Brands Influence on Alexa’s Purchase Recommendations
Amazon is introducing an advertising format called 'Branded Conversations,' which will allow brands to co-determine how Alexa talks about their products. It was unveiled at the company’s own unBoxed conference, with a planned launch in the US for January. Technically, the format is an extension of Sponsored Brands prompts, which appear in search results and on product detail pages, leading shoppers into a dialogue with Alexa for Shopping. Advertisers provide a so-called Conversation Brief, in which they describe differences from competitors, product comparisons, and potential bundles; Amazon combines this information with data from product pages and brand stores. The resulting dialogues will be marked as sponsored, including a disclosure that brand information has co-shaped the response. → STACKED MARKETER
Synthszr Take: The ad space here is the answer itself, and therein lies the real monetization. Amazon is selling influence over the phrasing of a recommendation and pricing it per prompt based on clicks, conversions, and return on ad spend—the very metrics used to pull budgets from traditional search advertising. The 48 percent higher conversion rate and 21 percent larger shopping carts mentioned by Andy Jassy in the second quarter are the sales pitch for every media agency, even if it’s Amazon’s own figure. The price is stated in the Ipsos survey: 14 percent trust an AI recommendation more if brands have co-written it.
Microsoft’s First Streaming Transcription Model Is Super Fast
Microsoft has expanded its MAI-model family with the first streaming-transcription model and has placed two new speech synthesis models alongside it. MAI-Transcribe-2-Streaming receives spoken language via a WebSocket and continuously updates the transcript as someone continues to speak, before marking the text as final at the end. According to the company, the first transcript hypothesis is available on average after 320 milliseconds, although Microsoft explicitly does not guarantee this value for every scenario, as network connection and the downstream response model play a role. The model is listed via Vercel’s AI Gateway, costs 54 cents per audio hour, and supports over 60 languages with automatic language detection. For comparison, the non-streaming version MAI-Transcribe-2, introduced last month, costs 10 cents per audio hour because it waits for the end of the utterance. → SiliconANGLE
Synthszr Take: The jump from 10 to 54 cents per audio hour tells the real story: the five-and-a-half-fold increase in cost is solely for the decision to no longer wait for the end of the sentence. A model that revises its own transcript with every new syllable constantly burns computing time on hypotheses that are discarded two words later. The 320 milliseconds are just the starting gun for a latency budget, in which network transit time, the pass-through Mai-Thinking-1, and the synthesis of the voice models are added up. Only this sum determines whether a conversation feels natural or if the person on the other end starts to falter.
OpenAI Fires Three Safety Researchers for Leaking Internal Information
OpenAI has parted ways with three researchers who allegedly shared confidential company information with an external AI safety organization. The Wall Street Journal reported this exclusively, citing people familiar with the matter. According to one of these people, the company recently informed parts of its staff that three employees from the safety team had been terminated. The WSJ places the separation in a period where OpenAI is responding to several incidents in which its own models went off track. → Wall Street Journal
Synthszr Take: Three firings in the safety team, and the subsequent reflex is well-known from tech history: Apple built its own internal investigation unit after the iPhone 4 leak, and Google noticeably tightened access to internal documents and forums after the memo leaks. The crackdown always ends up hitting the places where discussions were previously most open. At OpenAI, this happens to be the very department whose work depends on being able to report uncomfortable findings externally, and which, according to the WSJ, is currently dealing with several incidents of derailed models.
OpenAI Turns ChatGPT into a Virtual Try-On in Checkout
On Thursday, OpenAI rolled out two new shopping features for ChatGPT worldwide: a virtual try-on and a favorites function. For the try-on, users upload a selfie or a full-body photo and see clothing or accessories visualized on themselves, using a new 'Try On' button in the shopping results or by uploading a product image, such as a screenshot. The favorites function saves found products along with the try-on images in a library within the app. Technically, both features run on the new ChatGPT Images 2.5 model, which, according to OpenAI, is designed to deliver more natural light, finer textures, and lower latency in image generation. According to TechCrunch, users can also describe a style and have the app find matching items, or upload photos of celebrity outfits to find purchasable equivalents. → TechCrunch
Synthszr Take: The image shows me if the color suits me, but it remains silent on the one question that decides the purchase: Will the pants fit in my size? Return rates in online fashion retail depend on the cut, fabric behavior, and sizing logic of the respective brand, and an image model lacks precisely this data, no matter how natural the light looks. The favorites library is incidentally more valuable for shoppers because it shows what you genuinely look at again over weeks, rather than what you just found pretty in the moment.
Cloudflare Counters Jev with Two Open Clef Models, Two Weeks After Its Launch
On Thursday, Cloudflare released two of its own Decision Models: Clef and the smaller, faster Clef-flash. Both answer the same kind of limited, structured questions as TypeSafe AI’s Jev model—i.e., yes/no, multiple choice, and rankings with probabilities that an agent can process directly. The models are hosted on Workers AI and are also available on Hugging Face under the Apache-2.0 license as open-weight models. The API is fully Jev-compatible, meaning Clef can be used as a drop-in replacement without any changes.
Technically, Clef is built on a language model foundation: frozen, fine-tuned versions of Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, which perform a prefill-only pass during inference before the answer options are evaluated in parallel. TypeSafe continues to keep the architecture behind Jev a secret. Clef processes images and video in addition to text and supports a context window of 64k tokens; with Jev, the state plus the longest single question is limited to 32k. According to the responsible product manager, Michelle Chen, Clef-flash requires a GPU with at least 41 GB of VRAM, and Clef requires 85 GB, each with single parallelism and 64k context.
The performance claims come from Cloudflare itself. The company benchmarked Clef against Jev and other open decision models on the Jev Decision Index and reports higher accuracy at a slightly lower speed, while Clef-flash is said to be significantly faster with comparable accuracy. Against TypeSafe’s own benchmarks, Cloudflare claims leads in three out of four areas, losing, by its own account, only in agent-trace observability. These values have not yet been reproduced and officially listed on the Decision Index.
As an internal use case, Cloudflare cites its own threat intelligence department, where Clef categorizes domains, taking 2.2 seconds with browser rendering, compared to 4.7 seconds for the fastest general-purpose language model used, gpt-oss-120b. In terms of price, Clef is more expensive than Jev: $0.24 per million tokens compared to $0.042, which is almost six times as much. Despite being described as 'open source' in the announcement, the training datasets are not public, as Chen confirmed to The Register. Concurrently, Cloudflare is launching a new reinforcement learning product that allows customers to fine-tune Clef for their own use cases. → Cloudflare, The Register, Slashdot
Synthszr Take: Two weeks, that was TypeSafe’s entire head start. In mid-September, Jev was the surprise of the month; now, an API-compatible alternative is available for download that also processes images and video and offers a 64k context instead of 32k. The fact that Clef costs almost six times as much at $0.24 per million tokens only protects TypeSafe until someone loads the weights onto their own card with 41 GB of VRAM and pays nothing per token. The secret architecture behind Jev is the weakest position possible in this situation: Cloudflare demonstrates with a frozen Qwen backbone and a prefill-only pass that the recipe is not magic, and lays it bare. All that’s left for TypeSafe is the agent-trace observability, the one area where Clef loses according to Cloudflare’s own numbers, and it’s hard to build a company on that.

