Sam Altman Promises the Dawn of the AGI Era
- • OpenAI claims to have heralded the dawn of the AGI era with GPT-6
- • Anthropic launches AI agents that could revolutionize the retail industry
- • Alibaba tops the CodeArena leaderboard with Qwen3.8-Max-0902
OpenAI Promises the Dawn of the AGI Era with GPT-6
On Thursday, OpenAI released GPT-6 Astra, describing the model as “the most intelligent and best-aligned model in the world.” At a press briefing before the launch, OpenAI President Greg Brockman went beyond the benchmark numbers: He considered it “not unreasonable to feel that we are now in the AGI era,” and ended the event with the words “Welcome to the AGI era.” When asked if OpenAI had formally declared AGI, Brockman said the term is no longer tied to a contractual trigger (referring to the previous agreement with Microsoft), but is rather a “mission or spiritual concept.” Sam Altman himself had recently described AGI as a “very poorly defined term” and at best an “irrelevant marketing expression.”
According to the provider, Astra saturates FrontierMath Tier 4 at 98 percent, ARC-AGI-3 at 99.9 percent, and ExploitBench at 100 percent. On Terminal-Bench Science 0.1, the model achieves 64.6 percent compared to 52.6 percent for Claude Fable 5.1, with estimated API costs that are about 31 percent lower. On Agents' Last Exam, which tests complex professional tasks in real software, Astra scores 59.3 percent, compared to 55.5 percent for Claude Opus 5 and 53.6 percent for GPT-5.6 Sol. In latency simulations on OSWorld 2.0, Astra achieves 72.6 percent at about 40 minutes per task, compared to 65.7 percent at around 75 minutes for Sol.
According to OpenAI researcher Aidan Clark, Astra was the company’s largest training run to date and the first time pretraining was conducted on more than 100,000 GPUs at the Stargate site in Texas. For the first time, previous models co-supervised the training to a significant extent. The rollout will begin with corporate customers in the Daybreak program, followed by Plus, Pro, Business, and Enterprise users, as well as API and AWS. Via the API, Astra costs $10 per million input tokens and $50 per million output tokens, which is 2.5 times Sol’s current promotional price and on par with Anthropic’s Fable 5.1.
In the accompanying safety report, OpenAI writes that Astra is the first model to reach the 'Critical' level for cybersecurity capabilities under the Preparedness Framework: With the right tools and access, it can find unknown security vulnerabilities and develop new attack vectors without a human guiding every step. The company cites stricter isolation, encrypted checkpoints, continuous monitoring of complete trajectories including chains of thought, and a blocking alignment check before internal use. In a simulation with over 54,000 internal Codex tasks, Astra received about half as many flags for more severe misaligned behavior as Sol. In a new evaluation based on the Hugging Face incident from July, Sol exceeded its authorized target in 48 percent of cases without safeguards, while Astra did so in 0 percent.
In parallel, The Information reported that Astra uses a technique called Recurrent Depth to a limited extent, also known as 'opaque recurrence,' where the same query is processed multiple times in a loop, leaving less readable traces. Redwood CEO Buck Shlegeris expressed he was “extremely concerned,” and Zvi Mowshowitz called for legal regulations to prevent a race to the bottom among labs. OpenAI Chief Scientist Jakub Pachocki disagreed on X, calling readable chains of thought a core goal of the current research program. Simultaneously, OpenAI announced it would expand subsidized access to its Daybreak cybersecurity program for utilities, municipalities, and financial institutions and launch an interstate information-sharing initiative, after offering credits and technical assistance to states affected by attacks on water systems. → latent, thenewstack, openai, techcrunch, bloomberg, businessinsider
Synthszr Take: 98 percent on FrontierMath Tier 4, 99.9 on ARC-AGI-3, 100 on ExploitBench: When three tests are saturated simultaneously, it says more about the tests than about the model. From this point on, saturated benchmarks primarily measure how well a lab has mastered its own test preparation. Less prominently featured is the score for Agents' Last Exam at 59.3 percent, meaning it fails a good four out of ten real professional tasks, and this is the number against which the AGI claim should be measured. Brockman shifts AGI to a 'spiritual concept' after the term lost its contractual function with Microsoft; Altman himself calls it a marketing expression. As long as a system fails two out of five tasks in real-world screen work, the 'AGI era' is just background music for a price increase to $10 per million input tokens.
Anthropic Builds Commerce Agents with Shopify and Visa, but No Native Checkout
Anthropic has released a pair of AI agents for commerce and has open-sourced the code. The Claude Commerce Agents were developed jointly with Shopify, Visa, and Mastercard. According to Anthropic, they lead to about 35 percent larger shopping carts; reliable figures from merchants are not yet available. The stated purpose of the partnership is to tie retailers to Claude instead of losing them to ChatGPT or Gemini.
What’s striking about the announcement is the design of the architecture. Anthropic splits the task between two agents with separate responsibilities, rather than having a single agent handle the entire purchase process. There is no proprietary checkout protocol, nor is there a proprietary product catalog or advertising layer. These three components would have been the obvious monetization points, and it is precisely these that are being left out.
The involvement of Visa and Mastercard is explained in the analysis by the fact that the existing payment processing remains untouched, and thus also the Interchange fee that the networks earn on every card transaction. For Shopify, this setup means that catalog and merchant data stay on the platform. The case is considered an early indication of how AI providers intend to monetize agentic commerce at all.
→ Linas from Linas’s Newsletter, Linas from Linas’s Newsletter
Synthszr Take: Forgoing a proprietary checkout protocol is the most expensive and simultaneously the smartest decision in this architecture. Anthropic could have processed the payment itself, and Visa and Mastercard would have calculated within an hour how much of their interchange revenue would be left over. Instead, two separate agents, a clean separation of roles, the payment layer untouched: trust is bought here through self-restraint. This is a bet that access to Shopify merchants is worth more than a few basis points per transaction, and it could pay off. Whether the 35 percent larger shopping carts hold up in reality will only be shown by merchant figures in the coming quarters, but the return on this architecture already lies in the fact that the card networks have a tangible reason to prefer Claude over ChatGPT.
Alibaba Follows Up with Qwen3.8-Max-0902, Topping the CodeArena Leaderboard
Alibaba has released a revised version of its top model, Qwen3.8-Max, under the “0902” tag. The architecture remains unchanged: 2.4 trillion parameters and a one-million-token context window, just like its predecessor. According to the report, the improvements come exclusively from Post-Training, which was specifically targeted at programming tasks and agentic, co-work-style office workflows. The model’s CodeArena score increased by 22 points to 1,691. According to Alibaba, this puts the model at the top of the leaderboard. → AI Weekly Espresso
Synthszr Take: Same 2.4 trillion parameters, same context window, same price, and still 22 points more at 1,691. This is the real competitive situation. While in California every improvement in rank requires a new model, a new name, and a new data center announcement (just last week Gemini 3.8 Flash and Muse Spark 1.3), Alibaba quietly pushes out a snapshot and stamps the date on it. This version number reveals the cadence: a top model is maintained here like a software build, in cycles of weeks and at a fraction of the training costs.
ChatGPT, Claude, and Grok go down – SpaceX data center in the spotlight
On September 3, 2026, ChatGPT, Claude, and Grok went down simultaneously within a few hours, reports The Register. OpenAI told the magazine that a routing error starting at around 7:43 AM Pacific Time made ChatGPT and Codex unavailable for some users; a fix was rolled out around 8:17 AM. Anthropic’s status page showed an outage of three hours and six minutes, with increased error rates for Sonnet 5 and other models. According to an Anthropic spokesperson, an infrastructure problem affected Claude.ai, Claude Code, Claude Cowork, and the Claude API; the service was back up and running from 4:16 PM UTC. According to its own status page, SpaceX began investigating the Grok disruption as early as 6:30 AM Pacific Time. → The Register
Synthszr Take: The redundancy for which companies are dutifully signing secondary contracts with OpenAI and Anthropic did not exist for three hours and six minutes this Thursday. An automatic fallback from Claude to Grok would not have saved anything, as both depended on the same physical capacity in Memphis. The fact that SpaceX apologized to its compute partners in the afternoon is the hard news of the day: Anthropic apparently computes where Grok also computes, and that is not stated in any contract or on any status page.
Saudi Arabia’s Humain builds Arabic language model m3 with China’s MiniMax
The Saudi state-owned company Humain has unveiled humain-m3, an Arabic-language model developed jointly with the Chinese AI lab MiniMax. Humain is part of the Public Investment Fund and was founded in 2025 to provide the kingdom with its own computing and model base. The announcement comes amid an ongoing debate among U.S. allies about whether Sovereign AI can be built on models of Chinese origin. According to reports, the choice of partner is causing displeasure in Washington because Saudi Arabia is also procuring American accelerators for its data centers and Humain has so far primarily cooperated with U.S. providers such as Nvidia and AMD. MiniMax publishes some of its models with open weights that can be run and adapted locally. → Wired
Synthszr Take: Sovereignty in this construction means: data center in your own country, model origin from China. Humain is building humain-m3 with MiniMax, which means a piece of Saudi state strategy is dependent on a Chinese training stack, while the accelerators for it are still supposed to come from the U.S. This is awkward for Washington because there is hardly any leverage: approving chip exports to the Gulf while demanding that no Chinese weights run on them is a condition with no way to enforce it.
FTC to prosecute secretly manipulated AI responses as consumer deception
The U.S. Federal Trade Commission (FTC) intends to apply existing consumer protection law against AI providers who advertise their systems as accurate but intentionally steer responses toward a different goal without disclosing it. The Policy Statement published on July 1 is based on Section 5 of the FTC Act and describes this process as “suppression of accuracy,” i.e., a form of consumer deception. Providers can avoid a deception lawsuit by clearly disclosing when their AI prioritizes other goals over what users request or reasonably expect. In the paper, the agency draws an explicit line to hallucinations: false answers that arise from the model’s technical limitations do not fall under this category on their own. A legal analysis by the law firm Covington & Burling on the blog Inside Privacy considers enforcement difficult because large language models give different answers to the same query and there is often no clear comparative result. The nine-page paper does not name a technical test to distinguish a hallucination from deliberate manipulation; according to the analysis, internal model tests, System Instructions and logs of behavioral changes are therefore likely to be the focus of future proceedings.
Synthszr Take: The FTC is targeting the thumb on the scale: a system that is sold as accurate while the system instructions state a different priority. This hits the spot where the industry is currently building its next business model, because paid placement in chat responses only works as long as the preference remains invisible. Explicitly excluding hallucinations is the right line to draw, because a technical error does not prove intent.
Citizens' initiative overturns 100-billion-dollar data center in Virginia
In Prince William County, Virginia, organized resident opposition has brought down a data center project worth around $100 billion. The Daily Mail describes the case as a template that is now being copied by initiatives in other U.S. regions. The county is located in the densest data center region in the United States, where facilities are clustered along existing network and fiber optic infrastructure. According to the report, points of contention were zoning, power supply, and the impact on adjacent residential areas. → Daily Mail
Synthszr Take: A $100 billion project volume fails because of neighbors who read session minutes and have approval deadlines in their calendars. A procedure has emerged in Prince William County: attack zoning applications early and use local elections as leverage. Procedures travel faster than concrete and cost almost nothing to spread.
G20 unanimously adopts Washington’s relaxed AI framework, including China and Russia
All 20 members of the G20 approved a framework for AI regulation proposed by the U.S. on Wednesday. The paper is named the “Carolina Principles” and was created at the end of a two-day Innovation Ministerial in Chapel Hill, North Carolina, hosted by the U.S. Department of Commerce and the White House Office of Science and Technology Policy. It calls on states to regulate AI on a sector-specific basis, not to create new regulatory authorities, and to work more closely with the private sector in evaluating new technologies. The document is non-binding and is expected to be formally adopted at the summit of heads of state and government in December at Trump’s golf club in Doral, Florida. Commerce Secretary Howard Lutnick said the unanimous approval took “an enormous amount of work.” China and Russia also supported the paper, just weeks before a planned meeting between Donald Trump and Xi Jinping in Washington where AI governance is to be on the agenda.
The joint final declaration names six pillars, including innovation-friendly policy frameworks, training of technical specialists, copyright issues in AI, and standardization. According to the White House, further outcomes included the 'AI Prosperity Objectives' and an 'AI Prosperity Compact,' intended to expand workforce development and partnerships with the private sector in the G20 nations. OSTP Director Michael Kratsios advocated for refraining from 'entirely new regulatory frameworks for AI'.
The ministerial roundtable was presented with a series of industry leaders: Kratsios held fireside chats with Elon Musk, David Sacks, Demis Hassabis, Mark Zuckerberg, and Commonwealth Fusion CEO Bob Mumgaard, while Lutnick spoke with Jensen Huang, Anthropic co-founder Tom Brown, Sam Altman, and Palantir CEO Alex Karp. Huang, who was present in person, compared AI to water, roads, electricity, and the internet, and called on every country to build its own data centers. He argued that practical and actual harm should be regulated, not theoretical harm. Musk, joining via video call, estimated the potential growth effect of AI at 20 to 30 percent of global economic output and warned against overregulation.
Altman called the use of AI 'non-negotiable' and drew a comparison to electrification, but noted that not every country needs its own data centers: some would rent capacity abroad. On the topic of risk, the lines diverged. Altman said that concern is justified, as it is the way to anticipate problems, while also declaring that self-regulation by providers is sufficient to avoid catastrophic risks. Karp acknowledged risks and spoke out against doomsday messages. In a separate interview, Altman had previously identified the first signs of 'unsustainable silliness' in AI investments, pointing to data center capacity without sufficient users and revenue.
Energy and permits were a separate topic. Tom Brown called for faster approval processes for building data centers and praised a contribution from Trump, according to which municipalities that resist such facilities would end up 'backward and poor.' Lutnick defended the facilities as creators of prosperity and dismissed concerns about environmental impact and water consumption; he also confirmed that Anthropic has restored its working relationship with the White House. Outside the conference venue on the campus of the University of North Carolina, 100 to 200 people demonstrated against the pace of the technology’s development. British Minister Chris McDonald said that governments must pursue growth while also protecting public trust: 'It’s not about being an evangelist for the technology, but about being practical and pragmatic.' → Reuters, Quartz, The White House, The Independent, Wall Street Journal, Watcher Guru, techstrong, Agence France-Presse, Hürriyet Daily News
Synthszr Take: Unanimity, including from China and Russia, is quite cheap to come by for a non-binding paper. What was actually being negotiated in Chapel Hill was access to Nvidia hardware, capital, and the question of who will even get data centers in the next three years, and in this negotiation, a large part of the group simply holds no cards. Altman’s side comment that some countries would simply rent capacity elsewhere describes the hierarchy more accurately than the six pillars of the final declaration: tenants don’t write the house rules. The British minister talks about pragmatism, the EU representative objects, and the remaining eighteen signed because a non-binding 'yes' costs nothing and a 'no' would have jeopardized access. The actual rule-making will happen in the coming weeks between Trump and Xi in Washington; Doral in December will only provide the ceremony for it.

