älter | home
Launch Day: Google Still in the Race, Meta Claims Top Spot in CodingSynthszr
synthszr #248 from Thursday, September 3, 2026

Launch Day: Google Still in the Race, Meta Claims Top Spot in Coding

  • • Google releases Gemini 3.8 Flash – the third model in just six weeks.
  • • Meta unveils Muse Spark 1.3 and sees itself in a race with OpenAI.
  • • Google Pics launches as a Canva alternative in the Workspace subscription for businesses.

Google ships Gemini 3.8 Flash: third Flash model in six weeks

On September 2, 2026, Google released Gemini 3.8 Flash and its security variant, Gemini 3.8 Flash Cyber, three weeks after its predecessor, 3.7 Flash. It is the third Flash release within six weeks and the fourth within four months. The introductory price remains at $0.75 per million input tokens and $3.75 per million output tokens; it expires on December 31, 2026, and will then double to $1.50 and $7.50, respectively. The model processes text, images, video, audio, and PDFs with a context window of 1,048,576 tokens and a maximum of 65,536 output tokens. Gemini 3.5 Pro, announced in June, has still not been released; in August, DeepMind CEO Demis Hassabis moved to the Chairman position, with Koray Kavukcuoglu taking over operational management as SVP.

According to Google, 3.8 Flash outperforms most larger frontier models on the DeepSWE v1.1 for long-running software engineering tasks at a fraction of the cost. On the Artificial Analysis Intelligence Index, the model achieves 59 points at a high reasoning level, three more than its predecessor, on par with GPT-5.6 Sol and Grok 4.6. Above it are GLM-5.3 and Kimi K3 with 60 points each, Grok 4.6 and GPT-5.6 Sol with 61 each, Claude Opus 5 with 63, and Claude Fable 5.1 with 66 points. In terms of cost, Artificial Analysis puts Flash 3.8 at $0.58 per index task, while Fable 5.1 costs $3.76. For agentic computer use in the OSWorld-2.0 test, the model improves over 3.7 Flash but remains significantly behind Claude Opus; on Terminal-Bench 4.0, it achieves 19.1 percent.

Google itself admits that 3.8 Flash consumes more tokens: The model performs additional reasoning steps and calls tools iteratively. In Artificial Analysis’s measurement, a task costs about 40 percent more than with the predecessor, despite the token price remaining unchanged, caused by 30 percent more output tokens and more turns per agent run. Developers can still choose between low, medium, and high reasoning levels or stick with 3.7 Flash.

The Cyber variant replaces Gemini 3.5 Flash Cyber and is exclusively accessible through the new Fairwind program, which includes around 650 partners, such as Accenture, CrowdStrike, Datadog, Palo Alto Networks, Snowflake, Wiz, and the Center for Internet Security. Participants also get access to the CodeMender agent. According to the provider’s internal data, the Chrome Security team recorded a 2.6-fold higher patch accuracy, the Cloud team found a critical vulnerability in two hours, and the pass rate on CWE-Bench is 47.2 percent. According to the manufacturer, both models include protective mechanisms against misuse in the areas of CBRN and offensive cyber operations.

In parallel, 3.8 Flash was added to the model menu of AI Mode in Google Search on its release day. Robby Stein, VP Product for Google Search, announced its availability for Google AI Pro and Ultra subscribers worldwide, selectable via the plus icon in the input bar. Users of the free tier have no choice of model and continue to get the default setting. With the two previous Flash generations, the new model later became the global default. → Search Engine Land, The New Stack, The Verge, Vals AI, Ars Technica, The Register, Search Engine Journal, blockchain, NPowerUser, MarkTechPost

Synthszr Take: Three Flash models in six weeks, with a 21-day gap between 3.7 and 3.8: No one maintains this pace voluntarily, but because their own flagship is missing. The promised Gemini 3.5 Pro for June never appeared, Hassabis moved to the Chairman position in August, and since then, Google has been maintaining visibility through the budget class where it can deliver. For everyone integrating these models into products, the cadence is the real cost factor: An evaluation that ran on 3.7 Flash three weeks ago is now measuring a model that no longer even appears in the menu, and the 40 percent higher cost per task only becomes apparent on the following month’s bill. Competitive pressure isn’t recognized by benchmark scores, but by how briefly a provider keeps its own model on the market. Another Flash is likely to arrive before New Year’s Eve, and on January 1, the price will double anyway.

Meta declares itself on par with OpenAI and Anthropic with Muse Spark 1.3

On September 2, 2026, Meta released Muse Spark 1.3, the most powerful model to date from Meta Superintelligence Labs, claiming to catch up with the leading labs. In a Bloomberg interview, AI chief Alexandr Wang described the model as the company’s biggest performance leap yet, stating it is 'competitive' with Claude Fable 5.1 and 'better than' GPT-5.6 Sol, especially in code generation. The model is available immediately via the Meta Model API and in Muse Code, with the rollout to Facebook, Instagram, and the Meta AI app following in the coming days. It is the fourth Spark release since the family’s debut in April 2026.

Independent measurement by Artificial Analysis partially supports this claim. The 'max' variant achieves 62 points on the Intelligence Index and ranks 6th out of 636 evaluated models, behind Claude Fable 5.1 and Claude Opus 5. The 'xhigh' tier, accessible to customers, scores 61 points, putting it on par with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), up from 57 points for Muse Spark 1.2 in August. The 'max' variant remains in limited preview for partners. Meta’s own comparison table was generated on its in-house Harness and is therefore not the same kind of evidence as the external index.

In terms of price per task, Muse Spark 1.3 costs $0.55 compared to $0.95 for GPT-5.6 Sol and $0.94 for Grok 4.6, which is 42 percent less than OpenAI for the same point score. No model with at least 59 points was cheaper in the September 2 comparison; the closest was Gemini 3.8 Flash with 59 points at $0.58. The Token prices remained unchanged at $1.25 per million input tokens, $4.25 for output, and $0.15 for cached input. However, compared to its own predecessor, the cost per task increased from $0.40 to $0.55 because about 57 percent more input tokens were measured per task. A 'Contributor' tier with lower rate limits and use of prompts for training costs $0.10 for input and $0.20 for output.

The biggest gains are in multi-step tool use: In the Tau3-Bench Banking, the score increased from 35 to 47 percent, and in Terminal-Bench 2.1 from 80 to 85 percent. Two measurements declined: AA-LCR from 83 to 79 percent, and Omniscience accuracy by three percentage points with a higher abstention rate. According to the company, the model asks clarifying questions for ambiguous prompts, involves the user when it gets stuck, and confirms consequential actions before execution. A 'max reasoning' mode is slated to be added after further security testing. On X, Mark Zuckerberg also announced that an open-weights version of Muse Spark will be released 'soon,' after Meta switched from the Llama line to a proprietary strategy with the Spark family. → Artificial Analysis, SiliconANGLE, implicator, unite, The Register

Synthszr Take: 'Caught up' is a term from the analyst call, and it barely holds up to their own measurement basis. A score of 62 versus 61 separates Muse Spark from the Claude models, and the 62-point variant is in limited preview with select partners, making it unverifiable for anyone who wants to put it into production. On Bloomberg, Wang calls the model 'better than' GPT-5.6 Sol, while Meta’s comparison table runs on Meta’s own Harness, and two independently measured values have dropped: AA-LCR from 83 to 79 percent. The reliable number is in the price line: $0.55 per task versus $0.95 for OpenAI, with an identical index score of 61. After over $14 billion for ScaleAI and Wang, Meta is selling cost leadership and calling it a tie.

Google launches Canva alternative: piggybacked in the Workspace subscription

Google is entering the market for everyday design tasks with its own image and design tool called Google Pics. The product will be part of Google Workspace for business customers and will also be available to subscribers of Google AI Pro and Ultra. According to TechCrunch, the rollout should reach most Workspace customers in the coming weeks. Technically, Google Pics runs on Google’s image model Nano Banana. It is intended for posters, social media posts, illustrations, and similar visuals—tasks that are typically handled today in Canva or Adobe Express. Unlike Canva, there is no marketplace where illustrators and photographers can post templates and earn royalties: The content is generated via Prompt from a model trained on the work of creatives. → Techpresso

Synthszr Take: From now on, Canva and Adobe Express will have to explain why a company should pay for a second design subscription when a comparable tool is included in the Workspace subscription they already pay for. This is exactly how Google made presentation software cheaper with Slides and video conferencing with Meet: through bundling, not by having the better product. The entry point via Docs and Slides is the weak spot, as this is where the everyday graphics are created, for which marketing teams currently book Canva seats.

Claude Fable 5.1 is clearly the best model in Cursor’s coding benchmark

Cursor has integrated Claude Fable 5.1 into its development environment, calling it the most powerful coding model they have tested to date. According to Cursor, the model achieves a score of 73.4 percent on CursorBench 3.2, their in-house benchmark for practical programming tasks, placing it ahead of all other tested models. The provider highlights a key new feature: Fable 5.1 reviews the code it has written, identifies its own errors, and continues with the task instead of stopping after generating output. According to Cursor, the model is designed for long, multi-step workflows that run without constant intermediate checks. Cache reads are said to be 75 percent cheaper than with Fable 5, significantly reducing the cost of repeatedly used context. → AlphaSignal

Synthszr Take: A model that makes it into the model selection of an IDE where developers spend half their working day has achieved more than one that simply tops a benchmark chart. The 73.4 percent is Cursor’s own measurement using its own testing procedure, which is fine, but it mainly speaks to the model’s suitability for this specific tool. The sentence that really matters comes at the end of the announcement: In teams and organizations, an admin must first enable Fable 5.1 in the dashboard.

Best Buy has AI do 80 percent of creative work, cuts 22 asset systems down to five

Best Buy is building an AI-enabled Content Supply Chain, according to the company, and has consolidated 22 separate asset systems into five. These systems manage approximately 1.4 million assets, each with its own metadata, usage rights, and approval status. The core of the transition is the division of labor in creative production: high-volume tasks are performed by artificial intelligence to about 80 percent completion, with designers handling the remaining 20 percent. The company states that this division is not viable without three conditions: governance, seamlessly connected workflows, and accompanying change management. → MyClaw Newsletter

Synthszr Take: 80 percent of the preliminary work done by machine, 20 percent finished by humans: this division sounds like a solid calculation, not just a boardroom slide. The final 20 percent is the expensive part, as it involves brand judgment, rights issues, and the decision of whether a creative can be released at all. The consolidation of 22 systems into five is therefore the real prerequisite: without clean metadata and up-to-date approval statuses, the human at the end lacks the foundation to make decisions in minutes instead of days.

US government backs OpenAI against The New York Times in copyright dispute

The US government has intervened for the first time in copyright lawsuits concerning AI training, supporting OpenAI’s position. In a brief filed Tuesday in federal court in Manhattan, it argues that training models on copyrighted material is “extraordinarily” transformative and thus covered by Fair Use gedeckt. Since filing its lawsuit in 2023, The New York Times has accused OpenAI and its largest backer, Microsoft, of using millions of newspaper articles without permission to train its chatbot; other newspapers have joined the suit. Such a brief is advisory and not legally binding, but it could provide a tailwind for the tech giants in the proceedings. Deputy Attorney General Stanley Woodward Jr. wrote on X that AI dominance is crucial for national security and prosperity, and the government will not let the country fall behind due to a “plainly incorrect understanding of copyright law.” → The Guardian

Synthszr Take: Legally, this brief binds no one, but politically, it reshuffles the deck. The statement that training is “extraordinarily” transformative will now loom over every one of the dozens of pending cases, including those against Anthropic and Meta. After two conflicting court rulings last year, a court that rules against OpenAI will now also be ruling against the declared position of its own executive branch—and judges are human.

Court overturns Hegseth’s contract ban, Lutnick declares Anthropic acceptable again

Commerce Secretary Howard Lutnick has placed Anthropic back on the “right side” of the Trump administration, reports Axios from an interview picked up by Reuters. This follows a dispute with the Department of Defense: Pete Hegseth had excluded Anthropic from certain military contracts after the company refused to approve the use of its Claude models for domestic surveillance and autonomous weapons systems. Anthropic subsequently sued in a California court. On August 27, a U.S. judge ruled in the company’s favor in this dispute. According to Reuters, Anthropic and the Commerce Department did not initially respond to requests for comment on Lutnick’s statement. → Reuters

Synthszr Take: A California court ruled on August 27, settling the Defense Secretary’s contract ban before Lutnick even had to find his formula for reconciliation. Anthropic refused to approve use for domestic surveillance and autonomous weapons systems, sued, and won: its own terms of service now hold up in court, not just in corporate communications. For any provider doing business with the government, this is a solid precedent because a single department head can no longer single-handedly block contracts in defiance of the contractual terms. The real win is on the procurement side: a military customer who knows that their provider’s terms are legally sound will calculate differently than one hoping for political goodwill.

Beijing forces Tencent, Alibaba, and ByteDance to open their walled gardens

China’s Ministry of Industry and Information Technology (MIIT) has ordered the country’s leading internet companies to stop blocking each other’s external links. At a meeting on September 9, executives from Tencent, Alibaba, ByteDance, and Baidu were summoned and required to submit implementation plans by September 17, reports Caixin. The practice of building so-called Walled Gardens has been common for years: a Taobao product link cannot be opened in WeChat, Douyin videos cannot be shared directly with WeChat contacts, and Alibaba, in turn, blocks WeChat Pay on Taobao and Tmall. Zhao Zhiguo, head of the responsible MIIT authority, stated at a press conference that link blocking disrupts the user experience and the market, and that the opening up should be gradual. The scale is significant: Tencent has nearly 1.9 billion monthly active users with WeChat and QQ, Alibaba has 1.6 billion via Taobao and Alipay, and Douyin has 640 million. → Hello China Tech

Synthszr Take: A link that can be opened is the cheapest form of competition policy a state can have. Tencent sits on almost 1.9 billion monthly active accounts, and a good part of the value of this number depends on a Taobao link in WeChat leading to nowhere. If this block falls, the cost of user acquisition will drop for everyone who doesn’t own a billion-user app, and the moat will shrink to what the product itself can deliver.

Adobe brings over 70 creative features to Slack chat

Adobe has launched an integration called 'Adobe for Slack,' which brings more than 70 features from Firefly, Express, Photoshop, Premiere, Acrobat, InDesign, Illustrator, Stock, and Lightroom directly into the messenger. Users write a task for Slackbot in natural language, and the bot selects the appropriate Adobe tool, accessing content from Slack channels and Canvases as well as assets in the Adobe Creative Cloud. Adobe lists use cases such as turning a table into a graphic, creating marketing and social content from templates, building variants of already approved content, or performing bulk edits on large sets of images.

The app is available immediately for Slack customers on the Business+ and Enterprise+ plans, on web, mobile, and desktop. An Adobe account is not mandatory; a paid subscription unlocks additional features and access to one’s own cloud assets. Technically, this ties into Slackbot’s Model Context Protocol support, which was introduced in June and initially connected Canva and Figma. Late last year, Adobe had already embedded Express tools in Slack, albeit with a much narrower range of functions.

The number of 'more than 70 tools' corresponds to the scope Adobe had previously stated for its ChatGPT plugin, and the implementation appears nearly identical on both platforms. Forrester analyst Joe Cicman categorizes the move as Adobe opening its tools to a new user group without having to write a new interface or alienate existing professional users; in practice, this could let marketers test many campaign ideas faster. Adobe itself argues in favor of Firefly that its image generation does not violate copyrights. Last week, Salesforce announced it would embed its own functions into Anthropic’s Claude, with an open beta later this month. → Engadget, TechTarget

Synthszr Take: The benefit lies in the context switches that are eliminated: no jumping to the browser, no downloading, no re-uploading, no 'send me the final version again.' Anyone who has coordinated approval loops for 200 product images in a channel knows that the work isn’t in Photoshop, but in the back-and-forth between the channel, folder, and tool. Adobe is capitalizing on this context switch and turning the thread itself into the workspace, including channel content and Canvases as source material. The fact that an Adobe account isn’t required to get started makes things even clearer operationally: The intern in the marketing channel produces the variant without anyone having to request a license. Friction reduction is unspectacular and rarely celebrated in press releases, but it determines whether an integration will be used in four weeks or end up as another unused creative dashboard.

Anthropic’s watermark in Claude Fable 5.1 reliably marks prose, but is incomplete for code

Anthropic has released Claude Fable 5.1, embedding a statistical signature into the generated texts, which is only partially effective for programming code. The method is based on SynthID-Text from Google DeepMind and adds neither metadata nor hidden characters, but rather alters the randomness in the selection of the next token. Over a sufficiently long response, this creates a pattern that can be detected with the appropriate key; copying does not remove it, but heavy rewriting does. This works well for natural language because there are usually multiple ways to phrase the same statement. The situation is different with code: A different variable name, operator, or function call can change the program’s behavior, which is why Anthropic omits the marking where a specific token is necessary for correctness. The signal may appear in less fixed parts like comments, and short answers may not contain enough substance for a reliable detection. The company announced the move on August 14, after signing the EU Code of Practice on the Transparency of AI-Generated Content along with about 190 other signatories; the marking is being rolled out worldwide, with older Claude models to follow in the coming months.

Technically, Fable 5.1 is the successor to Fable 5 with the same input and output prices, a context window of one million tokens, and a maximum of 128,000 output tokens. Cache reads now cost only a quarter as much, at $0.25 per million tokens. Anthropic estimates the savings for typical workloads at around 25 percent, and up to 45 percent for highly agentic workflows. In its in-house test for agentic scientific work, the company reports 52.6 percent, compared to 24.7 percent for Fable 5. The reason for the price change is well-known: One month after its market launch, only 6 percent of Anthropic tokens purchased by companies were for Fable, according to the Ramps AI Index.

For existing integrations, there are three breaking changes. Forced tool use is no longer supported and will return a 400 error for tool_choice with 'any' or a specific tool, because thinking is permanently active in these models and a forced call would skip it. Thinking blocks are tied to the model that generated them, and editing earlier conversation turns invalidates them. Additionally, Anthropic is restricting when new API accounts can carry over received thinking blocks into a modified conversation, explaining that this practice has been used for large-scale distillation of its own models. Additive changes include per-message effort control, turn-based system messages, and readable progress reports between tool calls, all in beta.

In its own prompting guide, Anthropic admits that at a low effort level, Fable 5.1 calls a search or retrieval tool less often than Fable 5 and more frequently answers from memory. As a countermeasure, the company suggests a higher effort level for individual conversation turns or a special system prompt that instructs the model to search for unknown or fast-changing names before answering. According to the provider, the safety classifiers have also been made less sensitive: 60 percent fewer false positives on cyber topics, and the biology guardrails intervene 85 percent less often on harmless queries. Fable 5.1 is said to be able to find security vulnerabilities but not develop exploits for them. With the Enterprise Frontier Safeguards, the company is also announcing an architecture where monitoring data remains in customer-controlled infrastructure; it is scheduled to be rolled out later this year. Mythos 5.1 offers the same capabilities but remains exclusive to participants in the Glasswing program from cybersecurity and life sciences. → Claude, The New Stack, PCWorld, ITPro, TechRadar, startup, The Neuron

Synthszr Take: A proof of origin that triggers in comments but goes silent in the functional logic is of little value in a repository. The statistical signal needs freedom of choice between tokens, which code rarely has: where a variable name or operator is fixed, the system deliberately refrains from marking. For teams, this means a negative result proves nothing, especially with short responses. Reliable provenance in code continues to come from commit metadata, logged prompts, and signed pipelines—in other words, from what you control yourself. The EU’s transparency code is satisfied with this marking, but the question of who wrote the 300 lines in the payment module remains open.

Mentioned in this article

Search is about rankings, AI is not.

RAIDAR (may update)

Search is about rankings, AI is not.

From a ranking, you can't tell which audience sees which answer, which sources the models trust, or which areas no one has claimed yet. RAIDAR maps all of it across every model, customer segment, and market, down to the sources that feed the answers. Not a ranking. A map that tells you where to move. For brands that want to know.

More about RAIDAR →

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.