älter | home
Musk Touts New Grok 4.6 Release: “Objectively #1”Synthszr
Apple Podcasts
Spotify
synthszr #227 from Thursday, August 13, 2026

Musk Touts New Grok 4.6 Release: “Objectively #1”

  • • Grok 4.6 surpasses Kimi K3 and scores 61 points on the AI Index
  • • Lovable secures $400 million and the EU invests
  • • Mistral plans one gigawatt of its own AI computing power in Europe by 2030

Grok 4.6 better than Kimi K3, cheaper than OpenAI

SpaceXAI, formerly xAI, released Grok 4.6 on Wednesday, less than a month after Grok 4.5. The model achieves 61 points on the Artificial Analysis Intelligence Index, a composite score from nine benchmarks, placing it ahead of the Chinese open-weights model Kimi K3 from Moonshot and on par with OpenAI’s GPT-5.6 Sol Max. Only Anthropic’s Claude Opus 5 with 63 points and Claude Fable 5 with 62 points are ahead of Grok. This is a five-point increase compared to Grok 4.5 High, which scored 56.

The API price remains at $2 per million input tokens and $6 per million output tokens; both rates double starting at 200,000 prompt tokens. According to The Decoder's calculations, this makes Grok 4.6 over 60 percent cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). The model is available immediately via the API, in Cursor, in the in-house coding tool Grok Build (part of the SuperGrok subscription starting at $30 per month), and through partners like OpenRouter, Vercel, and Cloudflare. For the first week, SpaceXAI is offering double the included quota in Cursor and Grok Build.

Technically, SpaceXAI describes a longer supplementary training run than for Grok 4.5, using curated model-generated reasoning and technical data, as well as a modified optimizer. Subsequently, they reportedly used Grok 4.5 to regenerate the trajectories for Supervised Fine-Tuning across various reasoning stages, agent harnesses, and domains, and filtered out conspicuous progressions via model-based review. The reinforcement learning targeted agentic environments such as general coding, kernel optimization, web development, and CAD. According to the provider, the model more frequently checks its own intermediate work on long task chains before proceeding; this has not been independently verified.

According to The New Stack, the individual scores are mixed. On CursorBench v3.2, Grok 4.6 achieves 69.9 percent (Grok 4.5: 66.7), ahead of GPT-5.6 Sol Max at 67.2, but behind Fable 5 Max at 70.5 percent. On DeepSWE v1.1, the score increases from 54 to 65.9 percent, while Sol Max reaches 73 and Fable 5 Max 70 percent. On Terminal-Bench v3.0, Grok improves from 15.7 to 26 percent, remaining significantly behind Sol (34.6) and Fable (34.1). On the knowledge work benchmark GDPval-AA v2, the model ranks second with an Elo score of 1,753, behind Claude Opus 5, and requires about 53 steps for complex tasks, where Opus 5 needs around 103.

Elon Musk stated on X that Grok 4.6 is 'objectively #1 when considering intelligence, speed & cost'. Gizmodo points out that the return to the top group coincides with the acquisition of the coding provider Cursor, whose real-world usage data was reportedly already incorporated into the training of Grok 4.5. A day before the model’s release, SpaceXAI launched Grok Bot, a system for persistently running agents, into beta. Gizmodo also notes that Grok has publicly drawn attention primarily for the mass generation of non-consensual nude images and that several U.S. authorities have raised security concerns. → venturebeat, thenewstack, x, gizmodo, the-decoder

Synthszr Take: 61 points is an arithmetic compromise across nine benchmarks, and this compromise smooths over what makes a difference in everyday use. On Terminal-Bench v3.0, Grok 4.6 is at 26 percent, while Sol and Fable are over 34: In a setup that works heavily in the shell, that’s a one-third deficit, which is masked in the composite score by strong numbers on CursorBench and FrontierCode. The most telling figure of the entire release isn’t even in the index: 53 steps instead of 103 on comparable GDPval tasks, because the number of steps directly impacts the bill, while an index point costs no one anything. Musk’s 'objectively #1' is the logical consequence of a single number becoming a selling point, even though it is derived from nine very different tests. If your agent is optimizing kernels, a different model is ahead than if it’s building front-end prototypes, and neither of these cases is reflected in the rankings. Running ten real tickets from your own backlog against three models and measuring the cost per completed ticket: two afternoons of work, and the answer is more reliable than any leaderboard position.

EU Fund Invests in Lovable

Stockholm-based startup Lovable has raised $400 million in a Series C round, valuing it at $13.3 billion. The platform allows users to build software using natural language descriptions, a method now known as Vibe Coding. The Wall Street Journal first reported the news; previous rumors suggested a $300 million round at a $13.2 billion valuation. Co-leading the round with Menlo Ventures is the Scaleup Europe Fund, an EU investment vehicle managed by Swedish asset manager EQT, which, according to Bloomberg, has a €5 billion pot. Lovable is one of the first disclosed investments from this fund, whose stated purpose is to prevent European growth companies from relocating to America. The company states it is approaching an annual revenue run rate of nearly $600 million by the end of the month—almost triple the figure mentioned in December, according to its own statements—and plans to increase its workforce by 50 percent to 450 people. Its clients include Nvidia, Adidas, Hearst, and Deutsche Telekom. → Techpresso

Synthszr Take: The valuation has doubled in eight months, while the revenue run rate has nearly tripled in the same period. In December, $6.6 billion on a run rate of just over $200 million was more expensive than $13.3 billion on $600 million today: the revenue multiple has dropped from around 33x to about 22x, while people are talking about a bubble. So, the price is lagging behind the business, and that is the real signal of the Vibe Coding boom: the cost of building software has fallen so far that 1.2 million new projects are being created every week, meaning things are being built that simply weren’t profitable before. The weak spot lies in the half-life of this growth. A revenue run rate is a snapshot, and no one knows how many of the 60 million projects will still be running in a year, let alone being paid for. Zendesk is both a customer and a key witness to the skepticism, with its own product director publicly stating that the reliability is not sufficient at scale. You can buy certification and a policy from Lloyd’s, but not reliable code: Whether $600 million becomes a sustainable business depends on how the generated code holds up in production.

Mistral plans to build one gigawatt of its own AI computing power in Europe by 2030

Mistral has announced plans to build up to one gigawatt of computing capacity in Europe by 2030. The company justifies the move with the goal of sovereign infrastructure, i.e., an AI foundation that does not run on the data centers of American providers. According to TLDR IT, Mistral is already seeking long-term customer commitments to secure the expansion. So far, the French company operates its models predominantly on third-party infrastructure, including Microsoft Azure, Google Cloud, and CoreWeave. In early 2026, Mistral had raised €722 million to build its own data centers with NVIDIA hardware in France and Sweden. The one-gigawatt mark now mentioned is a target for the end of the decade, not an existing level of development. → TLDR IT

Synthszr Take: One gigawatt by 2030 is an industrial policy statement in Europe, and it’s intended as such. The number sounds large until you place it next to the hundreds of billions that US providers are investing in data centers in the next two years alone. More interesting than the capacity is the sentence behind it: Mistral is seeking long-term customer commitments before the bulldozers roll. Sovereignty thus becomes a contract that European companies must sign, and that is precisely where it will be decided whether the announcement turns into concrete. As long as procurement and IT departments continue to optimize for price per token from the hyperscalers, the gigawatt will remain a press release with a footnote. For an industrial company with machine, process, or design data, a European anchor tenancy is an affordable insurance policy against political contingencies—certainly much cheaper than a later migration under time pressure. The exciting question over the next twelve months is not whether Mistral can build, but who will sign the contracts.

Gemini reaches one billion users, ChatGPT got there weeks earlier

Google CEO Sundar Pichai announced on X that the Gemini app reaches one billion people monthly, making it the fastest-growing product in the company’s history. It is the 14th Google product to cross this threshold. OpenAI had already hit the mark: the announcement that 'more than a billion people use ChatGPT' appeared in an August 6 blog post about use cases, while external data, according to Reuters, already pointed to one billion in June. OpenAI spokesperson Lindsay McCallum says that Monthly Active Users surpassed one billion some time ago, and the weekly mark was passed in July; the company does not provide a specific monthly figure. In February, OpenAI reported 900 million weekly active users, while Google reported 750 million monthly Gemini users at the same time, and then 950 million at the end of July. Gemini lead Josh Woodward wrote that the app has over 100 million active users on iOS, which places the vast majority of the user base on Android. Anthropic does not publish user numbers; estimates are in the tens of millions. → MyClaw Newsletter

Synthszr Take: Pichai drops the billion figure in a post, OpenAI tucks it into a blog post about use cases. The volume is inversely proportional to the market position, and there’s a number for that: from 900 million in February to one billion in July is about eleven percent in five months. Gemini went from 750 to 950 million in the same window. For a product that was passed around in every presentation for years as the fastest-growing software of all time, its own curve is the most unpleasant news of the summer, hence the quiet tone. However, Google’s number comes with a disclaimer: a good 100 million active users on iOS, the rest on Android, where Gemini is pre-installed right out of the box. That’s distribution disguised as demand, and it explains pretty clearly why OpenAI has been investing in its own hardware since the summer. Both are now sitting on a billion people and the same open question of who among them is actually paying.

Y Combinator releases its internal agent framework QM under MIT license

Y Combinator has released QM (short for Quartermaster) on GitHub, a framework that the accelerator had previously used internally for months. The repo describes it as a multiplayer Agent Harness for work, usable in Slack and on the web. According to YC, the system was used internally in accounting, the legal department, the event team, and in engineering, with QM itself being developed using QM. The license is MIT; according to Linas’s Newsletter, a team only incurs costs for cloud hosting and model tokens. The release came ten days after Block’s open-source release Buzz, a Nostr-based agent workspace that aims to replace Slack and GitHub, whereas QM builds on top of the existing Slack. → Linas from Linas’s Newsletter

Synthszr Take: An MIT license means you can copy, modify, and use it commercially without asking anyone for permission. For everyone not training a model or raising a round, this exposes the coordination layer where most agent projects got stuck last year: multiple humans and multiple agents in one thread, traceable, with a clear division of tasks. The fact that QM hooks into the existing Slack lowers the entry barrier to an afternoon, while Block’s Buzz requires a team to migrate its entire collaboration. The 13,000 stars measure curiosity, nothing more. The real work lies in the permissions and in the decision of which tasks to hand over to an agent in the first place, and that remains with the team.

OpenAI brings ChatGPT with Codex natively to Linux, with access to local repos

OpenAI has released a preview of the official ChatGPT desktop app for Linux, after the program was only available for Mac and Windows for years. According to AlphaSignal, it is a native application, not a browser wrapper or a command-line tool. It bundles three components: the familiar ChatGPT assistant, ChatGPT Work for team and workplace tasks, and Codex, a Coding-Agent that can read local files, work in existing repositories, and control applications on the machine. It supports Ubuntu 24.04 and 26.04 LTS, Debian 13, and Fedora 43 and 44. Installation is done via .deb or .rpm packages for x64 and ARM64, with the installer also adding an OpenAI repository so that updates flow in via apt. The download is free, but individual features require a paid plan. → AlphaSignal

Synthszr Take: The agent moves onto the machine where the source code resides, and the permission management there is straight out of 1995: a Unix user who owns everything under their home directory. At the same time, researchers are showing that the encrypted reasoning blobs from Claude, GPT, and Gemini are portable. You take a blob from Opus, move it to the cheaper Haiku, and the model dutifully reads out the content, all without breaking the encryption. A scan of about 7,000 publicly shared traces uncovered 62 API keys, 33 email addresses, and 33 passwords. The labs have made improvements after the disclosure, but the timing remains skewed: access depth grows in weeks, protection mechanisms in quarters. In practice, this means giving the agent its own system user, granting access to repositories individually, and moving credentials from project files to a secret store; that can be done in an afternoon, while the .deb installation takes a minute. An agent with read access to the entire home directory is a choice, not a default setting.

Perplexity blocks Time’s ads for AI agents and threatens publishers with trust score deduction

Perplexity has confirmed that it is blocking the ads that Time recently started delivering in the Markdown versions of its web pages—the version that AI agents retrieve instead of human readers. Digiday had reported a good two weeks earlier that Time was the first publisher to sell such agent-targeted ads. Perplexity’s head of communications, Jesse Dwyer, described the format to Digiday as deceptive and stated that publishers using “deceptive advertising like markdown ads” risk being downgraded in the company’s search index, including a deduction from their Trust Score. The company did not answer how Perplexity defines “deceptive” or how the block works technically; Steven Liss of OpenAds.AI suggested to Digiday that the agents are simply instructed not to retrieve ad blocks. The ad product itself comes from the ad-tech firm Mobian, which generates Markdown content in an FAQ format from a brand briefing, inserts it into Time’s pages, and measures how often AI search engines pick it up. The first buyers were Ally Bank and the Project Management Institute; it’s unclear if their contracts will be adjusted. Mobian CEO Jonah Goodhart told Digiday that the response to the format has been positive so far but did not provide any figures. → MyClaw Newsletter

Synthszr Take: Time saw the obvious loophole and tried to plug it. The agent reads the page, no one sees the ad banner anymore, so you sell the message to the machine instead. This only works as long as the machine plays along, and Perplexity decided in less than two weeks that it wouldn’t. The asymmetry in the reasoning is interesting: What Time sells as brand-verified, labeled information, the index unilaterally classifies as deception, with the Trust Score right there as a pressure tactic. Ally Bank and the Project Management Institute paid for visibility, Time delivered the content, and the retrieval by the agent still costs Perplexity nothing. As long as this remains the case, any monetization via the page content itself is just an attempt to bypass the gatekeeper at the checkout. The robust alternative is a negotiated price per retrieval, with a contract instead of a hidden line of text. Anything else ends in a game of cat and mouse that the publisher will lose.

Beijing counts 140 trillion tokens a day and makes it an official economic indicator

The industry newsletter Hello China Tech argues that China’s AI market can be read more accurately through token volume, public tenders, and prices than through model benchmark rankings. This follows a decision from Beijing in March: The Token, the smallest processing unit into which language models break down text, was made an official economic metric. The author cites a baseline of around 140 trillion tokens processed daily in the Chinese market. This establishes a state-reported consumption figure for AI usage, comparable to metrics for electricity or freight traffic. The author raises the question of what this accounting framework represents and what it omits, and points out that he had already made this argument in November and expanded on it in April. The article does not provide figures on the composition of the volume or the collection method. → Hello China Tech

Synthszr Take: 140 trillion tokens a day is a consumption figure, about as meaningful as kilowatt-hours or freight tons. It shows how much is running through the data centers but says nothing about whether anyone solved a problem in the process. A model that retries three times because the first two answers were useless generates three times as many tokens as one that gets it right immediately: the failed attempt looks better in the index than the clean solution. Then there’s the price dynamic: if token prices fall faster than the volume grows, the indicator rises while revenues shrink. The tenders and price lists that Hello China Tech also follows are the harder currency, as they show who is actually putting money on the table. The index links the debate to actual demand rather than test results, something benchmark rankings have so far failed to do. It will become insightful the moment Beijing breaks down the figure by industry: then it will become clear whether the tokens are flowing into manufacturing and administration or into chatbots that nobody needs.

Spotify labels AI artists as 'AI Persona' and removes them from recommendations

Starting in mid-September, Spotify will label artist profiles whose identities are artificially generated with an “AI Persona” badge and will, by default, exclude their music from editorial and algorithmic recommendations. The company announced the change on Tuesday. Artists can identify themselves as an AI Persona, but according to the provider, Spotify will not rely solely on this self-declaration but will also check profiles to see if the name and visual material indicate a photorealistically generated identity. The check will start with profiles that have exceeded predefined listener thresholds, so that the most-listened-to accounts are covered first. The badge will appear in the profile banner, the “About” section, in search, and in the track lines of playlists. Music from an AI Persona will only land in personalized recommendations if a user actively follows the profile. The rule expands on Spotify’s AI policy from September 2025, which already labels tracks via an industry standard and prohibits unauthorized Voice Cloning and deepfakes. → AI Secret

Synthszr Take: For the listener, the real question is what the absence of the badge means. According to Spotify, they first check profiles above defined listener thresholds, which conversely means that for the vast majority below that, a missing label only indicates that no one has looked at it yet. A transparency notice that is read as a statement about authenticity, but is only a statement about the verification status, creates exactly the wrong kind of trust. The second part has a harsher impact than the badge: Anyone who is dropped from editorial and algorithmic recommendations loses the only distribution channel that truly generates reach on Spotify. From that point on, concealing the origin becomes financially worthwhile. Therefore, the interesting number won’t be the number of badges awarded, but the time between the upload of a new profile and its classification. If this remains in the range of weeks, the label is like a sign on an open door. A third state in the interface would make sense: verified, unverified, AI. Three states are more honest than a badge that only speaks in one direction.

CoEvoSkills lets AI agents write and verify their own skill packages

A research team led by Hanrong Zhang and Philip S. Yu has introduced CoEvoSkills, a method that allows agents to generate so-called Skills independently instead of having them handwritten. Skills are a concept introduced by Anthropic: While a Tool is a single, self-contained function, a Skill consists of a structured bundle of interdependent files for multi-step expert tasks. The authors argue that manually creating such packages is laborious and leads to discrepancies between human and machine task interpretation, which reduces agent performance in evaluations on SkillsBench. Existing self-improvement methods are tailored to individual Tools and cannot be directly applied to Skills due to their higher complexity. CoEvoSkills therefore couples a Skill Generator, which revises the packages step-by-step, with a Surrogate Verifier that is co-developed and provides feedback without knowing the actual test content. According to the authors, the method outperforms five comparison methods on SkillsBench using both Claude Code and Codex, and transfers to six other language models. The paper has been accepted at COLM, and the code is publicly available on GitHub. → Techpresso

Synthszr Take: The interesting part is in the Verifier, not the Generator. Fine-tuning shifts weights in a black box; no one can later read in a diff what has changed. A Skill is a bundle of files: readable, versionable, deletable. This is precisely why self-improvement is operationally manageable at this level, while it remains a trust issue within the model weights. The paper’s real technical bet is that a verifier who doesn’t see the ground-truth tests can still provide useful feedback and doesn’t start to simply talk up its own generator during co-development. The fact that this holds up across six additional language models is a stronger signal than the victory over five baselines. In practical terms, this means the library of one’s own Skills becomes an asset that is maintained like code, complete with review, rollback, and ownership. When agents write their own work instructions, the engineering effort shifts to the question of what makes a Skill good, and these criteria cannot be supplied by any paper—they come from one’s own domain. In twelve months, the more exciting metric won’t be how many Skills an agent has generated, but how many of them a human has thrown out again after a glance at the repo.

Mentioned in this article

Search is about rankings, AI is not.

RAIDAR (may update)

Search is about rankings, AI is not.

From a ranking, you can't tell which audience sees which answer, which sources the models trust, or which areas no one has claimed yet. RAIDAR maps all of it across every model, customer segment, and market, down to the sources that feed the answers. Not a ranking. A map that tells you where to move. For brands that want to know.

More about RAIDAR →

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.