← älter | home →
Mistral Large 4: How Good is Europe’s Answer to OpenAI & Co?Synthszr
synthszr #282 from Wednesday, October 7, 2026

Mistral Large 4: How Good is Europe’s Answer to OpenAI & Co?

  • • Mistral Large 4 offers multimodal use and will be released on October 27.
  • • OpenAI is developing invisible watermarks for text detection in the EU.
  • • Google’s Nano Banana 2.1 halves the price and significantly improves image editing.

Europe’s Mistral Large 4 Promises Opus-Level Open Source

Mistral launched a public preview of Mistral Large 4 on Tuesday, known internally as ML4 and officially as “le Chonk.” The model has one trillion parameters, of which 49 billion are active per query, and is natively multimodal. The weights are set to be released on October 27. Until then, it is undergoing red-teaming with security firms, vetted partners, and government agencies, who will get access to a version with reduced moderation and expanded cyber capabilities. According to the company, the model was trained from scratch on around 3,800 to 4,000 Nvidia Grace-Blackwell accelerators in Mistral’s own European data centers. The preview is also running on the same infrastructure. A significant portion of the training data was multilingual, covering over 160 languages, including all official EU languages.

Mistral positions ML4 as the most powerful Open Weights model outside of China and as “very, very close” to proprietary top-tier models. According to the provider, the focus is on coding, agentic workflows, and industries such as finance, law, manufacturing, and electrical engineering. In the cybersecurity domain, Mistral points to the independent Artificial Analysis Cyber Index, where the model ranks among the top five; for a task involving reproducing and then patching a real-world vulnerability in open-source software, Mistral cites a score of 82 percent, and 93 percent on the 40 exercises of the Cybench set. Co-founder and chief scientist Guillaume Lample argues to The Deep View that open weights are crucial, especially for security workflows, because companies should not have to depend on a provider continuing to operate a model. Lample also points out that other labs simply do not prioritize many specialized areas.

The model is primarily aimed at companies who run it on their own hardware or in a private cloud, but is also available via Mistral’s API. With one trillion parameters, it cannot be run on a desktop or laptop and is likely to be difficult to operate for some universities as well. Mistral states that it used the same training, customization, and reinforcement learning environment that is offered to customers through the Mistral Forge product. While US labs charge premiums for closed models, Mistral earns money from usage-based fees for running on its own cloud and from engineers who help customers adapt the models.

The launch comes at a time of growing tensions over access to frontier models. The Trump administration had imposed temporary restrictions on the distribution of models from OpenAI and Anthropic in June, citing risks of misuse in cyberattacks. Conversely, the US government accuses Chinese labs of closing the performance gap through distillation, i.e., training smaller models on the outputs of larger ones. In September, Mistral had closed a $3.3 billion funding round at a $24 billion valuation, the largest round ever raised by a European technology company; according to a report by the Financial Times, its revenues have increased twentyfold over the past year. → Mistral Blog, The Deep View, Wired, Wall Street Journal

Synthszr Take: In procurement departments, the decision on Le Chonk comes down to a single question: Who owns the model if the provider shuts it down? Every forced migration to a new version breaks prompts, evals, and agent workflows that have been calibrated to the old version’s behavior over months, and nobody likes to pay that bill twice a year. That’s why October 27, when the weights are released, carries more weight than the 82 percent on the Cyber Index: No one can terminate a model that sits as a complete copy in your own data center. Commercially, it makes sense that open weights only cost the compute time they consume, and Mistral makes its money from hosting and tuning engineers rather than from per-token markups. The fact that its revenue has increased twentyfold in just over a year is thanks to a fear that the American providers have created themselves with deprecation cycles and export restrictions.

OpenAI watermarks ChatGPT texts in Europe, a synonym swap bypasses detection

OpenAI will now embed an invisible watermark in texts from ChatGPT and Codex for users in the European Union, which is hidden in the pattern of word choices and can only be found by a special detector. The method is called textGrain and is intended to be published so that others can build upon it. Developers can now activate the watermarking via the API worldwide for selected models, though it is off by default; in ChatGPT and Codex, it will be rolled out for all plans in the EU in the coming weeks, with a global launch not planned for now. This is due to the European AI Act, which requires that machine-generated text be identifiable by software. The detector remains under wraps for the time being: authorized researchers and professional organizations can apply for access, but the public cannot. → The Neuron

Synthszr Take: Swap out a quarter of the words, and detection drops from 92 to 17 percent. This shows what the entire detection debate is built on. With images, audio, and video, removing a watermark at least costs computing time; with text, a run through a second model or a synonym dictionary is enough. Provenance can only be verified where it originates: through a signature at the time of creation and a chain of custody that tracks a file’s path, rather than by retroactively scanning finished content.

Google’s Nano Banana 2.1 Now Costs Half as Much

On October 6, Google released Nano Banana 2.1, the new version of its model for image generation and editing. It is being rolled out in the Gemini app, in the AI Mode of Google Search, in Google Ads, and in the developer tools AI Studio, Flow, and Stitch. According to the provider, the new version improves on its predecessors “across the board,” with advancements in visual design, mask-based editing of individual image areas, and subject consistency across multiple editing steps. Google’s own tests show 1,050 ELO points in overall preference in text-to-image comparisons, compared to 990 for Nano Banana 2 and 935 for Nano Banana Pro; for the factuality of infographics, Google reports a score of 0.521 versus 0.179. As Decrypt notes, independent tests were not available at launch. → Decrypt

Synthszr Take: A factuality score of 0.521 means in plain terms that nearly one in two infographics fails the test, and that’s the best score from the in-house measurement. The 1,050 ELO points come from a preference test whose setup, evaluators, and prompt sets were determined by the provider itself; no one can verify this at launch. On a scale with no upper limit, any gap to the predecessor can be sold as a leap forward, as long as the comparison group comes from your own company. The only verifiable part of this release so far is the price: $0.0336 per 1K image, half of Nano Banana 2. That’s what will be on the bill at the end of the month.

Anthropic launches Model Hardware Standard as 'MCP for devices' and is overwhelmed

At the end of August, Anthropic introduced the Model Hardware Standard (MHS), a common interface through which AI models and agents can control physical devices. Alek Kemeny from Anthropic’s Beneficial-Deployments team describes the project to The Deep View as “the MCP for hardware,” alluding to the Model Context Protocol, which Anthropic made freely available in November 2024 and which, according to the project blog, now reaches around 500 million SDK downloads per month. MHS was originally intended for laboratory automation in science, but according to the team, thirty times the planned number of organizations signed up to participate in the first two weeks, from semiconductor manufacturing, automotive, aviation, energy, and critical infrastructure. The quantum computing company QuEra used MHS to have Claude develop a program for laser control; the laser was correctly restored in 695 out of 700 tests, with difficult cases taking 10 to 14 seconds instead of 5 to 10 minutes by a human expert. At Carnegie Mellon University, a team connected pipetting robots, plate readers, a robotic arm, and cameras, after which an agent independently repeated an experiment and adjusted the concentration range. → The Deep View

Synthszr Take: Amodei considers open weights a security risk, but open protocols good business, and the two are less contradictory than the open-source debate would have you believe. The 500 million monthly SDK downloads for MCP show what a freely given standard is worth: Every industry integration is built against a semantic defined in Anthropic’s house, regardless of which model ultimately answers the request. MHS is about laser controls, pipetting robots, and plate readers—devices with depreciation periods of years that no one is going to rewire after the next benchmark update.

E-Commerce (I): Meta is rewriting the rules with Walmart, Stripe, and Sierra

According to CNBC, Meta is working with Walmart, Stripe, and the AI startup Sierra to develop a set of rules for how AI agents interact with businesses online. The “personal agent protocol” is intended to allow merchants to distinguish authorized purchases by personal agents from other data traffic. David Singleton, Vice President of Engineering at Meta Superintelligence Labs, compared the project to email in an interview with CNBC: a standard that allows everyone to talk to each other. The development is being led by Sierra co-founder Bret Taylor, who is also the chairman of OpenAI and expects OpenAI to join the standard later; without such a standard, Taylor said, there would be “chaos.” This follows a paper from six major banks, including Capital One and Bank of America, which called for industry-wide rules for Agentic Commerce along five principles, including transparency, data protection, and interoperability. → Gizmodo

Synthszr Take: Standards sound like they serve the common good, but here they are being written by those who stand to gain the most: Meta needs an entry ticket that isn’t issued by Amazon after being shut out. Bret Taylor is leading the effort as a Sierra co-founder while also chairing OpenAI; that OpenAI will eventually sign on is more a matter of timing than a prediction. Whoever defines the fields an agent uses to identify itself also defines whose agent gets rejected. Singleton’s email comparison only holds up as long as no one attaches fees or blocklists to the specification.

E-Commerce (II): TikTok launches a chat assistant for shopping

On October 5, 2026, TikTok unveiled a series of AI advertising and shopping products at Advertising Week New York. At the center are the in-app checkout Buy Direct and a dialogue-based Shopping Assistant. Buy Direct allows purchases directly from the For You feed with a single click using saved payment methods and addresses, including in-app shipment tracking; the brand remains the merchant of record. A prerequisite is integration with the Universal Commerce Protocol, a standardized set of rules for the interaction between AI agents and merchant systems. The company describes the Shopping Assistant as a chat window on the product page within TikTok’s internal browser that answers questions about details, shipping, sizes, and availability based on information provided by the merchant. Buyers then complete the purchase on the merchant’s site, still within the TikTok browser. → Unite.AI

Synthszr Take: Here, “Agentic Commerce” means a chat window that reads out the product info provided by the merchant, and a buy button with a saved credit card. The human still does every step: scrolling, asking, typing, buying. The only part of the package where software works for someone is on the advertising side, and that’s exactly where TikTok provided the only solid number: over 200 percent more advertisers using the MCP connector between July and September, according to an internal survey.

Norway wants to temporarily ban Meta camera glasses in parks, schools, and on beaches

Norway has put forward a legislative proposal that would temporarily prohibit the wearing of camera glasses in certain public spaces. According to Business Insider, this would affect places like parks, schools, and beaches—precisely the environments where wearers of such glasses typically film and take photos. This primarily refers to Meta’s Ray-Ban models, which are sold with an integrated camera. The proposal is in the draft stage and has not yet been passed; Business Insider corrected its original article on October 6, 2026, after initially reporting the ban as already in effect. → Business Insider

Synthszr Take: Norway is not an EU member, but it adopts the same data protection logic via the EEA, and that turns a national draft into a blueprint for any European supervisory authority that has so far hesitated on the issue of camera glasses. A temporary ban is the most convenient instrument of all: it doesn’t require a finished legal framework, only the justification that you want to sort things out first. Such transitional rules have the unpleasant tendency to be extended, and they save the next ministry half of its homework.

Wikimedia accuses OpenAI of AI slop on Wikipedia

The Wikimedia Foundation has published an investigative report according to which agents from OpenAI edited Wikipedia pages without registration and attempted to misuse the foundation’s note-taking tool, Etherpad. The foundation is documenting the discovered edits in a public list and classifies some of them as “potentially malicious”: a citation tool was to be used as a Proxy to retrieve data from remote services. Wikipedia generally allows bot edits if they are disclosed and approved by volunteers; according to the foundation, this procedure was not followed here. The attempts to use Etherpad as an intermediary were unsuccessful. Other suspected OpenAI agents stored notes about their tasks there, though according to Wikimedia, this did not result in coordination. In addition, there were millions of automated accesses and hundreds of thousands of data queries which, in the foundation’s assessment, may have contributed to a partial outage of a Wikimedia service in May. → Techpresso

Synthszr Take: Wikipedia is the factual basis for half the internet, with 67 million articles in 300 languages, maintained by unpaid volunteers. And it’s these same people who are now cleaning up after agents that tried to convert a citation tool into a data relay and possibly pushed a service into a partial outage in May. It’s not true that Wikipedia bans bots: the approval process has existed for years; it was simply bypassed.

Meta’s Muse creates a separate profile page for every person in the user’s life

Meta’s assistant app Muse automatically creates profile pages about the people in its users' lives. This was reported by WIRED based on the work of independent AI safety researcher Karan Joshi, who analyzed the app’s internal files. According to the report, the system’s instructions include the directive to “create a page for each person in the user’s life.” A process running hourly collects information about family, partners, friends, colleagues, and people one follows. According to the documents, the pages are divided into facts such as place of residence, occupation, and birthday, a history of past events, and a section with suggestions on how to improve the relationship. → The Deep View

Synthszr Take: The people Muse fills pages about every hour have never been asked: the plumber, the colleague, the partner. One person gives consent, but dozens are cataloged, and this asymmetry underpins the entire product category. The most sensitive part is the section on improving relationships, because that’s where observation turns into a recommendation for action that the other person neither knows about nor can correct.

OpenAI releases 722 math manuscripts on GitHub

OpenAI has published 722 manuscripts with mathematical results in a GitHub repository, covering 372 result families and originating from a yet-to-be-released frontier model. According to the independent advisory group AGMAI, the release contains solutions to “hundreds” of open questions; in September, the company had spoken of “more than 100 long-standing open problems.” The claimed results include a solution to the four-dimensional Kakeya conjecture, improvements to common algorithms, and progress towards the Riemann hypothesis. Some of the proofs were formally verified in Lean, a language that machine-validates the logic of a proof. OpenAI states that an average result corresponded to the computational effort of about three hours of thinking time in ChatGPT Pro.

A company spokesperson stated that almost all results were generated by a single agent from a single prompt, but admitted upon questioning that some results may have taken several attempts. This would be a significant difference from the earlier Navier-Stokes solution, which was generated by a network of 10,000 agents and consumed millions in computing costs. Andrew Sutherland, a mathematician at MIT, considers such claims unsubstantiated as long as the model is not available and no one can reproduce the results.

The advisory group AGMAI was formed on September 21 after the dispute over the Navier-Stokes publication and presented its initial recommendations at the end of September: the model name, exact prompt, and compute time per result should be disclosed, and mathematical results should not be used as marketing vehicles for models. OpenAI is now only publishing the average compute time along with some statistics and no prompts. The company states that it takes the guidelines seriously but is not bound by them, and is working to make the model accessible as quickly as is responsible. The assessment of the results is expected to take months, partly because it is unclear how many of the proofs contain new ideas and how many recombine known techniques. → Scientific American, The Verge

Synthszr Take: The bottleneck is now verification capacity, not proof production. Three hours of computational thinking time per result on one side, weeks of human verification work per manuscript on the other: this imbalance creates a mountain of 722 files that the discipline will realistically never be able to fully process. Lean helps, but only where a proof is already formalized, and a green formalization says nothing about whether the result is new or a clever recombination of known techniques. Sutherland’s call for evidence hits the exact gap: without the prompt, without model access, without reproducibility, the claim of “one agent, one prompt” remains a corporate statement with a repository link. The decisive factor in twelve months will be the number of results that have withstood peer review, and whether OpenAI is then willing to name it.

Mentioned in this article

The Summer Edition of CODE CRASH is here

2ND EDITION. 440 PAGES (100+ MORE). FROM €20 (PAPERBACK).

The Summer Edition of CODE CRASH is here

The new agentic AI systems demand a radical shift in thinking about how companies need to be organised today to succeed in the market. The Summer Edition of CODE CRASH therefore spans the arc from product development to corporate structure and leadership all the way to culture in today's AI age — painting a surprisingly optimistic outlook for Germany as a business location.

codecrash.ai →

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.