The Neuron Times

All the AI that's fit to print

N° 165 Édition du matinMorning EditionMorgenausgabeEdizione del mattinoEdizion del mattin · Genève DIMANCHE 14 JUIN 2026SUNDAY, 14 JUNE 2026SONNTAG, 14. JUNI 2026DOMENICA 14 GIUGNO 2026DOMENICA 14 GIUGNO 2026

À la Une · RégulationFront Page · RegulationSchlagzeilen · RegulierungPrima pagina · RegolamentazioneIn prima pagina · Regolazion

Le gouvernement américain ordonne le blocage mondial de Claude Fable 5 et Mythos 5U.S. government orders global block of Claude Fable 5 and Mythos 5US-Regierung ordnet weltweite Sperrung von Claude Fable 5 und Mythos 5 anIl governo USA ordina il blocco mondiale di Claude Fable 5 e Mythos 5El govern american el ordena el bloch global de Claude Fable 5 e Mythos 5

Une directive fédérale de contrôle des exportations force Anthropic à couper l'accès mondial à ses deux modèles les plus puissants, suscitant un débat sur la régulation des IA frontière.A federal export control directive forces Anthropic to cut global access to its two most powerful models, sparking debate over frontier AI regulation.Eine bundesstaatliche Exportkontrollverfügung zwingt Anthropic, den weltweiten Zugang zu seinen beiden leistungsfähigsten Modellen zu sperren, und entfacht eine Debatte über die Regulierung von Frontier-KI.Una direttiva federale di controllo delle esportazioni costringe Anthropic a tagliare l'accesso mondiale ai suoi due modelli più potenti, suscitando un dibattito sulla regolamentazione delle IA frontier.Ona direttiva federala de controll di esportazion la forza Anthropic a tajà l'access global ai sò du model pussee potent, fasend nass on disquàtt in su la regolazion di IA frontiera.

Anthropic a obtempéré tout en contestant publiquement la décision. « Nous ne sommes pas d'accord sur le fait qu'un jailbreak potentiel et étroit devrait justifier le retrait d'un modèle commercial déployé auprès de centaines de millions de personnes », a écrit l'entreprise dans un communiqué. La société souligne que des vulnérabilités similaires existent chez ses concurrents, notamment GPT-5.5 d'OpenAI, et prévient que cette décision pourrait créer un précédent gelant tout déploiement de modèles frontière. Tous les autres modèles Anthropic, dont Opus 4.8, restent disponibles.Anthropic complied while publicly contesting the decision. "We disagree that a potential and narrow jailbreak should justify the removal of a commercial model deployed to hundreds of millions of people," the company wrote in a statement. The company notes that similar vulnerabilities exist at its competitors, notably OpenAI's GPT-5.5, and warns that this decision could set a precedent freezing all frontier model deployments. All other Anthropic models, including Opus 4.8, remain available.Anthropic hat sich der Anordnung gefügt, sie aber öffentlich angefochten. « Wir sind nicht der Meinung, dass ein potenzieller und enger Jailbreak den Rückzug eines kommerziellen Modells rechtfertigen sollte, das bei Hunderten Millionen Menschen im Einsatz ist », schrieb das Unternehmen in einer Mitteilung. Die Firma betont, dass ähnliche Schwachstellen auch bei Konkurrenten wie GPT-5.5 von OpenAI bestünden, und warnt, dass dieser Entscheid einen Präzedenzfall schaffen könnte, der jede Bereitstellung von Frontier-Modellen einfriert. Alle anderen Anthropic-Modelle, darunter Opus 4.8, bleiben verfügbar.Anthropic ha obbedito contestando pubblicamente la decisione. « Non siamo d'accordo sul fatto che un potenziale jailbreak ristretto dovrebbe giustificare la rimozione di un modello commerciale distribuito a centinaia di milioni di persone », ha scritto l'azienda in un comunicato. La società sottolinea che vulnerabilità simili esistono presso i suoi concorrenti, in particolare GPT-5.5 di OpenAI, e avverte che questa decisione potrebbe creare un precedente congelando qualsiasi distribuzione di modelli frontier. Tutti gli altri modelli Anthropic, incluso Opus 4.8, rimangono disponibili.Anthropic l'ha obbedii, ma l'ha contestaa publicament la decision. « Nun sem minga d'acordi che on potenzial jailbreak strett el gh'ha de giustificà el retirà d'on model comercial dervii a centener de milion de personn », l'ha scrivuu l'azienda in d'on comunicaa. La società la sottolinea che di vulnerabilità simil hinn present anca in di sò concorrent, soratutt in del GPT-5.5 d'OpenAI, e la avvisa che questa decision la podaria creà on precedent che 'l gelea ogni despiegament de model frontiera. Tuti i olter model Anthropic, compres Opus 4.8, resten disponibil.

L'incident a provoqué une onde de choc dans l'industrie. Sur Hacker News, la communauté débat des implications pour la souveraineté numérique, tandis que des discussions émergent sur la création de réseaux torrent pour les modèles open source. En Inde, des responsables technologiques s'interrogent sur la dépendance du pays aux modèles américains, comme le rapporte TechCrunch. Le New York Times titre que l'administration Trump « rallume sa querelle avec Anthropic ».The incident has sent shockwaves through the industry. On Hacker News, the community debates the implications for digital sovereignty, while discussions emerge about creating torrent networks for open-source models. In India, technology officials question the country's dependence on American models, as reported by TechCrunch. The New York Times reports that the Trump administration "reignites its feud with Anthropic."Der Vorfall hat einen Schock durch die Branche geschickt. Auf Hacker News debattiert die Community über die Auswirkungen auf die digitale Souveränität, während Diskussionen über die Schaffung von Torrent-Netzwerken für Open-Source-Modelle aufkommen. In Indien fragen sich Technologieverantwortliche, wie abhängig das Land von amerikanischen Modellen ist, wie TechCrunch berichtet. Die New York Times titelt, die Trump-Administration « entfache ihren Streit mit Anthropic neu ».L'incidente ha provocato un'onda d'urto nel settore. Su Hacker News, la comunità discute le implicazioni per la sovranità digitale, mentre emergono discussioni sulla creazione di reti torrent per i modelli open source. In India, i responsabili tecnologici si interrogano sulla dipendenza del paese dai modelli americani, come riporta TechCrunch. Il New York Times titola che l'amministrazione Trump « riaccende la sua lite con Anthropic ».L'incident l'ha provocaa on'onda de s'cioch in de l'industria. Sora Hacker News, la comunità la disquàt in sui implicazion per la sovranità digitala, menter di discussion hinn dree a nass in su la creazion de red torrent per i model open source. In India, di responsabil tecnologich se interroguen in su la dipendenza del paes di model american, come 'l reporta TechCrunch. El New York Times el titola che l'aministrazion Trump « la torna a impicàss con Anthropic ».

Page 1 — Page 1 — Seite 1 — Pagina 1 — Pagina 1 — À la UneFront PageTitelgeschichteIn Primo PianoIn Prima Pagina

I. Modèles & FrontièreModels & FrontierModelle & FrontierModelli & FrontierModell & Frontiera

Benchmark

Benchmark

Benchmark

Benchmark

Benchmark

Fable 5 surpasse GPT-5.5 de 13 points sur FrontierMathFable 5 surpasses GPT-5.5 by 13 points on FrontierMathFable 5 übertrifft GPT-5.5 um 13 Punkte bei FrontierMathFable 5 supera GPT-5.5 di 13 punti su FrontierMathFable 5 el supera GPT-5.5 de 13 pont in su FrontierMath

Claude Fable 5 atteint 88 % de précision sur le niveau le plus difficile de FrontierMath, selon une analyse du Decoder publiée le 13 juin 2026. Ce score dépasse de 13 points celui de GPT-5.5 (environ 75 %) sur le même benchmark. Le bond est spectaculaire comparé à Opus 4.5, qui plafonnait sous les 10 % début 2026. Ces performances illustrent l'accélération des capacités mathématiques des modèles frontière, alors même que l'accès à Fable 5 est désormais suspendu.
Claude Fable 5 achieves 88% accuracy on the hardest level of FrontierMath, according to an analysis by The Decoder published on June 13, 2026. This score surpasses GPT-5.5 (approximately 75%) by 13 points on the same benchmark. The leap is dramatic compared to Opus 4.5, which plateaued below 10% in early 2026. These results illustrate the acceleration of mathematical capabilities in frontier models, even as access to Fable 5 is now suspended.
Claude Fable 5 erreicht 88 % Genauigkeit auf der schwierigsten Stufe von FrontierMath, wie eine Analyse des Decoder vom 13. Juni 2026 zeigt. Dieser Wert übertrifft GPT-5.5 (rund 75 %) im selben Benchmark um 13 Punkte. Der Sprung ist spektakulär im Vergleich zu Opus 4.5, das Anfang 2026 unter 10 % lag. Diese Leistungen veranschaulichen die Beschleunigung der mathematischen Fähigkeiten von Frontier-Modellen, und das, obwohl der Zugang zu Fable 5 nun ausgesetzt ist.
Claude Fable 5 raggiunge l'88% di precisione sul livello più difficile di FrontierMath, secondo un'analisi del Decoder pubblicata il 13 giugno 2026. Questo punteggio supera di 13 punti quello di GPT-5.5 (circa 75%) sullo stesso benchmark. Il balzo è spettacolare rispetto a Opus 4.5, che raggiungeva a malapena il 10% all'inizio del 2026. Queste prestazioni illustrano l'accelerazione delle capacità matematiche dei modelli frontier, proprio mentre l'accesso a Fable 5 è ora sospeso.
Claude Fable 5 el riva a 88 % de precision in sul livell pussee dificil de FrontierMath, segond on'analisi del Decoder publicada el 13 de giugn 2026. Quest score el supera de 13 pont quell de GPT-5.5 (circa 75 %) in sul midemm benchmark. El salt l'è spectacular confrontaa a Opus 4.5, che 'l se fermava sotta el 10 % al principi del 2026. Queste performance illustren l'acelerazion di capacità matematich di model frontiera, propi in del moment che l'access a Fable 5 l'è adess sospenduu.

Enquête

Investigation

Untersuchung

Indagine

Inchiesta

OpenAI visé par une enquête de procureurs généraux américainsOpenAI targeted by U.S. state attorneys general investigationOpenAI von US-Generalstaatsanwälten untersuchtOpenAI oggetto di un'indagine dei procuratori generali americaniOpenAI visaa d'on'inchiesta de procurator generai american

Un collectif de procureurs généraux d'États américains a ouvert une enquête sur OpenAI, selon un communiqué de l'entreprise rapporté par TechCrunch le 13 juin 2026. L'enquête porte sur un large éventail de pratiques : traitement des données utilisateurs, sécurité des mineurs, politiques publicitaires et gestion des données de santé. OpenAI a confirmé l'existence de l'enquête sans préciser quels États sont impliqués.
A coalition of U.S. state attorneys general has opened an investigation into OpenAI, according to a company statement reported by TechCrunch on June 13, 2026. The investigation covers a broad range of practices: user data handling, minor safety, advertising policies, and health data management. OpenAI confirmed the investigation without specifying which states are involved.
Ein Zusammenschluss von US-Generalstaatsanwälten hat eine Untersuchung gegen OpenAI eingeleitet, wie aus einer vom Unternehmen veröffentlichten Mitteilung hervorgeht, über die TechCrunch am 13. Juni 2026 berichtete. Die Untersuchung betrifft ein breites Spektrum an Praktiken: die Verarbeitung von Nutzerdaten, die Sicherheit Minderjähriger, Werberichtlinien und den Umgang mit Gesundheitsdaten. OpenAI hat die Existenz der Untersuchung bestätigt, ohne die beteiligten Bundesstaaten zu nennen.
Un collettivo di procuratori generali di stati americani ha aperto un'indagine su OpenAI, secondo un comunicato dell'azienda riportato da TechCrunch il 13 giugno 2026. L'indagine riguarda un'ampia gamma di pratiche: trattamento dei dati utente, sicurezza dei minori, politiche pubblicitarie e gestione dei dati sanitari. OpenAI ha confermato l'esistenza dell'indagine senza precisare quali stati sono coinvolti.
On coletiv de procurator generai de Stat american l'ha dervii on'inchiesta sora OpenAI, segond on comunicaa de l'azienda reportaa del TechCrunch el 13 de giugn 2026. L'inchiesta la riguarda on grand ventaj de pratich: el trattament di dat di utent, la sicurezza di minor, i politegh publicitari e la gestion di dat de salut. OpenAI l'ha confermaa l'esistenza de l'inchiesta senza precisà quai Stat hinn coinvolti.

Page 2 — Page 2 — Seite 2 — Pagina 2 — Pagina 2 — Le Cahier TechniqueThe Technical NotebookTechnisches BlattIl Quaderno TecnicoEl Carnet Tecnegh

II. Harnais, CLI & MoteursHarnesses, CLI & EnginesGeschirr, CLI & EnginesHarnais, CLI & MotoriHarnes, CLI & Motor

Moteur d'inférence

Inference engine

Inferenz-Engine

Motore di inferenza

Motor d'inferenza

llama.cpp b9628 ajoute le support SYCLllama.cpp b9628 adds SYCL supportllama.cpp b9628 fügt SYCL-Unterstützung hinzullama.cpp b9628 aggiunge il supporto SYCLllama.cpp b9628 el gionta el support SYCL

La version b9628 de llama.cpp a été publiée le 14 juin 2026. Cette mise à jour ajoute le support SYCL à l'étape de vérification des releases et propose des binaires pour macOS (Apple Silicon et Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16), Windows (CPU, CUDA 12/13, Vulkan, SYCL, HIP) et Android arm64.
Version b9628 of llama.cpp was released on June 14, 2026. This update adds SYCL support to the release verification step and offers binaries for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16), Windows (CPU, CUDA 12/13, Vulkan, SYCL, HIP), and Android arm64.
Die Version b9628 von llama.cpp wurde am 14. Juni 2026 veröffentlicht. Dieses Update fügt SYCL-Unterstützung im Release-Verifikationsschritt hinzu und bietet Binärdateien für macOS (Apple Silicon und Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16), Windows (CPU, CUDA 12/13, Vulkan, SYCL, HIP) und Android arm64.
La versione b9628 di llama.cpp è stata pubblicata il 14 giugno 2026. Questo aggiornamento aggiunge il supporto SYCL alla fase di verifica delle release e propone binari per macOS (Apple Silicon e Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16), Windows (CPU, CUDA 12/13, Vulkan, SYCL, HIP) e Android arm64.
La version b9628 de llama.cpp l'è stada publicada el 14 de giugn 2026. Questa misa a jorn la gionta el support SYCL a la fazion de verificazion di release e la propon di binari per macOS (Apple Silicon e Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL FP32/FP16), Windows (CPU, CUDA 12/13, Vulkan, SYCL, HIP) e Android arm64.

Outil agent

Agent tool

Agenten-Tool

Strumento agente

Oget agent

OpenCode v1.17.6 améliore la compatibilité MCPOpenCode v1.17.6 improves MCP compatibilityOpenCode v1.17.6 verbessert MCP-KompatibilitätOpenCode v1.17.6 migliora la compatibilità MCPOpenCode v1.17.6 el mejora la compatibilità MCP

OpenCode v1.17.6, publié le 13 juin 2026, améliore la compatibilité avec les serveurs MCP en déclarant les capacités client supportées. La mise à jour est disponible sur le dépôt GitHub du projet.
OpenCode v1.17.6, released on June 13, 2026, improves compatibility with MCP servers by declaring supported client capabilities. The update is available on the project's GitHub repository.
OpenCode v1.17.6, veröffentlicht am 13. Juni 2026, verbessert die Kompatibilität mit MCP-Servern, indem es die unterstützten Client-Fähigkeiten deklariert. Das Update ist im GitHub-Repository des Projekts verfügbar.
OpenCode v1.17.6, pubblicato il 13 giugno 2026, migliora la compatibilità con i server MCP dichiarando le capacità client supportate. L'aggiornamento è disponibile sul repository GitHub del progetto.
OpenCode v1.17.6, publicaa el 13 de giugn 2026, el mejora la compatibilità con i server MCP in del declarà i capacità client supportaa. La misa a jorn l'è disponibil in sul depòsit GitHub del proget.

Outil CLI

CLI tool

CLI-Tool

Strumento CLI

Oget CLI

Crush nightly du 13 juinCrush nightly of June 13Crush Nightly vom 13. JuniCrush nightly del 13 giugnoCrush nightly del 13 de giugn

La version nightly du 13 juin 2026 de Crush (Charmbracelet) est disponible. Cette pré-version inclut les correctifs et fonctionnalités accumulés depuis la dernière release stable.
The nightly build of June 13, 2026 for Crush (Charmbracelet) is available. This pre-release includes fixes and features accumulated since the last stable release.
Die Nightly-Version vom 13. Juni 2026 von Crush (Charmbracelet) ist verfügbar. Diese Vorabversion enthält die seit dem letzten stabilen Release angesammelten Korrekturen und Funktionen.
La versione nightly del 13 giugno 2026 di Crush (Charmbracelet) è disponibile. Questa pre-release include le correzioni e le funzionalità accumulate dall'ultima release stabile.
La version nightly del 13 de giugn 2026 de Crush (Charmbracelet) l'è disponibil. Questa pre-version la includ i corettiv e i funzionalità accumulaa de l'ultima release stabel.

Outil agent

Agent tool

Agenten-Tool

Strumento agente

Oget agent

ForgeCode v2.13.10 ajoute GLM-5.2ForgeCode v2.13.10 adds GLM-5.2ForgeCode v2.13.10 fügt GLM-5.2 hinzuForgeCode v2.13.10 aggiunge GLM-5.2ForgeCode v2.13.10 el gionta GLM-5.2

ForgeCode v2.13.10, publié le 13 juin 2026, ajoute le modèle GLM-5.2 de z.ai et des modèles Fireworks manquants à son fichier de configuration provider.json. La mise à jour corrige également des dépendances Rust, comme détaillé sur le dépôt GitHub.
ForgeCode v2.13.10, released on June 13, 2026, adds the GLM-5.2 model from z.ai and missing Fireworks models to its provider.json configuration file. The update also fixes Rust dependencies, as detailed on the GitHub repository.
ForgeCode v2.13.10, veröffentlicht am 13. Juni 2026, fügt das Modell GLM-5.2 von z.ai und fehlende Fireworks-Modelle zur Konfigurationsdatei provider.json hinzu. Das Update behebt zudem Rust-Abhängigkeiten, wie im GitHub-Repository beschrieben.
ForgeCode v2.13.10, pubblicato il 13 giugno 2026, aggiunge il modello GLM-5.2 di z.ai e i modelli Fireworks mancanti al suo file di configurazione provider.json. L'aggiornamento corregge anche dipendenze Rust, come dettagliato sul repository GitHub.
ForgeCode v2.13.10, publicaa el 13 de giugn 2026, el gionta el model GLM-5.2 de z.ai e di model Fireworks mancant al sò file de configürazion provider.json. La misa a jorn la coregg anca di dipendenze Rust, come detagliaa in sul depòsit GitHub.

Page 3 — Page 3 — Seite 3 — Pagina 3 — Pagina 3 — La RechercheResearchForschungLa RicercaLa Ricerca

III. Papers & LaboratoiresPapers & LaboratoriesPapiere & LaborePapers & LaboratoriPaper & Laboratori

Architecture

Architecture

Architektur

Architettura

Architettura

MiniMax dévoile une attention parcimonieuse pour contextes ultra-longsMiniMax unveils sparse attention for ultra-long contextsMiniMax enthüllt sparse Attention für ultra-lange KontexteMiniMax svela un'attenzione sparsa per contesti ultra-lunghiMiniMax el desvela on'attenzion parcimoniosa per contest ultra-longh

MiniMax a publié MiniMax Sparse Attention (MSA), une architecture d'attention parcimonieuse par blocs construite sur Grouped Query Attention. Un Index Branch léger sélectionne un sous-ensemble Top-k de blocs KV pour chaque groupe GQA, permettant une attention parcimonieuse spécifique au groupe. Sur un modèle de 109 milliards de paramètres, MSA réduit le calcul d'attention par token de 28,4× à 1 million de tokens de contexte, avec des accélérations de 14,2× en préfill et 7,6× en décodage sur GPU H800. Le noyau d'inférence est disponible sur GitHub.
MiniMax has published MiniMax Sparse Attention (MSA), a block-sparse attention architecture built on Grouped Query Attention. A lightweight Index Branch selects a Top-k subset of KV blocks for each GQA group, enabling group-specific sparse attention. On a 109-billion-parameter model, MSA reduces attention computation per token by 28.4× at 1 million tokens of context, with speedups of 14.2× in prefill and 7.6× in decoding on H800 GPUs. The inference kernel is available on GitHub.
MiniMax hat MiniMax Sparse Attention (MSA) veröffentlicht, eine blockweise sparse Attention-Architektur, die auf Grouped Query Attention aufbaut. Ein leichter Index Branch wählt für jede GQA-Gruppe ein Top-k-Subset von KV-Blöcken aus, was eine gruppenspezifische sparse Attention ermöglicht. Bei einem Modell mit 109 Milliarden Parametern reduziert MSA den Attention-Berechnungsaufwand pro Token um das 28,4-Fache bei 1 Million Token Kontext, mit Beschleunigungen von 14,2× beim Prefill und 7,6× beim Decoding auf H800-GPUs. Der Inferenzkern ist auf GitHub verfügbar.
MiniMax ha pubblicato MiniMax Sparse Attention (MSA), un'architettura di attenzione sparsa a blocchi costruita su Grouped Query Attention. Un Index Branch leggero seleziona un sottoinsieme Top-k di blocchi KV per ogni gruppo GQA, consentendo un'attenzione sparsa specifica per gruppo. Su un modello da 109 miliardi di parametri, MSA riduce il calcolo di attenzione per token di 28,4× a 1 milione di token di contesto, con accelerazioni di 14,2× in prefill e 7,6× in decodifica su GPU H800. Il kernel di inferenza è disponibile su GitHub.
MiniMax l'ha publicaa MiniMax Sparse Attention (MSA), ona architettura d'attenzion parcimoniosa per bloch costruida sora Grouped Query Attention. On Index Branch legger el selezziona on soto-insema Top-k de bloch KV per ogni grup GQA, permettend on'attenzion parcimoniosa specifica al grup. Sora on model de 109 miliard de parametri, MSA el ridus el calcol d'attenzion per token de 28,4× a 1 milion de token de contest, con di acelerazion de 14,2× in prefill e 7,6× in decodifega in su GPU H800. El kernel d'inferenza l'è disponibil in su GitHub.

Preuve mathématique

Mathematical proof

Mathematischer Beweis

Dimostrazione matematica

Prova matematega

MaxProof dépasse le seuil de la médaille d'or aux Olympiades de mathsMaxProof exceeds gold medal threshold at Math OlympiadMaxProof übertrifft Goldmedaillenschwelle bei Mathematik-OlympiadeMaxProof supera la soglia della medaglia d'oro alle Olimpiadi di matematicaMaxProof el supera el limit de la medaja d'or ai Olimpiad de matematega

Le framework MaxProof de MiniMax atteint 35/42 aux Olympiades Internationales de Mathématiques 2025 et 36/42 à l'USAMO 2026, dépassant le seuil de la médaille d'or humaine. MaxProof combine génération, vérification et réparation de preuves avec une recherche par tournoi au niveau population. Le modèle M3 sous-jacent intègre ces trois capacités entraînées par un vérificateur génératif à faible taux de faux positifs.
MiniMax's MaxProof framework achieves 35/42 on the 2025 International Mathematical Olympiad and 36/42 on the 2026 USAMO, surpassing the human gold medal threshold. MaxProof combines proof generation, verification, and repair with population-level tournament search. The underlying M3 model integrates these three capabilities trained by a generative verifier with a low false positive rate.
Das MaxProof-Framework von MiniMax erreicht 35/42 bei der Internationalen Mathematik-Olympiade 2025 und 36/42 bei der USAMO 2026 und übertrifft damit die Goldmedaillenschwelle menschlicher Teilnehmer. MaxProof kombiniert Generierung, Verifikation und Reparatur von Beweisen mit einer Turniersuche auf Populationsebene. Das zugrunde liegende Modell M3 integriert diese drei Fähigkeiten, die durch einen generativen Verifizierer mit niedriger Falsch-Positiv-Rate trainiert wurden.
Il framework MaxProof di MiniMax raggiunge 35/42 alle Olimpiadi Internazionali di Matematica 2025 e 36/42 all'USAMO 2026, superando la soglia della medaglia d'oro umana. MaxProof combina generazione, verifica e riparazione di dimostrazioni con una ricerca a torneo a livello di popolazione. Il modello M3 sottostante integra queste tre capacità addestrate da un verificatore generativo a basso tasso di falsi positivi.
El framework MaxProof de MiniMax el riva a 35/42 ai Olimpiad Internazional de Matematich 2025 e 36/42 a l'USAMO 2026, superand el limit de la medaja d'or umana. MaxProof el combina generazion, verificazion e reparazion de prov con ona ricerca per torneo a livell de popolazion. El model M3 soto-stant el integra queste tre capacità inzema a on verificator generativ a bass tass de fals positiv.

Agents

Agents

Agenten

Agenti

Agent

EvoArena : un benchmark pour agents en environnements dynamiquesEvoArena: a benchmark for agents in dynamic environmentsEvoArena: Ein Benchmark für Agenten in dynamischen UmgebungenEvoArena: un benchmark per agenti in ambienti dinamiciEvoArena: on benchmark per agent in ambient dinamegh

EvoArena, un benchmark développé au MIT, évalue les agents LLM dans des environnements dynamiques où les conditions évoluent progressivement. Les agents actuels n'atteignent que 39,6 % de précision moyenne. Le paradigme mémoire EvoMem, qui enregistre l'évolution sous forme d'historiques structurés, améliore la précision de 1,5 % sur EvoArena et de 6,1 % sur GAIA. L'analyse mécaniste montre qu'EvoMem améliore la capture des preuves dans la mémoire.
EvoArena, a benchmark developed at MIT, evaluates LLM agents in dynamic environments where conditions evolve progressively. Current agents achieve only 39.6% average accuracy. The EvoMem memory paradigm, which records evolution as structured histories, improves accuracy by 1.5% on EvoArena and 6.1% on GAIA. Mechanistic analysis shows EvoMem improves evidence capture in memory.
EvoArena, ein am MIT entwickelter Benchmark, bewertet LLM-Agenten in dynamischen Umgebungen, in denen sich die Bedingungen schrittweise verändern. Aktuelle Agenten erreichen nur eine durchschnittliche Genauigkeit von 39,6 %. Das Gedächtnisparadigma EvoMem, das die Entwicklung als strukturierte Verläufe aufzeichnet, verbessert die Genauigkeit um 1,5 % auf EvoArena und um 6,1 % auf GAIA. Eine mechanistische Analyse zeigt, dass EvoMem die Erfassung von Beweisen im Gedächtnis verbessert.
EvoArena, un benchmark sviluppato al MIT, valuta gli agenti LLM in ambienti dinamici dove le condizioni evolvono progressivamente. Gli agenti attuali raggiungono solo il 39,6% di precisione media. Il paradigma di memoria EvoMem, che registra l'evoluzione sotto forma di cronologie strutturate, migliora la precisione dell'1,5% su EvoArena e del 6,1% su GAIA. L'analisi meccanicistica mostra che EvoMem migliora la cattura delle prove nella memoria.
EvoArena, on benchmark desvilupaa al MIT, el valuta i agent LLM in di ambient dinamegh indove i condizion evolven progressivament. I agent d'incoeu riven domà a 39,6 % de precision media. El paradigma memoria EvoMem, che 'l registra l'evoluzion in forma de storegh strutturaa, el mejora la precision de 1,5 % in su EvoArena e de 6,1 % in su GAIA. L'analisi mecanista la mostra che EvoMem el mejora la cattura di prov in de la memoria.

Vision 3D

3D Vision

3D-Vision

Visione 3D

Vision 3D

SpatialClaw de NVIDIA : le code comme interface pour le raisonnement spatialNVIDIA's SpatialClaw: code as interface for spatial reasoningNVIDIAs SpatialClaw: Code als Schnittstelle für räumliches DenkenSpatialClaw di NVIDIA: il codice come interfaccia per il ragionamento spazialeSpatialClaw de NVIDIA: el codegh come interfaccia per el resonament spazial

SpatialClaw, développé par NVIDIA, est un framework sans entraînement pour le raisonnement spatial 3D/4D qui utilise le code comme interface d'action. Il maintient un noyau Python stateful préchargé avec des primitives de perception et de géométrie. Évalué sur 20 benchmarks, SpatialClaw atteint 59,9 % de précision moyenne, surpassant de 11,2 points l'agent spatial précédent, avec des gains constants sur six architectures VLM.
SpatialClaw, developed by NVIDIA, is a training-free framework for 3D/4D spatial reasoning that uses code as an action interface. It maintains a stateful Python kernel preloaded with perception and geometry primitives. Evaluated on 20 benchmarks, SpatialClaw achieves 59.9% average accuracy, surpassing the previous spatial agent by 11.2 points, with consistent gains across six VLM architectures.
SpatialClaw, entwickelt von NVIDIA, ist ein trainingsfreies Framework für 3D/4D-Raumdenken, das Code als Aktionsschnittstelle nutzt. Es unterhält einen vorab geladenen, zustandsbehafteten Python-Kern mit Wahrnehmungs- und Geometrie-Primitiven. Auf 20 Benchmarks evaluiert, erreicht SpatialClaw eine durchschnittliche Genauigkeit von 59,9 % und übertrifft den bisherigen räumlichen Agenten um 11,2 Punkte, mit konstanten Verbesserungen über sechs VLM-Architekturen hinweg.
SpatialClaw, sviluppato da NVIDIA, è un framework senza addestramento per il ragionamento spaziale 3D/4D che utilizza il codice come interfaccia d'azione. Mantiene un kernel Python stateful precaricato con primitive di percezione e geometria. Valutato su 20 benchmark, SpatialClaw raggiunge il 59,9% di precisione media, superando di 11,2 punti il precedente agente spaziale, con guadagni costanti su sei architetture VLM.
SpatialClaw, desvilupaa de NVIDIA, l'è on framework senza training per el resonament spazial 3D/4D che 'l dòvra el codegh come interfaccia d'azion. El manten on kernel Python stateful precaregaa con di primitiv de percezion e de geometria. Valutaa sora 20 benchmark, SpatialClaw el riva a 59,9 % de precision media, superand de 11,2 pont l'agent spazial precedent, con di guadagn costant in su ses architettur VLM.

Page 4 — Page 4 — Seite 4 — Pagina 4 — Pagina 4 — La Communauté & ÉditoCommunity & EditorialCommunity & EditorialLa Comunità & EditorialeLa Comunità & Editorial

IV. Communauté & IndustrieCommunity & IndustryCommunity & IndustrieComunità & IndustriaComunità & Industria

Recherche

Research

Forschung

Ricerca

Ricerca

Gemini-SQL2 de Google domine les benchmarks texte-SQLGoogle's Gemini-SQL2 dominates text-to-SQL benchmarksGoogles Gemini-SQL2 dominiert Text-zu-SQL-BenchmarksGemini-SQL2 di Google domina i benchmark testo-SQLGemini-SQL2 de Google el domina i benchmark text-SQL

Google Research a dévoilé Gemini-SQL2, un système de traduction langage naturel vers SQL basé sur Gemini 3.1 Pro. Il atteint 80,04 % de précision sur le benchmark BIRD, devançant largement OpenAI et Anthropic. Google indique que cette technologie pourrait améliorer les fonctionnalités en langage naturel dans ses services de données.
Google Research has unveiled Gemini-SQL2, a natural language to SQL translation system based on Gemini 3.1 Pro. It achieves 80.04% accuracy on the BIRD benchmark, far ahead of OpenAI and Anthropic. Google says this technology could improve natural language features in its data services.
Google Research hat Gemini-SQL2 vorgestellt, ein System zur Übersetzung von natürlicher Sprache in SQL, das auf Gemini 3.1 Pro basiert. Es erreicht eine Genauigkeit von 80,04 % auf dem BIRD-Benchmark und lässt OpenAI und Anthropic weit hinter sich. Google gibt an, dass diese Technologie die Funktionen für natürliche Sprache in seinen Datendiensten verbessern könnte.
Google Research ha svelato Gemini-SQL2, un sistema di traduzione da linguaggio naturale a SQL basato su Gemini 3.1 Pro. Raggiunge l'80,04% di precisione sul benchmark BIRD, superando ampiamente OpenAI e Anthropic. Google indica che questa tecnologia potrebbe migliorare le funzionalità in linguaggio naturale nei suoi servizi di dati.
Google Research l'ha desvelaa Gemini-SQL2, on sistema de traduzion lenguagg natural vers SQL fondaa in su Gemini 3.1 Pro. El riva a 80,04 % de precision in sul benchmark BIRD, lassand indree de bon longh OpenAI e Anthropic. Google l'indica che questa tecnologia la podaria migliorà i funzionalità in lenguagg natural in di sò servizzi de dat.

Optimisation

Optimization

Optimierung

Ottimizzazione

Ottimizazion

SkillOpt : un fichier Markdown pour booster les agents IASkillOpt: a Markdown file to boost AI agentsSkillOpt: Eine Markdown-Datei zur Steigerung von KI-AgentenSkillOpt: un file Markdown per potenziare gli agenti IASkillOpt: on file Markdown per potenzià i agent IA

Microsoft et trois universités chinoises ont développé SkillOpt, une méthode qui optimise les documents d'instruction pour agents IA. Un simple fichier Markdown suffit pour améliorer GPT-5.5 d'environ 23 points sur des tâches procédurales. Le même fichier se transfère entre modèles et environnements agents comme Codex et Claude Code.
Microsoft and three Chinese universities have developed SkillOpt, a method that optimizes instruction documents for AI agents. A simple Markdown file suffices to improve GPT-5.5 by approximately 23 points on procedural tasks. The same file transfers across models and agent environments such as Codex and Claude Code.
Microsoft und drei chinesische Universitäten haben SkillOpt entwickelt, eine Methode zur Optimierung von Instruktionsdokumenten für KI-Agenten. Eine einfache Markdown-Datei genügt, um GPT-5.5 bei prozeduralen Aufgaben um rund 23 Punkte zu verbessern. Dieselbe Datei ist zwischen Modellen und Agentenumgebungen wie Codex und Claude Code übertragbar.
Microsoft e tre università cinesi hanno sviluppato SkillOpt, un metodo che ottimizza i documenti di istruzione per agenti IA. Un semplice file Markdown è sufficiente per migliorare GPT-5.5 di circa 23 punti su compiti procedurali. Lo stesso file si trasferisce tra modelli e ambienti agente come Codex e Claude Code.
Microsoft e tre università cinei hann desvilupaa SkillOpt, ona metod che la ottimizza i document d'istruzion per agent IA. On semplic file Markdown el basta per migliorà GPT-5.5 de circa 23 pont in su di incarigh procedurai. El midemm file el se trasferiss in tra model e ambient agent come Codex e Claude Code.

Industrie

Industry

Industrie

Industria

Industria

Nadella admet être un « token-maxer » ; Meta gère ses coûts IANadella admits to being a "token-maxer"; Meta manages AI costsNadella gibt zu, ein « Token-Maxer » zu sein; Meta verwaltet KI-KostenNadella ammette di essere un « token-maxer »; Meta gestisce i suoi costi IANadella l'ammett de vess on « token-maxer »; Meta la gestiss i sò cost IA

Le PDG de Microsoft, Satya Nadella, a admis être lui-même un « token-maxer » — utilisateur qui jette les modèles les plus puissants sur chaque problème — lors d'une intervention rapportée par The Decoder le 13 juin 2026. Il a mis en garde contre cette pratique, soulignant que le coût marginal des gains de productivité doit correspondre au coût des tokens. Parallèlement, Meta aurait informé 6 000 employés que ses coûts internes d'IA atteignent des milliards et met en place un système de gestion des tokens via un tableau de bord central appelé « AI Gateway », comme le rapporte The Decoder.
Microsoft CEO Satya Nadella admitted to being a "token-maxer" himself — a user who throws the most powerful models at every problem — during an appearance reported by The Decoder on June 13, 2026. He warned against this practice, stressing that the marginal cost of productivity gains must match the cost of tokens. Meanwhile, Meta has reportedly informed 6,000 employees that its internal AI costs are in the billions and is implementing a token management system via a central dashboard called "AI Gateway," as reported by The Decoder.
Microsoft-CEO Satya Nadella hat eingeräumt, selbst ein « Token-Maxer » zu sein – ein Nutzer, der die leistungsfähigsten Modelle auf jedes Problem wirft –, wie The Decoder am 13. Juni 2026 berichtete. Er warnte vor dieser Praxis und betonte, dass die Grenzkosten von Produktivitätsgewinnen den Token-Kosten entsprechen müssten. Parallel dazu soll Meta 6000 Mitarbeiter informiert haben, dass die internen KI-Kosten Milliarden erreichen, und führt ein Token-Management-System über ein zentrales Dashboard namens « AI Gateway » ein, wie The Decoder berichtet.
L'amministratore delegato di Microsoft, Satya Nadella, ha ammesso di essere lui stesso un « token-maxer » — utente che lancia i modelli più potenti su ogni problema — durante un intervento riportato da The Decoder il 13 giugno 2026. Ha messo in guardia contro questa pratica, sottolineando che il costo marginale dei guadagni di produttività deve corrispondere al costo dei token. Parallelamente, Meta avrebbe informato 6.000 dipendenti che i suoi costi interni di IA raggiungono miliardi e sta implementando un sistema di gestione dei token tramite un cruscotto centrale chiamato « AI Gateway », come riporta The Decoder.
El CEO de Microsoft, Satya Nadella, l'ha ammess de vess lu midemm on « token-maxer » — utent che 'l trà i model pussee potent sora ogni problema — in d'on intervent reportaa del The Decoder el 13 de giugn 2026. L'ha avvisaa contra questa pratega, sottolineand che 'l cost marginal di guadagn de produttività el gh'ha de corrispond al cost di token. In del midemp temp, Meta la gh'ha informaa 6 000 impiegaa che i sò cost interni d'IA riven a di miliard e la mett in pee on sistema de gestion di token travers on quadrej central ciamaa « AI Gateway », come 'l reporta The Decoder.