The Neuron Times

All the AI that's fit to print

N° 2026-W39 Édition hebdomadaireWeekly EditionWochenausgabeEdizione settimanaleEdizion de la setemana · Genève SEMAINE DU 21–27 SEPTEMBRE 2026WEEK OF 21 – 27 SEPTEMBER 2026WOCHE VOM 21.–27. SEPTEMBER 2026SETTIMANA DEL 21–27 SETTEMBRE 2026SETEMANA DEL 21–27 SETTEMBRE 2026

À la Une · ÉcosystèmesFront Page · EcosystemsSchlagzeilen · ÖkosystemePrima pagina · EcosistemiIn prima pagina · Ecosistem

La semaine où les assistants sont devenus des plateformesThe week assistants became platformsDie Woche, in der Assistenten zu Plattformen wurdenLa settimana in cui gli assistenti sono diventati piattaformeLa setemana che i assistent hinn deventaa di piattaform

Plugins chez Anthropic, modèles de travail chez OpenAI, présence visuelle chez Google : la bataille s'est déplacée du modèle vers l'écosystème.Plugins at Anthropic, work models at OpenAI, visual presence at Google: the battle has shifted from the model to the ecosystem.Plugins bei Anthropic, Arbeitsmodelle bei OpenAI, visuelle Präsenz bei Google: Die Schlacht hat sich vom Modell zum Ökosystem verlagert.Plugin da Anthropic, modelli di lavoro da OpenAI, presenza visiva da Google: la battaglia si è spostata dal modello all'ecosistema.Plugin a ca' de Anthropic, modej de laoro a ca' de OpenAI, presenza visual a ca' de Google: la battaglia la s'è spostada del model vers l'ecosistema.

La leçon de la semaine ne tient pas dans un benchmark. Elle tient dans une stratégie : après des mois de course aux modèles, les grands laboratoires ont tous tiré dans la même direction — transformer l'assistant en surface d'intégration professionnelle. Anthropic ouvre Claude aux développeurs de plugins, OpenAI pousse Sol et Luna dans ChatGPT Work et Codex avec des études de cas chiffrées à l'appui, et Google DeepMind enchaîne TTS expressif et avatar temps réel pour occuper le terrain conversationnel.This week's lesson does not lie in a benchmark. It lies in a strategy: after months of racing over models, the major laboratories all pulled in the same direction — turning the assistant into a surface for professional integration. Anthropic opened Claude to plugin developers, OpenAI pushed Sol and Luna into ChatGPT Work and Codex backed by quantified case studies, and Google DeepMind rolled out expressive TTS followed by a real-time avatar to claim the conversational territory.Die Lehre dieser Woche steckt nicht in einem Benchmark, sondern in einer Strategie: Nach Monaten des Modellwettlaufs haben die grossen Labore alle in dieselbe Richtung gezogen — den Assistenten in eine Integrationsfläche für den professionellen Einsatz zu verwandeln. Anthropic öffnet Claude für Plugin-Entwickler, OpenAI bringt Sol und Luna in ChatGPT Work und Codex — flankiert von quantifizierten Fallstudien —, und Google DeepMind reiht expressives TTS und einen Echtzeit-Avatar aneinander, um das konversationelle Terrain zu besetzen.La lezione della settimana non sta in un benchmark. Sta in una strategia: dopo mesi di corsa ai modelli, i grandi laboratori hanno tutti tirato nella stessa direzione — trasformare l'assistente in superficie di integrazione professionale. Anthropic apre Claude agli sviluppatori di plugin, OpenAI spinge Sol e Luna in ChatGPT Work e Codex con studi di casi numerici a supporto, e Google DeepMind concatena TTS espressivo e avatar in tempo reale per occupare il terreno conversazionale.La lession de la setemana la sta no in d'on benchmark. La sta in d'ona strategia: dopo di mes de corsa ai modej, i grand laboratori hann tucc tiraa in la midemma direzion — trasformà l'assistént in superfis de integrazion professional. Anthropic el derva Claude ai sviluppador de plugin, OpenAI el sping Sol e Luna in ChatGPT Work e Codex con di studi de cas cifraa a sosten, e Google DeepMind el mett in fila TTS espressiv e avatar in temp real per occupà el terren conversazional.

Le mouvement s'explique par une pression économique devenue explicite. Anthropic présente Opus 5.5 comme un modèle au niveau de Fable 5.1 « sur la plupart des tâches professionnelles », pour 40 % de moins à l'exploitation qu'Opus 5. xAI promet pour Grok 4.7 une vitesse doublée à moitié prix. Quand la frontière se rapproche du bureau de travail, la différenciation ne se joue plus seulement sur la qualité brute du modèle, mais sur ce qui l'entoure : extensions, harnais, intégrations, données d'entreprise.The shift is explained by economic pressure that has become explicit. Anthropic presents Opus 5.5 as a model matching Fable 5.1 'on most professional tasks', at 40% lower operating cost than Opus 5. xAI promises Grok 4.7 with doubled speed at half the price. As the frontier moves closer to the desk, differentiation no longer hinges only on raw model quality, but on everything around it: extensions, harnesses, integrations, enterprise data.Die Bewegung erklärt sich durch einen inzwischen expliziten wirtschaftlichen Druck. Anthropic präsentiert Opus 5.5 als Modell auf dem Niveau von Fable 5.1 «bei den meisten professionellen Aufgaben», zu 40 % tieferen Betriebskosten als Opus 5. xAI verspricht für Grok 4.7 doppelt so hohe Geschwindigkeit zum halben Preis. Wenn sich die Spitze an den Arbeitsplatz heranschiebt, entscheidet die Differenzierung nicht mehr allein über die rohe Modellqualität, sondern über das Drumherum: Erweiterungen, Harnesses, Integrationen, Unternehmensdaten.Il movimento si spiega con una pressione economica diventata esplicita. Anthropic presenta Opus 5.5 come un modello al livello di Fable 5.1 « sulla maggior parte dei compiti professionali », per il 40% in meno di costi di esercizio rispetto a Opus 5. xAI promette per Grok 4.7 una velocità raddoppiata a metà prezzo. Quando la frontiera si avvicina alla postazione di lavoro, la differenziazione non si gioca più soltanto sulla qualità pura del modello, ma su ciò che lo circonda: estensioni, harness, integrazioni, dati aziendali.El moviment el se spiega con ona pression economica deventada esplicita. Anthropic el presenta Opus 5.5 'me on model al nivel de Fable 5.1 « sora la pupart di lavor professional », per el 40% de manch a l'esercizzi rispett a Opus 5. xAI el promett per Grok 4.7 ona velocità dobbiada a la metà del press. Quand la frontiera la se vesina al desch de laoro, la differenziazion la se gioega pu domà sora la qualità bruta del model, ma sora quell che gh'è intorna: estension, arnes, integrazion, dacc de impresa.

Cette course aux plateformes a aussi son revers. L'incident d'erreurs élevées sur plusieurs modèles Claude, l'enquête du Financial Times sur la fiabilité financière des chatbots ou la désignation d'Anthropic comme risque de chaîne d'approvisionnement rappellent que la confiance ne se décrète pas — elle se démontre, en public et dans la durée.This platform race also has a downside. A high-error-rate incident across several Claude models, a Financial Times investigation into chatbots' financial reliability, and Anthropic's designation as a supply-chain risk all serve as reminders that trust cannot be decreed — it must be demonstrated, publicly and over time.Dieser Plattform-Wettlauf hat auch eine Kehrseite. Der Vorfall mit erhöhten Fehlerraten bei mehreren Claude-Modellen, die Financial-Times-Recherche zur finanziellen Zuverlässigkeit von Chatbots oder die Einstufung Anthropics als Lieferkettenrisiko erinnern daran, dass Vertrauen sich nicht verordnen lässt — es wird öffentlich und über Zeit bewiesen.Questa corsa alle piattaforme ha anche il suo rovescio della medaglia. L'incidente di errori elevati su diversi modelli Claude, l'inchiesta del Financial Times sull'affidabilità finanziaria dei chatbot o la designazione di Anthropic come rischio per la catena di fornitura ricordano che la fiducia non si decreta — si dimostra, in pubblico e nel tempo.Questa corsa ai piattaform la gh'ha anca el sò revers. L'incident di error volt sora pussee modej Claude, l'indagen del Financial Times sora la fidabilità finanziaria di chatbot o la designazion de Anthropic 'me ris'c de cadena de forniment regorden che la fiducia la se decreta minga — la se demostra, in publegh e in del temp.

Les six prochains mois diront si l'ouverture de Claude aux plugins est le début d'un cycle d'applications, comme ce fut le cas pour les boutiques d'applications mobiles, ou une simple guerre de distribution entre trois ou quatre écosystèmes verrouillés. Une chose est sûre : le modèle seul n'est plus le produit.The next six months will tell whether opening Claude to plugins marks the start of an application cycle, as was the case with mobile app stores, or merely a distribution war between three or four locked-down ecosystems. One thing is certain: the model alone is no longer the product.Die nächsten sechs Monate werden zeigen, ob die Öffnung von Claude für Plugins der Beginn eines Anwendungszyklus ist, wie es die mobilen App-Stores waren, oder ein blosser Distributionskrieg zwischen drei oder vier geschlossenen Ökosystemen. Eines ist sicher: Das Modell allein ist nicht mehr das Produkt.I prossimi sei mesi diranno se l'apertura di Claude ai plugin è l'inizio di un ciclo di applicazioni, come è stato per i negozi di app mobili, o una semplice guerra di distribuzione tra tre o quattro ecosistemi chiusi. Una cosa è certa: il modello da solo non è più il prodotto.I ses mes che vegnen dirann se la vertura de Claude ai plugin l'è el principi de on ciclol de aplicazion, 'me l'è staa per i negozzi de aplicazion mobij, o domà ona guerra de distribuzion tra tri o quatter ecosistem saraa. Vuna cossa l'è segura: el model da sol el l'è pu el prodott.

Page 1 — Page 1 — Seite 1 — Pagina 1 — Pagina 1 — Rétro FrontièreFrontier RoundupFrontier-RückblickRetro FrontieraFrontiera

I.

xAI

xAI

xAI

xAI

xAI

Grok 4.7 : xAI mise sur le rapport performance-prixGrok 4.7: xAI bets on the performance-price ratioGrok 4.7: xAI setzt auf das Preis-Leistungs-VerhältnisGrok 4.7: xAI punta sul rapporto prestazioni-prezzoGrok 4.7: xAI el scommett sora el rapport prestazion-press

Le laboratoire d'Elon Musk a dévoilé Grok 4.7 avec une promesse de fiche de lancement tenue en deux chiffres : « deux fois plus rapide, à moitié prix des modèles comparables ». Le laboratoire le décrit comme son modèle le plus puissant pour le codage et le travail sur les connaissances — un positionnement frontal avec Claude et les modèles GPT orientés agents de développement. La discussion sur Hacker News a dépassé les seuils habituels d'une sortie de modèle.
Elon Musk's lab unveiled Grok 4.7 with a launch promise delivered in two figures: 'twice as fast, at half the price of comparable models'. The lab describes it as its most powerful model for coding and knowledge work — a head-on positioning against Claude and the GPT models geared toward development agents. The Hacker News discussion surpassed the usual thresholds of a model release.
Das Labor von Elon Musk hat Grok 4.7 vorgestellt, mit einem in zwei Zahlen gehaltenen Launch-Versprechen: «zweimal schneller, zum halben Preis vergleichbarer Modelle». Das Labor beschreibt es als sein stärkstes Modell für Coding und Wissensarbeit — eine frontale Positionierung gegen Claude und die auf Entwicklungsagenten ausgerichteten GPT-Modelle. Die Diskussion auf Hacker News übertraf die üblichen Schwellenwerte einer Modellveröffentlichung.
Il laboratorio di Elon Musk ha presentato Grok 4.7 con una promessa di scheda di lancio mantenuta in due cifre: « due volte più veloce, a metà prezzo dei modelli comparabili ». Il laboratorio lo descrive come il suo modello più potente per la programmazione e il lavoro sulle conoscenze — un posizionamento frontale con Claude e i modelli GPT orientati agli agenti di sviluppo. La discussione su Hacker News ha superato le soglie abituali di un lancio di modello.
El laboratori de l'Elon Musk l'ha desvelaa Grok 4.7 cont ona promessa de scheda de lancio tegnuda in du numer: « dò voeult pussee svelt, a la metà del press di modej paragonabij ». El laboratori el le descriv 'me el sò model pussee potent per el coding e 'l laoro sora i conoscenz — on posizionament frontal con Claude e i modej GPT orientaa ai agent de svilupp. La discussion sora Hacker News l'ha superaa i soej abitual de ona sortida de model.

Anthropic

Anthropic

Anthropic

Anthropic

Anthropic

Claude Opus 5.5 : Anthropic descend la frontière en gammeClaude Opus 5.5: Anthropic brings the frontier down-marketClaude Opus 5.5: Anthropic verlagert die Spitze nach unten in die PreisklasseClaude Opus 5.5: Anthropic porta la frontiera in fascia più bassaClaude Opus 5.5: Anthropic el porta giò la frontiera in gama

Anthropic a annoncé Claude Opus 5.5, qui « performe au niveau de Claude Fable 5.1 sur la plupart des tâches professionnelles » tout en coûtant « 40 % de moins à l'exploitation qu'Opus 5 ». Le positionnement est explicite : transférer les capacités du haut de gamme vers un palier de coût inférieur, là où vivent les usages professionnels. Quelques jours plus tard, le laboratoire ouvrait Claude aux plugins, complétant le mouvement du modèle vers la plateforme. Deux semaines intenses pour un laboratoire aussi confronté à une confirmation judiciaire de sa désignation comme risque de chaîne d'approvisionnement.
Anthropic announced Claude Opus 5.5, which 'performs at the level of Claude Fable 5.1 on most professional tasks' while costing '40% less to operate than Opus 5'. The positioning is explicit: transfer top-of-the-line capabilities to a lower cost tier, where professional usage lives. A few days later, the lab opened Claude to plugins, completing the move from model to platform. An intense fortnight for a lab also facing a judicial confirmation of its designation as a supply-chain risk.
Anthropic hat Claude Opus 5.5 angekündigt, das «auf dem Niveau von Claude Fable 5.1 bei den meisten professionellen Aufgaben» performt und dabei «40 % weniger im Betrieb als Opus 5» kostet. Die Positionierung ist explizit: die Fähigkeiten der Spitzenklasse auf eine tiefere Kostenstufe zu übertragen — dorthin, wo die professionelle Nutzung stattfindet. Wenige Tage später öffnete das Labor Claude für Plugins und vollendete so die Bewegung vom Modell zur Plattform. Zwei intensive Wochen für ein Labor, das zugleich eine gerichtliche Bestätigung seiner Einstufung als Lieferkettenrisiko hinnehmen musste.
Anthropic ha annunciato Claude Opus 5.5, che « rende al livello di Claude Fable 5.1 sulla maggior parte dei compiti professionali » pur costando « il 40% in meno di esercizio rispetto a Opus 5 ». Il posizionamento è esplicito: trasferire le capacità della gamma alta verso un livello di costo inferiore, là dove vivono gli usi professionali. Pochi giorni dopo, il laboratorio apriva Claude ai plugin, completando il movimento dal modello verso la piattaforma. Due settimane intense per un laboratorio anch'esso confrontato con una conferma giudiziaria della sua designazione come rischio per la catena di fornitura.
Anthropic l'ha anunziaa Claude Opus 5.5, che « 'l va al nivel de Claude Fable 5.1 sora la pupart di lavor professional » e in tant che 'l costa « el 40% de manch a l'esercizzi rispett a Opus 5 ». El posizionament l'è esplicit: trasferì i capacità de l'olta gama vers on palanch de cost inferior, là induve viven i us professional. Poch dì dopo, el laboratori 'l derveva Claude ai plugin, cont completà 'l moviment del model vers la piattaforma. Dò seteman intens per on laboratori anca lù confrontaa cont ona conferma giudiziaria de la soa designazion 'me ris'c de cadena de forniment.

OpenAI

OpenAI

OpenAI

OpenAI

OpenAI

GPT-6 Sol et Luna : OpenAI déploie deux modèles frontière dans le travail quotidienGPT-6 Sol and Luna: OpenAI deploys two frontier models into everyday workGPT-6 Sol und Luna: OpenAI rollt zwei Frontier-Modelle in die tägliche Arbeit ausGPT-6 Sol e Luna: OpenAI schiera due modelli frontiera nel lavoro quotidianoGPT-6 Sol e Luna: OpenAI el mett in camp du modej frontiera in del laoro quoridian

Publiés à 18h00 UTC via l'annonce officielle, les deux modèles visent à « apporter l'intelligence frontière au travail quotidien » dans ChatGPT Work et Codex, avec un écart de prix de 20x entre les deux paliers. OpenAI a appuyé la semaine avec des études de cas : Proaction revendique +60 % de ventes et plus de 75 heures économisées avec Codex et GPT-6 Astra. Autre fait d'armes médiatisé : GPT-6 Astra a contribué à casser un message Enigma irrésolu depuis 2005.
Released at 6:00 pm UTC via the official announcement, the two models aim to 'bring frontier intelligence to everyday work' in ChatGPT Work and Codex, with a 20x price gap between the two tiers. OpenAI backed up the week with case studies: Proaction reports +60% in sales and over 75 hours saved with Codex and GPT-6 Astra. Another widely publicized feat: GPT-6 Astra helped crack an Enigma message unsolved since 2005.
Um 18:00 UTC via der offiziellen Ankündigung veröffentlicht, zielen die beiden Modelle darauf ab, «Spitzenintelligenz in die tägliche Arbeit» zu bringen — in ChatGPT Work und Codex, mit einem Preisunterschied von 20x zwischen den beiden Stufen. OpenAI untermauerte die Woche mit Fallstudien: Proaction meldet +60 % Umsatz und über 75 eingesparte Stunden mit Codex und GPT-6 Astra. Weitere medienwirksame Leistung: GPT-6 Astra trug dazu bei, eine seit 2005 ungelöste Enigma-Nachricht zu knacken.
Pubblicati alle 18:00 UTC tramite l'annuncio ufficiale, i due modelli mirano a « portare l'intelligenza frontiera nel lavoro quotidiano » in ChatGPT Work e Codex, con un divario di prezzo di 20x tra i due livelli. OpenAI ha sostenuto la settimana con studi di casi: Proaction rivendica +60% di vendite e oltre 75 ore risparmiate con Codex e GPT-6 Astra. Altra impresa mediaticizzata: GPT-6 Astra ha contribuito a decifrare un messaggio Enigma irrisolto dal 2005.
Publicaa ai 18:00 UTC con l'annunzi offizial, i du modej voeuren « portà l'intelligenza frontiera in del laoro de tucc i dì » in ChatGPT Work e Codex, cont on descorde de press de 20x tra i du palanch. OpenAI l'ha sostegnuu la setemana cont di studi de cas: Proaction la revendega +60% de vend e pussee de 75 ore risparmiaa con Codex e GPT-6 Astra. On alter fatt armar mediatizzaa: GPT-6 Astra l'ha contribuii a romp on messagg Enigma minga risolt del 2005.

II.

Google DeepMind

Google DeepMind

Google DeepMind

Google DeepMind

Google DeepMind

Gemini 3.8 : la voix, puis le visageGemini 3.8: first the voice, then the faceGemini 3.8: erst die Stimme, dann das GesichtGemini 3.8: prima la voce, poi il voltoGemini 3.8: prima la vos, poeu la facia

Google DeepMind a enchaîné deux sorties en deux jours : d'abord les modèles text-to-speech Gemini 3.8 Flash TTS et Flash-Lite TTS, présentés comme ses modèles audio les plus expressifs à ce jour, puis Gemini 3.8 Live avec Live Avatar, qui couple cette voix expressive à un avatar animé synchronisé en quasi temps réel. Le laboratoire construit ainsi une présence conversationnelle complète, là où ses concurrents restent centrés sur le texte et le code. Google a aussi dévoilé Project Suncatcher, son projet d'infrastructures d'IA en orbite avec tests de survie du matériel spatial.
Google DeepMind strung together two releases in two days: first the Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models, presented as its most expressive audio models to date, then Gemini 3.8 Live with Live Avatar, which pairs that expressive voice with an animated avatar synchronized in near real time. The lab is thus building a complete conversational presence, where its rivals remain focused on text and code. Google also unveiled Project Suncatcher, its orbital AI infrastructure project with space hardware survival tests.
Google DeepMind hat zwei Veröffentlichungen an zwei Tagen hintereinander vorgelegt: zunächst die Text-to-Speech-Modelle Gemini 3.8 Flash TTS und Flash-Lite TTS, präsentiert als seine bisher ausdrucksvollsten Audio-Modelle, dann Gemini 3.8 Live mit Live Avatar, das diese expressive Stimme mit einem nahezu synchron in Echtzeit animierten Avatar koppelt. Das Labor baut damit eine vollständige konversationelle Präsenz auf, wo die Konkurrenten auf Text und Code zentriert bleiben. Google enthüllte zudem Project Suncatcher, sein Projekt für KI-Infrastrukturen im Orbit mit Überlebenstests der Weltraumhardware.
Google DeepMind ha concatenato due lanci in due giorni: prima i modelli text-to-speech Gemini 3.8 Flash TTS e Flash-Lite TTS, presentati come i suoi modelli audio più espressivi fino a oggi, poi Gemini 3.8 Live con Live Avatar, che abbina questa voce espressiva a un avatar animato sincronizzato in quasi tempo reale. Il laboratorio costruisce così una presenza conversazionale completa, là dove i concorrenti restano centrati sul testo e sul codice. Google ha anche svelato Project Suncatcher, il suo progetto di infrastrutture di IA in orbita con test di sopravvivenza dell'hardware spaziale.
Google DeepMind l'ha mettuu in fila dò sortid in du dì: prima i modej text-to-speech Gemini 3.8 Flash TTS e Flash-Lite TTS, presenta 'me i sò modej audio pussee espressiv fin adess, poeu Gemini 3.8 Live con Live Avatar, che 'l lega questa vos espressiva a on avatar animaa sincronizzaa in cuasi temp real. El laboratori el construiss inscì ona presenza conversazional completa, là induve i sò concorrent resten centraa sora el test e 'l codes. Google l'ha anca desvelaa Project Suncatcher, el sò progett de infrastruttur de IA in orbita cont di proeuve de sopravivenza del material spazial.

Open weights

Open weights

Open weights

Open weights

Open weights

MiMo v2.6, Hunyuan-A13B, Qwen Image 2.1 : le front chinois avance sur plusieurs tableauxMiMo v2.6, Hunyuan-A13B, Qwen Image 2.1: the Chinese front advances on several frontsMiMo v2.6, Hunyuan-A13B, Qwen Image 2.1: die chinesische Front rückt auf mehreren Brettern vorMiMo v2.6, Hunyuan-A13B, Qwen Image 2.1: il fronte cinese avanza su più tavoliMiMo v2.6, Hunyuan-A13B, Qwen Image 2.1: el front cinis el va innanz sora pussee tavoj

La Chine a placé trois pions cette semaine. Xiaomi a publié MiMo v2.6 le même jour que les grandes annonces frontière, s'invitant dans une course jusqu'ici réservée aux géants. Tencent a rendu public le rapport technique de Hunyuan-A13B : 80 milliards de paramètres, dont 13 milliards activés, en open source. Et Alibaba a relancé la course à la génération d'images open weights avec Qwen Image 2.1, dont la discussion Hacker News a dépassé 550 points et 160 commentaires en moins de 24 heures — un volume rare pour une annonce image. L'écart entre laboratoires chinois et occidentaux continue de se jouer en open weights.
China placed three pieces on the board this week. Xiaomi released MiMo v2.6 the same day as the big frontier announcements, inviting itself into a race so far reserved for the giants. Tencent published the Hunyuan-A13B technical report: 80 billion parameters, of which 13 billion active, in open source. And Alibaba reignited the open-weights image generation race with Qwen Image 2.1, whose Hacker News discussion topped 550 points and 160 comments in under 24 hours — a rare volume for an image announcement. The gap between Chinese and Western labs continues to play out in open weights.
China hat diese Woche drei Steine gesetzt. Xiaomi veröffentlichte MiMo v2.6 am selben Tag wie die grossen Frontier-Ankündigungen und drängte sich so in ein bislang den Giganten vorbehaltenes Rennen. Tencent machte den technischen Bericht zu Hunyuan-A13B öffentlich: 80 Milliarden Parameter, davon 13 Milliarden aktiviert, als Open Source. Und Alibaba belebte mit Qwen Image 2.1 den Wettlauf um die Open-Weights-Bildgenerierung neu — die Hacker-News-Diskussion übertraf in weniger als 24 Stunden 550 Punkte und 160 Kommentare, ein für eine Bildankündigung seltenes Volumen. Der Abstand zwischen chinesischen und westlichen Laboren wird weiterhin in Open Weights ausgetragen.
La Cina ha piazzato tre pedine questa settimana. Xiaomi ha pubblicato MiMo v2.6 lo stesso giorno dei grandi annunci frontiera, inserendosi in una corsa finora riservata ai giganti. Tencent ha reso pubblico il rapporto tecnico di Hunyuan-A13B: 80 miliardi di parametri, di cui 13 attivati, in open source. E Alibaba ha rilanciato la corsa alla generazione di immagini open weights con Qwen Image 2.1, la cui discussione su Hacker News ha superato 550 punti e 160 commenti in meno di 24 ore — un volume raro per un annuncio di immagine. Il divario tra laboratori cinesi e occidentali continua a giocarsi sugli open weights.
La Cina l'ha mettuu tri ponn questa setemana. Xiaomi l'ha publicaa MiMo v2.6 el midemm dì di grand annunzi de frontiera, involvendes in d'ona corsa fin adess riservada ai gigant. Tencent l'ha renduu publegh el rapport tecnegh de Hunyuan-A13B: 80 miliard de parameter, di quaj 13 miliard ativaa, in open source. E Alibaba l'ha relanziaa la corsa a la generazion de imagin open weights con Qwen Image 2.1, che la soa discussion sora Hacker News l'ha superaa 550 pont e 160 comment in manch de 24 or — on volum rar per on annunzi de imagin. El descorde tra laboratori cinis e ocidentai 'l va innanz a giocass in open weights.

Stratégie

Strategy

Strategie

Strategia

Strategia

Microsoft jette l'éponge sur le chatbot personnel avec un Copilot rebootMicrosoft throws in the towel on the personal chatbot with a Copilot rebootMicrosoft wirft beim persönlichen Chatbot das Handtuch — Copilot wird neu gestartetMicrosoft getta la spugna sul chatbot personale con un reboot di CopilotMicrosoft el tira la sponga al chatbot personal cont on reboot de Copilot

Selon Bloomberg, Microsoft abandonnerait la course au chatbot personnel avec un redémarrage complet de Copilot. Le signal est notable : à l'heure où Anthropic, OpenAI et Google se battent pour devenir la plateforme assistant par défaut, l'éditeur de Redmond choisirait de ne pas suivre. Après avoir pourtant été le premier à industrialiser un assistant grand public, Microsoft semblerait parier que la valeur se déplace ailleurs — dans les intégrations professionnelles et l'infrastructure.
According to Bloomberg, Microsoft would be dropping out of the personal chatbot race with a complete reboot of Copilot. The signal is notable: at a time when Anthropic, OpenAI and Google are fighting to become the default assistant platform, the Redmond giant appears to be choosing not to follow. Despite having been the first to industrialize a consumer assistant, Microsoft seems to be betting that value is shifting elsewhere — into professional integrations and infrastructure.
Laut Bloomberg will Microsoft das Rennen um den persönlichen Chatbot mit einem kompletten Neustart von Copilot aufgeben. Das Signal ist bemerkenswert: Während Anthropic, OpenAI und Google darum kämpfen, die Standard-Assistentenplattform zu werden, würde der Redmonder Hersteller auf ein Mitziehen verzichten. Nachdem er als Erster einen Massenmarkt-Assistenten industrialisiert hatte, scheint Microsoft zu setzen, dass sich der Wert anderswo verlagert — in professionelle Integrationen und Infrastruktur.
Secondo Bloomberg, Microsoft abbandonerebbe la corsa al chatbot personale con un riavvio completo di Copilot. Il segnale è notevole: nell'ora in cui Anthropic, OpenAI e Google si battono per diventare la piattaforma assistente predefinita, l'editore di Redmond sceglierebbe di non seguire. Pur essendo stato il primo a industrializzare un assistente per il grande pubblico, Microsoft sembrerebbe scommettere che il valore si sposta altrove — nelle integrazioni professionali e nell'infrastruttura.
Segond Bloomberg, Microsoft l'abbandonaria la corsa al chatbot personal cont on reinizzi total de Copilot. El segnal l'è notabil: a l'ora che Anthropic, OpenAI e Google se batten per deventà la piattaforma assistent de default, l'editor de Redmond el scegliria de segutà minga. Dopo vess staa tutavia el prim a industrializzà on assistent de grant publegh, Microsoft el par scommett che 'l valor el se sposta alter — in di integrazion professional e l'infrastruttura.

Page 2 — Page 2 — Seite 2 — Pagina 2 — Pagina 2 — Outils & PratiquesTools & PracticesWerkzeuge & PraxisStrumenti e praticheArnes & Pratega

III.

Retour d'expérience

Field Report

Erfahrungsbericht

Resoconto sul campo

Retor de esperienza

Quand le codage IA sature la CI : Linear raconte sa refonteWhen AI coding saturates CI: Linear tells the story of its overhaulWenn KI-Coding die CI überlastet: Linear erzählt von seinem UmbauQuando la programmazione IA satura la CI: Linear racconta la sua rifondazioneQuand el coding IA el satura la CI: Linear el conta la soa refada

Dans un retour d'expérience détaillé, Linear raconte sa refonte de CI : les agents de codage IA ont transformé son pipeline d'intégration continue en goulot d'étranglement, obligeant l'équipe à repenser l'architecture du processus. Le document fait déjà figure de référence pour toutes les équipes découvrant que la limite du codage agentique n'est plus le modèle, mais l'infrastructure qui l'entoure.
In a detailed retrospective, Linear recounts its CI overhaul: AI coding agents turned its continuous integration pipeline into a bottleneck, forcing the team to rethink the process architecture. The account is already shaping up as a reference for every team discovering that the limit of agentic coding is no longer the model, but the infrastructure around it.
In einem ausführlichen Erfahrungsbericht erzählt Linear von seinem CI-Redesign: KI-Coding-Agenten machten aus der Continuous-Integration-Pipeline einen Flaschenhals und zwangen das Team, die Prozessarchitektur neu zu denken. Das Dokument gilt bereits als Referenz für alle Teams, die feststellen, dass die Grenze des agentischen Codings nicht mehr das Modell ist, sondern die Infrastruktur drumherum.
In un resoconto dettagliato, Linear racconta la sua rifondazione della CI: gli agenti di programmazione IA hanno trasformato il suo pipeline di integrazione continua in un collo di bottiglia, obbligando il team a ripensare l'architettura del processo. Il documento figura già come riferimento per tutti i team che scoprono che il limite della programmazione agentica non è più il modello, ma l'infrastruttura che lo circonda.
In d'on retor de esperienza detailaa, Linear el conta la soa refada de CI: i agent de coding IA hann trasformaa el sò pipeline de integrazion continua in stortoeu, costring l'equipa a repensà l'architetura del process. El document l'è giamò diventaa on referiment per tucc i equip che scovren che 'l limit del coding agentich l'è pu el model, ma l'infrastruttura che gh'è intorna.

Inférence

Inference

Inferenz

Inferenza

Inferenza

tokenizers 1.0 et le pont Transformers–llama.cpptokenizers 1.0 and the Transformers–llama.cpp bridgetokenizers 1.0 und die Brücke zwischen Transformers und llama.cpptokenizers 1.0 e il ponte Transformers–llama.cpptokenizers 1.0 e 'l pont Transformers–llama.cpp

Deux pierres posées à l'édifice de l'interopérabilité : la bibliothèque tokenizers de Hugging Face est passée en version 1.0, avec des mesures de performance à l'appui, et Transformers exécute désormais nativement les quantifications llama.cpp. Ce dernier pont entre deux écosystèmes d'inférence qui s'ignoraient simplifie concrètement le déploiement de modèles quantifiés, du serveur au portable.
Two stones laid on the edifice of interoperability: the Hugging Face tokenizers library has reached version 1.0, with performance measurements to back it up, and Transformers now natively runs llama.cpp quantizations. This latest bridge between two inference ecosystems that used to ignore each other concretely simplifies deploying quantized models, from server to laptop.
Zwei Bausteine für die Interoperabilität: Die Bibliothek tokenizers von Hugging Face ist in Version 1.0 erschienen, begleitet von Leistungsmessungen, und Transformers führt nun nativ llama.cpp-Quantisierungen aus. Letztere Brücke zwischen zwei bisher einander ignorierenden Inferenz-Ökosystemen vereinfacht konkret das Deployment quantisierter Modelle, vom Server bis zum Laptop.
Due pietre posate sull'edificio dell'interoperabilità: la libreria tokenizers di Hugging Face è passata alla versione 1.0, con misurazioni di prestazioni a supporto, e Transformers esegue ormai nativamente le quantizzazioni llama.cpp. Quest'ultimo ponte tra due ecosistemi di inferenza che si ignoravano semplifica concretamente il deployment di modelli quantizzati, dal server al portatile.
Du prej poeust in su l'edifizzi de l'interoperabilità: la biblioteca tokenizers de Hugging Face l'è passada a la version 1.0, cont di misur de prestazion a sosten, e Transformers el eseguiss adess nativament i quantizzazion llama.cpp. Quest ultem pont tra du ecosistem de inferenza che seignoreven el semplifega concretament el despiegament di modej quantizzaa, del server al portabil.

Plateformes

Platforms

Plattformen

Piattaforme

Piattaform

Les Python Workers de Cloudflare passent en disponibilité généraleCloudflare's Python Workers reach general availabilityCloudflares Python Workers erreichen die allgemeine VerfügbarkeitI Python Workers di Cloudflare passano in disponibilità generaleI Python Workers de Cloudflare passen a la disponibilità general

Cloudflare a annoncé la disponibilité générale de ses Python Workers, ouvrant l'exécution de Python — le lingua franca du machine learning — à son réseau de périphérie. Pour les développeurs qui veulent rapprocher modèles et utilisateurs, l'option fait désormais partie du paysage standard, au même titre que Node.js.
Cloudflare announced the general availability of its Python Workers, opening up Python execution — the lingua franca of machine learning — to its edge network. For developers looking to bring models closer to users, the option is now part of the standard landscape, on a par with Node.js.
Cloudflare hat die allgemeine Verfügbarkeit seiner Python Workers angekündigt und öffnet damit die Ausführung von Python — der Lingua franca des Machine Learning — seinem Edge-Netzwerk. Für Entwickler, die Modelle und Nutzer einander näher bringen wollen, gehört diese Option inzwischen zur Standardlandschaft, gleichrangig mit Node.js.
Cloudflare ha annunciato la disponibilità generale dei suoi Python Workers, aprendo l'esecuzione di Python — la lingua franca del machine learning — alla sua rete periferica. Per gli sviluppatori che vogliono avvicinare modelli e utenti, l'opzione entra ormai nel panorama standard, al pari di Node.js.
Cloudflare l'ha anunziaa la disponibilità general di sò Python Workers, dervend l'esecuzion de Python — el lingua franca del machine learning — al sò red de periferia. Per i sviluppador che voeuren vesinà modej e utent, l'opzion la fa adess part del paesagg standard, 'me Node.js.

IV.

Mises en production

Production Deployments

Im Produktiveinsatz

Messa in produzione

Metr in produzion

Changer de modèle en production sans coupure : la méthode canarySwapping models in production without downtime: the canary methodModellwechsel im Produktivbetrieb ohne Unterbruch: die Canary-MethodeCambiare modello in produzione senza interruzione: il metodo canaryCambià de model in produzion senza taj: la metod canary

Dans un billet, Together AI détaille sa méthode de canary rollouts : monter en charge progressive d'un modèle en production sans coupure de service. À mesure que les mises à jour de modèles se rapprochent des mises à jour logicielles classiques, ces recettes d'exploitation deviennent un savoir-faire à part entière — et un avantage compétitif pour les plateformes d'inférence.
In a blog post, Together AI details its canary rollouts method: progressively ramping up a production model without service interruption. As model updates come to resemble classic software updates, these operations recipes are becoming a craft in their own right — and a competitive edge for inference platforms.
In einem Blogpost erläutert Together AI seine Methode der Canary-Rollouts: ein Modell im Produktivbetrieb progressiv hochfahren, ohne Serviceunterbruch. Je mehr Modell-Updates klassischen Software-Updates gleichen, werden solche Betriebsrezepte zu eigenständigem Know-how — und zu einem Wettbewerbsvorteil für Inferenz-Plattformen.
In un post, Together AI dettaglia il suo metodo di canary rollouts: messa in carico progressiva di un modello in produzione senza interruzione del servizio. Man mano che gli aggiornamenti dei modelli si avvicinano ai classici aggiornamenti software, queste ricette di esercizio diventano un know-how a tutti gli effetti — e un vantaggio competitivo per le piattaforme di inferenza.
In d'on billett, Together AI el detaja la soa metod di canary rollouts: và su de carga progressiva de on model in produzion senza tajà el servizzi. A misura che i agiornament di modej se vesinen ai agiornament software classigh, quest recett de esercizzi devenen on saoeu-fà de per lor — e on vantagg competitiv per i piattaform de inferenza.

Agents & IDE

Agents & IDEs

Agenten & IDE

Agenti e IDE

Agent & IDE

Whiteboard et Unreal Agent : les agents sortent du terminalWhiteboard and Unreal Agent: agents step out of the terminalWhiteboard und Unreal Agent: die Agenten verlassen das TerminalWhiteboard e Unreal Agent: gli agenti escono dal terminaleWhiteboard e Unreal Agent: i agent van foeu del terminal

Deux outils qui redessinent la façon de travailler avec les agents : Whiteboard, un IDE open source issu du lot YC W26 où humains et agents dessinent l'architecture ensemble (237 points sur Hacker News), et Unreal Agent, un agent dédié au moteur de jeu Unreal Engine. Signe des temps : les agents quittent le terminal pour des interfaces visuelles où la collaboration se voit.
Two tools reshaping how we work with agents: Whiteboard, an open-source IDE from the YC W26 batch where humans and agents sketch architecture together (237 points on Hacker News), and Unreal Agent, an agent dedicated to the Unreal Engine game engine. A sign of the times: agents are leaving the terminal for visual interfaces where collaboration is visible.
Zwei Werkzeuge, die die Zusammenarbeit mit Agenten neu zeichnen: Whiteboard, eine Open-Source-IDE aus dem YC-W26-Batch, in der Menschen und Agenten die Architektur gemeinsam skizzieren (237 Punkte auf Hacker News), und Unreal Agent, ein Agent für die Spiele-Engine Unreal Engine. Zeitgeist: Die Agenten verlassen das Terminal Richtung visueller Schnittstellen, in denen Zusammenarbeit sichtbar wird.
Due strumenti che ridisegnano il modo di lavorare con gli agenti: Whiteboard, un IDE open source uscito dal lotto YC W26 in cui umani e agenti disegnano insieme l'architettura (237 punti su Hacker News), e Unreal Agent, un agente dedicato al motore di gioco Unreal Engine. Segno dei tempi: gli agenti lasciano il terminale per interfacce visive in cui la collaborazione si vede.
Du arnes che redefinissen la manera de laorà con i agent: Whiteboard, on IDE open source vegnuu foeu del lott YC W26 induve omen e agent disen insema l'architetura (237 pont sora Hacker News), e Unreal Agent, on agent dedicaa al motor de gioeugh Unreal Engine. Segn di temp: i agent lassen el terminal per di interfacc visual induve la collaborazion la se ved.

Écosystème

Ecosystem

Ökosystem

Ecosistema

Ecosistema

Le créateur d'oMLX rejoint Hugging Face pour épauler la communauté MLXThe creator of oMLX joins Hugging Face to support the MLX communityDer Schöpfer von oMLX wechselt zu Hugging Face, um die MLX-Community zu unterstützenIl creatore di oMLX entra a far parte di Hugging Face per sostenere la comunità MLXEl creator d'oMLX el va insema a Hugging Face per dà ona man a la comunità MLX

Le créateur et mainteneur du moteur d'inférence MLX oMLX rejoint Hugging Face « pour soutenir la communauté MLX ». Le rapprochement consolide l'écosystème d'inférence locale sur Apple Silicon, à un moment où Apple elle-même publie des travaux de compression destinés à libérer la mémoire embarquée. L'inférence sur poste de travail n'est plus un bricolage de passionnés : elle s'industrialise.
The creator and maintainer of the MLX inference engine oMLX is joining Hugging Face 'to support the MLX community'. The move consolidates the local inference ecosystem on Apple Silicon, at a time when Apple itself is publishing compression work aimed at freeing up on-device memory. Workstation inference is no longer a hobbyist's hack: it is industrializing.
Der Schöpfer und Maintainer der MLX-Inferenz-Engine oMLX wechselt zu Hugging Face, «um die MLX-Community zu unterstützen». Die Annäherung festigt das Ökosystem der lokalen Inferenz auf Apple Silicon, in einem Moment, in dem Apple selbst Kompressionsarbeiten veröffentlicht, um den eingebetteten Speicher zu entlasten. Inferenz auf dem Arbeitsplatzrechner ist kein Bastler-Thema mehr: Sie wird industrialisiert.
Il creatore e manutentore del motore di inferenza MLX oMLX entra a far parte di Hugging Face « per sostenere la comunità MLX ». Il riavvicinamento consolida l'ecosistema di inferenza locale su Apple Silicon, in un momento in cui Apple stessa pubblica lavori di compressione destinati a liberare la memoria integrata. L'inferenza sulla postazione di lavoro non è più un arrangiamento da appassionati: si industrializza.
El creator e mantenitor del motor de inferenza MLX oMLX el va insema a Hugging Face « per sosten la comunità MLX ». El vesinament el consolida l'ecosistema de inferenza local sora Apple Silicon, in d'on moment che Apple medema la publica di lavor de compression destinà a liberà la memoria embarcaa. L'inferenza sora 'l posti de laoro l'è pu on pasticcio de appassionaa: la s'industrializza.

Protocoles

Protocols

Protokolle

Protocolli

Protocoi

MCP au banc des accusés, mais l'outillage continue de se construireMCP in the dock, but the tooling keeps being builtMCP auf der Anklagebank — doch der Werkzeugbau geht weiterMCP sul banco degli imputati, ma gli strumenti continuano a costruirsiMCP al banco di accusaa, ma l'arnes l'è adree a fass su ancamò

Une controverse a traversé la semaine : un billet intitulé « Pourquoi MCP a toujours été une mauvaise idée » a récolté 83 points et 83 commentaires sur Hacker News, tandis que des chercheurs publiaient en parallèle EvoOntology, une ontologie auto-évolutive servie via MCP pour les agents de données. Le protocole d'Anthropic est devenu assez central pour mériter ses détracteurs attitrés — sans doute le meilleur signe de son adoption.
A controversy ran through the week: a post titled 'Why MCP was always a bad idea' garnered 83 points and 83 comments on Hacker News, while researchers published EvoOntology in parallel, a self-evolving ontology served via MCP for data agents. Anthropic's protocol has become central enough to warrant its own dedicated critics — perhaps the best sign of its adoption.
Eine Kontroverse durchzog die Woche: Ein Beitrag mit dem Titel «Warum MCP von Anfang an eine schlechte Idee war» erzielte 83 Punkte und 83 Kommentare auf Hacker News, während Forschende parallel EvoOntology publizierten, eine sich selbst weiterentwickelnde Ontologie, über MCP bereitgestellt, für Datenagenten. Das Protokoll von Anthropic ist zentral genug geworden, um eingeschworene Kritiker zu verdienen — wohl das beste Zeichen seiner Verbreitung.
Una controversia ha attraversato la settimana: un post intitololato « Perché MCP è sempre stata una cattiva idea » ha raccolto 83 punti e 83 commenti su Hacker News, mentre dei ricercatori pubblicavano in parallelo EvoOntology, un'ontologia auto-evolutiva servita via MCP per gli agenti di dati. Il protocollo di Anthropic è diventato abbastanza centrale da meritarsi i suoi detrattori di rigore — senza dubbio il miglior segno della sua adozione.
Ona controversia l'ha traversaa la setemana: on billett titolaa « Perchè MCP l'è semper staa ona cativa idea » l'ha ricavaa 83 pont e 83 comment sora Hacker News, intant che di ricercador publicaven in paralel EvoOntology, vuna ontologia auto-evolutiva servida via MCP per i agent di dacc. El protocoll de Anthropic l'è deventaa assee central per merità i sò detractor titolar — fors el miglior segn de la soa adozion.

Page 3 — Page 3 — Seite 3 — Pagina 3 — Pagina 3 — RechercheResearchForschungRicercaRicerca

V.

Inférence

Inference

Inferenz

Inferenza

Inferenza

La famille des LLM à diffusion accélère sur tous les frontsThe diffusion LLM family accelerates on all frontsDie Familie der Diffusions-LLMs beschleunigt auf allen FrontenLa famiglia dei LLM a diffusione accelera su tutti i frontiLa familia di LLM a diffusion l'accelera sora tucc i front

Des chercheurs de MBZUAI publient Flash-dLLM, qui identifie les entrées redondantes dans les LLM à diffusion et annonce 11x d'accélération sans réentraînement. Dans le même temps, Apple ML Research présente le « probe guidance » pour guider ces mêmes modèles à diffusion sans surcoût, et Artificial Analysis mesure Mercury 2.5 à 770 tokens par seconde. L'architecture à diffusion, longtemps curiosité académique, s'impose comme la voie de la latence faible.
Researchers at MBZUAI publish Flash-dLLM, which identifies redundant inputs in diffusion LLMs and reports an 11x speedup without retraining. Meanwhile, Apple ML Research presents 'probe guidance' to steer those same diffusion models at no extra cost, and Artificial Analysis measures Mercury 2.5 at 770 tokens per second. The diffusion architecture, long an academic curiosity, is establishing itself as the low-latency path.
Forschende der MBZUAI veröffentlichen Flash-dLLM, das redundante Eingaben in Diffusions-LLMs identifiziert und eine 11-fache Beschleunigung ohne Nachtraining ankündigt. Zugleich präsentiert Apple ML Research die «Probe Guidance», um ebendiese Diffusionsmodelle ohne Mehraufwand zu steuern, und Artificial Analysis misst Mercury 2.5 bei 770 Tokens pro Sekunde. Die Diffusionsarchitektur, lange akademische Kuriosität, etabliert sich als Weg der niedrigen Latenz.
Dei ricercatori di MBZUAI pubblicano Flash-dLLM, che identifica gli input ridondanti nei LLM a diffusione e annuncia un'accelerazione di 11x senza riaddestramento. Nello stesso tempo, Apple ML Research presenta il « probe guidance » per guidare questi stessi modelli a diffusione senza costi aggiuntivi, e Artificial Analysis misura Mercury 2.5 a 770 token al secondo. L'architettura a diffusione, a lungo curiosità accademica, si impone come la via della bassa latenza.
Di ricercador de MBZUAI publichen Flash-dLLM, che 'l identifica i ingrèd ridondant in di LLM a diffusion e 'l nunzia 11x de accelerazion senza re-addestrament. In del midemm temp, Apple ML Research el presenta el « probe guidance » per guidà quej midemm modej a diffusion senza cost extra, e Artificial Analysis el misura Mercury 2.5 a 770 token al segond. L'architetura a diffusion, per tant temp curiosità academega, la s'impon 'me la via de la latenza bassa.

Interprétabilité

Interpretability

Interpretierbarkeit

Interpretabilità

Interpretabilità

Superposition linéaire : la preuve que les modèles superposent leurs conceptsLinear superposition: proof that models superpose their conceptsLineare Superposition: der Beweis, dass Modelle ihre Konzepte überlagernSovrapposizione lineare: la prova che i modelli sovrappongono i propri concettiSuperposizion linear: la prova che i modej superposen i sò concett

Le paper le plus plébiscité de la fin de semaine sur les Daily Papers de Hugging Face (57 votes) démontre une preuve de superposition linéaire : un Transformer peut tenir deux pensées à la fois dans ses activations. La publication donne un cadre formel à ce que les travaux d'interprétabilité observaient empiriquement depuis des années — et alimente le débat ouvert par la relecture de l'« effet LLMentalist », qui a de nouveau frappé les esprits cette semaine.
The most upvoted paper of the late week on Hugging Face's Daily Papers (57 votes) demonstrates a proof of linear superposition: a Transformer can hold two thoughts at once in its activations. The paper provides a formal framework for what interpretability work has observed empirically for years — and fuels the debate opened by the re-reading of the 'LLMentalist effect', which made waves again this week.
Das am Ende der Woche meistbejubelte Paper auf den Hugging-Face-Daily-Papers (57 Stimmen) erbringt den Nachweis der linearen Superposition: Ein Transformer kann in seinen Aktivierungen zwei Gedanken gleichzeitig halten. Die Publikation gibt dem, was Interpretierbarkeitsarbeiten seit Jahren empirisch beobachteten, einen formalen Rahmen — und nährt die Debatte, die die Neulektüre des «LLMentalist-Effekts» eröffnete, der diese Woche erneut die Geister schied.
Il paper più apprezzato di fine settimana sui Daily Papers di Hugging Face (57 voti) dimostra una prova di sovrapposizione lineare: un Transformer può tenere due pensieri allo stesso tempo nelle proprie attivazioni. La pubblicazione dà una cornice formale a ciò che i lavori di interpretabilità osservavano empiricamente da anni — e alimenta il dibattito aperto dalla rilettura dell'« effetto LLMentalist », che ha nuovamente colpito gli animi questa settimana.
El paper pussee plebiscitaa de la fin de setemana sora i Daily Papers de Hugging Face (57 vot) el demostra vuna prova de superposizion linear: on Transformer el po tegnì du penser in del midemm temp in di sò attivazion. La publicazion la dà on quadro formal a quell che i lavor de interpretabilitaa osservaven empiricament de agn — e la alimenta 'l dibattitt dervii de la rilettura de l'« effett LLMentalist », che l'ha ancamò colpii i cervell questa setemana.

Auto-amélioration

Self-improvement

Selbstverbesserung

Auto-miglioramento

Auto-migliorament

Quand les agents s'améliorent eux-mêmes : RRSI et AIDE²When agents improve themselves: RRSI and AIDE²Wenn Agenten sich selbst verbessern: RRSI und AIDE²Quando gli agenti migliorano se stessi: RRSI e AIDE²Quand i agent se miglioren de lor: RRSI e AIDE²

Deux papiers convergents sur l'auto-amélioration : une équipe de Google Research publie RRSI, qui régularise l'auto-amélioration récursive des harnais d'agents, et Weco AI présente AIDE², un agent de recherche qui réécrit son propre code et s'est amélioré pendant huit jours d'affilée. Le harnais — l'infrastructure qui entoure le modèle — devient lui-même objet d'apprentissage, une tendance que l'AgentKernel « trust-native » tente de encadrer côté sécurité.
Two converging papers on self-improvement: a Google Research team publishes RRSI, which regularizes the recursive self-improvement of agent harnesses, and Weco AI presents AIDE², a research agent that rewrites its own code and improved itself for eight days straight. The harness — the infrastructure surrounding the model — is itself becoming an object of learning, a trend the 'trust-native' AgentKernel attempts to keep in check on the safety side.
Zwei konvergierende Papers zur Selbstverbesserung: Ein Team von Google Research publiziert RRSI, das die rekursive Selbstverbesserung von Agenten-Harnesses regularisiert, und Weco AI stellt AIDE² vor, einen Forschungsagenten, der seinen eigenen Code umschreibt und sich acht Tage in Folge verbesserte. Der Harness — die Infrastruktur um das Modell — wird selbst zum Lernobjekt, eine Tendenz, die das «trust-native» AgentKernel sicherheitstechnisch zu rahmen versucht.
Due paper convergenti sull'auto-miglioramento: un team di Google Research pubblica RRSI, che regolarizza l'auto-miglioramento ricorsivo degli harness di agenti, e Weco AI presenta AIDE², un agente di ricerca che riscrive il proprio codice e si è migliorato per otto giorni di fila. L'harness — l'infrastruttura che circonda il modello — diventa esso stesso oggetto di apprendimento, una tendenza che l'AgentKernel « trust-native » tenta di inquadrare sul versante sicurezza.
Du papè convergent sora l'auto-migliorament: vuna squadra de Google Research la pubblica RRSI, che 'l regolarizza l'auto-migliorament recursiv di arnes di agent, e Weco AI el presenta AIDE², on agent de ricerca che 'l re-scriv el sò codes e l'è migliòraa per vott dì de fila. L'arnes — l'infrastruttura che gh'è intorna al model — el deventa lù medemm oggett de aprendiment, vuna tendenca che l'AgentKernel « trust-native » el proeuva a incornissà de la banda de la sicurezza.

IA & science

AI & Science

KI & Wissenschaft

IA e scienza

IA & scenza

StudentBench et la découverte : le tutorat IA au niveau humain, un Enigma cassé, une enzyme trouvéeStudentBench and discovery: human-level AI tutoring, an Enigma broken, an enzyme foundStudentBench und die Entdeckung: KI-Tutoring auf Menschenniveau, eine geknackte Enigma, ein gefundenes EnzymStudentBench e la scoperta: il tutoring IA a livello umano, un Enigma decifrato, un enzima trovatoStudentBench e la descoverta: el tutoring IA al nivel uman, on Enigma rott, on enzim descovert

Le papier StudentBench évalue le tutorat IA contre le tutorat humain au GRE : équivalence atteinte, pour 918 fois moins cher. Dans le même registre d'utilité mesurée, GPT-6 Astra a contribué à résoudre un message Enigma irrésolu depuis 2005, et Anthropic a annoncé que Claude a découvert un système enzymatique inédit à répétitions de type CRISPR — l'une des premières découvertes biologiques validées attribuées à un modèle commercial en production.
The StudentBench paper benchmarks AI tutoring against human GRE tutoring: parity achieved, at 918 times lower cost. In the same vein of measured usefulness, GPT-6 Astra contributed to solving an Enigma message unsolved since 2005, and Anthropic announced that Claude discovered a novel CRISPR-like repeat enzyme system — one of the first validated biological discoveries attributed to a commercial model in production.
Das Paper StudentBench vergleicht KI-Tutoring mit menschlichem Tutoring am GRE: Gleichwertigkeit erreicht, bei 918-mal tieferen Kosten. Im selben Register gemessenen Nutzens trug GPT-6 Astra zur Lösung einer seit 2005 ungelösten Enigma-Nachricht bei, und Anthropic gab bekannt, dass Claude ein neuartiges Enzymsystem mit CRISPR-artigen Wiederholungen entdeckt hat — eine der ersten validierten biologischen Entdeckungen, die einem kommerziellen Produktionsmodell zugeschrieben werden.
Il paper StudentBench valuta il tutoring IA contro il tutoring umano al GRE: equivalenza raggiunta, a un costo 918 volte inferiore. Nello stesso registro di utilità misurata, GPT-6 Astra ha contribuito a risolvere un messaggio Enigma irrisolto dal 2005, e Anthropic ha annunciato che Claude ha scoperto un sistema enzimatico inedito con ripetizioni di tipo CRISPR — una delle prime scoperte biologiche convalidate attribuite a un modello commerciale in produzione.
El paper StudentBench el valuta el tutoring IA contra el tutoring uman al GRE: equivalenza rivada, per 918 voeult de manch car. In del midemm registér de utilità misurada, GPT-6 Astra l'ha contribuii a resolv on messagg Enigma minga risolt del 2005, e Anthropic l'ha anunziaa che Claude l'ha descovert on sistema enzimategh inedit a ripetizion de tipo CRISPR — vuna di prim descovert biologich validaa atribuii a on model commercial in produzion.

Page 4 — Page 4 — Seite 4 — Pagina 4 — Pagina 4 — Édito hebdoWeekly EditorialWochenleitartikelEditoriale settimanaleEdito settimanal

VI. La semaine en perspectiveThe week in perspectiveDie Woche im ÜberblickLa settimana in prospettivaLa setemana in prospettiva

Édito

Editorial

Leitartikel

Editoriale

Edito

Le modèle n'est plus le produitThe model is no longer the productDas Modell ist nicht mehr das ProduktIl modello non è più il prodottoEl model l'è pu el prodott

Il y a des semaines qui se résument à une course, et d'autres qui révèlent un changement de terrain. Celle qui s'achève appartient à la seconde catégorie. Oui, les modèles se sont bousculés : Grok 4.7 « deux fois plus rapide, à moitié prix », Claude Opus 5.5 à niveau égal avec 40 % d'économie d'exploitation, GPT-6 Sol et Luna déployés dans ChatGPT Work, Gemini 3.8 qui gagne une voix expressive puis un avatar temps réel. Mais l'événement structurant est ailleurs : Anthropic a ouvert Claude aux plugins, et ce geste dit mieux que tous les benchmarks où va l'industrie.

Prenons la mesure du mouvement. Pendant deux ans, la concurrence s'est jouée sur le palier de capacité : tel modèle reasoning mieux, tel autre code plus vite. Cette semaine, trois des quatre grands laboratoires ont simultanément annoncé soit une baisse de coût à capacité constante, soit une extension de la surface produit autour du modèle. Anthropic descend son haut de gamme vers le palier professionnel. OpenAI pousse ses modèles directement dans les outils de travail et étaye le récit avec des chiffres clients — +60 % de ventes chez Proaction, plus de 75 heures économisées. Google construit une présence conversationnelle complète, voix et visage. Chacun cherche moins à avoir le meilleur modèle qu'à devenir l'endroit où le travail se fait.

Les plugins sont, historiquement, le geste qui transforme un produit en plateforme. C'est par eux que Facebook, puis les magasins d'applications mobiles, ont verrouillé des écosystèmes entiers. Ouvrir Claude aux extensions signifie qu'Anthropic accepte de dépendre de développeurs tiers — un pari sur l'effet de réseau plutôt que sur la seule supériorité du modèle. Ce pari ne paiera que si deux conditions sont réunies : une confiance suffisante pour que des entreprises confient leurs données à ces extensions, et une fiabilité à la hauteur. Or la semaine a montré les deux limites en même temps : un incident d'erreurs élevées sur plusieurs modèles Claude, et une enquête du Financial Tax— pardon, du Financial Times — rappelant que les chatbots se trompent « la plupart du temps » sur les questions financières.

C'est pourquoi le deuxième enseignement de la semaine est le retour de la confiance comme variable centrale, et pas comme slogan. OpenAI a déployé un historique de sécurité traçable dans ChatGPT et un groupe consultatif indépendant sur l'IA et les mathématiques, avec Terence Tao dans la boucle ; Sam Altman est allé plaider la coopération internationale devant le Conseil de sécurité de l'ONU. En face, une cour d'appel a confirmé la désignation d'Anthropic comme risque de chaîne d'approvisionnement, et les traces d'une attaque d'agents contre Hugging Face ont été rendues publiques. L'industrie découvre que la confiance se démontre en public — incidents listés, audits, benchmarks reproductibles comme ceux d'UK AISI et EvalEval — et non dans des communiqués.

Troisième enseignement, plus discret mais peut-être décisif : la frontière se démocratise par le bas. MiMo v2.6 chez Xiaomi, Hunyuan-A13B chez Tencent, Qwen Image 2.1 chez Alibaba : l'open weights chinois maintient une pression continue, pendant que Tim Dettmers appelle à faire tourner l'IA frontière sur du matériel ouvert et que l'écosystème d'inférence locale s'industrialise — tokenizers 1.0, le pont Transformers–llama.cpp, oMLX rejoignant Hugging Face. Quand un modèle à 80 milliards de paramètres n'en active que 13 et se publie en open source, la question « qui peut construire un assistant plateforme ? » cesse d'avoir quatre réponses.

Les six prochains mois diront si cette semaine fut le début du cycle des applications IA ou le moment où trois jardins clos ont commencé à s'observer en guerre froide. Le signe à surveiller ne sera pas le prochain benchmark, mais le premier plugin devenu incontournable — ce moment où un écosystème cesse de dépendre de son modèle fondateur. C'est exactement ce qui arrive quand une plateforme a gagné.
Some weeks come down to a race; others reveal a change of terrain. The one now ending belongs to the second category. Yes, models piled up: Grok 4.7 'twice as fast, at half the price', Claude Opus 5.5 at equal level with 40% operating savings, GPT-6 Sol and Luna deployed in ChatGPT Work, Gemini 3.8 gaining an expressive voice and then a real-time avatar. But the structural event lies elsewhere: Anthropic opened Claude to plugins, and that gesture says better than any benchmark where the industry is heading.

Let's take the measure of the shift. For two years, competition played out on the capability tier: one model reasoned better, another coded faster. This week, three of the four major labs simultaneously announced either a cost cut at constant capability, or an extension of the product surface around the model. Anthropic is bringing its high end down to the professional tier. OpenAI is pushing its models directly into work tools and backing the narrative with customer figures — +60% in sales at Proaction, more than 75 hours saved. Google is building a complete conversational presence, voice and face. Each is seeking less to have the best model than to become the place where work gets done.

Plugins are, historically, the gesture that turns a product into a platform. They are how Facebook, and later mobile app stores, locked in entire ecosystems. Opening Claude to extensions means Anthropic accepts depending on third-party developers — a bet on network effects rather than on the model's superiority alone. That bet only pays off if two conditions are met: enough trust for companies to entrust their data to these extensions, and reliability up to the task. Yet this week exposed both limits at once: a high-error-rate incident across several Claude models, and a Financial Times — not Financial Tax — investigation reminding us that chatbots get it wrong 'most of the time' on financial questions.

Which is why the week's second lesson is the return of trust as a central variable, not as a slogan. OpenAI rolled out a traceable safety history in ChatGPT and an independent advisory group on AI and mathematics, with Terence Tao in the loop; Sam Altman went to plead for international cooperation before the UN Security Council. On the other side, an appeals court confirmed Anthropic's designation as a supply-chain risk, and traces of an agent attack against Hugging Face were made public. The industry is learning that trust is demonstrated in public — incident logs, audits, reproducible benchmarks like those of UK AISI and EvalEval — not in press releases.

A third lesson, quieter but perhaps decisive: the frontier is being democratized from below. MiMo v2.6 at Xiaomi, Hunyuan-A13B at Tencent, Qwen Image 2.1 at Alibaba: Chinese open weights keep up steady pressure, while Tim Dettmers calls for running frontier AI on open hardware and the local inference ecosystem industrializes — tokenizers 1.0, the Transformers–llama.cpp bridge, oMLX joining Hugging Face. When an 80-billion-parameter model activates only 13 billion and is released in open source, the question 'who can build an assistant platform?' stops having four answers.

The next six months will tell whether this week was the start of the AI application cycle or the moment three walled gardens began eyeing each other in a cold war. The sign to watch will not be the next benchmark, but the first must-have plugin — the moment an ecosystem stops depending on its founding model. That is exactly what happens when a platform has won.
Es gibt Wochen, die sich als Wettlauf zusammenfassen lassen, und andere, die einen Terrainwechsel offenlegen. Die zu Ende gehende gehört zur zweiten Kategorie. Ja, die Modelle drängten sich: Grok 4.7 «zweimal schneller, zum halben Preis», Claude Opus 5.5 auf gleichem Niveau mit 40 % Betriebsersparnis, GPT-6 Sol und Luna in ChatGPT Work, Gemini 3.8 mit expressiver Stimme und dann Echtzeit-Avatar. Doch das strukturierende Ereignis liegt woanders: Anthropic hat Claude für Plugins geöffnet, und diese Geste sagt besser als alle Benchmarks, wohin die Industrie geht.

Messen wir die Bewegung aus. Zwei Jahre lang wurde der Wettbewerb auf der Fähigkeitsstufe ausgetragen: das eine Modell rezitiert besser, das andere codet schneller. Diese Woche haben drei der vier grossen Labore simultan entweder eine Kostensenkung bei gleicher Kapazität oder eine Erweiterung der Produktoberfläche um das Modell angekündigt. Anthropic senkt seine Spitzenklasse auf die professionelle Stufe herab. OpenAI bringt seine Modelle direkt in die Arbeitswerkzeuge und stützt die Erzählung mit Kundenzahlen — +60 % Umsatz bei Proaction, über 75 eingesparte Stunden. Google baut eine vollständige konversationelle Präsenz auf, Stimme und Gesicht. Jeder will weniger das beste Modell haben, als der Ort werden, an dem die Arbeit geschieht.

Plugins sind historisch jene Geste, die ein Produkt in eine Plattform verwandelt. Über sie haben Facebook und später die mobilen App-Stores ganze Ökosysteme eingeschlossen. Claude für Erweiterungen zu öffnen bedeutet, dass Anthropic bereit ist, von Drittanbieter-Entwicklern abzuhängen — eine Wette auf den Netzwerkeffekt statt bloss auf die Überlegenheit des Modells. Diese Wette zahlt sich nur aus, wenn zwei Bedingungen erfüllt sind: genügend Vertrauen, damit Unternehmen ihre Daten diesen Erweiterungen anvertrauen, und eine entsprechend hohe Zuverlässigkeit. Die Woche hat aber beide Grenzen gleichzeitig aufgezeigt: ein Vorfall mit erhöhten Fehlerraten bei mehreren Claude-Modellen und eine Financial-Times-Recherche — Entschuldigung, des Financial Times —, die daran erinnert, dass Chatbots bei Finanzfragen «die meiste Zeit» falsch liegen.

Deshalb ist die zweite Lehre der Woche die Rückkehr des Vertrauens als zentraler Variable — und nicht als Slogan. OpenAI hat eine rückverfolgbare Sicherheitshistorie in ChatGPT eingeführt und eine unabhängige Beratergruppe zu KI und Mathematik mit Terence Tao in der Schleife; Sam Altman plädierte vor dem UNO-Sicherheitsrat für internationale Kooperation. Auf der anderen Seite bestätigte ein Berufungsgericht die Einstufung Anthropics als Lieferkettenrisiko, und die Spuren eines Agenten-Angriffs auf Hugging Face wurden öffentlich. Die Industrie entdeckt, dass Vertrauen öffentlich bewiesen wird — mit gelisteten Vorfällen, Audits, reproduzierbaren Benchmarks wie denen von UK AISI und EvalEval — und nicht in Communiqués.

Die dritte Lehre, diskreter, aber womöglich entscheidend: Die Spitze demokratisiert sich von unten. MiMo v2.6 bei Xiaomi, Hunyuan-A13B bei Tencent, Qwen Image 2.1 bei Alibaba: Das chinesische Open-Weights-Lager hält den Druck konstant aufrecht, während Tim Dettmers dazu aufruft, Frontier-KI auf offener Hardware zu betreiben, und das Ökosystem der lokalen Inferenz industrialisiert wird — tokenizers 1.0, die Brücke zwischen Transformers und llama.cpp, oMLX bei Hugging Face. Wenn ein Modell mit 80 Milliarden Parametern nur 13 davon aktiviert und als Open Source erscheint, hört die Frage «Wer kann eine Assistenten-Plattform bauen?» auf, vier Antworten zu haben.

Die nächsten sechs Monate werden zeigen, ob diese Woche der Beginn des KI-Anwendungszyklus war oder der Moment, in dem drei geschlossene Gärten begannen, sich im Kalten Krieg zu beobachten. Das Zeichen, das es zu beobachten gilt, wird nicht der nächste Benchmark sein, sondern das erste unersetzliche Plugin — jener Moment, in dem ein Ökosystem aufhört, von seinem Gründungsmodell abzuhängen. Genau das geschieht, wenn eine Plattform gewonnen hat.
Ci sono settimane che si riassumono in una corsa, e altre che rivelano un cambio di terreno. Quella che si chiude appartiene alla seconda categoria. Sì, i modelli si sono accavallati: Grok 4.7 « due volte più veloce, a metà prezzo », Claude Opus 5.5 a livello pari con il 40% di risparmio di esercizio, GPT-6 Sol e Luna schierati in ChatGPT Work, Gemini 3.8 che guadagna una voce espressiva poi un avatar in tempo reale. Ma l'evento strutturante è altrove: Anthropic ha aperto Claude ai plugin, e questo gesto dice meglio di tutti i benchmark dove va l'industria.

Misuriamo il movimento. Per due anni, la concorrenza si è giocata sul livello di capacità: tale modello ragiona meglio, tale altro programma più velocemente. Questa settimana, tre dei quattro grandi laboratori hanno annunciato simultaneamente o un ribasso di costo a capacità costante, o un'estensione della superficie di prodotto attorno al modello. Anthropic porta la sua gamma alta verso il livello professionale. OpenAI spinge i propri modelli direttamente negli strumenti di lavoro e sorregge il racconto con cifre dei clienti — +60% di vendite da Proaction, oltre 75 ore risparmiate. Google costruisce una presenza conversazionale completa, voce e volto. Ciascuno cerca meno di avere il miglior modello che di diventare il luogo dove il lavoro si svolge.

I plugin sono, storicamente, il gesto che trasforma un prodotto in piattaforma. È tramite essi che Facebook, poi i negozi di app mobili, hanno chiuso interi ecosistemi. Aprire Claude alle estensioni significa che Anthropic accetta di dipendere da sviluppatori terzi — una scommessa sull'effetto rete più che sulla sola superiorità del modello. Questa scommessa pagherà solo se due condizioni saranno riunite: una fiducia sufficiente perché delle aziende affidino i propri dati a queste estensioni, e un'affidabilità all'altezza. Or la settimana ha mostrato i due limiti allo stesso tempo: un incidente di errori elevati su diversi modelli Claude, e un'inchiesta del Financial Tax— scusate, del Financial Times — che ricorda che i chatbot sbagliano « la maggior parte delle volte » sulle questioni finanziarie.

È per questo che il secondo insegnamento della settimana è il ritorno della fiducia come variabile centrale, e non come slogan. OpenAI ha schierato uno storico della sicurezza tracciabile in ChatGPT e un gruppo consultivo indipendente su IA e matematica, con Terence Tao nel loop; Sam Altman è andato a difendere la cooperazione internazionale davanti al Consiglio di sicurezza dell'ONU. Di fronte, una corte d'appello ha confermato la designazione di Anthropic come rischio per la catena di fornitura, e le tracce di un attacco di agenti contro Hugging Face sono state rese pubbliche. L'industria scopre che la fiducia si dimostra in pubblico — incidenti elencati, audit, benchmark riproducibili come quelli di UK AISI ed EvalEval — e non nei comunicati stampa.

Terzo insegnamento, più discreto ma forse decisivo: la frontiera si democratizza dal basso. MiMo v2.6 da Xiaomi, Hunyuan-A13B da Tencent, Qwen Image 2.1 da Alibaba: gli open weights cinesi mantengono una pressione continua, mentre Tim Dettmers chiama a far girare l'IA frontiera su hardware aperto e l'ecosistema di inferenza locale si industrializza — tokenizers 1.0, il ponte Transformers–llama.cpp, oMLX che entra in Hugging Face. Quando un modello da 80 miliardi di parametri ne attiva solo 13 e si pubblica in open source, la domanda « chi può costruire una piattaforma assistente? » smette di avere quattro risposte.

I prossimi sei mesi diranno se questa settimana è stata l'inizio del ciclo delle applicazioni IA o il momento in cui tre giardini recintati hanno cominciato a osservarsi in guerra fredda. Il segno da sorvegliare non sarà il prossimo benchmark, ma il primo plugin diventato imprescindibile — quel momento in cui un ecosistema smette di dipendere dal proprio modello fondatore. È esattamente ciò che accade quando una piattaforma ha vinto.
Gh'è di seteman che se resumen in d'ona corsa, e di alter che revelen on cambi de terren. Quella che la finiss la partegn a la seconda categoria. Sì, i modej s'hinn ingropaa: Grok 4.7 « dò voeult pussee svelt, a la metà del press », Claude Opus 5.5 a nivel pari cont el 40% de risparm d'esercizzi, GPT-6 Sol e Luna mettuu in ChatGPT Work, Gemini 3.8 che 'l guadagna vuna vos espressiva poeu on avatar in temp real. Ma l'event struturant l'è alter: Anthropic l'ha dervii Claude ai plugin, e quest gest el dis mej de tucc i benchmark induve che 'l va l'industria.

Prenemm la misura del moviment. Per du agn, la concorrenza la s'è giocada sora 'l palanch de capacità: on cert model 'l reasoning mej, on alter 'l codesa pussee svelt. Questa setemana, tri di quatter grand laboratori hann anunziaa in del midemm moment o ona sbassada de cost a capacità costanta, o vuna estension de la superfis prodott intorna al model. Anthropic el porta giò la soa olta gama vers el palanch professional. OpenAI el sping i sò modej diretament in di arnes de laoro e 'l sostegn el raccont cont di numer client — +60% de vend de Proaction, pussee de 75 ore risparmiaa. Google el costruiss ona presenza conversazional completa, vos e facia. Tucc e cercaven minga de avègh el miglior model ma de deventà 'l loeugh induve che 'l laoro 'l se fa.

I plugin hinn, storegament, el gest che 'l trasforma on prodott in piattaforma. L'è cont lor che Facebook, poeu i magasin de aplicazion mobij, hann saraa di ecosistem intregh. Dervì Claude ai estension el voeur dì che Anthropic l'acetta de dipend de sviluppador terz — ona scommessa sora l'effett de red plutost che sora la sola superiorità del model. Questa scommessa la pagherà domà se dò condizion hinn giontaa: ona fiducia assée perchè di impres afiden i sò dacc a quest estension, e vuna fidabilità a l'altessa. Ma la setemana l'ha mostraa i du limit in del midemm temp: on incident de error volt sora pussee modej Claude, e vuna indagen del Financial Tax— pardon, del Financial Times — che la regorda che i chatbot i se sbalien « la pupart del temp » sora i question finanziari.

L'è per quest che 'l segond insegnament de la setemana l'è 'l retorn de la fiducia 'me variabil central, e no 'me slogan. OpenAI l'ha despiegaa on storegh de sicurezza tracciabij in ChatGPT e on grup consultiv independent sora l'IA e la matematega, con Terence Tao in del gir; Sam Altman l'è andaa a difend la cooperazion internazional devant al Consili de sicurezza de l'ONU. De la banda opposta, vuna corte d'appell l'ha confermaa la designazion de Anthropic 'me ris'c de cadena de forniment, e i tracc de on attacch de agent contra Hugging Face hinn staa renduu publegh. L'industria la scovr che la fiducia la se demostra in publegh — incident elencaa, audit, benchmark riproducibij 'me quij de UK AISI ed EvalEval — e no in di comunegaa.

Terz insegnament, pussee discret ma fors decisiv: la frontiera la se democratizza del bass. MiMo v2.6 de Xiaomi, Hunyuan-A13B de Tencent, Qwen Image 2.1 de Alibaba: l'open weights cinis el mantegn vuna pression continua, intant che Tim Dettmers el ciamma a fà andà l'IA frontiera sora material dervii e che l'ecosistema de inferenza local el s'industrializza — tokenizers 1.0, el pont Transformers–llama.cpp, oMLX che 'l va insema a Hugging Face. Quand on model de 80 miliard de parameter en ativa domà 13 e 'l se publica in open source, la question « chi el po costruì on assistent piattaforma? » la smett de avègh quatter rispost.

I ses mes che vegnen dirann se questa setemana l'è stada el principi del ciclol di aplicazion IA o el moment che tri giardin saraa hann comenzà a vardass in guerra freggia. El segn de varda 'l sarà no 'l prossim benchmark, ma 'l prim plugin deventaa indispensabij — quel moment che on ecosistema el smett de dipend del sò model fundador. L'è propi quell che 'l succed quand vuna piattaforma l'ha vengiuu.