The Neuron Times

All the AI that's fit to print

N° 274 Édition du matinMorning EditionMorgenausgabeEdizione del mattinoEdizion del mattin · Genève JEUDI 1 OCTOBRE 2026THURSDAY, 1 OCTOBER 2026DONNERSTAG, 1. OKTOBER 2026GIOVEDÌ 1 OTTOBRE 2026GIOVEDÌ 1 OTTOBRE 2026

À la Une · FrontièreFront Page · FrontierSchlagzeilen · FrontierPrima pagina · FrontieraIn prima pagina · Frontiera

Gemini 4 Argon : Google DeepMind ouvre sa « nouvelle ère » d'intelligence de frontièreGemini 4 Argon: Google DeepMind opens its “new era” of frontier intelligenceGemini 4 Argon: Google DeepMind eröffnet seine «neue Ära» der Frontier-IntelligenzGemini 4 Argon: Google DeepMind apre la sua «nuova era» di intelligenza di frontieraGemini 4 Argon : Google DeepMind el derva la soa « noeuva era » d'intelligenza de frontiera

Google DeepMind dévoile son nouveau modèle frontière dédié au codage réel, au travail de connaissances en entreprise et à la cyberdéfense, avec un déploiement imminent.Google DeepMind unveils its new frontier model dedicated to real-world coding, enterprise knowledge work, and cyberdefense, with an imminent rollout.Google DeepMind enthüllt sein neues Frontier-Modell, gewidmet dem realen Programmieren, der Wissensarbeit im Unternehmen und der Cyberabwehr – mit unmittelbar bevorstehendem Rollout.Google DeepMind svela il suo nuovo modello di frontiera dedicato al coding reale, al lavoro di conoscenza in ambito aziendale e alla cyberdifesa, con un rilascio imminente.Google DeepMind el presentа el so noeuv model frontiera dedicàa al codaz real, al laurà de conossenze in azienda e a la ciberdefesa, con on deploy iminent.

Google DeepMind a annoncé le 30 septembre 2026, à 20h00 UTC, Gemini 4 Argon, présenté comme son modèle de frontière pour la prochaine ère. Le labo cible trois terrains d'application concrets : le « real-world coding », le travail de connaissances en entreprise et la cyberdéfense. Le déploiement est annoncé comme imminent (« rolling out soon »).Google DeepMind announced Gemini 4 Argon on Wednesday, 30 September 2026, at 20:00 UTC, presented as its frontier model for the next era. The lab targets three concrete application domains: “real-world coding,” enterprise knowledge work, and cyberdefense. Deployment is described as imminent (“rolling out soon”).Google DeepMind hat am 30. September 2026 um 20:00 UTC Gemini 4 Argon angekündigt, präsentiert als sein Frontier-Modell für die kommende Ära. Das Labor zielt auf drei konkrete Anwendungsfelder ab: «Real-World-Coding», Wissensarbeit im Unternehmen und Cyberabwehr. Der Rollout wird als unmittelbar bevorstehend angekündigt («rolling out soon»).Google DeepMind ha annunciato il 30 settembre 2026, alle 20:00 UTC, Gemini 4 Argon, presentato come il suo modello di frontiera per la prossima era. Il laboratorio punta su tre terreni d'applicazione concreti: il «real-world coding», il lavoro di conoscenza in ambito aziendale e la cyberdifesa. Il rilascio è annunciato come imminente («rolling out soon»).Google DeepMind l'ha anunziaa el 30 de settember 2026, ai 20:00 UTC, Gemini 4 Argon, presentaa come el so model de frontiera per la prossima era. El labo el mira a trii camp de aplicazion concrecc : el « real-world coding », el laurà de conossenze in azienda e la ciberdefesa. El deploy l'è staa anunziaa come iminent (« rolling out soon »).

L'ampleur de l'événement se mesure aussi côté communauté : la discussion Hacker News ouverte le 30 septembre 2026 a déjà rassemblé 1 145 points et 760 commentaires, avec un fil d'analyse parallèle intitulé « Gemini 4 Argon (High) : Intelligence, Performance and Price Analysis » (https://news.ycombinator.com/item?id=49914236). Ce niveau d'attention place Argon dans la même catégorie de moments-frontière que les sorties GPT-6 et Claude 5.5 de la semaine.The scale of the event can also be measured by the community's response: the Hacker News thread opened on 30 September 2026 has already gathered 1,145 points and 760 comments, alongside a parallel analysis thread titled “Gemini 4 Argon (High): Intelligence, Performance and Price Analysis” (https://news.ycombinator.com/item?id=49914236). This level of attention places Argon in the same category of frontier moments as this week's GPT-6 and Claude 5.5 releases.Das Ausmass des Ereignisses zeigt sich auch in der Community: Die am 30. September 2026 eröffnete Diskussion auf Hacker News hat bereits 1145 Punkte und 760 Kommentare versammelt, mit einem parallelen Analyse-Thread mit dem Titel «Gemini 4 Argon (High): Intelligence, Performance and Price Analysis» (https://news.ycombinator.com/item?id=49914236). Diese Aufmerksamkeit ordnet Argon in dieselbe Kategorie von Frontier-Momenten ein wie die Releases von GPT-6 und Claude 5.5 in dieser Woche.L'entità dell'evento si misura anche sul fronte della community: la discussione Hacker News aperta il 30 settembre 2026 ha già raccolto 1.145 punti e 760 commenti, con un thread di analisi parallelo intitolato «Gemini 4 Argon (High): Intelligence, Performance and Price Analysis» (https://news.ycombinator.com/item?id=49914236). Questo livello di attenzione colloca Argon nella stessa categoria di momenti-frontiera dei lanci GPT-6 e Claude 5.5 della settimana.L'ampiessa de l'event la se mizura anca de part de la comunità : la discussion Hacker News dervida el 30 de settember 2026 l'ha giamò cenii 1.145 pont e 760 coment, con on fil de analisi parallela intitolaa « Gemini 4 Argon (High) : Intelligence, Performance and Price Analysis » (https://news.ycombinator.com/item?id=49914236). Che nivel de atention el mett Argon in la stessa categuria de moment-frontiera di usidd de GPT-6 e Claude 5.5 de la setemana.

Page 1 — Page 1 — Seite 1 — Pagina 1 — Pagina 1 — À la UneFront PageSchlagzeilenPrima paginaIn Cover

I. Modèles & frontièreModels & FrontierModelle & FrontierModelli & frontieraModel & frontiera

Sécurité

Security

Sicherheit

Sicurezza

Sigurezza

SynthID Bio : DeepMind applique le watermarking aux protéines de synthèseSynthID Bio: DeepMind applies watermarking to synthetic proteinsSynthID Bio: DeepMind wendet Wasserzeichen auf synthetische Proteine anSynthID Bio: DeepMind applica il watermarking alle proteine di sintesiSynthID Bio : DeepMind el aplica el watermarking ai proteinn de sintesi

Google DeepMind a présenté le 30 septembre 2026 SynthID Bio, une preuve de concept qui étend sa technologie de tatouage numérique (watermarking) aux protéines générées par IA, en préservant leur fonction biologique. Détails et rapport de recherche complets sur le blog DeepMind.
Google DeepMind unveiled SynthID Bio on 30 September 2026, a proof of concept that extends its digital watermarking technology to AI-generated proteins while preserving their biological function. Full details and the complete research report are available on the DeepMind blog.
Google DeepMind hat am 30. September 2026 SynthID Bio vorgestellt, einen Proof of Concept, der seine Technologie für digitale Wasserzeichen auf KI-generierte Proteine ausdehnt und dabei deren biologische Funktion erhält. Details und der vollständige Forschungsbericht auf dem DeepMind-Blog.
Google DeepMind ha presentato il 30 settembre 2026 SynthID Bio, una proof of concept che estende la sua tecnologia di watermarking alle proteine generate dall'IA, preservandone la funzione biologica. Dettagli e rapporto di ricerca completi sul blog DeepMind.
Google DeepMind l'ha presentаa el 30 de settember 2026 SynthID Bio, ona preuva de concett che la slarga la soa tecnulogia de tatugg digital (watermarking) ai proteinn generaa de l'IA, preservand la soa funzion biulogica. Detall e rapport de ricerca complet sora el blog DeepMind.

Science appliquée

Applied Science

Angewandte Wissenschaft

Scienza applicata

Scienza aplicada

L'IA scientifique de Google n°1 pour prédire les hospitalisations grippalesGoogle's science AI ranks number one at predicting flu hospitalizationsGoogles Science-KI auf Platz 1 bei der Vorhersage grippebedingter HospitalisierungenL'IA scientifica di Google è numero 1 nel prevedere i ricoveri influenzaliL'IA scientifica de Google la number on per predì i ospedalizazion de infuensa

Le modèle de science-AI de Google s'est classé numéro 1 pour la prévision des hospitalisations liées à la grippe, selon une annonce des Centers for Disease Control relayée le 30 septembre 2026. Google souligne la portée opérationnelle de cette prévision pour la santé publique : annonce complète.
Google's science-AI model ranked number one for forecasting influenza-related hospitalizations, according to an announcement from the Centers for Disease Control relayed on 30 September 2026. Google emphasizes the operational reach of this forecasting for public health: full announcement.
Googles Science-KI-Modell belegte Platz 1 bei der Vorhersage grippebedingter Hospitalisierungen, wie eine am 30. September 2026 übermittelte Ankündigung der Centers for Disease Control zeigt. Google betont die operationelle Reichweite dieser Vorhersage für die öffentliche Gesundheit: vollständige Ankündigung.
Il modello di science-AI di Google si è classificato numero 1 per la previsione dei ricoveri legati all'influenza, secondo un annuncio dei Centers for Disease Control ripreso il 30 settembre 2026. Google sottolinea la portata operativa di questa previsione per la salute pubblica: annuncio completo.
El model de science-AI de Google l'è staa classificaa number on per la prevision di ospedalizazion ligaa a l'infuensa, segonda ona dichiarazion di Centers for Disease Control dada indree el 30 de settember 2026. Google el sottalinea l'importanza operativa de questa prevision per la sanità publica : dichiarazion completa.

II. Entreprise & écosystèmeEnterprise & EcosystemUnternehmen & ÖkosystemImpresa & ecosistemaAzienda & ecusistem

Infra & recherche

Infrastructure & Research

Infrastruktur & Forschung

Infra & ricerca

Infra & ricerca

Cohere lance Embed 5 et ouvre Compass en beta cloudCohere launches Embed 5 and opens Compass in cloud betaCohere lanciert Embed 5 und öffnet Compass in der Cloud-BetaCohere lancia Embed 5 e apre Compass in beta cloudCohere el lanza Embed 5 e el derva Compass in beta cloud

Cohere a dévoilé le 30 septembre 2026 deux lancements : Embed 5, une nouvelle famille de modèles d'embedding frontière disponible en déclinaisons Pro et Fast, et le passage en beta cloud de Compass, son moteur de recherche et de récupération state-of-the-art, jusqu'ici déployable uniquement en infrastructure dédiée. Annonces détaillées sur le blog Cohere et pour Compass cloud.
Cohere unveiled two launches on 30 September 2026: Embed 5, a new family of frontier embedding models available in Pro and Fast variants, and the move of Compass, its state-of-the-art search and retrieval engine, into cloud beta — until now deployable only on dedicated infrastructure. Detailed announcements on the Cohere blog and for Compass cloud.
Cohere hat am 30. September 2026 zwei Launches vorgestellt: Embed 5, eine neue Familie von Frontier-Embedding-Modellen, verfügbar in den Varianten Pro und Fast, sowie den Übergang von Compass, seiner State-of-the-Art-Such- und Retrieval-Engine, in die Cloud-Beta – bislang nur auf dedizierter Infrastruktur einsetzbar. Detaillierte Ankündigungen auf dem Cohere-Blog und für Compass Cloud.
Cohere ha svelato il 30 settembre 2026 due lanci: Embed 5, una nuova famiglia di modelli di embedding di frontiera disponibile nelle versioni Pro e Fast, e il passaggio in beta cloud di Compass, il suo motore di ricerca e retrieval state-of-the-art, finora distribuibile solo su infrastruttura dedicata. Annunci dettagliati sul blog Cohere e per Compass cloud.
Cohere l'ha presentаa el 30 de settember 2026 du lanzament : Embed 5, ona noeuva famiglia de model d'embedding frontiera disponibel in versiun Pro e Fast, e el passagg in beta cloud de Compass, el so motor de ricerca e de recuver state-of-the-art, fin adess lanciabel domà in infrastruttura dedicada. Anunzi detallaa sora el blog Cohere e per Compass cloud.

Page 2 — Page 2 — Seite 2 — Pagina 2 — Pagina 2 — Le Cahier TechniqueThe Technical BriefingDas Technik-DossierIl Quaderno TecnicoEl Quadern Tecnic

III. Harnais & moteursHarnesses & EnginesHarnesses & EnginesHarness & motoriHarnais & motor

Moteurs d'inférence

Inference Engines

Inferenz-Engines

Motori di inferenza

Motor d'inferenza

Magnitude : un moteur d'inférence qui s'accorde lui-même à votre matériel, jusqu'à 2x llama.cppMagnitude: an inference engine that tunes itself to your hardware, up to 2x llama.cppMagnitude: eine Inferenz-Engine, die sich selbst auf Ihre Hardware abstimmt – bis zu 2x llama.cppMagnitude: un motore di inferenza che si calibra da solo sul tuo hardware, fino a 2x llama.cppMagnitude : on motor d'inferenza che el s'accordà per lu cunt el voster hardware, fin a 2x llama.cpp

Magnitude, un moteur d'inférence open source (Apache 2.0) écrit en Rust par une équipe YC S25, se présente comme le premier moteur auto-optimisant pour agents locaux : ses kernels GPU se règlent sur la machine avant exécution. Sur Qwen 3.6 35B A3B (4 bit, 64k de contexte), il annonce un decode 92 % plus rapide que llama.cpp sur Mac M4 Pro 48 Go (30 → 57 tok/s), 19 % plus rapide sur DGX Spark en CUDA (49 → 58 tok/s), et environ 28 % de mémoire en moins par agent. Le projet s'accompagne d'une attention paginée hybride partageant les caches de préfixe entre sessions. Code et détails sur GitHub, discussion sur Hacker News (140 points, 63 commentaires).
Magnitude, an open-source (Apache 2.0) inference engine written in Rust by a YC S25 team, presents itself as the first self-optimizing engine for local agents: its GPU kernels tune themselves to the machine before execution. On Qwen 3.6 35B A3B (4-bit, 64k context), it reports decoding 92% faster than llama.cpp on a 48 GB Mac M4 Pro (30 → 57 tok/s), 19% faster on DGX Spark under CUDA (49 → 58 tok/s), and roughly 28% less memory per agent. The project also features a hybrid paged attention that shares prefix caches across sessions. Code and details on GitHub, discussion on Hacker News (140 points, 63 comments).
Magnitude, eine in Rust geschriebene Open-Source-Inferenz-Engine (Apache 2.0) eines YC-S25-Teams, präsentiert sich als erste selbstoptimierende Engine für lokale Agenten: Ihre GPU-Kernels kalibrieren sich vor der Ausführung auf der jeweiligen Maschine. Auf Qwen 3.6 35B A3B (4 Bit, 64k Kontext) verspricht sie ein um 92 % schnelleres Decoding als llama.cpp auf einem Mac M4 Pro mit 48 GB (30 → 57 tok/s), 19 % schneller auf DGX Spark mit CUDA (49 → 58 tok/s) und rund 28 % weniger Speicher pro Agent. Das Projekt umfasst zudem eine hybride, paginierte Attention, die Prefix-Caches über Sitzungen hinweg teilt. Code und Details auf GitHub, Diskussion auf Hacker News (140 Punkte, 63 Kommentare).
Magnitude, un motore di inferenza open source (Apache 2.0) scritto in Rust da un team YC S25, si presenta come il primo motore auto-ottimizzante per agenti locali: i suoi kernel GPU si calibrano sulla macchina prima dell'esecuzione. Su Qwen 3.6 35B A3B (4 bit, 64k di contesto), annuncia un decode del 92% più rapido rispetto a llama.cpp su Mac M4 Pro 48 GB (da 30 a 57 tok/s), del 19% più rapido su DGX Spark in CUDA (da 49 a 58 tok/s) e circa il 28% di memoria in meno per agente. Il progetto si accompagna a un'attenzione paginata ibrida che condivide le cache di prefisso tra sessioni. Codice e dettagli su GitHub, discussione su Hacker News (140 punti, 63 commenti).
Magnitude, on motor d'inferenza open source (Apache 2.0) scrivuu in Rust de ona squadra YC S25, el se presenta come el primm motor auto-utimizzant per agent lucal : i so kernel GPU se regolen sora la machina prima de l'ezecuzion. Sora Qwen 3.6 35B A3B (4 bit, 64k de contest), el anunzia on decode 92 % pussee svelt che llama.cpp sora Mac M4 Pro 48 Go (30 → 57 tok/s), 19 % pussee svelt sora DGX Spark in CUDA (49 → 58 tok/s), e circa 28 % de memoria de men per agent. El progett el ved anca ona atention paginada ibrida che la spartiss i cache de prefiss intra i session. Codiss e dettag sora GitHub, discussion sora Hacker News (140 pont, 63 coment).

Benchmarks

Benchmarks

Benchmarks

Benchmark

Benchmark

Hugging Face lance l'Open TTS Leaderboard pour la synthèse vocale multilingueHugging Face launches the Open TTS Leaderboard for multilingual speech synthesisHugging Face lanciert das Open TTS Leaderboard für mehrsprachige SprachsyntheseHugging Face lancia l'Open TTS Leaderboard per la sintesi vocale multilingueHugging Face la lanza l'Open TTS Leaderboard per la sintesi vus multilengov

Hugging Face a ouvert le 30 septembre 2026 l'Open TTS Leaderboard, un classement d'évaluation scalable pour la synthèse vocale multilingue et le clonage de voix. Objectif : offrir au domaine text-to-speech une boussole comparable aux classements existants pour les LLM et l'ASR. Le classement est consultable via l'annonce sur le blog Hugging Face.
Hugging Face opened the Open TTS Leaderboard on 30 September 2026, a scalable evaluation ranking for multilingual speech synthesis and voice cloning. The goal: give the text-to-speech field a compass comparable to existing leaderboards for LLMs and ASR. The ranking is available via the announcement on the Hugging Face blog.
Hugging Face hat am 30. September 2026 das Open TTS Leaderboard eröffnet, ein skalierbares Evaluations-Ranking für mehrsprachige Sprachsynthese und Voice-Cloning. Ziel: dem Text-to-Speech-Bereich einen Kompass zu geben, der mit den bestehenden Rankings für LLMs und ASR vergleichbar ist. Das Ranking ist über die Ankündigung im Hugging-Face-Blog einsehbar.
Hugging Face ha aperto il 30 settembre 2026 l'Open TTS Leaderboard, una classifica di valutazione scalabile per la sintesi vocale multilingue e il clonaggio della voce. Obiettivo: offrire al settore text-to-speech una bussola paragonabile alle classifiche esistenti per LLM e ASR. La classifica è consultabile tramite l'annuncio sul blog Hugging Face.
Hugging Face l'ha dervii el 30 de settember 2026 l'Open TTS Leaderboard, on classifiche de valutazion scalable per la sintesi vus multilengov e el cloning de vus. Ubiectiv : dà al camp text-to-speech ona bussoeula cumparabel ai classifegh esistent per i LLM e l'ASR. El classifiche el se pò consultà via l'anunzi sora el blog Hugging Face.

Page 3 — Page 3 — Seite 3 — Pagina 3 — Pagina 3 — La RechercheThe Research DeskDie ForschungLa RicercaLa Ricerca

IV. Papers du jourPapers of the DayPapers des TagesPaper del giornoPapers del dì

Architecture

Architecture

Architektur

Architettura

Architetura

Loop Scaling Laws : Meta unifie récurrence et MoE dans une même loi d'échelleLoop Scaling Laws: Meta unifies recurrence and MoE into a single scaling lawLoop Scaling Laws: Meta vereint Rekurrenz und MoE in einem gemeinsamen SkalierungsgesetzLoop Scaling Laws: Meta unifica ricorrenza e MoE in un'unica legge di scalaLoop Scaling Laws : Meta la uniffica recurrenza e MoE in l'istessa legg de scala

Une équipe d'AI at Meta publie le 30 septembre 2026 les premières lois d'échelle modélisant conjointement récurrence (looped transformers) et sparsité (MoE). Résultats chiffrés : la sparsité apporte ~3x d'efficacité en paramètres actifs, la récurrence ~2x en paramètres totaux sur le raisonnement, et à budget de calcul égal un MoE bouclé égale un MoE non-bouclé ~2x plus grand, jusqu'à l'échelle du trillion de tokens. Paper complet : arXiv:2609.40316.
An AI team at Meta published on 30 September 2026 the first scaling laws jointly modeling recurrence (looped transformers) and sparsity (MoE). Key figures: sparsity delivers ~3x efficiency in active parameters, recurrence ~2x in total parameters on reasoning, and at equal compute budget a looped MoE matches a non-looped MoE ~2x larger, up to trillion-token scale. Full paper: arXiv:2609.40316.
Ein KI-Team bei Meta veröffentlicht am 30. September 2026 die ersten Skalierungsgesetze, die Rekurrenz (looped transformers) und Sparsity (MoE) gemeinsam modellieren. Ergebnisse in Zahlen: Sparsity bringt rund 3x Effizienz bei aktiven Parametern, Rekurrenz rund 2x bei Gesamtparametern im Reasoning; bei gleichem Rechenbudget erreicht ein geloopter MoE einen rund 2x grösseren nicht-geloopten MoE – bis hinauf zur Skala von einer Billion Tokens. Vollständiges Paper: arXiv:2609.40316.
Un team di AI at Meta pubblica il 30 settembre 2026 le prime leggi di scala che modellano congiuntamente ricorrenza (looped transformers) e sparsità (MoE). Risultati numerici: la sparsità porta circa 3x di efficienza in parametri attivi, la ricorrenza circa 2x in parametri totali sul ragionamento e, a parità di budget di calcolo, un MoE looped eguaglia un MoE non-looped circa 2x più grande, fino alla scala del trilione di token. Paper completo: arXiv:2609.40316.
Ona squadra d'AI at Meta la publica el 30 de settember 2026 i primm legg de scala che modelen insema la recurrenza (looped transformers) e la sparsità (MoE). Resültài numerich : la sparsità la ghe dà ~3x de eficienza in parametr attiv, la recurrenza ~2x in parametr totaj sora el resanament, e a budgèt de calcol ugual on MoE buclaa el rivaa on MoE minga-buclaa ~2x pussee grand, fin a la scala del trillion de token. Paper complet : arXiv:2609.40316.

Auto-amélioration

Self-Improvement

Selbstverbesserung

Auto-miglioramento

Auto-megliurament

UniEvo-VL : les modèles multimodaux s'améliorent grâce à leurs propres critiquesUniEvo-VL: multimodal models improve through their own critiquesUniEvo-VL: multimodale Modelle verbessern sich dank ihrer eigenen KritikenUniEvo-VL: i modelli multimodali migliorano grazie alle proprie criticheUniEvo-VL : i model multimodaj se megliuren grazia ai so pròpri critiche

Stanford NLP propose UniEvo-VL (publié le 30 septembre 2026), une recette d'auto-distillation on-policy où un modèle multimodal unique joue tour à tour professeur et élève via ses propres critiques. Sur Qwen-image-2512, le gain est net : 0.747 → 0.808 sur GenEval et 32.97 → 35.53 sur GenEval2 Soft-TIFA, sans superviseur externe. Paper : arXiv:2609.38721 (15 upvotes sur HF Daily Papers).
Stanford NLP proposes UniEvo-VL (published 30 September 2026), an on-policy self-distillation recipe in which a single multimodal model takes turns as teacher and student via its own critiques. On Qwen-image-2512, the gain is clear: 0.747 → 0.808 on GenEval and 32.97 → 35.53 on GenEval2 Soft-TIFA, without any external supervisor. Paper: arXiv:2609.38721 (15 upvotes on HF Daily Papers).
Stanford NLP schlägt UniEvo-VL vor (veröffentlicht am 30. September 2026), ein On-Policy-Rezept zur Selbst-Destillation, bei dem ein einziges multimodales Modell abwechselnd Lehrer und Schüler spielt – anhand seiner eigenen Kritiken. Auf Qwen-image-2512 fällt der Zugewinn deutlich aus: 0.747 → 0.808 auf GenEval und 32.97 → 35.53 auf GenEval2 Soft-TIFA, ganz ohne externen Supervisor. Paper: arXiv:2609.38721 (15 Upvotes auf HF Daily Papers).
Stanford NLP propone UniEvo-VL (pubblicato il 30 settembre 2026), una ricetta di auto-distillazione on-policy in cui un unico modello multimodale svolge a turno il ruolo di insegnante e di allievo tramite le proprie critiche. Su Qwen-image-2512, il guadagno è netto: da 0,747 a 0,808 su GenEval e da 32,97 a 35,53 su GenEval2 Soft-TIFA, senza supervisione esterna. Paper: arXiv:2609.38721 (15 upvote su HF Daily Papers).
Stanford NLP el propos UniEvo-VL (pubblica el 30 de settember 2026), ona receta d'auto-distillazion on-policy indove on model multimodal unic el fa a torn professor e sculior via i so pròpri critiche. Sora Qwen-image-2512, el guadagn l'è ciar : 0.747 → 0.808 sora GenEval e 32.97 → 35.53 sora GenEval2 Soft-TIFA, senza supervisur ester. Paper : arXiv:2609.38721 (15 upvotes sora HF Daily Papers).

Agents

Agents

Agenten

Agenti

Agent

50 000 paires d'erreurs d'agents pour entraîner des modèles qui se relèvent50,000 pairs of agent errors to train models that get back up50'000 Agenten-Fehlerpaare, um Modelle zu trainieren, die wieder aufstehen50.000 coppie di errori di agenti per addestrare modelli che si rialzano50.000 cobbi de errur d'agent per imparà model che se repìjen

L'Agent Error Dataset (30 septembre 2026) rassemble 50 228 paires erreur-diagnostic issues de 9 961 tâches, 33 environnements, 19 familles de harness et 23 modèles : la plus large base de défaillances agents à ce jour. Sur 3 062 replays appariés, les corrections proposées font passer le taux de validation de 18,4 % à 51,1 % (+32,7 points). Paper : arXiv:2609.40111 (15 upvotes).
The Agent Error Dataset (30 September 2026) gathers 50,228 error-diagnosis pairs drawn from 9,961 tasks, 33 environments, 19 harness families, and 23 models: the largest base of agent failures to date. Across 3,062 matched replays, the proposed corrections raise the validation rate from 18.4% to 51.1% (+32.7 points). Paper: arXiv:2609.40111 (15 upvotes).
Das Agent Error Dataset (30. September 2026) versammelt 50'228 Fehler-Diagnose-Paare aus 9'961 Aufgaben, 33 Umgebungen, 19 Harness-Familien und 23 Modellen: die bislang grösste Datenbasis für Agenten-Fehlfunktionen. Bei 3'062 gepaarten Replays steigern die vorgeschlagenen Korrekturen die Validierungsrate von 18,4 % auf 51,1 % (+32,7 Punkte). Paper: arXiv:2609.40111 (15 Upvotes).
L'Agent Error Dataset (30 settembre 2026) raccoglie 50.228 coppie errore-diagnosi provenienti da 9.961 task, 33 ambienti, 19 famiglie di harness e 23 modelli: la più ampia base di guasti di agenti realizzata finora. Su 3.062 replay abbinati, le correzioni proposte fanno passare il tasso di validazione dal 18,4% al 51,1% (+32,7 punti). Paper: arXiv:2609.40111 (15 upvote).
L'Agent Error Dataset (30 de settember 2026) el regiusta 50.228 cobbi errur-diagnostich vegnud de 9.961 taegh, 33 ambiencc, 19 famii de harness e 23 model : la base pussee larga de defaii agent fin adess. Sora 3.062 replay aparegiaa, i curi proponid fan passà el tass de validazion dal 18,4 % al 51,1 % (+32,7 punt). Paper : arXiv:2609.40111 (15 upvotes).

Harnais

Harnesses

Harnesses

Harness

Harnais

AI4AI : un modèle apprend à concevoir les harness de ses pairsAI4AI: a model learns to design its peers' harnessesAI4AI: ein Modell lernt, die Harnesses seiner Kollegen zu entwerfenAI4AI: un modello impara a progettare gli harness dei suoi pariAI4AI : on model el inpàra a custruì i harness di so söcc

Signé Apodex et Heng Ji (UIUC), « Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI » (30 septembre 2026) montre qu'un modèle Builder figé peut apprendre à construire de meilleurs environnements d'exécution pour un modèle Target : +8,95 points de performance macro-moyenne sans compétences, +12,02 points par rapport à la remise directe du même socle au Target. Paper le plus upvoté du jour (34 upvotes) : arXiv:2609.38143.
Signed by Apodex and Heng Ji (UIUC), “Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI” (30 September 2026) shows that a frozen Builder model can learn to construct better execution environments for a Target model: +8.95 points of macro-average performance without skills, +12.02 points compared with directly handing the same foundation to the Target. Most upvoted paper of the day (34 upvotes): arXiv:2609.38143.
Von Apodex und Heng Ji (UIUC) zeigt «Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI» (30. September 2026), dass ein eingefrorenes Builder-Modell lernen kann, bessere Ausführungsumgebungen für ein Target-Modell zu bauen: +8,95 Punkte makro-gemittelter Performance ohne Skills, +12,02 Punkte gegenüber der direkten Übergabe derselben Basis an das Target. Das meistgeupvotete Paper des Tages (34 Upvotes): arXiv:2609.38143.
Firmato Apodex e Heng Ji (UIUC), «Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI» (30 settembre 2026) mostra che un modello Builder congelato può imparare a costruire ambienti di esecuzione migliori per un modello Target: +8,95 punti di performance macro-media senza competenze, +12,02 punti rispetto alla consegna diretta della stessa base al Target. Paper con più upvote della giornata (34 upvote): arXiv:2609.38143.
Firmaa Apodex e Heng Ji (UIUC), « Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI » (30 de settember 2026) el mostra che on model Builder cngrèzzaa pò imparà a custruì ambient d'ezecuzion mejj per on model Target : +8,95 punt de performance macro-media senza competenze, +12,02 punt rispet al consegnà diret del istess socc al Target. Paper pussee upvotaa del dì (34 upvotes) : arXiv:2609.38143.

Page 4 — Page 4 — Seite 4 — Pagina 4 — Pagina 4 — La Communauté & ÉditoCommunity & EditorialCommunity & EditorialLa Community & EditorialeLa Comunità & Edito

V. Signaux & tribuneSignals & OpinionSignale & TribüneSegnali & tribunaSignal & tribuna

Apple ML

Apple ML

Apple ML

Apple ML

Apple ML

Apple publie sur le contrôle des LLM et l'entraînement d'agents continusApple publishes on LLM control and continual agent trainingApple veröffentlicht zur Kontrolle von LLMs und zum Training kontinuierlich lernender AgentenApple pubblica sul controllo dei LLM e sull'addestramento di agenti continuiApple la publica sora el cuntroll di LLM e l'imparament d'agent cuntinuv

Apple ML Research a publié le 30 septembre 2026 deux travaux : une étude systématique du compromis efficacité-fluence dans le conditionnement des LLM (lire le paper), qui montre que les méthodes de steering efficaces dégradent souvent la qualité de génération, et SCLATE (détails ici), un substrat d'exécution pour entraîner et évaluer des agents à apprentissage continu sur des horizons multi-sessions — l'infrastructure que les benchmarks existants, incapables d'entrelacer événements de benchmark et événements d'agent, ne fournissaient pas.
Apple ML Research published two pieces of work on 30 September 2026: a systematic study of the effectiveness-fluency trade-off in LLM conditioning (read the paper), showing that effective steering methods often degrade generation quality, and SCLATE (details here), an execution substrate for training and evaluating continually learning agents over multi-session horizons — the infrastructure that existing benchmarks, unable to interleave benchmark events with agent events, did not provide.
Apple ML Research hat am 30. September 2026 zwei Arbeiten veröffentlicht: eine systematische Studie des Trade-offs zwischen Effektivität und Flüssigkeit beim Conditioning von LLMs (Paper lesen), die zeigt, dass effektive Steering-Methoden die Generierungsqualität oft verschlechtern, sowie SCLATE (Details hier), ein Ausführungssubstrat zum Training und zur Evaluation von Agenten mit kontinuierlichem Lernen über Multi-Session-Horizonte – eine Infrastruktur, die bestehende Benchmarks, die Benchmark-Ereignisse und Agenten-Ereignisse nicht verzahnen können, bislang nicht boten.
Apple ML Research ha pubblicato il 30 settembre 2026 due lavori: uno studio sistematico del compromesso efficacia-fluenza nel condizionamento dei LLM (leggere il paper), che mostra come i metodi di steering efficaci degradano spesso la qualità della generazione, e SCLATE (dettagli qui), un substrato di esecuzione per addestrare e valutare agenti ad apprendimento continuo su orizzonti multi-sessione — l'infrastruttura che i benchmark esistenti, incapaci di intrecciare eventi di benchmark ed eventi di agente, non fornivano.
Apple ML Research l'ha publica el 30 de settember 2026 du lavòr : on studio sistematic del compromiss eficenza-fuenza in el condizionament di LLM (lee el paper), che 'l mostra che i metod de steering eficenz menen spess via la qualità de generazion, e SCLATE (detall chì), on substrat d'ezecuzion per imparà e valutà agent a apprendiment cuntinuv sora orizzun multi-session — l'infrastruttura che i benchmark esistent, minga bonn de intrelaçà event de benchmark ed event d'agent, minga gh'even minga fornii.

Recherche appliquée

Applied Research

Angewandte Forschung

Ricerca applicata

Ricerca aplicada

Météo spatiale : Microsoft prédit les risques sur les réseaux électriques 30 à 60 minutes à l'avanceSpace weather: Microsoft predicts power grid risks 30 to 60 minutes aheadWeltraumwetter: Microsoft sagt Risiken für Stromnetze 30 bis 60 Minuten im Voraus vorausMeteo spaziale: Microsoft prevede i rischi sulle reti elettriche 30-60 minuti in anticipoMeteo spazial : Microsoft la prediss i ris'c sora i rej elettrich 30-60 minut prima

Microsoft Research a dévoilé le 30 septembre 2026 un système de machine learning capable de prédire où des dommages surviendront sur les réseaux électriques 30 à 60 minutes avant l'arrivée d'une tempête solaire — un enjeu qui touche aussi la précision GPS et les opérations satellitaires. Détails : annonce Microsoft Research.
Microsoft Research unveiled on 30 September 2026 a machine learning system capable of predicting where damage will occur on power grids 30 to 60 minutes before a solar storm arrives — a challenge that also affects GPS accuracy and satellite operations. Details: Microsoft Research announcement.
Microsoft Research hat am 30. September 2026 ein Machine-Learning-System vorgestellt, das 30 bis 60 Minuten vor Eintreffen eines Sonnensturms vorhersagen kann, wo auf Stromnetzen Schäden auftreten werden – ein Thema, das auch die GPS-Genauigkeit und Satellitenoperationen betrifft. Details: Ankündigung von Microsoft Research.
Microsoft Research ha svelato il 30 settembre 2026 un sistema di machine learning in grado di prevedere dove si verificheranno danni sulle reti elettriche 30-60 minuti prima dell'arrivo di una tempesta solare — una sfida che tocca anche la precisione GPS e le operazioni satellitari. Dettagli: annuncio Microsoft Research.
Microsoft Research l'ha presentаa el 30 de settember 2026 on sistema de machine learning bon de predì indove che di dagn succedaran sora i rej elettrich 30-60 minut prima de l'ariv d'on temporaa solar — on tema che 'l toca anca la precision GPS e i operazion satellitar. Dettag : dichiarazion Microsoft Research.