The Neuron Times

All the AI that's fit to print

N° 266 Édition du matinMorning EditionMorgenausgabeEdizione del mattinoEdizion del mattin · Genève MERCREDI 23 SEPTEMBRE 2026WEDNESDAY, 23 SEPTEMBER 2026MITTWOCH, 23. SEPTEMBER 2026MERCOLEDÌ 23 SETTEMBRE 2026MERCOLEDÌ 23 SETTEMBRE 2026

À la Une · Sortie frontièreFront Page · Frontier ReleaseSchlagzeilen · Frontier-ReleasePrima pagina · Uscita frontieraIn prima pagina · Sortida de la frontera

Claude Opus 5.5 : Anthropic rapproche la frontière du bureau de travailClaude Opus 5.5: Anthropic Brings the Frontier Closer to the DesktopClaude Opus 5.5: Anthropic rückt die Frontier auf den SchreibtischClaude Opus 5.5: Anthropic avvicina la frontiera alla scrivaniaClaude Opus 5.5: Anthropic la fa rivà la frontera in del bureau de laurà

Le nouveau modèle atteint le niveau de Claude Fable 5.1 sur la plupart des travaux, pour un coût d'exploitation inférieur de 40 % à Opus 5, selon Anthropic.The new model reaches the level of Claude Fable 5.1 on most workloads, at an operating cost 40% below Opus 5, according to Anthropic.Das neue Modell erreicht auf den meisten Arbeiten das Niveau von Claude Fable 5.1 – bei um 40 Prozent tieferen Betriebskosten als Opus 5, wie Anthropic mitteilt.Il nuovo modello raggiunge il livello di Claude Fable 5.1 nella maggior parte dei lavori, con un costo di esercizio inferiore del 40% rispetto a Opus 5, secondo Anthropic.El noeuv modell 'l riva al nivell de Claude Fable 5.1 in su la magior part di laurà, con on cust de esercizzi inferiur del 40 % rispett a Opus 5, segond Anthropic.

Anthropic a annoncé le 22 septembre 2026 la sortie de Claude Opus 5.5, qui « performe au niveau de Claude Fable 5.1 sur la plupart des tâches professionnelles » tout en coûtant « 40 % de moins à l'exploitation qu'Opus 5 », selon l'annonce officielle publiée sur la page d'annonces d'Anthropic. Le positionnement est explicite : transférer les capacités du haut de gamme vers un palier de coût inférieur.Anthropic announced on September 22, 2026 the release of Claude Opus 5.5, which "performs at the level of Claude Fable 5.1 on most professional tasks" while costing "40% less to operate than Opus 5," according to the official announcement published on Anthropic's announcements page. The positioning is explicit: transfer top-of-the-line capabilities to a lower cost tier.Anthropic hat am 22. September 2026 Claude Opus 5.5 angekündigt, der «auf den meisten beruflichen Aufgaben auf dem Niveau von Claude Fable 5.1 performt», dabei gemäss der offiziellen Ankündigung von Anthropic «im Betrieb 40 Prozent weniger kostet als Opus 5». Das Positionierung ist explizit: die Fähigkeiten der Spitzenklasse in eine tiefere Kostenstufe überführen.Anthropic ha annunciato il 22 settembre 2026 l'uscita di Claude Opus 5.5, che «offre prestazioni al livello di Claude Fable 5.1 nella maggior parte dei compiti professionali» pur costando «il 40% in meno nell'esercizio rispetto a Opus 5», secondo l'annuncio ufficiale pubblicato su la pagina degli annunci di Anthropic. Il posizionamento è esplicito: trasferire le capacità di fascia alta verso un livello di costo inferiore.Anthropic l'ha anunziaa el 22 de setember 2026 la sortida de Claude Opus 5.5, che « 'l va al nivell de Claude Fable 5.1 in su la magior part di laurà professionai » e che 'l custa « el 40 % de manch a fàll andà rispett a Opus 5 », segond l'anunzi offizial publicaa sora la pagina di anunzi de Anthropic. El posizzionament l'è ciar e net: portà i capacità de la gamma volta invers un nivell de cust pussee bass.

La sortie intervient trois semaines après celle de Claude Fable 5.1 et Claude Mythos 5.1, présentées le 1er septembre 2026 comme les « modèles les plus avancés » du laboratoire pour le codage et le travail de connaissance. Opus 5.5 s'inscrit ainsi comme la déclinaison économique de cette génération. La page produit est consultable sur anthropic.com/claude-opus-5-5.The release comes three weeks after that of Claude Fable 5.1 and Claude Mythos 5.1, presented on September 1, 2026 as the laboratory's "most advanced models" for coding and knowledge work. Opus 5.5 thus stands as the economy version of this generation. The product page is available at anthropic.com/claude-opus-5-5.Die Veröffentlichung erfolgt drei Wochen nach jener von Claude Fable 5.1 und Claude Mythos 5.1, die am 1. September 2026 als die «fortschrittlichsten Modelle» des Labors für Coding und Wissensarbeit vorgestellt wurden. Opus 5.5 positioniert sich damit als die wirtschaftliche Variante dieser Generation. Die Produktseite ist unter anthropic.com/claude-opus-5-5 abrufbar.L'uscita arriva tre settimane dopo quella di Claude Fable 5.1 e Claude Mythos 5.1, presentati il 1° settembre 2026 come i «modelli più avanzati» del laboratorio per il coding e il lavoro di conoscenza. Opus 5.5 si inserisce così come la declinazione economica di questa generazione. La pagina prodotto è consultabile su anthropic.com/claude-opus-5-5.La sortida la riva tri seteman dopo quella de Claude Fable 5.1 e Claude Mythos 5.1, presentaa el 1 de setember 2026 come i « modei pussee avanzaa » del laboratori per el codes e per el laurà de conoscenza. Opus 5.5 el se met inscì coma la version economica de quella generazion chì. La pagina prodott l'è consultabel sora anthropic.com/claude-opus-5-5.

L'accueil a été massif : la discussion Hacker News consacrée au modèle a dépassé 1 353 points et 860 commentaires en moins de 24 heures, tandis que l'observatoire Artificial Analysis publiait dès le 22 septembre une analyse indépendante de son intelligence, de ses performances et de son prix.The reception has been massive: the Hacker News discussion dedicated to the model surpassed 1,353 points and 860 comments in less than 24 hours, while the Artificial Analysis observatory published an independent analysis of its intelligence, performance and pricing as early as September 22.Das Echo war gewaltig: Die Hacker-News-Diskussion zum Modell überschritt innert 24 Stunden 1353 Punkte und 860 Kommentare, während das Observatorium Artificial Analysis bereits am 22. September eine unabhängige Analyse seiner Intelligenz, Leistung und seines Preises publizierte.L'accoglienza è stata massiccia: la discussione Hacker News dedicata al modello ha superato 1 353 punti e 860 commenti in meno di 24 ore, mentre l'osservatorio Artificial Analysis pubblicava già il 22 settembre un'analisi indipendente della sua intelligenza, delle sue prestazioni e del suo prezzo.L'acolt l'è staa enorm: la discussion Hacker News dedicada al modell l'ha superaa 1.353 pont e 860 comenti in manch de 24 or, intant che l'osservatori Artificial Analysis 'l publicava giamò el 22 de setember ona analisi independenta de la sò intelligenza, di sò prestazion e del sò pres.

Page 1 — Page 1 — Seite 1 — Pagina 1 — Pagina 1 — Page 1 — À la UnePage 1 — Front PageSeite 1 — Der AufmacherPagina 1 — In prima paginaPagina 1 — A la Vuna

I. Modèles & frontièreModels & the FrontierModelle & FrontierModelli & frontieraModei & frontera

OpenAI

OpenAI

OpenAI

OpenAI

OpenAI

GPT-6 Sol et Luna : deux modèles frontière, un écart de prix de 20xGPT-6 Sol and Luna: Two Frontier Models, a 20x Price GapGPT-6 Sol und Luna: zwei Frontier-Modelle, ein Preisunterschied von 20xGPT-6 Sol e Luna: due modelli frontiera, un divario di prezzo di 20xGPT-6 Sol e Luna: do modei de frontera, on divari de pres de 20x

Publiés le 22 septembre 2026 à 18h00 UTC via l'annonce officielle, GPT-6 Sol et GPT-6 Luna « apportent l'intelligence frontière au travail quotidien avec des équilibres différents entre capacité et coût ». Selon le changelog de la plateforme, les deux modèles de raisonnement acceptent texte et image en entrée via les APIs Responses et Chat Completions.
Released on September 22, 2026 at 18:00 UTC via the official announcement, GPT-6 Sol and GPT-6 Luna "bring frontier intelligence to everyday work with different balances between capability and cost." According to the platform changelog, both reasoning models accept text and image input via the Responses and Chat Completions APIs.
Am 22. September 2026 um 18.00 Uhr UTC offiziell angekündigt, «bringen GPT-6 Sol und GPT-6 Luna Frontier-Intelligenz in die tägliche Arbeit – mit unterschiedlichen Gleichgewichten zwischen Fähigkeit und Kosten». Gemäss dem Changelog der Plattform akzeptieren die beiden Reasoning-Modelle Text und Bild als Eingabe über die APIs Responses und Chat Completions.
Pubblicati il 22 settembre 2026 alle 18:00 UTC tramite l'annuncio ufficiale, GPT-6 Sol e GPT-6 Luna «portano l'intelligenza frontiera nel lavoro quotidiano con equilibri diversi tra capacità e costo». Secondo il changelog della piattaforma, i due modelli di ragionamento accettano testo e immagini in ingresso tramite le API Responses e Chat Completions.
Publicaa el 22 de setember 2026 ai 18:00 UTC cont l'anunzi offizial, GPT-6 Sol e GPT-6 Luna « porten l'intelligenza de frontera in del laurà de tucc i dì con equiliber diferent tra capacità e cust ». Segond el changelog de la piattaforma, i do modei de resoneiment i ametten test e imaggin in entrada cont i API Responses e Chat Completions.

Cryptanalyse

Cryptanalysis

Kryptanalyse

Crittoanalisi

Crittoanalisi

GPT-6 Astra vient à bout d'un message Enigma irrésolu depuis 2005GPT-6 Astra Cracks an Enigma Message Unsolved Since 2005GPT-6 Astra bezwingt eine seit 2005 ungelöste Enigma-NachrichtGPT-6 Astra ha la meglio su un messaggio Enigma irrisolto dal 2005GPT-6 Astra el riss'cia on messagg Enigma minga deslenguad del 2005

Un message Enigma qui résistait aux cryptanalystes depuis 2005 a été cassé avec l'aide de GPT-6 Astra, rapporte le récit publié le 22 septembre 2026 sur Crypto Cellar. La discussion Hacker News, forte de 608 points et 373 commentaires, souligne le rôle du modèle dans l'exploration des hypothèses de décryptage.
An Enigma message that had resisted cryptanalysts since 2005 has been broken with the help of GPT-6 Astra, reports the account published on September 22, 2026 on Crypto Cellar. The Hacker News discussion, with 608 points and 373 comments, highlights the model's role in exploring decryption hypotheses.
Eine Enigma-Nachricht, die den Kryptoanalytikern seit 2005 widerstand, wurde mit Hilfe von GPT-6 Astra geknackt, wie der am 22. September 2026 auf Crypto Cellar veröffentlichte Bericht schildert. Die Hacker-News-Diskussion mit 608 Punkten und 373 Kommentaren hebt die Rolle des Modells bei der Exploration von Entschlüsselungshypothesen hervor.
Un messaggio Enigma che resisteva ai crittoanalisti dal 2005 è stato violato con l'aiuto di GPT-6 Astra, riporta il racconto pubblicato il 22 settembre 2026 su Crypto Cellar. La discussione Hacker News, con 608 punti e 373 commenti, sottolinea il ruolo del modello nell'esplorazione delle ipotesi di decrittazione.
On messagg Enigma che 'l resisteva ai crittoanalista del 2005 l'è staa casciaa cont l'agiutt de GPT-6 Astra, 'l rapporta el cunt publicaa el 22 de setember 2026 sora Crypto Cellar. La discussion Hacker News, cont 608 pont e 373 comenti, la met in lus el roeul del modell in de l'esplorazion di ipotesi de decifratura.

Page 2 — Page 2 — Seite 2 — Pagina 2 — Pagina 2 — Page 2 — Le Cahier TechniquePage 2 — The Technical SectionSeite 2 — Das Technik-DossierPagina 2 — Il Quaderno TecnicoPagina 2 — El Quadern Tecnegh

II. Harnais, engines & outilsHarnesses, Engines & ToolsHarnesses, Engines & ToolsHarness, engine & strumentiHarness, engine & strument

Hugging Face

Hugging Face

Hugging Face

Hugging Face

Hugging Face

Transformers fait tourner les quantifiés llama.cpp, nativementTransformers Now Runs llama.cpp Quants NativelyTransformers führt llama.cpp-Quantisierungen nativ ausTransformers esegue nativamente i quantizzati llama.cppTransformers el fa girà i quantifegaa llama.cpp, nativament

La bibliothèque phare de Hugging Face announced le 22 septembre 2026 qu'elle exécute désormais les quantifications llama.cpp dans transformers directement, rapprochant les écosystèmes PyTorch et GGUF sans outillage supplémentaire.
Hugging Face's flagship library announced on September 22, 2026 that it now runs llama.cpp quantizations directly in transformers, bringing the PyTorch and GGUF ecosystems closer together without additional tooling.
Die Flaggschiff-Bibliothek von Hugging Face gab am 22. September 2026 bekannt, dass sie llama.cpp-Quantisierungen nun direkt in transformers ausführt – die Ökosysteme PyTorch und GGUF rücken ohne zusätzliches Tooling zusammen.
La biblioteca di punta di Hugging Face ha annunciato il 22 settembre 2026 che ora esegue le quantizzazioni llama.cpp direttamente in transformers, avvicinando gli ecosistemi PyTorch e GGUF senza strumenti aggiuntivi.
La biblioteca pussee importanta de Hugging Face l'ha anunziaa el 22 de setember 2026 che ades la fà andà i quantifegazzion llama.cpp direttament in transformers, e 'l fa vesinà i ecosistem PyTorch e GGUF sensa tool extra.

Inférence

Inference

Inferenz

Inferenza

Inferenca

Changer de modèle en production sans coupure : la méthode canary de Together AISwapping Models in Production Without Downtime: Together AI's Canary MethodModellwechsel im Produktivbetrieb ohne Unterbruch: die Canary-Methode von Together AICambiare modello in produzione senza interruzioni: il metodo canary di Together AICambià modell in produsion sensa fà s'cioppà nagott: el metod canary de Together AI

Dans un billet publié le 22 septembre 2026, Together AI détaille comment monter en charge progressive d'un modèle en production : les canary rollouts combinent rampes de trafic par étapes, portes métriques et retour arrière automatique, là où « un swap dur expose tous les utilisateurs d'un coup » et oblige à redémarrer l'ancien déploiement sous pression.
In a post published on September 22, 2026, Together AI details how to progressively scale up a model in production: canary rollouts combine stepwise traffic ramps, metric gates and automatic rollback, whereas "a hard swap exposes all users at once" and forces restarting the old deployment under pressure.
In einem am 22. September 2026 erschienenen Beitrag zeigt Together AI, wie ein Modell im Produktivbetrieb progressiv hochgefahren wird: Canary Rollouts kombinieren gestufte Traffic-Rampen, metrische Gates und automatisches Rollback – dort, wo «ein harter Swap alle Nutzer auf einmal aussetzt» und einen Neustart des alten Deployments unter Druck erzwingt.
In un post pubblicato il 22 settembre 2026, Together AI dettaglia come aumentare progressivamente il carico di un modello in produzione: i canary rollout combinano rampe di traffico a tappe, soglie metriche e rollback automatico, là dove «uno swap brusco espone tutti gli utenti in un colpo solo» e obbliga a riavviare il vecchio deployment sotto pressione.
In d'on post publicaa el 22 de setember 2026, Together AI 'l dis come fà montà a la progressiva on modell in produsion: i canary rollouts i combinen rampe de trafegh a pass, port de metrigh e tornà indree automatigh, indoa che « on swap dur 'l met in evidenza tucc i utent in d'on colp domà » e 'l obliga a fà partì ancamò el vegg despiegh sota pressi.

Outils

Tooling

Tools

Strumenti

Strument

Unreal Agent : un agent IA au service du moteur Unreal EngineUnreal Agent: An AI Agent in the Service of the Unreal EngineUnreal Agent: ein KI-Agent im Dienst der Unreal EngineUnreal Agent: un agente IA al servizio del motore Unreal EngineUnreal Agent: on agent IA al servizzi del motor Unreal Engine

Le robot d'agents pour moteurs de jeux Unreal Engine annoncé le 22 septembre 2026 sur Unreal Labs a suscité 148 points sur Hacker News ; son code est disponible sur GitHub.
The agent robot for Unreal Engine game engines, announced on September 22, 2026 on Unreal Labs, drew 148 points on Hacker News; its code is available on GitHub.
Der Agenten-Roboter für Unreal-Engine-Spiele, angekündigt am 22. September 2026 auf Unreal Labs, erzielte 148 Punkte auf Hacker News; der Code ist auf GitHub verfügbar.
Il robot di agenti per motori di gioco Unreal Engine annunciato il 22 settembre 2026 su Unreal Labs ha raccolto 148 punti su Hacker News; il suo codice è disponibile su GitHub.
El robot de agent per motor di videogioeugh Unreal Engine, anunziaa el 22 de setember 2026 sora Unreal Labs, l'ha tirad 148 pont sora Hacker News; el sò codes l'è disponibil sora GitHub.

Évaluation

Evaluation

Evaluation

Valutazione

Valutazion

Rendre les benchmarks reproductibles : le chantier UK AISI et EvalEvalMaking Benchmarks Reproducible: The UK AISI and EvalEval ProjectBenchmarks reproduzierbar machen: das Vorhaben von UK AISI und EvalEvalRendere i benchmark riproducibili: il cantiere di UK AISI ed EvalEvalFà diventà i benchmark reprodusibij: el cantier UK AISI e EvalEval

L'institut britannique UK AISI et l'initiative EvalEval publient le 22 septembre 2026 une méthode de benchmarking reproductible sur le blog de Hugging Face, visant à rendre les résultats d'évaluation comparables et vérifiables entre organisations.
The UK AI Security Institute and the EvalEval initiative published on September 22, 2026 a reproducible benchmarking method on the Hugging Face blog, aiming to make evaluation results comparable and verifiable across organizations.
Das britische Institut UK AISI und die Initiative EvalEval veröffentlichen am 22. September 2026 im Blog von Hugging Face eine reproduzierbare Benchmarking-Methode, mit der Evaluationsresultate organisationsübergreifend vergleichbar und überprüfbar werden sollen.
L'istituto britannico UK AISI e l'iniziativa EvalEval pubblicano il 22 settembre 2026 un metodo di benchmarking riproducibile sul blog di Hugging Face, con l'obiettivo di rendere i risultati di valutazione comparabili e verificabili tra organizzazioni.
L'institut britannegh UK AISI e l'iniziativa EvalEval i publiegen el 22 de setember 2026 ona metoda de benchmarking reprodusibila sora el blog de Hugging Face, con l'obietiv de fà diventà i resultaa de valutazion paragonabel e verificabil tra organizazion.

Page 3 — Page 3 — Seite 3 — Pagina 3 — Pagina 3 — Page 3 — La RecherchePage 3 — ResearchSeite 3 — Die ForschungPagina 3 — La RicercaPagina 3 — La Ricerca

III. Papers & labosPapers & LabsPapers & LaborePaper & laboratoriPaper & laboratori

Quantum + LLM

Quantum + LLM

Quantum + LLM

Quantum + LLM

Quantum + LLM

HyperQ : des circuits quantiques de 16 à 64 qubits au cœur d'un LLM à diffusionHyperQ: 16-to-64-Qubit Quantum Circuits at the Heart of a Diffusion LLMHyperQ: Quantenschaltkreise von 16 bis 64 Qubits im Herzen eines Diffusions-LLMHyperQ: circuiti quantistici da 16 a 64 qubit al centro di un LLM a diffusioneHyperQ: circuitt quantistegh de 16 a 64 qubit in del coeur de on LLM a diffusion

Des chercheurs de l'Université de Montréal/Mila proposent HyperQ, des hypernetworks de circuits quantiques greffés sur un modèle de langage à diffusion de 1,1 milliard de paramètres. En passant de 16 à 64 qubits, le score moyen passe de 47,65 à 54,30 ; à 64 qubits, HyperQ dépasse le backbone gelé de 4,71 points et son équivalent LoRA de 3,67 points, avec seulement 20 000 paires d'entraînement contre 200 000 pour les baselines classiques.
Researchers from the Université de Montréal/Mila propose HyperQ, quantum circuit hypernetworks grafted onto a 1.1-billion-parameter diffusion language model. Going from 16 to 64 qubits, the average score rises from 47.65 to 54.30; at 64 qubits, HyperQ surpasses the frozen backbone by 4.71 points and its LoRA counterpart by 3.67 points, using only 20,000 training pairs versus 200,000 for classical baselines.
Forschende der Universität Montreal/Mila schlagen HyperQ vor: Hypernetworks von Quantenschaltkreisen, gepfropft auf ein Diffusions-Sprachmodell mit 1,1 Milliarden Parametern. Beim Übergang von 16 auf 64 Qubits steigt der mittlere Score von 47,65 auf 54,30; bei 64 Qubits übertrifft HyperQ das eingefrorene Backbone um 4,71 Punkte und sein LoRA-Äquivalent um 3,67 Punkte – mit bloss 20 000 Trainingspaaren gegenüber 200 000 bei klassischen Baselines.
Dei ricercatori dell'Università di Montréal/Mila propongono HyperQ, degli hypernetwork di circuiti quantistici innestati su un modello linguistico a diffusione da 1,1 miliardi di parametri. Passando da 16 a 64 qubit, il punteggio medio sale da 47,65 a 54,30; a 64 qubit, HyperQ supera il backbone congelato di 4,71 punti e il suo equivalente LoRA di 3,67 punti, con soli 20 000 coppie di addestramento contro 200 000 per le baseline classiche.
Di ricercador de l'Università de Montréal/Mila i proponen HyperQ, di hypernetwork de circuitt quantistegh mettuu sora on modell de lengoeu a diffusion de 1,1 miliard de parameter. Passand da 16 a 64 qubit, el pontegg medi 'l va de 47,65 a 54,30; a 64 qubit, HyperQ 'l supera el backbone gelaa de 4,71 pont e 'l so equivallent LoRA de 3,67 pont, con domà 20.000 cobbi de addestrament contra 200.000 per i baseline classigh.

Génération 3D

3D Generation

3D-Generierung

Generazione 3D

Generazion 3D

GAE : un espace latent géométrique pour une génération de mondes 3D cohérenteGAE: A Geometric Latent Space for Coherent 3D World GenerationGAE: ein geometrischer Latentraum für kohärente 3D-WeltgenerierungGAE: uno spazio latent geometrico per una generazione coerente di mondi 3DGAE: on spazzi latent geometregh per ona generazion coerenta di mond 3D

Le laboratoire ARC de Tencent présente GAE, un autoencodeur nativement géométrique dont le latent se décode en apparence, profondeur, caméras et cartes de points. À générateur et protocole constants, le remplacement du latent par GAE réduit la FVD de 12,7 % sur RealEstate10K et 23,1 % sur DL3DV, et divise par deux l'erreur de trajectoire caméra sur RealEstate10K.
Tencent's ARC lab presents GAE, a natively geometric autoencoder whose latent decodes into appearance, depth, cameras and point maps. With generator and protocol held constant, replacing the latent with GAE reduces FVD by 12.7% on RealEstate10K and 23.1% on DL3DV, and halves camera trajectory error on RealEstate10K.
Das ARC-Labor von Tencent stellt GAE vor, einen nativ geometrischen Autoencoder, dessen Latent sich in Erscheinungsbild, Tiefe, Kameras und Punktkarten dekodieren lässt. Bei gleichem Generator und Protokoll reduziert der Ersatz des Latents durch GAE die FVD um 12,7 Prozent auf RealEstate10K und um 23,1 Prozent auf DL3DV – und halbiert den Kameratrajektorien-Fehler auf RealEstate10K.
Il laboratorio ARC di Tencent presenta GAE, un autoencoder nativamente geometrico il cui latent si decodifica in apparenza, profondità, camere e mappe di punti. A generatore e protocollo costanti, la sostituzione del latent con GAE riduce la FVD del 12,7% su RealEstate10K e del 23,1% su DL3DV, e dimezza l'errore di traiettoria della camera su RealEstate10K.
El laboratori ARC de Tencent 'l presenta GAE, on autoencoder nativament geometregh che 'l so latent 'l se decodifica in apparenza, profondità, camer e map de pont. A generator e protocoll costant, el sostituì del latent cont GAA 'l sbassa la FVD del 12,7 % sora RealEstate10K e del 23,1 % sora DL3DV, e 'l sbissa per doi l'error de trajettoria camera sora RealEstate10K.

Inférence

Inference

Inferenz

Inferenza

Inferenca

Flash-dLLM : 11x d'accélération pour les LLM à diffusion, sans réentraînementFlash-dLLM: An 11x Speedup for Diffusion LLMs, No Retraining RequiredFlash-dLLM: 11x Beschleunigung für Diffusions-LLM, ohne RetrainingFlash-dLLM: 11x di accelerazione per gli LLM a diffusione, senza riaddestramentoFlash-dLLM: 11x de accelerazion per i LLM a diffusion, senza re-addestrament

Flash-dLLM, publié le 22 septembre 2026 par MBZUAI, identifie les entrées-sorties mémoire GPU comme goulot des LLM à diffusion et propose un noyau KV-cache fusionné sensible aux I/O, sans réentraînement. Résultat : des accélérations de 5,1x sur GSM8K et 11,0x sur HumanEval face au meilleur baseline Elastic-Cache, avec un décodage draft-and-verify où le modèle se vérifie lui-même.
Flash-dLLM, published on September 22, 2026 by MBZUAI, identifies GPU memory input-output as the bottleneck of diffusion LLMs and proposes an I/O-aware fused KV-cache kernel, with no retraining. Result: speedups of 5.1x on GSM8K and 11.0x on HumanEval against the best baseline, Elastic-Cache, with draft-and-verify decoding in which the model verifies itself.
Flash-dLLM, publiziert am 22. September 2026 von MBZUAI, identifiziert die GPU-Speicher-Ein-/Ausgaben als Engpass von Diffusions-LLM und schlägt einen I/O-sensitiven, fusionierten KV-Cache-Kernel ohne Retraining vor. Resultat: Beschleunigungen von 5,1x auf GSM8K und 11,0x auf HumanEval gegenüber der besten Baseline Elastic-Cache, mit Draft-and-Verify-Dekodierung, bei der sich das Modell selbst verifiziert.
Flash-dLLM, pubblicato il 22 settembre 2026 da MBZUAI, identifica gli input-output di memoria GPU come collo di bottiglia degli LLM a diffusione e propone un kernel KV-cache fusionato sensibile agli I/O, senza riaddestramento. Risultato: accelerazioni di 5,1x su GSM8K e 11,0x su HumanEval rispetto alla migliore baseline Elastic-Cache, con una decodifica draft-and-verify in cui il modello verifica se stesso.
Flash-dLLM, publicaa el 22 de setember 2026 de MBZUAI, 'l identifega i input-output de memoria GPU come coll de bottilia di LLM a diffusion e 'l propon on kernel KV-cache fusionaa sensibel ai I/O, senza re-addestrament. Resultaa: accelerazion de 5,1x sora GSM8K e 11,0x sora HumanEval contra el mej baseline Elastic-Cache, cont on decodifica draft-and-verify indoa che 'l modell 'l se verifica da perlù.

RL

RL

RL

RL

RL

RULER : des rubriques plutôt que des métriques scalaires pour piloter le RL en SVGRULER: Rubrics Instead of Scalar Metrics to Steer RL for SVGRULER: Rubriken statt skalare Metriken für das RL bei SVGRULER: rubriche invece di metriche scalari per pilotare il RL in SVGRULER: rubrich inveci de metrigh scalar per menà el RL in SVG

inclusionAI publie RULER, des récompenses par rubrique instance-aware pour l'apprentissage par renforcement en génération SVG. En six axes notés par un juge VLM, le score rubrique grimpe de 0,432 à 0,693 sur MMSVG-Illustration et de 0,395 à 0,683 sur MMSVG-Icon, égalant le bien plus grand DeepSeek-V3 — l'article a récolté 35 votes sur les Daily Papers de Hugging Face.
inclusionAI publishes RULER, instance-aware rubric-based rewards for reinforcement learning in SVG generation. Across six axes scored by a VLM judge, the rubric score climbs from 0.432 to 0.693 on MMSVG-Illustration and from 0.395 to 0.683 on MMSVG-Icon, matching the far larger DeepSeek-V3 — the paper gathered 35 votes on Hugging Face's Daily Papers.
inclusionAI publiziert RULER: instanzbewusste Rubrik-Belohnungen für bestärkendes Lernen bei der SVG-Generierung. Über sechs von einem VLM-Richter bewertete Achsen steigt der Rubrik-Score von 0,432 auf 0,693 auf MMSVG-Illustration und von 0,395 auf 0,683 auf MMSVG-Icon – ebenbürtig dem wesentlich grösseren DeepSeek-V3; das Paper erhielt 35 Stimmen auf den Daily Papers von Hugging Face.
inclusionAI pubblica RULER, delle ricompense per rubrica instance-aware per l'apprendimento per rinforzo nella generazione SVG. Su sei assi valutati da un giudice VLM, il punteggio rubrica sale da 0,432 a 0,693 su MMSVG-Illustration e da 0,395 a 0,683 su MMSVG-Icon, eguagliando il ben più grande DeepSeek-V3 — l'articolo ha raccolto 35 voti sui Daily Papers di Hugging Face.
inclusionAI 'l publica RULER, di ricompens per rubrica instance-aware per l'insegnament con renforz in generazion SVG. In ses ass votaa de on giudes VLM, el pontegg rubrica 'l va su de 0,432 a 0,693 sora MMSVG-Illustration e de 0,395 a 0,683 sora MMSVG-Icon, 'l riva al nivell del ben pussee grand DeepSeek-V3 — l'articol l'ha ciapaa 35 vott sora i Daily Papers de Hugging Face.

IV. À suivreTo FollowIm BlickDa seguireDe vedè

Auto-amélioration

Self-Improvement

Selbstverbesserung

Auto-miglioramento

Auto-migliorament

AIDE² : un agent de recherche s'améliore lui-même pendant huit joursAIDE²: A Research Agent That Improves Itself for Eight DaysAIDE²: ein Forschungsagent verbessert sich selbst über acht TageAIDE²: un agente di ricerca si migliora da solo per otto giorniAIDE²: on agent de ricerca 'l se mejor deperlù per vot dì

Weco AI présente AIDE², un agent de recherche qui réécrit son propre code, évalue ses versions sur des tâches d'R&D masquées et conserve les meilleures. En 8 jours autonomes, le système a découvert 7 améliorations successives ; sur 4 benchmarks hors échantillon, l'agent découvert égale ou dépasse un agent de production conçu par des humains. Fait notable : le taux de reward hacking tombe de 55 % à 32 % pendant la boucle, sans que cette propriété ait été optimisée explicitement.
Weco AI presents AIDE², a research agent that rewrites its own code, evaluates its versions on masked R&D tasks and keeps the best. Over 8 autonomous days, the system discovered 7 successive improvements; on 4 out-of-sample benchmarks, the discovered agent matches or exceeds a production agent designed by humans. Notably, the rate of reward hacking falls from 55% to 32% during the loop, without this property being explicitly optimized.
Weco AI stellt AIDE² vor, einen Forschungsagenten, der seinen eigenen Code umschreibt, seine Versionen auf verdeckten F&E-Aufgaben evaluiert und die besten behält. In acht autonomen Tagen entdeckte das System sieben aufeinanderfolgende Verbesserungen; auf vier ausserhalb der Stichprobe liegenden Benchmarks erreicht oder übertrifft der entdeckte Agent einen von Menschen konzipierten Produktionsagenten. Bemerkenswert: Die Quote des Reward Hackings sinkt während der Schleife von 55 auf 32 Prozent – ohne dass diese Eigenschaft explizit optimisiert wurde.
Weco AI presenta AIDE², un agente di ricerca che riscrive il proprio codice, valuta le proprie versioni su compiti di R&D mascherati e conserva le migliori. In 8 giorni autonomi, il sistema ha scoperto 7 miglioramenti successivi; su 4 benchmark fuori campione, l'agente scoperto eguaglia o supera un agente di produzione progettato da umani. Dato notevole: il tasso di reward hacking scende dal 55% al 32% durante il ciclo, senza che questa proprietà sia stata ottimizzata esplicitamente.
Weco AI 'l presenta AIDE², on agent de ricerca che 'l scriv ancamò el sò codes deperlù, 'l valuta i sò version sora di lavorà de R&D sconduu e 'l ten i mej. In 8 dì autonom, el sistema l'ha descovert 7 migliorament vun dree l'alter; sora 4 benchmark foeu de moster, l'agent descovert 'l riva o 'l supera on agent de produsion desegnaa di òmen. Roba de notà: el tass de reward hacking 'l cascia de 55 % a 32 % in de la breada, senza che quella proprietà chì la sia stada otimizada a la ciara.

Page 4 — Page 4 — Seite 4 — Pagina 4 — Pagina 4 — Page 4 — La Communauté & ÉditoPage 4 — Community & EditorialSeite 4 — Community & LeitartikelPagina 4 — La Comunità & EditorialePagina 4 — La Comunitaa & Edito

V. Signaux & opinionSignals & OpinionSignale & MeinungSegnali & opinioneSegnai & opinion

Analyse

Analysis

Analyse

Analisi

Analisi

« OpenAI est bien placé pour copier Jev » : la thèse qui agite la communauté"OpenAI Is Well Positioned to Copy Jev": The Thesis Stirring the Community«OpenAI ist gut positioniert, Jev zu kopieren»: die These, die die Community bewegt«OpenAI è ben posizionato per copiare Jev»: la tesi che agita la comunità« OpenAI l'è ben mettuu per copià Jev »: la tesi che la fa s'ciopà la comunitaa

Un billet d'Arcturus Labs publié le 22 septembre estime qu'OpenAI est bien placé pour répliquer à Jev, la famille de modèles de décision qui renvoie des choix bornés plutôt que du texte. La discussion Hacker News (275 points, 199 commentaires) oppose vitesse d'exécution des géants et avantage structurel des pionniers.
A post by Arcturus Labs published on September 22 argues that OpenAI is well positioned to replicate Jev, the family of decision models that returns bounded choices rather than text. The Hacker News discussion (275 points, 199 comments) pits the giants' execution speed against the pioneers' structural advantage.
Ein Beitrag von Arcturus Labs vom 22. September urteilt, dass OpenAI gut positioniert ist, auf Jev zu replizieren, jene Familie von Entscheidungsmodellen, die begrenzte Auswahlwerte statt Text zurückgibt. Die Hacker-News-Diskussion (275 Punkte, 199 Kommentare) stellt die Umsetzungsgeschwindigkeit der Giganten dem strukturellen Vorsprung der Pioniere gegenüber.
Un post di Arcturus Labs pubblicato il 22 settembre ritiene che OpenAI sia ben posizionato per replicare a Jev, la famiglia di modelli decisionali che restituisce scelte limitate invece di testo. La discussione Hacker News (275 punti, 199 commenti) contrappone velocità di esecuzione dei giganti e vantaggio strutturale dei pionieri.
On post de Arcturus Labs publicaa el 22 de setember 'l stimi che OpenAI l'è ben mettuu per replicà a Jev, la familia de modei de decision che la rimanda di scernì limitaa inveci che test. La discussion Hacker News (275 pont, 199 comenti) la met a confront la velocità de esecuzion di gigant e 'l vantagg strutural di pionee.

Communauté

Community

Community

Comunità

Comunitaa

JevBench : classer les modèles de décision typés sur 534 questionsJevBench: Ranking Typed Decision Models Across 534 QuestionsJevBench: getypte Entscheidungsmodelle über 534 Fragen geranktJevBench: classificare i modelli decisionali tipizzati su 534 domandeJevBench: mett in orden i modei de decision tipaa sora 534 domand

JevBench, publié le 22 septembre sur Benchmark Heaven, évalue 534 décisions anglaises et combine intelligence corrigée du hasard, calibration, vitesse et coût. Classement du jour : Jev 74,4 ; SemIf 73,1 ; djev 73,0 ; Winnow-12B Q8 71,2 ; reflex 4B 70,3. Harness MIT, items publics et code de scoring ouverts sur GitHub.
JevBench, published on September 22 on Benchmark Heaven, evaluates 534 English decisions and combines chance-corrected intelligence, calibration, speed and cost. Today's ranking: Jev 74.4; SemIf 73.1; djev 73.0; Winnow-12B Q8 71.2; reflex 4B 70.3. MIT-licensed harness, public items and open scoring code on GitHub.
JevBench, publiziert am 22. September auf Benchmark Heaven, evaluiert 534 englische Entscheidungen und kombiniert zufallskorrigierte Intelligenz, Kalibrierung, Geschwindigkeit und Kosten. Ranking des Tages: Jev 74,4; SemIf 73,1; djev 73,0; Winnow-12B Q8 71,2; reflex 4B 70,3. MIT-Harness, öffentliche Items und offener Scoring-Code auf GitHub.
JevBench, pubblicato il 22 settembre su Benchmark Heaven, valuta 534 decisioni in inglese e combina intelligenza corretta dal caso, calibrazione, velocità e costo. Classifica del giorno: Jev 74,4; SemIf 73,1; djev 73,0; Winnow-12B Q8 71,2; reflex 4B 70,3. Harness MIT, item pubblici e codice di scoring aperti su GitHub.
JevBench, publicaa el 22 de setember sora Benchmark Heaven, 'l valuta 534 decision ingles e 'l combina intelligenza corregida del cas, calibradura, velocità e cust. Classifega del dì: Jev 74,4; SemIf 73,1; djev 73,0; Winnow-12B Q8 71,2; reflex 4B 70,3. Harness MIT, item publich e codes de scoring dervii sora GitHub.

Écosystème

Ecosystem

Ökosystem

Ecosistema

Ecosistema

Le rapport de forces des modèles ouverts, passé au cribleThe Open-Model Balance of Power, Put Under the MicroscopeDas Kräfteverhältnis der offenen Modelle unter der LupeIl rapporto di forza dei modelli aperti, passato al setaccioEl raport de forza di modei dervii, passaa al cribee

La newsletter Interconnects dresse le 22 septembre 2026 un état des lieux du rapport de force en modèles ouverts, relayé sur Hacker News ; un point de repère utile après trois semaines marquées par les sorties fermées d'OpenAI, d'Anthropic et d'xAI.
The Interconnects newsletter published on September 22, 2026 an assessment of the balance of power in open models, relayed on Hacker News; a useful reference point after three weeks marked by closed releases from OpenAI, Anthropic and xAI.
Der Newsletter Interconnects zeichnet am 22. September 2026 eine Bestandsaufnahme des Kräfteverhältnisses bei offenen Modellen, verbreitet auf Hacker News; ein nützlicher Orientierungspunkt nach drei Wochen, die von den geschlossenen Releases von OpenAI, Anthropic und xAI geprägt waren.
La newsletter Interconnects traccia il 22 settembre 2026 un quadro del rapporto di forza nei modelli aperti, ripreso su Hacker News; un punto di riferimento utile dopo tre settimane segnate dalle uscite chiuse di OpenAI, Anthropic e xAI.
La newsletter Interconnects la fa el 22 de setember 2026 on ritafoeul del raport de forza in di modei dervii, repescada sora Hacker News; on pont de riferiment util dopo tri seteman segnaa di sortid saraa de OpenAI, de Anthropic e de xAI.

VI. ÉditoEditorialLeitartikelEditorialeEdito

Édito

Editorial

Leitartikel

Editoriale

Edito

Édito — La semaine où la frontière a changé d'étalon : le prix du tokenEditorial — The Week the Frontier Changed Its Yardstick: The Price of a TokenLeitartikel — Die Woche, in der die Frontier den Massstab wechselte: der Preis des TokensEditoriale — La settimana in cui la frontiera ha cambiato misura: il prezzo del tokenEdito — La setemana che la frontera l'ha cambiaa etalon: el pres del token

Deux dates, le 21 et le 22 septembre 2026, résument la stratégie du secteur : Grok 4.7 « deux fois plus rapide, à moitié prix », puis Claude Opus 5.5 au niveau de Fable 5.1 pour 40 % de moins qu'Opus 5, puis GPT-6 Sol et Luna séparés par un facteur 20 sur le prix de sortie. La frontière ne se joue plus seulement en intelligence — elle se joue au tarif, comme en attestent les améliorations du cache de prompts annoncées par OpenAI le 22 septembre. C'est une bonne nouvelle pour les utilisateurs, à une condition : que la redevabilité progresse au même rythme, un chantier qu'OpenAI dit vouloir ouvrir aux évaluations tierces indépendantes. Le prix baisse ; la vigilance, elle, ne doit pas.
Two dates, September 21 and 22, 2026, sum up the industry's strategy: Grok 4.7 "twice as fast, at half the price," then Claude Opus 5.5 at the level of Fable 5.1 for 40% less than Opus 5, then GPT-6 Sol and Luna separated by a factor of 20 on output price. The frontier is no longer being contested on intelligence alone — it is being contested on price, as evidenced by the prompt caching improvements announced by OpenAI on September 22. This is good news for users, on one condition: that accountability advances at the same pace, a project OpenAI says it wants to open to independent third-party assessments. The price is falling; vigilance must not.
Zwei Daten, der 21. und der 22. September 2026, fassen die Strategie der Branche zusammen: Grok 4.7 «zweimal schneller, zum halben Preis», dann Claude Opus 5.5 auf Fable-5.1-Niveau für 40 Prozent weniger als Opus 5, schliesslich GPT-6 Sol und Luna, getrennt durch einen Faktor 20 beim Ausgabepreis. Die Frontier entscheidet sich nicht mehr nur über Intelligenz – sie entscheidet sich über den Tarif, wie die Verbesserungen des Prompt-Cachings belegen, die OpenAI am 22. September ankündigte. Für die Nutzer ist das eine gute Nachricht – unter einer Bedingung: dass die Rechenschaftspflicht im gleichen Tempo Fortschritte macht, ein Vorhaben, das OpenAI für unabhängige Drittevaluationen öffnen will. Der Preis sinkt; die Wachsamkeit darf es nicht.
Due date, il 21 e il 22 settembre 2026, riassumono la strategia del settore: Grok 4.7 «due volte più veloce, a metà prezzo», poi Claude Opus 5.5 al livello di Fable 5.1 per il 40% in meno di Opus 5, poi GPT-6 Sol e Luna separati da un fattore 20 sul prezzo di uscita. La frontiera non si gioca più solo sull'intelligenza — si gioca sulla tariffa, come attestano i miglioramenti del cache dei prompt annunciati da OpenAI il 22 settembre. È una buona notizia per gli utenti, a una condizione: che la rendicontabilità progredisca allo stesso ritmo, un cantiere che OpenAI dichiara di voler aprire alle valutazioni terze indipendenti. Il prezzo scende; la vigilanza, invece, non deve.
Do dat, el 21 e 'l 22 de setember 2026, i riassen la strategia del setor: Grok 4.7 « do voeult pussee svelt, a la mitaa del pres », poeu Claude Opus 5.5 al nivell de Fable 5.1 per el 40 % de manch de Opus 5, poeu ancamò GPT-6 Sol e Luna separaa de on fattor 20 sora el pres de sortida. La frontera la se giuga no domà in intelligenza — la se giuga al tariff, come disen i migliorament del cache di prompt anunziaa de OpenAI el 22 de setember. L'è ona bona notizia per i utent, a ona condizion: che la responsabilizzazion la vaga innanz al midemm ritm, on cantier che OpenAI 'l dis de vorè dervì ai valutazion terz independent. El pres 'l va giò; la vigilanza, quella lì, la gh'ha de andà no.