The Neuron Times

All the AI that's fit to print

N° 253 Édition du matinMorning EditionMorgenausgabeEdizione del mattinoEdizion del mattin · Genève JEUDI 10 SEPTEMBRE 2026THURSDAY, 10 SEPTEMBER 2026DONNERSTAG, 10. SEPTEMBER 2026GIOVEDÌ 10 SETTEMBRE 2026GIOVEDÌ 10 SETTEMBRE 2026

À la Une · Frontière des modèlesFront Page · Model FrontierSchlagzeilen · ModellgrenzePrima pagina · Frontiera dei modelliIn prima pagina · Frontiera di modei

GPT-6 Astra : OpenAI déploie sa nouvelle génération de modèles pour le travailGPT-6 Astra: OpenAI Rolls Out Its New Generation of Work ModelsGPT-6 Astra: OpenAI rollt seine neue Modellgeneration für die Arbeit ausGPT-6 Astra: OpenAI distribuisce la sua nuova generazione di modelli per il lavoroGPT-6 Astra: OpenAI el desplega la soa noeuva generazion de modei per el laurà

Le laboratoire dévoile son modèle le plus capable pour l'entreprise, avec raisonnement avancé, contrôle d'ordinateur et meilleur jugement rédactionnel — et l'étend immédiatement à ChatGPT Voice.The lab unveils its most capable enterprise model, with advanced reasoning, computer control and better editorial judgment — and extends it immediately to ChatGPT Voice.Das Labor stellt sein fähigstes Modell für Unternehmen vor – mit fortgeschrittenem Reasoning, Computernutzung und besserem redaktionellen Urteilsvermögen – und dehnt es umgehend auf ChatGPT Voice aus.Il laboratorio svela il suo modello più capace per l'impresa, con ragionamento avanzato, controllo del computer e miglior giudizio redazionale — e lo estende immediatamente a ChatGPT Voice.El laboratori el descovriss el sò modell pussee capable per l'impresa, con raisoneiment avanzaa, contròll del computer e on giudezi redazzional mej — e el d'estend subet a ChatGPT Voice.

Le 9 septembre 2026, OpenAI a annoncé GPT-6 Astra, présenté par le laboratoire comme son modèle le plus capable pour les usages professionnels. La version, décrite dans une publication officielle, combine un raisonnement avancé, l'usage d'ordinateur (computer use) et un jugement renforcé en rédaction et en conception visuelle. Cette sortie confirme la stratégie d'OpenAI de viser explicitement le segment « entreprise », là où se joue une part croissante de la concurrence frontière.On September 9, 2026, OpenAI announced GPT-6 Astra, presented by the lab as its most capable model for professional use. The release, described in an official publication, combines advanced reasoning, computer use, and strengthened editorial and visual-design judgment. The launch confirms OpenAI's strategy of explicitly targeting the enterprise segment, where a growing share of the frontier competition is now being played out.Am 9. September 2026 kündigte OpenAI GPT-6 Astra an, das vom Labor als sein fähigstes Modell für professionelle Anwendungen präsentiert wird. Die Version, beschrieben in einer offiziellen Veröffentlichung, kombiniert fortgeschrittenes Reasoning, Computernutzung (Computer Use) sowie ein gestärktes Urteilsvermögen in Redaktion und visuellem Design. Diese Freigabe bestätigt die Strategie OpenAIs, explizit auf das Segment «Enterprise» zu zielen, wo ein wachsender Teil des Wettbewerbs an der Modellgrenze ausgetragen wird.Il 9 settembre 2026 OpenAI ha annunciato GPT-6 Astra, presentato dal laboratorio come il suo modello più capace per gli usi professionali. La versione, descritta in una pubblicazione ufficiale, combina un ragionamento avanzato, l'uso del computer (computer use) e un giudizio rafforzato nella scrittura e nel design visivo. Questo lancio conferma la strategia di OpenAI di mirare esplicitamente al segmento «enterprise», dove si gioca una parte crescente della concorrenza di frontiera.El 9 de setember 2026, OpenAI l'ha anunziaa GPT-6 Astra, presentaa del laboratori come el sò modell pussee capable per i us professionai. La version, descrivuda in d'ona publegazion ufizial, la combina on raisoneiment avanzaa, l'us del computer (computer use) e on giudezi rinforzaa in redazion e in concezion visual. 'Sta sortida chì la conferma la strategia de OpenAI de mirà esplicittament al segment « entreprise », indova che se giuga ona part semper pussee granda de la competizion de frontiera.

L'annonce s'accompagne d'une extension immédiate dans les produits grand public : toujours le 9 septembre 2026, les notes de version de ChatGPT indiquent que ChatGPT Voice peut désormais recourir à GPT-5.6 ou à GPT-6 Astra lorsqu'il doit rechercher ou raisonner sur des questions difficiles, avec les mêmes contrôles de modèle et d'effort de raisonnement que le chat textuel.The announcement comes with an immediate extension into consumer products: also on September 9, 2026, ChatGPT release notes indicate that ChatGPT Voice can now use GPT-5.6 or GPT-6 Astra when it needs to search or reason through difficult questions, with the same model and reasoning-effort controls as the text chat.Die Ankündigung geht mit einer sofortigen Ausweitung auf die Konsumentenprodukte einher: Ebenfalls am 9. September 2026 zeigen die Release Notes von ChatGPT, dass ChatGPT Voice nun auf GPT-5.6 oder GPT-6 Astra zurückgreifen kann, wenn es schwierige Fragen recherchieren oder durchdenken muss – mit denselben Kontrollen für Modell und Reasoning-Aufwand wie im textbasierten Chat.L'annuncio è accompagnato da un'estensione immediata nei prodotti per il grande pubblico: sempre il 9 settembre 2026, le note di rilascio di ChatGPT indicano che ChatGPT Voice può ormai ricorrere a GPT-5.6 o a GPT-6 Astra quando deve ricercare o ragionare su questioni difficili, con gli stessi controlli di modello e di sforzo di ragionamento della chat testuale.L'anunzi el va insema a 'n'estension imediata in di prodott per el grand pubblegh: semper el 9 de setember 2026, i not de version de ChatGPT disen che ChatGPT Voice el pò ades doperaa GPT-5.6 o GPT-6 Astra quand che el gh'ha de cercà o de raisonee sora di domann difficii, con i midemm contròll de modell e de sfòrz de raisoneiment del chat testual.

Le lancement intervient dans une actualité dense pour le laboratoire : le même jour, Chris Lehane y publiait une tribune appelant à agir « pendant que la fenêtre politique de l'IA est ouverte », tandis que Paul Christiano, figure de l'alignement, rejoignait le conseil d'administration de la OpenAI Foundation et son comité sécurité. Capacités produit et cadrage institutionnel avancent désormais au même rythme.The launch comes amid a dense news cycle for the lab: the same day, Chris Lehane published an essay there calling for action "while the AI policy window is open," while Paul Christiano, a leading alignment figure, joined the board of the OpenAI Foundation and its safety committee. Product capabilities and institutional framing are now advancing in lockstep.Der Launch fällt in eine dichte Nachrichtenlage des Labors: Am selben Tag veröffentlichte Chris Lehane dort einen Meinungsbeitrag mit dem Aufruf zu handeln, «solange das politische Fenster der KI offen ist», während Paul Christiano, eine zentrale Figur des Alignment, in den Verwaltungsrat der OpenAI Foundation sowie in deren Sicherheitskomitee eintrat. Produktkapazitäten und institutionelle Rahmensetzung bewegen sich inzwischen im selben Takt.Il lancio arriva in un'attualità fitta per il laboratorio: lo stesso giorno Chris Lehane vi pubblicava un'opinione che invita ad agire «finché la finestra politica dell'IA è aperta», mentre Paul Christiano, figura centrale dell'allineamento, entrava nel consiglio di amministrazione della OpenAI Foundation e nel suo comitato sicurezza. Capacità di prodotto e inquadramento istituzionale avanzano ormai allo stesso ritmo.El lanç el riva in d'on moment dens per el laboratori: el midemm dì, Chris Lehane l'ha publicaa lee on'opinion che la ciamava a agì « intanta che la fenestra politega de l'IA l'è dervida », intant che Paul Christiano, figura de l'alignment, el rivava in del consei de ministrazzion de la OpenAI Foundation e in del sò comitaa sicurezza. Capacità de prodott e quader istituzzional ades van innanz al midemm ritm.

Page 1 — Page 1 — Seite 1 — Pagina 1 — Pagina 1 — À la Une — Modèles & FrontièreFront Page — Models & FrontierFrontseite — Modelle & GrenzePrima pagina — Modelli & FrontieraIn Prima Pagina — Modei & Frontiera

I. FrontièreFrontierGrenzeFrontieraFrontiera

Politique & sécurité

Policy & Security

Politik & Sicherheit

Politica & sicurezza

Politega & sicurezza

OpenAI : « la fenêtre politique de l'IA est ouverte, il faut agir »OpenAI: "The AI policy window is open — we must act"OpenAI: «Das politische Fenster der KI ist offen, jetzt gilt es zu handeln»OpenAI: «la finestra politica dell'IA è aperta, bisogna agire»OpenAI: « la fenestra politega de l'IA l'è dervida, gh'è de fà »

Chris Lehane plaide le 9 septembre 2026 que des capacités IA plus fortes exigent des preuves de sécurité plus solides, des standards partagés et une action politique durable « pendant que la fenêtre est ouverte ». La publication se veut un appel aux décideurs plus qu'un bilan technique : l'essai complet relie explicitement accélération des modèles et nécessité de garde-fous institutionnels.
Chris Lehane argues on September 9, 2026 that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action "while the window is open." The publication is meant as a call to policymakers more than a technical review: the full essay explicitly links accelerating models with the need for institutional guardrails.
Chris Lehane argumentiert am 9. September 2026, dass stärkere KI-Fähigkeiten solidere Sicherheitsnachweise, gemeinsame Standards und dauerhaftes politisches Handeln erfordern – «solange das Fenster offen ist». Die Veröffentlichung versteht sich eher als Appell an Entscheidungsträger denn als technische Bilanz: der vollständige Essay verknüpft explizit die Beschleunigung der Modelle mit der Notwendigkeit institutioneller Guardrails.
Chris Lehane sostiene il 9 settembre 2026 che capacità IA più forti richiedono prove di sicurezza più solide, standard condivisi e un'azione politica duratura «finché la finestra è aperta». La pubblicazione vuole essere un appello ai decisori più che un bilancio tecnico: il saggio completo collega esplicitamente l'accelerazione dei modelli e la necessità di garanzie istituzionali.
Chris Lehane el sostegn el 9 de setember 2026 che di capacità IA pussee fort esigen di prov de sicurezza pussee solida, di standard condivisuu e 'n'azion politega durabella « intanta che la fenestra l'è dervida ». La publegazion la se voeur vess on'insorgia ai decisòr pussee che on bilanci tecnegh: el saggi complet el lega esplicittament accelerazion di modei e necessità di garanzia istituzzionai.

Gouvernance

Governance

Governance

Governance

Governanza

Paul Christiano entre au conseil de la OpenAI FoundationPaul Christiano Joins the OpenAI Foundation BoardPaul Christiano tritt in den Rat der OpenAI Foundation einPaul Christiano entra nel consiglio della OpenAI FoundationPaul Christiano el va dent in del consei de la OpenAI Foundation

Paul Christiano, chercheur central de l'alignement, rejoint le 9 septembre 2026 le conseil d'administration de la OpenAI Foundation ainsi que son comité Sécurité et sécurité (Safety and Security Committee), y apportant son expérience en alignement et standards. Un signal de réinjection d'expertise sécurité au plus haut niveau de la gouvernance du laboratoire, détaillé dans l'annonce officielle.
Paul Christiano, a central alignment researcher, joins on September 9, 2026 the board of the OpenAI Foundation as well as its Safety and Security Committee, bringing his expertise in alignment and standards. A signal of safety expertise being reinjected at the highest level of the lab's governance, detailed in the official announcement.
Paul Christiano, zentraler Alignment-Forscher, tritt am 9. September 2026 dem Verwaltungsrat der OpenAI Foundation sowie deren Safety and Security Committee bei und bringt seine Erfahrung in Alignment und Standards ein. Ein Signal für die Rückführung von Sicherheitsexpertise auf die höchste Governance-Ebene des Labors, wie die offizielle Ankündigung ausführt.
Paul Christiano, ricercatore centrale dell'allineamento, entra il 9 settembre 2026 nel consiglio di amministrazione della OpenAI Foundation e nel suo Safety and Security Committee, portandovi la sua esperienza in allineamento e standard. Un segnale di reiniezione di competenze in materia di sicurezza ai vertici della governance del laboratorio, dettagliato in l'annuncio ufficiale.
Paul Christiano, resercador central de l'alignment, el se gionta el 9 de setember 2026 al consei de ministrazzion de la OpenAI Foundation e al sò comitaa Safety and Security Committee, portandgh la soa esperienza in alignment e standard. On segnall de re-iniezion de competenza sicurezza al nivell pussee volt de la governanza del laboratori, detaliaa in l'anunzi ufizial.

Science appliquée

Applied Science

Angewandte Wissenschaft

Scienza applicata

Scienza aplicada

Google Research et NASA JPL cartographient le méthane depuis l'espaceGoogle Research and NASA JPL Map Methane From SpaceGoogle Research und NASA JPL kartieren Methan aus dem AllGoogle Research e NASA JPL mappano il metano dallo spazioGoogle Research e NASA JPL mappen el metan dal spazzi

From the Wires — 9 septembre 2026

From the Wires — 9 September 2026

From the Wires — 9. September 2026

Dagli studi — From the Wires — 9 settembre 2026

From the Wires — 9 settembre 2026

Google et le JPL de la NASA ont développé un modèle d'apprentissage profond qui cartographie et quantifie les émissions mondiales de méthane depuis l'espace à partir de l'instrument EMIT, annonce le 9 septembre 2026 le blog Google Research. Une illustration concrète des modèles fondation appliqués au suivi climatique satellitaire.
Google and NASA's JPL have developed a deep learning model that maps and quantifies global methane emissions from space using the EMIT instrument, the Google Research blog announced on September 9, 2026. A concrete illustration of foundation models applied to satellite climate monitoring.
Google und das JPL der NASA haben ein Deep-Learning-Modell entwickelt, das auf Basis des Instruments EMIT weltweite Methanemissionen aus dem All kartiert und quantifiziert, meldet der Google-Research-Blog am 9. September 2026. Ein konkretes Beispiel für Foundation Models im satellitengestützten Klimamonitoring.
Google e il JPL della NASA hanno sviluppato un modello di apprendimento profondo che mappa e quantifica le emissioni globali di metano dallo spazio a partire dallo strumento EMIT, annuncia il 9 settembre 2026 il blog Google Research. Un'illustrazione concreta dei modelli di fondazione applicati al monitoraggio climatico satellitare.
Google e 'l JPL de la NASA hann desvilupaa on modell de deep learning che 'l mappa e 'l quantifega i emission mondiai de metan dal spazzi a partì de l'istrument EMIT, el nunzia el 9 de setember 2026 el blog Google Research. Ona ilustrazion concreta di modei fondazion aplicaa al monitoragg clamategh satellitar.

Création

Creation

Kreation

Creazione

Creazion

« Love, Rendered » : 70 ans d'histoire d'amour régénérés image par image"Love, Rendered": 70 Years of a Love Story Regenerated Frame by Frame«Love, Rendered»: 70 Jahre Liebesgeschichte Bild für Bild regeneriert«Love, Rendered»: 70 anni di storia d'amore rigenerati immagine per immagine« Love, Rendered »: 70 ann de storia d'amor regeneraa imagin per imagin

From the Wires — 9 septembre 2026

From the Wires — 9 September 2026

From the Wires — 9. September 2026

Dagli studi — From the Wires — 9 settembre 2026

From the Wires — 9 settembre 2026

Des cinéastes et Google DeepMind ont utilisé l'IA pour recréer image par image le passé non filmé d'un couple, sur 70 ans d'histoire d'amour, dans le court-métrage « Love, Rendered », raconte Google le 9 septembre 2026. Un cas d'école de la génération vidéo appliquée à la mémoire documentaire.
Filmmakers and Google DeepMind used AI to recreate, frame by frame, the unfilmed past of a couple across 70 years of love story, in the short film "Love, Rendered," Google reports on September 9, 2026. A case study in video generation applied to documentary memory.
Filmemacher und Google DeepMind haben mit KI die nicht gefilmte Vergangenheit eines Paares Bild für Bild über 70 Jahre Liebesgeschichte nachempfunden – im Kurzfilm «Love, Rendered», berichtet Google am 9. September 2026. Ein Lehrstück für die Anwendung von Videogenerierung auf dokumentarische Erinnerung.
Dei registi e Google DeepMind hanno usato l'IA per ricreare immagine per immagine il passato non filmato di una coppia, su 70 anni di storia d'amore, nel cortometraggio «Love, Rendered», racconta Google il 9 settembre 2026. Un caso di scuola della generazione video applicata alla memoria documentaria.
Di regista e Google DeepMind hann doperaa l'IA per ricreà imagin per imagin el passaa minga filmaa d'ona cobbia, sora 70 ann de storia d'amor, in del cort « Love, Rendered », el cunta Google el 9 de setember 2026. On cas de scola de la generazion video aplicada a la memoria documentaria.

Page 2 — Page 2 — Seite 2 — Pagina 2 — Pagina 2 — Le Cahier TechniqueThe Technical DeskDas Technische DossierIl Quaderno TecnicoEl Quadern Tecnegh

II. Harnais & OutilsHarnesses & ToolsHarness & WerkzeugeHarness & StrumentiHarnass & Strument

Étude de cas

Case Study

Fallstudie

Caso di studio

Studi de cas

Mistral fait migrer 40 000 lignes de Fortran 77 vers C++ par agentsMistral Migrates 40,000 Lines of Fortran 77 to C++ with AgentsMistral migriert 40 000 Zeilen Fortran 77 nach C++ mit AgentenMistral fa migrare 40 000 righe di Fortran 77 verso C++ tramite agentiMistral el fa migrà 40 000 lign de Fortran 77 vers C++ con di agent

L'expérience de migration publiée le 9 septembre 2026 montre comment Mistral a aidé un opérateur énergétique européen à migrer 40 000 lignes de Fortran 77 vers C++ à l'aide d'agents IA. Le récit détaille la méthode et les leçons retenues pour la modernisation du code historique, un des chantiers les plus coûteux de l'informatique industrielle.
The migration experiment published on September 9, 2026 shows how Mistral helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ using AI agents. The account details the method and lessons learned for modernizing legacy code, one of the costliest undertakings in industrial computing.
Das am 9. September 2026 veröffentlichte Migrationsexperiment zeigt, wie Mistral einem europäischen Energieversorger half, 40 000 Zeilen Fortran 77 nach C++ zu migrieren – mithilfe von KI-Agenten. Der Bericht schildert Methode und Erkenntnisse für die Modernisierung von Altsystemen, eines der teuersten Vorhaben der industriellen Informatik.
L'esperimento di migrazione pubblicato il 9 settembre 2026 mostra come Mistral ha aiutato un operatore energetico europeo a migrare 40 000 righe di Fortran 77 verso C++ con l'aiuto di agenti IA. Il racconto dettaglia il metodo e le lezioni apprese per la modernizzazione del codice storico, uno dei cantieri più costosi dell'informatica industriale.
L'esperiment de migrazion publicaa el 9 de setember 2026 el mostra come Mistral l'ha juttaa on operator energedegh europee a migrà 40 000 lign de Fortran 77 vers C++ con l'ajutt de agent IA. El cunt el detalia el metod e i lession ciapaa per la modernizzazion del codes storigh, vun di cantee pussee car de l'infarmatega industrial.

Modèles open-weight

Open-Weight Models

Open-Weight-Modelle

Modelli open-weight

Modei open-weight

IBM libère Granite Time Series PatchTST-r2, SOTA sous licence commercialeIBM Releases Granite Time Series PatchTST-r2, SOTA Under a Commercial LicenseIBM gibt Granite Time Series PatchTST-r2 frei – SOTA unter kommerzieller LizenzIBM rilascia Granite Time Series PatchTST-r2, SOTA sotto licenza commercialeIBM el libera Granite Time Series PatchTST-r2, SOTA sota lizenza comerciai

From the Wires — 9 septembre 2026

From the Wires — 9 September 2026

From the Wires — 9. September 2026

Dagli studi — From the Wires — 9 settembre 2026

From the Wires — 9 settembre 2026

IBM publie le 9 septembre 2026 Granite Time Series PatchTST-r2, un modèle de séries temporelles présenté comme de l'état de l'art, sous licence compatible usage commercial, annonce le blog IBM Research sur Hugging Face. Un pas de plus pour la prévision open-weight en environnement d'entreprise.
IBM released on September 9, 2026 Granite Time Series PatchTST-r2, a time-series model presented as state of the art, under a license compatible with commercial use, the IBM Research blog on Hugging Face announced. Another step forward for open-weight forecasting in enterprise settings.
IBM veröffentlicht am 9. September 2026 Granite Time Series PatchTST-r2, ein Zeitreihenmodell mit beanspruchtem State of the Art unter einer kommerziell nutzbaren Lizenz, wie das IBM-Research-Blog auf Hugging Face meldet. Ein weiterer Schritt für Open-Weight-Forecasting in Unternehmensumgebungen.
IBM pubblica il 9 settembre 2026 Granite Time Series PatchTST-r2, un modello di serie temporali presentato come allo stato dell'arte, sotto licenza compatibile con l'uso commerciale, annuncia il blog IBM Research su Hugging Face. Un passo in più per la previsione open-weight negli ambienti aziendali.
IBM el publega el 9 de setember 2026 Granite Time Series PatchTST-r2, on modell de seri tempora presentaa come de l'art, sota lizenza compatible con l'us comerciai, el nunzia el blog IBM Research sora Hugging Face. On pass pussee inanz per la prevision open-weight in ambient d'impresa.

Ingénierie agents

Agent Engineering

Agenten-Engineering

Ingegneria degli agenti

Ingegneria di agent

Co-évoluer harnais et modèles : l'imitation de l'expert peut coûter 30 pointsCo-Evolving Harness and Models: Imitating the Expert Can Cost 30 PointsHarness und Modelle ko-evolvieren: Imitation des Experten kann 30 Punkte kostenCo-evolvere harness e modelli: l'imitazione dell'esperto può costare 30 puntiCo-evolv harnass e modei: l'imitazion de l'espert la po costà 30 pont

Une analyse de Salesforce confrontée le 10 septembre 2026 sur HF Daily Papers révèle une tension centrale des systèmes agents : faire évoluer le harnais avec le modèle faible améliore les performances, mais imiter les trajectoires complètes d'un expert fait régresser les 7 tâches testées de 4 à 30 points sur Qwen3-Coder et Gemma 4. Les auteurs proposent une correction on-policy qui réécrit uniquement le tour défaillant du rollout, préservant le style de planification natif du modèle.
A Salesforce analysis surfaced on HF Daily Papers on September 10, 2026 reveals a central tension in agentic systems: co-evolving the harness with the weak model improves performance, but imitating an expert's full trajectories causes regressions of 4 to 30 points on all 7 tasks tested on Qwen3-Coder and Gemma 4. The authors propose an on-policy correction that rewrites only the failing turn of the rollout, preserving the model's native planning style.
Eine am 10. September 2026 auf HF Daily Papers diskutierte Analyse von Salesforce legt eine zentrale Spannung von Agentensystemen offen: Das Weiterentwickeln des Harnesses zusammen mit dem schwächeren Modell verbessert die Leistung, doch das Imitieren vollständiger Experten-Trajektorien lässt die 7 getesteten Aufgaben um 4 bis 30 Punkte zurückfallen auf Qwen3-Coder und Gemma 4. Die Autoren schlagen eine On-Policy-Korrektur vor, die ausschliesslich den fehlerhaften Zug des Rollouts neu schreibt und den natürlichen Planungsstil des Modells bewahrt.
Un'analisi di Salesforce esaminata il 10 settembre 2026 su HF Daily Papers rivela una tensione centrale dei sistemi agentici: far evolvere l'harness insieme al modello debole migliora le prestazioni, ma imitare le traiettorie complete di un esperto fa regredire i 7 compiti testati da 4 a 30 punti su Qwen3-Coder e Gemma 4. Gli autori propongono una correzione on-policy che riscrive solo la mossa difettosa del rollout, preservando lo stile di pianificazione nativo del modello.
On'analisi de Salesforce confrontada el 10 de setember 2026 sora HF Daily Papers la revela ona tenszion central di sistema agent: fà andà innanz el harnass cont el modell deboi el mejora i prestazion, ma imità i trajettori complet d'on espert el fa regredì i 7 prov testaa de 4 a 30 pont sora Qwen3-Coder e Gemma 4. I autor proponen 'na correzzion on-policy che la rescriv domà el turn fallii del rollout, conservand el stil de pianificazion nativ del modell.

Page 3 — Page 3 — Seite 3 — Pagina 3 — Pagina 3 — La RechercheThe Research DeskDie ForschungLa RicercaLa Ricerca

III. Papers du jourPapers of the DayPapers des TagesPaper del giornoPaper del dì

Robotique

Robotics

Robotik

Robotica

Robotega

Show-Harness : un simple agent VLM peut « jouer » aux robotsShow-Harness: A Simple VLM Agent Can "Play" RobotsShow-Harness: Ein einfacher VLM-Agent kann Roboter «spielen»Show-Harness: un semplice agente VLM può «giocare» ai robotShow-Harness: on sempligh agent VLM el po « giugà » ai robot

Show-Harness, 29 votes sur HF Daily Papers le 10 septembre 2026, équipe les VLM d'une interface d'actions sémantiques que des interpréteurs déterministes traduisent en commandes robotiques. Le résultat : des VLM frontière fermés pilotent des robots en zéro-shot, et des petits VLM ouverts s'adaptent en quelques GPU-heures de fine-tuning, surpassant les paradigmes VLA testés — sans pré-entraînement spécifique à l'incarnation.
Show-Harness, 29 votes on HF Daily Papers on September 10, 2026, equips VLMs with a semantic action interface that deterministic interpreters translate into robotic commands. The result: closed frontier VLMs drive robots zero-shot, and small open VLMs adapt with a few GPU-hours of fine-tuning, outperforming the VLA paradigms tested — without embodiment-specific pre-training.
Show-Harness, am 10. September 2026 mit 29 Stimmen auf HF Daily Papers, stattet VLMs mit einer Schnittstelle semantischer Aktionen aus, die deterministische Interpreter in Roboterbefehle übersetzen. Das Ergebnis: Geschlossene Frontier-VLMs steuern Roboter Zero-shot, und kleine offene VLMs passen sich in wenigen GPU-Stunden Feintuning an – besser als die getesteten VLA-Paradigmen, ohne Inkarnation-spezifisches Vortraining.
Show-Harness, 29 voti su HF Daily Papers il 10 settembre 2026, dota i VLM di un'interfaccia di azioni semantiche che interpreti deterministici traducono in comandi robotici. Il risultato: VLM di frontiera chiusi pilotano robot in zero-shot, e piccoli VLM aperti si adattano in poche GPU-ore di fine-tuning, superando i paradigmi VLA testati — senza pre-addestramento specifico all'incarnazione.
Show-Harness, 29 vòt sora HF Daily Papers el 10 de setember 2026, el forniss ai VLM ona interfaccia d'azion semantegh che di interpret deterministegh tradden in comand robotigh. El resultaa: di VLM de frontiera saraa i guiden di robot in zéro-shot, e di piscinitt VLM vert i se adatten in quaj or de GPU de fine-tuning, superand i paradigma VLA testaa — senza pre-addestrament specifegh a l'incarnazion.

World models

World Models

World Models

World models

World models

Programmable World Model : un état du monde persistant et exécutableProgrammable World Model: A Persistent, Executable World StateProgrammable World Model: Ein persistenter, ausführbarer WeltzustandProgrammable World Model: uno stato del mondo persistente ed eseguibileProgrammable World Model: on stat del mund persistent e eseguibii

Programmable World Model (20 votes) découple l'évolution de l'état du monde de la génération visuelle : un agent traduit le langage naturel en programmes exécutables, et un moteur léger maintient un état global persistant rendu par un modèle vidéo pré-entraîné. Sur le benchmark CombatStateBench, la méthode atteint 94 % d'exactitude de comptage et 98 % d'exactitude d'état, permettant de créer des jeux jouables aux mécaniques prédéfinies.
Programmable World Model (20 votes) decouples world-state evolution from visual generation: an agent translates natural language into executable programs, and a lightweight engine maintains a persistent global state rendered by a pre-trained video model. On the CombatStateBench benchmark, the method reaches 94% counting accuracy and 98% state accuracy, making it possible to create playable games with predefined mechanics.
Programmable World Model (20 Stimmen) entkoppelt die Evolution des Weltzustands von der visuellen Generierung: Ein Agent übersetzt natürliche Sprache in ausführbare Programme, und eine leichtgewichtige Engine hält einen persistenten globalen Zustand aufrecht, den ein vortrainiertes Videomodell rendert. Auf dem Benchmark CombatStateBench erreicht die Methode 94 % Zählgenauigkeit und 98 % Zustandsgenauigkeit – spielbare Spiele mit vordefinierten Mechaniken werden möglich.
Programmable World Model (20 voti) disaccoppia l'evoluzione dello stato del mondo dalla generazione visiva: un agente traduce il linguaggio naturale in programmi eseguibili, e un motore leggero mantiene uno stato globale persistente reso da un modello video pre-addestrato. Sul benchmark CombatStateBench, il metodo raggiunge il 94% di accuratezza nel conteggio e il 98% di accuratezza di stato, permettendo di creare giochi giocabili con meccaniche predefinite.
Programmable World Model (20 vòt) el desligua l'evoluzion de l'ambient del mund de la generazion visual: on agent el tradus el lengoeu natural in program eseguibii, e on motor leger el mantegn on stat global persistent renduu de on modell video pre-adestraa. Sora el benchmark CombatStateBench, el metod el riva a 94 % de precision in del contà e 98 % de precision de stat, permettend de creà di zögh giugabil con di mecanegh definii prima.

Méthodologie

Methodology

Methodik

Metodologia

Metodologia

DCP : les scores seuls ne prouvent pas la découverteDCP: Scores Alone Do Not Prove DiscoveryDCP: Scores allein belegen keine EntdeckungDCP: i punteggi da soli non provano la scopertaDCP: i scor domà i pròven minga la scuperta

Des chercheurs de Carnegie Mellon proposent le Discovery Certification Protocol, 11 votes le 10 septembre 2026 : des tests exécutables de récupération et de feedback pour auditer les agents de recherche IA. Deux audits contrôlés (optimisation SQLite, contrôle de catalyseurs) ont produit zéro récupération en 96 épisodes, borne supérieure 0,0468, avec un vérificateur déterministe sans LLM.
Researchers at Carnegie Mellon propose the Discovery Certification Protocol, 11 votes on September 10, 2026: executable retrieval and feedback tests to audit AI research agents. Two controlled audits (SQLite optimization, catalyst screening) produced zero recoveries across 96 episodes, upper bound 0.0468, using a deterministic LLM-free verifier.
Forschende der Carnegie Mellon University schlagen das Discovery Certification Protocol vor, am 10. September 2026 mit 11 Stimmen: ausführbare Tests für Retrieval und Feedback, um KI-Forschungsagenten zu auditieren. Zwei kontrollierte Audits (SQLite-Optimierung, Katalysatorkontrolle) ergaben null Funde in 96 Episoden, obere Schranke 0,0468 – mit deterministischem Verifikator ohne LLM.
Dei ricercatori di Carnegie Mellon propongono il Discovery Certification Protocol, 11 voti il 10 settembre 2026: test eseguibili di recupero e di feedback per verificare gli agenti di ricerca IA. Due audit controllati (ottimizzazione SQLite, controllo di catalizzatori) hanno prodotto zero recuperi in 96 episodi, limite superiore 0,0468, con un verificatore deterministico senza LLM.
Di resercador de Carnegie Mellon i proponen el Discovery Certification Protocol, 11 vòt el 10 de setember 2026: di test eseguibii de recuver e de feedback per audità i agent de ricerca IA. Dò audit contròllaa (optimizazion SQLite, contròll de catalisador) i hann produu zèro recuver in 96 episodi, bound superior 0,0468, cont on verificator deterministegh senza LLM.

Interprétabilité

Interpretability

Interpretierbarkeit

Interpretabilità

Interpretabilitaa

Des agents IA peuvent-ils devenir chercheurs en interprétabilité ?Can AI Agents Become Interpretability Researchers?Können KI-Agenten zu Interpretationsforschern werden?Gli agenti IA possono diventare ricercatori in interpretabilità?I agent IA i pòden devegnì resercador in interpretabilitaa?

SAEScientist-Bench (5 votes) évalue si des agents peuvent mener de la recherche d'interprétabilité mécaniste autonome : sur 10 configurations et 20 tâches, les agents frontière explorent un dictionnaire Gemma Scope de 131K+ features dans Gemma-2-9B-IT et approchent le niveau expert sur la sélectivité conceptuelle, mais restent nettement en retrait sur le steering causal et interprètent mal certaines mesures expérimentales.
SAEScientist-Bench (5 votes) assesses whether agents can conduct autonomous mechanistic interpretability research: across 10 configurations and 20 tasks, frontier agents explore a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT and approach expert level on concept selectivity, but still lag clearly behind on causal steering and misinterpret some experimental measures.
SAEScientist-Bench (5 Stimmen) prüft, ob Agenten selbstständige mechanistische Interpretationsforschung betreiben können: Bei 10 Konfigurationen und 20 Aufgaben erkunden Frontier-Agenten ein Gemma-Scope-Wörterbuch mit über 131 000 Features in Gemma-2-9B-IT und erreichen beim Konzeptselektivitätsniveau beinahe Expertenniveau, bleiben jedoch beim kausalen Steering deutlich zurück und deuten einzelne experimentelle Messwerte falsch.
SAEScientist-Bench (5 voti) valuta se degli agenti possano condurre ricerca di interpretabilità meccanicistica autonoma: su 10 configurazioni e 20 compiti, gli agenti di frontiera esplorano un dizionario Gemma Scope di 131K+ feature in Gemma-2-9B-IT e si avvicinano al livello esperto nella selettività concettuale, ma restano nettamente indietro sullo steering causale e interpretano male certe misure sperimentali.
SAEScientist-Bench (5 vòt) el valuta se di agent i pòden menà de la ricerca de interpretabilitaa mecanistega autonoma: sora 10 configurazion e 20 compit, i agent de frontiera i esploren on dizzionari Gemma Scope de 131K+ feature in Gemma-2-9B-IT e i se vesinen al nivell espert sora la selettività concettual, ma i resten nettement in dree sora el steering causal e i interpreten mal quaj misura sperimentala.

Page 4 — Page 4 — Seite 4 — Pagina 4 — Pagina 4 — La Communauté & ÉditoCommunity & EditorialCommunity & LeitartikelLa Comunità & EditorialeLa Comunitaa & Editòrial

IV. Signaux communautéCommunity SignalsCommunity-SignaleSegnali dalla comunitàSegnai de la comunitaa

« Claude, change the "Claude, change the «Claude, change the «Claude, change the« Claude, change the

« Claude, change the “Add to Cart” button to blue » — le 9 septembre 2026, cette critique du design des interfaces agents accumule 1 075 points et 415 commentaires sur Hacker News, la plus grosse discussion IA de la journée. Le texte, publié sur opusfived.dev, pointe la régression ergonomique des outils d'édition IA et touche une corde chez les développeurs.
"Claude, change the “Add to Cart” button to blue" — on September 9, 2026, this critique of agent interface design amassed 1,075 points and 415 comments on Hacker News, the biggest AI discussion of the day. The piece, published on opusfived.dev, targets the ergonomic regression of AI editing tools and strikes a chord with developers.
«Claude, change the “Add to Cart” button to blue» – am 9. September 2026 sammelt diese Kritik am Design von Agenten-Interfaces 1 075 Punkte und 415 Kommentare auf Hacker News, die grösste KI-Diskussion des Tages. Der auf opusfived.dev veröffentlichte Text kritisiert die ergonomische Regression von KI-Bearbeitungswerkzeugen und trifft einen Nerv bei den Entwicklern.
«Claude, change the “Add to Cart” button to blue» — il 9 settembre 2026, questa critica del design delle interfacce agentiche accumula 1 075 punti e 415 commenti su Hacker News, la più grande discussione IA della giornata. Il testo, pubblicato su opusfived.dev, denuncia la regressione ergonomica degli strumenti di editing IA e tocca un nervo scoperto tra gli sviluppatori.
« Claude, change the “Add to Cart” button to blue » — el 9 de setember 2026, 'sta critica del design di interfass agent la mett insema 1 075 pont e 415 comment sora Hacker News, la discussión IA pussee granda del dì. El test, publicaa sora opusfived.dev, el ponta a la regressione ergonomiga di strument de edizion IA e 'l tocca ona corda sensibii in di desvilupador.

Observatoire

Observatory

Observatorium

Osservatorio

Osservatori

Qwen 3.8 et les prefills de raisonnement de GPT-5.5 Pro font débatQwen 3.8 and GPT-5.5 Pro's Reasoning Prefills Spark DebateQwen 3.8 und die Reasoning-Prefills von GPT-5.5 Pro sorgen für DebatteQwen 3.8 e i prefill di ragionamento di GPT-5.5 Pro fanno discutereQwen 3.8 e i prefill de raisoneiment de GPT-5.5 Pro i fann discùt

Le 9 septembre 2026, une note publiée sur Hacker News observe que Qwen 3.8 suit les prefills de raisonnement de GPT-5.5 Pro (196 points, 77 commentaires), et l'analyse détaillée documente le comportement — relançant le débat sur la convergence des conventions de raisonnement entre modèles ouverts et fermés.
On September 9, 2026, a Hacker News post observed that Qwen 3.8 follows GPT-5.5 Pro's reasoning prefills (196 points, 77 comments), and the detailed analysis documents the behavior — reigniting the debate over converging reasoning conventions between open and closed models.
Am 9. September 2026 beobachtet eine Notiz auf Hacker News, dass Qwen 3.8 den Reasoning-Prefills von GPT-5.5 Pro folgt (196 Punkte, 77 Kommentare); die Detailanalyse dokumentiert das Verhalten – und belebt die Debatte über die Konvergenz der Reasoning-Konventionen zwischen offenen und geschlossenen Modellen neu.
Il 9 settembre 2026 una nota pubblicata su Hacker News osserva che Qwen 3.8 segue i prefill di ragionamento di GPT-5.5 Pro (196 punti, 77 commenti), e l'analisi dettagliata documenta il comportamento — rilanciando il dibattito sulla convergenza delle convenzioni di ragionamento tra modelli aperti e chiusi.
El 9 de setember 2026, ona nota publicada sora Hacker News la osserva che Qwen 3.8 el seguiss i prefill de raisoneiment de GPT-5.5 Pro (196 pont, 77 comment), e l'analisi detaliaa la documenta el comportament — cont el fait tornà sù el dibattit sora la convergenza di convenzion de raisoneiment tra modei vert e saraa.

Analyse

Analysis

Analyse

Analisi

Analisi

Raschka décortique GPT-6 Astra : transformers bouclés et raisonnement cachéRaschka Dissects GPT-6 Astra: Looped Transformers and Hidden ReasoningRaschka seziert GPT-6 Astra: Loop-Transformer und verborgenes ReasoningRaschka analizza GPT-6 Astra: transformer a ciclo e ragionamento nascostoRaschka el desquatta GPT-6 Astra: transformer a lòss e raisoneiment sconduu

L'analyse de Sebastian Raschka sur GPT-6 Astra, les transformers bouclés et le raisonnement caché recueille 376 points et 131 commentaires sur Hacker News le 9 septembre 2026. Le texte, lisible sur son magazine, décortique l'architecture présumée du nouveau modèle d'OpenAI au lendemain de son annonce.
Sebastian Raschka's analysis of GPT-6 Astra, looped transformers and hidden reasoning gathered 376 points and 131 comments on Hacker News on September 9, 2026. The piece, readable on his magazine, dissects the presumed architecture of OpenAI's new model in the wake of its announcement.
Sebastian Raschkas Analyse zu GPT-6 Astra, Loop-Transformern und verborgenem Reasoning erhält am 9. September 2026 376 Punkte und 131 Kommentare auf Hacker News. Der Text, lesbar in seinem Magazin, seziert die vermutete Architektur des neuen OpenAI-Modells unmittelbar nach dessen Ankündigung.
L'analisi di Sebastian Raschka su GPT-6 Astra, i transformer a ciclo e il ragionamento nascosto raccoglie 376 punti e 131 commenti su Hacker News il 9 settembre 2026. Il testo, leggibile sul suo magazine, analizza l'architettura presunta del nuovo modello di OpenAI il giorno dopo il suo annuncio.
L'analisi de Sebastian Raschka sora GPT-6 Astra, i transformer a lòss e 'l raisoneiment sconduu la recuj 376 pont e 131 comment sora Hacker News el 9 de setember 2026. El test, legibii sora el sò magazine, el desquatta l'architettura presumuda del noeuv modell de OpenAI el dì depos de l'anunzi.

Édito

Editorial

Leitartikel

Editoriale

Editòrial

Édito — Le jour où les capacités, la politique et l'ergonomie ont convergéEditorial — The Day Capabilities, Politics and Ergonomics ConvergedLeitartikel — Der Tag, an dem Kapazitäten, Politik und Ergonomie konvergiertenEditoriale — Il giorno in cui capacità, politica ed ergonomia sono convergentiEditòrial — El dì che capacità, politega e ergonomia i hinn confluii insema

Ce que la journée du 9 septembre 2026 dit de l'industrie : la frontière avance désormais sur trois fronts simultanés. Le produit d'abord, avec GPT-6 Astra déployé chez OpenAI du jour au lendemain jusqu'à la voix. La gouvernance ensuite, avec la tribune de Chris Lehane sur la « fenêtre politique » et l'arrivée de Paul Christiano au conseil de la OpenAI Foundation — capacités et garde-fous négociés dans le même cycle de publication. L'écosystème enfin, où un billet de blog sur la couleur d'un bouton réunit plus d'un millier de points sur Hacker News : preuve que la question n'est plus seulement ce que les agents savent faire, mais comment les humains travaillent avec eux. Le modèle devient une actualité politique, ergonomique et institutionnelle autant que technique. C'est précisément cette densité que ce journal s'attachera à suivre.
What September 9, 2026 says about the industry: the frontier is now advancing on three fronts simultaneously. First, product, with GPT-6 Astra deployed overnight by OpenAI all the way to voice. Next, governance, with Chris Lehane's essay on the "policy window" and Paul Christiano's arrival on the OpenAI Foundation board — capabilities and guardrails negotiated within the same release cycle. Finally, the ecosystem, where a blog post about the color of a button draws more than a thousand points on Hacker News: proof that the question is no longer only what agents can do, but how humans work with them. The model is becoming a political, ergonomic and institutional story as much as a technical one. It is precisely this density that this newspaper will strive to follow.
Was der 9. September 2026 über die Branche aussagt: Die Grenze bewegt sich inzwischen auf drei Fronten gleichzeitig. Zunächst das Produkt, mit GPT-6 Astra, das OpenAI über Nacht bis hin zur Stimme ausgerollt hat. Sodann die Governance, mit Chris Lehanes Beitrag zum «politischen Fenster» und Paul Christianos Eintritt in den Rat der OpenAI Foundation – Kapazitäten und Guardrails, verhandelt im selben Veröffentlichungszyklus. Schliesslich das Ökosystem, wo ein Blogeintrag über die Farbe eines Buttons über tausend Punkte auf Hacker News versammelt: ein Beleg, dass die Frage nicht mehr nur lautet, was Agenten können, sondern wie Menschen mit ihnen arbeiten. Das Modell wird zur politischen, ergonomischen und institutionellen Nachricht ebenso wie zur technischen. Genau dieser Dichte wird sich diese Zeitung widmen.
Cosa dice della industria la giornata del 9 settembre 2026: la frontiera avanza ormai su tre fronti simultanei. Prima il prodotto, con GPT-6 Astra distribuito da OpenAI dal giorno alla notte fino alla voce. Poi la governance, con l'opinione di Chris Lehane sulla «finestra politica» e l'arrivo di Paul Christiano nel consiglio della OpenAI Foundation — capacità e garanzie negoziate nello stesso ciclo di pubblicazione. Infine l'ecosistema, dove un post sul colore di un pulsante riunisce più di mille punti su Hacker News: la prova che la questione non è più solo ciò che gli agenti sanno fare, ma come gli umani lavorano con loro. Il modello diventa un'attualità politica, ergonomica e istituzionale oltre che tecnica. È proprio questa densità che questo giornale si impegherà a seguire.
Quel che 'l dì del 9 de setember 2026 el dis de l'industria: la frontiera ades la va innanz sora tri front simultani. El prodott prima, con GPT-6 Astra desplegaa a OpenAI de 'n dì a l'alter fin a la vos. La governanza depos, con la tribuna de Chris Lehane sora la « fenestra politega » e 'l rivà de Paul Christiano in del consei de la OpenAI Foundation — capacità e garanzia negoziaa in del midemm ciclo de publegazion. L'ecosistema a la fin, indova che 'n post de blog sora 'l color d'on boton el reuniss pussee de milla pont sora Hacker News: prova che la quistion l'è pu domà quel che i agent i san fà, ma come i hom i laura con lor. El modell el deventa ona noeuva politega, ergonomiga e istituzzionala tant quant tecnega. L'è propi 'sta densità chì che 'sto giornal el se proponn de seguì.