The Neuron Times

All the AI that's fit to print

N° 264 Édition du matinMorning EditionMorgenausgabeEdizione del mattinoEdizion del mattin · Genève LUNDI 21 SEPTEMBRE 2026MONDAY, 21 SEPTEMBER 2026MONTAG, 21. SEPTEMBER 2026LUNEDÌ 21 SETTEMBRE 2026LUNEDÌ 21 SETTEMBRE 2026

À la Une · Modèles de générationFront Page · Generative modelsSchlagzeilen · GenerationsmodellePrima pagina · Modelli di generazioneIn prima pagina · Modej de generazion

Qwen Image 2.1 relance la course à la génération d'images côté open weightsQwen Image 2.1 reignites the open-weights image generation raceQwen Image 2.1 belebt das Rennen um die Bildgenerierung im Open-Weights-Lager neuQwen Image 2.1 rilancia la corsa alla generazione di immagini open weightsQwen Image 2.1 el fa tornà la corsa a la generazion de imagin in del camp open weights

La nouvelle itération image de Qwen, révélée le 20 septembre 2026, provoque une discussion de grande ampleur dans la communauté du machine learning.The new Qwen image iteration, unveiled on September 20, 2026, has sparked a wide-ranging discussion across the machine learning community.Die neue Bild-Iteration von Qwen, am 20. September 2026 vorgestellt, löst eine breit angelegte Diskussion in der Machine-Learning-Gemeinschaft aus.La nuova iterazione immagine di Qwen, rivelata il 20 settembre 2026, scatena una discussione di grande portata nella comunità del machine learning.La noeuva iterazion imagin de Qwen, svelada el 20 de setember 2026, la provoca ona discussion de granda ampieza in de la comunitaa del machine learning.

Alibaba a publié le 20 septembre 2026 la version 2.1 de son modèle de génération d'images Qwen, annoncée sur le blog officiel du laboratoire. La discussion ouverte sur Hacker News autour de cette sortie a dépassé les 550 points et 160 commentaires en moins de 24 heures, un volume rare pour une annonce de modèle image open weights.On September 20, 2026, Alibaba released version 2.1 of its Qwen image generation model, announced on the lab's official blog. The Hacker News discussion that opened around the release surpassed 550 points and 160 comments in less than 24 hours — a rare volume for an open-weights image model announcement.Alibaba hat am 20. September 2026 die Version 2.1 seines Bildgenerierungsmodells Qwen veröffentlicht, angekündigt im offiziellen Blog des Labors. Die auf Hacker News eröffnete Diskussion zu dieser Veröffentlichung hat in weniger als 24 Stunden über 550 Punkte und 160 Kommentare überschritten – ein seltenes Volumen für die Ankündigung eines Open-Weights-Bildmodells.Il 20 settembre 2026 Alibaba ha pubblicato la versione 2.1 del suo modello di generazione di immagini Qwen, annunciata sul blog ufficiale del laboratorio. La discussione aperta su Hacker News in merito a questa uscita ha superato i 550 punti e i 160 commenti in meno di 24 ore, un volume raro per un annuncio di modello immagine open weights.Alibaba l'ha publicaa el 20 de setember 2026 la version 2.1 del sò modej de generazion de imagin Qwen, anunziada in sul blog offizial del laboratori. La discussión dervida in su Hacker News intorna de questa sortida l'ha superaa i 550 pont e i 160 comenter in manch de 24 or, on volum rar per ona anunzia de modej imagin open weights.

Cette sortie s'inscrit dans la stratégie soutenue de Qwen sur le terrain multimodal, après une série de déclinaisons de sa famille 3.8. La communauté s'est surtout penchée sur la qualité des rendus, les capacités d'édition et la comparaison avec les modèles propriétaires concurrents — le débat entre praticabilité locale et services fermés structure désormais chaque sortie majeure d'image.The release fits into Qwen's sustained multimodal strategy, following a series of iterations of its 3.8 family. The community focused above all on render quality, editing capabilities and comparisons with rival proprietary models — the debate between local practicality and closed services now frames every major image release.Diese Veröffentlichung reiht sich in die konsequente Multimodal-Strategie von Qwen ein, nach einer Serie von Ablegern der Familie 3.8. Die Gemeinschaft konzentrierte sich vor allem auf die Renderqualität, die Bearbeitungsfähigkeiten und den Vergleich mit konkurrierenden proprietären Modellen – die Debatte zwischen lokaler Praktikabilität und geschlossenen Diensten strukturiert inzwischen jede grössere Bildveröffentlichung.Questa uscita si inserisce nella strategia costante di Qwen sul fronte multimodale, dopo una serie di declinazioni della sua famiglia 3.8. La comunità si è concentrata soprattutto sulla qualità dei risultati, sulle capacità di editing e sul confronto con i modelli proprietari concorrenti — il dibattito tra fattibilità locale e servizi chiusi struttura ormai ogni uscita importante nel campo delle immagini.Questa sortida la se met dent in de la strategia sostegnuda de Qwen in sul terren multimodal, despoeu de ona serie de derivazion de la soa fameja 3.8. La comunitaa l'è andada soratutt sora la qualitaa di render, i capacità de modifega e el confront con i modej proprietari concorrent — el dibattit tra praticabilitaa local e servizzi saraa el strutura adess ogni sortida importanta de imagin.

Pour les observateurs de l'écosystème, le signal est double : la cadence de Qwen reste soutenable malgré la concurrence frontière, et la génération d'images open weights conserve son public auprès des développeurs qui veulent exécuter, auditer et personnaliser leurs modèles.For ecosystem observers, the signal is twofold: Qwen's cadence remains sustainable despite frontier competition, and open-weights image generation retains its audience among developers who want to run, audit and customize their models.Für Beobachter des Ökosystems ist das Signal doppelt: Das Tempo von Qwen bleibt trotz Frontale-Konkurrenz aufrecht erhalten, und die Open-Weights-Bildgenerierung behält ihr Publikum bei Entwicklern, die ihre Modelle selbst ausführen, auditieren und anpassen wollen.Per gli osservatori dell'ecosistema, il segnale è duplice: il ritmo di Qwen resta sostenibile nonostante la concorrenza di frontiera, e la generazione di immagini open weights conserva il proprio pubblico tra gli sviluppatori che vogliono eseguire, verificare e personalizzare i propri modelli.Per i osservador de l'ecosistema, el segnal l'è doppi: el ritm de Qwen el resta sostenibil anca se la concurenza de frontiera, e la generazion de imagin open weights la conserva el sò pubblegh in tra i svilupador che voeuren eseguì, verificà e personalizà i sò modej.

Page 1 — Page 1 — Seite 1 — Pagina 1 — Pagina 1 — À la UneFront PageFrontseitePrima PaginaIn Prima Pagina

I. Annonces des labos & écosystèmeLab announcements & ecosystemAnkündigungen der Labore & ÖkosystemAnnunci dei laboratori ed ecosistemaAnunzi di laboratori & ecosistema

Agents

Agents

Agenten

Agenti

Agent

AX : Google présenterait un orchestrateur agentique ouvertAX: Google reportedly set to unveil an open agentic orchestratorAX: Google soll einen offenen agentischen Orchestrator vorstellenAX: Google presenterebbe un orchestratore agentico openAX : Google el presentariss on orchestrador agentich dervii

Un projet nommé AX, présenté comme un orchestrateur agentique ouvert de Google et arrivé sur la page d'accueil de Hacker News le 20 septembre 2026, a récolté 334 points et 129 commentaires en quelques heures. La discussion ouvert sur le fil porte sur la place d'un orchestrateur générique dans un écosystème déjà dense en frameworks d'agents. L'initiative est consultable sur sa page officielle.
A project named AX, presented as an open agentic orchestrator from Google, reached the Hacker News front page on September 20, 2026, gathering 334 points and 129 comments within hours. The discussion on the thread centers on where a generic orchestrator fits in an ecosystem already dense with agent frameworks. The initiative is available on its official page.
Ein Projekt namens AX, präsentiert als offener agentischer Orchestrator von Google, hat am 20. September 2026 die Startseite von Hacker News erreicht und in wenigen Stunden 334 Punkte und 129 Kommentare gesammelt. Die im Diskussionsstrang geführte Debatte dreht sich um den Platz eines generischen Orchestrators in einem Ökosystem, das an Agent-Frameworks bereits dicht besetzt ist. Die Initiative ist auf ihrer offiziellen Seite einsehbar.
Un progetto denominato AX, presentato come un orchestratore agentico open di Google e arrivato in homepage su Hacker News il 20 settembre 2026, ha raccolto 334 punti e 129 commenti in poche ore. La discussione aperta sul thread verte sul ruolo di un orchestratore generico in un ecosistema già denso di framework per agenti. L'iniziativa è consultabile sulla sua pagina ufficiale.
On progett ciammaa AX, presentaa come on orchestrador agentich dervii de Google e rivaa in sora la pagina principala de Hacker News el 20 de setember 2026, l'ha tiraa insema 334 pont e 129 comenter in quaj or. La discussión dervida in sul fil la gira intorna al post de on orchestrador generich in d'on ecosistema giamò dens de framework de agent. L'iniziativa l'è consultabil in sora la soa pagina offiziala.

Inference locale

Local inference

Lokale Inferenz

Inferenza locale

Inferenzia local

Laya tourne hors ligne sur Mac M4 via CoreMLLaya runs offline on M4 Mac via CoreMLLaya läuft offline auf dem Mac M4 via CoreMLLaya gira offline su Mac M4 tramite CoreMLLaya el gira foeu de linia in su Mac M4 via CoreML

Un utilisateur a documenté le 20 septembre 2026 l'exécution locale de Laya, le modèle de décision qui renonce à générer des tokens et dont la sortie a fait le buzz la veille, sur un Mac M4 via CoreML en complète autonomie. Le guide publié sur GitHub a déjà réuni 147 points sur Hacker News, signe que l'approche zéro-token intrigue au-delà de la curiosité initiale.
A user documented on September 20, 2026 the local execution of Laya — the decision model that forgoes token generation and whose release made waves the day before — on an M4 Mac via CoreML, fully offline. The guide published on GitHub has already drawn 147 points on Hacker News, a sign that the zero-token approach intrigues beyond initial curiosity.
Ein Nutzer hat am 20. September 2026 die lokale Ausführung von Laya dokumentiert – jenes Entscheidungsmodell, das auf die Token-Generierung verzichtet und dessen Veröffentlichung am Vortag für Aufsehen sorgte – auf einem Mac M4 via CoreML in vollständiger Eigenständigkeit. Der auf GitHub veröffentlichte Guide hat bereits 147 Punkte auf Hacker News versammelt – ein Zeichen, dass der Zero-Token-Ansatz über die anfängliche Neugier hinaus interessiert.
Un utente ha documentato il 20 settembre 2026 l'esecuzione locale di Laya, il modello decisionale che rinuncia a generare token e la cui uscita aveva fatto scalpore il giorno prima, su un Mac M4 tramite CoreML in completa autonomia. La guida pubblicata su GitHub ha già raccolto 147 punti su Hacker News, segno che l'approccio zero-token incuriosisce oltre la curiosità iniziale.
On utilizzador l'ha documentaa el 20 de setember 2026 l'esecuzion local de Laya, el modej de decision che 'l renunzia a generà token e che la soa sortida l'ha faa parlà el dì prima, in su on Mac M4 via CoreML in completa autonomia. La guida publicada in su GitHub l'ha giamò regolt 147 pont in sora Hacker News, segnal che l'approcc zerotoken el fa curiosaa pussee in là de la curiositaa iniziala.

Communauté

Community

Gemeinschaft

Comunità

Comunitaa

Pirate Face veut sauver les LLM de la suppressionPirate Face wants to save LLMs from deletionPirate Face will die LLMs vor der Löschung rettenPirate Face vuole salvare gli LLM dalla cancellazionePirate Face el voeur salvà i LLM de la sopression

Le projet Pirate Face, arrivé le 20 septembre 2026 en tête des discussions techniques avec 494 points et 142 commentaires, se présente comme une solution pour préserver les modèles de langage menacés de suppression des plateformes d'hébergement. La discussion communautaire s'est polarisée entre archivage numérique et contournement de retraits pour cause de licence, un débat récurrent à chaque takedown de poids ouverts. Le service est accessible sur pirateface.co.
Pirate Face, which topped technical discussions on September 20, 2026 with 494 points and 142 comments, presents itself as a solution for preserving language models threatened with removal from hosting platforms. The community discussion split between digital archiving and circumventing license-related takedowns — a recurring debate every time significant open-weight models are pulled. The service is accessible at pirateface.co.
Das Projekt Pirate Face, das am 20. September 2026 mit 494 Punkten und 142 Kommentaren die technischen Diskussionen anführte, präsentiert sich als Lösung zur Bewahrung von Sprachmodellen, die von der Löschung durch Hosting-Plattformen bedroht sind. Die gemeinschaftliche Diskussion hat sich zwischen digitaler Archivierung und der Umgehung lizenzbedingter Entfernungen polarisiert – eine Debatte, die bei jedem grösseren Takedown offener Gewichte erneut aufflammt. Der Dienst ist unter pirateface.co erreichbar.
Il progetto Pirate Face, arrivato il 20 settembre 2026 in cima alle discussioni tecniche con 494 punti e 142 commenti, si presenta come una soluzione per preservare i modelli linguistici minacciati di rimozione dalle piattaforme di hosting. La discussione comunitaria si è polarizzata tra archiviazione digitale e aggiramento dei ritiri per motivi di licenza, un dibattito ricorrente a ogni takedown di pesi aperti. Il servizio è accessibile su pirateface.co.
El progett Pirate Face, rivaa el 20 de setember 2026 in testa ai discusson tecnich con 494 pont e 142 comenter, el se presenta come ona soluzion per conservà i modej de lengoe minazaa de toltavia di piataform de hosting. La discussion comunitaria l'è vàda polarizada tra archiviazion digitala e girà intorna ai retir per motiv de licensa, on dibattit che 'l torna semper a ogni takedown di pess dervii. El servizzi l'è accessibil in sora pirateface.co.

Page 2 — Page 2 — Seite 2 — Pagina 2 — Pagina 2 — Le Cahier TechniqueThe Technical SectionDer Technische TeilIl Quaderno TecnicoEl Quadern Tecnegh

II. Outils & systèmesTools & systemsWerkzeuge & SystemeStrumenti e sistemiStrument & sistema

Agents & données

Agents & data

Agenten & Daten

Agenti e dati

Agent & dacc

EvoOntology : une ontologie auto-évolutive servie via MCP pour les agents de donnéesEvoOntology: a self-evolving ontology served via MCP for data agentsEvoOntology: eine selbst-evolvierende Ontologie via MCP für DatenagentenEvoOntology: un'ontologia auto-evolutiva servita via MCP per gli agenti di datiEvoOntology : ona ontologia autoevolutiva servida via MCP per i agent de dacc

Des chercheurs de l'université Renmin (RUC-DataLab) publient EvoOntology, une couche d'ontologie auto-évolutive pour agents de données, mise en avant le 21 septembre 2026 sur les Daily Papers de Hugging Face avec 33 votes. Le système encapsule l'ontologie comme un serveur MCP à trois couches — schéma, contenu, outils — avec un agent constructeur et une boucle d'auto-évolution validée sur trois benchmarks et quatre LLM dorsaux. Le code est disponible sur GitHub et le papier sur arXiv.
Researchers at Renmin University (RUC-DataLab) have released EvoOntology, a self-evolving ontology layer for data agents, featured on September 21, 2026 on Hugging Face's Daily Papers with 33 upvotes. The system packages the ontology as a three-layer MCP server — schema, content, tools — with a builder agent and a self-evolution loop validated on three benchmarks and four backbone LLMs. Code is available on GitHub and the paper on arXiv.
Forschende der Renmin-Universität (RUC-DataLab) veröffentlichen EvoOntology, eine selbst-evolvierende Ontologieschicht für Datenagenten, am 21. September 2026 auf den Daily Papers von Hugging Face mit 33 Stimmen hervorgehoben. Das System kapselt die Ontologie als dreischichtigen MCP-Server – Schema, Inhalt, Werkzeuge – mit einem Konstruktor-Agenten und einer Selbst-Evolutions-Schleife, validiert auf drei Benchmarks und vier LLM-Rückgraten. Der Code ist auf GitHub verfügbar, das Paper auf arXiv.
Ricercatori dell'università Renmin (RUC-DataLab) pubblicano EvoOntology, un livello di ontologia auto-evolutivo per agenti di dati, messo in evidenza il 21 settembre 2026 sui Daily Papers di Hugging Face con 33 voti. Il sistema incapsula l'ontologia come server MCP a tre livelli — schema, contenuto, strumenti — con un agente costruttore e un ciclo di auto-evoluzione validato su tre benchmark e quattro LLM di base. Il codice è disponibile su GitHub e il paper su arXiv.
Di resercador de l'universitaa Renmin (RUC-DataLab) publichen EvoOntology, on strat de ontologia autoevolutiva per agent de dacc, mettuu in evidenza el 21 de setember 2026 in sui Daily Papers de Hugging Face con 33 vot. El sistema elInParameterisless encapsula l'ontologia come on server MCP a trii strat — schema, contegnuu, strument — con on agent costrutor e on circuit de autoevoluzion validaa in su trii benchmark e quatter LLM dorsaj. El codes l'è disponibil in sora GitHub e 'l paper in sora arXiv.

Protocoles

Protocols

Protokolle

Protocolli

Protocoi

« Pourquoi MCP a toujours été une mauvaise idée » : la controverse du week-end"Why MCP was always a bad idea": the weekend controversy«Warum MCP von Anfang an eine schlechte Idee war»: die Wochenend-Kontroverse«Perché MCP è sempre stata una cattiva idea»: la polemica del weekend« Perchè MCP l'è semper staa ona cattiva idea » : la polémica del weekend

Un billet publié le 20 septembre 2026 et abondamment discuté (83 points, 83 commentaires) attaque frontalement le protocole Model Context Protocol. L'auteur avance que MCP était dès l'origine une mauvaise idée, argument que la communauté HN a largement nuancé, plusieurs commentateurs pointant l'adoption industrielle massive du protocole comme preuve de son utilité malgré ses défauts.
A post published on September 20, 2026 and heavily debated (83 points, 83 comments) takes direct aim at the Model Context Protocol. The author argues that MCP was a bad idea from the start, an argument the HN community has largely qualified, with several commenters pointing to the protocol's massive industrial adoption as proof of its usefulness despite its flaws.
Ein am 20. September 2026 veröffentlichter und reichlich diskutierter Beitrag (83 Punkte, 83 Kommentare) greift das Model Context Protocol frontal an. Der Autor argumentiert, dass MCP von Anfang an eine schlechte Idee war – ein Argument, das die HN-Gemeinschaft weitgehend relativiert hat, wobei mehrere Kommentatoren auf die massive industrielle Adoption des Protokolls als Beweis seines Nutzens trotz seiner Mängel verweisen.
Un post pubblicato il 20 settembre 2026 e ampiamente discusso (83 punti, 83 commenti) attacca frontalmente il protocollo Model Context Protocol. L'autore sostiene che MCP fosse fin dall'origine una cattiva idea, argomento che la comunità HN ha ampiamente ridimensionato, con diversi commentatori che indicano la massiccia adozione industriale del protocollo come prova della sua utilità nonostante i difetti.
On post publicaa el 20 de setember 2026 e discuss a longh (83 pont, 83 comenter) el tacca de front el protocoll Model Context Protocol. L'autor el sostegn che MCP l'era giamò de l'inizzi ona cattiva idea, argument che la comunitaa HN l'ha minga acetaa inscì, con divers comenter che fann vedè l'adozion industriala massiccia del protocoll come proeuva de la soa utilitaa anca con i sò defecc.

Expérimentation

Experimentation

Experimentelles

Sperimentazione

Sperimentazion

Jev devient chatbot — et son créateur assume que le résultat soit mauvaisJev becomes a chatbot — and its creator owns up to how bad it isJev wird zum Chatbot – und sein Schöpfer steht dazu, dass das Ergebnis schlecht istJev diventa chatbot — e il suo creatore si assume che il risultato sia scarsoJev el deventa chatbot — e 'l sò creator el acet che 'l resultaa sia cativ

Un développeur a converti Jev — le petit modèle dont l'architecture minimaliste avait suscité l'intérêt des cercles NLP — en chatbot, en assumant le caractère « médiocre » du résultat. Le dépôt GitHub, arrivé le 20 septembre 2026 sur la une de Hacker News avec 109 points, illustre la fascination persistante pour les modèles de quelques paramètres censés capturer l'essentiel du raisonnement langagier.
A developer has converted Jev — the small model whose minimalist architecture drew interest in NLP circles — into a chatbot, openly owning the "mediocre" result. The GitHub repository, which hit the front page of Hacker News on September 20, 2026 with 109 points, illustrates the enduring fascination with models of just a few parameters supposedly capturing the essence of language reasoning.
Ein Entwickler hat Jev – das kleine Modell, dessen minimalistische Architektur das Interesse der NLP-Kreise geweckt hatte – in einen Chatbot verwandelt, wobei er den «mittelmässigen» Charakter des Ergebnisses in Kauf nimmt. Das GitHub-Repository, das am 20. September 2026 mit 109 Punkten auf die Titelseite von Hacker News gelangte, illustriert die anhaltende Faszination für Modelle mit wenigen Parametern, die das Wesentliche des sprachlichen Vernunftschlusses einfangen sollen.
Uno sviluppatore ha convertito Jev — il piccolo modello la cui architettura minimalista aveva suscitato l'interesse dei circoli NLP — in chatbot, assumendosi la «mediocrità» del risultato. Il repository GitHub, arrivato il 20 settembre 2026 in homepage su Hacker News con 109 punti, illustra la fascinazione persistente per i modelli di pochi parametri che dovrebbero catturare l'essenza del ragionamento linguistico.
On svilupador l'ha convertii Jev — el picoeul modej che la soa architettura minimalista l'aveva faa sorgì interess in di ambient NLP — in chatbot, acetand el caràter « mediocer » del resultaa. El deposet GitHub, rivaa el 20 de setember 2026 in prima pagina de Hacker News con 109 pont, el mostra la fascinazion che la va innanz per i modej de quaj parameter che sarissen bon de ciapà l'essenzial del resonament lengovestegh.

Page 3 — Page 3 — Seite 3 — Pagina 3 — Pagina 3 — La RechercheResearchDie ForschungLa RicercaLa Ricerca

III. Papers & labosPapers & labsPapers & LaborePaper e laboratoriPaper & laboratori

RL & agents

RL & agents

RL & Agenten

RL e agenti

RL & agent

CodeMidas : transformer le code lui-même en environnements RL pour agents de codageCodeMidas: turning code itself into RL environments for coding agentsCodeMidas: den Code selbst in RL-Umgebungen für Coding-Agenten verwandelnCodeMidas: trasformare il codice stesso in ambienti RL per agenti di codingCodeMidas : trasformà el codes midemm in ambient RL per agent de codes

Le papier CodeMidas, porté par Xiaomi MiMo et mis en avant le 21 septembre 2026 sur les Daily Papers de Hugging Face (29 votes), décrit un pipeline agentique qui transforme le code source en environnements d'entraînement RL : 5 545 tâches extraites de 3 185 dépôts couvrant 23 langages. Entraîné avec GRPO sur ces tâches, MiMo-V2.5 gagne 11,7 % sur DeepSWE, 17 % sur ProgramBench et 8,5 % sur Terminal-Bench v2.1. Détails sur arXiv.
The CodeMidas paper, led by Xiaomi MiMo and featured on September 21, 2026 on Hugging Face's Daily Papers (29 upvotes), describes an agentic pipeline that turns source code into RL training environments: 5,545 tasks extracted from 3,185 repositories spanning 23 languages. Trained with GRPO on these tasks, MiMo-V2.5 gains 11.7% on DeepSWE, 17% on ProgramBench and 8.5% on Terminal-Bench v2.1. Details on arXiv.
Das Paper CodeMidas, getragen von Xiaomi MiMo und am 21. September 2026 auf den Daily Papers von Hugging Face hervorgehoben (29 Stimmen), beschreibt eine agentische Pipeline, die Quellcode in RL-Trainingsumgebungen verwandelt: 5 545 Aufgaben, extrahiert aus 3 185 Repositories und 23 Sprachen abdeckend. Mit GRPO auf diesen Aufgaben trainiert, gewinnt MiMo-V2.5 11,7 % auf DeepSWE, 17 % auf ProgramBench und 8,5 % auf Terminal-Bench v2.1. Details auf arXiv.
Il paper CodeMidas, promosso da Xiaomi MiMo e messo in evidenza il 21 settembre 2026 sui Daily Papers di Hugging Face (29 voti), descrive una pipeline agentica che trasforma il codice sorgente in ambienti di addestramento RL: 5 545 task estratti da 3 185 repository che coprono 23 linguaggi. Addestrato con GRPO su questi task, MiMo-V2.5 guadagna l'11,7% su DeepSWE, il 17% su ProgramBench e l'8,5% su Terminal-Bench v2.1. Dettagli su arXiv.
El paper CodeMidas, portaa innanz de Xiaomi MiMo e mettuu in evidenza el 21 de setember 2026 in sui Daily Papers de Hugging Face (29 vot), el descriev ona pipeline agentica che la transforma el codes sorgent in ambient de adrezzament RL : 5 545 lavorà ciappaa foeura de 3 185 deposet che quatten 23 lengoeu. Adreezaa con GRPO in su 'sti lavorà chì, MiMo-V2.5 el guadagna 11,7 % in su DeepSWE, 17 % in su ProgramBench e 8,5 % in su Terminal-Bench v2.1. Detali in sora arXiv.

Benchmarks

Benchmarks

Benchmarks

Benchmark

Benchmark

RecreationWorld : recréer des applications pour mesurer les agents hybridesRecreationWorld: recreating applications to measure hybrid agentsRecreationWorld: Anwendungen nachbauen, um hybride Agenten zu messenRecreationWorld: ricreare applicazioni per misurare gli agenti ibridiRecreationWorld : tornà a fà su aplicazion per misurà i agent ibrid

L'équipe Qwen publie RecreationWorld, un framework d'évaluation pour agents hybrides du computer use opérant sur cinq plateformes (Ubuntu, macOS, Windows, Android, Web), mis en avant le 21 septembre 2026 avec 27 votes. Le benchmark RecreationBench compte 250 tâches ; GPT-6 Astra domine avec 58,1 % de réussite globale, mais ne passe la totalité des tests programmatiques que sur 2,8 % des tâches — les agents reproduisent mieux la structure statique des interfaces que leurs interactions. Papier et code sur arXiv et GitHub.
The Qwen team has released RecreationWorld, an evaluation framework for hybrid computer-use agents operating across five platforms (Ubuntu, macOS, Windows, Android, Web), featured on September 21, 2026 with 27 upvotes. The RecreationBench benchmark comprises 250 tasks; GPT-6 Astra leads with a 58.1% overall success rate but passes all programmatic tests on only 2.8% of tasks — agents reproduce the static structure of interfaces better than their interactions. Paper and code on arXiv and GitHub.
Das Qwen-Team veröffentlicht RecreationWorld, ein Evaluationsframework für hybride Computer-Use-Agenten auf fünf Plattformen (Ubuntu, macOS, Windows, Android, Web), am 21. September 2026 mit 27 Stimmen hervorgehoben. Der Benchmark RecreationBench umfasst 250 Aufgaben; GPT-6 Astra dominiert mit 58,1 % Gesamterfolgsquote, besteht jedoch die Gesamtheit der programmatischen Tests nur bei 2,8 % der Aufgaben – die Agenten reproduzieren die statische Struktur der Schnittstellen besser als deren Interaktionen. Paper und Code auf arXiv und GitHub.
Il team Qwen pubblica RecreationWorld, un framework di valutazione per agenti ibridi di computer use operanti su cinque piattaforme (Ubuntu, macOS, Windows, Android, Web), messo in evidenza il 21 settembre 2026 con 27 voti. Il benchmark RecreationBench conta 250 task; GPT-6 Astra domina con il 58,1% di successo globale, ma supera la totalità dei test programmatici solo sul 2,8% dei task — gli agenti riproducono meglio la struttura statica delle interfacce che le loro interazioni. Paper e codice su arXiv e GitHub.
L'equipa Qwen la publica RecreationWorld, on framework de valutazion per agent ibrid de computer use che lavoren in su cinch piataform (Ubuntu, macOS, Windows, Android, Web), mettuu in evidenza el 21 de setember 2026 con 27 vot. El benchmark RecreationBench el cuntà 250 lavorà ; GPT-6 Astra el domina con 58,1 % de success global, ma el passa no la totalitaa di proeuve programmategh domà in su 'l 2,8 % di lavorà — i agent reprodusen mej la struttura statica di interfac che i sò interazion. Paper e codes in sora arXiv e GitHub.

Agents visuels

Visual agents

Visuelle Agenten

Agenti visivi

Agent visiv

MintAct : Apple unifie les agents visuels en un seul modèle, de 2B à 8BMintAct: Apple unifies visual agents into a single model, from 2B to 8BMintAct: Apple vereinheitlicht visuelle Agenten in einem einzigen Modell, von 2B bis 8BMintAct: Apple unifica gli agenti visivi in un unico modello, da 2B a 8BMintAct : Apple la unifica i agent visiv in d'on modej domà, de 2B a 8B

Apple publie MintAct, une famille de vision-language models unifiant l'ancrage UI, la navigation multi-étapes (mobile, bureau, web) et l'usage d'outils visuels, entraînée aux échelles 2B, 4B et 8B. Soumis le 21 septembre 2026 aux Daily Papers, le modèle atteint 48,9 sur OSWorld-Verified, un état de l'art à taille de modèle comparable, grâce à une infrastructure RL asynchrone servant des centaines d'instances d'environnements concurrentes. Résumé sur arXiv.
Apple has released MintAct, a family of vision-language models unifying UI grounding, multi-step navigation (mobile, desktop, web) and visual tool use, trained at 2B, 4B and 8B scales. Submitted on September 21, 2026 to the Daily Papers, the model reaches 48.9 on OSWorld-Verified, a state of the art at comparable model size, thanks to an asynchronous RL infrastructure serving hundreds of concurrent environment instances. Summary on arXiv.
Apple veröffentlicht MintAct, eine Familie von Vision-Language-Modellen, die UI-Ankern, mehrstufige Navigation (mobil, Desktop, Web) und visuelle Werkzeugnutzung vereinheitlicht, trainiert in den Grössen 2B, 4B und 8B. Am 21. September 2026 bei den Daily Papers eingereicht, erreicht das Modell 48,9 auf OSWorld-Verified – ein Stand der Technik bei vergleichbarer Modellgrösse – dank einer asynchronen RL-Infrastruktur, die hunderte parallele Umgebungsinstanzen bedient. Zusammenfassung auf arXiv.
Apple pubblica MintAct, una famiglia di vision-language models che unifica l'ancoraggio UI, la navigazione multi-step (mobile, desktop, web) e l'uso di strumenti visivi, addestrata alle scale 2B, 4B e 8B. Presentato il 21 settembre 2026 ai Daily Papers, il modello raggiunge 48,9 su OSWorld-Verified, uno stato dell'arte a parità di dimensioni del modello, grazie a un'infrastruttura RL asincrona che serve centinaia di istanze di ambienti in concorrenza. Sintesi su arXiv.
Apple la publica MintAct, ona fameja de visionlanguage models che la unifica l'ancoragg UI, la navigazion multipass (mobil, desktop, web) e l'uso de strument visiv, adrezzada a scala 2B, 4B e 8B. Consignaa el 21 de setember 2026 ai Daily Papers, el modej el rivà a 48,9 in su OSWorld-Verified, on stat de l'art a dimension de modej confrontabil, grassie a ona infrastrutura RL asincrona che la serv centenaa de instanz de ambient in parallell. Sintesi in sora arXiv.

Génération d'images

Image generation

Bildgenerierung

Generazione di immagini

Generazion de imagin

Paint-Anything : imposer n'importe quelle couleur hexadécimale dans la génération d'imagesPaint-Anything: enforcing any hexadecimal color in image generationPaint-Anything: beliebige Hexadezimalfarbe in der Bildgenerierung erzwingenPaint-Anything: imporre qualsiasi colore esadecimale nella generazione di immaginiPaint-Anything : mett giò qualsessia colòr esadecimal in de la generazion de imagin

ByteDance Seed propose Paint-Anything, un contrôle couleur unifié pour la génération et l'édition d'images via interface en valeurs hexadécimales 24 bits, mis en avant le 21 septembre 2026 avec 15 votes. Le pipeline construit un corpus Paint-500K à partir d'images réelles et introduit le benchmark ACBench : sur FLUX.2-4B, Paint-Anything améliore les scores ACBench-T2I et ACBench-Edit de 85,3 % et 28,3 % respectivement. Papier complet sur arXiv.
ByteDance Seed proposes Paint-Anything, a unified color control for image generation and editing via a 24-bit hexadecimal interface, featured on September 21, 2026 with 15 upvotes. The pipeline builds a Paint-500K corpus from real images and introduces the ACBench benchmark: on FLUX.2-4B, Paint-Anything improves ACBench-T2I and ACBench-Edit scores by 85.3% and 28.3% respectively. Full paper on arXiv.
ByteDance Seed schlägt Paint-Anything vor, eine einheitliche Farbsteuerung für Generierung und Bearbeitung von Bildern über eine Schnittstelle mit 24-Bit-Hexadezimalwerten, am 21. September 2026 mit 15 Stimmen hervorgehoben. Die Pipeline erstellt einen Korpus Paint-500K aus realen Bildern und führt den Benchmark ACBench ein: Auf FLUX.2-4B verbessert Paint-Anything die Scores ACBench-T2I und ACBench-Edit um 85,3 % bzw. 28,3 %. Vollständiges Paper auf arXiv.
ByteDance Seed propone Paint-Anything, un controllo colore unificato per la generazione e l'editing di immagini tramite interfaccia in valori esadecimali a 24 bit, messo in evidenza il 21 settembre 2026 con 15 voti. La pipeline costruisce un corpus Paint-500K a partire da immagini reali e introduce il benchmark ACBench: su FLUX.2-4B, Paint-Anything migliora i punteggi ACBench-T2I e ACBench-Edit rispettivamente dell'85,3% e del 28,3%. Paper completo su arXiv.
ByteDance Seed el propon Paint-Anything, on contròll colòr unificaa per la generazion e la modifega de imagin via interfaccia in valor esadecimaj de 24 bit, mettuu in evidenza el 21 de setember 2026 con 15 vot. La pipeline la costruiss on corpus Paint-500K a partì de imagin reai e la introduc el benchmark ACBench : in su FLUX.2-4B, Paint-Anything el mejora i pontegg ACBench-T2I e ACBench-Edit de 85,3 % e 28,3 % respettivament. Paper complet in sora arXiv.

Page 4 — Page 4 — Seite 4 — Pagina 4 — Pagina 4 — La Communauté & ÉditoCommunity & EditorialGemeinschaft & LeitartikelLa Comunità ed EditorialeLa Comunitaa & Editorial

IV. Signaux & opinionSignals & opinionSignale & MeinungSegnali e opinioneSegnai & opinion

Signaux

Signals

Signale

Segnali

Segnai

L'« effet LLMentalist » revient hanter la discussion publiqueThe "LLMentalist effect" returns to haunt public debateDer «LLMentalist-Effekt» kehrt zurück, um die öffentliche Debatte zu heimsuchenL'«effetto LLMentalist» torna a tormentare il dibattito pubblicoL'« effett LLMentalist » el torna a visità la discussion pubblega

Une lettre de 2023 redécouverte sur « l'effet LLMentalist » — la propension des utilisateurs à surestimer la compréhension des modèles conversationnels — a trusté le haut de Hacker News le 20 septembre 2026, avec 169 points et 252 commentaires. Le texte trouve trois ans plus tard une résonance nouvelle à mesure que les agents s'invitent dans des décisions professionnelles à fort enjeu.
A rediscovered 2023 letter on the "LLMentalist effect" — users' propensity to overestimate conversational models' understanding — dominated the top of Hacker News on September 20, 2026, with 169 points and 252 comments. Three years on, the text resonates anew as agents insert themselves into high-stakes professional decisions.
Ein wiederentdeckter Brief aus dem Jahr 2023 über den «LLMentalist-Effekt» – die Neigung von Nutzern, das Verständnis konversationeller Modelle zu überschätzen – hat am 20. September 2026 die Spitze von Hacker News belegt, mit 169 Punkten und 252 Kommentaren. Der Text findet drei Jahre später eine neue Resonanz, da Agenten zunehmend in berufliche Entscheidungen mit hohen Einsätzen einbezogen werden.
Una lettera del 2023 riscoperta sull'«effetto LLMentalist» — la propensione degli utenti a sovrastimare la comprensione dei modelli conversazionali — ha occupato la vetta di Hacker News il 20 settembre 2026, con 169 punti e 252 commenti. Il testo ritrova tre anni dopo una risonanza nuova man mano che gli agenti entrano nelle decisioni professionali ad alto rischio.
Ona lettera del 2023 tornada a vess trovada in su « l'effett LLMentalist » — la propension di utilizzador a suravalutà la comprension di modej conversazionaj — l'ha ocupaa la part alta de Hacker News el 20 de setember 2026, con 169 pont e 252 comenter. El test el troeuva trii ann despoeu ona resonanza noeuva a man man che i agent se metten dent in di decision professionaj degrant ris'c.

Communauté

Community

Gemeinschaft

Comunità

Comunitaa

Une compétition de petits réseaux de neurones qui jouent aux jeux de stratégieA competition of small neural networks playing strategy gamesEin Wettbewerb kleiner neuronaler Netze, die Strategiespiele spielenUna competizione di piccole reti neurali che giocano a giochi di strategiaOna competizion de picoei red de neuroni che ghe geven ai gioeugh de strategia

Tiny Brains, une compétition de petits réseaux de neurones qui s'affrontent à des jeux de stratégie, a été présentée le 20 septembre 2026 sur Hacker News. Le site du concours, dont la présentation a recueilli 48 points, propose de couronner l'intelligence compacte plutôt que le brute-force paramétrique — dans la lignée des concours d'IA de jeu classiques.
Tiny Brains, a competition of small neural networks competing against each other at strategy games, was presented on September 20, 2026 on Hacker News. The competition site, whose introduction drew 48 points, aims to crown compact intelligence rather than parametric brute force — in the lineage of classic game AI contests.
Tiny Brains, ein Wettbewerb kleiner neuronaler Netze, die sich in Strategiespielen messen, wurde am 20. September 2026 auf Hacker News vorgestellt. Die Website des Wettbewerbs, deren Präsentation 48 Punkte gesammelt hat, will die kompakte Intelligenz krönen statt die parametrische Brute Force – in der Tradition klassischer KI-Spielwettbewerbe.
Tiny Brains, una competizione di piccole reti neurali che si sfidano in giochi di strategia, è stata presentata il 20 settembre 2026 su Hacker News. Il sito del concorso, la cui presentazione ha raccolto 48 punti, propone di incoronare l'intelligenza compatta piuttosto che il brute-force parametrico — sulla scia dei classici concorsi di IA ludica.
Tiny Brains, ona competizion de picoei red de neuroni che se scontren in di gioeugh de strategia, l'è stada presentada el 20 de setember 2026 in su Hacker News. El sitt del concors, che la soa presentazion l'ha regolt 48 pont, el propon de coronà l'intelligenza compatta pussee che la forza bruta di paramiter — in de la linea di concors classich de IA de gioeugh.

Édito

Editorial

Leitartikel

Editoriale

Editorial

Édito — La taille ne fait plus la uneEditorial — Size no longer makes the front pageLeitartikel — Grösse macht keine Titelseite mehrEditoriale — La taglia non fa più notiziaEditorial — La dimension la fa pu la prima pagina

Un week-end dominical où la Une appartient à une sortie open weights plutôt qu'à une balle de laboratoire fermé, où l'article le plus débattu du site d'agrégation le plus influent demande à quoi servent encore les mathématiciens humains, et où la sortie d'un modèle de 2,8 Mo tutoie les mastodontes : le signal est cohérent. La frontière de l'IA utile se déplace moins vite que la frontière paramétrique, et c'est dans l'écart entre les deux que se joue l'actualité. Les éditeurs du Neuron Times maintiennent leur exigence : mesurer, dater, sourcer — et se méfier, comme l'effet LLMentalist nous le rappelle, de la fluidité qui ressemble à la compréhension.
A Sunday weekend where the front page belongs to an open-weights release rather than a closed-lab bombshell, where the most debated article on the most influential aggregation site asks what human mathematicians are still for, and where a 2.8 MB model release rubs shoulders with the giants: the signal is consistent. The frontier of useful AI is moving more slowly than the parametric frontier, and it is in the gap between the two that the news is made. The editors of The Neuron Times stand by their standards: measure, date, source — and be wary, as the LLMentalist effect reminds us, of fluency that looks like understanding.
Ein sonntägliches Wochenende, an dem die Titelseite einer Open-Weights-Veröffentlichung gehört statt einer Kugel aus einem geschlossenen Labor, an dem der meistdiskutierte Artikel des einflussreichsten Aggregators fragt, wozu menschliche Mathematiker noch nütze sind, und an dem die Veröffentlichung eines Modells von 2,8 MB den Mastodonten auf Augenhöhe begegnet: Das Signal ist kohärent. Die Frontline der nützlichen KI verschiebt sich langsamer als die parametrische Frontline, und in der Lücke zwischen den beiden spielt sich die Aktualität ab. Die Redaktion des Neuron Times bleibt bei ihrem Anspruch: messen, datieren, Quellen angeben – und sich, wie der LLMentalist-Effekt uns in Erinnerung ruft, vor der Gewandtheit hüten, die wie Verständnis aussieht.
Un weekend domenicale in cui la Prima Pagina appartiene a un'uscita open weights piuttosto che a un colpo di un laboratorio chiuso, in cui l'articolo più discusso del sito di aggregazione più influente si chiede a cosa servano ancora i matematici umani, e in cui l'uscita di un modello da 2,8 MB gareggia con i colossi: il segnale è coerente. La frontiera dell'IA utile si sposta più lentamente della frontiera parametrica, ed è nel divario tra le due che si gioca l'attualità. I redattori del Neuron Times mantengono la loro esigenza: misurare, datare, citare le fonti — e diffidare, come ci ricorda l'effetto LLMentalist, della fluidità che assomiglia alla comprensione.
On weekend domenegh indova che la prima pagina la apartegn a ona sortida open weights pussee che a ona balla de laboratori saraa, indova che 'l post pussee discuss del sitt de agregazion pussee influent el domanda a chiff ghe serven ancamò i matemategh uman, e indova che la sortida de on modej de 2,8 Mo la sta adree ai mastodont : el segnal l'è coerent. La frontiera de l'IA util la se sposta minga inscì svelt come la frontiera parametrega, e l'è in del voeucc tra i dò che se giuga l'actualitaa. I editor del Neuron Times manten el sò regòrd : misurà, dà la data, dà i font — e dìgh no a la fluiditaa che la someja comprension, come che l'effett LLMentalist el ne regorda.