The Neuron Times

All the AI that's fit to print

N° 182 Édition du matinMorning EditionMorgenausgabeEdizione del mattinoEdizion del mattin · Genève MERCREDI 1 JUILLET 2026WEDNESDAY, 1 JULY 2026MITTWOCH, 1. JULI 2026MERCOLEDÌ 1 LUGLIO 2026MERCOLEDÌ 1 LUGLIO 2026

À la Une · FrontièreFront Page · FrontierSchlagzeilen · GrenzbereichPrima pagina · FrontieraIn prima pagina · Frontiera

Anthropic lance Claude Sonnet 5, le département du Commerce lève les restrictions sur Fable 5 et Mythos 5Anthropic launches Claude Sonnet 5, Commerce Department lifts restrictions on Fable 5 and Mythos 5Anthropic lanciert Claude Sonnet 5, Handelsministerium hebt Beschränkungen für Fable 5 und Mythos 5 aufAnthropic lancia Claude Sonnet 5, il Dipartimento del Commercio rimuove le restrizioni su Fable 5 e Mythos 5Anthropic el lancia Claude Sonnet 5, el Departament del Comerç el leva i restrizion sora Fable 5 e Mythos 5

Anthropic dévoile son modèle Sonnet 5 à prix réduit, obtient la levée des restrictions sur ses modèles les plus puissants et lance un environnement de travail pour les scientifiques.Anthropic unveils its lower-cost Sonnet 5 model, secures the lifting of restrictions on its most powerful models, and launches a work environment for scientists.Anthropic enthüllt sein günstigeres Sonnet-5-Modell, erreicht die Aufhebung der Beschränkungen für seine leistungsstärksten Modelle und lanciert eine Arbeitsumgebung für Wissenschaftler.Anthropic svela il suo modello Sonnet 5 a prezzo ridotto, ottiene la rimozione delle restrizioni sui suoi modelli più potenti e lancia un ambiente di lavoro per gli scienziati.Anthropic la presenta el sò model Sonnet 5 a prezzi ridott, la otegn la levada di restrizion in su i sò model pussee potent e la lancia on ambient de lavorà per i scientifegh.

Anthropic a dévoilé le 30 juin 2026 Claude Sonnet 5, son modèle le plus performant de la gamme Sonnet, avec des capacités de codage, d'agents et de travail professionnel qui rivalisent avec celles de l'Opus 4.8, pour un coût par token nettement inférieur. Selon les benchmarks publiés par Anthropic, Sonnet 5 obtient un score de 53,4 sur l'AA Intelligence Index et de 71,5 sur l'indice de codage, tout en proposant un prix de 2 $ par million de tokens en entrée et 10 $ par million en sortie — soit environ cinq fois moins cher que l'Opus 4.8. Le modèle est d'ores et déjà disponible sur l'API Anthropic, AWS, Google Cloud et Microsoft Foundry, et devient le modèle par défaut de Claude Code dès la version 2.1.197.Anthropic unveiled Claude Sonnet 5 on June 30, 2026, its most performant model in the Sonnet lineup, with coding, agentic, and professional work capabilities rivaling those of Opus 4.8 at a significantly lower per-token cost. According to benchmarks published by Anthropic, Sonnet 5 scores 53.4 on the AA Intelligence Index and 71.5 on the coding index, while offering a price of $2 per million input tokens and $10 per million output tokens — roughly five times cheaper than Opus 4.8. The model is already available on the Anthropic API, AWS, Google Cloud, and Microsoft Foundry, and becomes the default model for Claude Code starting with version 2.1.197.Anthropic hat am 30. Juni 2026 Claude Sonnet 5 vorgestellt, das leistungsstärkste Modell der Sonnet-Reihe mit Code-, Agenten- und professionellen Arbeitsfähigkeiten, die mit denen von Opus 4.8 konkurrieren, bei deutlich niedrigeren Kosten pro Token. Laut den von Anthropic veröffentlichten Benchmarks erreicht Sonnet 5 einen Wert von 53,4 auf dem AA Intelligence Index und 71,5 auf dem Code-Index, bei einem Preis von 2 $ pro Million Tokens für Eingabe und 10 $ pro Million für Ausgabe – etwa fünfmal günstiger als Opus 4.8. Das Modell ist ab sofort über die Anthropic-API, AWS, Google Cloud und Microsoft Foundry verfügbar und wird ab Version 2.1.197 zum Standardmodell von Claude Code.Anthropic ha svelato il 30 giugno 2026 Claude Sonnet 5, il suo modello più performante della gamma Sonnet, con capacità di coding, agenti e lavoro professionale che rivaleggiano con quelle dell'Opus 4.8, a un costo per token nettamente inferiore. Secondo i benchmark pubblicati da Anthropic, Sonnet 5 ottiene un punteggio di 53,4 sull'AA Intelligence Index e di 71,5 sull'indice di coding, offrendo al contempo un prezzo di 2 $ per milione di token in input e 10 $ per milione in output — circa cinque volte meno dell'Opus 4.8. Il modello è già disponibile sull'API Anthropic, AWS, Google Cloud e Microsoft Foundry, e diventa il modello predefinito di Claude Code a partire dalla versione 2.1.197.Anthropic l'ha presentaa el 30 de giugn 2026 Claude Sonnet 5, el sò model pussee performant de la gamma Sonnet, con capacità de codifica, d'agents e de lavorà professional che ghe fann concorrenza a quei de l'Opus 4.8, per on cost per token nettament inferior. Segond i benchmark publicaa de Anthropic, Sonnet 5 el gh'ha on score de 53,4 in su l'AA Intelligence Index e de 71,5 in su l'indice de codifica, cont on prezzi de 2 $ per milion de token in entrata e 10 $ per milion in sortida — circa cinch voeult pussee bon mercaa de l'Opus 4.8. El model l'è giamò disponibil in su l'API Anthropic, AWS, Google Cloud e Microsoft Foundry, e 'l deventa el model de default de Claude Code a partì de la version 2.1.197.

En parallèle, Anthropic a également lancé Claude Science, un environnement de travail dédié aux chercheurs scientifiques. L'outil intègre plus de 60 compétences préconfigurées couvrant des domaines comme la génomique et la chimie computationnelle, avec un agent de vérification qui contrôle automatiquement les citations et les calculs. Claude Science peut fonctionner localement ou sur des clusters HPC, garantissant que les données sensibles ne quittent jamais l'infrastructure du laboratoire.In parallel, Anthropic also launched Claude Science, a dedicated work environment for scientific researchers. The tool integrates over 60 preconfigured skills covering domains such as genomics and computational chemistry, with a verification agent that automatically checks citations and calculations. Claude Science can run locally or on HPC clusters, ensuring sensitive data never leaves the lab's infrastructure.Parallel dazu hat Anthropic auch Claude Science lanciert, eine Arbeitsumgebung für wissenschaftliche Forscher. Das Tool integriert über 60 vorkonfigurierte Fähigkeiten in Bereichen wie Genomik und computergestützter Chemie, mit einem Verifizierungsagenten, der automatisch Zitate und Berechnungen überprüft. Claude Science kann lokal oder auf HPC-Clustern betrieben werden und stellt sicher, dass sensible Daten die Laborinfrastruktur nie verlassen.In parallelo, Anthropic ha anche lanciato Claude Science, un ambiente di lavoro dedicato ai ricercatori scientifici. Lo strumento integra oltre 60 competenze preconfigurate che coprono ambiti come la genomica e la chimica computazionale, con un agente di verifica che controlla automaticamente le citazioni e i calcoli. Claude Science può funzionare localmente o su cluster HPC, garantendo che i dati sensibili non lascino mai l'infrastruttura del laboratorio.In parallelo, Anthropic l'ha anca lanciaa Claude Science, on ambient de lavorà dedicaa ai ricercator scientifegh. L'utensil l'integra pussee de 60 competenze preconfigurade che quatten di camp come la genomica e la chimica computazionala, cont on agent de verifica che 'l controlla automaticament i citazion e i càlcol. Claude Science el pò fonzionà localment o in su cluster HPC, garantend che i dati sensibij non lassien mai l'infrastruttura del laboratori.

Page 1 — Page 1 — Seite 1 — Pagina 1 — Pagina 1 — À la UneFront PageTitelgeschichteIn Primo PianoA la Vuna

I. Modèles & FrontièreModels & FrontierModelle & GrenzbereichModelli & FrontieraModell & Frontiera

Anthropic

Anthropic

Anthropic

Anthropic

Anthropic

Claude Sonnet 5 : performances de pointe à prix réduitClaude Sonnet 5: Top-tier performance at a reduced priceClaude Sonnet 5: Spitzenleistung zu reduziertem PreisClaude Sonnet 5: prestazioni di punta a prezzo ridottoClaude Sonnet 5: performance de ponta a prezzi ridott

Anthropic a publié le 30 juin 2026 Claude Sonnet 5, un modèle qui atteint des performances de pointe en codage, agents et travail professionnel tout en restant bien moins coûteux que l'Opus 4.8. Sur le benchmark GDPval-AA v2, Sonnet 5 devance même l'Opus 4.8 avec un score de 1 618. Le modèle est disponible sur l'API Anthropic au tarif de 2 $/M tokens en entrée et 10 $/M tokens en sortie, avec un cache de lecture à 0,20 $/M tokens. Il supporte un contexte natif d'un million de tokens et des niveaux de raisonnement adaptatifs (low, medium, high, max).
Anthropic released Claude Sonnet 5 on June 30, 2026, a model that achieves state-of-the-art performance in coding, agents, and professional work while remaining far less expensive than Opus 4.8. On the GDPval-AA v2 benchmark, Sonnet 5 even surpasses Opus 4.8 with a score of 1,618. The model is available on the Anthropic API at $2/M input tokens and $10/M output tokens, with a read cache at $0.20/M tokens. It supports a native context of one million tokens and adaptive reasoning levels (low, medium, high, max).
Anthropic hat am 30. Juni 2026 Claude Sonnet 5 veröffentlicht, ein Modell, das Spitzenleistungen beim Codieren, bei Agenten und bei der professionellen Arbeit erzielt und dabei deutlich günstiger ist als Opus 4.8. Im Benchmark GDPval-AA v2 übertrifft Sonnet 5 sogar Opus 4.8 mit einem Wert von 1'618. Das Modell ist über die Anthropic-API zu einem Preis von 2 $/M Tokens für Eingabe und 10 $/M Tokens für Ausgabe erhältlich, mit einem Read-Cache von 0,20 $/M Tokens. Es unterstützt einen nativen Kontext von einer Million Tokens und adaptive Reasoning-Stufen (low, medium, high, max).
Anthropic ha pubblicato il 30 giugno 2026 Claude Sonnet 5, un modello che raggiunge prestazioni di punta in coding, agenti e lavoro professionale pur rimanendo molto meno costoso dell'Opus 4.8. Sul benchmark GDPval-AA v2, Sonnet 5 supera persino l'Opus 4.8 con un punteggio di 1 618. Il modello è disponibile sull'API Anthropic al prezzo di 2 $/M token in input e 10 $/M token in output, con una cache di lettura a 0,20 $/M token. Supporta un contesto nativo di un milione di token e livelli di ragionamento adattivi (low, medium, high, max).
Anthropic l'ha publicaa el 30 de giugn 2026 Claude Sonnet 5, on model che 'l riva a di performance de ponta in codifica, agent e lavorà professional, restand però assosenn manch costos de l'Opus 4.8. In sul benchmark GDPval-AA v2, Sonnet 5 el supera fina l'Opus 4.8 cont on score de 1 618. El model l'è disponibil in su l'API Anthropic al tariff de 2 $/M token in entrata e 10 $/M token in sortida, cont on cache de lettura a 0,20 $/M token. El supporta on contest nativ de on milion de token e di nivell de resonament adattiv (low, medium, high, max).

Google DeepMind

Google DeepMind

Google DeepMind

Google DeepMind

Google DeepMind

Google lance Nano Banana 2 Lite et Gemini Omni FlashGoogle launches Nano Banana 2 Lite and Gemini Omni FlashGoogle lanciert Nano Banana 2 Lite und Gemini Omni FlashGoogle lancia Nano Banana 2 Lite e Gemini Omni FlashGoogle el lancia Nano Banana 2 Lite e Gemini Omni Flash

Google a dévoilé le 30 juin 2026 deux nouveaux modèles de génération : Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) et Gemini Omni Flash. Nano Banana 2 Lite, présenté comme le modèle d'image le plus rapide et le plus économique de Google, génère une image en quatre secondes pour 0,034 $ pièce, avec un prix d'appel de 0,25 $ par million de tokens en entrée. Gemini Omni Flash apporte pour la première fois la génération et l'édition vidéo par prompts textuels via l'API. Google recommande d'enchaîner les deux modèles pour passer d'une image fixe à une vidéo animée.
Google unveiled two new generation models on June 30, 2026: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) and Gemini Omni Flash. Nano Banana 2 Lite, touted as Google's fastest and most economical image model, generates an image in four seconds for $0.034 each, with an entry price of $0.25 per million input tokens. Gemini Omni Flash brings video generation and editing via text prompts through the API for the first time. Google recommends chaining the two models to go from a static image to an animated video.
Google hat am 30. Juni 2026 zwei neue Generationsmodelle vorgestellt: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) und Gemini Omni Flash. Nano Banana 2 Lite, präsentiert als Googles schnellstes und günstigstes Bildmodell, generiert ein Bild in vier Sekunden für 0,034 $ pro Stück, mit einem Einstiegspreis von 0,25 $ pro Million Tokens für Eingabe. Gemini Omni Flash bringt erstmals Videogenerierung und -bearbeitung durch Text-Prompts über die API. Google empfiehlt, die beiden Modelle zu verketten, um von einem Standbild zu einem animierten Video zu gelangen.
Google ha svelato il 30 giugno 2026 due nuovi modelli di generazione: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) e Gemini Omni Flash. Nano Banana 2 Lite, presentato come il modello d'immagine più veloce ed economico di Google, genera un'immagine in quattro secondi a 0,034 $ l'una, con un prezzo d'ingresso di 0,25 $ per milione di token in input. Gemini Omni Flash porta per la prima volta la generazione e l'editing video tramite prompt testuali via API. Google consiglia di concatenare i due modelli per passare da un'immagine fissa a un video animato.
Google l'ha presentaa el 30 de giugn 2026 duu noeuv model de generazion: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) e Gemini Omni Flash. Nano Banana 2 Lite, presentaa come 'l model d'imagin pussee svelt e pussee economegh de Google, el genera ona imagin in quatter second per 0,034 $ l'una, cont on prezz d'apell de 0,25 $ per milion de token in entrata. Gemini Omni Flash el porta per la prima voeulta la generazion e l'edizion video per moeud de test via l'API. Google la recomanda de metter in fila i duu model per passà de ona imagin fissa a on video animaa.

Page 2 — Page 2 — Seite 2 — Pagina 2 — Pagina 2 — Le Cahier TechniqueTech NotebookDas Technische HeftIl Quaderno TecnicoEl Cader Tecnegh

II. Harnais & MoteursHarnesses & EnginesGeschirre & MotorenHarnais & MotoriHarnes & Motor

Claude Code

Claude Code

Claude Code

Claude Code

Claude Code

Claude Code passe à Sonnet 5 par défautClaude Code switches to Sonnet 5 by defaultClaude Code wechselt standardmässig zu Sonnet 5Claude Code passa a Sonnet 5 come predefinitoClaude Code el passa a Sonnet 5 de default

La version 2.1.197 de Claude Code, publiée le 30 juin 2026, intègre Claude Sonnet 5 comme modèle par défaut, avec un contexte natif d'un million de tokens et un tarif promotionnel de 2 $/10 $ par million de tokens jusqu'au 31 août. La mise à jour est accessible dès maintenant. Par ailleurs, un article du blog Claude détaille l'utilisation des boucles (loops) dans Claude Code, une fonctionnalité clé pour l'exécution de tâches agentiques répétitives.
Claude Code version 2.1.197, released on June 30, 2026, integrates Claude Sonnet 5 as the default model, with a native context of one million tokens and a promotional rate of $2/$10 per million tokens until August 31. The update is available now. Additionally, a Claude blog post details the use of loops in Claude Code, a key feature for executing repetitive agentic tasks.
Version 2.1.197 von Claude Code, veröffentlicht am 30. Juni 2026, integriert Claude Sonnet 5 als Standardmodell, mit einem nativen Kontext von einer Million Tokens und einem Aktionspreis von 2 $/10 $ pro Million Tokens bis zum 31. August. Das Update ist ab sofort verfügbar. Zudem beschreibt ein Blogbeitrag von Claude die Verwendung von Schleifen (loops) in Claude Code, eine Schlüsselfunktion für die Ausführung repetitiver agentischer Aufgaben.
La versione 2.1.197 di Claude Code, pubblicata il 30 giugno 2026, integra Claude Sonnet 5 come modello predefinito, con un contesto nativo di un milione di token e una tariffa promozionale di 2 $/10 $ per milione di token fino al 31 agosto. L'aggiornamento è accessibile da subito. Inoltre, un articolo del blog Claude illustra l'utilizzo dei loop in Claude Code, una funzionalità chiave per l'esecuzione di compiti agentici ripetitivi.
La version 2.1.197 de Claude Code, publicada el 30 de giugn 2026, l'integra Claude Sonnet 5 come model de default, cont on contest nativ de on milion de token e on tariff promozional de 2 $/10 $ per milion de token fina al 31 de agost. L'azornament l'è accessibil giamò adess. D'oltra banda, on articol del blog Claude el detaja l'usagg di boucles (loop) in Claude Code, ona fonzionalità ciav per l'esecuzion de compit agentigh repetitiv.

Ollama

Ollama

Ollama

Ollama

Ollama

Ollama 0.31.1 : Gemma 4 jusqu'à 90 % plus rapide sur Apple SiliconOllama 0.31.1: Gemma 4 up to 90% faster on Apple SiliconOllama 0.31.1: Gemma 4 bis zu 90 % schneller auf Apple SiliconOllama 0.31.1: Gemma 4 fino al 90% più veloce su Apple SiliconOllama 0.31.1: Gemma 4 fina a 90% pussee svelta in su Apple Silicon

Ollama a publié le 30 juin 2026 la version 0.31.1, qui accélère significativement Gemma 4 sur Apple Silicon. La mise à jour exploite la prédiction multi-tokens (MTP) pour générer des tokens près de 90 % plus rapidement en moyenne sur un benchmark d'agents de codage. Ollama ajuste automatiquement le nombre de tokens draftés à l'exécution, sans configuration requise et sans modifier les sorties du modèle. La version met également à jour le moteur MLX et le moteur llama.cpp sous-jacent vers le build 9840.
Ollama released version 0.31.1 on June 30, 2026, which significantly accelerates Gemma 4 on Apple Silicon. The update leverages multi-token prediction (MTP) to generate tokens nearly 90% faster on average on a coding agent benchmark. Ollama automatically adjusts the number of drafted tokens at runtime, with no configuration required and without altering model outputs. The version also updates the MLX engine and the underlying llama.cpp engine to build 9840.
Ollama hat am 30. Juni 2026 Version 0.31.1 veröffentlicht, die Gemma 4 auf Apple Silicon deutlich beschleunigt. Das Update nutzt Multi-Token Prediction (MTP), um Tokens in einem Benchmark für Code-Agenten im Durchschnitt fast 90 % schneller zu generieren. Ollama passt die Anzahl der entworfenen Tokens zur Laufzeit automatisch an, ohne dass eine Konfiguration erforderlich ist und ohne die Modellausgaben zu verändern. Die Version aktualisiert zudem die MLX-Engine und die zugrunde liegende llama.cpp-Engine auf Build 9840.
Ollama ha pubblicato il 30 giugno 2026 la versione 0.31.1, che accelera significativamente Gemma 4 su Apple Silicon. L'aggiornamento sfrutta la predizione multi-token (MTP) per generare token quasi il 90% più velocemente in media su un benchmark di agenti di coding. Ollama regola automaticamente il numero di token draftati in fase di esecuzione, senza configurazione richiesta e senza modificare gli output del modello. La versione aggiorna inoltre il motore MLX e il motore llama.cpp sottostante al build 9840.
Ollama l'ha publicaa el 30 de giugn 2026 la version 0.31.1, che l'accelera significativament Gemma 4 in su Apple Silicon. L'azornament el doperà la prevision multi-token (MTP) per generà di token pressapoch 90% pussee svelt in media in su on benchmark d'agents de codifica. Ollama el regola automaticament el numer di token dervii a l'esecuzion, senza configurazion necessaria e senza modifegà i sortid del model. La version la met anca a giorn el motor MLX e 'l motor llama.cpp sotta al build 9840.

llama.cpp

llama.cpp

llama.cpp

llama.cpp

llama.cpp

llama.cpp b9851 corrige des erreurs CUDA dans l'attention flashllama.cpp b9851 fixes CUDA errors in flash attentionllama.cpp b9851 behebt CUDA-Fehler im Flash-Attentionllama.cpp b9851 corregge errori CUDA nell'attention flashllama.cpp b9851 el corregg di error CUDA in l'attenzion flash

La version b9851 de llama.cpp, publiée le 30 juin 2026, corrige des erreurs de troncature et de débordement d'entiers dans le noyau flash_attn_mask_to_KV_max lors de l'utilisation de strides de masque KQ sur CUDA. La release propose des binaires pour macOS (Apple Silicon et Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL), Windows (CPU, CUDA 12/13, Vulkan) et Android arm64.
llama.cpp version b9851, released on June 30, 2026, fixes truncation errors and integer overflows in the flash_attn_mask_to_KV_max kernel when using KQ mask strides on CUDA. The release offers binaries for macOS (Apple Silicon and Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL), Windows (CPU, CUDA 12/13, Vulkan), and Android arm64.
Version b9851 von llama.cpp, veröffentlicht am 30. Juni 2026, behebt Kürzungs- und Ganzzahlüberlauffehler im Kernel flash_attn_mask_to_KV_max bei Verwendung von KQ-Masken-Strides auf CUDA. Das Release bietet Binärdateien für macOS (Apple Silicon und Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL), Windows (CPU, CUDA 12/13, Vulkan) und Android arm64.
La versione b9851 di llama.cpp, pubblicata il 30 giugno 2026, corregge errori di troncamento e overflow di interi nel kernel flash_attn_mask_to_KV_max durante l'utilizzo di stride della maschera KQ su CUDA. La release offre binari per macOS (Apple Silicon e Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL), Windows (CPU, CUDA 12/13, Vulkan) e Android arm64.
La version b9851 de llama.cpp, publicada el 30 de giugn 2026, la corregg di error de troncatura e de sora-bordament d'intregh in del kernel flash_attn_mask_to_KV_max doperand di stride de maschera KQ in su CUDA. La release la propon di binari per macOS (Apple Silicon e Intel), Linux (CPU, Vulkan, ROCm 7.2, OpenVINO, SYCL), Windows (CPU, CUDA 12/13, Vulkan) e Android arm64.

Page 3 — Page 3 — Seite 3 — Pagina 3 — Pagina 3 — La RechercheResearchDie ForschungLa RicercaLa Ricerca

III. Papers & LabosPapers & LabsPapiere & LaborePapers & LaboratoriPaper & Laboratori

University of Washington / Meta

University of Washington / Meta

University of Washington / Meta

University of Washington / Meta

University of Washington / Meta

Agentic Abstention : le défi de savoir quand s'arrêterAgentic Abstention: The challenge of knowing when to stopAgentic Abstention: Die Herausforderung, zu wissen, wann man aufhören mussAgentic Abstention: la sfida di sapere quando fermarsiAgentic Abstention: la sfida de savè quand fermàss

Des chercheurs de l'Université de Washington et Meta Research ont publié le 30 juin 2026 Agentic Abstention, une étude sur la capacité des agents LLM à savoir quand cesser d'agir plutôt que de persister inutilement. En évaluant 13 systèmes LLM-as-agent et 2 scaffolds sur plus de 28 000 tâches couvrant le shopping web, les environnements terminaux et le question-réponse, les auteurs montrent que le principal défi n'est pas seulement la capacité à s'abstenir, mais le moment où ils le font. Certains agents ne s'abstiennent jamais quand ils le devraient, d'autres seulement après de nombreuses interactions superflues. La méthode CONVOLVE, qui distille les trajectoires d'interaction en règles d'arrêt réutilisables, améliore le taux de rappel opportun de Llama-3.3-70B de 26,7 à 57,4 sur WebShop.
Researchers from the University of Washington and Meta Research published Agentic Abstention on June 30, 2026, a study on the ability of LLM agents to know when to stop acting rather than persisting unnecessarily. Evaluating 13 LLM-as-agent systems and 2 scaffolds on over 28,000 tasks covering web shopping, terminal environments, and question-answering, the authors show that the main challenge is not just the ability to abstain, but the timing of when they do so. Some agents never abstain when they should, others only after many superfluous interactions. The CONVOLVE method, which distills interaction trajectories into reusable stopping rules, improves the timely recall rate of Llama-3.3-70B from 26.7 to 57.4 on WebShop.
Forscher der University of Washington und von Meta Research haben am 30. Juni 2026 Agentic Abstention veröffentlicht, eine Studie über die Fähigkeit von LLM-Agenten zu wissen, wann sie aufhören sollten zu handeln, anstatt unnötig weiterzumachen. Bei der Evaluierung von 13 LLM-as-Agent-Systemen und 2 Scaffolds über mehr als 28'000 Aufgaben, die Web-Shopping, Terminal-Umgebungen und Frage-Antwort umfassen, zeigen die Autoren, dass die grösste Herausforderung nicht nur die Fähigkeit zur Enthaltung ist, sondern der Zeitpunkt, zu dem sie erfolgt. Manche Agenten enthalten sich nie, wenn sie es sollten, andere erst nach vielen überflüssigen Interaktionen. Die CONVOLVE-Methode, die Interaktionsverläufe in wiederverwendbare Stoppregeln destilliert, verbessert die rechtzeitige Erinnerungsrate von Llama-3.3-70B von 26,7 auf 57,4 bei WebShop.
Ricercatori dell'Università di Washington e Meta Research hanno pubblicato il 30 giugno 2026 Agentic Abstention, uno studio sulla capacità degli agenti LLM di sapere quando smettere di agire piuttosto che persistere inutilmente. Valutando 13 sistemi LLM-as-agent e 2 scaffold su oltre 28 000 compiti che coprono shopping web, ambienti terminali e question-answering, gli autori mostrano che la sfida principale non è solo la capacità di astenersi, ma il momento in cui lo fanno. Alcuni agenti non si astengono mai quando dovrebbero, altri solo dopo numerose interazioni superflue. Il metodo CONVOLVE, che distilla le traiettorie di interazione in regole di arresto riutilizzabili, migliora il tasso di richiamo tempestivo di Llama-3.3-70B da 26,7 a 57,4 su WebShop.
Di ricercator de l'Università de Washington e de Meta Research hann publicaa el 30 de giugn 2026 Agentic Abstention, on studi in su la capacità di agent LLM de savè quand fermàss de agì inveci de perseverà inutilment. Valutand 13 sistema LLM-as-agent e 2 scaffold in su pussee de 28 000 compit che quatten el shopping web, i ambient terminal e 'l domanda-resposta, i autor mostren che 'l sfida principal l'è minga domà la capacità de astegniss, ma 'l moment in che lor el fann. Certi agent non se astegnen mai quand che ghe vorress, alter domà dopo tante interazion superflue. La metod CONVOLVE, che la destilla i traietori d'interazion in regoll de fermada doperabil de noeuv, la mejora el tass de ciamada oportuna de Llama-3.3-70B de 26,7 a 57,4 in su WebShop.

ByteDance

ByteDance

ByteDance

ByteDance

ByteDance

Dockerless : vérifier les correctifs sans exécutionDockerless: Verifying patches without executionDockerless: Patches ohne Ausführung verifizierenDockerless: verificare le patch senza esecuzioneDockerless: verificà i correzion senza esecuzion

ByteDance a publié le 30 juin 2026 Dockerless, un vérificateur de correctifs agentique sans environnement qui évalue les patches de code sans les exécuter. Sur un benchmark de vérification, Dockerless surpasse le meilleur vérificateur open-source de 14,3 points d'AUC. Utilisé comme filtre de trajectoires SFT et comme récompense RL, il permet d'atteindre des taux de résolution de 62,0 %, 50,0 % et 35,2 % sur SWE-bench Verified, Multilingual et Pro, surpassant le baseline Qwen3.5-9B de 2,4, 8,7 et 2,9 points.
ByteDance published Dockerless on June 30, 2026, an environment-free agentic patch verifier that evaluates code patches without executing them. On a verification benchmark, Dockerless surpasses the best open-source verifier by 14.3 AUC points. Used as an SFT trajectory filter and as an RL reward, it achieves resolution rates of 62.0%, 50.0%, and 35.2% on SWE-bench Verified, Multilingual, and Pro, surpassing the Qwen3.5-9B baseline by 2.4, 8.7, and 2.9 points.
ByteDance hat am 30. Juni 2026 Dockerless veröffentlicht, einen agentischen Patch-Verifizierer ohne Umgebung, der Code-Patches bewertet, ohne sie auszuführen. In einem Verifizierungs-Benchmark übertrifft Dockerless den besten Open-Source-Verifizierer um 14,3 AUC-Punkte. Als SFT-Trajektorienfilter und RL-Belohnung eingesetzt, ermöglicht es Lösungsraten von 62,0 %, 50,0 % und 35,2 % auf SWE-bench Verified, Multilingual und Pro und übertrifft die Baseline Qwen3.5-9B um 2,4, 8,7 und 2,9 Punkte.
ByteDance ha pubblicato il 30 giugno 2026 Dockerless, un verificatore di patch agentico senza ambiente che valuta le correzioni di codice senza eseguirle. Su un benchmark di verifica, Dockerless supera il miglior verificatore open-source di 14,3 punti di AUC. Utilizzato come filtro di traiettorie SFT e come ricompensa RL, consente di raggiungere tassi di risoluzione del 62,0%, 50,0% e 35,2% su SWE-bench Verified, Multilingual e Pro, superando il baseline Qwen3.5-9B di 2,4, 8,7 e 2,9 punti.
ByteDance l'ha publicaa el 30 de giugn 2026 Dockerless, on verificador de correzion agentich senza ambient che 'l valuta i patch de codegh senza eseguij. In su on benchmark de verificazion, Dockerless el supera el miglior verificador open-source de 14,3 pont de AUC. Doperad come filter de traietori SFT e come ricompensa RL, el permet de rivà a di tass de risoluzion del 62,0%, 50,0% e 35,2% in su SWE-bench Verified, Multilingual e Pro, superand el baseline Qwen3.5-9B de 2,4, 8,7 e 2,9 pont.

Tencent / UPenn

Tencent / UPenn

Tencent / UPenn

Tencent / UPenn

Tencent / UPenn

ReFreeKV : compression du cache KV sans seuilReFreeKV: Threshold-free KV cache compressionReFreeKV: KV-Cache-Kompression ohne SchwelleReFreeKV: compressione della cache KV senza sogliaReFreeKV: compression del cache KV senza soglia

Des chercheurs de Tencent et de l'Université de Pennsylvanie ont proposé ReFreeKV, une méthode de compression du cache KV sans seuil prédéfini. Contrairement aux approches existantes qui nécessitent un seuil spécifique au domaine pour déterminer le budget du cache KV, ReFreeKV alloue adaptativement le budget de compression tout en préservant les performances du cache complet. Les expériences menées sur 13 jeux de données de longueurs et types variés démontrent son efficacité.
Researchers from Tencent and the University of Pennsylvania proposed ReFreeKV, a threshold-free KV cache compression method. Unlike existing approaches that require a domain-specific threshold to determine the KV cache budget, ReFreeKV adaptively allocates the compression budget while preserving full-cache performance. Experiments conducted on 13 datasets of varying lengths and types demonstrate its effectiveness.
Forscher von Tencent und der University of Pennsylvania haben ReFreeKV vorgeschlagen, eine Methode zur KV-Cache-Kompression ohne vordefinierte Schwelle. Im Gegensatz zu bestehenden Ansätzen, die eine domänenspezifische Schwelle zur Bestimmung des KV-Cache-Budgets benötigen, weist ReFreeKV das Kompressionsbudget adaptiv zu, während die Leistung des vollständigen Caches erhalten bleibt. Experimente mit 13 Datensätzen unterschiedlicher Längen und Typen belegen seine Wirksamkeit.
Ricercatori di Tencent e dell'Università della Pennsylvania hanno proposto ReFreeKV, un metodo di compressione della cache KV senza soglia predefinita. A differenza degli approcci esistenti che richiedono una soglia specifica per dominio per determinare il budget della cache KV, ReFreeKV alloca adattivamente il budget di compressione preservando le prestazioni della cache completa. Gli esperimenti condotti su 13 dataset di lunghezze e tipi variati ne dimostrano l'efficacia.
Di ricercator de Tencent e de l'Università de Pennsylvania hann proponuu ReFreeKV, ona metod de compression del cache KV senza soglia predefinida. A diferenza di approcc esistent che gh'hann besogn de ona soglia specifica al domini per determinà el budget del cache KV, ReFreeKV el alloca adattativament el budget de compression, preservand i performance del cache complet. I esperiment faa in su 13 set de dati de longhezze e tip vari hann dimostraa la soa efficacia.

NVIDIA

NVIDIA

NVIDIA

NVIDIA

NVIDIA

Nemotron-Labs-Diffusion-Image : diffusion discrète masquéeNemotron-Labs-Diffusion-Image: Masked discrete diffusionNemotron-Labs-Diffusion-Image: Diskrete maskierte DiffusionNemotron-Labs-Diffusion-Image: diffusione discreta mascherataNemotron-Labs-Diffusion-Image: diffusion discreta mascherada

NVIDIA a dévoilé le 30 juin 2026 Nemotron-Labs-Diffusion-Image, un modèle de diffusion discrète masquée pour la synthèse texte-image haute résolution. Le modèle introduit un mécanisme d'édition de tokens permettant de réviser dynamiquement les tokens déjà démasqués pendant l'inférence, et une fonction de perte Grouped Cross-Entropy (GCE) qui atténue la rareté du signal d'apprentissage. Il atteint des scores de 0,90 sur GenEval, 86,9 sur DPG et 10,76 sur HPSv3.
NVIDIA unveiled Nemotron-Labs-Diffusion-Image on June 30, 2026, a masked discrete diffusion model for high-resolution text-to-image synthesis. The model introduces a token editing mechanism that dynamically revises already-unmasked tokens during inference, and a Grouped Cross-Entropy (GCE) loss function that mitigates the sparsity of the learning signal. It achieves scores of 0.90 on GenEval, 86.9 on DPG, and 10.76 on HPSv3.
NVIDIA hat am 30. Juni 2026 Nemotron-Labs-Diffusion-Image vorgestellt, ein diskretes maskiertes Diffusionsmodell für die hochauflösende Text-zu-Bild-Synthese. Das Modell führt einen Token-Editierungsmechanismus ein, der während der Inferenz dynamisch bereits aufgedeckte Tokens überarbeiten kann, sowie eine Grouped Cross-Entropy (GCE)-Verlustfunktion, die die Spärlichkeit des Lernsignals abmildert. Es erreicht Werte von 0,90 auf GenEval, 86,9 auf DPG und 10,76 auf HPSv3.
NVIDIA ha svelato il 30 giugno 2026 Nemotron-Labs-Diffusion-Image, un modello di diffusione discreta mascherata per la sintesi testo-immagine ad alta risoluzione. Il modello introduce un meccanismo di editing dei token che consente di revisionare dinamicamente i token già demascherati durante l'inferenza, e una funzione di perdita Grouped Cross-Entropy (GCE) che attenua la scarsità del segnale di apprendimento. Raggiunge punteggi di 0,90 su GenEval, 86,9 su DPG e 10,76 su HPSv3.
NVIDIA l'ha presentaa el 30 de giugn 2026 Nemotron-Labs-Diffusion-Image, on model de diffusion discreta mascherada per la sintesi test-imagin a alta risoluzion. El model l'introdux on mecanism d'edizion de token che 'l permet de revidà dinamicament i token giamò desmascheraa durant l'inferenza, e ona fonzion de perdita Grouped Cross-Entropy (GCE) che la smorza la scarsità del segnal d'aprendiment. El riva a di score de 0,90 in su GenEval, 86,9 in su DPG e 10,76 in su HPSv3.

Page 4 — Page 4 — Seite 4 — Pagina 4 — Pagina 4 — La Communauté & ÉditoCommunity & EditorialDie Gemeinschaft & EditorialLa Comunità & EditorialeLa Comunità & Editorial

IV. Communauté & ÉditoCommunity & EditorialGemeinschaft & EditorialComunità & EditorialeComunità & Editorial

Débat

Debate

Debatte

Dibattito

Dibattit

Claude Code marque stéganographiquement ses requêtesClaude Code steganographically marks its promptsClaude Code markiert seine Anfragen steganografischClaude Code marchia steganograficamente le sue richiesteClaude Code el marca steganograficament i sò richiest

Un développeur a découvert que Claude Code marque stéganographiquement ses requêtes, une pratique qui a suscité un vif débat sur Hacker News (1 589 points, 457 commentaires). L'article technique détaille comment les prompts générés par Claude Code contiennent des marqueurs cachés, soulevant des questions sur la transparence et la traçabilité des interactions avec les agents de codage.
A developer discovered that Claude Code steganographically marks its prompts, a practice that sparked a heated debate on Hacker News (1,589 points, 457 comments). The technical article details how prompts generated by Claude Code contain hidden markers, raising questions about transparency and traceability in interactions with coding agents.
Ein Entwickler hat entdeckt, dass Claude Code seine Anfragen steganografisch markiert, eine Praxis, die auf Hacker News eine lebhafte Debatte ausgelöst hat (1'589 Punkte, 457 Kommentare). Der technische Artikel beschreibt detailliert, wie die von Claude Code generierten Prompts versteckte Marker enthalten, was Fragen zur Transparenz und Rückverfolgbarkeit von Interaktionen mit Code-Agenten aufwirft.
Uno sviluppatore ha scoperto che Claude Code marchia steganograficamente le sue richieste, una pratica che ha suscitato un acceso dibattito su Hacker News (1 589 punti, 457 commenti). L'articolo tecnico dettaglia come i prompt generati da Claude Code contengano marcatori nascosti, sollevando domande sulla trasparenza e la tracciabilità delle interazioni con gli agenti di coding.
On desvilupador l'ha descovert che Claude Code el marca steganograficament i sò richiest, ona pratega che l'ha suscitaa on viv discussion in su Hacker News (1 589 pont, 457 comentari). L'articol tecnegh el detaja come i prompt generaa de Claude Code gh'hinnen denter di marcador sconduu, sollevand di question in su la trasparenza e la trazzabilità di interazion con i agent de codifica.

Meta AI

Meta AI

Meta AI

Meta AI

Meta AI

Brain2Qwerty v2 : décoder la parole à 61 % sans chirurgieBrain2Qwerty v2: Decoding speech at 61% without surgeryBrain2Qwerty v2: Sprache mit 61 % ohne Chirurgie entschlüsselnBrain2Qwerty v2: decodificare la parola al 61% senza chirurgiaBrain2Qwerty v2: decodificà la parolla al 61% senza chirurgia

Meta AI a publié le 30 juin 2026 Brain2Qwerty v2, un pipeline de décodage cerveau-texte non invasif utilisant la magnétoencéphalographie (MEG) qui atteint 61 % de précision au niveau des mots. Le modèle, qui a fait la une de Hacker News avec 130 points, permet de transcrire des phrases tapées mentalement sans intervention chirurgicale, ouvrant la voie à des interfaces de communication pour les personnes souffrant de paralysie sévère.
Meta AI published Brain2Qwerty v2 on June 30, 2026, a non-invasive brain-to-text decoding pipeline using magnetoencephalography (MEG) that achieves 61% word-level accuracy. The model, which made the front page of Hacker News with 130 points, can transcribe mentally typed sentences without surgery, paving the way for communication interfaces for people with severe paralysis.
Meta AI hat am 30. Juni 2026 Brain2Qwerty v2 veröffentlicht, eine nicht-invasive Gehirn-zu-Text-Dekodierungs-Pipeline mittels Magnetoenzephalographie (MEG), die eine Wortgenauigkeit von 61 % erreicht. Das Modell, das es mit 130 Punkten auf die Titelseite von Hacker News schaffte, ermöglicht die Transkription von gedanklich getippten Sätzen ohne chirurgischen Eingriff und ebnet den Weg für Kommunikationsschnittstellen für Menschen mit schweren Lähmungen.
Meta AI ha pubblicato il 30 giugno 2026 Brain2Qwerty v2, una pipeline di decodifica cervello-testo non invasiva che utilizza la magnetoencefalografia (MEG) e raggiunge il 61% di precisione a livello di parole. Il modello, che ha fatto notizia su Hacker News con 130 punti, consente di trascrivere frasi digitate mentalmente senza intervento chirurgico, aprendo la strada a interfacce di comunicazione per persone affette da paralisi grave.
Meta AI l'ha publicaa el 30 de giugn 2026 Brain2Qwerty v2, on pipeline de decodifica cervell-test minga invasiva che la doperà la magnetoencefalografia (MEG) e che la riva al 61% de precision a nivell di paroll. El model, che l'ha faa la prima pagina de Hacker News con 130 pont, el permet de trascriv di fras scrivud mentalment senza intervent chirurgich, dervend la strada a di interfacc de comunicazion per i personn che soffren de paralisi severa.

Benchmark

Benchmark

Benchmark

Benchmark

Benchmark

OSWorld 2.0 : les agents butent sur les tâches longuesOSWorld 2.0: Agents stumble on long-horizon tasksOSWorld 2.0: Agenten scheitern an langen AufgabenOSWorld 2.0: gli agenti si bloccano sui compiti lunghiOSWorld 2.0: i agent butten in sui compit longh

Le benchmark OSWorld 2.0, publié le 30 juin 2026, évalue les agents d'utilisation d'ordinateur sur 108 tâches longues représentant des workflows réels d'une durée médiane d'environ 1,6 heure pour un humain. Claude Opus 4.8 avec raisonnement maximal et appels d'outils par lots ne complète que 20,6 % des tâches, tandis que GPT-5.5 plafonne à environ 13 %. Les agents échouent non pas sur le contrôle GUI de base, mais en perdant la trace des contraintes, en manquant des informations qui arrivent en cours de tâche et en devinant plutôt qu'en demandant à l'utilisateur.
The OSWorld 2.0 benchmark, published on June 30, 2026, evaluates computer-use agents on 108 long-horizon tasks representing real-world workflows with a median duration of approximately 1.6 hours for a human. Claude Opus 4.8 with maximum reasoning and batched tool calls completes only 20.6% of tasks, while GPT-5.5 plateaus at around 13%. Agents fail not on basic GUI control, but by losing track of constraints, missing information that arrives mid-task, and guessing rather than asking the user.
Der Benchmark OSWorld 2.0, veröffentlicht am 30. Juni 2026, bewertet Computer-Nutzungs-Agenten anhand von 108 langen Aufgaben, die reale Arbeitsabläufe mit einer medianen Dauer von etwa 1,6 Stunden für einen Menschen darstellen. Claude Opus 4.8 mit maximalem Reasoning und Batch-Tool-Aufrufen schliesst nur 20,6 % der Aufgaben ab, während GPT-5.5 bei etwa 13 % stagniert. Die Agenten scheitern nicht an der grundlegenden GUI-Steuerung, sondern verlieren die Spur von Randbedingungen, übersehen Informationen, die während der Aufgabe eintreffen, und raten, anstatt den Benutzer zu fragen.
Il benchmark OSWorld 2.0, pubblicato il 30 giugno 2026, valuta gli agenti di utilizzo del computer su 108 compiti lunghi che rappresentano flussi di lavoro reali con una durata mediana di circa 1,6 ore per un umano. Claude Opus 4.8 con ragionamento massimo e chiamate di strumenti in batch completa solo il 20,6% dei compiti, mentre GPT-5.5 si attesta a circa il 13%. Gli agenti falliscono non sul controllo GUI di base, ma perdendo traccia dei vincoli, mancando informazioni che arrivano durante il compito e indovinando piuttosto che chiedere all'utente.
El benchmark OSWorld 2.0, publicaa el 30 de giugn 2026, el valuta i agent d'usagg del computer in su 108 compit longh che rappresenten di fluss de lavorà reai cont ona durada mediana de circa 1,6 ore per on uman. Claude Opus 4.8 cont el resonament massim e i ciamad d'utensil per lott el completa domà el 20,6% di compit, menter GPT-5.5 el riva a circa el 13%. I agent fallissen minga in sul controll GUI de bas, ma perden la traccia di vincol, se perden di informazion che riven durant el compit e indovinen inveci de domandà a l'utent.