À la Une · Agents & InférenceFront Page · Agents & InferenceSchlagzeilen · Agenten & InferenzPrima pagina · Agenti & InferenzaIn prima pagina · Agents & Inferenza
Google étend les Managed Agents avec tâches de fond et MCP distant ; NVIDIA dévoile un modèle tri-mode qui unifie AR, diffusion et décodage spéculatifGoogle expands Managed Agents with background tasks and remote MCP; NVIDIA unveils tri-mode model unifying AR, diffusion and speculative decodingGoogle erweitert Managed Agents um Hintergrundaufgaben und Remote-MCP; NVIDIA stellt Tri-Mode-Modell vor, das AR, Diffusion und spekulatives Decoding vereintGoogle estende gli Managed Agents con attività in background e MCP remoto; NVIDIA svela un modello tri-mode che unifica AR, diffusione e decodifica speculativaGoogle el slarga i Managed Agents con incarigh de fond e MCP distant; NVIDIA el presenta on modell tri-mod che unifica AR, diffusion e decodifega speculativa
Deux annonces majeures redessinent le paysage des agents et de l'inférence : Google ajoute l'exécution persistante et les connexions MCP distantes à ses Managed Agents, tandis que NVIDIA publie Nemotron-Labs-Diffusion, un modèle capable de basculer entre trois modes de génération.Two major announcements reshape the landscape of agents and inference: Google adds persistent execution and remote MCP connections to its Managed Agents, while NVIDIA releases Nemotron-Labs-Diffusion, a model capable of switching between three generation modes.Zwei bedeutende Ankündigungen verändern die Landschaft der Agenten und der Inferenz: Google erweitert seine Managed Agents um persistente Ausführung und Remote-MCP-Verbindungen, während NVIDIA Nemotron-Labs-Diffusion vorstellt, ein Modell, das zwischen drei Generierungsmodi umschalten kann.Due annunci importanti ridisegnano il panorama degli agenti e dell'inferenza: Google aggiunge l'esecuzione persistente e le connessioni MCP remote ai suoi Managed Agents, mentre NVIDIA pubblica Nemotron-Labs-Diffusion, un modello in grado di passare tra tre modalità di generazione.Dò anunzi important ridessegnen el paesagg di agent e de l'inferenza: Google el gionta l'esecuzion persistenta e i connession MCP distanti ai sò Managed Agents, menter NVIDIA la pubblica Nemotron-Labs-Diffusion, on modell bon de passà tra tri mod de generazion.
De la rédaction — 8 juillet 2026From the editorial desk — 8 July 2026Von der Redaktion — 8. Juli 2026Dalla redazione — 8 luglio 2026De la redazzion — 8 luglio 2026
Google a annoncé le 7 juillet 2026 une expansion significative de ses Managed Agents dans l'API Gemini, introduisant trois capacités clés : l'exécution de tâches en arrière-plan (background tasks), la connexion à des serveurs MCP (Model Context Protocol) distants, et un nouveau mode de déploiement "agent-as-service". Les développeurs peuvent désormais lancer des agents qui continuent de s'exécuter même après la déconnexion du client, une fonctionnalité critique pour les workflows de production. La prise en charge du MCP distant permet aux agents d'accéder à des outils et sources de données hébergés sur des serveurs externes, élargissant considérablement leur périmètre d'action au-delà des outils locaux.Google announced on July 7, 2026 a significant expansion of its Managed Agents in the Gemini API, introducing three key capabilities: background task execution, connection to remote MCP (Model Context Protocol) servers, and a new "agent-as-service" deployment mode. Developers can now launch agents that continue running even after the client disconnects, a critical feature for production workflows. Support for remote MCP allows agents to access tools and data sources hosted on external servers, significantly broadening their scope beyond local tools.Google hat am 7. Juli 2026 eine bedeutende Erweiterung seiner Managed Agents in der Gemini-API angekündigt und drei Schlüsselfunktionen eingeführt: die Ausführung von Hintergrundaufgaben (Background Tasks), die Verbindung zu entfernten MCP-Servern (Model Context Protocol) und einen neuen Bereitstellungsmodus «Agent-as-Service». Entwickler können nun Agenten starten, die auch nach der Trennung des Clients weiterlaufen – eine kritische Funktion für Produktions-Workflows. Die Unterstützung von Remote-MCP ermöglicht es Agenten, auf Tools und Datenquellen auf externen Servern zuzugreifen, was ihren Aktionsradius erheblich über lokale Werkzeuge hinaus erweitert.Google ha annunciato il 7 luglio 2026 un'espansione significativa dei suoi Managed Agents nell'API Gemini, introducendo tre capacità chiave: l'esecuzione di attività in background (background tasks), la connessione a server MCP (Model Context Protocol) remoti e una nuova modalità di deployment "agent-as-service". Gli sviluppatori possono ora lanciare agenti che continuano a eseguirsi anche dopo la disconnessione del client, una funzionalità critica per i workflow di produzione. Il supporto per MCP remoto consente agli agenti di accedere a strumenti e fonti di dati ospitati su server esterni, ampliando considerevolmente il loro raggio d'azione oltre gli strumenti locali.Google l'ha anunziaa el 7 de luj 2026 ona espansion significativa di sò Managed Agents in l'API Gemini, introduxend tri capacità fondamentai: l'esecuzion de incarigh in background (background tasks), la connession a server MCP (Model Context Protocol) distant, e on noeuv mod de despiegament "agent-as-service". I sviluppador poden adess lanzà di agent che continuen a eseguiss anca dopo la disconnession del client, ona funzionalità critica per i workflow de produzion. El support del MCP distant el permet ai agent de acced a istrument e sorgent de dat ospitaa su server esterni, slargand considerevolment el sò perimetr d'azzion oltra i istrument locai.
Parallèlement, les chercheurs de NVIDIA ont publié le 7 juillet Nemotron-Labs-Diffusion, une famille de modèles de langage "tri-mode" (3B, 8B et 14B paramètres) qui unifie au sein d'une même architecture le décodage autorégressif (AR), la génération par diffusion et l'auto-décodage spéculatif. Entraîné avec un objectif conjoint AR-diffusion, le modèle peut changer de mode à la volée pour optimiser le débit selon la charge du serveur. En mode auto-spéculatif, la diffusion génère un brouillon que l'AR vérifie, surpassant les méthodes MTP (multi-token prediction) traditionnelles : le modèle 8B décode 6 fois plus de tokens par passage avant que Qwen3-8B, avec une précision comparable, et atteint un débit 4 fois supérieur sur SPEED-Bench avec SGLang sur un GPU GB200.Meanwhile, NVIDIA researchers published on July 7 Nemotron-Labs-Diffusion, a family of "tri-mode" language models (3B, 8B and 14B parameters) that unify within a single architecture autoregressive decoding (AR), diffusion-based generation and self-speculative decoding. Trained with a joint AR-diffusion objective, the model can switch modes on the fly to optimize throughput based on server load. In self-speculative mode, diffusion generates a draft that AR verifies, outperforming traditional MTP (multi-token prediction) methods: the 8B model decodes 6 times more tokens per pass than Qwen3-8B, with comparable accuracy, and achieves 4 times higher throughput on SPEED-Bench with SGLang on a GB200 GPU.Parallel dazu haben NVIDIA-Forscher am 7. Juli Nemotron-Labs-Diffusion veröffentlicht, eine Familie von «Tri-Mode»-Sprachmodellen (3B, 8B und 14B Parameter), die autoregressives Decoding (AR), Diffusionsgenerierung und auto-spekulatives Decoding in einer einzigen Architektur vereinen. Trainiert mit einem gemeinsamen AR-Diffusion-Ziel, kann das Modell im laufenden Betrieb den Modus wechseln, um den Durchsatz je nach Serverlast zu optimieren. Im auto-spekulativen Modus erzeugt die Diffusion einen Entwurf, den der AR überprüft, und übertrifft damit traditionelle MTP-Methoden (Multi-Token Prediction): Das 8B-Modell decodiert 6-mal mehr Token pro Durchlauf als Qwen3-8B bei vergleichbarer Genauigkeit und erreicht auf SPEED-Bench mit SGLang auf einem GB200-GPU einen 4-fach höheren Durchsatz.Parallelamente, i ricercatori di NVIDIA hanno pubblicato il 7 luglio Nemotron-Labs-Diffusion, una famiglia di modelli linguistici "tri-mode" (3B, 8B e 14B parametri) che unifica all'interno di una stessa architettura la decodifica autoregressiva (AR), la generazione per diffusione e l'auto-decodifica speculativa. Addestrato con un obiettivo congiunto AR-diffusione, il modello può cambiare modalità al volo per ottimizzare il throughput in base al carico del server. In modalità auto-speculativa, la diffusione genera una bozza che l'AR verifica, superando i metodi MTP (multi-token prediction) tradizionali: il modello 8B decodifica 6 volte più token per passaggio rispetto a Qwen3-8B, con una precisione comparabile, e raggiunge un throughput 4 volte superiore su SPEED-Bench con SGLang su una GPU GB200.In del medemm temp, i ricercator de NVIDIA hann publicaa el 7 de luj Nemotron-Labs-Diffusion, ona fameja de modell de lenguagg "tri-mod" (3B, 8B e 14B parametri) che unifica denter de l'istessa architettura la decodifega autoregressiva (AR), la generazion per diffusion e l'auto-decodifega speculativa. Allenaa con on obietiv congiunt AR-diffusion, el modell el pò cambià de mod a la svoeulta per ottimizzà el débit segond la càrega del server. In mod auto-speculativ, la diffusion la genera on bozz che l'AR el verifega, superand i metod MTP (multi-token prediction) tradizionai: el modell 8B el decodifega 6 voeult pussee de token per passagg prima che Qwen3-8B, con ona precision compagna, e el riva a on débit 4 voeult superior su SPEED-Bench con SGLang su on GPU GB200.
Ces deux annonces illustrent une tendance convergente : d'un côté, les plateformes agentiques gagnent en autonomie et en portée grâce à l'infrastructure MCP et aux tâches persistantes ; de l'autre, l'inférence elle-même devient modulable, avec des modèles capables d'ajuster dynamiquement leur stratégie de génération. NVIDIA démontre également que les objectifs AR et diffusion sont complémentaires — la diffusion améliore la planification prospective (lookahead planning) tandis que l'AR fournit les priorités linguistiques gauche-droite — et qu'une analyse "speed-of-light" suggère un potentiel de 76,5 % de tokens supplémentaires par passage avant sous un échantillonneur optimal.These two announcements illustrate a converging trend: on one hand, agentic platforms are gaining autonomy and reach thanks to MCP infrastructure and persistent tasks; on the other, inference itself is becoming modular, with models capable of dynamically adjusting their generation strategy. NVIDIA also demonstrates that AR and diffusion objectives are complementary — diffusion improves lookahead planning while AR provides left-to-right linguistic priors — and that a "speed-of-light" analysis suggests a potential of 76.5% additional tokens per pass under an optimal sampler.Diese beiden Ankündigungen veranschaulichen einen konvergierenden Trend: Einerseits gewinnen Agentenplattformen durch die MCP-Infrastruktur und persistente Aufgaben an Autonomie und Reichweite; andererseits wird die Inferenz selbst modular, mit Modellen, die ihre Generierungsstrategie dynamisch anpassen können. NVIDIA zeigt zudem, dass AR- und Diffusionsziele komplementär sind – die Diffusion verbessert die vorausschauende Planung (Lookahead Planning), während der AR die links-rechts-Sprachprioritäten liefert – und dass eine «Speed-of-Light»-Analyse unter einem optimalen Sampler ein Potenzial von 76,5 % zusätzlichen Token pro Durchlauf nahelegt.Questi due annunci illustrano una tendenza convergente: da un lato, le piattaforme agentiche guadagnano autonomia e portata grazie all'infrastruttura MCP e alle attività persistenti; dall'altro, l'inferenza stessa diventa modulabile, con modelli capaci di regolare dinamicamente la propria strategia di generazione. NVIDIA dimostra inoltre che gli obiettivi AR e diffusione sono complementari — la diffusione migliora la pianificazione prospettica (lookahead planning) mentre l'AR fornisce le priorità linguistiche sinistra-destra — e che un'analisi "speed-of-light" suggerisce un potenziale del 76,5% di token aggiuntivi per passaggio sotto un campionatore ottimale.Ste dò anunzi illustren ona tendenza convergent: de ona part, i piattaform agenteghe guadagnen in autonomia e in portada grazzia a l'infrastruttura MCP e ai incarigh persistent; de l'oltra, l'inferenza istessa la deventa modulabela, con di modell bon de adattà dinamicament la soa strategia de generazion. NVIDIA la dimostra anca che i obietiv AR e diffusion hinn complementar — la diffusion la mejora la pianificazion prospettiva (lookahead planning) menter l'AR el forniss i priorità lenguistegh de manzina a dritta — e che ona analisi "speed-of-light" la suggeriss on potenzial del 76,5% de tokens supplementar per passagg sotta on campionador ottimal.