Architecture
Architecture
Architektur
Architettura
Architettura
Qwen3.8-Flash-Next : 125B de paramètres pour un neuvième des FLOPs d'entraînementQwen3.8-Flash-Next: 125B Parameters for One-Ninth of the Training FLOPsQwen3.8-Flash-Next: 125B Parameter für ein Neuntel der Trainings-FLOPsQwen3.8-Flash-Next: 125B di parametri per un nono dei FLOPs di addestramentoQwen3.8-Flash-Next: 125B de parameter per on noven di FLOPs de addestrament
Le paper d'architecture soumis le 1er septembre 2026 sur Hugging Face Daily Papers détaille Qwen3.8-Flash-Next : un MoE creux de 125B de paramètres, 6B activés par token, plus 51B de tables d'embeddings n-gram hors accélérateur. Sur 14 benchmarks de pré-entraînement, il devance son prédécesseur 397B-A17B sur 8 et le suit au pire de 2,6 points — pour 1/3 des paramètres activés, 1/3 des tokens d'entraînement et ~1/9 des FLOPs. Trois choix structurels : hybride Gated DeltaNet/attention (une couche full-attention sur quatre, remplacée en continued-pretraining par Qwen Sparse Attention), un Gated Residual à quatre branches, et l'optimiseur Muon qui déplace le learning rate optimal vers le haut —
arXiv.
The architecture paper featured on Tuesday, 1 September 2026 on Hugging Face Daily Papers details Qwen3.8-Flash-Next: a sparse MoE of 125B parameters, 6B activated per token, plus 51B of n-gram embedding tables held off-accelerator. Across 14 pre-training benchmarks, it beats its predecessor 397B-A17B on 8 and trails by at most 2.6 points — for one-third of the activated parameters, one-third of the training tokens and ~1/9 of the FLOPs. Three structural choices: a Gated DeltaNet/attention hybrid (one full-attention layer in four, replaced during continued pre-training by Qwen Sparse Attention), a four-branch Gated Residual, and the Muon optimizer, which shifts the optimal learning rate upward —
arXiv.
Das am 1. September 2026 auf Hugging Face Daily Papers eingereichte Architektur-Paper beschreibt Qwen3.8-Flash-Next: ein spärliches MoE mit 125B Parametern, davon 6B pro Token aktiviert, plus 51B an n-Gramm-Embedding-Tabellen ausserhalb des Beschleunigers. Von 14 Pre-Training-Benchmarks übertrifft es seinen Vorgänger 397B-A17B auf 8 und liegt im schlechtesten Fall bloss 2,6 Punkte dahinter – bei einem Drittel der aktivierten Parameter, einem Drittel der Trainings-Tokens und rund einem Neuntel der FLOPs. Drei strukturelle Entscheidungen: eine hybride Gated-DeltaNet/Attention-Architektur (eine von vier Schichten mit voller Attention, im Continued Pretraining durch Qwen Sparse Attention ersetzt), ein Gated Residual mit vier Zweigen sowie der Optimierer Muon, der die optimale Lernrate nach oben verschiebt —
arXiv.
Il paper di architettura presentato il 1° settembre 2026 su Hugging Face Daily Papers descrive nel dettaglio Qwen3.8-Flash-Next: un MoE sparso da 125B di parametri, 6B attivati per token, più 51B di tabelle di embedding n-gram fuori dall'acceleratore. Su 14 benchmark di pre-addestramento supera il predecessore 397B-A17B su 8 e lo segue nel caso peggiore di 2,6 punti — con 1/3 dei parametri attivati, 1/3 dei token di addestramento e ~1/9 dei FLOPs. Tre scelte strutturali: ibrido Gated DeltaNet/attenzione (un livello di full-attention su quattro, sostituito in continued-pretraining da Qwen Sparse Attention), un Gated Residual a quattro rami e l'ottimizzatore Muon che sposta il learning rate ottimale verso l'alto —
arXiv.
El paper de architettura presentaa el primm de settember 2026 sora Hugging Face Daily Papers el detaja Qwen3.8-Flash-Next: on MoE sbujaa de 125B parameter, 6B ativaa per token, pussee 51B de tabell d'embedding n-gram foeura de l'acelerator. Sora 14 benchmark de pre-addestrament, el passa el sò predecessur 397B-A17B in su 8 e ghe va adree al pesg de 2,6 pont — per 1/3 di parameter ativaa, 1/3 di token de addestrament e ~1/9 di FLOPs. Trii scerni struturai: ibrid Gated DeltaNet/attention (on strat full-attention su quatter, sostituii in continued-pretraining de Qwen Sparse Attention), on Gated Residual a quatter ram, e l'otimizador Muon che 'l moeuv el learning rate òptim vers alt —
arXiv.
Multimodal
Multimodal
Multimodal
Multimodale
Multimodal
DreamX-Creator : génération audio-vidéo native et synchronisée en 2KDreamX-Creator: Native, Synchronized 2K Audio-Video GenerationDreamX-Creator: native und synchronisierte Audio-Video-Generierung in 2KDreamX-Creator: generazione audio-video nativa e sincronizzata in 2KDreamX-Creator: generazion audio-video nativa e sincronizada in 2K
Publié le 31 août 2026 et bien reçu (39 upvotes sur HF Daily Papers), DreamX-Creator 1.0 est un générateur audio-vidéo natif de 7B paramètres : débruitage conjoint des flux audio et vidéo via Gated Cross-Modal Attention, RL post-entraînement avec feedback multimodal par modalité, et un raffineur autoregressif 1-step pour la résolution 2K —
arXiv. Les auteurs publient le générateur 7B et le 2K Refiner pour démocratiser la génération audio-vidéo synchronisée.
Published on 31 August 2026 and well received (39 upvotes on HF Daily Papers), DreamX-Creator 1.0 is a 7B-parameter native audio-video generator: joint denoising of the audio and video streams via Gated Cross-Modal Attention, post-training RL with per-modality multimodal feedback, and a 1-step autoregressive refiner for 2K resolution —
arXiv. The authors release the 7B generator and the 2K Refiner to democratize synchronized audio-video generation.
Am 31. August 2026 veröffentlicht und gut aufgenommen (39 Upvotes auf HF Daily Papers): DreamX-Creator 1.0 ist ein nativer Audio-Video-Generator mit 7B Parametern – gemeinsames Denoising der Audio- und Video-Ströme via Gated Cross-Modal Attention, RL-Post-Training mit modalitätsspezifischem multimodalem Feedback sowie ein autoregressiver 1-Schritt-Verfeinerer für die 2K-Auflösung —
arXiv. Die Autoren veröffentlichen den 7B-Generator und den 2K Refiner, um die synchronisierte Audio-Video-Generierung zu demokratisieren.
Pubblicato il 31 agosto 2026 e ben accolto (39 upvote su HF Daily Papers), DreamX-Creator 1.0 è un generatore audio-video nativo da 7B di parametri: denoising congiunto dei flussi audio e video tramite Gated Cross-Modal Attention, post-addestramento RL con feedback multimodale per modalità e un raffinatore autoregressivo a 1 step per la risoluzione 2K —
arXiv. Gli autori pubblicano il generatore 7B e il 2K Refiner per democratizzare la generazione audio-video sincronizzata.
Publicaa el 31 de agost 2026 e ben ricevuu (39 upvote sora HF Daily Papers), DreamX-Creator 1.0 l'è on generator audio-video nativ de 7B parameter: denoising congiunt di fluss audio e video via Gated Cross-Modal Attention, RL post-addestrament con feedback multimodal per modaletaa, e on raffinator autoregressiv 1-step per la resoluzion 2K —
arXiv. I autor i publichen el generator 7B e 'l 2K Refiner per democratizzà la generazion audio-video sincronizada.
Embodied AI
Embodied AI
Embodied AI
Embodied AI
Embodied AI
Lucida reconstruit des scènes intérieures éditables pour la robotiqueLucida Reconstructs Editable Interior Scenes for RoboticsLucida rekonstruiert editierbare Innenraum-Szenen für die RobotikLucida ricostruisce scene interne editabili per la roboticaLucida el recostruess scen de denter editabel per la robotech
Le 1er septembre 2026, un paper de ByteDance Seed (32 upvotes) décrit Lucida, pipeline de modélisation composable real-to-Sim : parsing en graphe de scène, génération d'assets par instance, puis placement par GizmoAct, une politique VLM qui manipule les gizmos en boucle fermée. Gains chiffrés : +69% de mAP sur R2S-Scene face à Boxer, ADD-SB@0.05 de 57,8% à 83,4% sur CA-1M, et scene F-Score de 0,794 à 0,924 face à SAM3D —
arXiv.
On 1 September 2026, a paper from ByteDance Seed (32 upvotes) described Lucida, a composable real-to-Sim modeling pipeline: scene-graph parsing, per-instance asset generation, then placement via GizmoAct, a VLM policy that manipulates gizmos in a closed loop. Quantified gains: +69% mAP on R2S-Scene over Boxer, ADD-SB@0.05 up from 57.8% to 83.4% on CA-1M, and scene F-Score from 0.794 to 0.924 over SAM3D —
arXiv.
Am 1. September 2026 beschreibt ein Paper von ByteDance Seed (32 Upvotes) Lucida, eine komponierbare Real-to-Sim-Modellierungspipeline: Parsing in einen Szenengraphen, Asset-Generierung pro Instanz und anschliessende Platzierung durch GizmoAct, eine VLM-Policy, die Gizmos in geschlossener Regelung manipuliert. Bezifferte Zuwächse: +69% mAP auf R2S-Scene gegenüber Boxer, ADD-SB@0.05 von 57,8% auf 83,4% auf CA-1M und Scene-F-Score von 0,794 auf 0,924 gegenüber SAM3D —
arXiv.
Il 1° settembre 2026 un paper di ByteDance Seed (32 upvote) descrive Lucida, una pipeline di modellazione componibile real-to-Sim: parsing in grafo di scena, generazione di asset per istanza, poi posizionamento tramite GizmoAct, una policy VLM che manipola i gizmo in loop chiuso. Guadagni quantificati: +69% di mAP su R2S-Scene rispetto a Boxer, ADD-SB@0,05 dal 57,8% all'83,4% su CA-1M e scene F-Score da 0,794 a 0,924 contro SAM3D —
arXiv.
El primm de settember 2026, on paper de ByteDance Seed (32 upvote) el descriv Lucida, pipeline de modellazion componibila real-to-Sim: parsing a graf de scena, generazion d'asset per istanza, poeu piazzament con GizmoAct, 'na politega VLM che la manipola i gizmo in esamidor serraa. Guadagn quantificaa: +69% de mAP sora R2S-Scene contra Boxer, ADD-SB@0.05 del 57,8% al 83,4% sora CA-1M, e scene F-Score de 0,794 a 0,924 contra SAM3D —
arXiv.
Fine-tuning
Fine-tuning
Fine-tuning
Fine-tuning
Fine-tuning
NoRA stabilise l'entraînement LoRA par simple normalisationNoRA Stabilizes LoRA Training Through Simple NormalizationNoRA stabilisiert das LoRA-Training durch schlichte NormalisierungNoRA stabilizza l'addestramento LoRA tramite semplice normalizzazioneNoRA el stabilizza l'addestrament LoRA cont la sola normalizzazion
Publié le 31 août 2026, NoRA normalise les matrices de down-projection pendant l'entraînement LoRA : convergence accélérée, meilleure stabilité et mitigation de l'oubli catastrophique en pré-entraînement, SFT et RL — sans paramètres supplémentaires ni coût en inférence —
arXiv. À lire aussi sur HF Daily Papers ce 1er septembre : CAST, supervision critique-aware pour agents tool-calling qui dépasse GPT-OSS-120B de plus de 10% en pass^4 sur Retail (
arXiv), et AutoSciRub, induction automatique de rubriques pour agents de recherche scientifique, +16,8 points en moyenne sur un sous-ensemble d'AstaBench (
arXiv).
Published on 31 August 2026, NoRA normalizes the down-projection matrices during LoRA training: faster convergence, better stability and mitigation of catastrophic forgetting in pre-training, SFT and RL — with no extra parameters and no inference cost —
arXiv. Also worth reading on HF Daily Papers this 1 September: CAST, critique-aware supervision for tool-calling agents that exceeds GPT-OSS-120B by more than 10% in pass^4 on Retail (
arXiv), and AutoSciRub, automatic rubric induction for scientific research agents, +16.8 points on average on a subset of AstaBench (
arXiv).
Am 31. August 2026 veröffentlicht: NoRA normalisiert die Down-Projektions-Matrizen während des LoRA-Trainings – schnellere Konvergenz, bessere Stabilität und Milderung des katastrophalen Vergessens bei Pre-Training, SFT und RL – ohne zusätzliche Parameter und ohne Inferenzkosten —
arXiv. Ebenfalls am 1. September auf HF Daily Papers: CAST, kritik-bewusste Supervision für Tool-Calling-Agenten, das GPT-OSS-120B auf Retail in pass^4 um mehr als 10% übertrifft (
arXiv), sowie AutoSciRub, automatische Induktion von Bewertungsrubriken für Agenten der wissenschaftlichen Recherche, im Schnitt +16,8 Punkte auf einer Teilmenge von AstaBench (
arXiv).
Pubblicato il 31 agosto 2026, NoRA normalizza le matrici di down-projection durante l'addestramento LoRA: convergenza accelerata, maggiore stabilità e mitigazione dell'oblio catastrofico in pre-addestramento, SFT e RL — senza parametri aggiuntivi né costi in inferenza —
arXiv. Da leggere sempre su HF Daily Papers questo 1° settembre: CAST, supervisione critica consapevole per agenti tool-calling che supera GPT-OSS-120B di oltre il 10% in pass^4 su Retail (
arXiv), e AutoSciRub, induzione automatica di rubriche per agenti di ricerca scientifica, +16,8 punti in media su un sottoinsieme di AstaBench (
arXiv).
Publicaa el 31 de agost 2026, NoRA el normalizza i matris de down-projection durant l'addestrament LoRA: convergenza pussee svelta, stabilitaa mej e mitigazion de l'oblid catastrofegh in pre-addestrament, SFT e RL — senza parameter pussee né cost in inferenza —
arXiv. De legg ancamò sora HF Daily Papers 'sto primm de settember: CAST, supervision critic-aware per agent tool-calling che 'l passa GPT-OSS-120B de pussee de 10% in pass^4 sora Retail (
arXiv), e AutoSciRub, induzion automatiga de rubrigh per agent de ricerca scentifega, +16,8 pont in media sora on sottainsiema de AstaBench (
arXiv).