model

Qwen3.6-35B-A3B-NVFP4

huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 ↗

9864413 downloads·507 likes·text-generation·Model Optimizer

from the model card

Model Overview Description: The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. Third-Party Community Consideration This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.6-35B-A3B) Model Card from Alibaba. References NVIDIA Model Optimizer: https://github.com/NVIDIA/Model-Optimizer License/Terms of Use: Apache license 2.0 Deployment Geography: Global Use Case: Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI-powered applications. Release Date: Hugging Face on 05/28/2026 via https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 Model Architecture: Architecture Type: Transformers Network Architecture: Mixture-of-Experts (MoE) with Hybrid Attention Number of Model Parameters: 35B in total and 3B activated Input: Input Type(s): Text, Image, Video Input Format(s): String, Red, Green, Blue (RGB), Video (MP4/WebM) Input Parameters: One-Dimensional (1D), Two-Dimensional …

discussions

recent items

← all models