model

GLM-5.2-NVFP4

huggingface.co/nvidia/GLM-5.2-NVFP4 ↗

365499 downloads·246 likes·text-generation·Model Optimizer

from the model card

Model Overview Description: The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention (with an IndexShare indexer) to support a long context. For more information, please check here. The NVIDIA GLM-5.2 NVFP4 model is quantized with Model Optimizer. This model is ready for commercial or non-commercial use. License/Terms of Use: GOVERNING TERMS: Use of the model is governed by the MIT License, same as the base model. Deployment Geography: Global Use Case: Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI-powered applications. Release Date: Hugging Face 06/25/2026 via https://huggingface.co/nvidia/GLM-5.2-NVFP4 References Nvidia Model Optimizer: https://github.com/NVIDIA/Model-Optimizer Model Architecture: Architecture Type: Transformers Network Architecture: GLM-5.2 (GlmMoeDsaForCausalLM) Number of Model Parameters: 753B in total and 40B activated Input: Input Type(s): Text Input Format(s): String Input Parameters: One-Dimensional (1D) Other Properties Related to Input: Context length up to 1M Output: Output Type(s): Text Output Format: String Output Parameters: 1D (One-Dimensional): Sequen…

discussions

recent items

← all models