model

GLM-5.3-Flash

huggingface.co/zai-org/GLM-5.3-Flash ↗

826875 downloads·2200 likes·image-text-to-text·transformers

from the model card

GLM-5.3-Flash 👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report. 📍 Use GLM-5.3-Flash API services on Z.ai API Platform. Introduction We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute. Serve GLM-5.3-Flash Locally GLM-5.3-Flash supports deployment with the following frameworks. Feel free to try them out: SGLang — see cookbook vLLM — see recipes TokenSpeed — see here Transformers — see transformers docs KTransformers — see tutorial Unsloth — see guide Note GLM-5.3-Flash supports controlling…

discussions

recent items

← all models