GLM-5.3-FlashX Officially Launches on B.AI, Now Available on Both API and Web Chat
Odaily News: The B.AI platform has announced the official launch of GLM-5.3-FlashX, a speed-optimized native multimodal large model developed by Z.AI, now simultaneously available for experience on both API and Web Chat. Built on an efficient sparse architecture with 320B total parameters and 18B activated parameters, the model boosts generation speed to up to 200 tokens/s while maintaining its original level of intelligence — approximately a 5x speedup compared to GLM-5.3-Flash. The model supports a 1M-token ultra-long context and full-dimensional native video/file understanding. Combined with interleaved tool-calling reasoning capabilities, it is purpose-built for high-frequency interactive coding and agent workflows. Starting today, developers can directly access GLM-5.3-FlashX through the B.AI platform. Log in to chat.b.ai/chat to experience it now.
