TL;DR: Z.ai revealed GLM-5.3-Flash (Ox Alpha), the first native multimodal model in the GLM-5 series, which became the largest model ever on OpenRouter by processing over 20 trillion tokens in six days.
Summary: Z.ai announced GLM-5.3-Flash, codenamed Ox Alpha, as the first native multimodal model in the GLM-5 series. The model was the biggest ever on OpenRouter, processing more than 20 trillion tokens in six days, and remains available via OpenRouter under z-ai/glm-5.3-flash. It is positioned as a high-throughput, cost-efficient multimodal option for builders.
Why it matters: For AI builders, GLM-5.3-Flash signals a strong open-weights-style alternative for high-volume multimodal workloads on OpenRouter. Try it via the OpenRouter endpoint if you need fast, low-cost image-plus-text inference.
Source: x_com