← All stories
● Covered by 2 sources · 3 reportsLow impact3 neutral

GLM-5.3 (max) Model Benchmarked by Artificial Analysis

🔄 Updated 2d ago — new reporting from The New Stack
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • GLM-5.3 (max) scored 60 on Artificial Analysis Intelligence Index.
  • It has a 1M token context window and supports text input/output.
  • Pricing is $1.40/1M input tokens and $4.40/1M output tokens.
  • The model operates at 74 tokens per second, slower than average.
  • Z AI released GLM-5.3-Flash on August 26, 2026.
  • GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index.
  • GLM-5.3-Flash has a 400k token context window.
  • GLM-5.3-Flash pricing is $0.15/1M input tokens and $0.50/1M output tokens.
  • GLM-5.3-Flash generated 150M tokens during Intelligence Index evaluation.
  • It cost $138.02 to evaluate GLM-5.3-Flash on the Intelligence Index.
  • GLM-5.3-Flash is 3x faster than GLM-5.3.
  • GLM-5.3-Flash architecture cuts attention computation by 3x.
  • GLM-5.3-Flash costs $0.075/1M input tokens and $0.25/1M output tokens on OpenRouter.
  • GLM-5.3 costs $1.188/1M input tokens and $4.18/1M output tokens on OpenRouter.

Model Overview and Performance

Artificial Analysis conducted benchmarks on the GLM-5.3 (max) model, which was released on August 18, 2026, by Z AI. The model achieved a score of 60 on the Artificial Analysis Intelligence Index, surpassing the median score of 35 for comparable models. During evaluation, it generated 170 million tokens, which is higher than the median of 72 million tokens.

Pricing and Cost Efficiency

The pricing for GLM-5.3 (max) is set at $1.40 per 1 million input tokens and $4.40 per 1 million output tokens. These rates are considered moderately priced compared to the median rates of $1.75 for input and $10.00 for output tokens. The total cost to evaluate GLM-5.3 (max) on the Intelligence Index was $1238.50.

Speed and Context Window

GLM-5.3 (max) processes at a speed of 74 tokens per second, which is slower than the average speed for models in its class. The model supports text input and output and features a 1 million token context window.

Benchmarking Methodology

Models are compared based on their class, including non-reasoning, reasoning, open-weights (categorized by parameter size: Tiny, Small, Medium, Large), and proprietary models (categorized by price range). Proprietary models are compared across proprietary and open-weights models within the same price range, using a blended 3:1 input/output price ratio.

Updates

🕒 2026-09-01 · new reporting from The New Stack
  • GLM-5.3-Flash is 3x faster than GLM-5.3.
  • GLM-5.3-Flash architecture cuts attention computation by 3x.
  • GLM-5.3-Flash costs $0.075/1M input tokens and $0.25/1M output tokens on OpenRouter.
  • GLM-5.3 costs $1.188/1M input tokens and $4.18/1M output tokens on OpenRouter.
🕒 2026-08-26 · new reporting from Hacker News Front Page
  • Z AI released GLM-5.3-Flash on August 26, 2026.
  • GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index.
  • GLM-5.3-Flash has a 400k token context window.
  • GLM-5.3-Flash pricing is $0.15/1M input tokens and $0.50/1M output tokens.
  • GLM-5.3-Flash generated 150M tokens during Intelligence Index evaluation.
  • It cost $138.02 to evaluate GLM-5.3-Flash on the Intelligence Index.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~19 min · 16 stories · Sep 04

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

Z.AI released GLM-5.3-Flash, claiming it offers stronger intelligence at lower cost and 3x faster serving speed compared to GLM-5.3. Initial tests were conducted on coding, reasoning, and information extraction tasks to evaluate these claims and determine the practical value of the new model.

Z AI released GLM-5.3-Flash on August 26, 2026, a new AI model that scores 57 on the Artificial Analysis Intelligence Index and features a 400k token context window. The model offers competitive pricing at $0.15 per 1M input tokens and $0.50 per 1M output tokens.

Artificial Analysis has benchmarked the GLM-5.3 (max) model, finding it scores 60 on their Intelligence Index, placing it above the median for comparable models. The model is noted for its 1M token context window, moderate pricing, and slower-than-average speed.