← All stories
● Covered by 4 sources · 8 reportsLow impact8 neutral

Ox Alpha Reasoning Model Released by Anonymous Third-Party Provider via OpenRouter

🔄 Updated 8d ago — new reporting from VentureBeat
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Ox Alpha is a reasoning model for coding and agentic work.
  • It is developed by an anonymous third-party provider.
  • Available for free with a 1M context window.
  • Prompts and completions are not used for training.
  • Ox Alpha was released on OpenRouter.
  • Stripe CEO Patrick Collison described Ox Alpha as "very impressive."
  • Speculation about its origin includes GLM models from Z.ai and Microsoft's MAI.
  • Ox Alpha appeared on OpenRouter on August 20.
  • OpenCode debuted Ox Alpha the same day.
  • OpenCode announced a capacity of 100 trillion tokens a day.
  • Prompt injection and gzip-NCD compression identified Ox Alpha as GLM.
  • Ox Alpha's system prompt identifies it as "ox-alpha" developed by an undisclosed organization.
  • Ox Alpha's tokenizer matches the GLM-5.x vocabulary.
  • Ox Alpha censors seven topics, including domestic incidents and Xi Jinping.
  • Ox Alpha answers identically to American models on Xinjiang and Taiwan topics.
  • Z.ai confirmed it is the developer of Ox Alpha.
  • Ox Alpha is the newest iteration of Z.ai's GLM series.
  • Z.ai will release Ox Alpha weights on Wednesday.
  • Ox Alpha is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.
  • Z.ai released GLM-5.3 earlier this month.
  • Ox Alpha is Z.ai's GLM-5.3-Flash.
  • GLM-5.3-Flash is a 320 billion-parameter hybrid model.
  • GLM-5.3-Flash has 18 billion active parameters.
  • Z.ai trained GLM-5.3-Flash for ultra-low-cost inference.
  • Z.ai released GLM-5.3-Flash weights on Hugging Face.
  • GLM-5.3-Flash is released under an MIT license.
  • GLM-5.3-Flash costs $0.075 per million input tokens on OpenRouter.
  • GLM-5.3-Flash costs $0.25 per million output tokens on OpenRouter.
  • OpenRouter prices for GLM-5.3-Flash reflect a 50% discount.
  • GLM-5.3-Flash performs comparably to Claude Opus 4.8 and OpenAI's GPT-5.6 Terra.
  • GLM-5.3-Flash scores 57 points on the Artificial Analysis Intelligence Index.
  • GLM-5.3-Flash operates on Chinese chips.
  • Ox Alpha was free.
  • Community estimates for Ox Alpha usage ranged from single digits to over 20 trillion tokens daily.
  • Z.ai ran Ox Alpha on public traffic on purpose.
  • GLM-5.3-Flash list price is 15 cents per million input tokens.
  • GLM-5.3-Flash list price is 50 cents per million output tokens.

Introduction of Ox Alpha

Ox Alpha, a new reasoning model, has been made available. It is specifically designed for coding, sustained agentic work, and production workloads, excelling in long-horizon software engineering and complex reasoning tasks that integrate text and visual context.

Provider and Availability

The model is developed and operated by an anonymous third-party provider, with OpenRouter facilitating access. It was released on August 20, 2026, and is currently offered for free. The model provides a 1M context window.

Data Privacy and Performance

Prompts and completions submitted to Ox Alpha are retained by the provider but are not used for training purposes. Performance metrics indicate a throughput of 57 tokens per second and a latency of 1.78 seconds. The model boasts high uptime and availability, with OpenRouter implementing load balancing to ensure continuous service.

Updates

🕒 2026-08-27 · new reporting from VentureBeat
  • Ox Alpha was free.
  • Community estimates for Ox Alpha usage ranged from single digits to over 20 trillion tokens daily.
  • Z.ai ran Ox Alpha on public traffic on purpose.
  • GLM-5.3-Flash list price is 15 cents per million input tokens.
  • GLM-5.3-Flash list price is 50 cents per million output tokens.
🕒 2026-08-26 · new reporting from The New Stack
  • Ox Alpha is Z.ai's GLM-5.3-Flash.
  • GLM-5.3-Flash is a 320 billion-parameter hybrid model.
  • GLM-5.3-Flash has 18 billion active parameters.
  • Z.ai trained GLM-5.3-Flash for ultra-low-cost inference.
  • Z.ai released GLM-5.3-Flash weights on Hugging Face.
  • GLM-5.3-Flash is released under an MIT license.
  • GLM-5.3-Flash costs $0.075 per million input tokens on OpenRouter.
  • GLM-5.3-Flash costs $0.25 per million output tokens on OpenRouter.
  • OpenRouter prices for GLM-5.3-Flash reflect a 50% discount.
  • GLM-5.3-Flash performs comparably to Claude Opus 4.8 and OpenAI's GPT-5.6 Terra.
  • GLM-5.3-Flash scores 57 points on the Artificial Analysis Intelligence Index.
  • GLM-5.3-Flash operates on Chinese chips.
🕒 2026-08-26 · new reporting from TechCrunch
  • Z.ai confirmed it is the developer of Ox Alpha.
  • Ox Alpha is the newest iteration of Z.ai's GLM series.
  • Z.ai will release Ox Alpha weights on Wednesday.
  • Ox Alpha is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.
  • Z.ai released GLM-5.3 earlier this month.
🕒 2026-08-25 · new reporting from Hacker News Front Page
  • Ox Alpha's tokenizer matches the GLM-5.x vocabulary.
  • Ox Alpha censors seven topics, including domestic incidents and Xi Jinping.
  • Ox Alpha answers identically to American models on Xinjiang and Taiwan topics.
🕒 2026-08-25 · new reporting from Hacker News Front Page
  • Prompt injection and gzip-NCD compression identified Ox Alpha as GLM.
  • Ox Alpha's system prompt identifies it as "ox-alpha" developed by an undisclosed organization.
🕒 2026-08-24 · new reporting from The New Stack
  • Ox Alpha appeared on OpenRouter on August 20.
  • OpenCode debuted Ox Alpha the same day.
  • OpenCode announced a capacity of 100 trillion tokens a day.
🕒 2026-08-23 · new reporting from TechCrunch
  • Ox Alpha was released on OpenRouter.
  • Stripe CEO Patrick Collison described Ox Alpha as "very impressive."
  • Speculation about its origin includes GLM models from Z.ai and Microsoft's MAI.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~19 min · 16 stories · Sep 04

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

Z.ai has revealed that the previously anonymous Ox Alpha model, which gained significant traction for its performance and free availability, is GLM-5.3-Flash. This model is notable for being served entirely on Chinese chips and infrastructure, offering competitive performance at a significantly lower cost compared to US-based alternatives.

Z.ai has unmasked and released GLM-5.3-Flash, a 320 billion-parameter hybrid AI model, under an MIT license on Hugging Face. This model offers competitive performance with significantly lower inference costs, partly due to its operation on Chinese chips, making advanced AI more accessible.

Z.ai has been identified as the AI lab behind Ox Alpha, an open-weight AI model that has topped benchmarks on OpenRouter. This release adds to the growing competition from capable models from China, potentially impacting market share for established AI companies.

Analysis of the Ox Alpha language model, which appeared on OpenRouter, indicates it belongs to the GLM family based on tokenizer matching. The model exhibits a bimodal censorship profile, specifically censoring seven topics related to domestic incidents and Xi Jinping, while appearing uncensored on other sensitive topics like Xinjiang and Taiwan.

An anonymous AI model named Ox Alpha appeared on OpenRouter, prompting questions about its origin and how it handles user code. While the identity of its creator is a subject of speculation, the more critical concern for developers is the legal entity receiving their code and the terms under which it is processed.

Analysis using prompt injection and gzip-NCD compression techniques has identified Ox-Alpha, a previously anonymous large language model on OpenRouter, as GLM, developed by Z.ai. This discovery reveals the true origin of a model that had been gaining attention on leaderboards without disclosing its developer.

A new AI model named Ox Alpha was released on OpenRouter, described as a "reasoning model designed for coding, sustained agentic work, and production workload." The developer of Ox Alpha has chosen to remain anonymous, leading to speculation about its origin.

A new reasoning model named Ox Alpha, designed for coding and agentic work, has been released by an anonymous third-party provider through OpenRouter. The model is free, offers a 1M context window, and its prompts and completions are not used for training.