# An Anonymous Model Called Ox Alpha Tops an Independent Coding Benchmark With No Disclosed Owner

A developer's run of the DeepSWE agent benchmark put the stealth model at 80 percent against 65 percent for Claude Fable 5 and 52 percent for GPT-5.6 Sol, but no vendor has published the result or claimed the model.

- Published: 2026-08-24T07:20:23.472Z
- Canonical: https://polylog.news/ai/2026-08-24/an-anonymous-model-called-ox-alpha-tops-an-independent-codin
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [OpenRouter](https://openrouter.ai/stealth), [explainx.ai](https://explainx.ai/blog/openrouter-ox-alpha-stealth-model-august-2026), [Local AI Zone](https://local-ai-zone.github.io/blog/ox-alpha-stealth-model-comprehensive-analysis.html)

A model listed only as Ox Alpha appeared on [OpenRouter's stealth channel](https://openrouter.ai/stealth) and in OpenCode on 20 August, offered free, with a context window of 1,048,576 tokens, text, image and video input, and a maximum output of roughly 128,000 tokens. The provider is undisclosed. [Coverage of the listing](https://explainx.ai/blog/openrouter-ox-alpha-stealth-model-august-2026) notes only that a third party operates it and has chosen to stay anonymous during the preview.

The attention comes from a single independent run. The developer Ben Davis put Ox Alpha through DeepSWE, a coding-agent benchmark, and [reported 80 percent](https://local-ai-zone.github.io/blog/ox-alpha-stealth-model-comprehensive-analysis.html) against 65 percent for Claude Fable 5 and 52 percent for GPT-5.6 Sol. That is one test setup, one configuration, one runner, and neither OpenRouter nor the anonymous provider has published a benchmark card. Treat the ranking as something that needs independent reproduction, not as a settled result.

Community tokeniser probes have pointed toward Z.ai's GLM family, which released GLM-5.2 Turbo on 17 August, but nobody has confirmed the model's origin. The stealth-listing pattern itself is now routine: a lab ships an unbranded endpoint, collects real agentic traffic for free for about a week, then attaches a name and a price once it knows how the model behaves under load.

For engineers, the practical facts do not depend on who built the model. A million-token window with video input, free during preview, is a cheap way to test long-horizon agent workloads that would otherwise be expensive to evaluate. The free window is expected to close after roughly a week, and neither pricing nor continued availability has been confirmed.

## What this means

Stealth preview listings are becoming the standard distribution channel for pre-launch frontier coding models, and the party that gains is the router: OpenRouter and similar aggregators now sit between labs and the evaluation traffic that determines how a new model's launch is perceived. Anthropic and OpenAI are exposed through a specific channel: a free anonymous endpoint that beats their flagship models in even one credible independent run changes the terms of the price comparison before they can respond with their own numbers. The decisive question is whether an evaluation using multiple independent test setups confirms the 80 percent DeepSWE figure, or whether the score falls back toward other models' results once other testers try it, as usually happens with single-run leaderboard jumps.

## What to watch

- Whether an established evaluator reproduces the DeepSWE ordering on Ox Alpha with a published configuration, which is what separates a real capability step from a favourable test setup.
- The price the model carries when the free window ends and it gets a name, since a low launch price would mark another round in the coding-model price competition rather than a capability claim.
- Whether the provider turns out to be a Chinese lab, which would say something about who is now willing to compete for Western developer mindshare through anonymous distribution.
