What happened?

On 20 August 2026, an unsigned model landed on OpenRouter as stealth/ox-alpha and, minutes later, on OpenCode. The pitch: a reasoning model for coding, sustained agentic work, and production workloads. Specs the operator actually published: 1,048,576 context, 131,072 max output, text / image / video in, tool calling, JSON. Price: $0 / $0. Window: about a week (OpenCode’s follow-up: through ~27 August, and it does not burn Go quota). Capacity claim, from OpenCode: 100 trillion tokens per day. “Let’s see what you can do.”

OpenRouter’s stealth note is the part that matters more than the ninja emoji: the provider does not train on prompts or completions, but does retain them. OpenCode marketed zero data retention. Those two sentences are not the same policy. One provider. One anonymous name. Two marketing layers.

By 22 August, OpenRouter’s own activity panel on the model page showed 2.61T prompt tokens and 31.6B completion tokens. Cache hit rate in the 82–85% band. Tool-call error ~2%. Availability ~99.5%. Top public apps on the pipe: Hermes Agent (728B), Claude Code (433B), omp, DeepSeek Harness, ZCode. Patrick Collison’s review was one line: $ ori --model stealth/ox-alpha, “It’s very impressive.” Cline, ChatLLM, and OpenCode Go all turned it on for free. Nobody put a lab on the card.

The internet spent the weekend playing Guess The Lab. Tokenizer and video-encoder probes match GLM-5.3 / GLM-5V closely (same token counts across 25 prompts with a fixed ~+75 wrapper; same video-token curve as GLM-5V-Turbo; audio rejected the GLM way). Manifold priced Z.ai / Zhipu ~84%. Pliny’s agent said GLM-5.X. LobeHub’s Max For AI said 100% Chinese, North American servers because that is where the vendor sits, and the team will surprise people who think they already know. A Microsoft-tokenizer theory and a Pinduoduo/Temu rumor are also in the pile. No lab has claimed it. Treat every name as a bet, not a byline.

Why this is interesting

  • The free week is the launch: Not a blog, not a board, not a keynote. A stealth slug plus a week of uncapped agent traffic. Hermes, Claude Code, and ZCode were already the customers before anyone had an Elo. Pony / Hunter / Elephant / Owl did this first: unsigned animal, free flood, brand later. The lack of a name is the distribution.
  • The token mix is the tell: 2.61T in, 31.6B out in two days is not chat. It is long-horizon agents rereading the same repo with an 85% cache. OpenRouter currently prints 0 reasoning tokens; that is wrapper telemetry until proven otherwise, not proof the model does not think. Teortaxes’ napkin: 100T/day of output at DeepSeek-V4 rates is ~580k GPUs. 100T of “all tokens” with cache is a small cluster brag. Believe the second until someone shows the first.
  • Mystery is a feature, not a leak: A GLM fingerprint is the favorite because tokenizer math is hard to fake. A “surprise Chinese team” is the competing insider line. Both can be true (fine-tune on a GLM substrate; a lab you would not have in the pool). Neither is a reason to route client code. The unsigned card is how you collect the world’s agent traces without inheriting the brand’s enemies for seven days.
  • Harness gravity beat the leaderboard: No Artificial Analysis row at time of writing. The 80% DeepSWE screenshot was a 10-task subset. DeepSWE’s author, on a larger slice, landed ~63% at ~47k average output tokens, just shy of Grok 4.6, in the Gemini 3.7 Flash / DeepSeek V4 Pro band, “big model smell.” Ben Davis’s operator note after living in it: good voice, decent design, handles subagents, leaves dead code, feels slow at high reasoning, Sol-medium. Abacus posted a table that put it next to Kimi 2.6; X did not buy it. Bake on your harness. The subset is not a coronation.
  • This is the August price war with the label ripped off: Same week as GLM-5.3 (text, cyber hold, weights in two weeks), Grok 4.6 at $2/$6, 3.7 Flash at $0.75 intro. Ox Alpha’s list is $0 until Thursday. After that, the product is whatever rate card the still-unnamed lab is willing to own. Do not budget 2026 on a ninja.

What it is not

Not a confirmed Zhipu, Xiaomi, xAI, Google, Microsoft, or Pinduoduo model. Not proof it beats Fable 5 or Sol as a default. Not an 80% DeepSWE result, that was a toy slice; ~63% is the number from the person who wrote the bench. Not zero retention just because OpenCode said the words. Not 100T/day of generated tokens. Not a production dependency. Not a Swiss-safe pipe for client repos sitting in an anonymous North American endpoint.

Bottom line

Ox Alpha is a week-long inference subsidy attached to a nameless API. The interesting output is not the fluid-sim demos. It is 2.6 trillion prompt tokens of other people’s agent loops, harvested under a stealth ToS, while the timeline argues about oxen. Use it this week on work you would paste into a stranger’s GPU. Put a fallback behind it. On ~27 August watch three things: who steps forward, what the rate card is, and whether the Flash-that-fits-on-two-Sparks rumor survives contact with a Hugging Face card. Until then the lab is whoever can afford to give the internet a free gym.