right signal
8 contenders · as of 14 Sep 2026

On the radar

New and emerging models that could change the board. Everything here is tracked from the day it appears: unverified until independent numbers exist, early data once they do — and a formal challenger or title case when the evidence clears our bar.

10 Sep 2026
unverified
weights out

DeepSeek-V4.1-Flash

DeepSeek · could contend for best open-weight LLM. A multimodal open-weight release from DeepSeek, published under an MIT licence with image-text-to-text support and FP8/8-bit checkpoints ready for standard endpoints. The striking part is the permissive licence on a vision-capable Flash-class model shipped in FP8 from day one. No independent evaluation yet, and how it compares with the earlier V4 Flash vision experiment on reasoning or coding remains unknown.

seen at: [1]
8 Sep 2026
unverified
weights out

Nex-N2.5-Pro

Nex AGI · could contend for best open-weight LLM. A text-generation model from a comparatively unknown lab, released on Hugging Face under Apache 2.0. The licence is the notable fact: fully permissive for commercial use, with transformers-compatible weights and no gated access. Nothing else is verified — no published benchmarks, no parameter count in the listing, and no third-party testing.

seen at: [1]
8 Sep 2026
unverified
in product

ChatGPT Images 2.5

OpenAI · could contend for best image generation. OpenAI's image generator inside ChatGPT, pitched at turning sketches and reference photos into more personalised results. The striking part is the emphasis on reference-image fidelity rather than raw prompt following, which is where the current title holder scores. No independent arena or benchmark placement yet, and it is unclear how it relates to the GPT Image 2 API model.

seen at: [1]
6 Sep 2026
early data
API only

Claude Mythos 5.1

Anthropic · could contend for best overall LLM. A second Anthropic 5.1-generation line appearing alongside Fable 5.1, so far visible only through third-party evaluation. It scores 65% on Vellum's aggregate, level with Fable 5.1 and above Opus 5. Anthropic has published nothing about positioning, pricing or effort settings, and no other board has it yet.

seen at: [1]
2 Sep 2026
early data
API only

Gemini 3.8 Flash

Google DeepMind · could contend for best overall LLM. A new Flash-tier Gemini, announced alongside a cyber-specialised sibling, 3.8 Flash Cyber. Shipping a security-hardened variant at the same time as the general model is unusual for this tier. No published evaluations, pricing or context details in the announcement, and nothing on the public boards yet.

seen at: [1]
1 Sep 2026
unverified
weights out

K2-Horizon-MoVA-36B-A4B

IFM · could contend for best open-weight LLM. An open-weight mixture-of-experts model with a custom k2_horizon architecture, 36B total parameters and roughly 4B active. The sparsity ratio puts frontier-style routing within reach of a single workstation card. Untested outside its own claims: no independent benchmark placement, and the custom code path means inference support is limited for now.

seen at: [1]
31 Aug 2026
unverified
preview

hy4-preview

Tencent · could contend for best LLM for coding. A preview build of Tencent's next Hunyuan generation, first surfacing through blind web-development voting. It entered the WebDev arena top five on debut at 1631 Elo, ahead of several released frontier models. There is no model card, licence or serving detail yet, and it has since slipped down the board as newer entrants arrived.

seen at: [1]
24 Aug 2026
unverified
weights out

Spark-X2.5-4B

XHToken · could contend for best small / on-device LLM. A 4B dense chat model on its own spark2_5 architecture, released under Apache 2.0 with a matching base checkpoint. Publishing base and instruct together at this size is welcome and rare from a newer lab. Nothing independent yet on reasoning quality, and the custom architecture will need runtime support before it is practical on device.

seen at: [1]