Log in

We create the world's fastest and most efficient AI.

Score (%) vs Speed (tok/sec, log)
Claude
GPT
GLM
Llama
Rysana
80859095100101001k10k100k1M10MHaiku 4.5Sonnet 5Opus 55.6 Luna5.55.6 Sol5.6 Terra5.4 Nano5 Nano3.5 Turbo5.24.5 Air3.3 70B3.1 70B3.1 8B3 HaikuV2 BaseV2 SmallV2 MiniV2 NanoV1 Small

Benchmark score across over 20k varied real-world tasks with no chain-of-thought, against peak end-to-end speed in tokens per second per request. V2 from July '26 checkpoint, others via fastest available conventional API, measured mid-August '26.

V2 is our latest model family built for unprecedented performance. It can generate millions of tokens per second, up to 10,000× faster and 1000× more efficient than conventional alternatives, letting you deploy AI in all-new latency, cost, and volume regimes.

By requesting access, you agree to our Terms and Privacy Policy.

Natively multimodal, V2 supports text, audio, vision, code, documents, etc.We offer the richest, strongest support for structured outputs anywhere: schemas, grammars, programs, any format you want, with nanosecond overhead and perfect constraint.Constrain model behavior, not just output format. Dynamically program sampling, cost-speed-effort balance, and more at a sub-structural level: a whole new way to engineer AI.Wield the fastest inference in the world, at real production scale. We can support more volume than any other provider, from the fastest chipmakers to the largest clouds and labs.Need trillions of tokens per hour? Millisecond inference end to end? Perfect constraints for your custom grammar? We're opening access to V2. If you need it for what you're building, reach out.Made in America.

FAQ