Real AI models released across the industry in the last 30 days — 2 companies, 2 releases. This is not howfastai.com's own measured data — everywhere else on this site, a number is either a recorded run or a clearly-labelled placeholder from our own tests. Here, specs and pricing are as reported by the vendor or a release tracker, summarized in our own words, sourced on every card, and not independently verified by us.
Not a chat model: it returns a fixed-shape decision (a classification plus a confidence score) instead of writing text, aimed at tasks like tool-call routing and structured data extraction rather than open-ended prompts. Reported 20-200x faster and 40-400x cheaper than comparable LLMs on those tasks; $0.042/M input tokens, output cost described as too low to meaningfully measure. Early access via waitlist, no public API yet.
Smallest model in DeepSeek's new architecture family, with native multimodal (visual) input. 552B total parameters (MoE) using a new causal encoder-decoder design: 8B active parameters for input, 16B for output. KV cache cut to 1/4 the HBM and 1/8 the SSD storage of the previous generation, lowering cache-hit costs. Live now on the DeepSeek API as deepseek-flash (the retired v4-flash and v4-flash-vision-exp IDs route here temporarily); lower pricing than the prior generation, same peak/off-peak split. Open-weight, on Hugging Face.
Want numbers this site has actually measured itself? Full model comparison → · Which model should I use? → · How we measure → · Voice AI →