Why DeepSeek V4 Pro and Flash Redefine GPU Clusters?

Why DeepSeek V4 Pro and Flash Redefine GPU Clusters? - Featured Image

DeepSeek V4 is a new family of models built around compressed sparse attention and hierarchical compressed attention that deliver million token context at roughly 27 percent of the compute cost. In practice it holds huge code bases and very long conversations in active context while staying fast and memory efficient, and its pro variant at … Read more

ERNIE 5.1 Tested in Detail at Mona Vale Beach

ERNIE 5.1 Tested in Detail at Mona Vale Beach - Featured Image

Baidu ERNIE 5.1 beats DeepSeek on the hardest agent math tasks with tools, competes with Claude and Gemini, and was trained at only 6% of the usual budget. It sits in the global top tier on agents, math, knowledge, and search while cutting total parameters to about a third and active parameters to half. They … Read more

Zaya1 8B by Zyphra: Efficient Local Intelligence Run

Zaya1 8B by Zyphra: Efficient Local Intelligence Run - Featured Image

Zaya1 8B is an open-source 8.4 billion parameter Mixture-of-Experts model with only 760 million active parameters per token, trained entirely on AMD hardware, and it posts results that beat some much larger systems on hard math and reasoning. In head-to-head claims, it outperforms Claude 4.5 Sonnet and Gemini 2.5 Pro on hard evaluations, stays competitive … Read more