In April, the US Center for AI Standards and Innovation (CAISI) conducted an independent evaluation of DeepSeek V4 Pro. The results show this domestic open-source flagship model is about 8 months behind US frontier models in capability, but shows significant advantages in cost efficiency.

CAISI's evaluation covered five major domains — cybersecurity, software engineering, natural science, abstract reasoning, and mathematics — using 16 benchmarks and 35 reference models. The results show DeepSeek V4's overall capability is roughly equivalent to GPT-5, about 8 months behind GPT-5.5.

But DeepSeek V4 pulls back a city in cost efficiency: in 7 benchmarks, 5 are cheaper than GPT-5.4 mini, with cost differences ranging from 53% cheaper to 41% more expensive.

In software engineering, DeepSeek V4 scores 74% on SWE-Bench, second only to GPT-5.5 (81%) and Opus 4.6 (79%), leading GPT-5.4 mini's 73%. But on the cybersecurity benchmark CTF-Archive-Diamond, DeepSeek V4 only gets 32%, far below GPT-5.5's 71%.

More notably, there's significant divergence between DeepSeek's official self-assessment and CAISI's real testing. DeepSeek's self-claim is that V4 is on par with Opus 4.6 and GPT-5.4, but CAISI's evaluation shows its actual performance is closer to GPT-5 level. This reflects methodological divergence between self-evaluation and third-party evaluation in the current AI industry.

In the long run, the significance of DeepSeek V4 Pro is that an open-source model has for the first time approached the US frontier camp, which is itself a breakthrough. The trade-off between cost efficiency and capability also reflects the current reality of model optimization.