Bỏ qua đến nội dung chính
Back to home
AI Tech 2 min read

Kimi K3 Evaluation: US Government and LiveBench Skeptical of Chinese AI Capabilities

Industry observers express skepticism over China's Kimi K3 model after both the US government and LiveBench AI ranked it below leading US frontier models.

Tier 2 · sources 56% confidence Reviewed
Sources x.com

On July 23, 2026, debates surrounding the true capabilities of the Chinese large language model Kimi K3 intensified as the US government unexpectedly released comparative performance evaluations. Reports from the US administration indicated that Kimi K3 performs significantly worse than current frontier AI models developed by US technology companies.

Background & Causes

The US government's involvement in evaluating foreign large language models (LLMs) marks a major shift in the global technology race. According to tech expert Bindu Reddy on the social media platform X, the US government is now acting as a new benchmark authority for AI models. This move comes at a time when nations are tightening their oversight and evaluating the technological capabilities of their competitors. The US government's assessment is not isolated, as LiveBench AI, an independent evaluation platform, reached a similar conclusion regarding the Kimi model's capabilities just a few days prior. This has raised major doubts about the claims made by Chinese developers regarding catching up with Western technology.

Technical & Technological Analysis

From a technical perspective, Kimi K3 is known for being optimized for long-running tasks. However, analysts argue that its architecture and actual performance do not reach the level of "frontier intelligence." Instead, Kimi K3 is positioned on par with cheaper models such as Claude Sonnet or older Opus versions (around the 4.6 class). While optimizing operational costs for long tasks gives Kimi K3 a commercial advantage, it compromises on complex reasoning capabilities, which remain the strength of top US models. Benchmark tests from LiveBench AI highlighted these vulnerabilities when faced with intensive logic evaluations.

Expert Opinions & Assessments

Bindu Reddy pointed out that calling Kimi K3 a frontier AI model is inaccurate, classifying it instead as a low-cost alternative for long-running tasks. "Kimi is a very cheap sonnet / opus 4.6 class model for long running tasks - that’s not frontier intelligence," Reddy shared. The tech community also expressed skepticism over a government agency directly taking on the role of benchmarking commercial AI models. Many argue that independent benchmarks like LiveBench AI continue to provide more objective and reliable results.

Impact & Future

This development shows that the tech rivalry between the US and China is shifting from hardware restrictions to software quality control and AI benchmarking. For developers and users, understanding the true classification of models like Kimi K3 will help in making cost-effective decisions rather than falling for marketing hype. The trend of state regulators actively participating in technology benchmarking is expected to continue, likely leading to stricter compliance standards in the near future.