美国商务部测评Kimi K3称美国领先,但报告显示测试并非完全对等
Odaily Planet Daily News: The AI Standards and Innovation Center under the U.S. Department of Commerce, in collaboration with the UK AI Safety Institute, tested the cyberattack capabilities of Kimi K3 and emphasized that "the United States still leads."
However, the value of this assessment is disputed due to the limitations of the testing scope. Constrained by the hosting environment, Kimi K3 only participated in partial testing, with its overall cyber capabilities estimated primarily based on a benchmark of 41 exploit scenarios, while other models underwent more comprehensive testing. Consequently, the results for Kimi K3 have a larger margin of error.
In the vulnerability exploit test, Kimi K3 scored approximately 32%, higher than GLM-5.2's 24%, but lower than the average level of approximately 76% achieved by leading U.S. models. In the simulated attack chain test, Kimi K3 completed an average of 17 out of 32 attack chain steps and successfully breached the network once in 10 attempts, while the U.S. frontier models completed an average of 28.5 steps.
The report indicates that Kimi K3 already possesses a certain level of autonomous attack capability, and its safety guardrails did not prevent the model from developing exploits or executing attacks. However, the report also emphasizes the limited scope of the testing.
