Research · Academic & arXiv
Back to sweepResearch sweep · deep · 2025 – 2026
Security Research into Chinese Open-Weight Models
Independent security research into Chinese open-weight models (DeepSeek R1 and V3, Alibaba Qwen, Moonshot Kimi K2, Zhipu GLM, MiniMax, Baidu Ernie) from February 2025 to August 2026: who is testing them, what red-teaming and provenance methods they use, which results survive independent replication, and how regulated, defence and military buyers are assuring models whose training data and training objectives are never disclosed
- GPT-5.6-sol
- financial
- frontier
- academic
- blogs
- tech
Synthesised 2026-08-03
Narrative
Independent technical evaluation of Chinese open-weight models has followed two distinct tracks since February 2025: capability/safety benchmarking by evaluation labs and government bodies, and a separate theoretical literature on whether backdoors in released weights can be detected at all. On capability, METR ran its HCAST and RE-Bench task suites against DeepSeek-V3, DeepSeek-R1 and later a joint DeepSeek/Qwen batch, each time using the same Triframe scaffold applied to Western frontier models. METR's DeepSeek-V3 report found the model's performance worse than o1 and Claude 3.5 Sonnet (New) and significantly better than GPT-4o, meaning it isn't introducing a new level of dual-use capabilities