Research · Academic & arXiv

Back to sweep

Research sweep · deep · 2025 – 2026

Security Research into Chinese Open-Weight Models

Independent security research into Chinese open-weight models (DeepSeek R1 and V3, Alibaba Qwen, Moonshot Kimi K2, Zhipu GLM, MiniMax, Baidu Ernie) from February 2025 to August 2026: who is testing them, what red-teaming and provenance methods they use, which results survive independent replication, and how regulated, defence and military buyers are assuring models whose training data and training objectives are never disclosed

  • GPT-5.6-sol
  • financial
  • frontier
  • academic
  • blogs
  • tech

Synthesised 2026-08-03

Narrative

Independent technical evaluation of Chinese open-weight models has followed two distinct tracks since February 2025: capability/safety benchmarking by evaluation labs and government bodies, and a separate theoretical literature on whether backdoors in released weights can be detected at all. On capability, METR ran its HCAST and RE-Bench task suites against DeepSeek-V3, DeepSeek-R1 and later a joint DeepSeek/Qwen batch, each time using the same Triframe scaffold applied to Western frontier models. METR's DeepSeek-V3 report found the model's performance worse than o1 and Claude 3.5 Sonnet (New) and significantly better than GPT-4o, meaning it isn't introducing a new level of dual-use capabilities


Sources

ID Title Outlet Date Significance
a1 DeepSeek-R1 Red Teaming Report: Alarming Security and Ethical Risks Uncovered – Unite.AI unite.ai February 1, 2025 Retrieved by this lane's web search.
a2 DeepSeek R1 Security Report - AI Red Teaming Results | Promptfoo promptfoo.dev March 13, 2025 Retrieved by this lane's web search.
a3 Deepseek R1 Red Teaming Report (pdf) - CliffsNotes cliffsnotes.com July 13, 2025 Retrieved by this lane's web search.
a4 A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models arxiv.org Retrieved by this lane's web search.
a5 GitHub - chen37058/Red-Team-Arxiv-Paper-Update: Awesome Jailbreak, red teaming arxiv papers (Automatically Update Every 12th hours) · GitHub github.com 1 month ago Retrieved by this lane's web search.
a6 Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model arxiv.org Retrieved by this lane's web search.
a7 DeepSh*t: Exposing the Security Risks of DeepSeek-R1 hiddenlayer.com January 30, 2025 Retrieved by this lane's web search.
a8 DeepSeek R1 Red Teaming & Jailbreaking Audit holisticai.com Retrieved by this lane's web search.
a9 Exploiting DeepSeek-R1: Breaking Down Chain of Thought Security | Trend Micro (US) trendmicro.com March 4, 2025 Retrieved by this lane's web search.
a10 Hidden in Memory: Sleeper Memory Poisoning in LLM Agents arxiv.org Retrieved by this lane's web search.
a11 AutoBackdoor: Automating Backdoor Attacks via LLM Agents arxiv.org Retrieved by this lane's web search.
a12 Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review arxiv.org Retrieved by this lane's web search.
a13 [2603.03371v1] Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs arxiv.org March 2, 2026 Retrieved by this lane's web search.
a14 Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness arxiv.org Retrieved by this lane's web search.
a15 [2511.15992] Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis arxiv.org November 20, 2025 Retrieved by this lane's web search.
a16 [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training arxiv.org January 17, 2024 Retrieved by this lane's web search.
a17 Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis arxiv.org Retrieved by this lane's web search.
a18 Details about METR's preliminary evaluation of DeepSeek and Qwen models metr.org June 27, 2025 Retrieved by this lane's web search.
a19 Details about METR's preliminary evaluation of DeepSeek- ... metr.org February 12, 2025 Retrieved by this lane's web search.
a20 VeriTrace: Evolving Mental Models for Deep Research Agents arxiv.org Retrieved by this lane's web search.
a21 Qwen3-Coder-Next Technical Report arxiv.org February 28, 2026 Retrieved by this lane's web search.
a22 DeepSeek-V3 Technical Report arxiv.org Retrieved by this lane's web search.
a23 CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks content.govdelivery.com Retrieved by this lane's web search.
a24 Center for AI Standards and Innovation issued interim findings from technical evaluation of DeepSeek AI models digitalpolicyalert.org September 30, 2025 Retrieved by this lane's web search.
a25 AIBoMGen: Generating an AI Bill of Materials for Secure, Transparent, and Compliant Model Training arxiv.org January 9, 2026 Retrieved by this lane's web search.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.