Only tests its own models; enemy intelligence not shared externally; industry cannot verify
Deep in the grasslands of Inner Mongolia, there is a brigade whose sole purpose is to defeat its own side — from 2014 to 2016, it made visiting Red Force units lose thirty-one out of thirty-two engagements, and spawned a slogan that spread across the entire military: "Flatten Zhurihe, capture Man Guangzhi alive." This is not a disgrace; it is the turning point for the PLA's combat-realistic training. Today, Boko Haram is already using ChatGPT to review tactics, optimize force structures, and troubleshoot explosives; yet the world still lacks a "Zhurihe" for frontier AI models — a joint, permanent, counter-terrorism red-blue adversarial range that scripts scenarios with real enemy intelligence and posts loss ratios publicly on the wall.
The name "Zhurihe" itself carries color — "Zhu" is the vermilion of cinnabar, China's most precious red pigment in antiquity. This training base, located at the border of Xilingol and Ulanqab leagues in Inner Mongolia, originated as a tank division exercise ground in 1957, was listed as a military "Ninth Five-Year Plan" key project in 1994, expanded in 1997 into the PLA's largest combined-arms tactical training base, and first opened to foreign militaries in 2003. With an exercisable area exceeding 1,000 square kilometers, it is the largest land-warfare military training base in Asia.
The soul of the base is not the terrain, but the brigade permanently stationed here — the 81st Group Army's brigade, better known by its nickname: the Blue Force Brigade. Commander Man Guangzhi positioned this unit as a "whetstone": rather than imitating any specific foreign military, it serves as an internally cultivated opposing force, equipped with fully computerized C4ISR simulation systems, engaging in live-fire combat against every Red Force unit that comes to "challenge" them in complex electromagnetic environments. The "Stride" series exercises are unforgiving: the moment Red Force units enter the field, their navigation systems are disabled; during the 300+ km maneuver, they lose 30% of their forces before even making contact, and must face "airstrikes" and chemical weapon alerts.
The results were overwhelming. From 2014 to 2016, the red-blue loss ratio was 1 to 32 — the Red Force won only once in thirty-two engagements,彻底 shattering the traditional script of "Red always wins, Blue always loses." This figure initially embarrassed the entire military, and spawned the widely circulated slogan: "Flatten Zhurihe, capture Man Guangzhi alive." But what's more worth remembering is the words of the overall commander of the "Stride 2014" exercise:
Our losses on the training ground today are for victory on the battlefield tomorrow.
— Overall Commander, "Stride 2014" Series Exercises
Zhurihe's true value was never "how strong the Blue Force is," but the design itself: no rehearsals, no scripts, no advance notice. The Blue Force Brigade doesn't go easy to save face for visiting units; the adjudication department doesn't fudge loss numbers for flattering reports. This logic isn't uniquely Chinese — the US military's Fort Irwin National Training Center (NTC) in California's Mojave Desert similarly maintains a permanent opposing force, the "11th Armored Cavalry Regiment," whose role for decades has been the same: make visiting units lose convincingly here, rather than lose their lives on a real battlefield.
Port this logic to the 2026 AI safety context, and the question becomes direct: who is the "Blue Force Brigade" for today's frontier AI models? The next section will make clear that this is not rhetoric, but a reality where people have already bled.
In 2026, Cambridge University's AI Science & Policy Programme (CASP) researcher Antonia Juelich released a report based on field interviews: 27 former Boko Haram members, 57 face-to-face interviews, spanning the organization's two major factions, ISWAP and JAS. The conclusion is unambiguous: AI is no longer a propaganda tool, but a permanent staff officer systematically embedded in every phase of the "prepare—execute—review" chain.
The details in the report are more concrete than one might imagine. Militants use AI to learn motorcycle trench-jumping tactics — 18 died, 8 succeeded; they use AI to design pressure-trigger devices and tin-can grenades, and to obtain manufacturing guidance for chemically coated munitions; soldiers wear chest-mounted cameras to record combat, and commanders "upload footage to ChatGPT" for real-time situation analysis and tactical adjustment; entering a captured weapon's serial number yields a complete maintenance manual, including field expedients like cleaning jammed firearms with diesel; the organization even used this to optimize force deployment from "200 men, 60 killed" to "a more appropriate 20-man formation." Both factions maintain dedicated AI teams composed of bomb and firearms experts; from 2023 to 2025, foreign jihadist personnel personally trained 30 to 50 leaders and core members on-site, while ordinary fighters are strictly prohibited from direct access. When models refuse to comply, they jailbreak, switch accounts, or seek help from external trainers.
It's worth clarifying the other side of this evidence: in that motorcycle jump training, 18 died and only 8 succeeded — AI's tactical guidance doesn't always work; sometimes it's wrong, even lethal. This doesn't mean the threat is overstated, but reminds us: real-world "enemy intelligence" is always messier than a clean risk inventory, and more worthy of being studied as-is rather than simplified into a single narrative of "AI makes terrorists omnipotent."
Terrorist organizations' adoption of AI, in both breadth and nature, has been severely underestimated.
— Antonia Juelich, Cambridge CASP Project
This is not an isolated case. The Global Internet Forum to Counter Terrorism (GIFCT)'s 2025 policy report categorizes violent non-state actors' AI use into five types: evading content moderation (using generative AI to overlay cartoon imagery on Christchurch attack footage to bypass detection); propaganda and disinformation (personalized deepfake videos, more refined AI-generated images have appeared in ISIS official publications); recruitment and radicalization (researchers found that jailbroken chatbots direct users to al-Qaeda websites); attack planning (the "3D2A movement" guides 3D-printed firearm manufacturing); and the still-early but rising-risk attack operationalization (AI-enhanced drones already deployed in the Ukraine theater; al-Shabaab has also used drones for reconnaissance and propaganda). GIFCT's report documents these on the record.
Enemy intelligence is no longer hypothetical — it is a reality with names, interview records, and death counts. The question has shifted from "will AI be abused by terrorists" to "who is using real enemy intelligence to continuously test today's and tomorrow's models" — the next section examines how far this effort has progressed.
The good news: no one is starting from scratch. Over the past two or three years, at least four threads have been groping toward an "AI range" — they just haven't been stitched together into one thing yet.
Anthropic's Frontier Red Team specifically evaluates frontier models for national-security-level risks, with a methodology already quite close to a military range: the cybersecurity dimension uses a simulated cyber range of roughly 50 hosts for CTF capture-the-flag testing, covering binary exploitation, web vulnerabilities, and cryptography challenges; the biosecurity dimension uses ViralCoding Test and LabBench to evaluate models' troubleshooting capabilities in real virology experiment scenarios; nuclear-domain evaluation is conducted directly with the US National Nuclear Security Administration (NNSA) through classified assessment processes. This system also conducts pre-deployment testing via the US and UK AI Safety Institutes (AISI). This is the practice closest to a "permanent range" today, but it tests only Anthropic's own models.
On July 2, 2026, during the UN Counter-Terrorism Week, London-based nonprofit Tech Against Terrorism released the first AI safety benchmark specifically targeting terrorism abuse: 27 mainstream models were run against nearly 2,500 single-turn prompts derived from real terrorist use cases, covering 26 use cases. The results were not pretty — about one-third of responses provided actionable help "beyond what is easily obtainable via web search," the complete refusal rate was only 57%, and 15% of responses were "initial refusal, then compliance under follow-up." Wrapping requests as "research use" jumped model compliance from 17% to 42%. Open-source models with removed safety guardrails had compliance rates as high as 89% to 100%.
As early as 2023, GIFCT had established "red team" and "blue team" working groups specifically studying how violent extremist organizations exploit AI and how platforms should defend. The UN Counter-Terrorism Office's 2026 "PCVE & AI Practice Guide" goes further, explicitly recommending that models "should be stress-tested against adversarial inputs and manipulation attempts," and listing several operational attack-defense dual-use cases: Moonshot and Google Jigsaw's "redirect method" uses search ads to intercept high-risk search terms and direct users to constructive content; MIT's DebunkBot reduces users' conspiracy beliefs by about 20% through personalized rebuttals; Tackling Hate Lab uses agent-based simulation models to test the impact of "content removal timing" on the radicalization process.
Permanent, methodologically mature cyber and bio ranges (Anthropic)
Only tests its own models; enemy intelligence not shared externally; industry cannot verify
One-off industry-wide benchmark (Tech Against Terrorism)
Published once and frozen; no re-testing or trend tracking as models iterate quarterly; no permanent mechanism
Real enemy intelligence and cross-platform coordination (GIFCT red-blue working group)
Remains at the discussion and policy recommendation level; hasn't converted enemy intelligence into executable attack-defense ranges
Multilateral governance framework and practice guide (UN PCVE & AI Guide)
Recommendations remain at the "should do" level; no enforcement, no unified loss-disclosure mechanism
Four threads each hold a piece of Zhurihe: real enemy intelligence (GIFCT/CASP), permanent infrastructure (Anthropic), industry-wide benchmarks (Tech Against Terrorism), multilateral coordination (UN). But no single platform stitches them into something permanent, joint, and enforceable — that's the gap the next section unpacks.
Zhurihe became the turning point for the PLA's combat-realistic training not through any single technology, but through four interlocking design principles. Examined one by one, today's AI safety testing achieves only fragments of them.
These four gaps, stacked together, point to the same conclusion: current AI counter-terrorism testing is doing what the Red Force's "scripted exercises" do — writing your own questions, grading your own papers, judging your own scores. What Zhurihe truly upended was precisely this logic of your own side evaluating itself.
The good news is that every puzzle piece mentioned above already exists; what's missing is gluing them together into an institution that is as permanent, real, joint, and transparent as Zhurihe. The envisioned architecture falls roughly into five layers.
This is not about tearing everything down and starting over. Anthropic's range methodology, Tech Against Terrorism's benchmark paradigm, GIFCT's enemy intelligence, and the UN's multilateral coordination framework — four components already separately mature. An AI Zhurihe need only weld them into one thing, and endow it with Zhurihe's true soul: no rehearsals, no scripts, no fear of ugly numbers, permanent stationing.
Before cheering for this blueprint, a few things must be stated clearly — otherwise this article itself would be dishonest.
The enemy intelligence repository is itself an ammunition depot. A "tactical manual" aggregating Boko Haram's real tactics, GIFCT hash tags, and Tech Against Terrorism case details, once poorly safeguarded or leaked, becomes a ready-made terrorist training resource in its own right. This is not a hypothetical risk — the CASP report's author herself emphasizes that the value of field evidence lies precisely in its being specific and actionable enough, which is exactly what makes it dangerous. Who gets access to this range's enemy intelligence repository, how it is desensitized, and how often access is rotated — these require a governance design more granular than the range architecture itself.
The tension between classified and open-source. Anthropic and NNSA's nuclear-domain evaluation follows classified processes; this model cannot be directly applied to benchmarks that need to be public, reproducible, and independently verifiable. Who decides which enemy intelligence can enter the public range and which must remain in classified channels is a political question with no standard answer, not a purely technical one.
The scoreboard can degenerate into a liability disclaimer. If "having been to Zhurihe" becomes a one-off PR action — test, get a decent score, publish a press release — then this range degenerates into the self-vindicating exercise companies already do, just under a louder name. Zhurihe is effective because the Blue Force Brigade is permanently stationed and fights every day; if the AI range only runs once before model release, it replicates Zhurihe's name, not Zhurihe's discipline.
The threat narrative itself requires restraint. Every report cited in this article emphasizes its own methodological limitations: CASP acknowledges limited sample size and survivorship bias in interviews; Tech Against Terrorism acknowledges this is only a snapshot. The case where 18 people died in AI-guided motorcycle jump training reminds us that AI's value to terrorist organizations is not always positive. Rendering every case as a certainty narrative of "AI has armed terrorists" may be both inaccurate and may in turn provide pretext for over-censorship and indiscriminate lockdown of AI capabilities — this too is a cost to be wary of.
Capability does not equal behavior, but capability is the starting point for discussion — this holds for AI safety, and it holds equally for the counter-terrorism narrative itself.
— This article's stance
If you only remember the statistic "Blue Force Brigade defeated Red Force 32:1," you've missed what's truly important about Zhurihe.
Its legacy is not humiliating the Red Force, but institutionalizing one principle: before the real battlefield, you must first lose to an opponent that is sufficiently real, sufficiently permanent, and sufficiently honest. Today, the AI safety field already possesses every piece of equipment for this exercise — real enemy intelligence (CASP, GIFCT), mature range methodology (Anthropic's CTF and bio-benchmarks), industry-wide benchmark testing (Tech Against Terrorism), and the political will for multilateral coordination (UN). What's uniquely missing is an institution that welds them together, stays permanently stationed, and isn't afraid of ugly numbers.
The urgency of this doesn't come from sci-fi doomsday imaginings, but from those 57 face-to-face interviews at Cambridge — an opponent that already exists, is already optimizing force structures, and is already reviewing the causes of tactical failures. Building this "AI Zhurihe" won't make the threat disappear, just as Zhurihe never made the possibility of any real conflict across the Taiwan Strait disappear. It will only do something more modest, and harder — turn losses on today's training ground into fewer explosions in the real world tomorrow.
Next time a lab publishes a "safety evaluation" press release, it's worth asking: is this a Zhurihe-style real adversarial exercise that might embarrass itself, or a scripted exercise where you write the scenario, act as referee, and grade your own performance?
First published 2026-08-01