Testing the Limits: AI-Enabled Targeting, Model Evaluation and International Humanitarian Law

Testing the Limits: AI-Enabled Targeting, Model Evaluation and International Humanitarian Law

[Sarah Shoker is a Senior Non-Resident Fellow in artificial intelligence at U.C. Berkeley Risk and Security Lab and former Geopolitics Team Lead at OpenAI.

Jessica Dorsey is an Assistant Professor of International Law at Utrecht University School of Law and Managing Editor of Opinio Juris.]

What began as a quiet contract dispute between the U.S. Department of Defense and the AI firm Anthropic in February 2026 has become a high-stakes confrontation over the invisible lines companies draw around AI in warfare. Central to this is deciding who determines the legal, ethical, and safety limits for machines that shape conflict. The “red lines” promoted by leading AI companies may seem reassuring, but they set a remarkably low bar. By focusing on extreme cases, like fully autonomous weapons without human involvement, they divert attention from the systems they are helping build today: AI-enabled decision-support (AI-DSS) tools that help select, identify, prioritize, and validate targets.

These tools, developed by companies like Google, xAI, Palantir, OpenAI, and Anthropic, are transforming modern warfare. AI-DSS compress decision time, strain oversight, and are deployed without transparent testing, complicating compliance with international law. The U.S. credits these systems with generating roughly 1,000 targets per day in Iran, a record-breaking pace that yielded more strikes in 100 hours than the first six months of the conflict against the Islamic State. This tempo is reminiscent of the “shock and awe” campaign in Iraq in 2003, which led to a reported average of 300 civilian deaths a day. 

International humanitarian law (IHL) requires parties to a conflict to adhere to the principles and rules around distinction, precaution, and proportionality. If technology undermines compliance with these principles, then that technology may be legally restricted, or in some cases, prohibited from use in armed conflict. The burden should not be on the public to trust companies’ assurances; it should be on states to demonstrate, in a way that is explainable and verifiable, that their use of these tools complies with international law, with companies providing the transparency necessary for states to make that determination.

Mass Precision, Compressed Judgment

We are now in an era of “mass precision,” in which targeting happens faster and at greater scale under the promise of accuracy. However, the result is not necessarily less destruction, but differently organized destruction: faster, more distributed, and  potentially harder to scrutinize. Shortened decision cycles shrink opportunities for deliberation, large target volumes jeopardize the possibility for careful review, and the authoritative impression AI outputs give can encourage human deference.

These risks exist even without fully autonomous weapons and are evident in AI-DSS used on today’s battlefields, which only nominally preserve human control. Reports from the recent Israel-Hamas conflict show AI-DSS generating or processing vast numbers of targets under extreme time pressure. Some IDF soldiers were reported to have only 20 seconds to validate a target from a list of 37,000. In Gaza, over 80 percent of the buildings were either damaged or destroyed, even as many strikes were individually described as “precise.” The U.S. strikes on Iran similarly demonstrate the problem at speed and scale: more than 13,000 targets were struck in the first 38 days, with AI-DSS helping to generate and prioritize targets. At such a tempo, the practical capacity to independently verify targets, assess the surrounding circumstances, and take feasible precautions risks being overwhelmed by the very speed and scale that these systems are designed to enable.

The speed and scale of AI-DSS can erode the conditions for sound legal and ethical judgment. Beyond speed, these systems reshape decision-making itself: by converting complex, context-dependent assessments into numerical thresholds and predictions, AI-DSS use can narrow human judgment, steering commanders toward machine-driven choices and undermining the discretion that makes command an operational art rather than a mathematical science. The focus therefore needs to shift to assessing whether systems augmenting humans in warfare are fit for purpose under the conditions in which they are actually being used.

The Testing Gap

Large language models (LLMs) like Claude, Grok and ChatGPT are relative latecomers when compared to AI like machine vision, which the Department of Defense has deployed in the targeting cycle since 2021 and in Iraq, Syria and Yemen for target identification. In Iran, LLMs were used for target identification and prioritization. Yet despite being only one element in a larger AI ecosystem, it is worth dwelling on LLMs because the DoD’s AI Strategy calls special attention to “frontier models” for battlefield use and because the pace of their adoption has dramatically outpaced the ability to test these models for real-world risks.

The current ecosystem of LLM safety testing is both nascent and misaligned with real-world risks. Safety testing for LLMs is largely self-regulated; the frontier labs set most of their own safety standards, and domains related to international law and military decision-support are not part of the suite of tests that dictate whether a model is sufficiently safe to sell to consumers or the government. The U.S. government’s recent use of export control laws against Anthropic’s Fable appears to sit uneasily with its own executive order emphasizing voluntary compliance by companies. Meanwhile, bills such as Illinois’ SB315 require companies to evaluate broadly defined “catastrophic risks,” but these frameworks have historically excluded applications such as AI-DSS while prioritizing risks such as cyber misuse and CBRN. These legislative efforts build on voluntary commitments advanced by the frontier labs themselves under the Biden Administration. This raises a broader question: who is left out of the safety agenda when legislation closely tracks the risks and priorities already identified by the companies developing these systems?

Companies primarily evaluate their models to benchmark against competitors or to demonstrate capabilities in controlled, static environments. These tests tell us little about how systems behave in time-constrained, high-pressure, hierarchical settings like military operations. They do not adequately assess how humans interact with AI under stress, nor do they capture well-documented risks such as over-reliance, cognitive biases or anchoring, or the gradual deskilling of human operators.

Indeed, many LLMs are optimized for user engagement, which can encourage deference to AI outputs instead of critical scrutiny. Reinforcement learning processes depend on human feedback that can vary widely depending on annotators’ backgrounds, introducing subtle biases. Models remain brittle, vulnerable to manipulation and capable of producing confident but incorrect outputs. Yet these limitations are rarely tested where they have the highest consequences.

In military targeting, the relevant object of evaluation and interrogation is the sociotechnical system: the model, the surrounding data and interfaces, the human operators, the command structure, and the operational tempo in which outputs become decisions. A model that performs acceptably in a static benchmark may behave very differently when placed inside a high-pressure targeting workflow in which thousands of outputs must be processed and humans are expected to exercise judgment in seconds. Without context-specific testing, states cannot meaningfully establish whether an AI-DSS can reliably support the human judgment, verification, and precautions that IHL requires.

The Missing Safeguards

At the same time, government capacity to evaluate and oversee AI is shrinking. Defense evaluation increasingly adopts a “move fast and break things” Silicon Valley approach, sidelining rigorous legal oversight in high-stakes military contexts. Defense Secretary Pete Hegseth has characterized safety and legal practices as “operational risks…not simply bureaucratic inconveniences.” Amid the U.S. war on Iran, he has ordered a second “ruthless review” of the JAG Corps, while programs aimed at civilian protection, including the Civilian Harm Mitigation and Response (CHMR) program and the Civilian Protection Center of Excellence (CPCoE), have been dramatically weakened. The DoD’s own latest Congressionally mandated annual report on the CPCoE records an approximately two-thirds reduction in the Center’s staff during 2025, from roughly 40 people in January 2025 to just nine by the end of the year, alongside reductions, reorganizations, and terminations across the wider CHMR community.

In this report, DoD claims that CPCoE developed a “civilian environment layer”  for Maven Smart System and an AI assistant called “AskSage,” which reportedly reduced the time needed to validate civilian-environment information by 97 percent during an Army exercise. These claims have not been independently verified and any such verification would have to be highly context-dependent, given how difficult it is to assess whether the system performs reliably across different operating environments. But saving time is not equivalent to improving quality or accuracy in terms of civilian-harm mitigation, and the DoD emphasis on speed raises the concern that AI for civilian harm mitigation may be being pulled into the same time-compression logic driving the targeting cycle. Ultimately, if these tools really help protect civilians, why were they not already being used, and what evidence is there that they improve outcomes for civilians rather than simply speed up the targeting process?

For a technologically advanced military, this discussion has legal implications. Where technology makes additional precautions feasible, the capacity to take those precautions can raise, rather than lower, the standard against which conduct must be assessed. Operational choices to accelerate target generation and engagement without first developing a sufficiently full understanding of the civilian environment introduces more fundamental questions about whether, where, and how precautions were taken at all.

Comparing Precision Claims to Reality

Those who support AI-enabled targeting argue that these tools improve the protection of civilians through enhanced intelligence accuracy and enabling more precise strikes. However, reality seems to tell a very different story. In June, Airwars released a comprehensive overview of the human cost of the first 40 days of the war, documenting nearly 1,700 civilian deaths in Iran and widespread harm throughout the region. Airwars’ subsequent investigation of the DoD report further illustrates the gap between claims of precision and the civilian harm that is actually acknowledged and accounted for. Public evidence that AI systems, including LLM integration, systematically reduce civilian casualties is poor and continues to be contested. Without independent testing, transparent reporting, and evidence that civilian-environment data is actually incorporated into targeting decisions, claims of AI-enabled “precision” remain impossible to verify.

The debate over “red lines” risks focusing attention on hypothetical future scenarios while ignoring the realities of current warfare. AI companies are not peripheral actors; they are central participants in a broader private-sector ecosystem that is transforming how wars are fought. By articulating narrow prohibitions, they can signal responsibility and reap reputational benefits, even as their tools are integrated into workflows or systems that may strain or erode longstanding legal protections for civilians.

For IHL to remain relevant and robust in this new context, it has to be considered not only in the context of weapons themselves but also in the decision architectures guiding their use. That requires a lifecycle approach to military AI governance: from design through to rigorous, context-specific testing of AI systems before deployment; independent oversight capable of auditing both technical performance and human-machine interaction; and transparency sufficient to allow public scrutiny.

Most importantly, it requires a shift in the placement of the burden: away from asking whether there is evidence of harm after the fact, and toward requiring proof, prior to deployment, that these systems can operate within the constraints the law imposes. The legitimacy of military AI ultimately turns on whether the systems in which these technologies are embedded are fit for the purposes for which they are used, and whether states can demonstrate that they remain capable of supporting appropriate levels of human judgment and compliance with international law. 

Print Friendly, PDF & Email
Topics
Artificial Intelligence, Autonomous Weapons, Featured, General, International Human Rights Law, International Humanitarian Law, International Law, Middle East, National Security Law, North America, Public International Law, Technology, Use of Force

Leave a Reply

Please Login to comment
avatar
  Subscribe  
Notify of