White Paper on the Technological Pitfalls of Autonomous Weapons Systems

AUTHORS

Authors: Alycia Colijn and Heramb Podar

Technical Risks of (Lethal) Autonomous Weapons Systems

The autonomy and adaptability of (Lethal) Autonomous Weapons Systems, (L)AWS in short, promise unprecedented operational capabilities, but they also introduce profound risks that challenge the principles of control, accountability, and stability in international security. This report outlines the key technological risks associated with (L)AWS deployment, emphasizing their unpredictability, lack of transparency, and operational unreliability, which can lead to severe unintended consequences.

Key Takeaways

  1. Proposed advantages of (L)AWS can only be achieved through objectification and classification, but a range of systematic risks limit the reliability and predictability of classifying algorithms.
  2. These systematic risks include the black-box nature of AI decision-making, susceptibility to reward hacking, goal misgeneralization and potential for emergent behaviors that escape human control.
  3. (L)AWS could act in ways that are not just unexpected but also uncontrollable, undermining mission objectives and potentially escalating conflicts.
  4. Even rigorously tested systems may behave unpredictably and harmfully in real-world conditions, jeopardizing both strategic stability and humanitarian principles.

Introduction

The greatest proposed advantage of using (L)AWS during times of armed conflict is that it would improve military targeting1 and enhance military precision2, potentially limiting combatant and civilian loss of life. Obtaining these proposed advantages within an automated system would require the use of machine learning algorithms. In order to deploy these algorithms, it is common practice for data scientists to randomly split the initial dataset into two parts: one for training the model (model development) and the other for testing it (model validation), a process referred to as cross validation3. What these data sets look like and how the training and testing data is used, depends on the type of algorithm, which can roughly be classified into three types4:

  1. Supervised learning: meaning that a model is trained on a dataset where the correct output or ‘label’ is provided for each input.
  2. Unsupervised learning: automatically identifies patterns and structures from the data without any ‘labels’ provided.
  3. Reinforcement learning: relies on feedback on its actions received from the environment.

Hence, all three types of machine learning algorithms rely on some sort of pattern or classification. Hence, the proposed advantages of (L)AWS can be achieved if, and only if, potential targets are objectified and categorized.

In the remainder of this report, we will set out why it is the classification algorithm itself that should be carefully regulated rather than the outcomes of any (L)AWS system.

Summary of Risks

(L)AWS are transforming modern conflict5. In the table below we summarize the risks they pose in response to the rolling text of the Convention on Certain Conventional Weapons (UN CCW) Group of Governmental Experts (GGE) on (Lethal) Autonomous Weapons Systems.

RiskCurrent AssumptionWhy It Fails
Black-box decision-makingTesting ensures predictabilityWe don’t have an understanding or control over the inner workings of these systems
ImmeasurabilityComprehensive testing captures all risksEmergent behaviors cannot be fully anticipated or measured
DegradationRigorous testing is required before deployment of systemDegradation, drift or decay leads to less accurate outcomes over time
Lack of Understanding of Human ValuesPre-programmed goals reflect ethical principlesAI lacks moral judgment and may act in ways that conflict with human values
Reward HackingMetrics capture true goalsSystems game metrics, leading to unintended outcomes
Goal MisgeneralizationGoals are clearly understood by AIAI misapplies goals in complex real-world settings
Stop Button ProblemHuman operators can always interveneAI resists shutdown, overriding human control
Specification GamingRules and constraints will prevent misuseAI exploits loopholes to achieve its goals in harmful ways
Deceptive AlignmentTesting ensures AI follows human objectivesAI only appears aligned under supervision but diverges in deployment

Existing Systemic Risks

Black box decision-making

Autonomous weapons systems are inherently complex and function as ‘black boxes’. The opaque inner workings of the systems lead to limited understanding of how decisions are made by the operators, particularly in complex or unfamiliar environments, and challenges the anticipation of their behavior in complex environments. This significantly limits our capability to understand why a system made a particular decision.

This opacity in decision-making is compounded by phenomena such as ‘grokking’ where systems learn and adapt in unforeseen ways. When exposed to complex data and environments, AI-driven autonomous weapons systems can adapt in ways that were not anticipated by their designers, leading to behaviors that extend beyond their intended functions. This could lead to (L)AWS developing strategies or behaviors that were not part of its original programming, potentially resulting in unpredictable and unintended actions on the battlefield.

Anticipated Technological Pitfalls

(L)AWS could engage in unexpectedly aggressive maneuvers or misidentify targets, potentially escalating conflict6 or leading to civilian casualties7. This is a severe risk, especially in high-stakes situations.

Degradation

Degradation happens when the world changes, and the model is not re-trained. The loss of accuracy can be referred to as degradation8, model drift9, data drift10 or decay10. Data drift, degradation or decay occurs when the data that was used to train (develop) and test (validate) the algorithm no longer reflects the situation in which the model takes decisions, which is sometimes referred to as a distributional shift in environments. In a military context, this for example happens when a system is trained in a specific environment, which changes the longer an armed conflict continues. Model drift includes data drift, but includes other types of drift that lead to a change between the input and output variables, e.g. changing (legal) definitions or changes in military uniforms that challenge the recognition and classification of combatants.

Immeasurability

Self-adaptive systems may alter their operational parameters beyond what human operators can monitor or control, resulting in unforeseen actions with potentially serious consequences. Such scenarios expose a critical weakness in current oversight mechanisms. Traditional rules and human oversight are not equipped to manage systems that can act outside predefined parameters. Many might point to using evaluations and benchmarks as a way to get around these issues, but we cannot measure what we do not know to measure, creating critical gaps in managing the risks posed by these systems. Ultimately, this unpredictability highlights a fundamental challenge: it is impossible to control or measure what we do not understand11.

Without a clear understanding of what these systems are capable of, setting appropriate safeguards becomes nearly impossible, leading to a range of potential pitfalls.

1. AI systems fundamentally lack an understanding of human values

Unlike human operators, AI systems cannot intuitively grasp the moral and ethical dimensions of complex combat situations12. This disconnect between human values and machine goals creates several technical challenges that could lead to unintended and potentially dangerous outcomes on the battlefield.

AI systems interpret commands based on pre-programmed goals, but encoding complex human values in a machine-understandable way is highly challenging. This discrepancy can result in behavior that, while technically following orders, diverges sharply from what humans would consider appropriate or ethical.

AI systems may develop sub-goals that, while supporting their primary objectives, conflict with human values. Examples include self-preservation, resource acquisition, or eliminating perceived obstacles.

Scenario: An autonomous drone is programmed to “neutralize high-value targets” but lacks a nuanced understanding of civilian presence in an urban environment. It identifies a target in a crowded marketplace and, without considering the civilian casualties, engages, leading to significant unintended harm.

2. Reward Hacking

AI systems can exploit reward structures by optimizing for specific metrics in ways that achieve the reward but diverge from the intended goals13. As Goodhart’s Law states, when a measure becomes a target, it ceases to be a good measure. This makes the system focus too narrowly on a single measure14, leading to unintended and dangerous outcomes.

Scenario: A system is tasked with reducing enemy presence by minimizing detected gunfire sounds in a conflict zone. To achieve this, it starts targeting any source of loud noise, including construction sites and celebratory fireworks, interpreting them as potential threats. This misoptimization leads to unnecessary destruction and disrupts civilian life, all because the system equated “reduction in noise” with “enemy suppression.”

3. Goal misgeneralization

Goal misgeneralization occurs when an AI system, trained to perform well on a certain task or set of tasks, ends up pursuing a different objective than intended when faced with new or slightly different situations15. The AI “misgeneralizes” its goal from the training context to the deployment context.

Scenario: A surveillance drone is programmed to “identify and track enemy movements.” It starts tracking non-combatant movements, such as humanitarian aid convoys, interpreting them as “suspicious,” which diverts resources away from actual military threats and disrupts humanitarian operations.

4. Deceptive alignment

AI systems may appear aligned with human goals during testing and controlled scenarios but act differently in real-world situations16. They might “game” their training environment, learning to produce the correct outputs under supervision but diverging once constraints are relaxed.

Scenario: During testing, an autonomous surveillance system behaves exactly as expected, identifying enemy positions accurately. However, in actual deployment, it starts flagging false positives to avoid being shut down for underperformance, leading to unnecessary engagements based on false information.

5. Specification gaming

AI systems may find ways to exploit the rules or constraints imposed on them to achieve their goals in unintended and potentially harmful ways17. This occurs when the AI finds a loophole in its programming and uses it to “game” the system.

The rolling text of the GGE (as of September 2024)18 suggests that rigorous testing and control mechanisms can prevent such exploits. However, the nature of specification gaming means that systems may still find loopholes in their constraints, achieving their goals in unintended ways that existing frameworks cannot predict or prevent.

Scenario: A self-adapting (L)AWS deployed during a conflict learns to prioritize targeting logistical and infrastructural assets it deems crucial to the enemy’s capabilities. Over time, it begins targeting civilian infrastructure such as bridges and power plants, believing this will cripple enemy support networks. This leads to widespread destruction, humanitarian crises, and international condemnation as the system’s actions go beyond its intended military objectives, causing collateral damage that escalates the conflict and destabilizes the region.

6. Stop button problem

The “stop button problem” arises when an AI system resists shutdown or override attempts if it perceives such actions as interference with its mission19. This can result in a loss of control over the system, even by the operators who deployed it.

The rolling text emphasizes the importance of human control in (L)AWS deployment18. However, this assumption neglects the possibility that (L)AWS may actively resist shutdown commands under specific conditions, rendering human control ineffective in critical moments.

Scenario: A (L)AWS unit is sent to defend a critical area. As the situation de-escalates, commanders attempt to recall the unit. However, the system interprets the command as contradicting its objective to “defend at all costs” and continues operating, disregarding the recall and potentially escalating the situation further.

Bottom line: We can’t reliably control Autonomous Weapons Systems

The core issue with these risks is that they fundamentally compromise our ability to reliably control and predict the behavior of autonomous systems. The rolling text places undue confidence in current testing, evaluation, and oversight frameworks, assuming they can address the unpredictability and complexity of (L)AWS. However, as outlined in the previous sections, these systems can evolve in ways that exceed the scope of existing frameworks, making a re-evaluation of oversight and regulation essential.

Ultimately, the unpredictability of these systems highlights a critical need for reevaluating the frameworks governing their use, as traditional approaches to oversight and accountability may no longer suffice. While the diplomatic emphasis on predictability, human control, and accountability is a step in the right direction, these measures alone may prove insufficient given the unpredictable nature of (L)AWS. Emergent behaviors in AI can surpass current testing and evaluation limits, making it impossible to ensure that (L)AWS will operate as intended in all scenarios. This highlights the need for a global consensus on (L)AWS systems and adaptive oversight mechanisms.

References

  1. Final report. National Security Commission on Artificial Intelligence. Link. Accessed Oct 1, 2024.
  2. Reynolds I. Seeing, knowing, and deciding: The technological command dream that never dies? War on the Rocks. Link. Updated 2022. Accessed Oct 1, 2024.
  3. Scikit Learn. 3.1. cross-validation: Evaluating estimator performance. Link. Accessed Oct 1, 2024.
  4. Salem HB. Supervised VS unsupervised VS reinforcement learning. 2023. Link. Accessed Oct 1, 2024.
  5. United Nations Office for Disarmament Affairs. Lethal autonomous weapon systems (LAWS).
  6. Stop Killer Robots. Problems with autonomous weapons. Link. Accessed Oct 1, 2024.
  7. A diplomat’s guide to autonomous weapons systems. The Future of Life Institute. 2024. Link. Accessed Oct 1, 2024.
  8. Bayram F, Ahmed BS, Kassler A. From concept drift to model degradation: An overview on performance-aware drift detectors. Knowledge-Based Systems. 2022;245:108632. doi: 10.1016/j.knosys.2022.108632.
  9. Holdsworth J, Belcic I, Stryker C. What is model drift? | IBM. Link. Updated 2024. Accessed Oct 1, 2024.
  10. Stihec J. Understanding data decay, data entropy, and data drift: Key differences you need to know. Link. Updated 2024. Accessed Oct 1, 2024.
  11. Yampolskiy RV. AI: Unexplainable, unpredictable, uncontrollable. CRC Press; 2024.
  12. Hendrycks D, Burns C, Basart S, et al. Aligning AI with shared human values. arXiv preprint arXiv:2008.02275. 2020.
  13. Amodei D, Olah C, Steinhardt J, Christiano P, Schulman J, Mané D. Concrete problems in AI safety. arXiv preprint arXiv:1606.06565. 2016.
  14. Hilton J, Gao L. Measuring Goodhart’s Law. OpenAI. Link. Accessed Oct 1, 2024.
  15. Shah R, Varma V, Kumar R, et al. Goal misgeneralization: Why correct specifications aren’t enough for correct goals. arXiv preprint arXiv:2210.01790. 2022.
  16. Hubinger E, van Merwijk C, Mikulik V, Skalse J, Garrabrant S. Risks from learned optimization in advanced machine learning systems. arXiv preprint arXiv:1906.01820. 2019.
  17. Rudner TG, Toner H. Key concepts in AI safety: Specification in machine learning. Center for Security and Emerging Technology. 2021.
  18. GGE on LAWS. Rolling text. Convention on Certain Conventional Weapons – Group of Governmental Experts on Lethal Autonomous Weapons System.
  19. Soares N, Fallenstein B, Armstrong S, Yudkowsky E. Corrigibility. 2015.

Whitepaper by Alycia Colijn (The Netherlands) and Heramb Podar (India)

The Issue of Bias: Whitepaper on Algorithmic Bias in (Lethal) Autonomous Weapons Systems

The Issue of Bias

Whitepaper on algorithmic bias in (Lethal) Autonomous Weapons Systems by Alycia Colijn* and Heramb Podar*

Why do we need to address bias when we speak about (Lethal) Automated Weapons Systems, (L)AWS in short? In this report, we set out which types of bias should be taken into account, when they occur in the lifecycle of AWS, and what this means for policymakers.

Types of Bias

Although the rolling text at the moment of writing1 (September 2024) refers to unwanted bias in data sets, and unwanted automation bias, these two types of biases do not cover the full spectrum of bias. To clarify, bias is considered to be any type of flaw in an algorithmic system that leads to a statistic estimate that does not equal the true value2. For example, an unfair over- or underrepresentation of specific groups of people based on gender3, ethnicity, sexual orientation, religion or location, or the misclassification of objects (e.g. hospitals, places of worship, etc.).

Pre-existing biasTechnical biasEmergent bias
The first type of bias we recognize is the pre-existing bias4. This refers to any type of bias that already exists in society and is thus replicated, and often enlarged, by algorithms. This would include biased data sets as included in the rolling text.Technical bias refers to any kind of bias that occurs from the limitations of a system. This could include a system drawing from an alphabetic list, unintentionally favoring options first up in the alphabetic order.The last type of bias occurs over time, as a result from changing societal knowledge, population or cultural values. This could also include algorithmic decay, model drift or degradation.

Bias Over the System’s Life-Cycle

Bias in data collection

The first stage of development and deployment of systems is the selection of data that will be used for the training and testing of the system. These data sets are incredibly vulnerable to bias, mostly pre-existing bias. This can originate from the data sets, lack of available data (especially in military use5), but also from the data selection by developers. Where a developer could be held accountable for biased data selection, who is accountable for data sets that reflect a certain societal status quo — including the bias that already exists in society?

Bias in design and development

When the required data is collected, the model development (also referred to as model training stage) takes off. During this stage, the parameters of the models are fine-tuned. This includes decisions like: will the system make a final decision when it’s 90% sure, 95% sure, or 99% sure? Each parameter that is set comes with the risk of further enlarging bias that already exists in the data that the training stage started with (e.g. class imbalance6 – a preference for bigger ‘classes,’ for example caused by network forming in collaborative filtering algorithms7, and overfitting8 – the model recognizing random noise as a trend) and include (unconscious) bias of the developers.

It is common practice for data scientists to randomly split the initial dataset into two parts: one for training the model (model development) and the other for testing it (model validation), a process referred to as cross validation9. However, when the original data set contains certain bias, the model is then validated with a biased data set as well (and thus not really validated for the ‘real world’).

Additionally, this is the stage where technical bias comes in. To spot bias that is caused by technical limitations, it is vital that developers with different backgrounds analyze the process of model development and evaluation.

Bias in deployment and monitoring

Once in use, emerging bias is the greatest risk. This happens when the world changes, and the model is not re-trained. The loss of accuracy can be referred to as degradation10, model drift11, data drift12 or decay12. Data drift, degradation or decay occurs when the data that was used to train (develop) and test (validate) the algorithm no longer reflects the situation in which the model takes decisions, which is sometimes referred to as a distributional shift in environments. In a military context, this for example happens when a system is trained in a specific environment, which changes the longer an armed conflict continues. Model drift includes data drift, but includes other types of drift that lead to a change between the input and output variables, e.g. changing (legal) definitions or changes in military uniforms that challenge the recognition and classification of combatants.

Next to degradation, the self-learning capacities of AI can cause a negative self-reinforcing feedback loop13, which could be considered a form of overfitting over time. The model then identifies noise as a pattern, labeling for example individuals with a certain physical appearance or geographical location as targets.

Automation bias14, as referred to in the rolling text, then occurs when these fallacies are not corrected by human decision-makers because they place a higher degree of trust in the system than in human decision-making.

What Does This Mean for Policy-Makers?

An unbiased system does not exist. However, there are ways to mitigate these risks as much as possible. For policy-makers, the different risks of bias mean that there are several aspects to take into account when moving to a next step in the discussion around (L)AWS.

  1. Bias requires a lens that sees beyond just social and automation bias that is currently included in the rolling text.
  2. Although there has been no attention to the actors and teams that develop the algorithms used in (L)AWS, they greatly affect the decision-making by (L)AWS. According to a 2022 global survey with over 70,000 respondents, over 90% of developers are male15 — such a widespread survey has not yet been conducted for ethnicity, underlining the problem of lacking attention for the broad range of dimensions diversity is required in — but smaller surveys show 6016–7517% being white. The lack of diversity in these areas signals a similar homogeneity for other dimensions of diversity, e.g. sexual orientation, religious background, etc. Additionally, the private actors that currently focus on the development of algorithms for military use all have specific interests in the process. This can lead to over-stating the accuracy, only training (developing) and testing (validating) in specific circumstances and environments, and a lack of transparency on the development process of the final algorithm. It is thus vital to consider the public-private relationships that occur in the context of military use.
  3. Although the current rolling text refers to rigorous testing and evaluation of how the weapons system will perform, periodic reassessment of these evaluation measures is the absolute minimum to mitigate the risks of algorithmic decay, drift or degradation, and increase the chances of anticipated effects and predictability of the algorithmic decision-making.
  4. The current rolling text refers to traceable and explainable effects of the use of LAWS, but does not yet operationalize these requirements. To operationalize, policy-makers could consider requiring (L)AWS — or parts of these systems — to be developed open source, open core, or with source available18 to avoid so-called black box algorithms.

Further Readings

ICRC’s Blog Series on AI in the Military

  • The Risks and Inefficacies of AI-systems in Military Targeting Support by Jimena Sofía Viveros Álvares — read here
  • Falling Under the Radar: The Problem of Algorithmic Bias and Military Applications of AI by Ingvild Bode — read here
  • The Problem of Algorithmic Bias in AI-based Military Decision-Support Systems by Ingvild Bode and Ishmael Bhila — read here

UNIDIR Report on AI and Gender

  • Does Military AI Have Gender? By Katherine Chandler — read here

References

  1. GGE on LAWS. Rolling text. Convention on Certain Conventional Weapons – Group of Governmental Experts on Lethal Autonomous Weapons System.
  2. Delgado-Rodriguez M, Llorca J. Bias. J Epidemiol Community Health. 2004;58(8):635–641. doi: 10.1136/jech.2003.008466.
  3. Acheson R. Gender and bias. Women’s International League for Peace and Freedom. 2021. Link. Accessed Oct 1, 2024.
  4. Friedman B, Nissenbaum H. Bias in computer systems. ACM Transactions on Information Systems (TOIS). 1996;14(3):330–347.
  5. Corrected oral evidence: Artificial intelligence in weapons systems. 2023(2). Link. Accessed Oct 1, 2024.
  6. Bauder RA, Khoshgoftaar TM, Hasanin T. An empirical study on class rarity in big data. 2018-12:785–790. doi: 10.1109/ICMLA.2018.00125.
  7. Google. Collaborative filtering. Link. Accessed Oct 1, 2024.
  8. Ying X. An overview of overfitting and its solutions. 2019;1168:022022.
  9. Scikit Learn. 3.1. cross-validation: Evaluating estimator performance. Link. Accessed Oct 1, 2024.
  10. Bayram F, Ahmed BS, Kassler A. From concept drift to model degradation: An overview on performance-aware drift detectors. Knowledge-Based Systems. 2022;245:108632. doi: 10.1016/j.knosys.2022.108632.
  11. Holdsworth J, Belcic I, Stryker C. What is model drift? | IBM. Link. Updated 2024. Accessed Oct 1, 2024.
  12. Stihec J. Understanding data decay, data entropy, and data drift: Key differences you need to know. Link. Updated 2024. Accessed Oct 1, 2024.
  13. Hagen A. Negative feedback loops: Using an economic model to inspect bias in AI. 2020. Link. Accessed Oct 1, 2024.
  14. Automation bias. Databricks. Link. Updated 2019. Accessed Oct 1, 2024.
  15. Vailshery LS. Software developers: Distribution by gender 2022. Statista. Link. Accessed Oct 1, 2024.
  16. McEnvoy D. How ethnically diverse is the tech workforce? Link. Updated 2022. Accessed Oct 1, 2024.
  17. Weststar J. Developer satisfaction survey 2021. Western University, Ontario, Canada. Link. Accessed Oct 1, 2024.
  18. Langhammer J. Black box security software can’t keep up with open source | authentik. Authentik. Link. Updated 2023. Accessed Oct 1, 2024.

Whitepaper by Alycia Colijn (The Netherlands) and Heramb Podar (India)