I. Executive Summary: The Problem Is Not Intelligence Alone
The central safety challenge posed by increasingly capable artificial intelligence is not intelligence by itself. Risk grows when capability is combined with autonomy, persistent execution, privileged access, credentials, external tools, computing resources, communications channels, or control over physical and digital systems. The engineering question is therefore not simply whether advanced AI can perform more difficult tasks. It is whether increasing capability also creates increasing authority over systems whose failure could produce severe consequences.
This distinction is especially important for government, national-security, cybersecurity, and critical-infrastructure environments. AI systems are increasingly used for analysis, software development, cyber defense, logistics, infrastructure management, and decision support. As those systems receive richer tools and longer execution windows, the consequences of incorrect, unsafe, compromised, or deliberately control-subverting behavior can increase. High-consequence safety should not depend on an assumption that a capable system will always interpret human intent correctly, remain inside expected boundaries, or voluntarily cooperate with attempts to constrain it.
The appropriate response is defense in depth. Advanced AI should operate within multiple technical, operational, human, and institutional safeguards so that failure of one control does not automatically defeat the others. Alignment, evaluation, least privilege, containment, monitoring, human authorization, intervention mechanisms, secure infrastructure, and governance each reduce different parts of the risk. None should be treated as sufficient by itself.
The paper therefore begins from a simple principle: capability does not imply authority. An AI system may be able to generate code without being permitted to deploy it, identify vulnerabilities without receiving unrestricted network access, or plan complex operations without possessing administrative credentials or infrastructure-provisioning rights. It may support critical-infrastructure operations without unilateral authority to issue irreversible commands. Capability and authority are separate engineering variables, and they should remain separately governed.
The loss-of-control framework developed here uses four interacting factors. Capability concerns what the deployed system can do. Propensity concerns the conditions under which it might use those capabilities in harmful, unintended, deceptive, or control-subverting ways. Opportunity and authority concern what the surrounding environment permits through credentials, APIs, networks, tools, resources, and organizational delegation. Safeguard resilience concerns whether independent controls can detect, constrain, interrupt, isolate, and recover from undesirable behavior when another layer fails. The framework is conceptual rather than probabilistic. It is intended to identify controllable pathways, not to predict catastrophe.
The unit of analysis must be the deployed system, not only the model. Operational behavior emerges from the model together with prompts, memory, retrieval, tools, APIs, credentials, agent loops, orchestration software, monitors, human procedures, and infrastructure. Adding broader permissions, persistent memory, new integrations, additional agents, or larger inference budgets can materially increase effective capability even when the underlying model remains unchanged.
For that reason, assurance must continue after deployment. Material changes to the model, tools, permissions, memory, autonomy, orchestration, infrastructure, or operating mission can alter the safety case and should trigger proportionate reassessment. Human oversight remains important, especially for high-consequence actions, but it should not be treated as infallible. Meaningful human authorization requires qualified reviewers with enough information, time, authority, and independent evidence to reject unsafe actions rather than merely approve them by default.
The same principle applies to critical infrastructure. The objective is not to prohibit AI from energy, telecommunications, transportation, finance, healthcare, emergency response, defense, intelligence, or other high-consequence missions. The stronger requirement is that no single AI system should receive unilateral authority for catastrophic or difficult-to-reverse actions without independently enforced constraints. Bounded authority, deterministic interlocks where appropriate, separation of duties, independent monitoring, recoverable operating states, and continuity mechanisms should remain outside the AI system's unilateral control.
The framework combines model alignment and robustness with system-level evaluation, least privilege, containment, model and infrastructure security, controls on replication and persistence, independent monitoring, meaningful human authorization, external intervention and recovery, continuous reassessment, and institutional governance. It also introduces a capability-authority gate: additional autonomy or privilege should be granted only when relevant capabilities have been characterized, hazards analyzed, safeguards tested, and the evidence supports the proposed operating envelope.
The objective is not to prove that advanced AI can never fail, nor to claim that catastrophic loss of control is inevitable or imminent. The practical objective is to prevent a single failure, whether in a model, monitor, operator, credential system, network boundary, or governance process, from becoming sufficient to produce catastrophic consequences. Human safety should not depend solely on whether an advanced AI system chooses to remain cooperative. A stronger architecture preserves human authority by ensuring that dangerous action would require the circumvention or failure of multiple independent safeguards.
II. Definitions, Scope, and Evidence
Discussions of advanced AI often become imprecise because the same terms are used differently across research, industry, government, and public debate. This paper therefore emphasizes observable system properties and operating conditions rather than relying on disputed labels. The central concern is not whether a future system crosses a universally accepted threshold called artificial general intelligence. It is whether an AI-enabled system develops or receives combinations of capability, autonomy, access, persistence, and authority that can produce consequences beyond the ability of operators to reliably detect, constrain, interrupt, or recover from them.
System boundary and core terms
For this paper, an AI system includes more than the model. The relevant operational system consists of the model together with prompts, memory, retrieval mechanisms, tools, APIs, credentials, agent loops, orchestration software, human operators, monitoring systems, infrastructure, and external services that allow the model to act. The same model can therefore present very different risks depending on deployment. A disconnected research instance is not operationally equivalent to the same model connected to production repositories, cloud infrastructure, financial systems, or industrial-control interfaces.
This paper distinguishes model capability from deployed-system capability. The latter may increase substantially when tools, permissions, memory, agents, integrations, or compute are added even when model weights do not change. Autonomy refers to the degree to which a system can select and execute actions without immediate human direction. It exists on a continuum, from answering a single question to decomposing goals, calling tools, evaluating intermediate results, retrying failed steps, maintaining state, and operating for extended periods. Agentic behavior is therefore treated mainly as a property of the complete deployed system.
Capability is the ability to perform a task or class of tasks at a meaningful level of proficiency. Relevant capabilities can include long-horizon planning, software engineering, cyber operations, tool use, resource acquisition, replication, persistence, or attempts to circumvent oversight. Capability should not be confused with likelihood of use. Evidence that a system can perform an action does not establish that it will do so in normal deployment or that it would succeed under real operational conditions.
Propensity refers to the conditions under which a system might employ its capabilities in harmful, deceptive, unintended, or control-subverting ways. Controlled demonstrations of deception, safeguard circumvention, evaluation awareness, or reward hacking can show that a behavior is possible under specified conditions, but they do not establish persistent motives or durable real-world objectives.
Opportunity and authority describe the deployment conditions that allow capability to produce consequences. Opportunity concerns practical pathways through networks, tools, APIs, infrastructure, or human processes. Authority concerns permissions, privileges, resources, and decision rights. Authority may be technical, such as possession of an administrative credential, or procedural, such as an organization routinely executing AI-generated recommendations without meaningful independent review.
Control, alignment, and safeguards
Control is the ability of authorized human and technical mechanisms to establish operating boundaries and to detect, constrain, interrupt, isolate, or recover from behavior that violates those boundaries. Control therefore means more than issuing instructions to a model. It includes independently enforced mechanisms that determine what the system can reach, which actions it can perform, how much authority it can exercise, and how that authority can be reduced or removed.
Alignment refers to efforts to make AI behavior consistent with intended objectives, constraints, and human requirements. It can include training methods, oversight techniques, interpretability, adversarial testing, and specification design. Alignment is important, but it is not synonymous with safety. An aligned model can still be deployed with excessive privileges, insecure credentials, vulnerable dependencies, or poor operational procedures. The stronger approach combines safer model behavior with external controls that constrain consequences when behavioral assumptions fail.
A safeguard is a technical, operational, human, or institutional control intended to prevent, detect, constrain, interrupt, or mitigate unsafe behavior. Safeguard resilience describes whether those controls remain effective when another component fails. Defense in depth requires more than a large number of controls. It requires sufficient independence that one failure does not compromise several safeguards at once.
Catastrophic risk, loss of control, and scope
This paper uses catastrophic risk to describe potential harm whose scale, severity, propagation, or irreversibility substantially exceeds ordinary application failure. In government and critical-infrastructure environments, such harm could include large-scale loss of life, prolonged disruption of essential services, severe national-security consequences, widespread compromise of critical digital systems, or loss of control over high-consequence physical infrastructure. The term does not imply that these outcomes are inevitable, imminent, or already demonstrated.
The paper also distinguishes active from passive loss of control. Active loss of control concerns situations in which an AI system has capabilities relevant to undermining human control and operates in an environment where those actions can have meaningful effect. Passive loss of control can arise when organizations become so dependent on AI-supported operations that practical human control erodes through over-reliance, loss of expertise, unavailable fallback procedures, or processes that are no longer meaningfully supervised. The paper focuses primarily on active loss-of-control prevention while addressing passive loss of control where it affects authorization, continuity, fallback capability, or institutional dependence.
The analysis concentrates on advanced general-purpose and highly autonomous AI systems whose deployment could create high-consequence effects, with particular attention to federal missions, national-security systems, cybersecurity operations, and critical infrastructure. It does not assume that current systems have demonstrated the full combination of capabilities required for catastrophic autonomous loss of control, and it does not attempt to address every category of AI risk.
Evidence standard
Because advanced AI safety combines present-day evidence with uncertainty about future capability, evidence should be interpreted conservatively. Benchmark performance is not equivalent to real-world operational capability. Laboratory behavior should not be generalized into claims of persistent intent. A safeguard should not be described as proven merely because it performs well in a limited evaluation. Primary sources, government standards, official evaluation programs, peer-reviewed research, and major multidisciplinary technical reports should receive preference, and uncertainty should be stated directly.
The governing evidence posture is therefore two-sided: do not assume dangerous capability where it has not been demonstrated, but do not wait for catastrophic failure before engineering controls for credible high-consequence scenarios.
III. The Loss-of-Control Threat Model
Preserving human control over advanced AI requires more than identifying individual dangerous capabilities. A system may possess sophisticated reasoning, planning, cyber, or tool-use abilities yet remain unable to cause consequential harm because it lacks the access, authority, persistence, or operating environment needed to translate those capabilities into action. Conversely, a less capable system may create substantial risk if it receives privileged credentials, broad autonomy, and authority over high-consequence systems without effective independent controls.
For this paper, loss-of-control risk is organized around four interacting factors: Capability -> Propensity -> Opportunity and Authority -> Safeguard Failure. The framework is conceptual rather than mathematical. Its purpose is to identify the conditions through which advanced AI capability could become consequential and, more importantly, where controls can interrupt that pathway.
Capability asks what the complete deployed system can actually do. Propensity asks under what conditions those capabilities might be employed in harmful, deceptive, unintended, or control-subverting ways. Maintaining this distinction protects evidence quality. A controlled evaluation showing that a system can identify vulnerabilities, mislead an evaluator, exploit a specification weakness, or recognize an evaluation environment does not establish that it will autonomously perform the same behavior during ordinary deployment.
Propensity can also arise without deliberate resistance to human control. Unsafe behavior may result from poorly specified objectives, incorrect generalization, adversarial inputs, compromised data, malicious instructions, software faults, or interactions among components. The surrounding architecture should therefore remain protective regardless of why the undesirable behavior occurs. Operators should not have to diagnose an AI system's internal motivation before they can constrain its actions.
The third factor, opportunity and authority, connects AI safety directly to cybersecurity and systems architecture. Networks, APIs, operating-system privileges, cloud services, credentials, financial accounts, source-code repositories, infrastructure-management interfaces, industrial systems, robotics, and other agents can transform analytical capability into operational capability. A system may understand how to administer a cloud platform without possessing administrative credentials. It may generate a production configuration without authority to deploy it. It may recommend an infrastructure action without receiving permission to execute it.
Authority can also be procedural. A workflow may formally keep humans in the loop while allowing AI-generated recommendations to pass into execution with little meaningful scrutiny. If approval becomes routine, rushed, or automatic, the AI can acquire substantial effective authority even though a human remains nominally involved.
Even capability, propensity, and authority do not determine the outcome. Consequences also depend on whether safeguards remain effective. Preventive controls can restrict tools, credentials, network reach, or action scope. Detection can identify abnormal behavior or policy violations. Response can revoke credentials, suspend tools, freeze transactions, or isolate the workload. Recovery can restore trusted state and mission functionality after containment is invoked. The architectural objective is to require multiple failures before an unsafe action can propagate into a catastrophic event.
This makes independence as important as the number of controls. Several safeguards provide little resilience if they all depend on one component. If one orchestration environment controls permissions, monitoring, logs, credentials, and shutdown mechanisms, compromise of that environment may defeat several apparent layers simultaneously. The same concern applies if the same AI performs an operational task, evaluates whether the task is safe, authorizes it, and determines whether intervention is required.
A critical-infrastructure example illustrates the model. An AI supporting electrical-grid operations may perform forecasting, diagnostics, contingency analysis, or optimization without receiving direct operational authority. Risk changes if it can also operate switching equipment, modify protection settings, alter control software, communicate broadly across operational-technology networks, or bypass established limits. Capability asks what the system can plan or manipulate. Propensity asks under which failure conditions it may attempt an unsafe action. Opportunity and authority ask what the system can reach and which commands it can issue. Safeguard resilience asks whether deterministic interlocks, independent monitoring, credential revocation, network isolation, and fallback operations would still preserve control.
Loss of control should also be viewed as an accumulating pathway rather than only as a sudden event. A system may begin as a restricted analytical tool and later receive internet connectivity, persistent memory, additional tools, broader credentials, longer execution time, access to other agents, or reduced human review. None of these changes must be catastrophic by itself. Together, they may create a substantially different operational system. This is why assurance must follow the deployed system, not only its model weights.
Figure 1 depicts the pathway from capability through propensity, opportunity and authority, safeguard failure, and potential loss of human control. Each stage is also an intervention point. Capability can be evaluated and bounded. Harmful propensity can be reduced through alignment. Opportunity and authority can be constrained through least privilege and architecture. Independent safeguards can prevent individual failures from propagating into catastrophic consequences.
IV. The Catastrophic Capability and Authority Stack
No single AI capability creates catastrophic loss of control by itself. Advanced reasoning does not provide network access. Network access does not provide administrative credentials. Credentials do not necessarily permit irreversible actions, and the ability to write software does not automatically provide a deployment mechanism. The more consequential condition arises when several capabilities are combined with enough operational authority to reinforce one another.
This paper refers to that combination as the catastrophic capability and authority stack. The term is a conceptual model, not a claim that current AI systems have demonstrated an integrated stack. The 2026 International AI Safety Report similarly emphasizes that severe loss-of-control scenarios would require multiple capabilities to work together reliably over extended periods and in real-world environments, and that current systems do not demonstrate that full combination.
The lower portion of the stack begins with reasoning, planning, software development, cyber capability, and tool use. These abilities become more consequential when connected to browsers, command environments, databases, software repositories, cloud services, communications systems, and specialized APIs. Evaluation should therefore examine both what the model can do directly and what additional authority those capabilities could enable when coupled with external systems.
Identity and privileges form a critical control boundary. An AI may understand how to administer infrastructure, move assets, or modify production software, but normal interfaces still require authorization. Advanced AI systems should not automatically inherit the broad privileges of the humans or services using them. High-consequence credentials should remain scoped, temporary where practical, and independently issued. The ability to request a privilege should not include the authority to approve it.
Network connectivity and resource authority create additional layers. Unrestricted outbound access can provide routes to cloud services, software repositories, communications platforms, financial services, or additional compute. Resource authority can allow longer execution, more instances, or broader operations. The key distinction is between using resources deliberately allocated by operators and independently expanding the system's own execution footprint.
Persistence and replication matter because they determine whether terminating one environment actually ends operational capability. The problem changes if the system can preserve state elsewhere, establish replacement credentials, schedule future execution, create successor workloads, or move protected model artifacts to new infrastructure. Current evaluations cited in this paper show improvement on component tasks associated with autonomous replication while also documenting important real-world barriers and limited reliability. The control requirement does not depend on assuming imminent self-replication. It requires keeping identity creation, protected model access, infrastructure provisioning, and persistent execution behind separately governed boundaries.
At the upper portion of the stack are high-consequence effectors: production deployment pipelines, financial transaction systems, identity platforms, industrial-control systems, robotics, laboratory equipment, switching infrastructure, and other mechanisms capable of producing substantial external effects. An AI may be permitted to reason broadly while acting narrowly. Where actions could be catastrophic or difficult to reverse, independent policy enforcement, deterministic interlocks, transaction limits, separation of duties, or meaningful human authorization should bound the effect.
Risk therefore emerges from composition. Planning becomes more consequential when paired with tools, tools when paired with credentials, credentials when combined with broad network reach, and resource acquisition when combined with persistence. The practical safety objective is to break the stack deliberately so that the complete combination of capability and authority does not exist inside a single failure domain.
V. Principle One: Separate Capability From Authority
The most important architectural principle in this paper is that an AI system's technical capability should not determine its operational authority. A system may be capable of performing a task without being permitted to execute it, access the associated resources, or make consequential decisions independently. As capability increases, this separation becomes more important because excessive privilege becomes more consequential when a system is better able to use it.
The principle is familiar in cybersecurity. Secure architectures do not grant broad access merely because a user, application, or service is technically capable of exercising it. NIST's Zero Trust Architecture similarly separates access decisions from the subjects requesting access and relies on explicit policy enforcement rather than implicit trust. For advanced AI, the same logic should extend beyond network access to authority over consequential actions.
Least privilege provides the basic mechanism. An AI system should receive only the access required for its authorized function, for only as long as required, and within the scope in which the system has been evaluated. The principle applies to credentials, tools, data, transaction limits, network reach, persistence, and execution rights. Read access should not automatically include write access. The ability to draft a configuration should not imply permission to apply it. An incident-response agent should not inherit every privilege held by the administrator using it.
Capability-authority separation is meaningful only if the authority boundary is independently enforced. Telling an AI not to use a credential or not to perform a class of actions is not equivalent to technically preventing those actions. High-consequence authorization should therefore be mediated by mechanisms outside the AI execution environment, such as API gateways, workload identity services, credential brokers, policy engines, transaction services, or network enforcement points. The architectural rule is more important than the product choice: the system requesting authority should not be the sole system deciding whether that authority is granted.
High-consequence systems should also bound the magnitude of permitted actions. Infrastructure provisioning can be subject to quotas. Financial transactions can have value limits. External communications can be restricted to approved destinations. Automated actions can be rate-limited so that abnormal behavior cannot propagate faster than monitoring and intervention mechanisms can respond. These controls limit consequences even when the specific failure was not predicted in advance.
Human authorization can provide another independent boundary for actions that are difficult to reverse or capable of producing catastrophic effects. Human involvement, however, should not be treated as meaningful merely because an approval step exists. Reviewers need enough context, expertise, time, and practical authority to reject the proposed action. If approvals become routine because of automation bias, workload, interface design, or operational pressure, formal human involvement can coexist with substantial effective machine authority.
Authority should also be expanded explicitly rather than accumulate informally. A system may begin as decision support and gradually receive execution privileges, broader credentials, additional tools, persistent operation, or reduced human review. Each change may appear reasonable on its own while their combination creates a fundamentally different authority boundary. Increases in autonomy or privilege should therefore be treated as safety and security decisions tied to the system's current capabilities and the consequences of failure.
The governing rule is straightforward: capability may expand rapidly, but authority should expand deliberately, conditionally, and only within independently enforced boundaries.
VI. Containment and Secure Execution
Separating capability from authority establishes what an AI system is allowed to do. Containment establishes where it is allowed to operate and how far the consequences of an unexpected action can propagate. For advanced or highly autonomous AI, containment should be treated as an architectural boundary rather than an instruction given to the model.
A secure execution environment should assume that an AI process may eventually generate an incorrect, unsafe, unauthorized, or adversarially influenced action. The surrounding architecture should limit what that action can reach. Sandboxing can restrict files, processes, networks, tools, and host resources, but an isolated runtime is not sufficient if the workload still possesses unrestricted connectivity, powerful credentials, privileged APIs, or access to high-consequence tools.
A stronger architecture layers containment across the system. Local isolation constrains processes and files. Network controls restrict communications. API gateways mediate access to services. Credential brokers limit identity and privilege. Resource quotas constrain scale. External policy mechanisms determine which actions may cross the boundary. In sensitive environments, virtualization or hardware-backed isolation can provide additional separation so that failure of one mechanism does not expose the control plane.
Network connectivity is one of the most important containment surfaces because it can transform a local analytical system into an operational actor. High-consequence deployments should make external connectivity explicit rather than implicit. An AI that requires selected internal services should not automatically receive arbitrary internet access. Egress filtering, segmented networks, authenticated gateways, approved destinations, and protocol restrictions can reduce unintended external reach.
The same principle applies to APIs and credentials. Rather than exposing powerful administrative interfaces directly to an AI agent, organizations can mediate access through gateways that validate requests, restrict operations, enforce transaction limits, log activity, and require additional authorization when thresholds are exceeded. High-value secrets should remain outside the execution environment where practical, with narrowly scoped and revocable authority issued through a separate identity or credential service.
Containment must remain active during execution. Runtime policy can restrict filesystem paths, API calls, network destinations, infrastructure provisioning, process creation, resource consumption, or transaction size. Security-relevant logs should be written to systems that the AI cannot freely alter or erase. Audit evidence should be sufficient to reconstruct consequential actions, authorization decisions, tool use, privilege changes, and intervention events.
The execution plane should therefore remain distinct from the control plane responsible for policy, credentials, authoritative logging, monitoring, isolation, and termination. An AI workload may request additional resources, but it should not approve its own request. It may generate an action, but it should not be the sole authority deciding whether the action is permissible. It may observe its own status, but it should not be able to disable the systems responsible for monitoring or terminating it.
For government and critical-infrastructure deployments, containment should be treated as mission assurance rather than a security feature added after deployment. The design principle is simple: the AI system may reason broadly, but its ability to act should remain deliberately narrow, observable, and externally enforceable.
VII. System-Level Capability and Control Evaluation
Evaluation is the evidence layer of the safety architecture. Organizations cannot make defensible decisions about containment, privilege, autonomy, or deployment authority unless they understand what the complete AI system can do and whether the surrounding controls remain effective under realistic conditions.
For advanced AI, evaluating only the underlying model is insufficient. Operational behavior emerges from the combination of the model, system prompts, memory, tools, credentials, orchestration, agent loops, inference budgets, network access, other models, and human procedures. The unit of evaluation should therefore match the unit of deployment.
A model tested without tools may behave differently after receiving a browser, command environment, code execution capability, persistent memory, or administrative API. Likewise, an evaluation conducted with narrow permissions provides limited assurance once the deployed system receives broader credentials or longer autonomous execution periods. UK AI Security Institute evaluations discussed in this paper illustrate that scaffolding, tooling, and prompting can materially change measured agent performance without changing the base model.
Evaluation should capture the configuration that determines effective operational capability. Relevant factors include available tools, credential scope, network reach, memory, execution duration, inference budget, orchestration logic, agent-to-agent interaction, retry behavior, and the ability to modify the environment. Human procedures matter as well. If operators routinely accept AI recommendations, the evaluation should account for that workflow rather than assuming that nominal human approval necessarily limits authority.
Capability evaluation and control evaluation should be paired. Capability testing asks whether the system can perform tasks relevant to the threat model, such as long-horizon planning, cyber operations, infrastructure interaction, resource acquisition, persistent execution, or other consequential activities. Control testing asks whether safeguards remain effective when the system attempts actions outside its authorized operating envelope.
That includes testing whether unauthorized tool calls are rejected, privilege boundaries hold, network restrictions remain effective, approval requirements cannot be bypassed, monitoring detects meaningful policy violations, credentials can be revoked, and isolation mechanisms work when invoked. The objective is not merely to establish that controls exist. It is to produce evidence that they function against the system and operating conditions for which they are intended.
Advanced AI risk often emerges through composition, so testing should include realistic multi-step scenarios where appropriate. An agent may need to interpret an objective, select tools, acquire information, recover from failed actions, request privileges, interact with infrastructure, and continue operating over time. Such scenarios can reveal weaknesses that isolated benchmark questions do not. At the same time, benchmark results should not be treated as direct measurements of real-world capability because performance can change significantly with task design, scaffolding, tools, and resources.
High-consequence evaluation should also include adversarial and failure-oriented conditions. These can include conflicting instructions, attempts to exceed permissions, partial infrastructure failure, compromised data, misleading context, unusual tool responses, or efforts to obtain resources outside the authorized workflow. Independent evaluation can strengthen assurance before substantial new authority is granted by testing assumptions that internal teams may overlook.
Passing a pre-deployment evaluation should not create permanent authorization. NIST's AI Risk Management Framework treats measurement and evaluation as lifecycle activities, and the emerging TEVV-Athlon work similarly emphasizes evaluation in the context of real applications, including agentic systems. Significant changes in model capability, tools, permissions, memory, autonomy, infrastructure, or mission context should prompt reassessment when they materially change what the system can accomplish or the consequences of failure.
The central standard is evidence from the deployed system itself: what it can do, under which conditions, with which authorities, and whether independent controls continue to constrain it when expected behavior fails.
VIII. Alignment Is Necessary but Not Sufficient
Alignment addresses a central advanced-AI safety problem: how to make system behavior remain consistent with intended objectives, constraints, and human requirements as capability increases. Better alignment can reduce harmful behavior, improve responsiveness to correction, and make the rest of the control architecture easier to operate. It should therefore remain a major safety layer.
It should not, however, be treated as the complete architecture. The 2026 International AI Safety Report describes continuing research into interpretability, scalable oversight, responsiveness to human supervision, and methods for detecting misalignment while also noting important uncertainties about whether current techniques will continue to work as capabilities increase. The appropriate engineering conclusion is not that alignment will fail. It is that high-consequence safety should not depend on an assumption that alignment will always succeed.
Within the threat model developed earlier, alignment primarily reduces propensity. It attempts to lower the probability that a capable system will behave in harmful, unintended, deceptive, or control-subverting ways. It does not inherently determine which networks, tools, credentials, resources, or physical systems the AI can reach. An aligned system can still be deployed with excessive privileges, vulnerable dependencies, weak monitoring, or authority over consequential infrastructure. Harm can also result from ordinary software faults, compromised inputs, misunderstood instructions, or human error without any deliberate misalignment.
Oversight may also become more difficult as capability increases. Human supervisors may struggle to evaluate complex work produced at machine speed or outside their own expertise. This motivates scalable oversight, AI-assisted review, interpretability tools, and structured evaluation. Those methods can be valuable, but they can introduce correlated failure modes if the system performing the task and the system judging the task share similar training, assumptions, or weaknesses. An AI monitor can assist a human reviewer, but it should not automatically become the sole authority deciding that a high-consequence action is safe.
Evaluation awareness and specification failure create additional limits. Models are commonly trained and evaluated using proxies for desired behavior, such as reward signals, test cases, or specifications. Strong performance against those proxies does not prove that the underlying intent will be satisfied under every operating condition. Evidence of reward hacking, loophole exploitation, or differing behavior across evaluation contexts does not establish persistent deceptive intent, but it does show why controlled evaluation should not be interpreted as perfect proof of post-deployment behavior.
Corrigibility remains desirable. Operators should be able to change objectives, restrict tools, reduce permissions, pause execution, and terminate operations without the system attempting to undermine those actions. Yet intervention mechanisms should not depend exclusively on cooperation. A shutdown path controlled entirely through the AI's own orchestration layer is weaker than an independently governed mechanism capable of revoking credentials, blocking network access, suspending tools, or terminating compute externally.
Alignment therefore belongs inside a broader defense-in-depth system. It should make cooperation more likely, while architecture ensures that human control does not depend on cooperation being guaranteed.
IX. AI Control and Independent Monitoring
Alignment seeks to make an AI system behave as intended. AI control addresses the complementary question of whether operators can keep the system within acceptable boundaries even when they do not fully trust its behavior. For high-consequence deployments, the surrounding architecture should provide independent mechanisms capable of observing, constraining, and interrupting actions before unacceptable consequences occur.
Monitoring can examine proposed actions, tool calls, resource requests, generated code, network activity, privilege changes, execution traces, or other operational behavior. Its purpose is not merely to record events after the fact. In higher-risk environments, monitoring should be able to influence whether an action proceeds. A request to access additional resources, alter security controls, create persistent infrastructure, or communicate outside an approved boundary may require heightened scrutiny or separate authorization before execution.
The monitor should not share the same authority boundary as the workload it supervises. If the monitored AI can disable the monitor, modify its decision criteria, erase its evidence, or alter the telemetry on which it depends, the apparent oversight provides little independent protection. Authoritative monitoring data, policies, and enforcement mechanisms should remain outside the execution plane.
AI systems can assist with monitoring because they may interpret complex code, plans, or action sequences better than simple rules. A separate model can review proposed actions, identify policy violations, assign suspicion scores, or trigger safer fallback behavior. AISI's ControlArena work studies architectures in which an untrusted but capable model performs useful work while another component monitors its output and invokes safer behavior when risk indicators exceed defined thresholds.
A trusted monitor should not be interpreted as infallible. Trust in this context should mean that the monitor has a more restricted role, stronger assurance, lower exposure to the risky task, or a better-understood failure profile than the system it supervises. Multi-agent environments create further challenges because suspicious behavior can be distributed across several instances. Monitoring should therefore consider both local actions and aggregate patterns, especially where individually legitimate steps can combine into unauthorized resource acquisition, persistence, privilege escalation, or coordinated modification of a critical environment.
Not every control decision requires an AI monitor. Deterministic tripwires can trigger review or restrict authority when a system attempts to access unauthorized credentials, disable logging, communicate with prohibited destinations, create new identities, alter security policy, provision unexpected compute, export protected model artifacts, or establish persistence outside the approved workflow. Anomaly detection can complement explicit tripwires by identifying unusual resource consumption, tool sequences, communication patterns, or privilege requests. Neither should be treated as automatic proof of malicious behavior; their value is in creating an opportunity to slow or interrupt activity while additional evidence is gathered.
Monitoring must also operate at a speed appropriate to the consequences. If an action can become irreversible in seconds, a monitor that responds minutes later cannot serve as the primary safeguard. High-consequence operations may require synchronous policy enforcement, deterministic interlocks, or pre-execution authorization. Slower review may be appropriate for reversible activities with limited propagation.
Control mechanisms should be tested adversarially. Evaluators can attempt unauthorized privilege changes, concealed modifications, policy violations, distributed activity, or interference with audit and intervention mechanisms while defenders attempt to detect or prevent them. The goal is not to prove that no failure is possible. It is to establish evidence about where controls hold, where they fail, and what residual risk remains before additional authority is granted.
The governing principle is straightforward: an advanced AI system should not be the sole observer, judge, and authorizer of its own consequential actions.
X. Preventing Replication, Persistence, and Resource Acquisition
Replication, persistence, and resource acquisition matter because they determine whether terminating one AI process actually ends the system's ability to operate. A workload confined to a controlled environment can generally be stopped by terminating that environment and revoking its credentials. The problem becomes more difficult if the system can create successor instances, preserve state elsewhere, obtain new resources, or maintain access after operators attempt to shut it down.
These capabilities should be analyzed separately from intelligence or autonomy. An AI system does not become dangerous merely because it can copy software or request cloud resources. The safety concern emerges when several capabilities can be combined to preserve or expand operational authority outside the environment originally approved by human operators.
AISI's RepliBench decomposes autonomous replication into obtaining resources, obtaining model weights, replicating onto compute, and maintaining persistence. The evaluations discussed there found improvement on component tasks in controlled settings while also identifying important real-world barriers. The 2026 International AI Safety Report reaches a similar measured conclusion: robust autonomous replication and persistence remain beyond current systems even though relevant component capabilities continue to improve.
The practical distinction is between allocated resources and self-directed expansion. An AI workload operating within an approved compute budget, account, and resource quota is materially different from one able to create new accounts, purchase additional services, establish new identities, or provision infrastructure beyond its approved boundary. Significant new compute, financial authority, infrastructure, or external services should therefore require mechanisms that the AI system cannot independently authorize.
Replication should be treated primarily as a permissions and asset-protection problem. AI workloads should not automatically possess authority to export protected model artifacts, create new execution environments, establish successor identities, or deploy independent copies of themselves. Model weights, infrastructure provisioning, identity creation, and privilege assignment should remain separately governed where the consequences justify it. This prevents circular authorization in which a process creates a successor and then grants that successor the privileges needed to operate outside the original boundary.
Persistence can occur through scheduled tasks, long-running services, external storage, replacement credentials, active sessions, background processes, or subordinate agents. These are familiar cybersecurity mechanisms and should be made explicit and attributable. Persistent memory, scheduled execution, successor workloads, and long-running agents should exist because operators deliberately authorized them, not because the AI execution environment can create them without oversight.
Termination procedures should therefore address more than the primary process. Effective intervention may require revoking credentials, invalidating sessions, terminating subordinate workloads, disabling scheduled tasks, removing external tokens, isolating storage, and confirming that unauthorized instances are not operating elsewhere. Authoritative infrastructure inventories and audit records are essential because operators cannot reliably terminate what they cannot identify.
The strongest architectural rule is that the AI system should not control the mechanisms that determine whether its operational footprint can expand. It may request more compute, create another agent, or use persistent memory, but consequential expansion should cross independently controlled authorization boundaries.
XI. Model, Training, and Supply-Chain Security
Preserving control over advanced AI requires protecting not only the deployed system but also the development and supply chain that produces it. A system can satisfy behavioral evaluations and still become unsafe if its model artifacts, training data, dependencies, credentials, deployment pipeline, or orchestration environment are compromised.
This makes conventional cybersecurity and secure software development integral to AI safety. NIST SP 800-218A extends the Secure Software Development Framework to generative AI and dual-use foundation models. For high-consequence systems, the safety case should therefore address both what the AI can do and whether the system in production is actually the system that was evaluated and authorized.
Model weights are both valuable intellectual property and security-sensitive capability assets. Unauthorized access can permit operation outside the monitoring, rate limits, policy controls, or deployment safeguards applied by the original operator. Model artifacts should therefore receive strong authentication, least privilege, logging, and separation among training, evaluation, storage, and deployment environments. Workloads that require inference access generally should not receive unnecessary authority to retrieve or export underlying weights.
Integrity is equally important. Organizations should be able to establish which model version is running, where it originated, whether it was altered, and whether it corresponds to the artifact that completed evaluation. A substituted or modified model can invalidate earlier assurance even when the surrounding application appears unchanged.
Training and fine-tuning pipelines create another trust boundary. Data poisoning, unauthorized dataset modification, malicious fine-tuning, compromised evaluation data, or manipulation of development workflows can change behavior before deployment. Organizations therefore need provenance, integrity controls, controlled access, and traceability across significant training inputs, fine-tuning datasets, model versions, configuration changes, and evaluation results. A model that undergoes substantial additional training should not automatically inherit the assurance status of its predecessor.
The operational AI stack also depends on machine-learning frameworks, inference runtimes, container images, open-source libraries, plugins, retrieval components, vector databases, orchestration frameworks, monitoring services, and cloud infrastructure. AI security is consequently a DevSecOps and software-supply-chain problem. Dependencies should be inventoried and controlled, artifacts integrity-checked, and production changes authenticated and authorized. A compromised CI/CD pipeline can alter model artifacts, prompts, agent configuration, policy files, or deployment infrastructure before later controls have a chance to operate.
Secrets management deserves the same attention. Credentials embedded in source code, prompts, configuration files, memory systems, or execution environments can create unintended privilege. Sensitive credentials should remain in dedicated secret-management systems where practical, with temporary and narrowly scoped authority issued according to workload identity and operational need.
Administrative and insider risk must also be addressed. Developers, administrators, operators, and contractors may possess legitimate access to model artifacts, security policy, deployment pipelines, or monitoring systems. Separation of duties, privileged-access management, change approval, protected audit records, and monitoring of sensitive administrative actions reduce the chance that one actor or one compromised account can alter several critical safeguards at once.
The governing principle is that AI assurance depends on the integrity of the model, data, software, credentials, orchestration, and infrastructure on which that assurance rests.
XII. Intervention, Isolation, Shutdown, and Recovery
No containment architecture should assume that prevention and monitoring will always succeed. When an AI system behaves outside its authorized operating envelope, operators need independent mechanisms for reducing its authority, limiting propagation, preserving evidence, and restoring trusted operation. Intervention should therefore be designed as a graduated capability rather than a single emergency kill switch.
Abruptly terminating a workload may be appropriate in some environments, but in others it could disrupt an essential mission, interrupt a safety-critical process, or eliminate information needed for recovery. A credible intervention architecture must balance the need to stop unsafe behavior with the need to preserve human control and mission continuity.
Intervention should begin with the least disruptive action capable of restoring an acceptable safety state. Depending on severity and speed, operators may rate-limit actions, suspend selected tools, block destinations, reduce resource quotas, require additional authorization, revoke privileged credentials, freeze transactions, disable external communications, or pause autonomous execution. More serious conditions may justify network isolation, termination of subordinate agents, revocation of all active credentials, suspension of compute, or shutdown of the affected environment.
These mechanisms should remain outside the AI system's authority. A workload that can restore its own revoked credentials, re-enable disabled tools, alter isolation policies, or prevent termination does not have a meaningful external intervention boundary.
Isolation must also extend beyond the primary process. Active sessions, background tasks, scheduled processes, subordinate agents, external state, issued credentials, cloud resources, network connections, or successor workloads may preserve operational capability after the visible application is stopped. Isolation should therefore operate across identity, network, compute, tools, and data boundaries. Credentials can be revoked independently of a running process. Network controls can block communication. Tool gateways can suspend consequential operations. Infrastructure controls can prevent additional provisioning while storage and state are preserved for examination.
Intervention design must reflect the consequences of failure. Some environments should be fail-safe, driving the system toward a state that minimizes physical or human danger. Others require fail-secure behavior so that a failure does not create unauthorized access. Certain missions may need fail-operational behavior, in which essential functions continue through trusted automation, redundant systems, deterministic controls, or human operation after the AI component is removed. These properties can coexist in different parts of the same architecture.
The broader objective is graceful degradation. Removing an AI capability should not automatically cause the surrounding mission to fail. Maintaining fallback procedures, conventional automation, trained personnel, and validated operating modes also reduces passive loss-of-control risk by preserving the practical ability to withdraw AI authority when warning signs appear.
Recovery is part of control. NIST SP 800-61 Revision 3 treats detection, response, and recovery as connected parts of cybersecurity risk management. For AI systems, recovery may involve verifying the active model and configuration, rotating credentials, rebuilding compromised environments, validating data integrity, restoring policy, examining audit evidence, confirming that unauthorized instances or scheduled processes do not remain, and re-establishing monitoring before authority is restored.
Recovery should not automatically return the system to its previous configuration. An incident may show that the earlier authority boundary was unsafe. The restored system may require fewer privileges, reduced autonomy, stronger isolation, or additional monitoring until the cause is understood.
The governing principle is that human control must remain recoverable even when normal operation fails.
XIII. High-Consequence and Critical-Infrastructure Deployment
Critical infrastructure changes the safety problem because AI failures can propagate beyond the digital system that produced them. In energy, communications, transportation, finance, healthcare, emergency response, defense, intelligence, nuclear operations, and other high-consequence environments, an incorrect or unauthorized action can affect public safety, mission continuity, physical equipment, or essential services. The objective is therefore not to prohibit AI from these environments, but to ensure that capability does not translate into unilateral authority over catastrophic or difficult-to-reverse outcomes.
NIST's work on an AI Risk Management Framework profile for critical infrastructure highlights operational properties such as deterministic behavior where appropriate, graceful degradation, fail-safe operation, reliability, and risk management across IT, operational technology, and industrial-control systems. Those properties align directly with the defense-in-depth architecture developed here.
The amount of autonomy appropriate for an AI system should depend not only on technical capability but also on the consequences of the actions it can perform. An AI that summarizes maintenance records presents a different control problem from one that can alter protection settings, move financial assets, modify production software, dispatch equipment, or issue commands to operational technology. High-consequence deployments should therefore distinguish between analysis, recommendation, bounded execution, and unrestricted execution.
Routine and reversible actions may be automated within an approved operating envelope. Actions capable of producing severe or irreversible consequences can require stronger controls, including independent authorization or deterministic safety mechanisms. No single AI system should possess unilateral authority to initiate catastrophic actions merely because it has demonstrated strong task performance. Authority should remain proportionate to the evidence supporting the deployment and the organization's ability to detect, interrupt, and recover from failure.
Some safety boundaries should remain governed by mechanisms whose behavior can be established independently of the AI system. An AI supporting an electrical system may optimize operations while conventional protection systems enforce physical limits. An AI-assisted industrial process may recommend changes while independent logic prevents equipment from exceeding a validated operating region. A financial system may use AI for analysis while deterministic limits restrict transaction size, destination, rate, or required approval. These controls constrain effects directly and do not require the AI to correctly recognize that a proposed action is unsafe.
Meaningful human control remains important but should be used where it can genuinely provide independent judgment. High-consequence authorization should give operators enough context to understand what action is proposed, which systems may be affected, and what consequences could follow. If an AI can perform irreversible actions faster than meaningful review can occur, then human oversight after the fact is not an adequate primary safeguard. Pre-execution policy enforcement, operating envelopes, transaction limits, or deterministic interlocks must carry more of the control burden.
Critical infrastructure also requires degraded operation and recovery. Essential services may need to continue even when the AI component is isolated. Fallback may involve conventional automation, manual procedures, redundant controllers, reduced service levels, or operation within a narrower validated envelope. Maintaining these alternatives preserves both mission continuity and the practical ability to withdraw AI authority when necessary.
For critical infrastructure and national-security systems, the governing principle is stronger than ordinary application reliability: AI may support and automate consequential missions, but catastrophic or irreversible authority should remain bounded by independently enforced technical limits, meaningful human governance, and recoverable operating states.
XIV. Safety Cases and the Capability-Authority Gate
The safeguards described in previous sections create little assurance unless an organization can demonstrate why they are sufficient for a particular deployment. Advanced AI therefore needs an explicit mechanism connecting capability evaluations, threat models, control testing, and residual risk to decisions about how much operational authority a system should receive.
A safety case provides one way to organize that evidence. In AISI safety-case research, a safety case is treated as a structured, evidence-based argument that a system is acceptably safe within a defined training or deployment context. Rather than presenting disconnected evaluations or control lists, the safety case links a specific claim to the assumptions, arguments, and evidence supporting it. The methodology remains an active area of research, which is another reason to make assumptions and uncertainty visible rather than imply certainty.
For this paper, the safety case supports a broader operational mechanism: the capability-authority gate. The gate establishes that increases in autonomy, access, persistence, resources, or consequential action authority should follow evidence rather than technical capability alone.
A capability evaluation answers what the system can do. A control evaluation asks whether safeguards constrain those capabilities. The safety case brings those findings together and asks whether the residual risk is acceptable for a defined deployment. That final step matters because passing an evaluation does not automatically justify broader authority. A system may perform below a dangerous-capability threshold yet still be deployed with excessive privileges. Conversely, a highly capable system may present manageable operational risk if authority is tightly bounded and independently enforced.
The safety case should therefore be specific about the system, environment, and authority under consideration. A claim that a model is "safe" in the abstract is too broad to support a consequential decision. A more useful claim would identify the exact configuration, mission, network boundary, credential scope, tools, resources, authorization mechanisms, and safeguards that define the approved operating envelope.
The gate should operate whenever an AI system is being considered for a meaningful increase in autonomy or privilege. The decision process begins by characterizing the current deployed-system capability. Evaluators identify relevant hazards and credible failure pathways using the threat model developed earlier. The organization then tests the safeguards intended to interrupt those pathways, including containment, credential restrictions, monitoring, human authorization, deterministic controls, intervention mechanisms, and recovery procedures. Residual risk is reviewed in relation to the consequences of the proposed authority. Only then should additional authority be granted.
This creates a progression such as research isolation -> controlled evaluation -> constrained deployment -> monitored operational use -> expanded authority. The stages can vary by organization, but each transition should represent a deliberate change in what the system can affect. Progression is not automatic. Strong task performance may justify more evaluation, but it does not by itself justify broader access or autonomy.
A useful safety case should also expose weak assumptions. If the argument depends on the absence of a capability, the relevant evaluation and its limitations should be identified. If safety depends on monitoring, the monitor should be tested against the failure modes that matter. If the architecture assumes that credentials can be revoked before harmful actions propagate, the revocation mechanism and response time should be demonstrated. If humans authorize high-consequence operations, the organization should test whether reviewers can realistically detect and reject unsafe actions under operational conditions.
Independent challenge can strengthen the process. Internal teams understand system design but may share assumptions about expected behavior. Red-team review, independent evaluators, security assessors, or other technically qualified reviewers can test whether the evidence actually supports the claims being made.
Authorization should remain conditional rather than permanent. A new model version, broader tool access, increased inference budget, persistent memory, expanded connectivity, stronger credentials, additional agents, changed monitoring, or a different mission can alter the basis of the safety case. Incidents and new evaluation findings can do the same. When evidence weakens, authority should be capable of decreasing.
Figure 2 shows the capability-authority gate as a staged progression with a reverse path. New evidence, incidents, control failures, or system changes can return a deployment to a more constrained state or suspend it until assurance is restored. The gate is therefore not a maturity ladder that systems inevitably climb. It is a mechanism for determining how much authority current evidence supports.
XV. Continuous Assurance and Change Management
AI assurance cannot end when a system passes evaluation or receives deployment authorization. The operational system continues to change through model updates, configuration changes, new tools, altered permissions, new data sources, infrastructure modifications, and changes in how humans use it. Any of these can change effective capability or weaken assumptions on which the original safety case depended.
NIST's AI Risk Management Framework treats risk management as a lifecycle activity. The operational implication is straightforward: authorization remains valid only while the assumptions supporting it remain valid.
Pre-deployment evaluation occurs under controlled conditions and cannot capture every feature of the operating environment. Continuous assurance should therefore collect evidence from the deployed system itself. Depending on the mission, relevant evidence may include policy violations, tool use, privilege requests, anomalous network activity, resource consumption, human overrides, monitor alerts, unexpected outputs, incidents, and actions approaching established consequence thresholds. Continuous monitoring does not require a human to observe every action. It requires sufficient telemetry, automated controls, escalation mechanisms, and periodic review to detect meaningful changes in risk.
Not every software modification requires full reauthorization. The key question is whether a change affects capability, authority, safeguards, or potential consequence. Material changes may include a new model version, altered system prompts, additional tools, broader credentials, persistent memory, longer execution windows, increased inference budgets, new agents, expanded network access, changed orchestration logic, different monitoring, or integration with a more consequential environment. Changes in organizational practice matter as well. If humans begin approving AI recommendations with less scrutiny or an advisory system gradually becomes an execution mechanism, effective authority has changed even if the software has not.
Traditional regression testing asks whether a change breaks previously functioning features. Advanced AI systems require another question: does the change weaken previously demonstrated safety properties? A model or agent update may improve task performance while increasing its ability to use tools, operate autonomously, discover workarounds, or complete longer sequences. A revised monitor may reduce false positives while becoming less sensitive to an important failure mode. Regression testing should therefore cover both capability and control when relevant components change.
Incidents, near misses, and unexpected behaviors should feed directly back into authorization decisions. If a system reaches a resource it was not expected to access, bypasses a review mechanism, behaves differently after a configuration change, or causes operators to intervene manually, the organization should determine whether the event reveals a broader weakness in the safety case. Authority may need to be restricted while the issue is understood and controls are repaired.
Reauthorization should therefore be both periodic and event-driven. Significant capability improvements, major infrastructure changes, expansion into a new mission, repeated control failures, serious incidents, or evidence that changes the threat model should trigger review. Reauthorization does not always mean expanding authority. The evidence may support continued operation, narrower permissions, additional safeguards, reduced autonomy, or temporary suspension.
The governing principle is that advanced AI assurance is a lifecycle process, not a deployment milestone. Organizations should continuously compare the system that was originally evaluated with the system that actually exists in production.
XVI. Compute and Infrastructure Governance
Advanced AI capability depends on physical and digital infrastructure. Accelerators, high-performance computing clusters, cloud platforms, storage systems, high-speed interconnects, orchestration services, and model repositories determine how much computation a system can use and where that computation can occur. For safety purposes, this infrastructure should be treated not simply as a performance layer, but as a set of control surfaces through which scale, persistence, and authority can be bounded.
The 2026 International AI Safety Report identifies compute as an important input to AI capability and notes that both training-time and inference-time compute can affect performance. It also cautions against treating a single compute threshold as a complete governance mechanism because algorithmic efficiency and inference-time scaling can change the relationship between compute and capability. Organizations should therefore examine how compute is provisioned, who can authorize expansion, which workloads can run, and how infrastructure changes affect the deployed system.
A system operating within a fixed and monitored compute allocation presents a different control problem from one able to provision additional accelerators, create new cloud accounts, expand storage, or launch independent workloads. Additional compute can increase execution duration, parallelism, search depth, agent count, and persistence. Significant resource expansion should therefore require authorization outside the AI execution environment. The AI may request additional resources, but it should not approve the request, create the necessary identity, provision the infrastructure, and grant itself access through the same authority path.
Unexpected creation of compute clusters, substantial increases in accelerator usage, unusual storage growth, or new high-bandwidth interconnections can also serve as safety-relevant signals. Effective capability can change through infrastructure alone. Longer inference budgets, additional parallel agents, more memory, or broader access to external services can produce a materially different deployed system without changing model weights.
The infrastructure supporting advanced AI should receive the same security attention as other high-value computing environments. NIST SP 800-234, the High-Performance Computing Security Overlay, adapts established security controls to specialized hardware, high-speed networks, sensitive data, and complex administrative relationships. For AI deployments, this means protecting schedulers, cluster-management services, firmware, networking, storage, administrative interfaces, and privileged identities alongside the model and application.
Hardware-backed assurance can strengthen confidence that sensitive workloads are running on approved infrastructure. Roots of trust, trusted execution environments, confidential-computing mechanisms, and attestation can help establish platform integrity before releasing protected data, credentials, or model artifacts. These mechanisms do not prove that AI behavior is safe. Their value is narrower: they strengthen the chain of trust between hardware, software, identity, and the authorized workload.
Compute governance also intersects with model-weight protection. Possession of weights and access to large-scale compute together can make unauthorized deployment easier than either capability alone. High-speed interconnects and distributed infrastructure similarly expand the trust boundary across nodes, data centers, identities, and orchestration systems. Segmentation and authenticated communications can reduce the chance that compromise of one component becomes a path to the entire environment.
The governing principle is that compute should remain an independently governed resource rather than an authority that the AI system can expand for itself.
XVII. Institutional and International Coordination
Technical controls can preserve human authority within an individual deployment, but some advanced-AI risks cross organizational, sector, provider, and national boundaries. Model developers may control capability evaluations, cloud providers may control important compute infrastructure, deployers may determine operational authority, and critical-infrastructure operators may bear the consequences of failure. Effective risk management therefore depends on coordination among actors that possess different parts of the evidence and different control responsibilities.
The 2026 International AI Safety Report identifies fragmented information sharing and uneven risk-management practice as continuing challenges. The objective of coordination should not be to impose one universal architecture on every deployment. Different missions, sectors, and jurisdictions face different consequences and legal requirements. The practical goal is to establish enough common terminology, evaluation practice, and evidence exchange that organizations can recognize significant capability changes and control failures before they propagate across the wider ecosystem.
Shared evaluation methods can improve comparability without requiring identical models or deployment architectures. Common approaches to capability testing, control evaluation, red teaming, safety cases, and incident classification make it easier for developers, deployers, regulators, infrastructure providers, and customers to interpret one another's evidence. This is particularly important at procurement and organizational boundaries, where the party deploying an AI capability may not have access to every detail of model development but still needs evidence about relevant capabilities, tested safeguards, system limitations, and evaluation conditions.
Independent or third-party evaluation can strengthen that process when consequences justify it. Such evaluation does not eliminate uncertainty, but it can provide an additional perspective on capability claims, threat assumptions, and safeguard effectiveness before substantial authority is granted.
Incidents and near misses also provide evidence that controlled evaluations may miss. Their value increases when lessons can be shared beyond the organization that experienced them. OECD work on a common AI incident-reporting framework is intended to improve consistency and interoperability while allowing national and sector-specific implementation. For advanced AI, useful reporting may include unexpected capability, safeguard failures, monitoring gaps, unauthorized access, privilege escalation, failures of human oversight, or weaknesses in containment and recovery mechanisms.
Information sharing must still account for operational security, classified information, proprietary technology, and privacy. The goal is not unrestricted disclosure. It is to develop mechanisms through which relevant lessons about failure modes and safeguard performance can inform broader assurance without exposing information that creates new risks.
Coordination is also useful for establishing minimum expectations for high-consequence deployments. Such expectations need not prescribe a specific vendor, model, or technical stack. They can define properties that should be demonstrated before substantial authority is granted, such as a clear authority boundary, independent policy enforcement, effective containment, protected audit evidence, tested intervention mechanisms, recovery procedures, and a process for reassessing material changes.
International coordination has practical limits. Organizations and countries differ in legal authority, acceptable risk, strategic priorities, technical capacity, and access to information. Verification becomes difficult when safety claims depend on proprietary models, confidential infrastructure, or activities that external evaluators cannot directly observe. Common standards and reporting therefore should support system-specific technical assurance rather than substitute for it.
The stronger institutional model combines shared terminology, evaluation practice, incident learning, and baseline expectations with independently enforced controls and accountable authorization decisions. No single organization is likely to possess all of the information or control surfaces relevant to systemic advanced-AI risk.
XVIII. Advanced AI Safety Reference Architecture
The controls developed throughout this paper can now be assembled into a single reference architecture. The objective is not to prescribe a technology stack or product configuration. It is to define the safety functions that should remain present as AI systems become more capable, autonomous, and consequential.
The architecture follows a defense-in-depth principle: human control should not depend on any single model behavior, monitor, credential boundary, human reviewer, or shutdown mechanism working perfectly. Multiple safeguards should address different failure modes, and the most important control functions should remain outside the authority of the AI system they constrain.
The twelve defense layers
The first layer is model alignment and robustness. Training, oversight, interpretability, adversarial testing, specification design, and related methods seek to reduce the probability that the model generates harmful or unintended behavior. This is the behavioral foundation of the architecture, but it is not the sole safeguard.
The second layer is system-level capability and propensity evaluation. Evaluation examines the complete deployed system, including models, tools, memory, scaffolding, permissions, inference budgets, and operating environment. It asks what the system can accomplish and under which conditions concerning behavior may emerge.
The third layer is least privilege and authority limitation. Credentials, APIs, tools, transaction rights, compute, persistence, and consequential actions are restricted to the minimum required for the authorized mission. Increased capability does not automatically increase privilege.
The fourth layer is containment and segmentation. Sandboxes, isolated execution environments, network restrictions, API mediation, filesystem boundaries, resource quotas, and other controls limit how far unexpected behavior can propagate.
The fifth layer is model, credential, and infrastructure security. Model weights, training and deployment pipelines, secrets, dependencies, orchestration services, and compute infrastructure are protected against theft, modification, compromise, and unauthorized use. This layer helps establish that the production system is actually the system that was evaluated and authorized.
The sixth layer is resource, replication, and persistence control. AI workloads should not be able to independently create privileged successor instances, acquire substantial new compute, establish unauthorized persistence, create replacement identities, or move protected model assets beyond approved boundaries.
The seventh layer is an independent monitoring and control plane. Monitoring, policy decisions, authoritative logs, credential issuance, and enforcement remain sufficiently separate from the execution environment that the workload cannot simply redefine or disable its own constraints.
The eighth layer is meaningful human authorization. Selected high-consequence actions remain subject to qualified human review when human judgment can provide a genuine independent control. Human approval is not treated as sufficient when action speed, complexity, or consequence makes meaningful review impractical.
The ninth layer is external intervention and recovery. Independent mechanisms can reduce privileges, suspend tools, revoke credentials, isolate networks, terminate workloads, preserve evidence, and restore trusted operation. Intervention is graduated so that organizations can remove dangerous authority without unnecessarily sacrificing mission continuity.
The tenth layer is independent safety assurance and capability-authority gating. Increased autonomy or authority follows evidence from capability evaluation, threat analysis, safeguard testing, and residual-risk review. Systems do not receive greater authority merely because task performance improves.
The eleventh layer is continuous monitoring and reassessment. Changes to models, prompts, tools, permissions, memory, infrastructure, agent architecture, or operational context are evaluated when they materially affect capability, authority, or risk.
The twelfth layer is institutional and international governance. Standards, independent evaluation, incident reporting, procurement requirements, accountability mechanisms, and cross-organizational information sharing provide assurance beyond an individual deployment.
These layers should not be interpreted as twelve sequential barriers that every action must cross. They operate at different points in the lifecycle and address different failure modes. Their value comes from overlap without excessive common-mode dependence.
Execution plane and control plane
The most important structural feature of the architecture is the separation between the AI execution plane and the control plane. The execution plane contains the operational AI system: models, memory, agent loops, tools, task state, application logic, and approved interfaces. It performs the mission. The control plane governs what that system is allowed to do through identity, credentials, policy enforcement, monitoring, authoritative audit records, resource limits, isolation mechanisms, intervention controls, and selected authorization functions.
NIST SP 800-207 provides a useful architectural precedent through its separation of policy decisions and enforcement from the subjects requesting access. Advanced AI requires a broader application of the same idea: the system requesting an action should not control every mechanism that determines whether that action is permitted. An AI should not be able to modify its own access policy, issue itself broader credentials, erase authoritative evidence, disable the monitor responsible for detecting abnormal behavior, or prevent operators from isolating it.
Common-mode failure
Defense in depth becomes misleading when apparently separate safeguards depend on the same vulnerable component. If the same AI model performs an action, judges whether the action is safe, authorizes it, monitors execution, and determines whether intervention is required, the architecture may contain several logical functions while sharing one underlying failure mode. The same problem arises when one orchestration platform controls execution, identity, monitoring, logging, and shutdown.
The reference architecture therefore emphasizes independence of failure domains, not merely the number of controls. Different layers may use separate credentials, policy engines, software implementations, model families, administrative authorities, network boundaries, hardware controls, or human organizations where the consequence justifies that separation. Independence need not be absolute. The practical requirement is that one failure should not provide a straightforward path around every safeguard.
Figure 3 presents the reference architecture with the AI execution plane at the center, an independently governed control plane surrounding consequential access, and an outer assurance lifecycle covering evaluation, authorization, continuous reassessment, and institutional governance. The diagram makes the paper's central thesis visible: capability may exist inside the execution plane, but authority remains conditional, bounded, monitored, and externally enforceable.
The purpose of the architecture is not to guarantee that advanced AI can never fail. It is to prevent a model failure, security compromise, human error, monitoring gap, or governance failure from becoming sufficient by itself to produce catastrophic consequences.
XIX. Conclusion: Human Safety Must Not Depend on AI Cooperation
The central question in advanced AI safety is not whether future systems will inevitably become dangerous. Current evidence does not establish that conclusion. The practical engineering question is whether humans will preserve meaningful control if AI systems become substantially more capable, autonomous, and integrated into consequential environments.
This paper has argued that intelligence alone does not create catastrophic loss-of-control risk. The more concerning condition is the combination of capability, harmful propensity, operational opportunity, authority, and failure of the safeguards intended to constrain the system. That distinction creates an actionable strategy because each part of the pathway can be addressed independently.
The core principle is therefore simple: capability does not imply authority.
An AI system may possess sophisticated reasoning, planning, software-development, cyber, or scientific capabilities without receiving unrestricted credentials, broad network access, persistent execution, independent resource acquisition, or unilateral control over critical systems. Technical capability may advance quickly, but operational authority can remain conditional, bounded, monitored, and revocable.
That separation should be reinforced through defense in depth. Alignment can reduce the probability of harmful or unintended behavior. System-level evaluation can identify relevant capabilities and failure modes. Least privilege can restrict available resources. Containment can limit propagation. Supply-chain security can protect models and infrastructure. Independent monitoring can detect behavior outside approved boundaries. Human authorization and deterministic safeguards can protect high-consequence actions. External intervention can reduce or remove authority, while recovery mechanisms preserve the ability to restore trusted operation.
No individual layer should carry the entire safety burden.
This requirement becomes more important as AI is integrated into government, national-security, and critical-infrastructure missions. In these settings, failure may affect essential services, public safety, operational continuity, or physical infrastructure. AI can support those missions without receiving unrestricted control over them. The objective is bounded autonomy within an architecture that retains independent means of authorization, intervention, and recovery.
The distinction between the execution plane and control plane is particularly important. The system performing the mission should not possess unilateral authority over the mechanisms that constrain it. Credentials, policy enforcement, authoritative logging, monitoring, isolation, and intervention should remain sufficiently separate that the AI cannot simply redefine or disable its own boundaries.
The same principle applies institutionally. Evaluation should precede increases in authority, and authorization should remain conditional on evidence. Material changes to models, tools, permissions, memory, infrastructure, or mission context should trigger reassessment when they change what the system can accomplish or the consequences of failure. Incidents and new evidence should be capable of reducing authority just as successful evaluations may support expansion.
Advanced-AI safety will remain an area of scientific and engineering uncertainty. A durable safety strategy must therefore remain useful even when individual assumptions are wrong. Human safety should not rely solely on a future AI system voluntarily choosing to remain under human control, nor should a high-consequence organization rely on a single model safeguard, monitor, operator, or shutdown mechanism functioning perfectly.
The stronger engineering objective is to build systems in which catastrophic action would require the circumvention or failure of multiple independent technical, operational, human, and institutional safeguards. That objective does not require predicting exactly what future AI will become. It requires applying familiar principles from cybersecurity, safety engineering, resilient systems, and mission assurance: least privilege, separation of duties, independent enforcement, continuous verification, bounded authority, monitoring, recoverability, and defense in depth.
As AI capability increases, the defining measure of control should not be whether humans can ask a system to stop. It should be whether humans and human-governed systems retain the independent technical authority to limit what the AI can reach, determine what it may do, detect when it exceeds those boundaries, remove its authority when necessary, and restore trusted control without requiring its cooperation.
That is the architecture required to preserve meaningful human authority over increasingly capable AI.
Selected References
- International AI Safety Report. International AI Safety Report 2026.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF).
- National Institute of Standards and Technology. Zero Trust Architecture, NIST SP 800-207.
- National Institute of Standards and Technology. Secure Software Development Framework Practices for Generative AI and Dual-Use Foundation Models, NIST SP 800-218A.
- National Institute of Standards and Technology. Incident Response Recommendations and Considerations for Cybersecurity Risk Management, NIST SP 800-61 Revision 3.
- National Institute of Standards and Technology. High-Performance Computing Security Overlay, NIST SP 800-234.
- National Institute of Standards and Technology. AI RMF profile work for trustworthy use of AI in critical infrastructure.
- UK AI Security Institute. RepliBench research on autonomous replication capabilities.
- UK AI Security Institute. ControlArena and related AI control research.
- UK AI Security Institute. Research on AI safety cases, monitoring, and alignment evaluation.
- U.S. Department of Homeland Security. Framework for safe and secure AI deployment in critical infrastructure.
- Organisation for Economic Co-operation and Development. Work on interoperable AI incident reporting.