Read the proposed waiver in its proper scope
Senator Elissa Slotkin introduced S.4113 on March 17, 2026, when it was referred to the Senate Armed Services Committee. The discussion here concerns the introduced proposal; it should not be treated as an enacted acquisition requirement.
The bill’s text addresses three different areas: AI use to execute nuclear launch or detonation; monitoring or targeting people in the United States without the specified legal basis; and employment of lethal force by autonomous weapons without appropriate human judgment and supervision. Its waiver provision applies to the third category, not to the other two prohibitions.
The Secretary could issue the waiver without delegation for up to one year, or renew it for up to one year. The required certification would state both that extraordinary national-security circumstances require the waiver and that the probability of a result inconsistent with commander intent does not exceed the documented error rate of trained human operators performing equivalent functions under equivalent conditions.
The proposal also calls for congressional notification within five days of specified waiver events involving formal development, fielding, or substantial changes to algorithms, missions, environments, targets, or expected countermeasures. Its notification material includes realistic developmental and operational testing results, among other supporting information.
That scope changes the discussion. A debate about whether the waiver is too broad or the comparison too blunt is legitimate. It should begin with what the proposed text would actually require.
Equivalent conditions are the difficult part
A human comparison sounds simple until a test team must define it. A model score and an operator’s error rate are not comparable when they come from different tasks, inputs, or time constraints.
The evaluation plan should resolve:
- The function: detection, classification, recommendation, selection, or another defined action.
- The human baseline: training, experience, available information, tools, and workload.
- The operating conditions: sensor quality, communications, time pressure, clutter, and adversarial interference.
- The error categories: missed threats, incorrect identification, unintended actions, and departures from the commander’s intent.
- The statistical evidence: sample size, uncertainty, rare events, and performance across important subgroups or scenarios.
An aggregate accuracy score can conceal a dangerous failure mode. False positives and false negatives may have very different consequences, and neither is fully described by a single average. The test should make those consequences visible rather than selecting a metric merely because it produces an attractive comparison.
The introduced bill leaves substantial implementation work to the department. That creates an opportunity for industry to improve test design and evidence quality before a program reaches a consequential review.
Human supervision must work at the system’s speed
Human involvement can take several forms: approving an individual action, setting bounded operating conditions, monitoring execution, or retaining an effective ability to intervene. Which arrangement is appropriate depends on the system, mission, applicable policy, and operating environment.
A nominal override button does not establish meaningful supervision if the operator cannot understand the state, recognize a failure, or intervene in time. Equally, placing a human approval step in every low-level control action may be incompatible with some defensive functions. The engineering task is to connect authority with realistic information and response time.
The bill references the January 25, 2023 version of DoD Directive 3000.09. Subsequent policy work also matters: NSPM-11, issued in June 2026, directed updates to autonomy policy and standardized test, evaluation, verification, and validation methods. A direction to produce an update does not by itself establish that every implementing document has been issued.
Build the assurance evidence with the architecture
Verification becomes expensive when the system was not designed to expose the information needed for review. A stronger approach establishes the evidence chain alongside the software and interfaces.
- Document the intended behavior. Translate mission boundaries and command authority into testable conditions.
- Keep uncertainty visible. Preserve sensor quality and confidence information through the processing chain rather than hiding it behind a categorical answer.
- Exercise difficult cases. Include degraded inputs, spoofing or manipulation where relevant, unexpected combinations of conditions, and recovery from faults.
- Record the deployed configuration. Link results to software, model, sensor, and parameter versions.
- Define intervention and fallback. Test the human interface and the system’s response when confidence, communications, or equipment health deteriorate.
- Review material changes. Determine when an update invalidates earlier evidence and requires additional evaluation.
An explanation generated by a model can help a user, but it is not necessarily a faithful record of the model’s internal reasoning. Assurance should rest on observed behavior, traceable inputs and actions, bounded authority, and reproducible evaluation rather than persuasive explanations alone.
The policy discussion should improve the test
The proposed standard raises useful questions for legislators, acquisition officials, and engineers. How should equivalent conditions be defined? How should rare but severe failures be weighted? What changes should trigger renewed scrutiny? What evidence should Congress receive without exposing sensitive operational details?
Those questions remain relevant regardless of this bill’s legislative outcome. A vendor with a documented performance envelope, realistic tests, and a credible change-control process gives a customer more reason to trust its system. The objective is an autonomy capability whose authority is supported by evidence appropriate to the consequences of its use.
Sources and further reading
- GovInfo: S.4113 introduction and referral
- AI Guardrails Act of 2026, introduced text
- White House: NSPM-11, June 2026
Spartan X’s AI assurance and systems engineering capabilities connect policy expectations to evidence a program can use. Clear operating boundaries, representative tests, and accountable release decisions make autonomy more credible through development, review, and operational change.



