What the Army Set Out to Build
The Integrated Visual Augmentation System is a wearable heads-up display built on Microsoft's HoloLens platform, adapted for military use under a development contract awarded to Microsoft in late 2018. The program's premise was operationally sound: soldiers navigating complex terrain, managing fires coordination, and integrating live sensor data could benefit from a display that layered digital information — map overlays, target data, shared operational picture — directly on their field of view, without requiring them to look away from the environment.
AI integration was built into the program's design intent from the start. Automatic target recognition, thermal contrast enhancement, friend-or-foe overlay drawn from the command network, and machine-assisted navigation are all functions where AI processing adds value over purely manual methods. IVAS was conceived as a platform that would grow more capable as AI capabilities matured — the hardware and interface designed to receive software updates carrying new models and new mission applications.
The strategic case was real and recognized. Program advocates and Army officials described IVAS's potential in terms comparable to what GPS did for navigation — a foundational shift in individual situational awareness — though operational testing later surfaced the limits of that framing.
What the Tests Found
Initial operational testing conducted by the Army Test and Evaluation Command surfaced a set of problems that did not map cleanly onto the technical performance categories the program had been tracking. Soldiers reported eye strain and disorientation during extended wear. The headset's optical field of view was narrower than the effective visual field that soldiers relied on in close-combat maneuvers. Battery performance fell short of the duration required for sustained missions. The weight and profile of the device interfered with standard protective equipment and was difficult to use in prone firing positions.
The Government Accountability Office's annual weapons systems assessments documented IVAS among the programs with significant developmental risk and noted the human factors findings as schedule drivers. The Army Acquisition Executive acknowledged the findings and directed development of a revised configuration — IVAS 1.2 — to address the optical system, thermal management, battery, and physical form factor before authorizing production quantities.
These were not, in most cases, failures to meet a specification requirement. The system largely performed to contract as written. They were operational environment failures: the specifications had not adequately captured the demands of sustained wear in the conditions the program was fielding into.
Three Structural Causes
The IVAS experience is consistent with how human factors problems emerge in defense AI programs. Three structural causes recur:
Requirements are written against the design case, not the operational population. A requirement stating "field of view: X degrees" describes average performance on a test apparatus. A requirement stating "maintain effective visual integration at X degrees while wearing an ACH with the issued pad system, at the 95th percentile head circumference, with gloves" describes the operational population. The latter is harder to write and harder to test. In competitive development programs where contractors bid against technical requirements, the operational-population specificity that human factors engineers can provide is often absent from the original requirements document because it requires user population data that programs frequently don't invest in during pre-acquisition.
Human factors testing is scheduled late in development. The defense acquisition framework supports human factors engineering throughout the program, but the actual testing that surfaces operational failures — formative evaluation with representative users in representative conditions — is typically funded and scheduled as a late-development or production-entry activity. By that point, design decisions have been made, contracts have been awarded, and changes are expensive. The human factors findings from initial operational testing then do what IVAS demonstrated: they arrive at a point where the choices are accept a substandard system, terminate, or redesign.
AI subsystem performance is evaluated in laboratory conditions. When programs test AI components, they test against curated data sets in controlled environments. Thermal contrast enhancement might perform at 95% accuracy on a test set; it performs differently when the sensor is warm from body heat, when the soldier is moving, when the optical pathway is slightly degraded by sweat on the lens. The performance envelope that matters operationally is not the performance against a clean test set. Specifications that define AI accuracy without defining the operating conditions — temperature range, motion state, sensor fouling, degraded optics — allow laboratory pass rates that do not translate to the field.
What Buyers Should Require
The IVAS record is specific enough to extract concrete acquisition guidance:
Require a human factors test plan at Milestone B. Before a development contract is authorized, the program should have a human factors test plan that identifies the user population, the representative tasks, the full environmental conditions (including worst-case and 95th-percentile conditions), and the acceptance criteria. The plan should be produced or reviewed by a human factors engineer with defense-system experience. The test plan is inspectable evidence that the program knows where its human factors risks are.
Require sustained-use data before production quantities. Sustained use means the tasks and durations that characterize the actual mission: not a 30-minute laboratory session but the duration of a notional patrol, a mounted operation, a sustained defensive position. For any system that is worn, carried, or continuously operated, sustained-use testing with representative users in representative conditions is the primary mechanism for discovering the battery life, weight, fit, and fatigue failures that laboratory testing misses.
Budget human factors engineering as a named program line. When human factors engineering is part of a broad "systems engineering" line, it is invisible to oversight and easily absorbed when schedule and cost pressure arrive. A named line item with its own budget and milestone creates accountability. The investment is recoverable during development; the consequences of deferring it are not.
Write AI performance requirements against operational conditions. For any AI-enabled component, define performance against the conditions in which the system will actually operate: temperature range, motion state, sensor degradation modes, communication latency, data quality. A specification that only addresses clean-input performance is a specification that will not catch the failures that matter in the field.
Require failure mode documentation before production authorization. For AI components, require the program to produce and defend documentation of the operating conditions under which the AI model degrades or fails — the input conditions, data conditions, and environmental conditions that take the system outside its effective envelope. This documentation is evidence that the program has tested the edges of the performance space. Its absence is evidence that the program has not.
Evidence That Would Change This Analysis
This guidance assumes the IVAS pattern generalizes: that programs systematically underinvest in early human factors work, and that this produces late-program failures across AI-enabled system categories. That assumption has limits.
Programs with a strong human factors tradition — aviation crew system development, for example — have invested in operational environment testing for decades and have generally found problems earlier. The argument here is strongest for systems in categories where human factors engineering has historically been treated as a secondary concern: dismounted soldier systems, networked data systems, AI-enabled analytical tools. It is less urgent advice for programs already running mature human factors processes.
The IVAS 1.2 fielding experience is itself the relevant evidence update. If 1.2 sustains operational testing without significant human factors findings, it demonstrates that redesign can successfully recover a program — which strengthens the case for continuing even poorly-designed programs rather than terminating. If 1.2 finds its own human factors gaps, the argument for early investment becomes more urgent. Program managers should track 1.2's operational test results as this record develops.
The Next Generation of Systems
IVAS is one program. The pipeline behind it — AI-enabled decision aids, edge computing at the squad, autonomous systems in the logistics network, targeting AI integrated with crew-served weapons — represents dozens of programs that share the same structural risk. Each faces the same choice: fund human factors engineering early and discover operational problems in formative testing, or defer it and discover them at production entry.
The record suggests that deferral is more expensive. The question for acquisition professionals evaluating these programs is not whether to require human factors evidence but how to make that requirement specific enough to be useful. The IVAS test plan, requirements specification, and operational test reporting are the working documents from which that specificity can be drawn.
Sources and further reading
- U.S. Army, Program Executive Office Soldier — IVAS: https://peosoldier.army.mil/equipment/networking/ivas/
- Government Accountability Office, Weapon Systems Annual Assessments (series): https://www.gao.gov/weapons-systems-annual-assessment
- Department of Defense, *DoD Instruction 5000.02: Operation of the Adaptive Acquisition Framework*, January 23, 2020: https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/500002p.pdf
The AI accuracy dimension of the IVAS gap — AI subsystems tested against curated data in controlled conditions, with performance assumptions that did not carry to the operational environment — is where Spartan X's Arbiter practice applies: verifying AI subsystem outputs through multi-model consensus and adversarial review against the conditions that matter, not only the controlled conditions where models were trained. That AI assurance work is the evidence base that program managers need before production authorization, rather than after operational testing finds the gap.



