Four Years of the Software Acquisition Pathway: What AI Programs Are Learning
Back to Signal
AIDefenseGovernmentInnovation

Four Years of the Software Acquisition Pathway: What AI Programs Are Learning

October 5, 2026Jess Loban

What the Software Acquisition Pathway Actually Changes

The SWP replaced traditional Milestone A–B–C reviews with an iterative structure centered on a Mission Thread Workshop, a User Agreement, and a Capability Roadmap updated at each delivery cycle. The pathway supports programs using Continuous Delivery — releasing software directly into a live operational environment in frequent, small increments — and programs using periodic major releases, where defined capability increments are delivered on a scheduled cycle. MOSA (Modular Open Systems Approach) is a separate architectural mandate applied across DoD acquisition pathways, requiring open and modular interfaces to reduce vendor lock-in; it is a design requirement, not a delivery cadence. DAU Software Acquisition Pathway Guide

For a traditional software program — a logistics management system, a command-and-control interface, a communications platform — this structure solves a real problem. The previous model required that requirements be fully specified before development began, which was structurally impossible for software whose functionality depends on user feedback. The SWP acknowledges that users often cannot articulate requirements until they see working software, and that good development depends on iterating with actual users in operational environments.

For AI programs, the pathway's benefits are real but different. AI programs are not simply software programs that happen to use machine learning — they are systems whose behavior emerges from training data and model architecture, and whose performance can degrade in ways that traditional software testing does not detect. The SWP removes milestone delays that would prevent rapid model updates, which is valuable. But it does not define how program managers should govern model changes, validate that updates preserve intended behavior, or maintain the chain of accountability when a model's behavior shifts between delivery cycles.

Four Years of Use: What the Record Shows

The Government Accountability Office has published multiple oversight reviews of DoD software acquisition programs that found consistent gaps between what the SWP requires and what programs actually document: specifically, completing the Mission Thread Workshop, maintaining an updated User Agreement, and executing operational testing before each delivery. GAO defense acquisition oversight reports The finding was not that the pathway failed — it was that programs were skipping the governance steps the pathway requires while retaining its flexibility to deliver quickly. That combination produces programs with rapid output and weak oversight, which is the worst outcome the SWP was designed to prevent.

For AI-specific programs in the SWP, the oversight problems are compounded. The pathway's minimum governance artifacts — Mission Thread Workshop, User Agreement, Capability Roadmap — were designed for deterministic software. They do not address several questions specific to AI systems:

  • Model versioning and accountability. When a machine-learning model is updated between delivery cycles, does the program have a record that links each deployed model version to its training data, evaluation results, and approved scope? Without this, the User Agreement signed against a previous model version may not describe the system actually in operational use.
  • Evaluation currency. Software testing verifies that a function produces the correct output for a given input. AI evaluation is a different activity: it assesses model performance across a distribution of inputs, under conditions that may shift over time. A model that passed evaluation at delivery cycle 3 may perform differently by delivery cycle 6 if the operational environment has changed and the model has not been updated or re-evaluated.
  • Drift accountability. Machine-learning models can degrade in production when the distribution of real-world inputs diverges from the training distribution. The SWP's operational testing checkpoints can detect this if they include population-representative test sets and performance baselines — but the pathway does not mandate those requirements. Programs that define their operational testing only in terms of functional correctness will miss model drift until it affects mission outcomes.
  • Data governance. AI model performance is directly dependent on training data quality, provenance, and access controls. The SWP does not include a data governance requirement. Programs that treat data management as a detail to be resolved within the development team rather than as a formal program artifact lack the documentation needed to explain model behavior to oversight bodies or to recover from a data-quality incident.

What Good SWP AI Governance Looks Like

The programs that have used the SWP most effectively for AI have done several things that the pathway does not require but that the pathway's structure makes possible:

Model register embedded in the Capability Roadmap. Rather than treating model updates as implementation details, the best programs track each deployed model version in the Capability Roadmap as a formal program artifact. The register includes the training data version, evaluation results, approved scope, and any known limitations. This creates the accountability chain that formal milestone reviews would have provided, without the delay.

Evaluation thresholds written into the User Agreement. The User Agreement is the SWP's primary governance document — it specifies what the system must do, under what conditions, and for whom. Programs that write evaluation thresholds into the User Agreement (accuracy floor, false-positive ceiling, operational condition bounds) give themselves an objective criterion for determining whether a model update constitutes a material change requiring additional oversight. Programs that do not write these thresholds in the User Agreement leave the evaluation criteria implicit, which means they also leave the decision about when to seek additional oversight implicit.

Operational data sampling as a formal delivery artifact. Several mature SWP AI programs have established that a representative sample of recent operational data is a required input to each delivery cycle — not an optional post-delivery activity. This turns operational testing from a periodic event into an ongoing program function, which is the only architecture that can reliably detect model drift before it affects mission outcomes.

Designated AI assurance officer at the program level. The SWP's governance model assumes that the program manager has the organizational capacity to execute Mission Thread Workshops, maintain the Capability Roadmap, and conduct operational testing. For AI programs, this capacity needs to include someone with authority and expertise to evaluate model governance artifacts — training data provenance, evaluation methodology, scope limitations. Programs that assign this function to a software engineer without a formal program role produce inconsistent governance quality and inconsistent accountability when oversight bodies ask questions.

The Oversight Gap the Pathway Leaves Open

The SWP's Continuous Delivery model allows programs to release software updates without a formal review checkpoint — the review is embedded in the User Agreement and operational testing process. This is appropriate for deterministic software where a release either does or does not meet the specified functional requirements. For AI programs, it creates a specific oversight gap: updates that change model behavior in ways that are within the technical scope of the User Agreement but outside the intent of the operators can ship without triggering a formal review.

A hypothetical example illustrates the risk: a medical logistics AI program uses the SWP Continuous Delivery model to ship monthly model updates. The User Agreement specifies that the system should prioritize supply requests based on urgency and availability. An update at delivery cycle 8 shifts the model's prioritization logic in a way that is technically within the specified criteria but that produces different outcomes for edge-case supply categories that operators depended on. The update passes functional testing. No operator complaint surfaces for six weeks, at which point a supply shortage is traced to the model's behavior. The program has no formal record of what changed between cycles 7 and 8 that would allow reconstruction of the decision.

This scenario is not hypothetical in structure — the general pattern of iterative updates producing behavioral regressions not caught by functional testing has been documented in commercial AI deployments across healthcare, financial services, and recommendation systems, though the specific causal paths vary by domain. The SWP does not prevent this scenario because it does not require model-level versioning, behavioral baselines, or regression testing distinct from functional correctness. Programs that want to prevent this scenario need to build those requirements into their User Agreement and delivery processes explicitly.

The Credible Alternative: Hybrid Pathway

For AI programs where the iterative delivery benefits of the SWP are real but the oversight gaps are significant, a hybrid approach that uses the SWP's delivery structure with mandatory AI governance overlays is more defensible than either the traditional milestone path (too slow for effective AI development) or the unmodified SWP (insufficient oversight for systems with emergent behavior).

The hybrid approach uses the Mission Thread Workshop and User Agreement as the foundation but adds three AI-specific requirements:

  1. A model register as a formal program artifact, maintained at each delivery cycle.
  2. Behavioral baselines written into the User Agreement, with a formal review trigger when a model update produces baseline deviation above a specified threshold.
  3. A data governance plan as a required program document, maintained by a designated responsible official.

This does not require additional milestone reviews or statutory authority changes. It requires that program managers treat AI governance artifacts with the same formality they apply to the software architecture documentation the SWP already requires.

Questions for Program Managers Entering the SWP with an AI Program

  1. Does your User Agreement specify evaluation thresholds for the AI component, or only functional requirements for the system? If only functional requirements, model behavior changes between delivery cycles have no formal trigger for additional oversight.
  1. Does your Capability Roadmap include model versioning, or does it describe only software features? If only features, you cannot reconstruct the chain from a specific operational outcome to the model and data that produced it.
  1. What is your definition of a material model change that requires a new evaluation cycle? If this is not written into a program document, the decision will be made informally and inconsistently.
  1. Who in your program has designated responsibility for AI governance artifacts? If this function is distributed informally across the development team, you will produce inconsistent documentation quality and have no single point of accountability.
  1. What is your operational testing methodology for AI components? If your operational testing only evaluates functional correctness, you will not detect model drift between delivery cycles.

The SWP is a better framework than the traditional milestone path for most AI programs. Getting the most from it requires building the governance elements that the pathway does not mandate — starting with the first Mission Thread Workshop, not the last delivery cycle before a major review.

Spartan X's AI engineering and program management work directly addresses the governance gap between rapid AI delivery and accountable program oversight: model versioning architecture, evaluation methodology, and the documentation structure that makes AI behavior explainable to oversight bodies and recoverable after unexpected performance changes.

Sources and further reading

Share this article
LinkedIn

BUILD WITH US

Ready to Solve Hard Problems?

Spartan X builds AI systems, autonomous platforms, and cybersecurity solutions for defense and national security.