What a demonstration leaves unanswered
A software RFP can evaluate functional requirements, cost, delivery schedule, vendor stability and security posture. Keep those checks. Add evidence about the model, its data and its intended operating conditions. AI systems can produce repeatable outputs: a fitted decision tree, for example, applies learned decision rules to an input. Conventional software can also behave differently when dependencies, execution order or inputs change. Repeatability alone does not establish that either system is correct or that a demonstration covers the agency's workload.
Three evaluation questions deserve particular attention. The NIST AI RMF Playbook's Measure function recommends context-sensitive evaluation, assessment of whether results generalize, and monitoring of performance over time:
Distribution shift. A model may perform well in testing and less reliably when the operating environment differs: applicants submit different documents, a program rule changes, or a new region uses different record formats. A vendor may already have tested these variations; the procurement team should examine that evidence instead of assuming it exists or is absent. Ask how training, validation and final evaluation datasets were separated, which conditions were tested, and how those conditions compare with the state's current caseload. Measure 1.2 specifically calls for examining whether measurements generalize beyond the context in which they were obtained.
Performance gaps hidden by averages. An overall accuracy score can hide concentrated errors in a subset of cases. As a hypothetical example, a document-processing model might handle standard payroll records reliably but make more errors extracting income from self-employment records. That is a documentation difference, not a demographic characteristic. Testing supported languages, document types and operating conditions separately can reveal where additional validation or human review is needed. Demographic breakdowns may also help evaluate performance for the people an agency serves, where relevant to the use case and supported by appropriate data. A measured difference does not, by itself, establish its cause or justify different eligibility standards. The NIST AI RMF Playbook, Measure 1.3 explains why evaluations should look beyond averages for concentrated failures.
Performance changes over time. A deployed model can remain unchanged while its inputs or operating environment change. As hypothetical examples, a fraud-detection model could encounter different claims patterns after a disruption, or a document-processing model could encounter new fields after a rule change. The procurement question is whether the contract requires useful monitoring, reporting and corrective action when performance falls outside agreed limits. Check the actual terms; do not assume that monitoring is included or absent. Measure 2.4 addresses monitoring deployed AI systems in their production context.
What NIST's AI Risk Management Framework gives procurement teams
The NIST AI Risk Management Framework 1.0 (January 2023) organizes AI risk management around four functions: Govern, Map, Measure and Manage. NIST describes it as voluntary, non-sector-specific and use-case-agnostic. It supplies a structure for managing risk, rather than a legal certification or a mandatory checklist for every purchase. A state procurement team can use that structure to develop evaluation criteria appropriate to its program.
The RMF's most useful contribution to procurement is its treatment of the full deployment context: who the AI system affects, what happens when it produces an error, and how harm can be detected and corrected. That framing translates directly into RFP questions that vendors either can or cannot answer.
Four areas to tailor to the proposed system's role, consequences and procurement scope:
1. Validation documentation with useful breakdowns. Request the vendor's evaluation methods, results and limitations. Choose metrics that fit the task: precision and recall may help evaluate a classifier, while a document-extraction tool needs measures of field-level correctness and omissions. Specify relevant breakdowns by document type, supported language, operating condition and, where justified, demographic characteristics. Record sample sizes and uncertainty so a small subgroup is not mistaken for reliable evidence. Missing evidence creates a question to resolve through testing, additional documentation or a narrower deployment scope; it does not establish that every product requires the same demographic report. Measure 1.1 calls for methods appropriate to the risks being assessed.
2. Data lineage and scope. Request documentation of training and evaluation data sources, covered periods, relevant populations and known gaps, subject to applicable confidentiality and privacy constraints. For benefits eligibility and program-integrity tools, ask whether the evaluation represents current rules and case documentation. A different training period warrants investigation; it does not prove a particular error pattern. If a supplier cannot disclose underlying data, assess what independent testing, controlled access or other evidence can substantiate its performance claims, and record the remaining uncertainty.
3. Monitoring and drift detection plan. Require a written post-deployment plan covering how the vendor will detect model drift, what performance metrics will be reported to the state and on what cadence, and what thresholds trigger a mandatory revalidation event. This should be a contract requirement enforceable at renewal, not a verbal assurance.
4. Review, correction and appeals integration. For a system that influences benefits determinations, map its role to the program's actual notice, review, appeal and recordkeeping requirements. Specify who can correct an error, what authority they have, and what evidence they can inspect. Ask the vendor to demonstrate how its output enters the case record and supports agency review. These are procurement recommendations that must be tailored to the workflow; NIST does not impose a universal legal right to human override for every AI-assisted action.
The explainability requirement
Identify the notice and appeal rules for the specific program before writing an explanation requirement. For example, 42 CFR 431.210 specifies content for Medicaid notices, including the intended action, reasons, supporting regulations or a change in law, hearing rights and circumstances in which benefits continue pending a hearing. That program-specific rule should not be presented as a uniform model-explanation requirement across all state services.
For a benefits workflow, require a demonstration showing how the agency produces an accurate, understandable account of its decision and supports any required notice. A model-generated explanation alone does not establish compliance. The agency must connect the relevant case evidence, applicable program rules and decision process, with counsel confirming the requirements for that program.
"The algorithm determined..." does not tell a caseworker why an outcome is correct. A useful acceptance exercise should show: (a) which case facts and program rules support the decision, (b) how a reviewer can identify and correct an erroneous input or output, and (c) how the result integrates into the required notice. Internal model weights may not provide that explanation, and the cited Medicaid provision does not require their disclosure. Evaluate whether the complete workflow supports the agency's obligations before approving its intended use.
Pre-award technical evaluation checklist
Before shortlisting an AI vendor, a state procurement team should be able to answer these questions from vendor-provided documentation:
- Does the evaluation use task-appropriate metrics and relevant breakdowns, with sample sizes and limitations, to test the state's actual workloads?
- What is the training data time range and source, and how comparable is it to the state's current caseload?
- What is the vendor's post-deployment monitoring process, and what performance reporting does the state contractually receive?
- How does the system generate an explanation for an adverse determination, and what does that explanation include?
- What contractual provisions govern model updates — specifically, what triggers a notification requirement, and does the state retain the right to retest before a major model revision goes live?
- Who owns the model output records, and what data retention obligations does the vendor assume?
Use the answers to identify what is demonstrated, what needs an acceptance test, and what remains uncertain. An unanswered question can reflect a gap in the product, its documentation or the requested scope. Resolve consequential gaps before accepting the corresponding operational risk.
Red flags in vendor responses
Three response patterns that should pause evaluation before a vendor is shortlisted:
Proprietary system with no usable evaluation evidence. A supplier can protect intellectual property while providing meaningful evidence of performance. Seek validation methods and results, controlled testing or independent assessment appropriate to the purchase. If no usable evidence is available, the agency lacks a defensible basis for accepting the claimed model quality. Document the gap and its consequences instead of treating a proprietary architecture itself as proof of poor performance.
Aggregate accuracy without workload detail. An overall number is insufficient when it leaves a material risk in the intended workflow untested. Ask for the relevant breakdowns and enough information to interpret them. For a document-processing tool, that might include document formats and missing-field errors; for a resident-facing system, supported languages and relevant population characteristics may also matter. The required evidence follows the use case, rather than a single checklist imposed on all government AI.
Updates without agreed change controls. If model changes can affect consequential outputs, seek terms defining notification, impact assessment, testing, rollback and agency acceptance where warranted. Routine fixes and material behavioral changes may need different treatment. Establish that distinction before award so the agency and supplier know when additional validation is required.
The RFP language gap
An RFP is an opportunity to make evaluation expectations explicit before proposals arrive. A supplier may provide useful evidence without being asked, and requirements may also appear in other contract documents. The agency should nevertheless check that the complete procurement establishes the evidence, monitoring reports and remedies it needs. Leaving those matters unresolved creates avoidable uncertainty at acceptance and during operation.
The NIST AI RMF Playbook — the companion implementation guide — translates the RMF's functions into specific actions and recommended activities. State procurement teams don't need to develop AI evaluation criteria from first principles. They need to translate existing federal frameworks into the RFP language their contracting offices will act on.
One practical option is to designate a qualified evaluator to examine AI performance evidence separately from the functional demonstration, consistent with the agency's procurement rules and available expertise. Measure 1.3 recommends involving internal experts who were not the system's front-line developers or independent assessors in evaluation. Give that reviewer a defined scope, access to evidence and responsibility for documenting unresolved risks. This strengthens the technical basis for the decision without assuming every agency needs the same committee structure.
Sources and further reading
- NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0)," NIST AI 100-1, January 2023: doi.org/10.6028/NIST.AI.100-1
- NIST AI RMF Playbook, AI Risk Management Framework companion: NIST AI RMF Playbook
- Medicaid notice content requirements, 42 CFR 431.210: ecfr.gov/current/title-42/section-431.210
Spartan X's AI advisory, evaluation and governance capabilities can help an agency connect its mission needs with the technical questions a vendor should answer. That expertise can support a clearer evaluation before award or a focused assessment of an existing system.



