An AI demonstration is optimized for the happy path: curated data, a cooperative prompt, a responsive model, and a presenter who knows how to recover. Procurement owns the ordinary path, the failure path, the change path, and the exit path. The object being purchased is not a model in isolation. It is a changing service assembled from models, prompts, retrieval systems, data stores, safety controls, people, subprocessors, and contract terms.
The procurement dossier should therefore connect claims to evidence and evidence to enforceable obligations. A security questionnaire without workflow testing is incomplete. A benchmark without data terms is incomplete. A favorable pilot without a migration plan is incomplete.
Define the Use Case Before Comparing Vendors
Begin with the decision or workflow, not a feature list. Write down:
- who uses the system and who is affected by its output;
- what input data it receives and how that data is classified;
- which systems it can read from or write to;
- whether it recommends, drafts, decides, or acts;
- what a correct outcome looks like;
- what error is acceptable and what error is intolerable;
- who reviews or overrides it;
- and what business process continues when it is unavailable.
This scope determines diligence. A public copywriting assistant does not need the same evidence as an agent that can send payments, rank job candidates, modify production code, or process health records.
Use risk tiers rather than a single approved-vendor label. A provider may be acceptable for public information and prohibited for confidential source code. A model may be approved for drafting and not for autonomous action. Link approval to the exact service, region, data mode, integration, and use case; otherwise a low-risk pilot can become an undeclared high-risk deployment.
For a broader sourcing view, see AI in strategic procurement. The dossier described here focuses on acquiring AI itself.
Map the Real Service and Its Value Chain
Ask for an architecture and data-flow diagram that matches the proposed configuration. Identify:
- the contracting entity and service operator;
- base-model providers and model versions;
- cloud, observability, moderation, retrieval, and support subprocessors;
- customer-managed and vendor-managed data stores;
- fine-tuning, feedback, caching, abuse-monitoring, and logging paths;
- regions used for storage, inference, backup, and support access;
- open-source models, libraries, adapters, and data dependencies;
- and every external tool the system can invoke.
Names alone are not enough. Record purpose, data received, retention, location, replacement notice, and assurance evidence for each dependency. A vendor that calls another model API may inherit that provider’s outage, retention, safety, and change behavior.
The UK National Cyber Security Centre’s secure AI development guidance recommends assessing and monitoring the AI supply chain across the lifecycle, documenting models, data, prompts, and failure modes, and preparing failover for mission-critical systems. That is useful procurement logic even when the buyer is not building the model.
Inventory prompts, evaluation sets, user feedback, generated content, and logs as organizational assets. Clarify ownership and export rights. Logs may contain source documents, model outputs, tool results, personal data, credentials accidentally pasted by users, or sensitive reasoning about business decisions. “We retain logs for reliability” is not a complete control description.
Build an Evidence Ladder
Procurement teams often collect many documents without distinguishing their strength. Use an evidence ladder:
- Marketing statement: useful for discovery, not assurance.
- Policy or product documentation: specific enough to establish intended behavior, but generally controlled by the vendor.
- Completed control response: maps a claim to owner, implementation, scope, and exception.
- Independent attestation or audit: stronger when the report scope, period, service, exclusions, and exceptions match the purchase.
- Buyer-observed test: demonstrates behavior in the proposed configuration with representative inputs.
- Contractual commitment: defines notification, remedy, audit, deletion, and exit rights.
- Ongoing production evidence: proves the control continues after purchase and through model changes.
No single rung replaces the others. A SOC report may cover the cloud environment but not hallucination, prompt injection, training-data rights, or a new model release. A model card may describe a benchmark but not the buyer’s workflow. A successful test today does not create a notification duty tomorrow.
The NIST Generative AI Profile specifically recommends use-case-based supplier assessment, third-party inventories, contingency planning, and procurement diligence covering intellectual property, privacy, security, and value-chain risk. NIST’s profile is voluntary guidance, not a certification or warranty; use it to organize questions and evidence rather than as a pass badge.
Examine Data, Security, and Permissions
For every data class, request plain answers to these questions:
- Is customer content used to train, fine-tune, evaluate, or improve any shared model?
- Is opt-out the default for the contracted service, and does it cover subprocessors?
- What is retained in prompts, files, outputs, embeddings, caches, backups, abuse logs, and support systems?
- Can the buyer configure retention, legal hold, deletion, and regional processing?
- How is deletion verified, including derived stores and backups?
- Who can access content, under which approval, and how is access logged and reviewed?
- What encryption, key management, tenant isolation, and incident-response controls apply?
- How are prompt injection, data exfiltration, insecure output handling, and tool abuse tested?
- What is the vulnerability disclosure process and patch expectation?
An agent adds a second layer: authority. List each tool, identity, scope, credential, action, approval gate, rate limit, and rollback. Apply least privilege and short-lived credentials. The companion guide on AI tool-permission security explains why “can use the finance system” is too broad to be a meaningful control.
Test permissions yourself. Attempt cross-tenant access, unauthorized tool calls, prompt injection through retrieved content, and actions beyond the user’s role. Use synthetic or approved test data; a security evaluation should not create a privacy incident.
Demand Evaluation and Change-Control Evidence
Ask the vendor to define intended use, known limitations, prohibited use, and performance by relevant language, document type, user group, and difficulty. Request:
- evaluation-set construction and contamination controls;
- sample size, uncertainty, failure taxonomy, and subgroup results;
- red-team scope and unresolved findings;
- human-review procedures and reviewer qualifications;
- release gates for model, prompt, retrieval, and safety-control changes;
- rollback ability and version pinning;
- customer notification before material behavior changes;
- and production monitoring for drift, abuse, and incidents.
Do not accept a single aggregate accuracy score. A 95-percent average may conceal a catastrophic authorization failure or poor performance in Persian. Require severe errors to be reported separately.
Evaluate the whole configured system, not only the base model. Retrieval quality, system prompts, permissions, document parsing, tool schemas, and user interface often determine the outcome. Re-run the buyer’s protected evaluation after material changes. If the vendor cannot pin a version, negotiate a notice period and a right to test before migration.
Run a Pilot Designed to Fail Informatively
A pilot is an evidence-producing experiment, not free-form exploration. Pre-register:
- baseline process and current outcome;
- representative workload and protected holdout cases;
- success, pause, and stop criteria;
- severe errors that trigger immediate review;
- human-review design;
- maximum data sensitivity and tool authority;
- latency, volume, and cost envelope;
- and evidence required for expansion.
Include ordinary, boundary, adversarial, and unavailable-service cases. Test long documents, conflicting evidence, unsupported languages, ambiguous instructions, malicious retrieved content, missing permissions, model refusal, and downstream timeout.
Compare against the real baseline, including human time and correction. Measure end-to-end outcome, not how impressive the answer sounds. A summarizer that is 20 percent faster but doubles verification time has not delivered the claimed gain.
Keep the pilot reversible. Use isolated credentials, narrow data, a capped budget, and no irreversible autonomous actions. Expansion should be a new decision based on the dossier, not the default consequence of a popular demo.
Model Total Cost and Concentration Risk
Seat or token price is only one line. Estimate:
total cost = subscription + usage + integration + data preparation + retrieval + security + observability + human review + support + failures + change + exit
Run billing scenarios with realistic context length, reasoning tokens, retries, tool calls, embeddings, peak concurrency, storage, network egress, and support tier. Ask whether failed calls, cached tokens, batch work, and provider-initiated retries are billed.
Estimate cost per valid business outcome, not per generated token. Include correction, escalation, quality incidents, manual fallback, and customer remediation. Report ranges because model use is often heavy-tailed.
Concentration risk also has a price. Identify workflows that would stop if one model, region, identity provider, or vendor failed. Test export and fallback rather than relying on an architecture slide. The fallback does not need feature parity; it needs to preserve the essential function safely.
Put Operational Rights in the Contract
High-value questions become useful only when the agreement provides a response. Depending on risk and jurisdiction, address:
- permitted data use and an explicit prohibition on training shared models with customer content;
- subprocessor list and advance notice of material changes;
- security control schedule and vulnerability notification;
- incident definition, notification clock, investigation support, and evidence preservation;
- service levels for availability, latency, support, and recovery;
- evaluation, audit, or assurance rights proportionate to risk;
- model and control change notice, version options, and regression remedies;
- ownership and licenses for inputs, outputs, fine-tunes, prompts, and evaluation artifacts;
- intellectual-property claims and indemnity allocation;
- compliance assistance and documentation duties;
- export format, transition support, deletion deadline, and deletion evidence;
- termination rights for control failure or unacceptable model change;
- and limits of liability that match credible loss scenarios.
Contract language must be reviewed by qualified legal and privacy professionals. Requirements vary by country, sector, data type, and role.
The EU AI Act illustrates why roles matter. Article 25 of Regulation (EU) 2024/1689 addresses responsibilities along the AI value chain and circumstances in which a distributor, importer, deployer, or other third party can be treated as the provider of a high-risk system—for example after certain substantial modifications. Whether a particular purchase is in scope or high-risk is a legal classification question, not something a generic procurement checklist can decide.
Design the Exit Before Signing
Ask the vendor to demonstrate export and deletion during diligence. Identify which items can leave:
- source documents and structured records;
- prompts, system instructions, and templates;
- embeddings or a reproducible way to rebuild them;
- fine-tuning data and, where negotiated, weights or adapters;
- evaluations, labels, red-team cases, and score history;
- logs, traces, incidents, and audit evidence;
- generated assets and associated provenance;
- workflow definitions, tool schemas, and configuration.
Specify usable formats, rate limits, fees, transition period, technical assistance, and how identity and permissions are revoked. Verify deletion across primary storage, caches, derived data, support copies, and backups under the agreed schedule.
Portability is architectural as well as contractual. Keep business rules, evaluation cases, and authoritative data outside vendor-only prompt consoles where possible. Use an adapter boundary for model calls and tools. Avoid depending on undocumented behavior. Test a second provider or manual fallback before leverage is lost.
Use a Decision Scorecard with Gates
Scorecards help comparison only after hard gates are separated. A vendor that fails a non-negotiable data-use or security condition should not compensate with better demo quality.
Possible weighted categories include:
- workflow quality and severe-error rate;
- privacy and data governance;
- security and tool control;
- transparency and evaluation quality;
- operations, support, and resilience;
- legal and regulatory fit;
- total cost;
- portability and exit;
- and vendor viability.
For each score, cite an artifact and name its date, scope, and owner. Mark unknowns explicitly. “No evidence” should not become a neutral middle score.
The final decision should state approved use, prohibited use, conditions, residual risk owner, monitoring metrics, review date, and triggers for reassessment. Purchase is the start of vendor risk management, not its conclusion.
Frequently Asked Questions
Is an independent audit enough?
No. Check whether the audit covers the exact legal entity, service, region, period, and controls. It rarely answers all model-quality, data-rights, or workflow-specific questions.
Should a buyer demand model weights or training data?
Not always, and vendors may be unable to provide them. Ask what decision the evidence must support. Data provenance summaries, evaluation artifacts, audit rights, contractual restrictions, and workflow tests may be more actionable. High-risk or regulated uses can justify stronger transparency.
Can a small company do this?
Yes, by scaling diligence to risk. Start with use-case boundaries, data terms, permissions, a small protected test set, incident contacts, and export/deletion. Shared questionnaires can reduce effort, but the buyer must still test its own workflow.
How often should the dossier be refreshed?
On material model, subprocessor, data-use, control, contract, regulatory, or use-case changes—and at a scheduled interval based on risk. Monitor continuously where the vendor can change behavior without a new contract.
Source Notes — July 30, 2026
- NIST AI Risk Management Framework: Generative AI Profile, official voluntary guidance including procurement, supplier assessment, value-chain, and contingency practices.
- NIST AI Risk Management Framework 1.0, official cross-sector framework for Govern, Map, Measure, and Manage.
- Guidelines for Secure AI System Development, official guidance led by the UK NCSC with CISA and other international agencies.
- Regulation (EU) 2024/1689, the AI Act, official legal text; applicability and obligations require fact-specific legal analysis.
- FAR 39.102, Management of Risk, current US federal acquisition rule illustrating lifecycle IT risk, prototyping, measurement, and post-implementation review.
Standards, attestations, and regulatory classifications evolve. Verify the current text, effective dates, jurisdiction, and contract facts before making a legal or compliance determination.