AI Supply Chain Security: How to Assess Models, Datasets and Dependencies
How to secure the AI supply chain: model provenance and licences, safe model formats, signing, datasets and poisoning risk, package dependencies, third-party AI services and tools, AI bills of materials, vulnerability management and deployment controls.
Quick answer
Secure the AI supply chain by inventorying every model, dataset, package, container and AI service you depend on; obtaining models from official sources with checked hashes or signatures and safe formats such as safetensors; reviewing licences and data rights; protecting training and retrieval data from poisoning; scanning and pinning software dependencies; assessing third-party model APIs, tools and MCP servers as vendors; recording all of it in an AI bill of materials; and controlling what reaches production through review and signed artefacts.
Where This Fits
This article extends software supply chain practice to AI. Lineage of data is covered in AI data lineage, third-party tools in AI tool security and MCP security, self-hosted models in LLM self-hosting and the overall security picture in AI security for business applications.
What the AI Supply Chain Includes
| Component | Examples | Key risks |
|---|---|---|
| Pre-trained models | Open-weight LLMs, embedding and vision models | Tampered files, unsafe formats, licence terms, hidden behaviour |
| Datasets | Public corpora, purchased data, scraped content | Rights, consent, poisoning, quality |
| Software | ML frameworks, inference servers, SDKs, containers | Vulnerabilities, typosquatting, abandoned packages |
| Hosted AI services | Model APIs, embedding APIs, vector databases | Data handling, model changes, outages |
| Tools and integrations | Plugins, MCP servers, agent tools | Malicious descriptions, excess permissions, compromise |
Models: Provenance, Formats and Licences
Download models from official publisher accounts, verify hashes where published and record the exact revision. Prefer the safetensors format, which stores weights without executable code, over pickle-based formats that can run code when loaded; never load untrusted pickle files on systems with access to secrets. Emerging model signing, such as the OpenSSF model transparency project built on Sigstore, lets you verify that a model came from its claimed publisher.
Read licences carefully. Many popular models are open-weight but not open-source, with use restrictions, attribution requirements or thresholds; see LLM self-hosting for the distinction.
Data: Rights and Poisoning
For datasets used in training, fine-tuning, evaluation or retrieval, record source, licence or legal basis, collection date and processing. Poisoning risk applies wherever outsiders can influence data: public web data, user-generated content, shared documents indexed for retrieval. Restrict who can write to sources that feed AI systems, monitor changes, validate datasets before training and keep provenance so suspicious data can be traced and removed. CISA and partner agencies have published AI data security best practices covering these risks.
Bringing open models or third-party AI tools into production?
ZSpace Labs assesses AI dependencies and builds controlled deployment pipelines for models and tools. See AI engineering services.
Software Dependencies
AI projects pull in large dependency trees: ML frameworks, tokenizers, inference servers, vector database clients and agent frameworks, many moving fast. Apply normal software supply chain controls: pin versions, use lockfiles, scan for known vulnerabilities, verify package names to avoid typosquatting (AI coding assistants sometimes suggest non-existent packages), build containers from trusted bases and generate provenance for builds using frameworks such as SLSA.
Third-Party AI Services and Tools
Treat model API providers, AI SaaS vendors, plugins and MCP servers as suppliers. Assess data processing terms, retention, training use, regions, security certifications, incident notification and how model changes are communicated. For tools and MCP servers, review source or vendor reputation, required permissions and update practices, and watch for changed tool descriptions that could carry injected instructions.
AI Bill of Materials
An AI bill of materials lists models, datasets, software components and services with versions, sources and licences for each AI system. It speeds up response when a vulnerability or licence issue appears in a component, and it supports governance and customer questionnaires. The CycloneDX ML-BOM format is one standard way to express it; link it to your AI inventory in AI governance.
Deployment Controls
Serve models and packages to production only from internal registries populated through review, not directly from public hubs. Sign artefacts you build and verify signatures at deploy time. Run model serving with least privilege and restricted network egress. Monitor advisories for components in your bill of materials and re-evaluate models after updates.
Advantages and Limitations
Supply chain controls prevent some of the most damaging and least visible AI risks, from malicious model files to licence surprises. Tooling for model signing and AI BOMs is still maturing, and full provenance for large pre-trained models is often unavailable. Focus on verifiable steps you control: sources, formats, scanning, registries and records.
How to Secure the AI Supply Chain Step by Step
- 1. Inventory components for each AI system
- 2. Verify sources, hashes and signatures for models
- 3. Prefer safe formats and block untrusted pickle loading
- 4. Review licences and data rights
- 5. Scan and pin software dependencies
- 6. Assess AI vendors, tools and MCP servers
- 7. Use internal registries and an AI bill of materials
Model Evaluation as a Supply Chain Control
Verifying where a model came from does not tell you how it behaves. Before adopting a new model or version, run your evaluation suite, including safety, jailbreak and injection tests, and compare it with the model it replaces. Fine-tuned community models can carry altered behaviour that looks normal on common tasks. Re-evaluate after provider updates to hosted models too, since behaviour can change without notice; see LLM regression testing.
Vendor Assessment Questions
- What data is retained, for how long, and is it used for training?
- Where is data processed and stored, and which sub-processors are involved?
- How are model changes and retirements communicated, and with what notice?
- What security certifications and audit reports are available?
- How are incidents detected, handled and notified?
- What controls exist for access, logging and administration?
Example AI Bill of Materials Entry
An AI bill of materials does not have to start as a formal standard document; a structured record per system already answers most questions. Map it to a standard such as CycloneDX later.
system: claims-summary-assistant
models:
- name: <open-weight-model> version: <revision hash>
source: official publisher repository format: safetensors
licence: <licence name> verified_hash: true
- name: <embedding-model> provider: <hosted API> version: <pinned>
datasets:
- fine_tune: claims-summaries-v4 (internal, 3,200 reviewed examples, legal basis recorded)
- retrieval: policy-docs index v22
software:
- inference engine <name>@<version>, container sha256:...
- agent framework <name>@<version>
services:
- model gateway (internal), vector database (managed, EU region)
last_reviewed: 2026-09-29 owner: claims-platform-teamWorked Example
An illustrative scenario, not a client case: a team downloads a fine-tuned model from a community account because it scores well on a leaderboard. Review finds a pickle-based file from an unknown publisher and a licence that forbids commercial use. The team switches to the official base model in safetensors format, fine-tunes it internally with documented data, stores it in an internal registry and records it in the AI bill of materials.
Common Mistakes
- Loading models from unverified community uploads
- Using pickle-based model files from untrusted sources
- Assuming open-weight means unrestricted licence
- Installing AI-suggested packages without checking they exist and are legitimate
- No record of which models and datasets are in production
Need an AI bill of materials or dependency review?
Talk to ZSpace Labs about an AI supply chain assessment for models, data, packages and services.
Conclusion
AI systems inherit risk from every model, dataset, package and service they use. Verify sources, prefer safe formats, review licences and data rights, vet vendors and tools, and keep an AI bill of materials so you can respond quickly when something changes.
Common questions
Everything an AI system depends on that you did not build: pre-trained models and weights, datasets, ML frameworks and packages, containers, model APIs, third-party tools, plugins and MCP servers, and the vendors behind them.