Skip to content
AI & Automation

AI Supply Chain Security: How to Assess Models, Datasets and Dependencies

How to secure the AI supply chain: model provenance and licences, safe model formats, signing, datasets and poisoning risk, package dependencies, third-party AI services and tools, AI bills of materials, vulnerability management and deployment controls.

Quick answer

Secure the AI supply chain by inventorying every model, dataset, package, container and AI service you depend on; obtaining models from official sources with checked hashes or signatures and safe formats such as safetensors; reviewing licences and data rights; protecting training and retrieval data from poisoning; scanning and pinning software dependencies; assessing third-party model APIs, tools and MCP servers as vendors; recording all of it in an AI bill of materials; and controlling what reaches production through review and signed artefacts.

Where This Fits

This article extends software supply chain practice to AI. Lineage of data is covered in AI data lineage, third-party tools in AI tool security and MCP security, self-hosted models in LLM self-hosting and the overall security picture in AI security for business applications.

What the AI Supply Chain Includes

ComponentExamplesKey risks
Pre-trained modelsOpen-weight LLMs, embedding and vision modelsTampered files, unsafe formats, licence terms, hidden behaviour
DatasetsPublic corpora, purchased data, scraped contentRights, consent, poisoning, quality
SoftwareML frameworks, inference servers, SDKs, containersVulnerabilities, typosquatting, abandoned packages
Hosted AI servicesModel APIs, embedding APIs, vector databasesData handling, model changes, outages
Tools and integrationsPlugins, MCP servers, agent toolsMalicious descriptions, excess permissions, compromise

Models: Provenance, Formats and Licences

Download models from official publisher accounts, verify hashes where published and record the exact revision. Prefer the safetensors format, which stores weights without executable code, over pickle-based formats that can run code when loaded; never load untrusted pickle files on systems with access to secrets. Emerging model signing, such as the OpenSSF model transparency project built on Sigstore, lets you verify that a model came from its claimed publisher.

Read licences carefully. Many popular models are open-weight but not open-source, with use restrictions, attribution requirements or thresholds; see LLM self-hosting for the distinction.

Data: Rights and Poisoning

For datasets used in training, fine-tuning, evaluation or retrieval, record source, licence or legal basis, collection date and processing. Poisoning risk applies wherever outsiders can influence data: public web data, user-generated content, shared documents indexed for retrieval. Restrict who can write to sources that feed AI systems, monitor changes, validate datasets before training and keep provenance so suspicious data can be traced and removed. CISA and partner agencies have published AI data security best practices covering these risks.

Bringing open models or third-party AI tools into production?

ZSpace Labs assesses AI dependencies and builds controlled deployment pipelines for models and tools. See AI engineering services.

Start a Project

Software Dependencies

AI projects pull in large dependency trees: ML frameworks, tokenizers, inference servers, vector database clients and agent frameworks, many moving fast. Apply normal software supply chain controls: pin versions, use lockfiles, scan for known vulnerabilities, verify package names to avoid typosquatting (AI coding assistants sometimes suggest non-existent packages), build containers from trusted bases and generate provenance for builds using frameworks such as SLSA.

Third-Party AI Services and Tools

Treat model API providers, AI SaaS vendors, plugins and MCP servers as suppliers. Assess data processing terms, retention, training use, regions, security certifications, incident notification and how model changes are communicated. For tools and MCP servers, review source or vendor reputation, required permissions and update practices, and watch for changed tool descriptions that could carry injected instructions.

AI Bill of Materials

An AI bill of materials lists models, datasets, software components and services with versions, sources and licences for each AI system. It speeds up response when a vulnerability or licence issue appears in a component, and it supports governance and customer questionnaires. The CycloneDX ML-BOM format is one standard way to express it; link it to your AI inventory in AI governance.

Only components that pass verification and review should reach the internal registry production pulls from.

Deployment Controls

Serve models and packages to production only from internal registries populated through review, not directly from public hubs. Sign artefacts you build and verify signatures at deploy time. Run model serving with least privilege and restricted network egress. Monitor advisories for components in your bill of materials and re-evaluate models after updates.

Advantages and Limitations

Supply chain controls prevent some of the most damaging and least visible AI risks, from malicious model files to licence surprises. Tooling for model signing and AI BOMs is still maturing, and full provenance for large pre-trained models is often unavailable. Focus on verifiable steps you control: sources, formats, scanning, registries and records.

How to Secure the AI Supply Chain Step by Step

  • 1. Inventory components for each AI system
  • 2. Verify sources, hashes and signatures for models
  • 3. Prefer safe formats and block untrusted pickle loading
  • 4. Review licences and data rights
  • 5. Scan and pin software dependencies
  • 6. Assess AI vendors, tools and MCP servers
  • 7. Use internal registries and an AI bill of materials

Model Evaluation as a Supply Chain Control

Verifying where a model came from does not tell you how it behaves. Before adopting a new model or version, run your evaluation suite, including safety, jailbreak and injection tests, and compare it with the model it replaces. Fine-tuned community models can carry altered behaviour that looks normal on common tasks. Re-evaluate after provider updates to hosted models too, since behaviour can change without notice; see LLM regression testing.

Vendor Assessment Questions

  • What data is retained, for how long, and is it used for training?
  • Where is data processed and stored, and which sub-processors are involved?
  • How are model changes and retirements communicated, and with what notice?
  • What security certifications and audit reports are available?
  • How are incidents detected, handled and notified?
  • What controls exist for access, logging and administration?

Example AI Bill of Materials Entry

An AI bill of materials does not have to start as a formal standard document; a structured record per system already answers most questions. Map it to a standard such as CycloneDX later.

Example: AI-BOM entry (illustrative)
system: claims-summary-assistant
models:
  - name: <open-weight-model>  version: <revision hash>
    source: official publisher repository  format: safetensors
    licence: <licence name>  verified_hash: true
  - name: <embedding-model>  provider: <hosted API>  version: <pinned>
datasets:
  - fine_tune: claims-summaries-v4 (internal, 3,200 reviewed examples, legal basis recorded)
  - retrieval: policy-docs index v22
software:
  - inference engine <name>@<version>, container sha256:...
  - agent framework <name>@<version>
services:
  - model gateway (internal), vector database (managed, EU region)
last_reviewed: 2026-09-29  owner: claims-platform-team

Worked Example

An illustrative scenario, not a client case: a team downloads a fine-tuned model from a community account because it scores well on a leaderboard. Review finds a pickle-based file from an unknown publisher and a licence that forbids commercial use. The team switches to the official base model in safetensors format, fine-tunes it internally with documented data, stores it in an internal registry and records it in the AI bill of materials.

Common Mistakes

  • Loading models from unverified community uploads
  • Using pickle-based model files from untrusted sources
  • Assuming open-weight means unrestricted licence
  • Installing AI-suggested packages without checking they exist and are legitimate
  • No record of which models and datasets are in production

Need an AI bill of materials or dependency review?

Talk to ZSpace Labs about an AI supply chain assessment for models, data, packages and services.

Start a Project

Conclusion

AI systems inherit risk from every model, dataset, package and service they use. Verify sources, prefer safe formats, review licences and data rights, vet vendors and tools, and keep an AI bill of materials so you can respond quickly when something changes.

FAQ

Common questions

Everything an AI system depends on that you did not build: pre-trained models and weights, datasets, ML frameworks and packages, containers, model APIs, third-party tools, plugins and MCP servers, and the vendors behind them.

Related services
Relevant industries
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.

Keep exploring
AI & Automation
6 min read

AI Security Testing: A Practical Checklist for Testing AI Systems

A practical, defensive checklist for testing AI systems: authentication and authorization, prompt injection, data leakage, tool execution, output handling, logging and privacy, dependencies, resource limits and incident response readiness.

Read article
AI & Automation
7 min read

AI Data Lineage: How to Track the Origin and Transformation of AI Data

How to track lineage for AI systems: source tracking, transformation history, dataset and index versions, links to models, prompts and outputs, reproducibility, auditability, deletion and governance, with standards and tools.

Read article
AI & Automation
8 min read

LLM Self-Hosting: How to Run Open-Weight Models on Your Own Infrastructure

How to self-host open-weight language models: when it makes sense, open-weight vs open-source, licences, hardware selection, serving software, security, scaling, monitoring, maintenance and total cost of ownership compared with hosted APIs.

Read article