AI API Integration: How to Connect AI Models to Business Applications
How to integrate AI model APIs into business applications: backend architecture, key management, prompt and context building, structured outputs, streaming, rate limits, retries, fallbacks, costs and monitoring.
Quick answer
Integrate AI models through your backend. A dedicated AI service authenticates the user, builds the prompt and context (retrieved data, conversation, instructions) within a token budget, calls the model API with a server-side key, timeouts and retries, validates structured outputs against schemas and business rules, then stores and returns the result. Around it, track tokens and cost per feature, enforce rate limits, plan fallbacks for provider errors, version prompts and evaluate changes. Never call model APIs with secret keys from browsers or mobile apps.
Where This Fits
Once several apps or providers are involved, an LLM gateway centralizes access, and LLM routing picks models per task. Connecting AI clients to your systems (the reverse direction) is covered in MCP vs API. General API integration practice is in website API integration and mobile app API integration.
Reference Architecture
| Component | Responsibility |
|---|---|
| Client (web, mobile, internal tool) | Sends user requests; never holds model API keys |
| AI service in your backend | Auth, context building, model calls, validation |
| Context sources | Database records, retrieval index, user profile |
| Model provider(s) | Generation, embeddings, speech |
| Gateway (optional) | Shared routing, limits, logging across apps |
| Observability | Traces, token usage, cost, errors, evaluations |
Choosing the Right Provider API
Use each provider's current recommended interface: OpenAI recommends the Responses API (the Assistants API was retired on 26 August 2026); Anthropic offers the Messages API with tool use and structured outputs; Google offers the Gemini API. Wrap provider SDKs behind your own interface so switching or adding providers does not touch business code.
Building Prompts and Context
Separate the parts: system instructions (versioned), task input, retrieved context and conversation history. Set a token budget and trim lowest-value context first. Clearly mark untrusted content (user input, documents) as data. Prompt caching features offered by providers reduce cost and latency when large stable prefixes repeat; order prompts so stable content comes first.
Structured Outputs and Validation
When code consumes the result, use schema-constrained outputs (available from major providers) and validate values with business rules. On failure, retry once with the validation error or fall back to a review path. For free-text responses shown to users, apply content checks proportionate to the risk.
async function classifyTicket(ticket, user) {
authorize(user, "tickets:classify")
const input = buildPrompt({ instructions: PROMPT_V7, ticket: asUntrusted(ticket.text) })
const res = await withRetry(() => model.respond({
model: config.models.classify,
input,
output_schema: TicketClassification,
max_output_tokens: 300,
timeout_ms: 15000,
}), { retryOn: [429, 500, 503], maxAttempts: 3, backoff: "exponential+jitter" })
const result = TicketClassification.parse(res.output) // validation
enforceRules(result) // business rules
logUsage({ feature: "ticket_classify", tokens: res.usage, promptVersion: "v7" })
return result
}Adding AI features to an existing product?
ZSpace Labs builds backend AI services with validation, cost tracking and fallbacks, integrated with your web and mobile apps.
Streaming
Streaming tokens to the client makes chat and long text feel fast. Stream through your backend (for example with server-sent events) so you keep control of authentication and logging. For structured outputs that drive actions, validate the complete response before acting. Handle client disconnects so you stop paying for unused generation where the provider allows cancellation.
Rate Limits, Retries and Fallbacks
- Retry 429 and transient 5xx errors with exponential backoff and jitter
- Respect retry-after headers
- Set timeouts per call and per user request
- Queue non-urgent work and use batch APIs where available
- Apply per-user and per-tenant rate limits in your service
- Fall back to another model or provider for critical paths, with evaluated quality
- Degrade gracefully: show a clear message instead of a spinner that never ends
Security and Privacy
Keep keys in a secrets manager and rotate them. Authenticate and authorize every request. Minimize data sent to providers, check retention and training policies and regions, and use enterprise data controls where needed. Treat model output as untrusted when it is used in queries, HTML or commands, to avoid injection into downstream systems.
Monitoring and Cost Control
Log model, prompt version, tokens, latency, errors and cost for every call, attributed to feature and customer. Set budgets and alerts. Review the most expensive features monthly and apply cost optimization levers. Re-run evaluations when providers update models.
Advantages and Limitations
A clean AI integration layer lets you add AI features quickly, swap models as they improve and keep control of cost and data. The limits are provider dependence, variable latency, non-deterministic outputs and costs that scale with usage. Design for those from the start.
How to Integrate Step by Step
- 1. Define the feature's input, output schema and success criteria
- 2. Build a backend AI service with a provider-neutral interface
- 3. Implement context building with token budgets
- 4. Add structured outputs and validation
- 5. Add retries, timeouts and fallbacks
- 6. Add logging, cost tracking and alerts
- 7. Evaluate on real examples before launch and on every change
Testing AI Integrations
Unit tests can mock the model client to check prompt construction, validation and error handling deterministically. Evaluation tests call the real model with a fixed set of inputs and score outputs, run on every prompt or model change. Contract tests confirm that structured outputs still match the schema your code expects. Load tests check behaviour under rate limits. Record model versions in test reports, because provider updates can change behaviour without any code change on your side. See AI evaluation.
Web and Mobile Considerations
Clients should call your backend, which calls the model. For streaming in browsers, use server-sent events or WebSockets from your backend; for mobile apps, handle interrupted connections and resume gracefully. Show progress states for longer operations, let users cancel, and design for failure with clear messages and retry options. Cache results the user is likely to revisit. For realtime voice in apps, providers offer WebRTC-based options; see voice AI agent development.
Worked Example
An illustrative scenario, not a client case: a mobile app calls a model API directly with a key embedded in the app, which is extracted and abused. The team moves calls to a backend AI service with user authentication and per-user limits, rotates the key, adds schema validation for the app's structured responses and gains per-feature cost reporting for the first time.
Common Mistakes
- API keys in front-end or mobile code
- Parsing free text instead of using structured outputs
- No timeouts or retry strategy
- No cost attribution
- Unversioned prompts
- Using model output directly in SQL, HTML or shell commands
Need a dependable AI integration layer?
Talk to ZSpace Labs about backend and API development, AI features in mobile apps and AI automation.
Conclusion
Treat model APIs like any critical dependency: call them from the backend, validate outputs, handle failures, track cost and test changes. Related: LLM gateway, LLM routing and LLM cost optimization.
Common questions
Call the model provider's API from your backend, not from the browser or app, with your own service handling authentication, prompt and context building, structured output validation, error handling, logging and cost tracking.