Enterprise RAG Architecture: How to Build AI Systems With Company Data
How to architect RAG for an organization: source connectors, ingestion pipelines, access control sync, permission-aware retrieval, indexing, freshness, monitoring, governance and deployment.
Quick answer
Enterprise RAG architecture connects many company sources to AI applications safely. Connectors pull content and its access permissions; an ingestion pipeline parses, chunks, embeds and indexes it with metadata; retrieval resolves the user's identity and filters by permissions before hybrid search and reranking; a serving layer exposes grounded, cited answers to assistants and APIs; and an operations layer handles incremental sync, deletion, evaluation, monitoring and audit logs. Permissions and freshness are the two problems that separate enterprise RAG from a prototype.
Where This Fits
Core RAG concepts are in the RAG guide; the user-facing product in AI knowledge base. Retrieval quality techniques are in hybrid search and reranking; storage choices in vector databases.
Reference Architecture
| Layer | Components | Key concerns |
|---|---|---|
| Sources | Document stores, wikis, tickets, CRM, databases | Owners, formats, volume, change rate |
| Connectors | Content and ACL sync, change detection | Incremental updates, deletions, rate limits |
| Ingestion | Parsing, OCR, chunking, embedding, metadata | Structure preservation, cost, versioning |
| Index | Vector and keyword indexes, metadata store | Scale, filtering, multi-tenancy |
| Retrieval | Identity resolution, filters, hybrid search, rerank | Permissions, relevance, latency |
| Serving | Assistant UI, APIs, agents | Citations, refusals, rate limits |
| Operations | Evaluation, monitoring, audit, governance | Quality, cost, compliance |
Permission-Aware Retrieval
The non-negotiable rule: users must only receive answers built from content they could open in the source system. Sync access control lists with content (users, groups, sharing links), resolve the requesting user's identity and group memberships at query time, and apply them as retrieval filters. Re-sync permissions when they change, not only when content changes. Test with users who have different access, and log which sources contributed to each answer.
Connectors, Freshness and Deletion
Prefer change notifications or incremental sync over full re-crawls. Track each document's version and last sync. Deletions and permission removals must propagate quickly; an assistant that keeps quoting a withdrawn policy or a document someone lost access to is a real risk. Show the source date in answers so users can judge freshness.
Connector design, change detection and permission capture are covered in AI data ingestion, and near-real-time updates in real-time data for AI.
Connecting AI to company data without leaking it?
ZSpace Labs builds enterprise RAG with permission sync, incremental updates and audit logging across your document and business systems.
Indexing at Scale
Large corpora need batch and incremental ingestion, versioned embeddings (so you can re-embed when you change models) and an index that filters efficiently by metadata and permissions. Separate indexes by tenant or sensitivity when isolation requirements are strict. Plan for re-indexing: changing chunking or embedding models means reprocessing everything, so budget for it.
Security, Privacy and Data Residency
Classify sources by sensitivity and decide which can be indexed at all. Check where embedding and model providers process and retain data, and whether contractual terms meet your requirements. Encrypt indexes, restrict administrative access, keep audit logs of queries and sources, and treat retrieved content as untrusted input that cannot trigger actions without separate authorization.
Evaluation and Monitoring
- Question sets per department with expected sources
- Retrieval recall and ranking metrics by source
- Faithfulness and correctness of answers
- Permission tests with different user profiles
- Ingestion health: failures, lag, document counts
- Usage, feedback, latency and cost dashboards
Governance and Ownership
Assign owners to sources and content areas. Answers are only as good as the documents, so owners need reports on unanswered or poorly rated questions in their area. Define which sources are authoritative when documents conflict, and retire outdated content rather than leaving it searchable.
Build vs Buy
| Option | Fits | Trade-offs |
|---|---|---|
| Workplace AI features in existing suites | Content already in one suite | Limited control and customization |
| Enterprise search or RAG platforms | Many standard sources | Licence cost, connector coverage |
| Cloud building blocks | Teams with engineering capacity | More integration work |
| Custom build | Specialised sources or product-embedded RAG | Full control, full responsibility |
How to Implement Step by Step
- 1. Choose the first use case and its sources
- 2. Map permissions in each source and the identity system
- 3. Build connectors with content, ACL and deletion sync
- 4. Build ingestion and indexing with metadata and versioning
- 5. Implement permission-filtered hybrid retrieval and reranking
- 6. Evaluate, including permission tests
- 7. Launch to a pilot group with feedback and audit logging
- 8. Add sources and departments one at a time
Product-Embedded and Multi-Tenant RAG
When RAG is a feature of your product, serving many customers, tenant isolation becomes the top concern. Options include a separate index per tenant (strong isolation, more operational overhead), a shared index with mandatory tenant filters (efficient, but every query path must apply the filter), or a hybrid by tenant size. Enforce tenant scoping in the retrieval service, not in the application calling it, and test it with automated cross-tenant checks.
Agentic RAG in the Enterprise
Agents increasingly use retrieval as one tool among many, deciding when to search, which source to query and whether to search again. That improves complex answers but multiplies retrieval calls and permission checks. Give agents retrieval tools that enforce the user's permissions automatically, cap retrieval calls per run, and log which sources each step used. Keep retrieval-only assistants separate from agents that can take actions, to reduce prompt injection risk from indexed content.
Worked Example
An illustrative scenario, not a client case: a consulting firm wants an assistant over proposals, methodologies and client deliverables. Client folders have strict access, so the connector syncs folder permissions and group memberships; retrieval filters by the consultant's groups. A permission test suite runs nightly with test accounts for three roles, and an audit log records the sources behind each answer.
Common Mistakes
- Indexing everything with a single service account and no ACLs
- Filtering permissions after generation
- Ignoring deletions and permission changes
- No plan for re-embedding
- No content owners
Planning AI over your organization's knowledge?
Talk to ZSpace Labs about enterprise RAG development and connectors, APIs and deployment.
Conclusion
Enterprise RAG succeeds on permissions, freshness and ownership as much as on retrieval quality. Related: RAG guide, AI knowledge base and vector databases.
Common questions
Retrieval-augmented generation built for organizational use: many data sources, document-level permissions, large and changing content, audit requirements, multiple applications and production operations.