The same platform — multimodal ingestion, hybrid retrieval, cited answers, and enterprise governance — deployed inside your VPC, datacenter, or an air-gapped network. Your data and your models never leave your control.
When your data can’t go to the cloud, bring the platform to your data.
Your documents, prompts, and answers never leave your network. Keep data in-country and in-region to satisfy GDPR, data-residency, and sovereignty requirements — with no third-party processor in the loop.
Run fully disconnected from the internet for classified, defense, or isolated OT environments. The entire stack — models, index, and UI — operates with zero external calls.
The on-premise deployment ships the controls SOC 2, ISO 27001, and HIPAA require — encryption, access control, audit, key management — so you can certify it inside the environment you already own and audit.
Serve local LLMs (Ollama, vLLM, TGI) or your own OpenAI-compatible endpoints. No document or question is ever sent to an external AI provider — inference stays on your hardware.
Colocate the platform next to where your data already lives, use your own GPUs, and avoid per-token cloud pricing and egress. Predictable cost that scales with hardware, not usage.
Your VPC, your SSO/SCIM, your retention and moderation policy, your backups. The whole platform runs on infrastructure you own and can audit end to end.
Certification-ready for SOC 2, ISO 27001 & HIPAA.
The on-premise deployment ships the technical controls these frameworks require — encryption at rest, RBAC and ACLs, MFA and SSO/SCIM, immutable audit logging, key management, and backup/restore. Because it runs inside your own infrastructure, you certify it within the environment you already own and audit — we supply the control-to-framework mapping.
See the compliance & control mappingSize your rig in seconds — pick your data, users, and GPU.
A quick estimate of the hardware to run Enterprise RAG on your own infrastructure. Drag the sliders, choose an accelerator, and watch the spec build itself.
Desktop AI
RTX / Workstation
Datacenter
Other
NVIDIA A100 80GB — 80 GB HBM2e
Recommended deployment
Estimate only — real sizing depends on document mix, model choice, and latency targets. RAM is per node and holds the index working set; the full corpus lives on NVMe (sharded + quantized), never entirely in memory. Our team confirms the final spec.
Built for regulated and data-sensitive organizations.
If your data is subject to regulation, classification, or contractual controls that keep it off public clouds, the on-premise deployment gives you the full product without the compliance trade-off.
Your environment
Docker Compose or Kubernetes, in your VPC or datacenter — with an offline installer for air-gapped sites.
Your identity & policy
Your SSO (OIDC/SAML), SCIM provisioning, retention, moderation, and audit — wired into your systems.
Your models
Local or self-hosted LLMs and embedders; nothing routed to an external AI API.
Tell us where you want to run Enterprise RAG — air-gapped, VPC, or your own datacenter — and our team will reach out with a deployment plan tailored to your compliance and infrastructure needs.
Request on-premise