Deployment · On-premise

Run Enterprise RAG entirely on your own infrastructure.

The same platform — multimodal ingestion, hybrid retrieval, cited answers, and enterprise governance — deployed inside your VPC, datacenter, or an air-gapped network. Your data and your models never leave your control.

Why on-premise

When your data can’t go to the cloud, bring the platform to your data.

Data sovereignty & residency

Your documents, prompts, and answers never leave your network. Keep data in-country and in-region to satisfy GDPR, data-residency, and sovereignty requirements — with no third-party processor in the loop.

Air-gapped & offline

Run fully disconnected from the internet for classified, defense, or isolated OT environments. The entire stack — models, index, and UI — operates with zero external calls.

Certification-ready compliance

The on-premise deployment ships the controls SOC 2, ISO 27001, and HIPAA require — encryption, access control, audit, key management — so you can certify it inside the environment you already own and audit.

Your models, your GPUs

Serve local LLMs (Ollama, vLLM, TGI) or your own OpenAI-compatible endpoints. No document or question is ever sent to an external AI provider — inference stays on your hardware.

Performance & cost at scale

Colocate the platform next to where your data already lives, use your own GPUs, and avoid per-token cloud pricing and egress. Predictable cost that scales with hardware, not usage.

Full control, no lock-in

Your VPC, your SSO/SCIM, your retention and moderation policy, your backups. The whole platform runs on infrastructure you own and can audit end to end.

Compliance

Certification-ready for SOC 2, ISO 27001 & HIPAA.

The on-premise deployment ships the technical controls these frameworks require — encryption at rest, RBAC and ACLs, MFA and SSO/SCIM, immutable audit logging, key management, and backup/restore. Because it runs inside your own infrastructure, you certify it within the environment you already own and audit — we supply the control-to-framework mapping.

See the compliance & control mapping
SOC 2 ISO/IEC 27001 HIPAA

Sizing calculator

Size your rig in seconds — pick your data, users, and GPU.

A quick estimate of the hardware to run Enterprise RAG on your own infrastructure. Drag the sliders, choose an accelerator, and watch the spec build itself.

Your workload

Data to index500 GB
Approx. users100

Preferred accelerator

Desktop AI

RTX / Workstation

Datacenter

Other

NVIDIA A100 80GB80 GB HBM2e

Recommended deployment

Team

14B class (e.g. Qwen2.5 14B)
GPUs80 GB VRAM total
1× A100 80G
RAM / node
128 GB
128 GB across 1 node
CPU / node
24 cores
24 cores total
NVMe storage
2 TB
sharded · quantized · disk-backed
Nodes
1
1 GPU / node
Peak concurrent
5
Queries / min
58
Ingest pages / hr
100,800
Networking: 10 GbE
Get a tailored deployment plan

Estimate only — real sizing depends on document mix, model choice, and latency targets. RAM is per node and holds the index working set; the full corpus lives on NVMe (sharded + quantized), never entirely in memory. Our team confirms the final spec.

Who runs on-premise

Built for regulated and data-sensitive organizations.

If your data is subject to regulation, classification, or contractual controls that keep it off public clouds, the on-premise deployment gives you the full product without the compliance trade-off.

  • Financial services & banking
  • Healthcare & life sciences
  • Government & public sector
  • Defense & intelligence
  • Legal & professional services
  • Critical infrastructure & energy

What you get

Your environment

Docker Compose or Kubernetes, in your VPC or datacenter — with an offline installer for air-gapped sites.

Your identity & policy

Your SSO (OIDC/SAML), SCIM provisioning, retention, moderation, and audit — wired into your systems.

Your models

Local or self-hosted LLMs and embedders; nothing routed to an external AI API.

Talk to us about your environment

Tell us where you want to run Enterprise RAG — air-gapped, VPC, or your own datacenter — and our team will reach out with a deployment plan tailored to your compliance and infrastructure needs.

Request on-premise

← Back to Enterprise RAG