Private AI · on your hardware · behind your firewall
Enclave turns NVIDIA-accelerated hardware into a complete AI application: single sign-on, permission-aware retrieval, answers cited to their source documents, and a full audit trail. Deployed on your systems, in your building. Your documents never leave your network.
Why Enclave
Racked GPUs, validated drivers, the NVIDIA software stack installed and burned in — a good hardware partner handles that well. Then comes the question that stalls most private-AI projects: how do 400 employees actually use this — and how do you stop the wrong ones from reading HR files? Enclave is the answer to that question: the production application layer that ships on the box.
Your people get a chat application they sign into with the company credentials they already have — not a framework your team has to finish building.
Retrieval is scoped to what each person is allowed to read, enforced at the API — and the audit trail proves it, question by question.
Every response cites the documents it drew from, shows how it got there, and is accepted against a scored evaluation — not a demo.
The platform
Not a chatbot bolted to a folder. A governed answer engine your security team can sign off on — running as four hardened services on the delivered hardware.
What happens between your question and the answer: every turn is routed by a supervisor — to document retrieval, to a locked-down analysis sandbox, or to a connected system — and composed only when enough has been gathered. A greeting never triggers a search; a data question never gets a guess. The same trail your people see in "How I got this" is this loop, step by step:
UNDER THE HOOD · NVIDIA RAG Blueprint · NVIDIA NIM™ microservices · NVIDIA NeMo™ Retriever extraction, embedding + reranking · LangGraph agent orchestration · Elasticsearch hybrid vector search · OpenID Connect SSO · Postgres audit & state · Prometheus + OpenTelemetry · Docker Compose or Kubernetes delivery · mTLS between every service — see the full stack
Security & governance
Defense suppliers working in DFARS 252.204-7012 and NIST SP 800-171 CUI environments. State and local government. Community banks and credit unions. Law firms. If your policy says client data doesn't touch public AI services, Enclave is built for you: the entire pipeline — models, retrieval, storage, audit — runs on hardware you own, and an air-gapped installation with an offline runbook is available for disconnected environments.
Users and groups flow in from your identity provider. Administrators grant collections to groups; the retrieval API enforces the grants on every query — the UI never has to be trusted.
The audit log captures the permission scope each answer was computed under — which collections the asker was allowed, which were queried, which documents were cited.
The AI engine is never exposed directly: one authenticated edge, mTLS between services, sandboxed code execution with no network access, secrets outside the repo.
How you buy
No per-seat subscriptions, no metered tokens. You buy a deployment with a written scope and a scored acceptance evaluation — and an optional managed tier that keeps it good after go-live. Hardware is quoted separately by your integration partner; if you don't have one, we'll introduce you to one.
| Package | What you get | Price |
|---|---|---|
| Assessment first step |
Corpus survey, permission-model review, hardware sizing, written findings and a firm fixed quote | $2,500 · credited |
| Pilot working in 30 days |
Single system (2× NVIDIA RTX™ PRO 6000-class), one document corpus, branded chat application, single sign-on, up to 25 users, guardrails, admin training, scored acceptance demo | $25,000 fixed |
| Production | Everything in Pilot, plus SharePoint and file-share connectors with permission sync, ingestion tuning on your real corpus, an evaluation gold set built with your team, monitoring, backup/restore runbooks, admin console | $55–95k fixed-scope |
| Managed AI the ongoing tier |
Patching and upgrade cadence, model refreshes, index maintenance, quarterly quality report scored against your gold set, next-business-day support | $3.5–4.5k / mo |
Built on the NVIDIA RAG Blueprint with NVIDIA NIM microservices and NVIDIA NeMo Retriever, under an NVIDIA AI Enterprise subscription (licensed per GPU, resold with the deployment). For compliance-driven buyers who need the commercially supported stack.
Fully open-source substrate (MIT/Apache licensed), no per-GPU software licensing, comparable capability. For cost-driven buyers. Same application layer, same deployment services, same acceptance standard.
Acceptance is a scored evaluation on a question set built with your team — including questions the system should refuse — with written pass criteria. You sign off on a number, not a feeling.
HOW EVERY ENGAGEMENT ENDS
The foundation
Enclave productionizes the NVIDIA RAG Blueprint — NVIDIA's reference architecture for enterprise retrieval-augmented generation — and adds the application layer it deliberately leaves to independent software vendors: identity, permissions, connectors, agents, audit, and evaluation. We build on the Blueprint and track its releases; we never fork it.
Delivery
Enclave is software. We don't rack servers, design clusters, or sell hardware — your GPU integration partner does that, and does it well. We work alongside them: the platform is installed at integration time, so the system ships working, and your data onboarding happens after delivery, inside your network. Already have a hardware partner? We'll work with them. Need one? We'll introduce you.
Why us
Enclave's founder built a U.S. state treasury's private AI platform from scratch — multi-agent retrieval over sensitive financial data, local models, enterprise SSO and role-based access, zero-trust architecture, Kubernetes — the same architecture Enclave deploys.
Next step
A corpus survey, a permission-model review, a sizing workbook, and a firm fixed quote — $2,500, credited toward your build. If private AI isn't a fit for your environment, the findings will say so in writing.
Contact us