Private AI · on your hardware · behind your firewall

Private AI your whole organization can actually log into.

Enclave turns NVIDIA-accelerated hardware into a complete AI application: single sign-on, permission-aware retrieval, answers cited to their source documents, and a full audit trail. Deployed on your systems, in your building. Your documents never leave your network.

30 days
assessment to working pilot
Every answer
cited to its source documents
1,900+
automated tests behind each release
enclave — your-company.internal
The Enclave chat application answering a records-retention question with a cited answer and an expandable 'How I got this' reasoning trail showing the four steps and five sources consulted.
The Enclave application — a cited answer with its full reasoning trail. Demo environment, fictional aerospace-supplier corpus.

Why Enclave

The cluster is the easy part.

Racked GPUs, validated drivers, the NVIDIA software stack installed and burned in — a good hardware partner handles that well. Then comes the question that stalls most private-AI projects: how do 400 employees actually use this — and how do you stop the wrong ones from reading HR files? Enclave is the answer to that question: the production application layer that ships on the box.

An application, not a toolkit

Your people get a chat application they sign into with the company credentials they already have — not a framework your team has to finish building.

Permissions, provable

Retrieval is scoped to what each person is allowed to read, enforced at the API — and the audit trail proves it, question by question.

Accountable answers

Every response cites the documents it drew from, shows how it got there, and is accepted against a scored evaluation — not a demo.

The platform

What "working private AI" actually includes

Not a chatbot bolted to a folder. A governed answer engine your security team can sign off on — running as four hardened services on the delivered hardware.

For the people using it daily

A chat experience that shows its work

  • Answers with citations — every response links to the source documents it drew from, so users verify instead of trust.
  • Visible reasoning — an expandable trail of what was searched, which documents were consulted, and how the answer formed.
  • Projects — persistent workspaces that keep a team's documents, conversations, and context together.
  • Reads the whole document — tables, charts, and diagrams inside PDFs and slide decks, not just the paragraphs.
  • Analysis on request — computations and charts generated from your data files, in a locked-down sandbox, on the box.
For the security & compliance team

Governance that survives an audit

  • Single sign-on — your identity provider (Microsoft Entra ID and other OIDC providers), with role- and group-based access control.
  • Permission-aware retrieval — answers respect the source system's access rules. If a user can't open the file, the AI won't quote it to them.
  • Append-only audit trail — who asked, what was retrieved, which model answered, exportable to CSV.
  • Guardrails — topic control, content safety, and prompt-injection defense at the platform layer.
  • Service-to-service mTLS — every internal hop encrypted and mutually authenticated.
For the team that owns it after go-live

Engineering that keeps answers good

  • Live connectors — SharePoint and file shares with scheduled incremental sync, including deletions. The corpus stays current without re-uploading.
  • Retrieval tuned to your corpus — hybrid search with reranking, chunking calibrated to your actual documents, not defaults.
  • Measured before sign-off — acceptance is a scored evaluation on a question set built with your team, with negative controls. Sign-off is a number, not a vibe.
  • Sized to the workload — model selection and GPU allocation engineered for the stated user count.
  • Operable — metrics, tracing, token-usage accounting, backup and restore runbooks, and a managed-operations tier.
enclave — analysis agent
The Enclave analysis agent computing monthly scrap rates from an uploaded workbook, showing its planned approach, sandboxed execution steps, and downloadable chart and table artifacts.
The analysis agent: sandboxed computation over your files, with the chart and the working shown.

What happens between your question and the answer: every turn is routed by a supervisor — to document retrieval, to a locked-down analysis sandbox, or to a connected system — and composed only when enough has been gathered. A greeting never triggers a search; a data question never gets a guess. The same trail your people see in "How I got this" is this loop, step by step:

needs documents

needs computation

needs a live system

passages gathered

charts + tables gathered

results gathered

no retrieval needed

enough gathered

/v1/search

turn starts

supervisor
routes with a structured decision:
rag · code · tools · respond · synthesize

rag agent
permission-scoped search
via the Blueprint

code agent
plan → script → locked
Docker sandbox

tools agent
typed tools + MCP servers

respond

synthesize
composed answer with
inline citations

streamed to the user

NVIDIA RAG Blueprint
(Diagram 1, stock)

The application brain. The NVIDIA RAG Blueprint runs stock underneath as the retrieval engine — see the full architecture.

UNDER THE HOOD · NVIDIA RAG Blueprint · NVIDIA NIM™ microservices · NVIDIA NeMo™ Retriever extraction, embedding + reranking · LangGraph agent orchestration · Elasticsearch hybrid vector search · OpenID Connect SSO · Postgres audit & state · Prometheus + OpenTelemetry · Docker Compose or Kubernetes delivery · mTLS between every service — see the full stack

Security & governance

Built for buyers whose data can't leave

Defense suppliers working in DFARS 252.204-7012 and NIST SP 800-171 CUI environments. State and local government. Community banks and credit unions. Law firms. If your policy says client data doesn't touch public AI services, Enclave is built for you: the entire pipeline — models, retrieval, storage, audit — runs on hardware you own, and an air-gapped installation with an offline runbook is available for disconnected environments.

enclave — admin console · audit
The Enclave admin console audit view: an append-only log of every authentication, question, retrieval, and access change, filterable by actor and event, with per-turn detail including the collections the user was allowed to query and CSV export.
The audit trail: every question records who asked, what they were permitted to search, what was retrieved, and which model answered.
Access control

Users and groups flow in from your identity provider. Administrators grant collections to groups; the retrieval API enforces the grants on every query — the UI never has to be trusted.

Every turn on the record

The audit log captures the permission scope each answer was computed under — which collections the asker was allowed, which were queried, which documents were cited.

Deployment hardening

The AI engine is never exposed directly: one authenticated edge, mTLS between services, sandboxed code execution with no network access, secrets outside the repo.

How you buy

Fixed scope. Fixed price. Measured acceptance.

No per-seat subscriptions, no metered tokens. You buy a deployment with a written scope and a scored acceptance evaluation — and an optional managed tier that keeps it good after go-live. Hardware is quoted separately by your integration partner; if you don't have one, we'll introduce you to one.

PackageWhat you getPrice
Assessment
first step
Corpus survey, permission-model review, hardware sizing, written findings and a firm fixed quote $2,500 · credited
Pilot
working in 30 days
Single system (2× NVIDIA RTX™ PRO 6000-class), one document corpus, branded chat application, single sign-on, up to 25 users, guardrails, admin training, scored acceptance demo $25,000 fixed
Production Everything in Pilot, plus SharePoint and file-share connectors with permission sync, ingestion tuning on your real corpus, an evaluation gold set built with your team, monitoring, backup/restore runbooks, admin console $55–95k fixed-scope
Managed AI
the ongoing tier
Patching and upgrade cadence, model refreshes, index maintenance, quarterly quality report scored against your gold set, next-business-day support $3.5–4.5k / mo

NVIDIA edition

Built on the NVIDIA RAG Blueprint with NVIDIA NIM microservices and NVIDIA NeMo Retriever, under an NVIDIA AI Enterprise subscription (licensed per GPU, resold with the deployment). For compliance-driven buyers who need the commercially supported stack.

Open edition

Fully open-source substrate (MIT/Apache licensed), no per-GPU software licensing, comparable capability. For cost-driven buyers. Same application layer, same deployment services, same acceptance standard.

Acceptance is a scored evaluation on a question set built with your team — including questions the system should refuse — with written pass criteria. You sign off on a number, not a feeling.

HOW EVERY ENGAGEMENT ENDS

The foundation

Built on the NVIDIA AI stack

Enclave productionizes the NVIDIA RAG Blueprint — NVIDIA's reference architecture for enterprise retrieval-augmented generation — and adds the application layer it deliberately leaves to independent software vendors: identity, permissions, connectors, agents, audit, and evaluation. We build on the Blueprint and track its releases; we never fork it.

EnclaveSSO · RBAC · permission-aware retrieval · connectors · agents · chat application · audit · evals
NVIDIA RAG Blueprintmultimodal ingestion · hybrid retrieval · reranking · generation
NVIDIA NIM microservices · NVIDIA NeMo Retrieveraccelerated inference, embedding, extraction
NVIDIA AI Enterprisesupported, production-ready software platform — licensed per GPU
Your hardwareNVIDIA RTX PRO workstations to multi-GPU servers, sized in the Assessment

How Enclave and the NVIDIA stack fit together →

Delivery

Delivered with your hardware partner

Enclave is software. We don't rack servers, design clusters, or sell hardware — your GPU integration partner does that, and does it well. We work alongside them: the platform is installed at integration time, so the system ships working, and your data onboarding happens after delivery, inside your network. Already have a hardware partner? We'll work with them. Need one? We'll introduce you.

GPU integrator? See the partner program →

Why us

Built and proven where the stakes were highest

Enclave's founder built a U.S. state treasury's private AI platform from scratch — multi-agent retrieval over sensitive financial data, local models, enterprise SSO and role-based access, zero-trust architecture, Kubernetes — the same architecture Enclave deploys.

Next step

Start with the Assessment.

A corpus survey, a permission-model review, a sizing workbook, and a firm fixed quote — $2,500, credited toward your build. If private AI isn't a fit for your environment, the findings will say so in writing.

Contact us