White-label private AI · for GPU integrators

You deliver the cluster.
We deliver what runs on it.

Enclave is the white-label application layer that turns your hardware quote into a working private AI deployment — a fixed-price line item on your paper, delivered under your brand. You margin every deal. We never touch your customer relationship.

~49%
typical support-contract attach
0%
typical AI-services attach
30 days
quote to working deployment

The stall

Everyone's warranty says the same thing: software not included

That's not a flaw — it's the right call for a hardware business. But it leaves your customer with a $30–100k system and nobody to make it do the thing they bought it for.

What the customer wanted

"Private ChatGPT on our documents. Our hardware, our building. Nothing leaves."

What the free stack delivers

The NVIDIA RAG Blueprint gets them a strong pilot — by design. Its interface is a sample for evaluation, and identity, document permissions, and connectors are left to the application layer. Someone still has to build that layer.

What happens next

The box underdelivers, a freelancer half-finishes it, and the customer remembers who sold them the hardware.

Enclave is the layer that monetizes what your warranty disclaims.

Built, not promised

The platform exists. You can watch it run.

This isn't a services pitch with a deck. Enclave is a built platform — four hardened services, 1,900+ automated tests, a scored evaluation harness — and the standing demo runs the full story on a single machine in about two minutes: sign-in, document sync, permission-scoped cited answers, sandboxed data analysis, audit export. With the Wi-Fi visibly off.

enclave — your-customer.internal
The Enclave chat application answering a records-retention question with a cited answer and an expandable reasoning trail showing the steps and sources consulted.
The product your customer logs into — under your brand, on your delivered hardware. Demo environment, fictional corpus.
What ships in a pilot

Branded chat application, single sign-on against the customer's identity provider, role-based access, cited answers with a visible reasoning trail, guardrails, admin training.

What production adds

SharePoint and file-share connectors with permission sync, ingestion tuned on the real corpus, an evaluation gold set, monitoring, backup/restore runbooks, admin console with audit export.

Factory pre-image

The platform installs at your integration bench, so the box ships working. Data onboarding happens after delivery, remotely, inside the customer's network.

And it is more than a chat window on a vector search. The application brain is a LangGraph supervisor loop that routes every turn — retrieval, sandboxed computation, or live systems — and calls the stock NVIDIA RAG Blueprint as its retrieval engine. This is the layer between the cluster you deliver and the outcome your customer bought:

needs documents

needs computation

needs a live system

passages gathered

charts + tables gathered

results gathered

no retrieval needed

enough gathered

/v1/search

turn starts

supervisor
routes with a structured decision:
rag · code · tools · respond · synthesize

rag agent
permission-scoped search
via the Blueprint

code agent
plan → script → locked
Docker sandbox

tools agent
typed tools + MCP servers

respond

synthesize
composed answer with
inline citations

streamed to the user

NVIDIA RAG Blueprint
(Diagram 1, stock)

Runs on the hardware you quote, under your brand. The Blueprint stays stock — your NVIDIA story and your customer's NVIDIA AI Enterprise entitlement carry straight through. Full walkthrough on the Built on NVIDIA page.

The SKU

A line item your reps can type today

QUOTE LINE ITEMENC-PILOT-25
Private AI Pilot — branded chat-with-your-documents on the delivered system. One corpus · SSO · 25 users · working in 30 days from readiness.
$25,000 fixed  ·  requires ENC-ASSESS ($2,500, credited)  ·  never gates the hardware PO

Separable line item — the customer can strike it and the box still ships. Or attach it as a scripted follow-up two weeks after ship notification. Your reps choose per deal.

PackageWhat the customer getsPrice
Assessment
mandatory first step
Corpus survey, permission-model review, sizing, written findings + firm quote$2,500 (credited)
PilotSingle box (2× NVIDIA RTX™ PRO 6000-class), one corpus, branded chat UI, SSO, ≤25 users, guardrails, training, scored acceptance demo. Pre-priced Production option: 60 days, 50% pilot credit$25,000 fixed
Production+ SharePoint / file-share connectors with document-permission sync, ingestion tuning on the real corpus, evaluation gold set, monitoring, runbook, admin console$55–95k fixed-scope
Managed AIPatching, model refreshes, index maintenance, quarterly quality report, next-business-day support$3.5–4.5k / mo

NVIDIA edition

Built on the NVIDIA RAG Blueprint: NVIDIA NIM™ microservices, NVIDIA NeMo™ Retriever embedding and reranking, guardrails, GPU-accelerated hybrid search — under an NVIDIA AI Enterprise subscription. For compliance-driven buyers who need the supported-stack story — and you resell the per-GPU NVIDIA AI Enterprise licenses for additional margin on every deployment.

Open edition

Fully open-source substrate (MIT/Apache), no per-GPU licensing, comparable capability — for cost-driven buyers. Same services pricing, same white-label rights.

Partner economics

Roughly double your gross per box. Zero delivery labor.

$3,000
Your gross today: $30k workstation at 10% margin
+$3,750
Referral margin on the $25k pilot (15%) — you introduce, we deliver
$750
Flat SPIF to your rep, per registered services deal, paid on customer PO

"For buyers whose data can't touch the cloud — CUI and DFARS 7012 environments, no-cloud-AI policies, air-gapped networks — cloud assistants aren't on the menu. And a cloud assistant never sold anyone a box."

WHY THIS SKU EXISTS

Why us

Built and proven where the stakes were highest

Enclave's founder built a U.S. state treasury's private AI platform from scratch: multi-agent RAG over sensitive financial data, local models, enterprise SSO and role-based access control, zero-trust architecture, Kubernetes — the exact architecture this SKU deploys.

Next step

One call. Bring an open hardware deal — we'll show you the attach.

20 minutes with the founder, including the live demo. If it's not a fit for your quote flow, you'll know by minute ten.

Book a partner call