Frontier AI, from private cloud to air gap
Private models without the hardware
- Everything in Teams
- Pro & Ultra seat tiers
- Sovereign Presets on private infra
- Zero third-party subprocessors
- SAML SSO + SCIM provisioning
- Workspace model allowlisting
- Custom retention & legal holds
- Immutable audit logs & SIEM export
- SOC 2 Type II & HIPAA BAA package
- Dedicated CSM & security review help
Sovereign Box: the whole stack in one rack
- 4× RTX PRO 6000 Blackwell Server Edition, 384 GB VRAM
- Multiple mid-size open models, routed intelligently
- Up to 4 models served simultaneously
- From 150 concurrent streams and 1,000 users
- +50 streams per node beyond the third
- Full capacity with two nodes down
- Standard 110/220 V office power
- Self-contained, with database, cache, and storage on board
- Unlimited seats
- Quarterly offline model updates
- 8× NVIDIA HGX B300, 2.3 TB HBM3e
- Largest open-weight models resident in memory
- Headroom for smaller routing models
- From 600 concurrent streams and 4,000 users
- +200 streams per node beyond the third
- Full capacity with two nodes down
- 24/7 monitoring, 4-hour P1 response
- Unlimited seats
- Secure courier-delivered updates
One private endpoint for every AI tool
- Coding agents on private modelsPoint Claude Code or any other AI assistant at the private endpoint. Source code, prompts, and diffs stay on infrastructure the company controls.
- Standard, drop-in APIsIndustry-standard API compatibility means existing SDKs, IDE plugins, and internal tools work unchanged. Swap the base URL, keep the workflow.
- Internal automations, same perimeterNightly extraction jobs, CI reviews, and in-house agents call the same endpoint, so prompts, source code, and documents never leave the deployment.
As many agents as you want, for the same bill
- The meter is offYou pay for capacity: a box in your rack, or single-tenant infrastructure we operate. Whether it idles or runs all night, the invoice at the end of the month is the one you signed.
- Agents without a quotaCoding agents, extraction pipelines, CI reviews, and overnight evaluations run at once, around the clock. The tenth agent costs what the first one did.
- More tokens for the same moneyNobody trims context to save cents or thinks twice about a retry. The work gets the tokens it actually needs, and the number on the invoice stays where it was.
When the big providers go down, you don't
- Nobody else's traffic, nobody else's incidentThe cluster serves your users and no one else, so another company's spike or a provider's control-plane failure has no path to it. An air-gapped Box does not even need the internet to be up.
- Provider outages pass you byLarge providers go down for hours at a time. Work on a Box has carried on straight through those incidents, because nothing in the answer path belongs to a company you do not control.
- Full capacity with two nodes downEvery cluster ships as at least three nodes, so a failed node, or one pulled for maintenance, does not take the service with it. Updates land in your change window, not when a vendor ships.
Total cost of ownership
Marketing, sales & operations
Drafting, research, summarizing, everyday chat
Analysts, legal & research
Long documents, deep research, council answers
Engineers
Coding assistants and agentic workloads
How hard do your engineers push it?
Agentic workloads are the line that dominates a metered bill.
Sovereign Box support
Per managed node equivalent, with the first three included across the fleet.
Saved per month
$12,800
Saved per year
$153,600
Running 100 people on your own hardware, against the same org on metered Enterprise SaaS seats.
Enterprise SaaS
$25,000/mo
- Seats
- $5,000
- Agentic usage
- $20,000
3 × Sovereign Box Theta
$12,200/mo
3-node HA cluster, automatic failover- 3 nodes (hardware)
- $7,200
- Software license (cluster)
- $5,000
- Standard support
- Included
- Seats
- Unlimited
- Agentic usage
- $0 (marginal token is free)
Serving 100 people and ~30 agents
22 of 150 streams
Sizing is set by how many people generate at the same instant rather than by headcount, because most of an org is reading rather than waiting on tokens. At this mix the configuration has room for about 560 people before you need more capacity.
The bar shows whichever ceiling sets the node count, streams or population. Committed capacity here is 150 streams across 3 nodes at 50 each, with 1,000 rated seats. Each node commits a third of what it can physically serve, so 2 nodes' worth stays in reserve and losing any two of them serves this load without interruption.
Estimates only, for comparison. Seats priced at the published Pro ($30) and Ultra ($80) rates; appliances at Hardware-as-a-Service monthly pricing on a 36-month term, software license included. Agentic figures assume $1,000/engineer/month of metered model usage at standard intensity, and about 1.5 concurrent agents per engineer. Clusters start at 3 nodes, hold 2 nodes' worth of capacity in reserve, and are sized to stay under 80% of committed streams; every node bills the same monthly rate under a single cluster-wide software license. A Theta cluster scales horizontally until metered spend passes $131,000/month, where Omega and the largest open-weight models become the better step. Support is shown at the tier selected above, priced per managed node equivalent with the first 3 MNE included once across the fleet.
Frequently asked questions
Add-ons and services
Listed up front rather than averaged into the quote.
- Additional sites$25,000/yr beyond the first
- Three-year prepayment10% discount
- Hardware refresh20% trade-in credit at year three
- Pre-sale site survey$25,000, credited against purchase
- Sovereign model evaluation on in-house data$25,000, credited against purchase
- Theta installation and commissioningFrom $10,000
- Omega installation and commissioningFrom $50,000
- On-site spares cache$50,000
- RAG and data-source integrationFrom $50,000
- Domain distillation and model tuningFrom $75,000