Pricing
Pay per token. Not per seat.
Mandarum prices against GPU compute delivered — measured in tokens served, inferences completed, or GPU-hours consumed. Transparent tiers, monthly or annual.
Starter
Self-serve teams < 100 employees
$99/ month
- 1 workspace · 5 seats
- 2 integrations
- 1B tokens / month
- Standard models via NIM (Llama 3.1, Mistral)
- Community Slack support
Growth
Most popularMid-market · 100 – 1,000 employees
$2,400/ month
- Unlimited models & pipelines
- 10+ integrations · MCP servers
- 20B tokens / month
- SSO · SAML · SCIM · audit log
- Policy engine · approval routing
- Eval harness · golden datasets
- Slack support · 99.9% SLA
Enterprise
1,000+ employees or regulated
Custom
- Unlimited tokens · dedicated GPU capacity
- VPC / on-prem / BYOC
- NeMo fine-tunes + TensorRT-LLM compilation
- Forward-deployed engineering
- HIPAA · FedRAMP · ISO 42001 packs
- Data residency · BYOK
- 24×7 named CSM · 99.99% SLA
Credits
Universal Mandarum Credits.
One credit ≈ 1M tokens served. Complex multi-model pipelines consume 5–50 credits.
| Volume / month | Per-credit price | Equivalent |
|---|---|---|
| Up to 100k | $0.020 | ≈ 100B tokens |
| 100k – 1M | $0.015 | ≈ 1T tokens |
| 1M – 10M | $0.010 | ≈ 10T tokens |
| 10M+ | Negotiated | Custom GPU commitment |
Credits roll over one month. Overage is metered hourly. Budget caps and Slack alerts included in every plan.
Compare
What's included where.
| Starter | Growth | Enterprise | |
|---|---|---|---|
| Models & pipelines | Limited | Unlimited | Unlimited |
| Tokens / month | 1B | 20B | Unlimited |
| SSO / SAML / SCIM | — | Included | Included |
| Policy engine | Basic | Full | Full + custom |
| Eval harness | Sandbox | Production | Production + CI |
| VPC / on-prem | — | — | Included |
| NeMo fine-tunes | — | Add-on | Included |
| SLA | 99.5% | 99.9% | 99.99% |
| Support | Community | Slack · 4h | Named CSM · 24×7 |
FAQ
Pricing questions.
A "credit" covers approximately 1M tokens served across your models — whether via NIM microservices, Triton-served open-weight models, or hosted APIs.
Yes. Anthropic, OpenAI, Google, Azure OpenAI, AWS Bedrock, NIM microservices, Triton-served open weights, vLLM, and your own fine-tunes — billed per inference at your provider's price.
For Enterprise. We co-define the value metric (booked meetings, closed tickets, migrated modules) and share upside transparently.
60-day paid pilot at 50% of list, converts to annual if you meet your success metric.