Mandarum Twinex — Model the fleet before you buy it.
A digital twin of your GPU inference fleet, built from real DCGM telemetry and rendered in NVIDIA Omniverse. Run what-if scenarios — hardware swaps, Dynamo disaggregation, batching, autoscaling, routing — and predict p99, throughput, utilization, and $/token before anything touches production.
Why GPU decisions are still made on gut and guesswork.
GPU buys are a gamble
Teams size fleets on spreadsheets and hope. Twinex predicts the impact of adding Blackwell or doubling traffic on p99 and $/token.
Topology changes hit prod blind
Switching to disaggregated serving or new batching can backfire. Twinex tests the change in simulation first.
The GPU bill is a black box
Finance can't forecast inference spend. Twinex projects $/token across scenarios so capacity and budget decisions are grounded.
What Twinex does.
A calibrated twin of your fleet.
Twinex ingests DCGM traces and serving logs to build a digital twin that replays GPU kernel timing, KV-cache, and batching behavior, calibrated to your real hardware and rendered in Omniverse.
- → DCGM + serving-log calibration
- → GPU-accurate timing + KV-cache model
- → Omniverse-rendered fleet twin
twinex.simulate({ swap: "h100 → gb200", serving: "disaggregated", traffic: "2x" }); // predicts p99 · $/tok
H100 · static · p99 240ms
GB200 · disagg · p99 150ms
$/tok −38%
62% → 84%
What-if scenarios, safely.
Model hardware mix, NVIDIA Dynamo disaggregation, batch sizes, MIG partitioning, autoscaling, and routing policy — then compare predicted utilization, latency, and cost side by side before you commit.
- → Hardware mix + Dynamo + MIG scenarios
- → Predicted p99, throughput, utilization
- → Validate a Voltra policy before rollout
Plan capacity and budget with evidence.
Export predicted cost and capacity curves for finance and platform, with a target error band validated against a real run — so GPU purchases and topology changes are decisions, not bets.
- → Cost + capacity forecast curves
- → Prediction validated vs. real runs
- → Shareable plans for finance + platform
Twinex, in production.
One product. One platform.
Twinex predicts what Voltra will do under load. It simulates how Praxon policy changes affect fleet cost. It models Cuvex retrieval throughput at scale. It plans Edgeon fleet upgrades safely before OTA. Every product's future state, tested before you commit a dollar of compute.
Mandarum Voltra
Agentic scheduling and autoscaling that lifts GPU utilization and cuts $/token across H100/H200/Blackwell fleets.
Explore Voltra →Mandarum Praxon
Run autonomous agents with per-action identity, NeMo Guardrails, durable execution, and an immutable audit trail.
Explore Praxon →Mandarum Cuvex
GPU embedding, cuVS vector search, and reranking in one permission-aware recall path.
Explore Cuvex →Mandarum Edgeon
Serve governed AI on NVIDIA Jetson fleets with cloud-parity policy, telemetry, and safe OTA.
Explore Edgeon →Simulate first. Decide second. Regret never.
Model your first hardware swap or topology change before it touches production. Start free, or let the team build a calibrated twin from your DCGM traces.