Mandarum Edgeon — Your runtime, at the edge.
A compact runtime for NVIDIA Jetson that serves TensorRT-optimized models and NIM microservices on-device, enforces the same identity, policy, and audit as the cloud, streams DCGM and OpenTelemetry back to one control plane, and ships signed, staged OTA model updates via Fleet Command.
Why edge AI is the last ungoverned frontier.
Every device is a one-off
Edge AI today is bespoke per site, with no shared runtime. Edgeon brings one governed serving runtime to thousands of Jetson nodes.
No identity or audit off-cloud
Models at the edge usually escape policy and logging. Edgeon enforces Praxon-grade identity, guardrails, and audit on-device.
Model rollouts brick fleets
Pushing a bad model to the field is expensive. Edgeon does signed, staged OTA rollouts with instant rollback via Fleet Command.
What Edgeon does.
TensorRT + NIM, on the device.
Pull TensorRT-optimized models and NIM microservices and serve them locally with hardware-accelerated DeepStream vision pipelines — low-latency and offline-capable when the network drops.
- → TensorRT-compiled models on Jetson Orin/Thor
- → DeepStream + Metropolis vision pipelines
- → Offline-capable local inference
await edgeon.deploy({ model: "trt/yolov9-c", target: "jetson-orin-nx", sites: 1240 }); // signed OTA · staged
Cloud-parity policy and telemetry.
Edge nodes carry the same scoped identity, policy-as-code, and immutable audit as data-center serving, and stream DCGM plus OpenTelemetry back to one control plane for fleet-wide visibility.
- → Identity + policy + audit on-device
- → DCGM + OTel streamed to one plane
- → IGX Orin for regulated deployments
Safe OTA across the whole fleet.
Roll a new model to a canary cohort, watch live accuracy and latency, then promote or roll back across thousands of devices — signed end to end and managed through NVIDIA Fleet Command.
- → Signed, staged OTA via Fleet Command
- → Canary cohorts with live evals
- → One-click fleet-wide rollback
Edgeon, in production.
One product. One platform.
Edgeon brings the runtime to the edge. Voltra packs the fleet it runs on. Praxon governs every action it takes. Cuvex gives it GPU memory for on-device grounding. Twinex simulates OTA rollouts and hardware upgrades before they touch the field. Cloud to factory floor — one platform.
Mandarum Voltra
Agentic scheduling and autoscaling that lifts GPU utilization and cuts $/token across H100/H200/Blackwell fleets.
Explore Voltra →Mandarum Praxon
Run autonomous agents with per-action identity, NeMo Guardrails, durable execution, and an immutable audit trail.
Explore Praxon →Mandarum Cuvex
GPU embedding, cuVS vector search, and reranking in one permission-aware recall path.
Explore Cuvex →Mandarum Twinex
Simulate GPU inference fleets on Omniverse to plan cost, capacity, and serving topology before production.
Explore Twinex →Govern every edge node like it's the data center.
See how Edgeon governs your Jetson fleet with the same identity, policy, and audit as your H100 cluster — from first deploy to fleet-wide OTA. Start free, or schedule an edge pilot.