We are building the universal GPU runtime for production AI.
As enterprises move from AI demos to production inference at scale, an entire layer of GPU infrastructure management becomes critical. Mandarum is the platform that operates that layer — efficiently, observably, and at enterprise scale.
Make GPU-powered AI trustworthy in production.
We believe the inference era will be defined by efficiency and governance — not raw model capability. Our job is to give every enterprise the runtime, identity, and policy infrastructure to serve AI models like first-class citizens of their GPU fleet.
What we hold true.
Reliability is a feature.
Demos work on A100s; production is H100s under pressure. We invest in benchmarks, traces, and recovery.
Policy before action.
An inference pipeline without guardrails is a liability. Approval and NeMo Guardrails are first-class primitives.
Open by default.
Bring any model, any framework, any GPU. Lock-in lives in SLAs, not protocols.
Design is the moat.
In an era of commodity intelligence, craft is the differentiator.
Ship weekly.
Velocity compounds. Enterprises that adopt Mandarum get a new product every Friday.
Customer outcomes > ours.
We measure ourselves in tokens served, GPU utilization, and inference cost reduction for the people who pay us.
Seed-funded by Khudi Ventures.
Mandarum raised a $3M seed round led by Khudi Ventures to accelerate the GPU-native runtime for production AI — built on Triton, TensorRT-LLM, NIM, Dynamo, NeMo, and Jetson at the edge.
The engineers building the GPU-native AI runtime.
Abdullah Usama
CEO & Co-Founder
LLM & AI Engineer. Built production computer vision and RAG systems at RapidsAI. Expert in deep learning, machine learning, and GPU-accelerated inference pipelines.
NUST · B.Sc. Computer Software Engineering · ex-RapidsAI · ex-OneScreen
Safa Gul
CTO
Full-stack software and ML engineer. Specialises in agentic AI, LLMs, and production AI systems. Dean's Honor List 2023 & 2024. Currently building at FlyRank AI.
University of Lahore · B.E. Computer Software Engineering · ex-FlyRank AI
Shah Zeb
Key Engineer · Computer Vision & AI
AI Engineer with 3+ years building end-to-end ML and computer vision systems — object detection, OCR, facial recognition, real-time video analytics, and LLM deployment across cloud and edge.
COMSATS Islamabad · MSc Computer Engineering · ex-Elexoft (2.5 yrs)