About

We are building the universal GPU runtime for production AI.

As enterprises move from AI demos to production inference at scale, an entire layer of GPU infrastructure management becomes critical. Mandarum is the platform that operates that layer — efficiently, observably, and at enterprise scale.

Mission

Make GPU-powered AI trustworthy in production.

We believe the inference era will be defined by efficiency and governance — not raw model capability. Our job is to give every enterprise the runtime, identity, and policy infrastructure to serve AI models like first-class citizens of their GPU fleet.

Principles

What we hold true.

Reliability is a feature.

Demos work on A100s; production is H100s under pressure. We invest in benchmarks, traces, and recovery.

Policy before action.

An inference pipeline without guardrails is a liability. Approval and NeMo Guardrails are first-class primitives.

Open by default.

Bring any model, any framework, any GPU. Lock-in lives in SLAs, not protocols.

Design is the moat.

In an era of commodity intelligence, craft is the differentiator.

Ship weekly.

Velocity compounds. Enterprises that adopt Mandarum get a new product every Friday.

Customer outcomes > ours.

We measure ourselves in tokens served, GPU utilization, and inference cost reduction for the people who pay us.

Backed by

Seed-funded by Khudi Ventures.

Lead Investor · Seed Round Khudi Ventures $3,000,000 — Seed · 2026

Mandarum raised a $3M seed round led by Khudi Ventures to accelerate the GPU-native runtime for production AI — built on Triton, TensorRT-LLM, NIM, Dynamo, NeMo, and Jetson at the edge.

Team

The engineers building the GPU-native AI runtime.

We're hiring →
AU
Abdullah Usama

CEO & Co-Founder

LLM & AI Engineer. Built production computer vision and RAG systems at RapidsAI. Expert in deep learning, machine learning, and GPU-accelerated inference pipelines.

LLMs & RAG Computer Vision Deep Learning

NUST · B.Sc. Computer Software Engineering · ex-RapidsAI · ex-OneScreen

SG
Safa Gul

CTO

Full-stack software and ML engineer. Specialises in agentic AI, LLMs, and production AI systems. Dean's Honor List 2023 & 2024. Currently building at FlyRank AI.

Agentic AI LLMs Full-Stack

University of Lahore · B.E. Computer Software Engineering · ex-FlyRank AI

SZ
Shah Zeb

Key Engineer · Computer Vision & AI

AI Engineer with 3+ years building end-to-end ML and computer vision systems — object detection, OCR, facial recognition, real-time video analytics, and LLM deployment across cloud and edge.

Computer Vision PyTorch / TF MLOps & Docker

COMSATS Islamabad · MSc Computer Engineering · ex-Elexoft (2.5 yrs)