Magnitudeminds AI Implementation

Careers/Engineering

Senior ML/AI Platform Engineer

Build the ML platform that trains, serves and monitors our models. Own Ray/KubeRay clusters, model registries, inference infrastructure, and the operator experience for ML teams.

Hybrid · Ho Chi Minh City, VietnamFull Time6+ years

About the role

We are looking for a Senior ML/AI Platform Engineer to own the foundation that every model in the company runs on. This is the person our AI engineers and researchers depend on for reliable training jobs, fast inference and predictable cost. You will operate Ray clusters on Kubernetes, design model-serving infrastructure on top of vLLM and liteLLM, build the model registry and artifact pipeline, and shape the developer experience that lets AI teams move from experiment to production without a platform on-call interrupt. This role is open at Senior or Staff level depending on experience.

Responsibilities

Training & Compute Infrastructure

  • Operate KubeRay clusters (RayCluster, RayJob, RayService) for distributed training, fine-tuning, and batch inference.
  • Design GPU scheduling, node pool topology, and gang-scheduling with Kueue / NVIDIA KAI for multi-tenant fairness.
  • Build self-service training job submission: templates, quota, secrets, artifact output to a shared registry.
  • Instrument GPU utilisation, throughput and cost-per-job dashboards for AI teams and finance.

Inference Platform

  • Stand up and tune vLLM, SGLang and TGI deployments as the default serving layer.
  • Own the liteLLM gateway: routing, fallbacks, rate-limiting, cost-tracking across self-hosted and third-party models.
  • Define deployment patterns (canary, blue/green, shadow) for model rollouts with rollback guarantees.
  • Partner with AI Full-Stack engineers on latency, throughput and cost SLOs per endpoint.

Developer Experience & MLOps

  • Design the ML platform SDK (Python) that AI engineers use to launch training, register models and deploy endpoints.
  • Build a model registry + artifact pipeline (MLflow, custom store) with lineage, promotion workflows and approval gates.
  • Create reproducible training environments: container images, dependency pinning, CUDA compatibility matrices.
  • Own CI/CD for model code + prompts + evals with regression-aware promotion pipelines.

Reliability & Governance

  • Define SLOs for training job success rate, inference latency and availability; own post-mortems for platform incidents.
  • Implement budget caps and circuit breakers for runaway experiments or inference spend.
  • Design access control and audit trails for model artifacts, training data and inference logs.
  • Contribute to the platform roadmap and drive build-vs-buy decisions on ML infrastructure.

Qualifications

Must-Have Technical Expertise

  • 6+ years of platform / infrastructure engineering experience, with 2+ years operating ML workloads in production.
  • Strong Python and Kubernetes fundamentals (Helm, operators, custom resources, RBAC, networking).
  • Hands-on production experience with Ray / KubeRay: RayCluster, RayJob, RayService, autoscaler configuration.
  • Direct experience running vLLM or a comparable inference server in production (TGI, SGLang, Triton).
  • Working knowledge of liteLLM, Ray Serve, or another multi-model gateway.
  • Experience with GPU scheduling, NVIDIA drivers, and cost-aware capacity planning.
  • Proficiency with AI-assisted development tools (Cursor, Claude Code, GitHub Copilot, or similar).

Platform & ML Systems

  • Familiarity with model training frameworks: PyTorch, Hugging Face Accelerate, DeepSpeed.
  • Experience designing or operating a model registry and artifact store (MLflow, W&B Registry, or custom).
  • Understanding of parameter-efficient fine-tuning, quantisation (AWQ/GPTQ/GGUF) and their deployment tradeoffs.
  • Comfort owning on-call for GPU-backed infrastructure and debugging incidents under load.

Preferred/Bonus

  • Contributions to open-source ML infrastructure (Ray, vLLM, KubeRay, Kueue, liteLLM).
  • Experience with Kueue, Volcano or NVIDIA KAI scheduler for gang scheduling.
  • Experience with Databricks MLflow, Weights & Biases, or ClearML at platform scale.
  • Exposure to multi-cluster / multi-cloud GPU capacity management.
  • Strong Vietnamese and English communication skills.

Benefits

  • Competitive salary and performance incentives
  • Shape the ML platform standards every AI engineer in the company depends on
  • Direct work with open-source tooling (Ray, vLLM, KubeRay, liteLLM)
  • Flexible work arrangements
  • A collaborative, innovative engineering team environment

Apply

Apply for Senior ML/AI Platform Engineer.

Send a CV and a short note on what you have shipped. We read every application and reply to all of them.

Apply by email