Magnitudeminds AI Implementation

Careers/Engineering

LLM Inference Platform Engineer

Specialise in LLM serving at production scale. Tune vLLM and SGLang, design continuous batching and tensor-parallel topologies, and own the cost and latency of every token we serve.

Hybrid · Ho Chi Minh City, VietnamFull Time5+ years

About the role

We are hiring an LLM Inference Platform Engineer who lives at the intersection of model internals and serving infrastructure. This is a deep specialist role for someone who has spent time in the weeds of PagedAttention, continuous batching, KV-cache management and tensor parallelism. You will own the inference layer: the choice of serving engine, the tuning of every knob that matters, and the routing layer that sits in front. Your scorecard is tokens-per-second per GPU, p50/p95 time-to-first-token, and dollars-per-million-tokens, and you can move all three without a drop in quality.

Responsibilities

Inference Engine Ownership

  • Own the choice and tuning of the serving engine (vLLM, SGLang, TGI, TensorRT-LLM) across model families.
  • Configure continuous batching, PagedAttention, chunked prefill and speculative decoding for the right workload mix.
  • Design tensor-parallel, pipeline-parallel and data-parallel topologies for models that do not fit on a single GPU.
  • Profile and eliminate bottlenecks in tokenisation, attention kernels, KV-cache memory, and scheduler decisions.

Serving Gateway & Routing

  • Operate liteLLM (or an equivalent gateway) for multi-model, multi-provider routing with fallbacks and cost tracking.
  • Design rate-limit tiers, per-tenant quotas, priority queues and graceful degradation paths.
  • Implement observability for every request: TTFT, ITL, queue time, GPU wait, cache hit rate.
  • Build streaming patterns (SSE, WebSocket) with backpressure and cancellation semantics.

Performance & Cost Engineering

  • Run structured load tests and establish published SLOs (p50 TTFT, p95 ITL, throughput per GPU) per model endpoint.
  • Apply quantisation (AWQ, GPTQ, FP8) and model compression where quality budgets allow; measure impact rigorously.
  • Tune batch sizes, max-num-seqs, gpu-memory-utilisation and scheduler parameters against real traffic shapes.
  • Partner with finance on cost-per-token forecasting and commitment planning for GPU capacity.

Reliability & Evolution

  • Own production incidents for inference: capacity loss, degraded quality, cache pathologies, cold-starts.
  • Design canary, shadow and A/B harnesses for engine upgrades and model swaps without user-visible regressions.
  • Evaluate new serving engines and model architectures; publish benchmarks that inform platform decisions.
  • Collaborate with ML/AI Platform and AI Full-Stack teams on rollout, rollback and eval gating.

Qualifications

Must-Have Technical Expertise

  • 5+ years in systems engineering with 2+ years running LLM inference in production at meaningful scale.
  • Deep hands-on experience with vLLM (continuous batching, PagedAttention, tensor parallelism) or SGLang / TGI equivalents.
  • Strong Python; working C++ / CUDA fluency for reading kernels and profiling hot paths.
  • Expert-level Kubernetes familiarity for GPU workloads; experience with Ray Serve a plus.
  • Proficiency with load testing tools (vegeta, locust, k6) and inference-specific benchmarks.
  • Proficiency with AI-assisted development tools (Cursor, Claude Code, GitHub Copilot, or similar).

Performance & Systems

  • Demonstrated ability to move p50/p95 latency and throughput with evidence from real traffic.
  • Understanding of GPU memory layout, NCCL collectives, and distributed-inference collective costs.
  • Experience with quantisation formats and their accuracy/latency/throughput tradeoffs.
  • Comfort reading and contributing fixes to open-source inference engines.

Preferred/Bonus

  • Contributions to vLLM, SGLang, TGI, TensorRT-LLM or liteLLM.
  • Experience with speculative decoding, MoE routing or multi-LoRA serving.
  • Experience with NVIDIA Triton, Dynamo-Triton, or Kserve deployments.
  • Familiarity with Envoy AI Gateway or GKE Inference Gateway.
  • Strong Vietnamese and English communication skills.

Benefits

  • Competitive salary and performance incentives
  • Own production SLOs for every LLM token served
  • Direct engagement with the open-source inference community
  • Flexible work arrangements
  • A collaborative, innovative engineering team environment

Apply

Apply for LLM Inference Platform Engineer.

Send a CV and a short note on what you have shipped. We read every application and reply to all of them.

Apply by email