Careers/Engineering
LLM Inference Platform Engineer
Specialise in LLM serving at production scale. Tune vLLM and SGLang, design continuous batching and tensor-parallel topologies, and own the cost and latency of every token we serve.
About the role
We are hiring an LLM Inference Platform Engineer who lives at the intersection of model internals and serving infrastructure. This is a deep specialist role for someone who has spent time in the weeds of PagedAttention, continuous batching, KV-cache management and tensor parallelism.
You will own the inference layer: the choice of serving engine, the tuning of every knob that matters, and the routing layer that sits in front. Your scorecard is tokens-per-second per GPU, p50/p95 time-to-first-token, and dollars-per-million-tokens, and you can move all three without a drop in quality.
Responsibilities
Inference Engine Ownership
- Own the choice and tuning of the serving engine (vLLM, SGLang, TGI, TensorRT-LLM) across model families.
- Configure continuous batching, PagedAttention, chunked prefill and speculative decoding for the right workload mix.
- Design tensor-parallel, pipeline-parallel and data-parallel topologies for models that do not fit on a single GPU.
- Profile and eliminate bottlenecks in tokenisation, attention kernels, KV-cache memory, and scheduler decisions.
Serving Gateway & Routing
- Operate liteLLM (or an equivalent gateway) for multi-model, multi-provider routing with fallbacks and cost tracking.
- Design rate-limit tiers, per-tenant quotas, priority queues and graceful degradation paths.
- Implement observability for every request: TTFT, ITL, queue time, GPU wait, cache hit rate.
- Build streaming patterns (SSE, WebSocket) with backpressure and cancellation semantics.
Performance & Cost Engineering
- Run structured load tests and establish published SLOs (p50 TTFT, p95 ITL, throughput per GPU) per model endpoint.
- Apply quantisation (AWQ, GPTQ, FP8) and model compression where quality budgets allow; measure impact rigorously.
- Tune batch sizes, max-num-seqs, gpu-memory-utilisation and scheduler parameters against real traffic shapes.
- Partner with finance on cost-per-token forecasting and commitment planning for GPU capacity.
Reliability & Evolution
- Own production incidents for inference: capacity loss, degraded quality, cache pathologies, cold-starts.
- Design canary, shadow and A/B harnesses for engine upgrades and model swaps without user-visible regressions.
- Evaluate new serving engines and model architectures; publish benchmarks that inform platform decisions.
- Collaborate with ML/AI Platform and AI Full-Stack teams on rollout, rollback and eval gating.
Qualifications
Must-Have Technical Expertise
- 5+ years in systems engineering with 2+ years running LLM inference in production at meaningful scale.
- Deep hands-on experience with vLLM (continuous batching, PagedAttention, tensor parallelism) or SGLang / TGI equivalents.
- Strong Python; working C++ / CUDA fluency for reading kernels and profiling hot paths.
- Expert-level Kubernetes familiarity for GPU workloads; experience with Ray Serve a plus.
- Proficiency with load testing tools (vegeta, locust, k6) and inference-specific benchmarks.
- Proficiency with AI-assisted development tools (Cursor, Claude Code, GitHub Copilot, or similar).
Performance & Systems
- Demonstrated ability to move p50/p95 latency and throughput with evidence from real traffic.
- Understanding of GPU memory layout, NCCL collectives, and distributed-inference collective costs.
- Experience with quantisation formats and their accuracy/latency/throughput tradeoffs.
- Comfort reading and contributing fixes to open-source inference engines.
Preferred/Bonus
- Contributions to vLLM, SGLang, TGI, TensorRT-LLM or liteLLM.
- Experience with speculative decoding, MoE routing or multi-LoRA serving.
- Experience with NVIDIA Triton, Dynamo-Triton, or Kserve deployments.
- Familiarity with Envoy AI Gateway or GKE Inference Gateway.
- Strong Vietnamese and English communication skills.
Benefits
- Competitive salary and performance incentives
- Own production SLOs for every LLM token served
- Direct engagement with the open-source inference community
- Flexible work arrangements
- A collaborative, innovative engineering team environment
Apply
Apply for LLM Inference Platform Engineer.
Send a CV and a short note on what you have shipped. We read every application and reply to all of them.