Sarvam AI
Performance Engineer, Inference
Full-time5+ yrsBengaluruNot disclosedApply by 9 Oct 2026
Overview
PERFORMANCE ENGINEER, INFERENCE Part of Sarvam's Performance Engineering team. We are hiring two specialized performance roles Inference (this posting) and Kernels (companion posting).
What you'll do
- You should be source-level fluent in at least one of SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM
- able to read and modify it where stock behavior does not fit our workloads
- and you should operate and extend a distributed-serving stack at depth: disaggregated prefill-decode across nodes, distributed KV/cache transfer, and the routing and scheduling that span them.
- You will integrate artifacts from the model and kernel teams into a running multi-node, multi-tenant stack, and you will build and train your own speculators
- draft models, distillation from the target, acceptance-rate tuning against the live serving distribution
- rather than only wiring in stock implementations.
- You will produce, and defend, the latency and throughput numbers the company plans against, and you will spend significant time in cross-team work with architecture co-design, the kernels team, the model team, and SRE.
- Your scoreboard: TTFT (p50 / p95 / p99), TPOT, throughput, GPU utilization, and cost per million tokens.
Requirements
- Practical experience with Node, LLM, RAG, C++.
- Experience level: 5+ yrs.
- Strong written and verbal communication in English.
- Comfortable working on-site in Bengaluru.
Skills
NodeLLMRAGC++SRE