Sarvam AI

Performance Engineer, Inference

Full-time5+ yrsBengaluruNot disclosedApply by 9 Oct 2026
Apply now
WA

Overview

PERFORMANCE ENGINEER, INFERENCE Part of Sarvam's Performance Engineering team. We are hiring two specialized performance roles Inference (this posting) and Kernels (companion posting).

What you'll do

  • You should be source-level fluent in at least one of SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM
  • able to read and modify it where stock behavior does not fit our workloads
  • and you should operate and extend a distributed-serving stack at depth: disaggregated prefill-decode across nodes, distributed KV/cache transfer, and the routing and scheduling that span them.
  • You will integrate artifacts from the model and kernel teams into a running multi-node, multi-tenant stack, and you will build and train your own speculators
  • draft models, distillation from the target, acceptance-rate tuning against the live serving distribution
  • rather than only wiring in stock implementations.
  • You will produce, and defend, the latency and throughput numbers the company plans against, and you will spend significant time in cross-team work with architecture co-design, the kernels team, the model team, and SRE.
  • Your scoreboard: TTFT (p50 / p95 / p99), TPOT, throughput, GPU utilization, and cost per million tokens.

Requirements

  • Practical experience with Node, LLM, RAG, C++.
  • Experience level: 5+ yrs.
  • Strong written and verbal communication in English.
  • Comfortable working on-site in Bengaluru.

Skills

NodeLLMRAGC++SRE

Related roles