Sarvam AI
Performance Engineer, Kernels
Full-time5+ yrsBengaluruNot disclosedApply by 9 Oct 2026
Overview
PERFORMANCE ENGINEER, KERNELS Part of Sarvam's Performance Engineering team. We are hiring two specialized performance roles Kernels (this posting) and Inference (companion posting).
What you'll do
- cuBLAS, cuDNN, FlashAttention, out-of-the-box Triton
- leave performance on the table, you will author the custom CUDA, DSL-based, and PTX kernels that close the gap.
- This is a hard, narrow, high-leverage role.
- We hire engineers who have shipped kernels that beat published baselines on real workloads, not engineers who have used kernels.
- When your code lands, production p99 moves, and you own the explanation of why.
- 5+ years in ML systems, with 2+ years authoring production CUDA kernels.
- You have a kernel in production that beat the prior baseline by a measurable margin.
- CUDA at kernel-authoring level: thread-block sizing, shared-memory layout, warp primitives, async copies (cp.async, TMA), and MMA selection.
Requirements
- Practical experience with Node, LLM, RAG, SRE.
- Experience level: 5+ yrs.
- Strong written and verbal communication in English.
- Comfortable working on-site in Bengaluru.
Skills
NodeLLMRAGSRE