Sarvam AI

Performance Engineer, Kernels

Full-time5+ yrsBengaluruNot disclosedApply by 9 Oct 2026
Apply now
WA

Overview

PERFORMANCE ENGINEER, KERNELS Part of Sarvam's Performance Engineering team. We are hiring two specialized performance roles Kernels (this posting) and Inference (companion posting).

What you'll do

  • cuBLAS, cuDNN, FlashAttention, out-of-the-box Triton
  • leave performance on the table, you will author the custom CUDA, DSL-based, and PTX kernels that close the gap.
  • This is a hard, narrow, high-leverage role.
  • We hire engineers who have shipped kernels that beat published baselines on real workloads, not engineers who have used kernels.
  • When your code lands, production p99 moves, and you own the explanation of why.
  • 5+ years in ML systems, with 2+ years authoring production CUDA kernels.
  • You have a kernel in production that beat the prior baseline by a measurable margin.
  • CUDA at kernel-authoring level: thread-block sizing, shared-memory layout, warp primitives, async copies (cp.async, TMA), and MMA selection.

Requirements

  • Practical experience with Node, LLM, RAG, SRE.
  • Experience level: 5+ yrs.
  • Strong written and verbal communication in English.
  • Comfortable working on-site in Bengaluru.

Skills

NodeLLMRAGSRE

Related roles