Sarvam AI
Infrastructure SRE - HPC
Full-time5+ yrsBengaluruNot disclosedCloses in 7 days · 25 Aug 2026
Overview
About Sarvam Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures.
What you'll do
- Keeping both healthy at once is a hard, specialized reliability problem, and it is the problem this team exists to solve.
- This is not a Kubernetes administration role.
- We assume Kubernetes fluency as a baseline.
- The difficulty lies above and below it
- in parallel filesystems under heavy checkpoint load, in RDMA fabrics that degrade quietly, in NCCL hangs whose root cause may be the network or the kernel, in driver and firmware drift across heterogeneous hardware, and in distributed training failures that masquerade as infrastructure faults.
- We are hiring a team of specialists rather than a set of identical generalists.
- This posting covers five areas of focus.
Requirements
- When you apply, please indicate the area of focus that best matches your experience.
- Strong generalists are welcome; we will place you where your depth is most useful.
- Operate the GPU fleet end to end across training and serving
- provisioning, observability, capacity, and fleet health.
- Hold a meaningful on-call rotation, write runbooks that hold up under pressure, and drive postmortems that produce durable fixes.
- Build the internal tooling the team relies on, rather than operating off-the-shelf systems alone.
- Partner with ML and platform teams to keep large runs alive and serving latency predictable.
- 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.*
Skills
NodePythonGoKubernetesRAG