Clarity2Cloud Technology Private Limited

AI/ML Developer

Full-time1-3 yrsonline₹420000Apply by 24 Aug 2026
Apply now
WA

Overview

What You'll Own: Inference infrastructure: Stand up and operate the inference engine, deploying vLLM as the primary solution and evaluating TGI where model coverage requires it, across GPU capacity from India-region partners and burst/global providers. API gateway: Build and maintain the OpenAI-compatible API surface, including /v1/chat/completions, /v1/embeddings, and /v1/models, so existing Open

What you'll do

  • What You'll Own: Inference infrastructure: Stand up and operate the inference engine, deploying vLLM as the primary solution and evaluating TGI where model coverage requires it, across GPU capacity from India-region partners and burst/global providers.
  • API gateway: Build and maintain the OpenAI-compatible API surface, including /v1/chat/completions, /v1/embeddings, and /v1/models, so existing OpenAI SDK code can run against MIRA with a one-line endpoint change.
  • Smart routing: Build MIRA's core differentiator, including the complexity classifier that scores incoming requests and routes them across Tier 1-3 models, the post-generation quality validation step, and the escalation path to frontier fallback providers such as Anthropic/OpenAI.
  • Implement the logging required to tune routing thresholds using real traffic.
  • Model catalogue: Deploy and manage the v1 model catalogue, including Llama, Qwen, Deepseek, Mistral, Mixtral, nomic-embed-text, and BGE-M3 embeddings, and evaluate Indic models such as Sarvam-1 for the roadmap.
  • Platform integration: Integrate MIRA with auth.thq.digital SSO and the unified THQ credit wallet, including per-key usage metering for input/output tokens, model, tier, and escalation flag, which will serve as the billing source of truth.
  • Observability & unit economics: Implement per-request logging, latency and error dashboards, and per-customer GPU cost attribution from day one.
  • This data will support sustainable pricing decisions.
  • Pilot delivery: Support benchmark data collection for the MeiTY 90-day pilot, including latency P50/P95, INR cost-per-million-tokens compared with AWS Bedrock/OpenRouter/Together AI, and OpenAI conformance testing.
  • Contribute to the public Day-90 report.

Requirements

  • Practical experience with Python, Go, FastAPI, AWS.
  • Experience level: 1-3 yrs.
  • Strong written and verbal communication in English.
  • Comfortable working on-site in online.

Skills

PythonGoFastAPIAWSGCPKubernetesDockerLLMRAGScala

Related roles