Clarity2Cloud Technology Private Limited
AI/ML Developer
Full-time1-3 yrsonline₹420000Apply by 24 Aug 2026
Overview
What You'll Own: Inference infrastructure: Stand up and operate the inference engine, deploying vLLM as the primary solution and evaluating TGI where model coverage requires it, across GPU capacity from India-region partners and burst/global providers. API gateway: Build and maintain the OpenAI-compatible API surface, including /v1/chat/completions, /v1/embeddings, and /v1/models, so existing Open
What you'll do
- What You'll Own: Inference infrastructure: Stand up and operate the inference engine, deploying vLLM as the primary solution and evaluating TGI where model coverage requires it, across GPU capacity from India-region partners and burst/global providers.
- API gateway: Build and maintain the OpenAI-compatible API surface, including /v1/chat/completions, /v1/embeddings, and /v1/models, so existing OpenAI SDK code can run against MIRA with a one-line endpoint change.
- Smart routing: Build MIRA's core differentiator, including the complexity classifier that scores incoming requests and routes them across Tier 1-3 models, the post-generation quality validation step, and the escalation path to frontier fallback providers such as Anthropic/OpenAI.
- Implement the logging required to tune routing thresholds using real traffic.
- Model catalogue: Deploy and manage the v1 model catalogue, including Llama, Qwen, Deepseek, Mistral, Mixtral, nomic-embed-text, and BGE-M3 embeddings, and evaluate Indic models such as Sarvam-1 for the roadmap.
- Platform integration: Integrate MIRA with auth.thq.digital SSO and the unified THQ credit wallet, including per-key usage metering for input/output tokens, model, tier, and escalation flag, which will serve as the billing source of truth.
- Observability & unit economics: Implement per-request logging, latency and error dashboards, and per-customer GPU cost attribution from day one.
- This data will support sustainable pricing decisions.
- Pilot delivery: Support benchmark data collection for the MeiTY 90-day pilot, including latency P50/P95, INR cost-per-million-tokens compared with AWS Bedrock/OpenRouter/Together AI, and OpenAI conformance testing.
- Contribute to the public Day-90 report.
Requirements
- Practical experience with Python, Go, FastAPI, AWS.
- Experience level: 1-3 yrs.
- Strong written and verbal communication in English.
- Comfortable working on-site in online.
Skills
PythonGoFastAPIAWSGCPKubernetesDockerLLMRAGScala