Senior AI Platform Engineer – MLOps & LLM Serving - Chennai (Hybrid)

Chennai, Tamil Nadu / Remote (Global)6-12 yrsPermanentHybridINR 23 - 25 LPA

Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.

Role: Senior AI Platform Engineer - MLOps & LLM Serving - Chennai (Hybrid)

Positions: 1

Experience: 6 to 12 years

Location(s): Remote (Global), Chennai

Type: Hybrid / Permanent

Salary: Up to INR 25 LPA

Notice Period: Immediate to 30 days


Key Responsibilities:

• Install, configure and operate OpenShift, NVIDIA GPU operator, OpenShift AI, and NIM microservices on 12× RTX PRO 6000 across two servers; single-node and HA control- plane topologies

• Serving configuration and tuning: quantized model deployment (FP8/FP4), replica balancing, batching, KV-cache and context management

• Azure GPU build environments: provisioning, cost control, parity with the on-prem stack via pinned container/model versions; cloud-to-factory migration with parity regression

• GitOps CI/CD, container registry, artifact/model versioning, environment promotion;observability and audit wiring (Splunk, Prometheus/Grafana)

• Benchmark automation: load harness, p50/p95/p99 latency, tokens/sec, GPU utilization; the capacity report data pipeline

• Platform upgrade procedure with evaluation-regression gates; deployment runbook as a first-class deliverable


Technical Skills:

• 6+ years infrastructure/platform engineering with 3+ years production Kubernetes; OpenShift experience strongly preferred

• Hands-on GPU inference serving in production: NIM, Triton, vLLM, or TensorRT-LLM you have sized, deployed, and tuned LLM serving on real GPUs and can talk memory- bandwidth trade-offs

• GitOps fluency (ArgoCD/Flux), infrastructure-as-code, container internals; comfortable in air-gapped/proxy-restricted enterprise networks

• Observability depth: metrics, traces, log pipelines; has built performance test harnesses, not just run them

• Azure or AWS GPU compute operations experience


Strongly preferred

• NVIDIA GPU operator and AI Enterprise stack specifics; KServe; Milvus or pgvector operations; VAST/NFS/S3 storage integration; banking or other regulated-environment delivery


Skills

ArgoCDAWS GPU InstancesAzure GPUDockerGPU ComputingGrafanaKubernetesLLM Inference & Performance TuningLLM Model ServingLLM Serving FrameworksNVIDIA GPUOpenShiftPrometheusRegulated Environments ComplianceSplunkTensorRT-LLMTerraformTriton Inference Server

Posted October 6, 2026