Senior Gen AI Engineer- RAG & Full Stack - Chennai

Chennai, Tamil Nadu / Remote (Global)6-10 yrsPermanentRemote, HybridINR 20 - 25 LPA

Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.

Role: Senior Gen AI Engineer- RAG & Full Stack - Chennai

Positions: 1

Experience: 6 to 10 years

Location(s): Chennai (Hybrid)

Type: Hybrid, Remote / Permanent

Salary: Up to INR 25 LPA (Best as per the fitment)

Notice Period: Immediate to 15 days


Key Responsibilities:

• RAG pipeline end to end: ingestion integration, hybrid retrieval with reranking, prompt/context strategy, citation resolution, refusal behavior

• Permission-aware retrieval: source ACL mapping (SharePoint/Entra, Confluence) to fail-closed retrieval filters; zero-leakage test suite partnership with QA

• Vector database design and operations (Milvus or pgvector): schema, metadata filters,sync, performance

• Full-stack product build: React/TypeScript chat and citation experience, Python/FastAP services, REST APIs, SSO/OIDC integration, admin configuration UI

• Evaluation-driven development: retrieval precision, faithfulness, and citation-accuracy metrics as the daily working loop; A/B testing of retrieval and prompt variants

• Latency engineering to the 3–5s first-token / ~15s complete-answer targets at concurrency


Technical Skills:

• 5+ years software engineering with 2+ years building RAG/LLM applications in production with real users and real quality metrics, not notebooks

• Deep retrieval craft: chunking strategy, embeddings, hybrid search, rerankers; you canexplain why retrieval fails and how you measured the fix

• Genuine full-stack evidence: shipped React/TypeScript front ends AND Python back-end services in production; API design; OIDC/SAML integration

• Vector database production experience (Milvus, pgvector, Weaviate, or equivalent) including permission/metadata filtering

• Evaluation fluency: has built or operated a retrieval/answer quality harness with numeric thresholds


Strongly Preferred:

• Permission-aware/multi-tenant retrieval specifically; Microsoft Graph API; NIM/OpenAI-compatible serving endpoints; streaming UX; enterprise design systems; banking content domains

Skills

ConfluenceEmbeddingsFastAPIFull Stack DevelopmentGenAILarge Language Models (LLM)Microsoft Entra IDMicrosoft Graph APIPythonRAG (Retrieval-Augmented Generation)RAG EvaluationReactRerankingSharePointStreamingTypeScriptVector DB

Posted October 5, 2026