Job Title: Technology & Transformation | Senior Consultant | AI Engineer | Bangalore, Delhi, Mumbai | DigiGov

Technology & Transformation | Senior Consultant | AI Engineer | Bangalore, Delhi, Mumbai | DigiGov
• Job requisition ID : 107512
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
The team
Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information that can support fact-driven decision-making and generate an insight-driven advantage.
Your work profile
We are looking for a highly skilled AI Engineer to design, build, and operationalize AI-assisted developer productivity solutions. The ideal candidate will have strong expertise in Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), AI Agents, and enterprise-scale AI deployments.
You will be responsible for building secure, scalable, and production-ready AI solutions, with a focus on code generation, developer productivity, knowledge retrieval, and sovereign AI deployments for regulated industries
Key skills required:
- Experience, with at least 3+ years specifically on LLM /generative-AI systems in production
- Experience deploying AI dev tooling in regulated environments like BFSI, government, defence, healthcare
- Familiarity with sovereign / air-gapped AI, on-prem Ollama / vLLM clusters
- Experience with India-specific concerns: data localisation under DPDP Act 2023, MeitY / CERT-In guidelines
- AI Engineer to design, build, and operationalise AI-assisted developer productivity solutions.
- Evaluate AI coding assistants across accuracy, latency, cost, IP / data posture, and IDE coverage
- Design reference architectures for code-AI in enterprise, including air-gapped / sovereign deployments behind VPC or on-prem
- Fine-tune / LoRA-adapt open-weight code models (Qwen2.5-Coder, DeepSeek-Coder, StarCoder2, CodeLlama) on internal codebases
- Build retrieval-augmented pipelines over private code repos using tree-sitter, AST-aware chunking, and vector search (Qdrant, Weaviate, pgvector)
- Design context-engineering strategies: repo maps, dependency graphs, semantic code search, MCP-based context provisioning
- Build coding agents that plan, execute, test, and self-correct using frameworks like LangGraph, CrewAI, AutoGen, or custom orchestration
Technical Skills
- Strong grasp of transformer architecture, attention mechanisms, tokenisation
- Fine-tuning: full FT, LoRA / QLoRA, DPO, SFT pipelines using PEFT, TRL, Axolotl, Unsloth
- Evaluation methodology: pass@k, exact-match vs functional correctness, LLM-as-judge with calibration
- Tree-sitter, LSP, AST manipulation, code embeddings (CodeBERT, UniXcoder, Voyage-code, Jina-code)
- Fill-in-the-middle (FIM) training and inference patterns
- Static analysis integration (Semgrep, CodeQL, SonarQube)
- PyTorch, Hugging Face Transformers, Datasets, Accelerate
- LangChain, LlamaIndex, DSPy, Haystack
- vLLM, SGLang, TGI for serving; Ray / Ray Serve for distributed workloads
- MCP (Model Context Protocol) for tool / context integration
Qualifications
- B.E. / B.Tech / M.Tech / MS in CS, AI / ML, or equivalent
Location and Way of Working:
Base location: Bengaluru, Delhi, Mumbai
