Job Title: Deputy Manager | AI Engineer | Bengaluru | Sustainability & Emerging Assurance

Deputy Manager | AI Engineer | Bengaluru | Sustainability & Emerging Assurance
• Job requisition ID : 107821
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
The team
Enterprise technology has to do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth. Learn more about ET&P
Your work profile
We are seeking a highly skilled Generative AI Engineer to design, develop, and deploy intelligent AI-powered applications using Large Language Models (LLMs). As a key member of our team, you will build scalable AI solutions leveraging Retrieval-Augmented Generation (RAG), LangChain, LangGraph, prompt engineering, and caching strategies to deliver accurate, efficient, and production-ready AI systems. You will collaborate with cross-functional teams to develop AI agents, integrate enterprise data sources, optimize LLM performance, and create robust backend services that power next-generation AI experiences.
Job Responsibilities:
- LLM Application Development: Design, develop, and integrate AI-powered applications using Large Language Models (LLMs) to automate workflows, enhance user experiences, and solve complex business problems.
- RAG Pipeline Development: Build and optimize Retrieval-Augmented Generation (RAG) pipelines by integrating vector databases and enterprise knowledge sources to deliver accurate, context-aware, and grounded AI responses.
- AI Agent Development: Develop intelligent, multi-step AI agents using LangGraph to orchestrate reasoning, tool usage, decision-making, and workflow automation for complex business processes.
- LangChain Integration: Leverage LangChain to build scalable LLM workflows by connecting language models with external APIs, databases, documents, tools, and memory components for end-to-end AI solutions.
- Prompt Engineering & Optimization: Design, test, and refine prompts and prompting strategies to maximize response quality, consistency, accuracy, and reliability across diverse AI use cases.
- Caching & Performance Optimization: Implement caching mechanisms for prompts, embeddings, and LLM responses to reduce latency, optimize API utilization, lower operational costs, and improve application scalability.
- AI Solution Deployment: Develop and deploy production-ready AI services using Python and modern backend frameworks, ensuring secure, scalable, and maintainable architectures across cloud environments.
- Collaboration & Innovation: Partner with product managers, software engineers, data scientists, and cross-functional teams to deliver enterprise-grade Generative AI solutions while continuously evaluating emerging LLM technologies and best practices.
