Job Title:  Deputy Manager | AI Engineer | Bengaluru | Sustainability & Emerging Assurance

Deputy Manager | AI Engineer | Bengaluru | Sustainability & Emerging Assurance
Job requisition ID : 107821 
Location: Bengaluru
Entity: Deloitte Touche Tohmatsu India LLP 

The Team

Assurance is about much more than just the numbers. It’s about attesting to accomplishments and challenges and helping to assure strong foundations for future aspirations. Deloitte exemplifies what, how, and why of change so you’re always ready to act ahead. Learn more about Audit & Assurance Practice.

 

Your work profile

 

We are seeking a highly skilled Generative AI Engineer to design, develop, and deploy intelligent AI-powered applications using Large Language Models (LLMs). As a key member of our team, you will build scalable AI solutions leveraging Retrieval-Augmented Generation (RAG), LangChain, LangGraph, prompt engineering, and caching strategies to deliver accurate, efficient, and production-ready AI systems. You will collaborate with cross-functional teams to develop AI agents, integrate enterprise data sources, optimize LLM performance, and create robust backend services that power next-generation AI experiences.

 

Job Responsibilities:

 

  • LLM Application Development: Design, develop, and integrate AI-powered applications using Large Language Models (LLMs) to automate workflows, enhance user experiences, and solve complex business problems.
  • RAG Pipeline Development: Build and optimize Retrieval-Augmented Generation (RAG) pipelines by integrating vector databases and enterprise knowledge sources to deliver accurate, context-aware, and grounded AI responses.
  • AI Agent Development: Develop intelligent, multi-step AI agents using LangGraph to orchestrate reasoning, tool usage, decision-making, and workflow automation for complex business processes.
  • LangChain Integration: Leverage LangChain to build scalable LLM workflows by connecting language models with external APIs, databases, documents, tools, and memory components for end-to-end AI solutions.
  • Prompt Engineering & Optimization: Design, test, and refine prompts and prompting strategies to maximize response quality, consistency, accuracy, and reliability across diverse AI use cases.
  • Caching & Performance Optimization: Implement caching mechanisms for prompts, embeddings, and LLM responses to reduce latency, optimize API utilization, lower operational costs, and improve application scalability.
  • AI Solution Deployment: Develop and deploy production-ready AI services using Python and modern backend frameworks, ensuring secure, scalable, and maintainable architectures across cloud environments.
  • Collaboration & Innovation: Partner with product managers, software engineers, data scientists, and cross-functional teams to deliver enterprise-grade Generative AI solutions while continuously evaluating emerging LLM technologies and best practices.