Job Title:  Senior Consultant | ITSM | Bengaluru | Engineering | Platform Development & Integration

Senior Consultant | ITSM | Bengaluru | Engineering | Platform Development & Integration
Job requisition ID : 109992 
Location: Bengaluru
Entity: Deloitte Touche Tohmatsu India LLP 

Senior Consultant |  Engineering | Platform Development & Integration | ITSM 

 

Location: Bangalore 

 

The team

Our Enterprise Technology & Performance team helps organizations build, operate, and optimize resilient cloud-native platforms. We are looking for an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate should possess strong expertise in production support, Site Reliability Engineering (SRE), cloud technologies, and ITIL-based service management, preferably with experience in Banking & Financial Services, specifically Cards & Payments and Mobile Applications. Enterprise technology has to do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth. Learn more about: Customer 

 

Role Summary

We are seeking an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, major incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate will be responsible for ensuring secure, scalable, highly available, and cost-effective cloud operations while driving operational excellence, service reliability, and continuous improvement across mission-critical enterprise applications. This role requires strong expertise in ITIL-based service management, Site Reliability Engineering (SRE), AWS cloud technologies, production support, automation, and stakeholder management. Experience in the Banking & Financial Services domain, particularly Cards & Payments, Mobile Applications, and Cloud-Native Solutions, will be highly advantageous.

  • Lead enterprise-wide Incident, Problem, and Change Management activities aligned with ITIL best practices.

  • Own the end-to-end lifecycle of production incidents (P1–P4), ensuring timely identification, escalation, communication, resolution, and closure.

  • Drive service recovery, Root Cause Analysis (RCA), Post Incident Reviews (PIR), and Corrective & Preventive Actions (CAPA) to improve service reliability and reduce MTTR.

  • Lead 24x7 production support operations, managing L2/L3 application and infrastructure support across Linux-based environments.

  • Oversee production deployments, release management, infrastructure operations, and Site Reliability Engineering (SRE) initiatives.

  • Manage Disaster Recovery (DR), High Availability (HA) architecture, and Active-Active/Active-Passive failover strategies.

  • Drive automation, cloud infrastructure optimization, capacity planning, and operational excellence using AWS, Kubernetes, Docker, and CI/CD pipelines.

  • Monitor and optimize production environments using observability tools including Grafana, Kibana, ELK, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7.

  • Support REST API-based applications, perform production troubleshooting using SQL, and collaborate with engineering teams to ensure highly resilient and secure production environments.

  • Lead and mentor SRE and Production Support teams while ensuring SLA, SLO, KPI, and customer satisfaction targets are consistently achieved.

 

Key Skills Required

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • 10–15+ years of overall IT experience with at least 8+ years specializing in Enterprise Production Support, Incident Management, Cloud Infrastructure, and Site Reliability Engineering (SRE).
  • Strong experience in Banking & Financial Services, preferably supporting Cards & Payments, Mobile Applications, and Cloud-Native solutions.
  • Proven expertise in Incident, Problem & Change Management (ITIL), Major Incident Management, Service Recovery, Root Cause Analysis (RCA), and Production Operations.
  • Experience leading 24x7 production support teams and managing L2/L3 application and infrastructure support across Linux environments.
  • Strong knowledge of AWS services including EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, CloudWatch, and secure cloud architecture.
  • Experience designing highly available, scalable, secure, and disaster recovery-enabled cloud solutions.
  • Hands-on experience with Docker, Kubernetes, Jenkins, CI/CD pipelines, Infrastructure as Code (Terraform, AWS CloudFormation, Ansible), and automation using Python or Bash.
  • Strong understanding of Linux administration, networking, load balancing, SQL/NoSQL databases, REST APIs, and cloud security best practices.
  • Experience with monitoring and observability tools including Grafana, Kibana, ELK Stack, Splunk, CloudWatch, Prometheus, Nagios, Zenduty, and Site24x7.
  • Proven ability to drive SLA/SLO/KPI compliance, reduce MTTR, improve service reliability, and implement automation and continuous improvement initiatives.
  • Strong leadership, stakeholder management, client communication, analytical, and problem-solving skills with the ability to coordinate cross-functional teams during critical incidents.
  • ITIL Foundation Certification is preferred.
  • AWS Certified Solutions Architect – Associate or Professional certification is highly preferred.