Job Title:  Manager | ITSM | Bengaluru | Engineering | Platform Development & Integration

Manager | ITSM | Bengaluru | Engineering | Platform Development & Integration
Job requisition ID : 109995 
Location: Bengaluru
Entity: Deloitte Touche Tohmatsu India LLP 

Lead Production IM | Enterprise Technology & Performance | Cloud Operations (AWS)

 

Location: Bangalore

 

The team

Our Enterprise Technology & Performance team helps organizations build, operate, and optimize resilient cloud-native platforms. We are looking for an experienced Lead Production Incident Manager (IM) to lead enterprise production operations, incident management, and cloud infrastructure reliability across large-scale AWS environments. The ideal candidate should possess strong expertise in production support, Site Reliability Engineering (SRE), cloud technologies, and ITIL-based service management, preferably with experience in Banking & Financial Services, specifically Cards & Payments and Mobile Applications. Enterprise technology has to do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth. Learn more about: Customer 

Role Summary

We are seeking an experienced Lead Production IM responsible for managing secure, scalable, and highly available cloud infrastructure on AWS while leading enterprise production support and incident management. The role involves driving service reliability, production governance, automation, disaster recovery, and operational excellence across mission-critical applications.

  • Lead enterprise-wide Incident, Problem, and Change Management activities following ITIL best practices.
  • Own end-to-end lifecycle management of production incidents (P1–P4), ensuring timely resolution, communication, escalation, and closure.
  • Lead service recovery activities and restore critical business services within agreed SLAs.
  • Act as the primary communication bridge between technical teams, business stakeholders, and clients during production outages.
  • Conduct war rooms, bridge calls, and cross-functional coordination during major incidents.
  • Drive Root Cause Analysis (RCA), Post Incident Reviews (PIR), and Corrective & Preventive Actions (CAPA) to improve service reliability.
  • Identify recurring incidents and collaborate with engineering teams to reduce MTTR, operational toil, and incident recurrence.
  • Lead 24x7 production support operations for enterprise applications and cloud infrastructure.
  • Manage L2/L3 Infrastructure and Application Support across Linux-based environments.
  • Oversee application deployments, release management, production support, and infrastructure operations.
  • Drive Site Reliability Engineering (SRE) initiatives focused on automation, monitoring, observability, and platform reliability.
  • Manage and mentor teams of SREs and Production Support Engineers while ensuring SLA, KPI, SLO, and MTTR compliance.
  • Lead Disaster Recovery (DR) planning, Active-Active and Active-Passive failover activities, and High Availability architecture.
  • Manage Kubernetes and Docker-based container platforms for deployment, scaling, and workload orchestration.
  • Monitor production environments using ELK, Kibana, Grafana, CloudWatch, Splunk, Prometheus, Nagios, Zenduty, and Site24x7.
  • Support REST API-based applications and perform deployment validation activities including API verification, cache management, and health checks.
  • Execute complex SQL (DDL/DML) queries for production troubleshooting and application support.
  • Drive cloud infrastructure optimization, automation, capacity planning, and cost optimization initiatives.
  • Monitor customer satisfaction metrics (CSAT/NPS) and implement continuous service improvement initiatives.
  • Collaborate with development, infrastructure, cloud, networking, and business teams to deliver highly resilient production environments.

 

Key skills required

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.

  • 10–15+ years of overall IT experience with at least 8+ years of experience in Enterprise Production Support, Incident Management, Cloud Infrastructure, and Site Reliability Engineering (SRE).

  • Strong experience in Banking & Financial Services domain, preferably Cards & Payments, Mobile Applications, and Cloud Native solutions.

  • Hands-on expertise in Incident, Problem & Change Management (ITIL) and Major Incident Management.

  • Strong knowledge of AWS cloud services including EC2, S3, RDS, Lambda, VPC, IAM, DynamoDB, and CloudWatch.

  • Experience designing and managing highly available, scalable, secure, and disaster recovery-enabled cloud architectures.

  • Hands-on experience with Docker, Kubernetes, Jenkins, CI/CD pipelines, Infrastructure as Code (Terraform, CloudFormation, Ansible), and automation using Python or Bash.

  • Strong understanding of Linux administration, networking, load balancing, SQL/NoSQL databases, REST APIs, and cloud security best practices.

  • Experience with monitoring and observability tools including Grafana, Kibana, ELK Stack, Splunk, Prometheus, Nagios, Zenduty, and Site24x7.

  • Proven expertise in L2/L3 Production Support, Root Cause Analysis (RCA), Post Incident Reviews (PIR), Service Recovery, and Production Operations.

  • Experience leading 24x7 support teams, driving SLA/SLO/KPI compliance, MTTR reduction, automation initiatives, and operational excellence.

  • Strong stakeholder management, client communication, leadership, analytical, and problem-solving skills with the ability to coordinate cross-functional teams during critical incidents.

  • ITIL Foundation Certification is preferred.

  • AWS Certified Solutions Architect – Associate or Professional certification is highly preferred.