Job Title: Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering

Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering
• Job requisition ID : 111843
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
Team Lead | Engineering, AI & Data – Engineering | SRE-GCP Devops
Location: Bangalore
The team
Engineering helps empower and drive mission-critical solutions whether we need to modernize existing systems or implement new technology products and platforms. Through innovation, we improve financial performance, accelerate new digital businesses and fuel growth. Learn more about Engineering, AI and Data
Your work profile
We are looking for a highly skilled DevOps Team Lead to manage and scale mission-critical, production-grade distributed systems running on GCP. The ideal candidate will focus on reliability, automation, observability, and operational excellence while leading a team of engineers to minimize toil and improve system availability The goal is maintaining and improving uptime from 4 Nines to 5 Nines through engineering efforts.
This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration, and strong troubleshooting capabilities for distributed systems As a Team Lead, you will participate in the overall lifecycle management of mission-critical banking services with a 24x7 operations mode, while also mentoring team members and driving strategic initiatives.
- Team Management: Lead, mentor, and grow a team of DevOps/SRE engineers (Added for Lead role).
- Operational Excellence: Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and process enhancements
- Accountability: Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission-critical services with full accountability for upholding the SLOs
- Collaboration: Partner with engineering, operations, and cloud management teams to deliver highly reliable services in a timely manner
- Cloud & Infrastructure (GCP)
- Infrastructure Management: Design, deploy, and manage infrastructure on GCP
- Core Services: Work extensively on Kubernetes Engine, Compute, networking, IAM, Load Balancers, TLS Certs, BigQuery, Pub/Sub, cloud logging, metrics, and logs analysis
- IaC: Implement and manage infrastructure using Terraform (Infrastructure as Code)
- Orchestration: Deploy and manage containerized workloads using Kubernetes (GKE)
- Troubleshooting: Resolve issues related to pods, nodes, networking, storage, and services on an ongoing basis
- Deployments: Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green)
- Pipelines: Build and maintain CI/CD pipelines using Jenkins (pipeline-based, Groovy/Shell/Python scripting)
- Version Control: Strong experience in using GitHub as a PowerUser
- Scripting: Develop automation using Python and Shell scripting to reduce operational toil
- Tooling: Implement and manage monitoring systems using Dynatrace, Grafana, logs, and metrics explorer
- Proactive Monitoring: Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively
- Incident Management: Work alongside operations teams to identify, fix production incidents, and own problem resolution
- Tenure: 4-6 years of relevant and progressive experience in SRE / DevOps / Cloud Engineering
- Leadership: Prior experience leading or mentoring technical teams
- Production: Hands-on experience managing production-grade systems
- Cloud Platform: Strong expertise in GCP (Kubernetes Engine, VPC, IAM, Load Balancing, KMS, logs, metrics, BigQuery, Pub/Sub)
- Infrastructure as Code: Strong hands-on experience with Terraform; ability to write and debug code from scratch
- Containers: Deep expertise in Kubernetes (GKE) and Docker with strong troubleshooting experience
- CI/CD: Hands-on experience with Jenkins and GitHub
- Scripting: Strong skills in Python (preferred) and Shell scripting
- Observability: Experience with Dynatrace / Grafana and log/metrics/trace-based monitoring
- Programming: Working knowledge of Java and/or Golang applications
- Networking: Strong Linux fundamentals and deep understanding of TCP/IP networking
Key Skills Required
Education-Any Bachelors
- Solid understanding of SLI, SLO, SLA, and Error Budgets
- Demonstrable experience improving MTTR and MTTA
- Experience handling the incident management lifecycle
Soft Skills
- Strong analytical and troubleshooting mindset
- Excellent communication and stakeholder management
- Ability to work in high-pressure production environments
- Ownership-driven and proactive approach
Preferred Qualifications
- Experience in high-scale distributed systems
- Exposure to banking/financial domain
- Understanding of security and compliance practices
- Experience with deployment strategies: Canary, Blue-Green
Ideal Candidate Profile
- Strong GCP + Kubernetes + Terraform core
- Hands-on production troubleshooting expert
- Good at automation + reducing toil
- Deep understanding of SRE principles
- Comfortable in 24x7 production environments
- Leadership: Ability to drive reliability engineering practices and culture across the team
