Job Title: Manager | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering

Manager | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering
• Job requisition ID : 111848
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
Delivery Manager | Engineering, AI & Data – Engineering | SRE
Location: Bangalore
The team
Engineering helps empower and drive mission-critical solutions whether we need to modernize existing systems or implement new technology products and platforms. Through innovation, we improve financial performance, accelerate new digital businesses and fuel growth. Learn more about Engineering, AI and Data
Your work profile
The SRE Operations Lead will be responsible for defining and establishing the SRE governance framework and operational model for the client’s environment. The role will work closely with client service delivery, engineering, operations and SRE stakeholders to assess current ways of working, identify gaps and opportunities, and design a scalable, standardized and outcome-oriented SRE operating model.
The successful candidate will combine SRE technical expertise, IT operations leadership, service management and consulting skills. They will be expected to translate the client's current operational challenges into a practical target operating model covering governance, roles and responsibilities, service reliability, incident/problem management, operational readiness, metrics and continual improvement.
The role will also provide senior leadership during the transition from the existing model to the proposed SRE-led operating model.
- 12+ years of experience in IT operations, SRE, DevOps, platform engineering or related disciplines.
- 5+ years of experience in SRE/DevOps/reliability leadership roles.
- Demonstrable experience designing an SRE governance framework and operating model for a large or complex organization.
- Proven experience conducting current-state assessments of client operating models/ways of working and translating findings into a target operating model.
- Experience defining SRE governance covering SLOs, SLIs, error budgets, reliability metrics, incident management and problem management.
- Strong experience in major incident management, problem management and RCA.
- Experience designing centralized, federated or hybrid SRE/Operations models.
- Experience working with multiple application/product teams and coordinating distributed SRE/engineering stakeholders.
- Experience in service transition and operational readiness.
- Strong understanding of ITIL/ITSM processes and how they integrate with modern SRE practices.
- Experience establishing operational KPIs and reliability dashboards.
- Strong stakeholder management experience, including working with senior client leadership.
Key skills required:
- Education-Bachelors in Engineering
- Experience with large-scale operations consolidation or managed-services transformation.
- Experience defining AI/AIOps and automation strategies for IT operations.
- Experience identifying and quantifying operational toil and automation opportunities.
- Experience with cloud platforms such as AWS, Azure or GCP.
- Experience with observability platforms and technologies such as Prometheus, Grafana, Datadog, Dynatrace, Splunk, OpenTelemetry, etc.
- Experience with Kubernetes/container platforms.
- Experience with CI/CD and infrastructure-as-code.
- Experience establishing enterprise SRE Centres of Excellence.
- SRE certification or relevant cloud/DevOps certifications.
- Key Competencies
- The ideal candidate should demonstrate a combination
- Deep understanding of SRE principles
- Reliability engineering
- Observability
- Incident/problem management
- SLO/SLI/error-budget practices
- Automation and toil reduction
- Cloud and modern application architectures
