Job Title: Senior Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineerin
Senior Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineerin(109996)
Role Overview
We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on AWS. The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts.
This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration and TCP/IP fundamentals, and strong troubleshooting capabilities for distributed systems. The candidate needs to participate in the overall lifecycle management of mission critical banking services with a 24x7 operations mode in an rotational on-call basis. The job requires the candidate to have strong troubleshooting skills in a distributed environment spread across multiple cloud environments. The bare minimum ask would be to maintain high level of agility, learnability and adaptability in different scenarios. An engineer with a zeal to learn fast and having a bias for action would be the best fit for the role.
Key Responsibilities
Reliability & Operations
- Own end-to-end production systems reliability, availability, scalability, cost and performance.
- Troubleshooting EKS from SRE perspective
- Route53, Connectivity, Networking, RDS, S3
- Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
- Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
- Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.
Cloud & Infrastructure
- Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
- Work extensively on:
- GKE (Kubernetes Engine)
- Compute, networking, IAM, Load Balancers, TLS Certs
- BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
- Implement and manage infrastructure using Terraform (Infrastructure as Code).
Kubernetes & Containers
- Deploy and manage containerized workloads using Kubernetes (GKE).
- Troubleshoot issues related to:
- Pods, nodes, networking, storage, services on an ongoing basis
- Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).
Automation & CI/CD
- Build and maintain CI/CD pipelines using:
- Jenkins (pipeline-based, Groovy / Shell / Python scripting)
- Strong experience in using GitHub as a PowerUser
- Develop automation using Python and Shell scripting.
- Reduce operational toil through automation initiatives.
Observability & Monitoring
- Implement and manage monitoring systems using:
- Dynatrace, Grafana, logs and metrics explorer
- Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively.
- Define alerting strategies based on system behaviour and SLOs and create runbooks.
- Work alongside operations teams to identify, fix the production incidents and own the problem resolution.
- Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents.
System & Application Troubleshooting
- Perform deep troubleshooting for:
- Distributed systems
- Microservices-based architectures on containerised workloads
- Java and Golang applications
- Strong debugging of:
- Application issues
- Infrastructure issues
- Network-related problems
Plan and execute continuous improvement
- Identify and eliminate repetitive manual tasks.
- Drive reliability engineering practices and culture. (DRY – Don’t Repeat Yourself)
- Collaborate with development teams to improve system design and resilience.
Job Requirements
- 3–5 years of relevant and progressive experience in SRE / DevOps / Cloud Engineering
- Hands-on experience managing production-grade systems (24x7 environments)
Technical Skills
- Troubleshooting EKS from SRE perspective
- Route53, Connectivity, Networking, RDS, S3
Infrastructure as Code
- Strong hands-on experience with Terraform
- Ability to write and debug Terraform code from scratch
Containers & Orchestration
- Deep expertise in:
- Kubernetes (GKE)
- Docker
- Strong troubleshooting experience in Kubernetes environments
CI/CD & Automation
- Hands-on experience with:
- Jenkins (pipeline-based CI/CD)
- GitHub
- Strong scripting skills:
- Python (preferred)
- Shell scripting
- Experience with automation frameworks and tooling
Observability
- Experience with:
- Dynatrace / Grafana
- Log, metrics, and trace-based monitoring
Programming & Debugging
- Working knowledge of:
- Java and/or Golang applications
- Strong debugging skills across application and infrastructure layers
Linux & Networking
- Strong Linux fundamentals
- Deep understanding of TCP/IP networking
- Ability to debug network issues in distributed systems
Reliability Engineering Skills
- Solid understanding of:
- SLI, SLO, SLA, Error Budgets
- Demonstrable and Proven Experience improving:
- MTTR, MTTA
- Experience handling incident management lifecycle
Soft Skills
- Strong analytical and troubleshooting mindset
- Excellent communication and stakeholder management
- Ability to work in high-pressure production environments
- Ownership-driven and proactive approach
Preferred Qualifications
- Experience in high-scale distributed systems
- Exposure to banking/financial domain (optional but valuable)
- Understanding of security and compliance practices
- Experience with deployment strategies:
- Canary, Blue-Green
Ideal Candidate Profile in summary would be like
- Strong AWS + Kubernetes + Terraform core
- Hands-on production troubleshooting expert
- Good at automation + reducing toil
- Deep understanding of SRE principles
- Comfortable in 24x7 production environments