Job Title:  Consultant | Site Reliability Engineering | Bengaluru | Engineering | Hybrid Cloud Engineering

Consultant | Site Reliability Engineering | Bengaluru  

 

Role Overview

We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on AWS. The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts.

 

This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration and TCP/IP fundamentals, and strong troubleshooting capabilities for distributed systems. The candidate needs to participate in the overall lifecycle management of mission critical banking services with a 24x7 operations mode in an rotational on-call basis. The job requires the candidate to have strong troubleshooting skills in a distributed environment spread across multiple cloud environments. The bare minimum ask would be to maintain high level of agility, learnability and adaptability in different scenarios. An engineer with a zeal to learn fast and having a bias for action would be the best fit for the role.

 

Key Responsibilities

Reliability & Operations

  • Own end-to-end production systems reliability, availability, scalability, cost and performance.
  • Troubleshooting EKS from SRE perspective
  • Route53, Connectivity, Networking, RDS, S3
  • Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
  • Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
  • Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.

Cloud & Infrastructure

  • Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
  • Work extensively on:
  • GKE (Kubernetes Engine)
  • Compute, networking, IAM, Load Balancers, TLS Certs
  • BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
  • Implement and manage infrastructure using Terraform (Infrastructure as Code).

Kubernetes & Containers

  • Deploy and manage containerized workloads using Kubernetes (GKE).
  • Troubleshoot issues related to:
  • Pods, nodes, networking, storage, services on an ongoing basis
  • Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).

Automation & CI/CD

  • Build and maintain CI/CD pipelines using:
  • Jenkins (pipeline-based, Groovy / Shell / Python scripting)
  • Strong experience in using GitHub as a PowerUser
  • Develop automation using Python and Shell scripting.
  • Reduce operational toil through automation initiatives.

Observability & Monitoring

  • Implement and manage monitoring systems using:
  • Dynatrace, Grafana, logs and metrics explorer
  • Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively.
  • Define alerting strategies based on system behaviour and SLOs and create runbooks.
  • Work alongside operations teams to identify, fix the production incidents and own the problem resolution.
  • Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents.

System & Application Troubleshooting

  • Perform deep troubleshooting for:
  • Distributed systems
  • Microservices-based architectures on containerised workloads
  • Java and Golang applications
  • Strong debugging of:
  • Application issues
  • Infrastructure issues
  • Network-related problems

Plan and execute continuous improvement

  • Identify and eliminate repetitive manual tasks.
  • Drive reliability engineering practices and culture. (DRY – Don’t Repeat Yourself)
  • Collaborate with development teams to improve system design and resilience.

 

Job Requirements

Experience

  • 3–5 years of relevant and progressive experience in SRE / DevOps / Cloud Engineering
  • Hands-on experience managing production-grade systems (24x7 environments)

 

Technical Skills

  • Troubleshooting EKS from SRE perspective
  • Route53, Connectivity, Networking, RDS, S3

Infrastructure as Code

  • Strong hands-on experience with Terraform
  • Ability to write and debug Terraform code from scratch

Containers & Orchestration

  • Deep expertise in:
  • Kubernetes (GKE)
  • Docker
  • Strong troubleshooting experience in Kubernetes environments

CI/CD & Automation

  • Hands-on experience with:
  • Jenkins (pipeline-based CI/CD)
  • GitHub
  • Strong scripting skills:
  • Python (preferred)
  • Shell scripting
  • Experience with automation frameworks and tooling

Observability

  • Experience with:
  • Dynatrace / Grafana
  • Log, metrics, and trace-based monitoring

Programming & Debugging

  • Working knowledge of:
  • Java and/or Golang applications
  • Strong debugging skills across application and infrastructure layers

Linux & Networking

  • Strong Linux fundamentals
  • Deep understanding of TCP/IP networking
  • Ability to debug network issues in distributed systems

 

Reliability Engineering Skills

  • Solid understanding of:
  • SLI, SLO, SLA, Error Budgets
  • Demonstrable and Proven Experience improving:
  • MTTR, MTTA
  • Experience handling incident management lifecycle

 

Soft Skills

  • Strong analytical and troubleshooting mindset
  • Excellent communication and stakeholder management
  • Ability to work in high-pressure production environments
  • Ownership-driven and proactive approach

 

Preferred Qualifications

  • Experience in high-scale distributed systems
  • Exposure to banking/financial domain (optional but valuable)
  • Understanding of security and compliance practices
  • Experience with deployment strategies:
  • Canary, Blue-Green

 

Ideal Candidate Profile in summary would be like :

  • Strong AWS + Kubernetes + Terraform core
  • Hands-on production troubleshooting expert
  • Good at automation + reducing toil
  • Deep understanding of SRE principles
  • Comfortable in 24x7 production environments