Job Title:  Senior Associate | Site Reliability Engineer (SRE) | Bengaluru | Engineering as a Service/ Operate

Job Title: Senior Site Reliability Engineer

JOB DESCRIPTION

The Infrastructure Reliability Engineering (IRE) team is a new group focused on ensuring the stability, performance, and scalability of our company's infrastructure. We are a team of engineers who build tools to test and verify our infrastructure, automate issue remediation, and create self-healing systems. Our work is critical to providing a seamless and reliable experience for our customers.

Job Summary:

Describe what the person will do in the role - how he/she will impact the organization.

As a Senior Site Reliability Engineer on the Infrastructure Reliability team, you will design, build, and maintain the software and tools that test and verify the performance and reliabilityof our infrastructure. You will work closely with other engineers to identify potential issues, develop automated tests, and create solutions that ensure our systems are robust and resilient.

In this role, you will own the solutions you build, collaborating with cross-functional teams in a fast-paced environment to ensure successful implementation. You will be responsible for continuously improving service resiliency through automation, identifying and resolving performance issues, and conducting capacity planning.

Additionally, you will participate in incident reviews, assist in root cause analysis, and deliver SRE solutions in a globally distributed, multi-cloud hybrid environment (AWS, GCP, and On-prem) to ensure the highest level of uptime and Quality of Service (QoS) for internal customers.

Responsibilities and Duties of the Role:

Design and develop software tools for infrastructure testing and verification.

Create and maintain automated testing frameworks for our infrastructure.

Collaborate with infrastructure and development teams to identify and resolve reliability issues.

Participate in code reviews and contribute to the team's high standards for software quality.

Required Education, Experience/Skills/Training:

•Bachelor’s degree in computer science or a related field, or equivalent experience.

•8+ years of software development experience.

•Proficiency in at least one programming language (e.g., Python, Go, Java).

•Proficiency in Kubernetes administration, modern CI/CD techniques and Infrastructure as Code (IaC).

•Deep understanding of Linux operating systems and TCP/IP fundamentals.

•Experience with software testing methodologies and tools.

•Proficient in monitoring, metrics gathering, APM, container management, and log collection tools.

•Creative problem solver with excellent debugging skills and great documentation abilities.Preferred Qualifications

•Experience with performance and chaos engineering.

•Experience with CI/CD pipelines.•Understands complex system architectures and infrastructures.•Passion for automation, scalability, and building reliable systems from the ground up.

•Familiarity with cloud infrastructure (AWS, Azure, GCP).