Job Title:  Manager | DevOps | Hyderabad | Engineering as a Service/ Operate

About the Role

We are looking for a highly skilled and motivated Senior DevOps Engineer to join our platform engineering team. In this role, you will be a key contributor to the design, delivery, and evolution of our cloud infrastructure and Kubernetes platforms. You are someone who thrives in ambiguity, brings a growth mindset to every challenge, and communicates clearly across engineering and business stakeholders.

Key Responsibilities

  • Design, build, and maintain scalable, secure, and highly available infrastructure across multi-account, multi-region, and multi-AZ AWS environments
  • Lead Infrastructure as Code practices using Terraform — authoring reusable modules, managing workspaces, and enforcing standards across teams
  • Administer and evolve Amazon EKS clusters including upgrades, secondary networking configurations, and cluster lifecycle management
  • Write and maintain Helm charts including umbrella charts and integration of OSS/community charts; troubleshoot complex Helm deployment issues
  • Develop and maintain custom Kubernetes controllers and operators using standard tooling
  • Define and champion multi-cloud architecture patterns and best practices
  • Analyze open-ended infrastructure and platform problems, propose well-reasoned solutions, and drive them to completion with minimal supervision
  • Collaborate closely with development, security, and architecture teams — communicating complex technical concepts clearly to diverse audiences
  • Mentor junior engineers and contribute to a culture of continuous learning and improvement

Required Skills & Experience

Core Competencies

  • Strong communication skills — written and verbal — with the ability to present technical proposals to both engineers and non-technical stakeholders
  • Proven ability to analyze ambiguous, open-ended problems and independently propose and deliver pragmatic solutions
  • Strong problem-solving mindset with a structured, methodical approach to debugging and root-cause analysis
  • Demonstrated growth mindset — actively learns from incidents, seeks feedback, and keeps up with evolving industry practices

Terraform

  • hands-on experience writing production-grade Terraform configurations and reusable modules
  • experience with Terraform workspaces for environment and account segregation
  • experience using Terraform providers to provision and manage AWS resources across multi-account, multi-region, and multi-AZ topologies
  • experience provisioning and managing EKS clusters, VPC networking, and related AWS services via HashiCorp Terraform modules
  • experience configuring EKS secondary networking (e.g., VPC CNI custom networking) via Terraform
  • Good understanding of multi-cloud infrastructure architecture and portability considerations

Kubernetes & EKS

  • hands-on experience in Amazon EKS administration — cluster provisioning, IAM integration, node group management, and day-2 operations
  • experience as a Kubernetes practitioner — workload management, RBAC, networking, storage, autoscaling, and observability
  • understanding of Custom Resource Definitions (CRDs) and consumption of OSS Helm charts and Kubernetes operators
  • experience managing and executing EKS cluster version upgrades with minimal disruption
  • Hands-on experience writing custom controllers and operators using frameworks such as controller-runtime or Kubebuilder
  • hands-on experience with AWS Load Balancer Controller — ingress and service configuration on EKS
  • hands-on experience with Kubernetes Gateway API — HTTPRoute, GatewayClass, and migration from Ingress

Helm

  • experience authoring Helm charts for complex, production workloads
  • experience building and managing umbrella charts for multi-component application deployments
  • understanding of CRDs and OSS Helm charts — including evaluation, customisation, and maintenance
  • experience troubleshooting Helm release failures, hook issues, and chart rendering problems

Istio

  • hands-on experience deploying and operating Istio in production Kubernetes environments
  • experience with both Istio Sidecar (proxy) and Ambient mesh deployment modes
  • understanding of Istio core concepts and traffic management — VirtualService, DestinationRule, Gateway, AuthorizationPolicy, mTLS, and observability integration

AWS

  • hands-on experience designing and operating multi-region, multi-AZ AWS architectures
  • Deep working knowledge across core AWS services, including:
  • EKS — cluster administration, networking, and integrations
  • VPC — subnets, route tables, security groups, VPC peering, Transit Gateway
  • EC2 — instance types, launch templates, Auto Scaling Groups
  • S3 — bucket policies, lifecycle rules, replication
  • AWS API Gateway
  • Networking — NAT, IGW, PrivateLink, VPN, Direct Connect concepts
  • ACM (AWS Certificate Manager) — certificate provisioning, renewal, and integration with AWS services
  • PCA (AWS Private Certificate Authority) — private CA management for internal TLS
  • Route 53 — DNS management, health checks, routing policies
  • Strong understanding of TLS — certificate chains, mTLS, termination strategies, and PKI fundamentals

Nice to Have

  • Experience with GitOps tooling (e.g., ArgoCD, GitHub Actions, Cloudbees Jenkins)
  • Familiarity with observability stacks (e.g., Prometheus, Grafana, OpenTelemetry)
  • Experience with CI/CD pipelines (e.g., GitHub Actions, Jenkins, Atlantis)

Exposure to service mesh technologies (e.g., Istio, Linkerd)