Job Title: Manager | DevOps | Hyderabad | Engineering as a Service/ Operate
About the Role
We are looking for a highly skilled and motivated Senior DevOps Engineer to join our platform engineering team. In this role, you will be a key contributor to the design, delivery, and evolution of our cloud infrastructure and Kubernetes platforms. You are someone who thrives in ambiguity, brings a growth mindset to every challenge, and communicates clearly across engineering and business stakeholders.
Key Responsibilities
- Design, build, and maintain scalable, secure, and highly available infrastructure across multi-account, multi-region, and multi-AZ AWS environments
- Lead Infrastructure as Code practices using Terraform — authoring reusable modules, managing workspaces, and enforcing standards across teams
- Administer and evolve Amazon EKS clusters including upgrades, secondary networking configurations, and cluster lifecycle management
- Write and maintain Helm charts including umbrella charts and integration of OSS/community charts; troubleshoot complex Helm deployment issues
- Develop and maintain custom Kubernetes controllers and operators using standard tooling
- Define and champion multi-cloud architecture patterns and best practices
- Analyze open-ended infrastructure and platform problems, propose well-reasoned solutions, and drive them to completion with minimal supervision
- Collaborate closely with development, security, and architecture teams — communicating complex technical concepts clearly to diverse audiences
- Mentor junior engineers and contribute to a culture of continuous learning and improvement
Required Skills & Experience
Core Competencies
- Strong communication skills — written and verbal — with the ability to present technical proposals to both engineers and non-technical stakeholders
- Proven ability to analyze ambiguous, open-ended problems and independently propose and deliver pragmatic solutions
- Strong problem-solving mindset with a structured, methodical approach to debugging and root-cause analysis
- Demonstrated growth mindset — actively learns from incidents, seeks feedback, and keeps up with evolving industry practices
Terraform
- hands-on experience writing production-grade Terraform configurations and reusable modules
- experience with Terraform workspaces for environment and account segregation
- experience using Terraform providers to provision and manage AWS resources across multi-account, multi-region, and multi-AZ topologies
- experience provisioning and managing EKS clusters, VPC networking, and related AWS services via HashiCorp Terraform modules
- experience configuring EKS secondary networking (e.g., VPC CNI custom networking) via Terraform
- Good understanding of multi-cloud infrastructure architecture and portability considerations
Kubernetes & EKS
- hands-on experience in Amazon EKS administration — cluster provisioning, IAM integration, node group management, and day-2 operations
- experience as a Kubernetes practitioner — workload management, RBAC, networking, storage, autoscaling, and observability
- understanding of Custom Resource Definitions (CRDs) and consumption of OSS Helm charts and Kubernetes operators
- experience managing and executing EKS cluster version upgrades with minimal disruption
- Hands-on experience writing custom controllers and operators using frameworks such as controller-runtime or Kubebuilder
- hands-on experience with AWS Load Balancer Controller — ingress and service configuration on EKS
- hands-on experience with Kubernetes Gateway API — HTTPRoute, GatewayClass, and migration from Ingress
Helm
- experience authoring Helm charts for complex, production workloads
- experience building and managing umbrella charts for multi-component application deployments
- understanding of CRDs and OSS Helm charts — including evaluation, customisation, and maintenance
- experience troubleshooting Helm release failures, hook issues, and chart rendering problems
Istio
- hands-on experience deploying and operating Istio in production Kubernetes environments
- experience with both Istio Sidecar (proxy) and Ambient mesh deployment modes
- understanding of Istio core concepts and traffic management — VirtualService, DestinationRule, Gateway, AuthorizationPolicy, mTLS, and observability integration
AWS
- hands-on experience designing and operating multi-region, multi-AZ AWS architectures
- Deep working knowledge across core AWS services, including:
- EKS — cluster administration, networking, and integrations
- VPC — subnets, route tables, security groups, VPC peering, Transit Gateway
- EC2 — instance types, launch templates, Auto Scaling Groups
- S3 — bucket policies, lifecycle rules, replication
- AWS API Gateway
- Networking — NAT, IGW, PrivateLink, VPN, Direct Connect concepts
- ACM (AWS Certificate Manager) — certificate provisioning, renewal, and integration with AWS services
- PCA (AWS Private Certificate Authority) — private CA management for internal TLS
- Route 53 — DNS management, health checks, routing policies
- Strong understanding of TLS — certificate chains, mTLS, termination strategies, and PKI fundamentals
Nice to Have
- Experience with GitOps tooling (e.g., ArgoCD, GitHub Actions, Cloudbees Jenkins)
- Familiarity with observability stacks (e.g., Prometheus, Grafana, OpenTelemetry)
- Experience with CI/CD pipelines (e.g., GitHub Actions, Jenkins, Atlantis)
Exposure to service mesh technologies (e.g., Istio, Linkerd)