Job Title: Senior Associate | Databricks | Bengaluru | Engineering as a Service/ Operate
Job Summary:
We’re looking for a Sr Product Development Engineer I to support and maintain the cloud infrastructure that powers Ad Technology data platforms. This role is focused on the day-to-day stability and performance of systems running Airflow (MWAA), Spark, Flink (on EKS), and Databricks in AWS. You’ll be hands-on in responding to incidents, troubleshooting issues, maintaining infrastructure, and implementing improvements that enhance reliability and efficiency. This is a great opportunity for an experienced engineer who thrives in production environments, enjoys solving complex infrastructure challenges, and is passionate about ensuring the stability of the platforms that drive Disney’s advertising workflows.
Responsibilities:
• Hands-on Platform Support: Provide direct support for production EKS workloads, CI/CD pipelines, and AWS infrastructure.
• Infrastructure Automation: Build and maintain Terraform modules, scripts, and monitoring tools to support platform operations diagnostics, pod health, and deployment issues.
• Troubleshooting & Root Cause Analysis: Respond to issues in EKS, Istio, and CI/CD systems, supporting recovery and documenting resolution.
• Collaboration & Delivery: Partner with application and DevOps teams on shared projects, onboarding, and delivery tasks.
• Documentation & Standards: Contribute to operational runbooks, infrastructure standards, and best practices for platform reliability.
Basic Qualifications:
• 4+ years of experience in infrastructure, platform engineering, or DevOps roles
• Background in data systems support
• Solid operational experience with AWS infrastructure (IAM, VPC, EKS, ALB/NLB).
• Hands-on knowledge of Data infrastructure platforms, MWAA, Databricks, running Spark and Flink on Kubernetes in production environments.
• Experience supporting CI/CD platforms such as Jenkins and Spinnaker, including pipelinetroubleshooting.
• Familiarity with cloud security and networking concepts relevant to EKS, including IAM, TLS, VPCs, DNS, and service-to-service access.
• Strong collaboration and solid communication skills, with the ability to align across distributed engineering teams.
• Working knowledge of infrastructure-as-code (Terraform) and scripting languages (Python, Go, etc.)
• Experience participating in on-call rotations or incident resolution processes across globally distributed teams
• Experience participating in high-severity incidents and driving root cause resolution across globally distributed teams
Preferred Qualifications:
• Experience supporting data engineering cloud-based workloads using Spark and Flink
• Exposure to Azure or GCP alongside AWS
• Experience with Datadog, Prometheus, or similar observability tools
• Exposure to regulated environments (SOX, CCPA) or large-scale production support
• Experience in support tooling, automation, and AIOps concepts
• Familiarity with cost optimization and usage monitoring in AWS and Databricks
Required Education:
• Bachelor’s degree in Computer Science, Engineering, or related technical field (or equivalent experience).