Job Title:  Assistant Manager | Machine Learning Operations (MLOps) | Bengaluru | Engineering as a Service/ Oper

RoleSenior ML Ops Engineer

Experience: 6+ years

Location: Bangalore / Pune

Requirement:

1.     8+ years of experience in Data Engineering, Machine Learning, and MLOps.

 

2.     Strong expertise in Azure Databricks for data processing, feature engineering, ML pipeline development, and model deployment.

 

3.     Hands-on experience with GitHub for source control, branching strategies, code reviews, and collaborative development.

 

4.     Extensive experience in GitHub Actions for building and managing automated CI/CD pipelines

 

  1. Strong understanding of MLOps best practices, including:
  • MLflow experiment tracking and model registry
  • Model lifecycle management
  • Automated model deployment
  • Model monitoring and observability
  • Performance tracking and drift detection

6.     Candidate should be responsible for building enterprise-scale ML platforms and deploying production-grade machine learning solutions on Azure.

 

7.     Knowledge of Azure cloud services such as Azure Machine Learning, Azure Data Factory, Azure Kubernetes Service (AKS), Azure DevOps, and Azure Storage is preferred.




Role Summary

This role is responsible for building and operating enterprise-scale ML platforms and deploying production-grade machine learning solutions on Azure. The candidate will own the full ML lifecycle — from feature engineering and CI/CD through deployment, monitoring, and drift detection — and will work at the intersection of MLOps and DevOps/platform engineering.

Key Requirements

1. Experience Profile

•     8–10 years of experience across Data Engineering, Machine Learning, and MLOps.

•     At least 4–5 years focused specifically on production ML platform / infrastructure work — not just model building.

2. Azure Databricks Expertise

•     Deep, hands-on expertise in PySpark/Spark internals.

•     Delta Lake and medallion architecture (bronze/silver/gold).

•     Feature engineering pipelines and Databricks Workflows/Jobs orchestration.

•     Databricks Feature Store and Unity Catalog governance experience strongly preferred.

3. Source Control & Collaborative Engineering

•     Strong command of GitHub branching strategies (GitFlow / trunk-based development).

•     PR review discipline and codebase structuring for multi-team ML repositories.

4. CI/CD for ML Systems

•     Proven experience building and maintaining CI/CD pipelines using GitHub Actions and/or Azure DevOps YAML pipelines, including:

–    Automated model validation gates (accuracy/bias thresholds) before promotion.

–    Retraining triggers based on schedule or drift signals.

–    Blue-green / canary or shadow deployment strategies for models.

5. Full MLOps Lifecycle Ownership

•     MLflow experiment tracking and model registry management.

•     Model versioning and staging-to-production promotion workflows.

•     Reproducibility practices — environment pinning, data versioning.

6. Production Model Deployment

•     Experience across batch and real-time serving patterns — Azure ML endpoints, Databricks Model Serving, or containerized deployment on AKS.

•     Strong Docker skills required.

•     Ability to author and maintain Kubernetes manifests from scratch (Deployments, Services, HPA, resource limits/requests, ConfigMaps/Secrets) — not just operate or debug existing ones.

7. Model Monitoring & Observability

•     This is treated as a distinct specialization, not a checkbox item.

•     Data drift detection (PSI, KS-test, or equivalent statistical methods).

•     Prediction / output drift monitoring.

•     Feedback-loop design for ground-truth capture and performance decay tracking.

•     Tooling exposure: Evidently AI, Azure Monitor, or custom drift jobs with alerting.

8. Infrastructure-as-Code

•     Working knowledge of Terraform or Bicep for provisioning ML infrastructure.

•     This should not be a portal-clicking role.

9. Azure Cloud Foundation

•     Azure Machine Learning

•     Azure Data Factory

•     Azure Kubernetes Service (AKS)

•     Azure DevOps

•     Azure Storage

•     Entra ID / RBAC for access governance

10. Bonus / Differentiators

•     Governance and lineage experience (model cards, audit trails) — especially valuable in regulated-industry engagements.

•     Secrets management (Key Vault) and PII handling in feature pipelines.

•     Cost optimization experience for ML compute at scale.