Job Title:  Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering

Role mandate

The Director will architect, engineer and operate production NVIDIA GPU infrastructure on public cloud and connect it with private AI estates. The role covers GPU landing zones, cluster fabrics, cloud-native orchestration, performance, security, reliability and AI infrastructure FinOps.


Key responsibilities

·      Design GPU landing zones, accounts/subscriptions/projects, network topology, private connectivity, identity, encryption, policy and observability.

·      Select NVIDIA GPU instances and cluster patterns for distributed training, fine-tuning, batch inference and low-latency serving.

·      Engineer cloud GPU clusters using managed Kubernetes or HPC schedulers, placement/topology controls and high-performance network adapters.

·      Design high-throughput object, file and block storage, data ingestion, checkpointing, cache and cross-region/data-centre movement patterns.

·      Build hybrid connectivity and workload portability between private GPU clusters and public cloud.

·      Implement Terraform, image pipelines, CI/CD/GitOps, autoscaling, quota automation, reservations/capacity blocks and environment promotion.

·      Integrate cloud ML services where appropriate while retaining infrastructure controls for custom NVIDIA-based workloads.

·      Establish GPU availability, utilization, token, latency, throughput, reliability and cost observability.

·      Implement AI infrastructure FinOps covering commitments, spot/preemptible usage, idle detection, rightsizing, storage/egress and showback/chargeback.

·      Engineer security for images, drivers, model artifacts, secrets, endpoints, data residency and software supply chain.

·      Own market proposition, account strategy, pipeline, strategic alliances and executive relationships.

·      Shape multi-year AI infrastructure transformations and lead commercial negotiations, quality, risk and economics.

·      Build and retain a nationally recognized team of AI data-centre, GPU, network, storage, cloud and platform specialists.

Required experience and skills

·      Typically 18+ years in infrastructure, cloud, data-centre, HPC or platform engineering, with substantial leadership of production AI/GPU estates.

·      Proven business development, executive advisory, alliance and large-programme governance experience.

·      NVIDIA GPU architecture and systems including DGX/HGX or equivalent certified platforms.

·      GPU cluster design across compute, high-speed network, storage, control plane and management plane.

·      AI workload characteristics across distributed training, fine-tuning, RAG, batch inference, real-time inference and HPC.

·      Kubernetes/OpenShift and/or Slurm; GPU scheduling, partitioning, quotas, isolation and multi-tenancy.

·      Linux, containers, CUDA ecosystem, NCCL, drivers, firmware and GPU observability fundamentals.

·      Security, resilience, capacity, performance, automation and day-2 operations for production AI infrastructure.

·      Deep expertise in at least one of AWS, Microsoft Azure or Google Cloud and working awareness of the other platforms.

·      Experience with cloud GPU capacity, high-performance networking, managed Kubernetes/HPC, IaC and cloud cost optimization.

Preferred profile

·      Experience in hyperscalers, GPU cloud/neocloud providers, GCC cloud platform teams or large enterprise cloud AI estates.

·      Professional cloud, NVIDIA, Kubernetes, Terraform, network or FinOps certification.

·      Experience with hybrid AI, sovereign cloud, reserved GPU capacity and production-scale inference.

Success outcomes

·      Production-ready, secure and scalable GPU infrastructure.

·      High GPU utilization, predictable application performance and reduced time to onboard AI workloads.

·      Reliable capacity expansion, platform automation and measurable cost/energy efficiency.

·      Strong client satisfaction, delivery quality and reusable engineering accelerators.