Job Title: Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering

Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering
• Job requisition ID : 111027
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
The team
The team must do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth.
Key responsibilities
- Own end-to-end architecture across NVIDIA GPU compute, CPU/DPU, management nodes, control plane, high-speed fabrics, storage, backup and security.
- Design and deploy DGX/HGX or NVIDIA-certified GPU clusters and select fit-for-purpose GPU platforms for training, fine-tuning, visualization, batch inference and real-time inference.
- Architect InfiniBand, Spectrum-X Ethernet and RoCE networks, including topology, rail optimization, congestion control, OOB management, telemetry and performance engineering.
- Define high-performance file and object storage architectures covering checkpointing, data staging, metadata, throughput, IOPS, resiliency and data lifecycle requirements.
- Engineer AI data-centre requirements covering rack layout, GPU density, power, cooling, cabling, physical resilience, capacity expansion and DCIM/monitoring integration.
- Implement Kubernetes/OpenShift and/or Slurm-based platforms, including GPU scheduling, NVIDIA GPU/Network Operators, MIG, quotas, tenancy, workload isolation and policy enforcement.
- Integrate and operationalize NVIDIA AI Enterprise, CUDA, NCCL, NIM, NeMo, Triton, DCGM and approved AI model/runtime services.
- Establish automated provisioning, firmware and driver lifecycle management, image management, validation, burn-in, acceptance testing and performance benchmarking.
- Operate the platform using SRE principles, with comprehensive observability, incident and problem management, capacity planning, patching, disaster recovery and security controls.
- Optimize GPU utilization, job throughput, tokens per second, workload performance, energy efficiency and cost per workload.
- Own the market proposition, account strategy, pipeline development, strategic alliances and executive relationships for AI infrastructure services.
- Shape multi-year AI infrastructure transformation programmes and lead commercial negotiations, governance, delivery quality, risk management and business economics.
- Build, mentor and retain a nationally recognized team of AI data-centre, GPU, network, storage, cloud and platform engineering specialists.
Required skills
- Typically 18+ years of experience across infrastructure, cloud, data-centre, HPC or platform engineering, including substantial leadership of production AI/GPU environments.
- Proven experience in business development, executive advisory, strategic alliances and governance of large-scale technology programmes.
- Strong expertise in NVIDIA GPU architectures and systems, including DGX/HGX or equivalent certified platforms.
- Proven experience designing GPU clusters across compute, high-speed networking, storage, control plane and management plane.
- Strong understanding of AI workload characteristics, including distributed training, fine-tuning, RAG, batch inference, real-time inference and HPC.
- Hands-on architecture and leadership experience with Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, workload isolation and multi-tenancy.
- Strong Linux, container, CUDA ecosystem, NCCL, GPU driver, firmware and GPU observability fundamentals.
- Proven experience with production AI infrastructure covering security, resilience, capacity management, performance optimization, automation and day-2 operations.
- Strong understanding of high-density data-centre environments, including power, cooling, physical networking, rack-scale deployment and facility dependencies.
- Experience designing and operating high-performance storage and InfiniBand/Ethernet fabrics for distributed AI and HPC workloads.
- Strong commercial, stakeholder-management and executive-communication capabilities.
- Experience across AI factories, GPU clouds, HPC centres, enterprise AI data centres, GCCs, OEMs or colocation facilities.
- Exposure to NVIDIA DGX/HGX, NVIDIA AI Enterprise, Spectrum-X, InfiniBand, BasePOD, SuperPOD or NVIDIA validated reference architectures.
- Experience with NVIDIA-certified infrastructure and production-scale AI platform deployments.
- Relevant certifications such as CKA/CKS, Red Hat, NVIDIA, Linux, networking, storage or data-centre certifications.
- Experience building AI infrastructure practices, centres of excellence or nationally scaled specialist engineering teams.
- Production-ready, secure, resilient and scalable NVIDIA GPU infrastructure.
- High GPU utilization, predictable application performance and reduced time to onboard AI workloads.
- Reliable capacity expansion with automated provisioning, validation and lifecycle management.
- Measurable improvements in workload throughput, energy efficiency and cost per workload.
- Strong platform reliability, observability, security and operational maturity.
- High client satisfaction, consistent delivery quality and reusable engineering accelerators.
- Strong market positioning, strategic partnerships and sustainable AI infrastructure pipeline.
- A high-performing, nationally recognized AI infrastructure engineering organization.
