Job Title:  Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering

Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering
Job requisition ID : 111027 
Location: Bengaluru
Entity: Deloitte Touche Tohmatsu India LLP 

The team

The team must do much more than keep the wheels turning; it is the engine that drives functional excellence and the enabler of innovation and long-term growth.

Key responsibilities

  • Own end-to-end architecture across NVIDIA GPU compute, CPU/DPU, management nodes, control plane, high-speed fabrics, storage, backup and security.
  • Design and deploy DGX/HGX or NVIDIA-certified GPU clusters and select fit-for-purpose GPU platforms for training, fine-tuning, visualization, batch inference and real-time inference.
  • Architect InfiniBand, Spectrum-X Ethernet and RoCE networks, including topology, rail optimization, congestion control, OOB management, telemetry and performance engineering.
  • Define high-performance file and object storage architectures covering checkpointing, data staging, metadata, throughput, IOPS, resiliency and data lifecycle requirements.
  • Engineer AI data-centre requirements covering rack layout, GPU density, power, cooling, cabling, physical resilience, capacity expansion and DCIM/monitoring integration.
  • Implement Kubernetes/OpenShift and/or Slurm-based platforms, including GPU scheduling, NVIDIA GPU/Network Operators, MIG, quotas, tenancy, workload isolation and policy enforcement.
  • Integrate and operationalize NVIDIA AI Enterprise, CUDA, NCCL, NIM, NeMo, Triton, DCGM and approved AI model/runtime services.
  • Establish automated provisioning, firmware and driver lifecycle management, image management, validation, burn-in, acceptance testing and performance benchmarking.
  • Operate the platform using SRE principles, with comprehensive observability, incident and problem management, capacity planning, patching, disaster recovery and security controls.
  • Optimize GPU utilization, job throughput, tokens per second, workload performance, energy efficiency and cost per workload.
  • Own the market proposition, account strategy, pipeline development, strategic alliances and executive relationships for AI infrastructure services.
  • Shape multi-year AI infrastructure transformation programmes and lead commercial negotiations, governance, delivery quality, risk management and business economics.
  • Build, mentor and retain a nationally recognized team of AI data-centre, GPU, network, storage, cloud and platform engineering specialists.

Required skills

  • Typically 18+ years of experience across infrastructure, cloud, data-centre, HPC or platform engineering, including substantial leadership of production AI/GPU environments.
  • Proven experience in business development, executive advisory, strategic alliances and governance of large-scale technology programmes.
  • Strong expertise in NVIDIA GPU architectures and systems, including DGX/HGX or equivalent certified platforms.
  • Proven experience designing GPU clusters across compute, high-speed networking, storage, control plane and management plane.
  • Strong understanding of AI workload characteristics, including distributed training, fine-tuning, RAG, batch inference, real-time inference and HPC.
  • Hands-on architecture and leadership experience with Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, workload isolation and multi-tenancy.
  • Strong Linux, container, CUDA ecosystem, NCCL, GPU driver, firmware and GPU observability fundamentals.
  • Proven experience with production AI infrastructure covering security, resilience, capacity management, performance optimization, automation and day-2 operations.
  • Strong understanding of high-density data-centre environments, including power, cooling, physical networking, rack-scale deployment and facility dependencies.
  • Experience designing and operating high-performance storage and InfiniBand/Ethernet fabrics for distributed AI and HPC workloads.
  • Strong commercial, stakeholder-management and executive-communication capabilities.
  • Experience across AI factories, GPU clouds, HPC centres, enterprise AI data centres, GCCs, OEMs or colocation facilities.
  • Exposure to NVIDIA DGX/HGX, NVIDIA AI Enterprise, Spectrum-X, InfiniBand, BasePOD, SuperPOD or NVIDIA validated reference architectures.
  • Experience with NVIDIA-certified infrastructure and production-scale AI platform deployments.
  • Relevant certifications such as CKA/CKS, Red Hat, NVIDIA, Linux, networking, storage or data-centre certifications.
  • Experience building AI infrastructure practices, centres of excellence or nationally scaled specialist engineering teams.
  • Production-ready, secure, resilient and scalable NVIDIA GPU infrastructure.
  • High GPU utilization, predictable application performance and reduced time to onboard AI workloads.
  • Reliable capacity expansion with automated provisioning, validation and lifecycle management.
  • Measurable improvements in workload throughput, energy efficiency and cost per workload.
  • Strong platform reliability, observability, security and operational maturity.
  • High client satisfaction, consistent delivery quality and reusable engineering accelerators.
  • Strong market positioning, strategic partnerships and sustainable AI infrastructure pipeline.
  • A high-performing, nationally recognized AI infrastructure engineering organization.