Job Title: T&T | EAD | Associate Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering

T&T | EAD | Associate Director | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering
• Job requisition ID : 111058
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
The team
Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information that can support fact-driven decision-making and generate an insight-driven advantage. Our practice addresses the continuum of opportunities in business intelligence & visualization, data management, performance management and next-generation analytics and technologies, including big data, cloud, cognitive and machine learning.
Your work profile
Key skills required
-
Strong expertise in NVIDIA GPU infrastructure, architecture and AI data-centre strategy.
-
Ability to translate AI workload demand into GPU, compute, network, storage, power and cooling requirements.
-
Strong expertise in AI infrastructure architecture across private cloud, public cloud and hybrid environments.
-
Strong understanding of GPU cluster design, high-performance networking, storage and AI platform technologies.
-
Expertise in AI infrastructure economics, GPU tokenomics, TCO/ROI modelling and capacity planning.
-
Strong understanding of Kubernetes/OpenShift and/or Slurm, including GPU scheduling, quotas, partitioning, isolation and multi-tenancy.
-
Strong client leadership, architecture authority, commercial estimation and multidisciplinary delivery skills.
Experience
15+ Years
Technologies
-
GPU & AI: NVIDIA GPU architecture, DGX/HGX or equivalent certified platforms, CUDA, NCCL, NVIDIA software and GPU observability.
-
Cloud: Private cloud, public cloud, hybrid cloud, cloud GPU services and specialist GPU cloud providers.
-
Platform: Kubernetes, OpenShift, Slurm, GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
-
Networking: InfiniBand, NVIDIA Spectrum-X, RoCE and high-performance networking architectures.
-
Storage: High-performance object, file and block storage, data pipelines and AI workload storage architectures.
-
Data Centre: High-density rack architectures, power chain, cooling, floor loading, cabling, resilience and operational technology integration.
-
Operating Model: GPU-as-a-Service, service catalogue, tenancy, quota management, SLOs, governance and showback/chargeback.
-
FinOps: GPU acquisition/consumption economics, commitments, utilization, rightsizing, energy, infrastructure cost, refresh cycles and TCO/ROI.
-
Security & Resilience: Infrastructure security, supply-chain controls, sovereignty, business continuity, operational controls and sustainability.
Key Responsibilities
-
Assess AI use-case demand and translate model size, concurrency, tokens, latency, data volume and growth assumptions into compute, GPU memory, fabric, storage, power and cooling requirements.
-
Develop workload placement strategies across enterprise data centres, colocation facilities, public cloud GPU services and specialist GPU clouds.
-
Define AI infrastructure strategy, target architecture, sourcing model, AI factory/data-centre roadmap and phased capacity plan.
-
Build hybrid cloud business cases covering GPU acquisition/consumption, facilities, network, storage, software, operations, energy and refresh cycles.
-
Lead AI data-centre readiness assessments covering rack density, power chain, cooling, floor loading, cabling, resilience and operational technology integration.
-
Evaluate DGX/HGX, certified systems, cloud GPU instances, InfiniBand, Spectrum-X, RoCE, high-performance storage, Kubernetes/OpenShift, Slurm and NVIDIA software options.
-
Define GPU-as-a-Service operating models covering service catalogue, tenancy, quota, showback/chargeback, SLOs and governance.
-
Create migration and onboarding waves for AI applications, training pipelines and inference services.
-
Define AI infrastructure security, supply-chain, sovereignty, continuity, sustainability, FinOps and operational controls.
-
Own solution architecture and delivery governance from assessment through production operations.
-
Lead OEM, NVIDIA, hyperscaler, colocation, networking and storage partner discussions and technical selections.
-
Lead proposals, estimates, architecture boards and development of reusable reference architectures.
Required Qualifications & Skills
-
15+ years of relevant experience in infrastructure, platform engineering, HPC, cloud or AI infrastructure.
-
Strong hands-on and strategic experience with NVIDIA GPU infrastructure and production AI/GPU platforms.
-
Demonstrated experience in GPU cluster architecture, high-performance networking, storage and AI infrastructure.
-
Strong understanding of distributed AI workloads, GPU scheduling, Kubernetes/OpenShift and/or Slurm.
-
Strong understanding of AI data-centre infrastructure, including power, cooling, rack density, networking and resilience.
-
Proven ability to develop defensible capacity models, TCO/ROI cases and executive decision papers.
-
Strong understanding of security, resilience, capacity, performance, automation and day-2 operations for production AI infrastructure.
-
Strong client-facing leadership, stakeholder management, architecture governance and commercial estimation skills.
-
Ability to lead multidisciplinary teams and complex technology transformation programmes.
Education
-
Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering or a related technical discipline.
-
Relevant professional certifications in cloud, NVIDIA, Kubernetes, Terraform, networking or FinOps are preferred.
Location and Way of Working
Pan India | Hybrid
