Job Title: T&T | EAD | Manager | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering

T&T | EAD | Manager | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering
• Job requisition ID : 111055
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
The team
Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information that can support fact-driven decision-making and generate an insight-driven advantage. Our practice addresses the continuum of opportunities in business intelligence & visualization, data management, performance management and next-generation analytics and technologies, including big data, cloud, cognitive and machine learning.
Your work profile
Key skills required:
- Strong expertise in NVIDIA GPU infrastructure, architecture and production AI platforms.
- Cloud GPU infrastructure design across compute, networking, storage, security and observability.
- Kubernetes/OpenShift and/or Slurm, including GPU scheduling, quotas, partitioning, isolation and multi-tenancy.
- Linux, containers, CUDA ecosystem, NCCL, GPU drivers, firmware and observability.
- Infrastructure as Code, particularly Terraform, along with CI/CD and GitOps.
Experience:
11-15 Years
Technologies:
- GPU & AI: NVIDIA GPU architecture, DGX/HGX or equivalent certified platforms, CUDA, NCCL, GPU drivers and firmware.
- Cloud: Strong expertise in at least one of AWS, Microsoft Azure or Google Cloud, with working awareness of the other major cloud platforms.
- Platform: Kubernetes, OpenShift, Slurm, managed Kubernetes/HPC, GPU scheduling and multi-tenancy.
- Infrastructure: Terraform, Infrastructure as Code, image pipelines, CI/CD and GitOps.
- Storage & Networking: High-throughput object, file and block storage, high-performance network adapters, private connectivity and hybrid networking.
Key Responsibilities:
- Design GPU landing zones covering accounts/subscriptions/projects, network topology, private connectivity, identity, encryption, policy and observability.
- Select NVIDIA GPU instances and cluster patterns for distributed training, fine-tuning, batch inference and low-latency serving.
- Engineer cloud GPU clusters using managed Kubernetes or HPC schedulers, placement/topology controls and high-performance network adapters.
- Design high-throughput object, file and block storage, data ingestion, checkpointing, caching and cross-region/data-centre data movement patterns.
Required Qualifications & Skills
- 11–15 years of relevant experience in infrastructure, cloud, SRE, HPC or platform engineering.
- Strong hands-on experience with NVIDIA GPU infrastructure and production AI/GPU platforms.
- Demonstrated experience in GPU cluster architecture, cloud infrastructure and hybrid AI environments.
- Strong understanding of distributed AI workloads, GPU scheduling, high-performance networking and storage.
- Proficiency in Kubernetes/OpenShift and/or Slurm.
- Strong Linux, container, CUDA and NCCL fundamentals.
- Proven experience with infrastructure automation using Terraform and modern CI/CD/GitOps practices.
- Deep expertise in at least one hyperscale cloud platform: AWS, Microsoft Azure or Google Cloud.
Education
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering or a related technical discipline.
- Relevant professional certifications in cloud, NVIDIA, Kubernetes, Terraform, networking or FinOps are preferred.
Location and Way of Working:
Pan India and Hybrid
