Job Title: T&T | EAD | Senior Consultant/Manager | AI Infra GPU | PAN India

T&T | EAD | Senior Consultant/Manager | AI Infra GPU | PAN India
• Job requisition ID : 111054
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
The Team
Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information that can support fact-driven decision-making and generate an insight-driven advantage. Our practice addresses the continuum of opportunities in business intelligence & visualization, data management, performance management and next-generation analytics and technologies, including big data, cloud, cognitive and machine learning.
Your work profile :
- Strong expertise in NVIDIA GPU infrastructure, architecture and production AI platforms.
- Strong implementation and automation skills across cloud GPU infrastructure.
- Experience with GPU cluster design, high-performance networking, storage and cloud-native orchestration.
- Strong understanding of Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
- Strong Linux, containers, CUDA ecosystem, NCCL, drivers, firmware and GPU observability fundamentals.
- Experience with Infrastructure as Code, Terraform, CI/CD and GitOps.
- Strong troubleshooting, benchmarking, testing and operational handover capabilities.
Experience:
- 2–12 Years
Key Responsibilities:
- Design GPU landing zones covering accounts/subscriptions/projects, network topology, private connectivity, identity, encryption, policy and observability.
- Select NVIDIA GPU instances and cluster patterns for distributed training, fine-tuning, batch inference and low-latency serving.
- Engineer cloud GPU clusters using managed Kubernetes or HPC schedulers, placement/topology controls and high-performance network adapters.
- Design high-throughput object, file and block storage, data ingestion, checkpointing, caching and cross-region/data-centre movement patterns.
- Build hybrid connectivity and workload portability between private GPU clusters and public cloud.
- Implement Terraform, image pipelines, CI/CD/GitOps, autoscaling, quota automation, reservations/capacity blocks and environment promotion.
- Integrate cloud ML services where appropriate while retaining infrastructure controls for custom NVIDIA-based workloads.
- Establish observability for GPU availability, utilization, tokens, latency, throughput, reliability and cost.
- Implement AI infrastructure FinOps covering commitments, spot/preemptible usage, idle
Required Qualifications & Skills:
- 2–12 years of relevant experience in infrastructure, SRE, HPC or platform engineering.
- Hands-on experience with GPU or accelerated-computing environments.
- Strong implementation, automation, testing, troubleshooting and technical-documentation skills.
- Strong understanding of NVIDIA GPU architecture and systems, including DGX/HGX or equivalent certified platforms.
- Demonstrated experience in GPU cluster design across compute, high-speed networking, storage, control plane and management plane.
- Understanding of AI workloads including distributed training, fine-tuning, RAG, batch inference, real-time inference and HPC.
- Proficiency in Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
- Strong Linux, container, CUDA, NCCL, driver, firmware and GPU observability fundamentals.
- Experience with security, resilience, capacity, performance, automation and day-2 operations for production AI infrastructure.
- Deep expertise in at least one of AWS, Microsoft Azure or Google Cloud, with working awareness of the other major platforms.
- Experience with cloud GPU capacity, high-performance networking, managed Kubernetes/HPC, Infrastructure as Code
Education:
- Any Degree, Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering or a related technical discipline is preferred.
Location and Way of Working:
- Pan India | Hybrid
