Job Title: T&T | EAD Senior Consultant | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering

T&T | EAD Senior Consultant | Hybrid cloud | Bengaluru | Engineering | Hybrid Cloud Engineering
• Job requisition ID : 111054
• Location: Bengaluru
• Entity: Deloitte Touche Tohmatsu India LLP
The team
Deloitte’s Technology & Transformation practice can help you uncover and unlock the value buried deep inside vast amounts of data. Our global network provides strategic guidance and implementation services to help companies manage data from disparate sources and convert it into accurate, actionable information that can support fact-driven decision-making and generate an insight-driven advantage. Our practice addresses the continuum of opportunities in business intelligence & visualization, data management, performance management and next-generation analytics and technologies, including big data, cloud, cognitive and machine learning.
Your work profile
Key skills required
-
Strong expertise in NVIDIA GPU infrastructure, architecture and production AI platforms.
-
Strong implementation and automation skills across cloud GPU infrastructure.
-
Experience with GPU cluster design, high-performance networking, storage and cloud-native orchestration.
-
Strong understanding of Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
-
Strong Linux, containers, CUDA ecosystem, NCCL, drivers, firmware and GPU observability fundamentals.
-
Experience with Infrastructure as Code, Terraform, CI/CD and GitOps.
-
Strong troubleshooting, benchmarking, testing and operational handover capabilities.
-
Understanding of AI infrastructure security, resilience, performance, capacity and FinOps.
-
Ability to independently own engineering stories and deliver infrastructure components.
Experience
6–10 Years
Technologies
-
GPU & AI: NVIDIA GPU architecture, DGX/HGX or equivalent certified platforms, CUDA, NCCL, NVIDIA drivers, firmware and GPU observability.
-
Cloud: AWS, Microsoft Azure or Google Cloud, with deep expertise in at least one platform and working awareness of the others.
-
Platform: Kubernetes, OpenShift, Slurm, managed Kubernetes/HPC, GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
-
Infrastructure: Terraform, Infrastructure as Code, image pipelines, CI/CD, GitOps, autoscaling and quota automation.
-
Networking: High-performance network adapters, private connectivity, cluster fabrics and hybrid networking.
-
Storage: High-throughput object, file and block storage, data ingestion, checkpointing, caching and cross-region/data-centre data movement.
-
Observability: GPU availability, utilization, tokens, latency, throughput, reliability and infrastructure cost monitoring.
-
FinOps: Commitments, spot/preemptible usage, idle detection, rightsizing, storage/egress optimization and showback/chargeback.
-
Security: Image security, driver security, model artifacts, secrets, endpoints, data residency and software supply-chain controls.
Key Responsibilities
-
Design GPU landing zones covering accounts/subscriptions/projects, network topology, private connectivity, identity, encryption, policy and observability.
-
Select NVIDIA GPU instances and cluster patterns for distributed training, fine-tuning, batch inference and low-latency serving.
-
Engineer cloud GPU clusters using managed Kubernetes or HPC schedulers, placement/topology controls and high-performance network adapters.
-
Design high-throughput object, file and block storage, data ingestion, checkpointing, caching and cross-region/data-centre movement patterns.
-
Build hybrid connectivity and workload portability between private GPU clusters and public cloud.
-
Implement Terraform, image pipelines, CI/CD/GitOps, autoscaling, quota automation, reservations/capacity blocks and environment promotion.
-
Integrate cloud ML services where appropriate while retaining infrastructure controls for custom NVIDIA-based workloads.
-
Establish observability for GPU availability, utilization, tokens, latency, throughput, reliability and cost.
-
Implement AI infrastructure FinOps covering commitments, spot/preemptible usage, idle detection, rightsizing, storage/egress and showback/chargeback.
-
Engineer security for images, drivers, model artifacts, secrets, endpoints, data residency and software supply chain.
-
Build and automate assigned infrastructure components and independently own engineering stories.
-
Execute validation, benchmarking, troubleshooting, upgrades and operational handover.
-
Create as-built documentation, test evidence, runbooks and reusable infrastructure modules.
-
Contribute to technical design reviews, troubleshooting and continuous improvement of AI infrastructure platforms.
Required Qualifications & Skills
-
6–10 years of relevant experience in infrastructure, DevOps, SRE, HPC or platform engineering.
-
Hands-on experience with GPU or accelerated-computing environments.
-
Strong implementation, automation, testing, troubleshooting and technical-documentation skills.
-
Strong understanding of NVIDIA GPU architecture and systems, including DGX/HGX or equivalent certified platforms.
-
Demonstrated experience in GPU cluster design across compute, high-speed networking, storage, control plane and management plane.
-
Understanding of AI workloads including distributed training, fine-tuning, RAG, batch inference, real-time inference and HPC.
-
Proficiency in Kubernetes/OpenShift and/or Slurm, including GPU scheduling, partitioning, quotas, isolation and multi-tenancy.
-
Strong Linux, container, CUDA, NCCL, driver, firmware and GPU observability fundamentals.
-
Experience with security, resilience, capacity, performance, automation and day-2 operations for production AI infrastructure.
-
Deep expertise in at least one of AWS, Microsoft Azure or Google Cloud, with working awareness of the other major platforms.
-
Experience with cloud GPU capacity, high-performance networking, managed Kubernetes/HPC, Infrastructure as Code
Education
-
Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering or a related technical discipline.
-
Relevant professional certifications in cloud, NVIDIA, Kubernetes, Terraform, networking or FinOps are preferred.
Location and Way of Working
Pan India | Hybrid
