Job Title:  Consultant | Data Analytics | Delhi | Operations, Industry & Domain Solutions | DigiGov

Consultant | Data Analytics | Delhi | Operations, Industry & Domain Solutions | DigiGov
Job requisition ID : 111595 
Location: Delhi
Entity: Deloitte Touche Tohmatsu India LLP 

Your work profile

We are looking for a Data Engineer with 3–5 years of experience to design, develop, and maintain scalable data pipelines and ETL workflows. The ideal candidate should have strong hands-on experience with PySpark, Python, SQL, and data engineering frameworks, along with a good understanding of distributed data processing.

 

Key Responsibilities

  • Design, develop, and maintain ETL/ELT pipelines for processing large volumes of data.
  • Build scalable and efficient PySpark data pipelines for batch and/or streaming workloads.
  • Develop data transformation, cleansing, validation, and aggregation processes.
  • Work with data lakes, data warehouses, and cloud-based data platforms.
  • Optimize Spark jobs and pipelines for performance, scalability, and reliability.
  • Write complex SQL queries for data extraction, transformation, and analysis.
  • Implement data quality checks, error handling, logging, and monitoring.
  • Collaborate with data analysts, data scientists, and application teams to understand data requirements.
  • Troubleshoot pipeline failures and proactively improve pipeline reliability and performance.

 

Required Skills

  • 3–5 years of hands-on Data Engineering experience
  • Strong experience with PySpark / Apache Spark
  • Strong Python programming skills
  • Advanced SQL and experience working with relational databases
  • Hands-on experience developing ETL/ELT pipelines
  • Good understanding of data warehousing and data lake concepts
  • Experience with distributed data processing and large datasets
  • Familiarity with Git and CI/CD practices
  • Experience with workflow orchestration tools such as Apache Airflow, Databricks Workflows, or similar
  • Exposure to cloud platforms such as AWS, Azure, or GCP

 

Key Responsibilities

  • B.Tech/B.E./Relevant bachelor's degree

 

Good to Have

  • Experience with Databricks
  • Experience with Delta Lake
  • Knowledge of Kafka / streaming pipelines
  • Experience with AWS Glue, S3, EMR, or equivalent cloud data services
  • Familiarity with Docker/Kubernetes
  • Experience with data modeling and dimensional modeling
  • Knowledge of monitoring and data quality tools