Job Title: Consultant | Data Analytics | Delhi | Operations, Industry & Domain Solutions | DigiGov

Consultant | Data Analytics | Delhi | Operations, Industry & Domain Solutions | DigiGov
• Job requisition ID : 111595
• Location: Delhi
• Entity: Deloitte Touche Tohmatsu India LLP
Your work profile
We are looking for a Data Engineer with 3–5 years of experience to design, develop, and maintain scalable data pipelines and ETL workflows. The ideal candidate should have strong hands-on experience with PySpark, Python, SQL, and data engineering frameworks, along with a good understanding of distributed data processing.
Key Responsibilities
- Design, develop, and maintain ETL/ELT pipelines for processing large volumes of data.
- Build scalable and efficient PySpark data pipelines for batch and/or streaming workloads.
- Develop data transformation, cleansing, validation, and aggregation processes.
- Work with data lakes, data warehouses, and cloud-based data platforms.
- Optimize Spark jobs and pipelines for performance, scalability, and reliability.
- Write complex SQL queries for data extraction, transformation, and analysis.
- Implement data quality checks, error handling, logging, and monitoring.
- Collaborate with data analysts, data scientists, and application teams to understand data requirements.
- Troubleshoot pipeline failures and proactively improve pipeline reliability and performance.
Required Skills
- 3–5 years of hands-on Data Engineering experience
- Strong experience with PySpark / Apache Spark
- Strong Python programming skills
- Advanced SQL and experience working with relational databases
- Hands-on experience developing ETL/ELT pipelines
- Good understanding of data warehousing and data lake concepts
- Experience with distributed data processing and large datasets
- Familiarity with Git and CI/CD practices
- Experience with workflow orchestration tools such as Apache Airflow, Databricks Workflows, or similar
- Exposure to cloud platforms such as AWS, Azure, or GCP
Key Responsibilities
- B.Tech/B.E./Relevant bachelor's degree
Good to Have
- Experience with Databricks
- Experience with Delta Lake
- Knowledge of Kafka / streaming pipelines
- Experience with AWS Glue, S3, EMR, or equivalent cloud data services
- Familiarity with Docker/Kubernetes
- Experience with data modeling and dimensional modeling
- Knowledge of monitoring and data quality tools
