Job Title: Senior Consultant | AWS Data Engineering | Mumbai | Engineering | Data Modernization & Migration
AWS Data Engineer
Job Title
AWS Data Engineer
Experience
3–8 Years
Location
Hybrid / Remote / Onsite
Employment Type
Full-Time
Graduation - B.E/B.Tech
Job Summary
We are looking for an experienced AWS Data Engineer to design, develop, and maintain scalable data pipelines and data lake solutions on AWS. The ideal candidate should have hands-on experience with AWS data services, SQL, Spark, ETL development, and modern data lake architectures. Experience in Banking and Financial Services (BFSI) is highly preferred.
The candidate will work closely with business analysts, data architects, and reporting teams to build reliable, high-performance data pipelines that support enterprise analytics and regulatory reporting.
Key Responsibilities
* Design, develop, and maintain scalable ETL/ELT pipelines on AWS.
* Build and manage enterprise Data Lakes using Amazon S3 and Apache Iceberg.
* Develop and optimize Spark applications using PySpark on Amazon EMR.
* Create and optimize SQL queries using Amazon Athena.
* Design and maintain dimensional data models and data marts.
* Implement incremental loading, CDC, and data reconciliation processes.
* Develop data ingestion frameworks from multiple source systems.
* Optimize Spark jobs for performance, scalability, and cost efficiency.
* Monitor and troubleshoot data pipelines, job failures, and production issues.
* Implement data quality checks and validation frameworks.
* Collaborate with business users to understand reporting and analytics requirements.
* Participate in code reviews, unit testing, integration testing, and deployment activities.
* Create technical documentation, mapping documents, and operational runbooks.
* Follow data governance, security, and compliance standards.
Required Technical Skills
AWS Services
* Amazon S3
* AWS Glue
* AWS EMR
* Amazon Athena
* AWS Lambda
* IAM
* CloudWatch
* EventBridge
* AWS Step Functions
* Glue Data Catalog
Big Data Technologies
* Apache Spark
* PySpark
* Apache Iceberg
* Hive
* Parquet
Database Technologies
* SQL
* PostgreSQL
* Oracle
* Teradata
* Amazon Redshift (preferred)
Programming
* Python
* SQL
* Shell Scripting
Data Engineering
* ETL/ELT Development
* Data Warehousing
* Data Lake Architecture
* Data Modeling
* Star Schema
* Snowflake Schema
* Slowly Changing Dimensions (SCD)
* Partitioning
* Data Quality Frameworks
Reporting
Exposure to one or more of:
* Amazon QuickSight
* Tableau
* Power BI
Version Control
* Git
Required Experience
* 5–10 years of experience in Data Engineering.
* Strong experience with AWS Data Services.
* Hands-on experience in PySpark development on EMR.
* Strong SQL development and query optimization skills.
* Experience working with large datasets.
* Experience developing Data Lakes and Data Marts.
* Experience with Apache Iceberg tables is highly preferred.
* Experience building production-grade ETL pipelines.
* Good understanding of data reconciliation and validation techniques.
* Experience with Agile/Scrum methodologies.
Preferred Qualifications
* Banking or Financial Services domain experience.
* Experience with Finacle or Core Banking data models.
* Experience in Enterprise Data Warehouse (EDW) migration projects.
* Knowledge of regulatory reporting.
* AWS Certified Data Engineer – Associate or AWS Certified Solutions Architect.
* Databricks or Apache Spark certifications are a plus.
Soft Skills
* Strong analytical and problem-solving abilities.
* Excellent communication and collaboration skills.
* Ability to work independently and take ownership of deliverables.
* Strong debugging and troubleshooting skills.
* Good stakeholder management and documentation skills.
* Ability to work under tight deadlines and manage multiple priorities.
Nice to Have
* Experience with CI/CD pipelines.
* Knowledge of Apache Airflow or workflow orchestration tools.
* Experience with Terraform or CloudFormation.
* Familiarity with Data Governance and Metadata Management tools.
* Exposure to streaming technologies such as Apache Kafka or Amazon Kinesis.
* Experience implementing Lakehouse architecture.