Job Summary
We are looking for an experienced
AWS Data Engineer with strong expertise in
Python, PySpark, Databricks, and AWS to design, develop, and optimize scalable data pipelines and cloud-based data solutions. The ideal candidate should have hands-on experience with big data processing, data warehousing, and ETL development on AWS.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Python and PySpark.
- Develop and optimize data processing workflows using Databricks.
- Build and manage cloud-based data solutions on AWS.
- Work with structured and unstructured data from multiple sources.
- Optimize Spark jobs for performance and cost efficiency.
- Develop reusable data engineering frameworks and automation.
- Ensure data quality, governance, security, and reliability.
- Collaborate with data analysts, data scientists, and business stakeholders.
- Troubleshoot production issues and implement performance improvements.
Required Skills
- 6+ years of experience in Data Engineering.
- Strong programming skills in Python.
- Hands-on experience with PySpark.
- Experience working with Databricks.
- Strong knowledge of AWS services such as S3, Glue, EMR, Lambda, Redshift, IAM, EC2, CloudWatch, and Athena.
- Strong SQL skills and experience with relational databases.
- Experience building scalable ETL/ELT pipelines.
- Knowledge of Delta Lake, Spark optimization, and performance tuning.
- Experience with Git and CI/CD pipelines.
- Good understanding of data modeling and data warehousing concepts.
Skills: databricks,aws,python