AWS Glue, Python
- Posted 20 hours ago
- Be among the first 10 applicants
Job Description
Job Description:
- About the job
- Join a collaborative team where you ll help transform raw data into trusted analytics ready assets that power smarter decisions
- In this role you will design and optimize cloud based data integration workflows using AWS Glue and Python partnering closely with engineers analysts and stakeholders to deliver reliable pipelines at scale
- You ll work in an environment that values clean engineering continuous improvement and shared ownership where your ideas on performance automation and data quality are welcomed
- If you enjoy solving real world data challenges building robust ETL processes and learning modern AWS data services while contributing to impactful outcomes this opportunity offers the right mix of hands on development and problem solving responsibility
Key Responsibilities:
- Key Responsibilities
- Design build and maintain ETL ELT pipelines using AWS Glue and Python for batch and incremental data processing
- Develop and optimize Glue Jobs PySpark Python including job parameters bookmarks retries and performance tuning
- Implement data ingestion transformation and validation logic to ensure accuracy completeness and consistency of datasets
- Integrate pipelines with AWS services e
- g
- S3 IAM CloudWatch to enable secure observable and scalable workflows
- Troubleshoot job failures analyze logs metrics and implement fixes to improve stability and runtime efficiency
- Collaborate with cross functional teams to gather requirements define data mappings and deliver well documented solutions
- Follow engineering best practices including code reviews version control and reusable modular coding patterns
- Minimum Qualifications
- Bachelor s degree or equivalent e
- g
- BE BTech MSc MCA MTech
- 3 5 years of experience in data engineering ETL development or data integration roles
- Strong hands on experience with AWS Glue and Python for building production grade data pipelines
- Working knowledge of core AWS concepts including security basics IAM storage patterns and monitoring
- Ability to debug data pipeline issues and deliver reliable solutions with clear documentation
Technical Requirements:
- Primary skills AWS Glue Technology Cloud Platform Amazon Webservices DevOps Technology Cloud Platform AWS Data Analytics AWS Glue DataBrew Technology OpenSystem Python OpenSystem
Additional Responsibilities:
- Preferred Qualifications
- Experience with PySpark and distributed data processing patterns within AWS Glue
- Strong SQL skills and experience working with structured semi structured datasets CSV JSON Parquet
- Exposure to orchestration and scheduling patterns for ETL workflows and dependency management
- Familiarity with data quality checks schema evolution handling and building resilient pipelines
- Experience collaborating in Agile teams and contributing to CI CD or automated deployment practices for data jobs
- Good to have skills
- PySpark Amazon S3 AWS IAM Amazon CloudWatch SQL
