Job Title: Data Engineer – GCP, PySpark & ScalaExperience
6+ Years
Job Summary
We are looking for a skilled Data Engineer with strong expertise in GCP, PySpark, and Scala to design, develop, deploy, and optimize scalable data pipelines. The ideal candidate should have hands-on experience building batch ETL workflows, implementing Medallion Architecture, and working with GCP data services. The role requires strong problem-solving skills and experience in performance optimization, workflow orchestration, and data quality.
Key Responsibilities
- Design, develop, deploy, monitor, and optimize batch ETL/data pipelines using Scala, PySpark, and GCP services.
- Build and maintain scalable data processing workflows using Google Cloud Platform technologies such as BigQuery, Dataproc, and Cloud Storage.
- Develop and manage Airflow DAGs, ensuring proper dependency handling, scheduling, retry mechanisms, and idempotent execution.
- Implement Medallion Architecture (Bronze, Silver, Gold) with robust data quality checks and data summarization processes.
- Optimize ETL jobs and Spark applications for performance, scalability, and cost efficiency.
- Monitor production pipelines, create observability dashboards, and develop runbooks for operational support and continuous optimization.
- Troubleshoot production issues and collaborate with cross-functional teams to deliver reliable data solutions.
- Follow coding standards, best practices, and CI/CD processes for data engineering solutions.
- Support data governance initiatives and ensure compliance with enterprise data management standards.
Required Skills
- 6+ years of experience in Data Engineering.
- Strong hands-on experience with Scala and PySpark.
- Expertise in Google Cloud Platform (GCP) services, including BigQuery, Dataproc, Cloud Storage, and related data services.
- Experience developing, deploying, and optimizing batch ETL workflows.
- Strong knowledge of Apache Airflow, including workflow orchestration, dependency management, retries, and idempotent execution.
- Experience implementing Medallion Architecture and data quality frameworks.
- Strong SQL skills and experience with large-scale data processing.
- Experience in performance tuning, monitoring, observability, and operational support.
- Good understanding of distributed data processing and Spark optimization techniques.
Preferred Skills
- Exposure to data governance and metadata management.
- Experience with CI/CD pipelines and Git.
- Familiarity with Agile/Scrum development methodologies.
- Knowledge of data security and cloud best practices.
Skills: gcp,pyspark,airproc,etl,scala