Job Description
We are looking for a GCP Data Engineer with strong expertise in Python and PySpark to join our team. The ideal candidate will have hands-on experience in building scalable ETL/ELT pipelines, processing large datasets, and working with Google Cloud Platform (GCP) services. The candidate should possess excellent communication skills and be available to join immediately.
Mandatory Skills
- Google Cloud Platform (GCP)
- Python
- PySpark
- ETL/ELT Pipeline Development
- BigQuery
- Google Cloud Storage (GCS)
- SQL
- Data Engineering concepts
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines on GCP.
- Develop data processing solutions using Python and PySpark.
- Build and optimize data pipelines using BigQuery, GCS, and other GCP services.
- Write efficient SQL queries for data extraction, transformation, and analysis.
- Process structured and semi-structured data from multiple sources.
- Monitor, troubleshoot, and optimize data pipelines for performance and reliability.
- Collaborate with cross-functional teams to understand business requirements and deliver data solutions.
- Ensure data quality, security, and governance across data platforms.
- Support deployment, automation, and operational activities within GCP environments.
Required Experience
- 4+ years of experience in Data Engineering.
- Hands-on experience with Google Cloud Platform (GCP).
- Strong programming experience in Python and PySpark.
- Experience working with BigQuery, Google Cloud Storage (GCS), and SQL.
- Good understanding of data warehousing and ETL concepts.
- Experience in performance tuning and optimization of data pipelines.
Preferred Skills
- Experience with Dataproc, Airflow, Pub/Sub, or Cloud Composer.
- Knowledge of Spark SQL and DataFrames.
- Familiarity with CI/CD and version control tools.
- Experience working in Agile environments.