Job Description: GCP Data Engineer
Experience: 7–10 Years
Roles & Responsibilities
- Design, develop, and implement scalable data processing solutions using Big Data technologies and distributed systems
- Strong hands-on experience with Python and PySpark for developing efficient data pipelines and applications
- Develop and optimize complex SQL queries, data models, and database solutions
- Analyze and improve performance of Apache Spark workloads through optimization techniques such as partitioning, caching, tuning, and resource optimization
- Design and implement various data processing patterns including Batch Processing, Near Real-Time (NRT), and Real-Time data pipelines
- Understand business requirements and translate them into scalable data architecture and technical solutions
- Work with cloud-based data platforms and services, preferably Google Cloud Platform (GCP)
- Hands-on experience with GCP managed services such as: Google Cloud Storage, Dataproc, Dataflow, Pub/Sub, Cloud Functions, Cloud Composer, BigQuery, BigTable
.
Required Skills
- 7–10 years of experience in Data Engineering / Big Data Engineering
- Strong programming experience in Python, PySpark, SQL
- Hands-on experience with Spark, Hadoop ecosystem, and distributed computing
- Experience designing scalable data pipelines and cloud-based solution
- Strong exposure to GCP data services and cloud architecture
- Good understanding of data modeling, ETL/ELT, and data integration pattern
- Experience with performance tuning and optimization of large-scale data processing workloads
Good to Have
- GCP certification(s)
- Experience in enterprise cloud migration projects
- Knowledge of DevOps practices and CI/CD automation
- Experience working in Agile development environments