Search by job, company or skills

zorba ai

Data Engineer with GCP + Pyspark, Scala

Save
  • Posted 5 days ago
  • Be among the first 10 applicants
Early Applicant

Job Description

Job Title: Data Engineer – GCP, PySpark & ScalaExperience

6+ Years

Job Summary

We are looking for a skilled Data Engineer with strong expertise in GCP, PySpark, and Scala to design, develop, deploy, and optimize scalable data pipelines. The ideal candidate should have hands-on experience building batch ETL workflows, implementing Medallion Architecture, and working with GCP data services. The role requires strong problem-solving skills and experience in performance optimization, workflow orchestration, and data quality.

Key Responsibilities

  • Design, develop, deploy, monitor, and optimize batch ETL/data pipelines using Scala, PySpark, and GCP services.
  • Build and maintain scalable data processing workflows using Google Cloud Platform technologies such as BigQuery, Dataproc, and Cloud Storage.
  • Develop and manage Airflow DAGs, ensuring proper dependency handling, scheduling, retry mechanisms, and idempotent execution.
  • Implement Medallion Architecture (Bronze, Silver, Gold) with robust data quality checks and data summarization processes.
  • Optimize ETL jobs and Spark applications for performance, scalability, and cost efficiency.
  • Monitor production pipelines, create observability dashboards, and develop runbooks for operational support and continuous optimization.
  • Troubleshoot production issues and collaborate with cross-functional teams to deliver reliable data solutions.
  • Follow coding standards, best practices, and CI/CD processes for data engineering solutions.
  • Support data governance initiatives and ensure compliance with enterprise data management standards.

Required Skills

  • 6+ years of experience in Data Engineering.
  • Strong hands-on experience with Scala and PySpark.
  • Expertise in Google Cloud Platform (GCP) services, including BigQuery, Dataproc, Cloud Storage, and related data services.
  • Experience developing, deploying, and optimizing batch ETL workflows.
  • Strong knowledge of Apache Airflow, including workflow orchestration, dependency management, retries, and idempotent execution.
  • Experience implementing Medallion Architecture and data quality frameworks.
  • Strong SQL skills and experience with large-scale data processing.
  • Experience in performance tuning, monitoring, observability, and operational support.
  • Good understanding of distributed data processing and Spark optimization techniques.

Preferred Skills

  • Exposure to data governance and metadata management.
  • Experience with CI/CD pipelines and Git.
  • Familiarity with Agile/Scrum development methodologies.
  • Knowledge of data security and cloud best practices.

Skills: gcp,pyspark,airproc,etl,scala

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151279125

Similar Jobs

Bengaluru, India

Skills:

DockerScalaAWSAkkaMicroservicesPLAY

Bengaluru, India

Skills:

UnixData ModelingCassandraPysparkKafkaELTOoziePythonAWSScalaApache SparkEmrSqlDevopsGitSpark StreamingGcpLinuxDatabricksMongoDBAzureEtlAirflowStructured StreamingDelta Lake

Bengaluru, India

Skills:

JavaHadoopScalaApache SparkKafkaApache AirflowData QualityDockerHelmKubernetesSpark Structured StreamingData Validation

Bengaluru, India

Skills:

KafkaDjangoNosqlFault ToleranceScalaPostgreSQLAkkaDistributed SystemsKubernetesPythonDockerCatsmicroservice architecturesnon-relational databasesEvent SourcingCats EffectPekkoAWS cloud managementevent-driven architectureRelational Databasescloud-native developmentAgile product developmentCQRSobservabilityFast API

Bengaluru, India

Skills:

react.js snowflake DatabricksSQL ServerDesign PatternsAWSApplication SecurityContinuous IntegrationNode.jsKubernetesAzureDockerScalaPostgreSQLSparkcode refactoringdesign-driven developmentperformance optimization tools