Search by job, company or skills

Data Engineer Scala | Kafka | Spark | GCP

  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

We are looking for an experienced Data Engineer with 6+ years of experience in designing, developing, deploying, monitoring, and optimizing scalable batch and streaming data pipelines. The ideal candidate should have strong hands-on expertise in Scala, Apache Spark, Kafka, Airflow, BigQuery, Dataproc, and GCP.

Key Responsibilities

  • Design, develop, deploy, monitor, and optimize ETL/ELT workflows for both batch and real-time/streaming data processing.
  • Develop scalable data processing solutions using Scala and Apache Spark.
  • Build and manage streaming pipelines using Kafka, including handling high-volume data and back-pressure scenarios.
  • Develop and maintain Airflow DAGs for complex data orchestration, dependency management, scheduling, retries, and idempotent execution.
  • Work extensively with Google Cloud Platform (GCP) services including BigQuery and Dataproc.
  • Implement Medallion Architecture (Bronze, Silver, Gold) with appropriate data quality validations, transformations, aggregations, and summarization.
  • Perform performance tuning and optimization of Spark jobs, Kafka pipelines, BigQuery queries, and ETL workflows.
  • Implement monitoring and observability mechanisms to identify and troubleshoot pipeline failures, performance issues, and data-quality problems.
  • Create and maintain runbooks and operational documentation for production support and continuous optimization.
  • Troubleshoot runtime data issues independently and ensure timely resolution of production incidents.
  • Follow best practices for data governance, security, access control, and information security.

Required Skills

  • 6+ years of experience as a Data Engineer.
  • Strong hands-on experience with Scala and Apache Spark.
  • Strong knowledge of Kafka and streaming data processing.
  • Experience with Apache Airflow and workflow orchestration.
  • Hands-on experience with Google Cloud Platform (GCP).
  • Strong experience with BigQuery and Dataproc.
  • Experience building both batch and streaming ETL pipelines.
  • Good understanding of Medallion Architecture and data quality frameworks.
  • Strong understanding of Spark performance optimization and distributed data processing.
  • Knowledge of streaming concepts such as back-pressure handling, fault tolerance, checkpointing, and recovery.
  • Experience with monitoring, observability, troubleshooting, and production support.
  • Good understanding of idempotency, dependency handling, retries, and failure recovery in Airflow.

Good to Have

  • Exposure to data governance and information security.
  • Experience with GCP data services beyond BigQuery and Dataproc.
  • Experience implementing automated data quality checks and reconciliation.
  • Knowledge of CI/CD and infrastructure-as-code practices.
  • Experience creating production runbooks and operational standards.

Ideal Candidate

A strong candidate should be capable of owning end-to-end data pipelines, from development and deployment through production monitoring, troubleshooting, performance optimization, and continuous improvement, across both batch and real-time data environments.

Skills: kafka,gcp,scala,spark

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153795277

Similar Jobs

Noida, India

Skills:

BigQueryScalaApache SparkKafkaDataprocAirflow

Beware of Scammers

We don’t charge money for job offers