Search by job, company or skills

Data Engineer (Spark/Scala)

Early Applicant
  • Posted 2 days ago
  • Be among the first 40 applicants

Job Description

We are looking for an experienced Data Engineer (Spark/Scala) with strong hands-on expertise in Apache Spark, Databricks, Scala, PySpark, Python, and SQL.

The role involves designing and developing large-scale data pipelines across on-premises and cloud environments, working with multiple file systems and data formats, modernizing legacy workflows, and supporting hybrid data architectures.

Key Responsibilities

  • Design, develop, and maintain scalable data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark.
  • Build and support complex on-premises data workflows and hybrid on-prem-to-cloud data integration solutions.
  • Integrate data across HDFS, NAS, on-prem file shares, Amazon S3, and other storage platforms.
  • Work with multiple data formats including JSON, Parquet, CSV, Avro, Fixed-Length, and Excel.
  • Develop optimized SQL queries for data extraction, transformation, and loading.
  • Connect to multiple relational and non-relational databases and implement performance-efficient data extraction strategies.
  • Develop and maintain workflow orchestration using Apache Airflow or similar scheduling tools.
  • Write clean, production-grade Python code for data processing, automation, and engineering utilities.
  • Develop unit, integration, and data-quality tests for data pipelines.
  • Troubleshoot pipeline failures, performance bottlenecks, data quality issues, and complex multi-system integration problems.
  • Support migration and modernization of legacy on-premises data processes to hybrid/cloud environments.
  • Collaborate with Data Scientists, Analysts, Application Engineers, and other stakeholders.
  • Create technical documentation covering pipelines, data flows, architecture, and data lineage.
  • Support cloud integration initiatives, particularly across Azure environments.
  • Leverage coding assistants and AI agents to improve development productivity and automate engineering tasks.

Required Skills

  • Strong hands-on experience with Apache Spark and Databricks.
  • Strong experience with Scala/Spark Scala and PySpark.
  • Strong Python programming skills.
  • Strong SQL, including complex joins, query optimization, and performance tuning.
  • Hands-on experience with Amazon S3.
  • Experience working with HDFS, NAS, on-prem file systems, and cloud storage.
  • Strong experience handling JSON, Parquet, CSV, Avro, Fixed-Length, and Excel data formats.
  • Experience extracting data efficiently from multiple databases.
  • Strong understanding of complex on-premises data workflows and multi-system integrations.
  • Experience building hybrid on-prem/cloud data pipelines.
  • Strong troubleshooting and production support skills.

Secondary Skills

  • Azure cloud services.
  • Apache Airflow or similar workflow orchestration tools.
  • Automated unit and integration testing.
  • Data quality validation and monitoring.
  • Data lineage and technical documentation.

Good to Have

  • Working knowledge of Java.
  • Experience with Prefect.
  • Familiarity with React for internal tools or dashboards.
  • Experience using AI coding assistants and AI agents.
  • PBM / Pharmacy Benefit Management / Healthcare domain experience.

Preferred Candidate Profile

Candidates with strong experience in Scala + Spark + Databricks + PySpark, combined with on-premises data engineering and hybrid cloud integration, will be preferred.

Key Skills: Apache Spark, Scala, Spark Scala, Databricks, PySpark, Python, SQL, Amazon S3, HDFS, Airflow, Azure, On-Premise Data Engineering, ETL, Data Pipelines.

Skills: cloud,apache spark,scala

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152502817

Beware of Scammers

We don’t charge money for job offers