Search Jobs

Search by job, company or skills

Pyspark
  • Posted 8 hours ago
  • Be among the first 10 applicants

Job Description

Good to have skills: SQL, Hadoop, Hive, Kafka, Airflow

Key Responsibilities:

  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
  • Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
  • Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
  • Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
  • Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
  • Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
  • Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
  • Contribute to code reviews and follow engineering best practices to improve quality and maintainability. Minimum Qualifications:
  • Education: BTECH, MTECH, MCA, MSC.
  • 2–3 years of experience in data engineering or large-scale data processing roles.
  • Strong hands-on experience with PySpark for building data pipelines and transformations.
  • Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
  • Ability to debug Spark applications and resolve data/job issues effectively. Preferred Qualifications:
  • Experience optimizing Spark workloads (tuning partitions, managing skew, memory/executor settings) for performance and cost efficiency.
  • Exposure to building end-to-end data pipelines with strong data quality checks and automated validations.
  • Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development.
  • Experience collaborating in agile teams, participating in code reviews, and improving engineering standards for data pipelines.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

2-5 yrs
Bengaluru, India
Skills:
Hadoop, Pyspark, Scala, Kafka, Data Modeling, HBase, ELT, Hive, Etl, Airflow, HDFS, Spark optimization
3-5 yrs
Bengaluru, India
Skills:
Github, Data Modelling, Bitbucket, Pyspark, Spark, Data Warehousing, Databricks, Sql, Python, Aws S3, AI capabilities
5-10 yrs
Bengaluru, India
Skills:
Spark SQL, Pyspark, Banking Domain Knowledge, Sql, Performance Tuning, Logging and monitoring, Dataframes, ETL Estimation, Unix Scheduling Tool, Generic process development, Advanced Data transformations, Error handling, UDFS, Spark Functions, Teradata BTEQ
5-8 yrs
Bengaluru, India
Skills:
Spark SQL, Agile Methodologies, Pyspark, Performance Tuning, Apache Spark, Data Warehousing, Azure Databricks, Sql, Query Optimization, Spark Core, Git, Hive, Hadoop Ecosystem, Python, AWS EMR, RDDs, HDFS, DataFrames