Search Jobs

Search by job, company or skills

Pyspark
  • Posted 11 hours ago
  • Be among the first 10 applicants

Job Description

  • Primary skills:Technology->Big Data - Data Processing->PySpark

Key Responsibilities

  • Design, develop, and maintain scalable batch data pipelines using PySpark and Apache Spark for large datasets.
  • Perform data ingestion, transformation, and enrichment while ensuring accuracy, completeness, and consistency of outputs.
  • Optimize Spark jobs for performance (partitioning, caching, joins, shuffles) and improve runtime efficiency and resource utilization.
  • Implement robust error handling, logging, and monitoring to ensure reliable pipeline execution and faster issue resolution.
  • Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver well-documented solutions.
  • Conduct code reviews, follow engineering best practices, and contribute to reusable components and standards.
  • Troubleshoot production issues, perform root-cause analysis, and drive corrective and preventive actions. Minimum Qualifications:
  • Bachelor's degree (or equivalent) in Engineering/Computer Science/IT or related field.
  • 3–5 years of experience in data engineering or big data development roles.
  • Strong hands-on experience with PySpark and Apache Spark for building data processing workflows.
  • Solid understanding of distributed data processing concepts and performance tuning fundamentals.
  • Ability to translate business requirements into technical implementations and deliver within timelines. Preferred Qualifications:
  • Experience building end-to-end Spark applications including job orchestration, dependency management, and production support readiness.
  • Strong data transformation skills with a focus on data quality checks, reconciliation, and pipeline reliability.
  • Exposure to designing modular, reusable Spark components and maintaining clean, maintainable codebases.
  • Familiarity with structured and semi-structured data formats and efficient processing patterns in Spark.
  • Proven ability to collaborate effectively across teams, communicate clearly, and contribute to continuous improvement initiatives. Good to have skills: Spark SQL, Delta Lake, Databricks, Airflow, Hadoop

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

About Company

Similar Jobs

2-5 yrs
Bengaluru, India
Skills:
Hadoop, Pyspark, Scala, Kafka, Data Modeling, HBase, ELT, Hive, Etl, Airflow, HDFS, Spark optimization
6-10 yrs
Bengaluru, India
Skills:
Hld, BigQuery, Unit Testing, Hadoop, Pyspark, Apache Spark, Google Cloud, User Acceptance Testing, Jenkins, Git, Hive, Tdd, System Testing, Shell scripting, Python Programming, Cloudera, AWS ecosystem, Hortonworks Data Platform, CDC operations, UNIX operating system concepts, Agile delivery model, AWS S3 Filesystem operations
3-5 yrs
Bengaluru, India
Skills:
Github, Data Modelling, Bitbucket, Pyspark, Spark, Data Warehousing, Databricks, Sql, Python, Aws S3, AI capabilities
5-10 yrs
Bengaluru, India
Skills:
Spark SQL, Pyspark, Banking Domain Knowledge, Sql, Performance Tuning, Logging and monitoring, Dataframes, ETL Estimation, Unix Scheduling Tool, Generic process development, Advanced Data transformations, Error handling, UDFS, Spark Functions, Teradata BTEQ
6-8 yrs
Bengaluru, India
Skills:
Spark SQL, Spark Core, Hive, Spark Streaming, Hadoop, Pyspark, Data Warehousing, Unix Shell Scripting, Etl, HDFS, Batch Processing