Search by job, company or skills

PySpark, Spark Developer

PySpark, Spark Developer

Infosys
5-8 Years
Not Disclosed

This job is no longer accepting applications

Job Description

Required Skills Strong experience in PySpark, Apache Spark, and Python. Good understanding of Spark Core, Spark SQL, DataFrames, and RDDs. Experience with Hive, HDFS, SQL, and data warehousing concepts. Knowledge of Azure Databricks, AWS EMR, or Hadoop Ecosystem. Experience with performance tuning and query optimization. Hands-on experience with Git and Agile methodologies. Strong analytical and problem-solving skills.

ob Title: PySpark / Spark Developer Experience: 5-8Years Design, develop, and maintain scalable data processing solutions using Apache Spark and PySpark. Build and optimize ETL/ELT pipelines for large-scale data processing. Develop Spark applications for batch and real-time data processing. Analyze, transform, and load structured and unstructured data from multiple sources. Tune Spark jobs for performance, scalability, and reliability. Work with distributed computing frameworks and big data technologies. Collaborate with Data Engineers, Data Architects, and Business Analysts to understand requirements. Troubleshoot production issues and provide efficient solutions. Implement data quality checks and monitoring mechanisms. Follow coding standards, version control, and CI/CD best practices.

More Info

Job Type:
Industry:
Employment Type:

About Company

Similar Jobs

5-10 yrs
Bengaluru, India
Skills:
Spark SQL, Pyspark, Sql, Performance Tuning, Logging, Dataframes, Unix Scheduling Tool, Advanced Data transformations, Error handling, UDFs, Spark Functions, Monitoring
5-8 yrs
Bengaluru, India
Skills:
Spark SQL, Agile Methodologies, Pyspark, Performance Tuning, Apache Spark, Data Warehousing, Azure Databricks, Sql, Query Optimization, Spark Core, Git, Hive, Hadoop Ecosystem, Python, AWS EMR, RDDs, HDFS, DataFrames
2-5 yrs
Bengaluru, India
Skills:
Hadoop, Pyspark, Scala, Kafka, Data Modeling, HBase, ELT, Hive, Etl, Airflow, HDFS, Spark optimization