Key Responsibilities:
- ob Title PySpark Spark Developer
- Experience 5 8Years
- Design develop and maintain scalable data processing solutions using Apache Spark and PySpark
- Build and optimize ETL ELT pipelines for large scale data processing
- Develop Spark applications for batch and real time data processing
- Analyze transform and load structured and unstructured data from multiple sources
- Tune Spark jobs for performance scalability and reliability
- Work with distributed computing frameworks and big data technologies
- Collaborate with Data Engineers Data Architects and Business Analysts to understand requirements
- Troubleshoot production issues and provide efficient solutions
- Implement data quality checks and monitoring mechanisms
- Follow coding standards version control and CI CD best practices
Technical Requirements:
- Required Skills
- Strong experience in PySpark Apache Spark and Python
- Good understanding of Spark Core Spark SQL DataFrames and RDDs
- Experience with Hive HDFS SQL and data warehousing concepts
- Knowledge of Azure Databricks AWS EMR or Hadoop Ecosystem
- Experience with performance tuning and query optimization
- Hands on experience with Git and Agile methodologies
- Strong analytical and problem solving skills
Preferred Skills:
Technology->Big Data - Data Processing->PySpark,Technology->Big Data - Data Processing->Spark