Job Description:
- We are seeking a skilled Pyspark Developer with 5 9 years of experience in building scalable solutions on Pyspark
- The ideal candidate should have strong expertise in Pyspark with Spark Hadoop SQL exposure to Hive Sqoop UNIX shell scripting
Key Responsibilities:
- At least 6 years of experience in designing and developing large scale distributed data processing pipelines using PySpark Hadoop and related technologies
- Having expertise in Pyspark Spark Core Spark SQL Batch processing and Spark Streaming
- Experience with Hadoop HDFS Hive and other BigData technologies
- Familiarity with Data warehousing and ETL concepts and techniques
- UNIX shell scripting will be an added advantage in scheduling running application jobs
- At least 5 years of experience in Project development life cycle activities and development maintenance projects
- Work with business stakeholders and other SMEs to understand high level business requirements
- Work with the Solution Designers and contribute to the development of project plans by participating in the scoping and estimating of proposed project
- Apply technical background understanding business knowledge system knowledge in the elicitation of Systems Requirements for projects
- Possess good knowledge on Spark architecture and transformations using Spark and PySpark
- Work in an Agile environment and participation in scrum daily standups sprint planning reviews and retrospectives
- Understand project requirements and translate them into technical solutions which meets the project quality standards
- Ability to work in team in diverse multiple stakeholder environment and collaborate with upstream downstream functional teams to identify troubleshoot and resolve data issues
- Strong problem solving and Good Analytical skills
- Excellent verbal and written communication skills
- Experience and desire to work in a Global delivery environment
- Stay up to date with new technologies and industry trends in Development
Technical Requirements:
- Primary skills Pyspark Spark Hadoop SQL
Preferred Skills:
Technology->Big Data - Data Processing->PySpark