Job Summary
We are seeking an experienced Data Engineer (PySpark) to join our Data Engineering team. The ideal candidate will be responsible for designing, developing, and maintaining large-scale data pipelines, data marts, and analytical solutions. The role requires strong expertise in Python, PySpark, SQL, Data Warehousing, and Big Data technologies, along with hands-on experience across the end-to-end Software Development Life Cycle (SDLC).
Key Responsibilities
- Design, develop, and maintain scalable ETL pipelines and data processing frameworks.
- Build and support Data Mart solutions to meet business and analytical requirements.
- Develop high-quality, maintainable, and efficient code using Python and PySpark.
- Participate in the complete Software Development Life Cycle (SDLC), including:
- Requirement Analysis
- Development
- Unit Testing
- UAT Support
- Defect Fixing
- Production Deployment
- Post-Production Support
- Perform complex data analysis and troubleshoot data-related issues.
- Debug and optimize PySpark applications for performance and reliability.
- Write and optimize complex Oracle SQL queries.
- Work with structured, semi-structured, and unstructured datasets.
- Implement data quality, validation, and monitoring processes.
- Collaborate with cross-functional teams to resolve dependencies and ensure timely project delivery.
- Follow software engineering best practices, coding standards, testing methodologies, and CI/CD processes.
Required Skills & Experience
Data Engineering & Big Data
- PySpark
- Apache Spark
- Hadoop Ecosystem
- MapReduce
- Hive
- ETL Development
- Data Pipeline Development
- Data Mart Development
- Data Warehousing
Programming
- Python
- Pandas
- Jupyter Notebook
Database Technologies
- Oracle SQL
- SQL
- NoSQL Databases
- Query Optimization
- Data Analysis
Workflow & Automation
- Apache Airflow
- Oozie
- Jenkins
- Git
- CI/CD Pipelines
Data Engineering Techniques
- Data Cleansing
- Data Linking
- Feature Engineering
- Data Imputation Techniques
- Data Validation and Quality Checks
Preferred Experience
- Banking or Financial Services domain experience.
- Experience supporting enterprise-scale data platforms.
- Exposure to production support and deployment activities.
- Experience working in Agile environments.
Desired Competencies
- Strong analytical and problem-solving skills.
- Excellent debugging and troubleshooting capabilities.
- Ability to lead technical initiatives with ownership and accountability.
- Strong collaboration and stakeholder management skills.
- Excellent verbal and written communication skills.
- Ability to communicate effectively with both technical and non-technical stakeholders.
- Ability to prioritize tasks and perform effectively in a fast-paced environment.