About The Opportunity
A fast-scaling technology services firm operating in the Data Engineering and Cloud Analytics space, we partner with enterprises to build scalable data pipelines, automate ETL workflows, and unlock real-time insights from structured and unstructured datasets. Our Python PySpark developers architect and deploy high-performance data solutions on cloud platforms—enabling faster decision-making, predictive modeling, and regulatory compliance for global clients.
Role & Responsibilities
- Design, develop, and optimize PySpark ETL pipelines for large-scale batch and streaming data processing.
- Collaborate with data scientists and analysts to ingest, transform, and model data for ML training and reporting use cases.
- Implement data quality checks, error handling, and logging frameworks within Spark jobs for production-grade reliability.
- Integrate PySpark workflows with cloud data lakes (AWS S3, Azure Data Lake), metastores (Hive, Glue), and orchestration tools (Airflow, Luigi).
- Performance-tune Spark applications—partitioning, caching, broadcasting, and resource allocation—to reduce job runtimes and costs.
- Write clean, modular, testable Python code with unit/integration tests and document architecture decisions for team scalability.
Skills & Qualifications
Must-Have
- PySpark
- Python
- Apache Spark
- ETL Development
- SQL
- Data Lake Architecture
- AWS S3
- Apache Airflow
Preferred
- Azure Data Lake
- Databricks
- Delta Lake
Benefits & Culture Highlights
- On-site collaborative workspace in major Indian tech hubs with modern infrastructure and R&D labs.
- Opportunities to upskill in cloud-native data platforms and GenAI-powered data pipelines.
- Fast-track career growth with cross-functional exposure to data science, ML engineering, and cloud architecture teams.
Skills: data,python,lake,spark,cloud