I
Pyspark
I
Pyspark
Infosys3-5 Years
- Posted 11 hours ago
- Be among the first 10 applicants
Job Description
- Primary skills:Technology->Big Data - Data Processing->PySpark
- Design, develop, and maintain scalable batch data pipelines using PySpark and Apache Spark for large datasets.
- Perform data ingestion, transformation, and enrichment while ensuring accuracy, completeness, and consistency of outputs.
- Optimize Spark jobs for performance (partitioning, caching, joins, shuffles) and improve runtime efficiency and resource utilization.
- Implement robust error handling, logging, and monitoring to ensure reliable pipeline execution and faster issue resolution.
- Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver well-documented solutions.
- Conduct code reviews, follow engineering best practices, and contribute to reusable components and standards.
- Troubleshoot production issues, perform root-cause analysis, and drive corrective and preventive actions. Minimum Qualifications:
- Bachelor's degree (or equivalent) in Engineering/Computer Science/IT or related field.
- 3–5 years of experience in data engineering or big data development roles.
- Strong hands-on experience with PySpark and Apache Spark for building data processing workflows.
- Solid understanding of distributed data processing concepts and performance tuning fundamentals.
- Ability to translate business requirements into technical implementations and deliver within timelines. Preferred Qualifications:
- Experience building end-to-end Spark applications including job orchestration, dependency management, and production support readiness.
- Strong data transformation skills with a focus on data quality checks, reconciliation, and pipeline reliability.
- Exposure to designing modular, reusable Spark components and maintaining clean, maintainable codebases.
- Familiarity with structured and semi-structured data formats and efficient processing patterns in Spark.
- Proven ability to collaborate effectively across teams, communicate clearly, and contribute to continuous improvement initiatives. Good to have skills: Spark SQL, Delta Lake, Databricks, Airflow, Hadoop

