I
Databricks
I
Databricks
Infosys3-5 Years
- Posted 10 hours ago
- Be among the first 10 applicants
Job Description
- Primary skills:Technology->Data Engineering->Databricks
- Develop and maintain scalable ETL pipelines using Databricks and PySpark to process large datasets efficiently.
- Implement data transformations, cleansing, and enrichment logic aligned to business and analytics requirements.
- Optimize Spark jobs for performance and cost by tuning partitions, caching, and cluster configurations where applicable.
- Build reusable notebooks/jobs and support scheduling/orchestration of workloads within the Databricks environment.
- Perform data validation, reconciliation, and quality checks to ensure accuracy and reliability of curated datasets.
- Troubleshoot pipeline failures, analyze logs, and resolve issues to maintain stable production operations.
- Collaborate with cross-functional teams to gather requirements, provide estimates, and deliver enhancements iteratively.
- Maintain clear technical documentation for pipelines, transformations, and operational runbooks. Minimum Qualifications:
- Bachelor's degree (or equivalent) in Engineering/Technology/Computer Science or related field (BTech/BE/MSc or equivalent).
- 3–5 years of experience in data engineering or related roles with hands-on Databricks experience.
- Strong hands-on development experience with PySpark for distributed data processing.
- Proven experience building and supporting ETL pipelines in production environments.
- Ability to analyze data issues, debug Spark/ETL jobs, and implement reliable fixes. Preferred Qualifications:
- Experience designing end-to-end data workflows on Databricks including job scheduling, monitoring, and operational support.
- Strong understanding of data modeling concepts and building curated datasets for analytics use cases.
- Familiarity with Delta Lake concepts (ACID tables, incremental processing, upserts/merges) and best practices for lakehouse implementations.
- Experience with performance tuning techniques for Spark workloads and handling large-scale datasets efficiently.
- Exposure to CI/CD practices for data pipelines and maintaining code quality through reviews and standards. Good to have skills: Delta Lake, Spark SQL, Data Modeling, Workflow Orchestration, Performance Tuning
More Info
Key Skills
Delta Lake
ETL pipelines
Workflow Orchestration


