Search Jobs

Search by job, company or skills

Databricks

Databricks

Infosys
  • Posted 10 hours ago
  • Be among the first 10 applicants

Job Description

  • Primary skills:Technology->Data Engineering->Databricks

Key Responsibilities:

  • Develop and maintain scalable ETL pipelines using Databricks and PySpark to process large datasets efficiently.
  • Implement data transformations, cleansing, and enrichment logic aligned to business and analytics requirements.
  • Optimize Spark jobs for performance and cost by tuning partitions, caching, and cluster configurations where applicable.
  • Build reusable notebooks/jobs and support scheduling/orchestration of workloads within the Databricks environment.
  • Perform data validation, reconciliation, and quality checks to ensure accuracy and reliability of curated datasets.
  • Troubleshoot pipeline failures, analyze logs, and resolve issues to maintain stable production operations.
  • Collaborate with cross-functional teams to gather requirements, provide estimates, and deliver enhancements iteratively.
  • Maintain clear technical documentation for pipelines, transformations, and operational runbooks. Minimum Qualifications:
  • Bachelor's degree (or equivalent) in Engineering/Technology/Computer Science or related field (BTech/BE/MSc or equivalent).
  • 3–5 years of experience in data engineering or related roles with hands-on Databricks experience.
  • Strong hands-on development experience with PySpark for distributed data processing.
  • Proven experience building and supporting ETL pipelines in production environments.
  • Ability to analyze data issues, debug Spark/ETL jobs, and implement reliable fixes. Preferred Qualifications:
  • Experience designing end-to-end data workflows on Databricks including job scheduling, monitoring, and operational support.
  • Strong understanding of data modeling concepts and building curated datasets for analytics use cases.
  • Familiarity with Delta Lake concepts (ACID tables, incremental processing, upserts/merges) and best practices for lakehouse implementations.
  • Experience with performance tuning techniques for Spark workloads and handling large-scale datasets efficiently.
  • Exposure to CI/CD practices for data pipelines and maintaining code quality through reviews and standards. Good to have skills: Delta Lake, Spark SQL, Data Modeling, Workflow Orchestration, Performance Tuning

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

6-8 yrs
Bengaluru, India
Skills:
Pyspark, ELT, Azure Synapse, Azure Functions, Terraform, Python, Power Bi, Sql, Azure Data Factory, Databricks, Mosaic AI, Cost Optimization for Databricks, Performance Optimisation of DLT Pipelines, Medallion Architecture, Delta Live Tables, CI CD, Asset Bundle, LLMs, Logic Apps, AI GenAI-driven Innovation, Gen AI Genie, Agents, Structured Live Streaming, GitHub CoPilot, Unity Catalog, ADLS, Photon Serverless
5-15 yrs
Bengaluru, India
Skills:
snowflake , Denodo, Adf, Databricks, Sql, Python
6-8 yrs
Bengaluru, India
Skills:
Spark SQL, Sftp, Pyspark, Amazon S3, Databricks, Rest Apis, Sql, Python, Delta Lake
3-5 yrs
Bengaluru, India
Skills:
Networking, Prometheus, Grafana, Python Scripting, Storage, Cloudwatch, Terraform, Iam, Alerting, Cloud platforms, Databricks Administration, Compute, CI/CD, Monitoring
5-7 yrs
Bengaluru, India
Skills:
snowflake , Data Modeling, Sql, ELT, Apache Airflow, Git, MLops, Docker, Data Architecture, Gitlab, Databricks, Rest Apis, Azure, Kubernetes, Python, Etl, AWS, cdc, Airbyte, AI Coding Assistants