Search Jobs

Search by job, company or skills

Spark-Scala, Databricks

Spark-Scala, Databricks

Infosys
  • Posted 11 hours ago
  • Be among the first 10 applicants

Job Description

Primary skills:Technology->Big Data - Data Processing->Spark Technology->Data Engineering->Databricks Technology->Functional Programming->Scala

Key Responsibilities: Data Engineering & ETL Development

  • Design, develop, and maintain ETL pipelines using Spark-Scala on Databricks for batch and incremental processing.
  • Implement data transformations, joins, aggregations, and validations to ensure accurate and consistent outputs.
  • Build reusable Spark components and follow best practices for modular, maintainable code. Performance, Reliability & Operations
  • Tune Spark jobs (partitioning, caching, shuffle optimization) to improve performance and cost efficiency.
  • Monitor job execution, troubleshoot failures, and provide timely production support with root-cause analysis.
  • Implement logging, error handling, and data quality checks to improve pipeline reliability. Collaboration & Delivery
  • Work with cross-functional teams to gather requirements and translate them into technical solutions.
  • Participate in code reviews, documentation, and knowledge sharing to uplift team standards.
  • Support release cycles by validating outputs, ensuring backward compatibility, and maintaining deployment readiness. Minimum Qualifications:
  • Bachelor's/Master's degree (BE/BTech/MSc/MCA/MTech or equivalent).
  • 3–5 years of experience in data engineering or ETL development roles.
  • Strong hands-on experience with Spark using Scala and working on Databricks.
  • Solid understanding of ETL concepts, data transformations, and pipeline troubleshooting.
  • Ability to write clean, testable code and collaborate effectively with technical and non-technical stakeholders. Preferred Qualifications:
  • Experience building end-to-end pipelines on Databricks including notebooks, jobs/workflows, and cluster configuration basics.
  • Strong SQL skills and experience integrating Spark pipelines with structured data sources and curated datasets.
  • Familiarity with data quality frameworks, reconciliation strategies, and automated validation checks.
  • Exposure to CI/CD practices for data engineering (version control, automated testing, release management).
  • Proven ability to optimize distributed workloads and deliver measurable improvements in runtime and stability. Good to have skills: SQL, Delta Lake, Apache Airflow, Azure Data Lake Storage (ADLS), Git

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

6-8 yrs
Bengaluru, India
Skills:
Pyspark, ELT, Azure Synapse, Azure Functions, Terraform, Python, Power Bi, Sql, Azure Data Factory, Databricks, Mosaic AI, Cost Optimization for Databricks, Performance Optimisation of DLT Pipelines, Medallion Architecture, Delta Live Tables, CI CD, Asset Bundle, LLMs, Logic Apps, AI GenAI-driven Innovation, Gen AI Genie, Agents, Structured Live Streaming, GitHub CoPilot, Unity Catalog, ADLS, Photon Serverless
5-15 yrs
Bengaluru, India
Skills:
snowflake , Denodo, Adf, Databricks, Sql, Python
6-8 yrs
Bengaluru, India
Skills:
Spark SQL, Sftp, Pyspark, Amazon S3, Databricks, Rest Apis, Sql, Python, Delta Lake
3-5 yrs
Bengaluru, India
Skills:
Networking, Prometheus, Grafana, Python Scripting, Storage, Cloudwatch, Terraform, Iam, Alerting, Cloud platforms, Databricks Administration, Compute, CI/CD, Monitoring
5-7 yrs
Bengaluru, India
Skills:
snowflake , Data Modeling, Sql, ELT, Apache Airflow, Git, MLops, Docker, Data Architecture, Gitlab, Databricks, Rest Apis, Azure, Kubernetes, Python, Etl, AWS, cdc, Airbyte, AI Coding Assistants