AT&T -Data Engineers
Job Description
Overview
We are looking for a Databricks Engineer to join a data modernization program focused on building new data capabilities on Databricks and migrating existing data and application features from Palantir Foundry to the Databricks platform.
The candidate will work closely with data architects, application teams, business stakeholders, and AI/ML engineers to understand existing Palantir functionality, design the equivalent or improved capabilities in Databricks, and develop scalable, production-ready data pipelines and solutions.
The role requires strong hands-on engineering skills in Databricks, Apache Spark, Python, SQL and Delta Lake, along with the ability to understand and translate existing Palantir Foundry implementations into Databricks-native solutions.
Responsibilities
Databricks
We are looking for a Databricks Engineer to join a data modernization program focused on building new data capabilities on Databricks and migrating existing data and application features from Palantir Foundry to the Databricks platform.
The candidate will work closely with data architects, application teams, business stakeholders, and AI/ML engineers to understand existing Palantir functionality, design the equivalent or improved capabilities in Databricks, and develop scalable, production-ready data pipelines and solutions.
The role requires strong hands-on engineering skills in Databricks, Apache Spark, Python, SQL and Delta Lake, along with the ability to understand and translate existing Palantir Foundry implementations into Databricks-native solutions.
Responsibilities
- Databricks Development
- Design and develop scalable data engineering solutions using Databricks.
- Build batch and near-real-time data ingestion and transformation pipelines.
- Develop data processing workflows using PySpark, Spark SQL and Python.
- Implement data models using Delta Lake and follow Medallion Architecture principles.
- Develop reusable and optimized data pipelines for large datasets.
- Implement data quality, validation, error handling and reconciliation processes.
- Optimize Spark jobs, SQL queries, Delta tables and cluster configurations for performance and cost.
- Palantir Foundry to Databricks Migration
- Analyze existing Palantir Foundry pipelines, datasets, transformations and business logic.
- Map existing Foundry capabilities to appropriate Databricks technologies and patterns.
- Re-engineer Foundry data pipelines and transformations using Databricks, PySpark, SQL and Delta Lake.
- Identify opportunities to simplify, modernize and improve existing functionality rather than performing a one-to-one migration.
- Perform source-to-target data mapping and migration validation.
- Compare outputs between Palantir and Databricks to ensure functional and data accuracy.
- Support migration of complex data transformations and large-scale data comparison processes.
- Work with architects and business teams to resolve gaps between the existing Foundry implementation and the target Databricks solution.
- Data Engineering & Platform Capabilities
- Implement data pipelines using Databricks Lakeflow / Spark Declarative Pipelines, where appropriate.
- Work with Unity Catalog for data governance, access control, lineage and discovery.
- Implement incremental processing, CDC and SCD patterns where required.
- Build data quality and observability capabilities.
- Implement CI/CD and deployment practices for Databricks workloads.
- Work with Databricks Asset Bundles and Git-based development practices.
- Develop and maintain Databricks Workflows/Jobs and production operational processes.
- Integration & Application Support
- Integrate Databricks with enterprise data sources, APIs, databases, cloud storage and downstream applications.
- Work with structured and unstructured data sources.
- Support data requirements for dashboards, analytics and downstream applications.
- Collaborate with AI/ML teams where data pipelines are required to support AI and GenAI use cases.
- Performance & Production Support
- Analyze and resolve production data pipeline issues.
- Perform root-cause analysis of data quality and processing failures.
- Optimize workloads for performance, reliability and cost.
- Establish appropriate monitoring, logging and alerting mechanisms.
- Participate in code reviews and enforce engineering standards.
Databricks
- Strong hands-on experience with Databricks
- Apache Spark / PySpark
- Spark SQL
- Delta Lake
- Databricks Workflows / Jobs
- Unity Catalog
- Lakeflow / Spark Declarative Pipelines
- Databricks Asset Bundles
- Performance tuning and optimization
- Strong Python
- Strong SQL
- Good understanding of software engineering principles and object-oriented programming
- ETL/ELT pipeline development
- Data modelling
- Medallion Architecture
- Batch and incremental processing
- CDC
- SCD Type 1/2
- Data quality and reconciliation
- Large-scale data processing
- Experience with Azure cloud platform
- Experience with cloud storage and cloud-based data platforms
More Info
Key Skills
Incremental processing
ETL ELT pipeline development
Medallion Architecture
Unity Catalog
Databricks Asset Bundles
Lakeflow Spark Declarative Pipelines
Large-scale data processing
Azure cloud platform
Databricks Workflows Jobs
Delta Lake
Data quality and reconciliation
SCD Type 1 2
