

Search by job, company or skills

Responsibilities Include
- Design and maintain DBT models that produce trusted datasets, features, and metrics
the Data Science team relies on for analysis, experimentation, ML, and reporting.
- Build and operate pipelines in Databricks - PySpark jobs and Delta/Iceberg tables - that
turn raw operational events into analysis-ready data.
- Develop deep familiarity with operations so the datasets, schemas, and models
you ship reflect how the business actually works.
- Orchestrate end-to-end data workflows in Airflow (and Prefect where it fits), with SLAs
the DS team can count on for daily models, dashboards, and operational decisions.
- Participate in peer code reviews and raise the bar on data quality, testing, and
documentation across the team's models.
- Partner with data scientists to scope, design, and productionize feature pipelines and the
model-supporting data behind them.
- Optimize DBT and Spark workloads for cost, performance, and reliability as data
volume grows.
- Learning new technologies quickly - nobody comes into this role knowing every piece of
the stack. You'll lean on peers, documentation, and experimentation to grow.
What You BringBuild a strong business
- 5+ years of experience building, testing, and deploying data engineering systems.
- Experience with at least one distributed data system, and the ability to reason about
consistency, latency, throughput, and fault tolerance.
- Strong SQL and proficiency with at least one of (py)Spark, DBT, or Airflow in production.
- Experience with Infrastructure-as-Code systems such as Terraform, AWS CDK, or
Pulumi.
- Understanding or strong interest in supply chain and the data challenges it creates.
- A self-starter who takes initiative, moves fast, and ships while collaborating on big
challenges.
- Enthusiastic about working closely with team members to develop creative solutions to
novel problems.
- Excellent written and oral communication in English.
- Tech: DBT, Databricks, (py)Spark, Airflow, Prefect, SQL, Iceberg/Delta; familiarity
with the broader stack (Kinesis, EMR, Sigma, Pulumi) a plus.
- Comfortable using modern AI coding assistants (Claude Code, Cursor, Copilot, or
similar) and experienced with AI-native workflows - prompting, agentic tooling,
evaluations, retrieval - and willing to bring them into your data engineering practice.
- A plus: Demonstrated, measurable success building with LLMs, evaluating model
outputs, or integrating AI into data pipelines or internal tooling.
Job ID: 151513117
Skills:
S3, Data Transformation, Data Modelling, Api Development, PostgreSQL, SQL Server, Kafka, Schema Design, Redshift, Sql, Git, MySQL, MongoDB, Google Analytics, Python, AWS, user behaviour tracking, Looker, Snowplow, dbt, AWS Kinesis
Skills:
Numpy, Pandas, Amazon Redshift, Agile, Rest Apis, Sql, Python, ELT, AWS, Etl
Skills:
Python, Sql, Java, Spark, AWS, Docker, Kubernetes, Etl, Data Modeling, Machine Learning
Skills:
S3, BigQuery, Pyspark, Dataproc, Redshift, Sql, Gcp, DataFlow, Python, AWS, Airflow, MWAA, Cloud Composer, Glue, GCS, Athena
Skills:
Pyspark, Netezza, Data Modeling, Bi Reporting Tools, Cloud Storage, Data Visualization, automation, Python, BigQuery, SQL Server, Dataproc, Sql, Git, DataFlow, Etl, Airflow, CI CD, MPP databases, performance cost and runtime optimization, Pub Sub, Cloud Composer, GCP services, Cloud Functions, GCP certifications, Teradata, Architect, ELT pipeline design, Professional Data Engineer, GCP APIs