

Search by job, company or skills

A leading financial institution is seeking a Senior Data Engineer to build scalable lakehouse platforms, governed data products and AI‑ready data assets in close collaboration with architecture, governance, analytics and business stakeholders.
You will deliver batch & real‑time data pipelines, enforce data quality and observability, and partner with AI/data science teams to power NLP, RAG and agent‑driven analytics for banking risk and fraud use‑cases.
Responsibilities:
• Design metadata‑driven ingestion frameworks for structured and unstructured data, covering batch, CDC, streaming, API and file sources
• Develop reusable Spark / Java / Python components; build Bronze‑Silver‑Gold lakehouse layers and implement data modelling, lineage, SLA‑enabled data products
• Tune distributed workloads, implement data quality controls and production observability
• Apply DevOps, CI/CD and IaC on Docker / Kubernetes; operationalise ML models alongside data scientists
• Architect data foundations supporting RAG, agent‑AI and multi‑modal unstructured‑data workflows
• Build internal engineering tooling and align data implementations with financial regulatory requirements
• Troubleshoot production workloads and drive platform reliability improvements
Requirements:
• 10+ years hands‑on data‑engineering delivery experience, banking / financial services background preferred
• Strong skills in PySpark, Java, Python, advanced SQL; solid knowledge of data modelling, lakehouse architecture, ETL/ELT and CDC
• Practical experience building ingestion pipelines and data products; working familiarity supporting RAG / NLP / agent‑AI workloads
• Hands‑on with Docker, Kubernetes, CI/CD and observability tooling
• Able to translate business & risk requirements into data‑platform designs; good cross‑team communication
• Bachelor / Master's in Computer Science or quantitative‑related discipline
Desirable:
• Airflow, Terraform and cloud‑native data‑platform experience
• Experience with Iceberg, Kafka, Flink, NiFi or Cloudera Machine Learning
Please contact Xuemin at [Confidential Information] for a discussion.
EA License no.: 16S8066 Rep no.: R25157345
Only successful applicants will be notified.
Job ID: 153578899
Skills:
snowflake , BigQuery, Postgres, FastAPI, Python, Sql, Airflow, Dagster, Anthropic APIs, Prefect
Skills:
Hadoop, Apache Spark, Angular, React, Shell scripting, Data Visualization, Scrum, Python, data pipelines, MSSQL databases, APIs within microservice architectures, Agile environments, AWS cloud-native services, data models, ETL processes, AWS SageMaker, data frameworks, Java Spring Boot, Kanban
Skills:
Spark SQL, S3, Power Bi, Cloudformation, AWS Glue, Tableau, Emr, Redshift, Sql, Nosql, Lambda, Kinesis, Terraform, Databricks, Python, Informatica Data Management Cloud, Databricks Delta Lake, MLflow, IDMC, R, Athena
Skills:
Java, Ranger, Hadoop, Scala, Prometheus, Bitbucket, Grafana, OpenShift Container Platform, Jenkins, Hive, Docker, Ansible, Spark, Splunk, Python, Kubernetes, HDFS, Quantexa
Skills:
workflow management , Data Modelling, Pyspark, Apache Spark, Data Warehousing, Automated Testing, Google Cloud, Sql, Azure, Python, AWS, data quality practices, copilots, Monitoring, modern data architecture, AI-enabled data engineering practices, AI agents, data orchestration, Validation, observability