

Search by job, company or skills

Design & Build ingestion pipelines using ELT and schema on read in Databricks Delta
lake;
Design & Develop Transformations aspects using ELT framework to modernize the
Ingestion pipelines and build data transformations at scale;
Provide technical expertise in the areas of design and implementation of Ratings Data
Ingestion pipelines with modern AWS cloud and other technologies such as S3, Hive,
Databricks, Scala, Python, and Spark;
Build and maintain a data environment for speed, accuracy, consistency, and up time;
Work closely with other data teams & Data Science team and participate in the
development of ingestion pipelines;
Ensure data governance principles are adopted, data quality checks and data lineage are
implemented in each hop of the data;
Be in tune with emerging trends Big data and cloud technologies and participate in the
evaluation of new technologies;
Ensure compliance through adopting enterprise standards and promoting best
practices/guiding principles aligned with organization standards.
What you'll bring
A minimum of 5+ years of significant experience in application development;
Previous experience in the areas of design and implementation of Ratings Data Ingestion
pipelines with modern AWS cloud and other technologies such as S3, Hive, Databricks,
Scala, Python, and Spark;
Development, design, and architecture exposure & the ability to ensure quality across
various technology components that are developed by geographically diversified software
engineers with superior knowledge of system architecture, object-oriented design, and
design patterns;
Proficient with software development lifecycle (SDLC) methodologies like Agile and
test-driven development;
Job ID: 127017697
Skills:
Unix, Apache Airflow, Docker, Shell scripting, Python, AWS, Java, Apache Hadoop, Flume, Scala, HBase, Google Cloud, Big Data Technologies, Devops, Jenkins, Hive, Linux, Ansible, Spark, Apache Kafka, Azure, Kubernetes, Flink, Management Skills, database integration, ETL skills
Skills:
snowflake , Databricks, Sql, Hadoop Mapreduce, Apache Nifi, Java, Openshift, Docker Swarm, Apache Spark, AWS, Talend, Kubernetes, Python, Informatica, Azure Synapse, Azure, Amazon Kinesis, Docker, Scala, Google Cloud Platform, containerd, Podman, Apache Pulsar, Dask