Search by job, company or skills

Senior Data Engineer

Early Applicant
  • Posted 12 hours ago
  • Be among the first 10 applicants

Job Description

AI/ML, No Code and Lo Code are the fastest-growing engineering concepts in the software industry. These systems allow customers to get the information they need to make critical strategic decisions in a fast-paced rapidly changing business world.

For this position, we are looking for a Software Developer with experience in distributed systems. You will get an opportunity to build and operate a suite of massive scale, integrated engineering platforms in a distributed, multi-tenant cloud environment.

Job Qualifications:

  • 5+ years of coding experience in Java or Scala

  • 5+ years of experience in building micro services/APIs using Spring Boot

  • 5+ years of distributed computing experience like Spark, MapReduce, or Hive

  • 5+ year experience with building high-performance, resilient, scalable, and well-engineered systems

  • Good understanding of full-stack software, object-oriented design, distributed systems, multi-threading and databases

  • Experience in CI/CD and development best practices

  • Experience with cloud data storage technologies likeHDFS, S3 etc preferred

  • Experience with instrumentation, logging systems

  • Preferred Multi cloud networking experience

  • B.E in Computer Science, Computer Engineering

We hire people with a broad set of technical skills and ready to take on some of technology's greatest challenges to make an impact on millions of customers.

We value engineering bestpractices, partnership,and empathy. We measure our success in the accomplishment and enablement of others.

Key Responsibilities

Data Processing & Pipelining - Data Requirements, Collection, and Infrastructure:

  • Identifies data requirements and business objectives of a project or initiative in collaboration with cross-functional teams.

  • Designs and builds proper data pipelines required for optimal data processing from a variety of data sources.

  • Independently analyzes, designs, and troubleshoots data flows based on business needs.

  • Analyzes business requirements and translates that into a technical specification.

  • Adjusts data collection processes that involve indexing and query optimizations, for optimal performance.

  • Builds generic data pipelines to support efficient data collection and extraction, independently.

  • Analyzes data sources for profiling the data to ensure pipeline build success.

  • Defines success and failure thresholds for data collection pipelines.

Data Processing & Pipelining - Data Governance:

  • Implements data governance policies and procedures for data handling (e.g., data retention) to manage data consistency, integrity, accuracy, and reliability throughout the data lifecycle.

  • Redacts Personally Identifiable Information (PII) and Protected Health Information (PHI) data to ensure compliance with data privacy and security standards.

  • Follows data security measures to protect data from unauthorized access, use, disclosure, alteration, or destruction.

  • Ensures data compliance with relevant laws, regulation, and industry standards.

Data Processing & Pipelining - Data Validation & Quality Assurance:

  • Implements rigorous data validation and integrity checks, identifying and addressing any data quality issues that could impact data pipeline and model performance.

  • Defines data annotation and labeling processes to ensure data quality, independently.

  • Designs and implements automation of data validation and governance.

Data Pipeline and Solutions Engineering - Pipeline Design:

  • Independently designs, develops, and optimizes automated and scalable data pipeline architectures to build reusable data products.

  • Implements appropriate data storage solutions to store the processed data to be used in a scalable and optimized way for access and analysis.

  • Manages the flow of data pipeline and storage day-to-day operations.

Data Pipeline and Solutions Engineering - Data Solutions Engineering:

  • Works independently and collaboratively in an agile development environment with other engineers to develop, maintain, and debug data solutions that are scalable, efficient, cost effective, and reliable.

  • Writes runnable code and performs testing and debugging of data solutions, independently.

Core Responsibilities

Planning & Execution:

  • Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements.

  • Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.

Collaboration & Partnership:

  • Collaborates across teams to align on expectations and achieve shared objectives.

  • Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships.

  • Actively listens to diverse perspectives and asks questions to ensure understanding of others.

Problem Solving:

  • Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate.

  • Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors.

  • Contributes to knowledge sharing and best practices.

Continuous Learning:

  • Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices.

  • Seeks out and leverages feedback and training to improve skills.

  • Contributes to a culture of continuous learning and knowledge sharing with team members.

Continuous Improvement:

  • Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team.

  • Seeks input from team members on alternative approaches and methods for improving work.

Career Level - IC3

More Info

About Company

Oracle Corporation is an American multinational computer technology corporation headquartered in Austin, Texas.In 2020, Oracle was the second-largest software company in the world by revenue and market capitalization.The company sells database software and technology (particularly its own brands), cloud engineered systems, and enterprise software products, such as enterprise resource planning (ERP) software, human capital management (HCM) software, customer relationship management (CRM) software (also known as customer experience), enterprise performance management (EPM) software, and supply chain management (SCM) software.

Job ID: 151664653

Similar Jobs

Bengaluru, India

Skills:

Spark SQLGitAdfPysparkSqlDelta LakeAutoloaderUnity Catalog

Bengaluru, India

Skills:

snowflake JavaUnixCPower BiOBIEEInformatica EtlPythonAWSAirflowOASCI CD pipelinesOracle PL-SQL

Bengaluru, India

Skills:

data engineering CassandraBig DataKafkaData WarehouseSqlCloudSparkPostgresPythonAirflowFlink

Bengaluru, India

Skills:

snowflake S3CassandraPysparkPostgreSQLApache AirflowLambdaMySQLShell scriptingPythonAWSApache FlinkDynamodbEmrRedshiftGitHiveIamApache KafkaSparkMongoDBAdvanced SqlApache IcebergGoogle BigQueryParquetStep FunctionsGlueAthena

Bengaluru, India

Skills:

Spark SQLSparkData WarehousingPythonSqlELTEtlStructured DataSemi-Structured DataRelational DatabasesBatch Processing