Senior Data Engineer
Recro- Posted 17 hours ago
- Be among the first 10 applicants
Job Description
Data Engineer JD
Location - BLR (5 days WFO)
EXP - 5+ years
Job Role:
Develop long-term vision for a highly scalable data platform, data management and Data Ops practices.
Design and architect data flows, data management in Hadoop or Cloud environment which are scalable, repeatable and eliminate time consuming steps Promote Data Ops approach to automate the provision of data, testing and monitoring and to shorten development cycles and increase deployment frequency.
Establish development and data governance processes to build mature data pipelines, CI/CD, test coverages, etc.
Evaluate, provide insights and recommendations on tools and technology strategy for analytics data platforms and applications in conjunction with Enterprise Architecture team
Ability to lead data engineering workstreams with a product mindset
Who are we looking for
Bachelor's or master's degree in computer science, Information Systems or equivalent field.
At least 5+ years of experience in building data flows and data management on modern big data tech stack Data Strategy: Understands, articulates, and applies principles of the defined strategy to routine business problems that involve a single function.
Data Transformation and Integration: Extracts data from identified databases. Creates data pipelines and transform data to a structure that is relevant to the problem by selecting appropriate techniques.
Develops knowledge of current analytics trends.
Data Source Identification: Supports the understanding of the priority order of requirements and service level agreements.
Helps identify the most suitable source for data that is fit for purpose.
Demonstrates expertise in writing complex, highly optimized queries across large data sets Strong experience in using ETL framework (eg. Airflow, Oozie, Jenkins etc.) to build and deploy production-quality ETL pipelines.
Experience in ingesting and transforming structured and unstructured data from internal and third-party sources into dimensional models.
Knowledge of data structures and distributed computing.
Should be comfortable in manipulation and analysis of high-volume data from variety of internal and third-party sources.
Experience in one or more programming languages like Python or PySpark and moderate knowledge on unix scripting.
More Info
Key Skills
Cloud environment
Data Ops
CI/CD
Data pipelines
Data transformation and integration





