Senior Data Engineer – Data Pipelines & Cloud Databases
Location: Pune | Experience: 4+ years | Type: Full-Time
Role Overview
Design, build, manage enterprise data pipelines on Azure and Databricks. Own schema design, API development, and data infrastructure for analytics and intelligence products.
Key Responsibilities
Data Pipeline Development
- Implement robust, scalable data pipelines using Microsoft Azure and Databricks stack
- Build reusable data pipeline components and frameworks
- Design and optimize data workflows for performance and reliability
- Monitor pipeline health and implement automated alerting
Database Architecture
- Design relational and non-relational database schemas
- MongoDB: schema design, aggregation pipelines, indexing, sharding, replica sets, performance tuning
- Azure Cosmos DB: multi-model access, partitioning, consistency levels, throughput management (RU/s), global distribution
- Azure Cosmos DB Gremlin API: graph data modeling, traversals, vertex/edge design, relationship analytics
API Development
- Build RESTful APIs using FastAPI with async endpoint design
- Implement Pydantic models, middleware, dependency injection, background tasks
- Generate and maintain auto-generated OpenAPI/Swagger documentation
- Configure Azure API Management (APIM) for API gateway and security
Project & Stakeholder Management
- Collect progress updates from squads regularly
- Consolidate updates into weekly project status reports
- Develop L3-level project plans with task details, milestones, dependencies
- Participate in early-stage design and feature definition
- Communicate complex data insights to non-technical stakeholders
Collaboration & Integration
- Work across multiple engineering teams on prototype integration
- Support integration of proven prototypes into core intelligence products
- Strong team collaboration and cross-functional communication
- Knowledge sharing and documentation
Required Experience:
Data Engineering & Databases
- 4+ years data engineering or data platform experience
- Relational and non-relational database expertise
- MongoDB: advanced schema design, aggregation, indexing, sharding, performance optimization
- Azure Cosmos DB: multi-model APIs, partitioning, consistency, throughput management
- Graph databases: Cosmos DB Gremlin API, graph modeling, traversals, relationship analytics
API & Backend Development
- FastAPI proficiency: async endpoints, Pydantic models, middleware, dependency injection
- Background tasks and job scheduling
- RESTful API design and best practices
- OpenAPI/Swagger documentation
Cloud Platform
- Azure fundamentals and hands-on experience
- Azure Data Factory or Databricks for ETL/ELT
- Azure Cosmos DB multi-region setup
- Azure API Management (APIM) configuration
Data Pipeline Skills
- ETL/ELT pipeline design and development
- Data quality validation and monitoring
- Schema design for analytics and reporting
- Performance optimization and scalability
Preferred Experience
- Databricks Delta Lake experience
- Azure Synapse Analytics
- Python for data engineering
- Spark SQL optimization
- Real-time data streaming
- Data governance and metadata management
- Agile/Scrum development model
- CI/CD pipeline experience