Role Summary
We are seeking a skilled Data Engineer to join our team. The candidate will be responsible for designing, building, and maintaining robust data infrastructure that powers PocketFM's recommendation systems, analytics, and business intelligence capabilities. This role offers an exciting opportunity to work with large-scale data systems that directly impact millions of users audio entertainment experience.
Key Responsibilities
Data Infrastructure & Pipeline Development
- Design, develop, and maintain scalable ETL/ELT pipelines to process large volumes of user interaction data, content metadata, and streaming analytics
- Build and optimize data warehouses and data lakes to support both real-time and batch processing requirements
- Implement data quality monitoring and validation frameworks to ensure data accuracy and reliability
- Develop automated data ingestion systems from various sources including mobile apps, web platforms, and third-party integrations
Analytics & Reporting Infrastructure
- Create and maintain data models that support business intelligence, user analytics, and content performance metrics
- Build self-service analytics platforms enabling stakeholders to access insights independently
- Implement real-time dashboards and alerting systems for key business metrics
- Support A/B testing frameworks and experimental data analysis requirements
Data Architecture & Optimization
- Collaborate with software engineers to optimize database performance and query efficiency
- Design data storage solutions that balance cost, performance, and accessibility requirements
- Implement data governance practices including data cataloging, lineage tracking, and access controls
- Ensure GDPR and data privacy compliance across all data systems
Collaboration & Support
- Work closely with data scientists, product managers, and analysts to understand data requirements
- Participate in code reviews and maintain high standards of code quality and documentation
- Mentor junior team members and contribute to knowledge sharing initiatives
Required Qualifications
Technical Skills
- Programming Languages: Proficiency in Python, SQL, and at least one of: Java, Scala, or Go
- Big Data Technologies: Hands-on experience with Apache Spark, Kafka, Airflow, and distributed computing frameworks
- Cloud Platforms: Strong experience with AWS, GCP, or Azure data services (S3, BigQuery, Redshift, etc.)
- Database Systems: Expertise in both SQL (PostgreSQL, MySQL) and NoSQL (MongoDB, Cassandra, Redis) databases
- Data Warehousing: Experience with modern data warehouse solutions like Snowflake, BigQuery, or Databricks
- Containerization: Proficiency with Docker and Kubernetes for deploying data applications
Experience Requirements
- 2-4 years of experience in data engineering or related roles
- Proven track record of building and maintaining production data pipelines at scale
- Experience with streaming data processing and real-time analytics systems
- Strong understanding of data modeling, schema design, and data architecture principles
- Experience with version control systems (Git) and CI/CD pipelines
Preferred Qualifications (Good to Have)
- Machine Learning & Model Operations
- Model Deployment: Experience deploying machine learning models to production environments using frameworks like MLflow, Kubeflow, or SageMaker
- MLOps Practices: Familiarity with ML pipeline automation, model versioning, and continuous integration for machine learning
Advanced Technical Skills
- Experience with Vector Database, graph databases and knowledge graphs
- Understanding of data mesh architecture and domain-driven data design
- Experience with data privacy and security implementations