Job description
PRINCIPLE RESPONSIBILITIES:
1. AI/ML Platform Development
- Implement and maintain components of the centralized enterprise AI/ML platform, including model serving, inference APIs, and shared ML libraries.
- Build and operate robust, production-grade ML pipelines covering data ingestion, feature engineering, model training, deployment, and inference.
- Contribute to the development of shared feature stores, model registries, and inference services used by multiple product teams.
2. MLOps & Model Lifecycle
- Develop and maintain CI/CD pipelines for ML model training, evaluation, and deployment.
- Implement automated model monitoring, drift detection, and alerting mechanisms.
- Build tooling to support model versioning, lineage tracking, and rollback capabilities.
3. Performance & Reliability Engineering
- Optimize inference pipelines for latency, throughput, and resource efficiency.
- Contribute to scalability improvements to handle growing model counts and inference request volumes.
- Investigate and resolve production issues in the AI/ML platform.
4. Reusability & Inner-Sourcing
- Develop well-documented, reusable ML components and libraries for consumption by product
engineering teams.
- Actively participate in and contribute to the inner-source AI/ML asset ecosystem.
- Write technical documentation, runbooks, and integration guides for shared services.
5. Collaboration & Engineering Excellence
- Work closely with data scientists, ML engineers, and product architects to translate model requirements
into production implementations.
- Participate in architectural and code reviews, upholding high standards for quality, security, and
maintainability.
- Mentor junior engineers on AI/ML engineering best practices.
REQUIRED KNOWLEDGE, SKILLS AND EXPERIENCE:
Experience and Expertise
- 5+ years of software engineering experience, with at least 3 years focused on building and operating AI/ML systems in production.
- Demonstrated experience building and maintaining ML pipelines and model serving infrastructure at scale.
- Hands-on experience with MLOps practices including CI/CD for ML, model monitoring, and lifecycle management.
- Proven ability to build reliable, maintainable software components used by multiple teams.
Technical Skills
- Strong proficiency in Python; familiarity with Go or Java is a plus.
- Hands-on experience with major AI/ML frameworks: TensorFlow and/or PyTorch.
- Practical knowledge of MLOps tools: MLflow, Kubeflow, or equivalent (Vertex AI, SageMaker, Azure ML).
- Solid experience with containerization (Docker) and orchestration (Kubernetes) for ML workloads.
- Experience with data processing technologies (Spark, Kafka) and data storage (data lakes, feature stores, data warehouses).
- Familiarity with at least one public cloud AI/ML ecosystem (AWS, GCP, or Azure).
- Experience designing and consuming RESTful or gRPC APIs for inference services.
Analytical and Problem-Solving Skills
- Strong debugging and root-cause analysis skills for distributed AI/ML systems.
- Ability to profile and optimize ML inference for latency and cost.
Collaboration & Communication
- Clear written and verbal communication; ability to document technical systems for a broad audience.
- Collaborative working style with data scientists, engineers, and architects.
Educational Background
- Bachelor's degree in Computer Science, Artificial Intelligence, Machine Learning, or related technical field.
Additional Skills
- Experience in the telecom domain or with OSS (Operations Support Systems) data and use cases.
- Familiarity with GPU/TPU-based inference optimization.
- Knowledge of responsible AI practices (fairness, explainability, privacy).
- Contributions to open-source MLOps or AI infrastructure projects.
- Experience with streaming/real-time inference architectures.
RAKUTEN SHUGI PRINCIPLES:
Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
- Always improve, always advance. Only be satisfied with complete success - Kaizen.
- Be passionately professional. Take an uncompromising approach to your work and be determined to be the best.
- Hypothesize - Practice - Validate - Shikumika. Use the Rakuten Cycle to success in unknown territory.
- Maximize Customer Satisfaction. The greatest satisfaction for workers in a service industry is to see their customers smile.
- Speed!! Speed!! Speed!! Always be conscious of time. Take charge, set clear goals, and engage your team.