About Albertsons Companies Inc. (ACI):
As a leading food and drug retailer in the United States, Albertsons Companies, Inc. operates over 2,200 stores across 35 states and the District of Columbia. Our well-known banners across the United States, including Albertsons, Safeway, Vons, Jewel-Osco and others, serve more than 36 million U.S customers each week.
We build and shape technology solutions that solve customers problems every day, making things easier for them when they shop with us online or in a store. We have made bold, strategic moves to migrate and modernize our core foundational capabilities, positioning ourselves as the first fully cloud-based grocery tech company in the industry.
Our success is built on a one-team approach, driven by the desire to understand and enhance the customer experience. By constantly pushing the boundaries of retail, we are transforming shopping into an experience that is easy, efficient, fun and engaging.
About Albertsons India Capability Center:
At Albertsons India Capability Center, we're not just pushing the boundaries of technology and retail innovation, we're cultivating a space where ideas flourish and careers thrive. Our workplace in India is a vital extension of the Albertsons Companies Inc. workforce and important to the next phase in the company's technology journey to support millions of customers lives every day.
At the Albertsons India Capability Center, we are raising the bar to grow across Technology & Engineering, AI, Digital and other company functions, and transform a 165-year-old American retailer. At Albertsons India Capability Center, associates collaborate directly with international teams, enhancing decision-making processes and organizational agility through exciting and pivotal projects. Your work will make history and help millions of lives each day come together around the joys of food and inspire their well-being.
We are searching for someone with the following skills:
- 10+ years of industry experience in software engineering, machine learning engineering, AI systems development, or production data systems.
- Strong experience designing, building, deploying, and operating AI/ML systems for real-world production use cases at enterprise scale.
- Deep hands-on experience with Python and common ML frameworks and libraries such as scikit-learn, XGBoost, PyTorch, TensorFlow, Pandas, NumPy, and related production ML tooling.
- Proven experience building machine learning solutions for forecasting, anomaly detection, prediction, classification, clustering, ranking, recommendation, prioritization, or decision-support use cases.
- Strong understanding of time-series modeling techniques for forecasting, anomaly detection, capacity planning, operational prediction, and service health intelligence.
- Experience building or leading ML solutions for alert classification, incident prediction, event deduplication, signal correlation, prioritization, noise reduction, or operational intelligence.
- Strong knowledge of causal ML, causal inference, graph-based reasoning, dependency-aware analysis, and statistical modeling techniques for RCA and impact analysis.
- Hands-on experience building LLM-powered applications using frameworks such as LangChain, LangGraph, or similar agentic AI frameworks.
- Experience designing AI agents or multi-agent systems for reasoning, summarization, task orchestration, troubleshooting, knowledge retrieval, decision support, or assistant workflows.
- Strong understanding of prompt engineering, RAG architectures, embeddings, vector databases, tool calling, memory handling, context management, guardrails, and agent evaluation techniques.
- Experience integrating LLM systems with enterprise tools, APIs, knowledge repositories, observability platforms, ticketing systems, automation tools, and operational workflows.
- Strong experience building backend services, APIs, microservices, model-serving systems, inference services, and agent orchestration platforms.
- Good understanding of observability data, including logs, metrics, traces, topology, alerts, incidents, service maps, deployment events, and operational metadata.
- Strong data engineering knowledge, including feature engineering, data preprocessing, model pipelines, batch inference, streaming inference, data quality validation, and production data workflows.
- Familiarity with graph databases such as Neo4j and their use in dependency mapping, knowledge graphs, causal analysis, impact analysis, and knowledge-driven AI systems.
- Strong experience with REST APIs, microservices architecture, Docker, Kubernetes, cloud-native deployment patterns, distributed systems, and scalable architecture design.
- Experience with CI/CD, MLOps, model lifecycle management, experiment tracking, model versioning, feature stores, model monitoring, drift detection, and automated deployment practices.
- Knowledge of OpenTelemetry, monitoring systems, observability platforms, incident management systems, and SRE operating models is highly desirable.
- Strong understanding of software engineering fundamentals, system design, distributed systems, reliability engineering, testing strategies, and scalable architecture patterns.
- Proven ability to decompose ambiguous operational problems into measurable AI/ML solutions with clear success metrics and production impact. Excellent communication, collaboration, and technical leadership skills,
- with the ability to influence SREs, platform engineers, architects, product owners, and business stakeholders. Self-driven mindset with strong curiosity, innovation, ownership, and the ability to evaluate and apply
- emerging AI techniques effectively in production environments.
We believe the successful candidate has these qualifications and experience:
- Bachelor's degree in Computer Science, Engineering, Data Science, Artificial Intelligence, Information Systems, or a related field, or equivalent practical experience.
- 10+ years of overall industry experience in software engineering, machine learning engineering, AI system development, data platforms, or production-grade distributed systems.
- 6+ years of hands-on experience building, deploying, and operating machine learning systems in production environments.
- Experience designing or building AI/ML systems for observability, monitoring, AIOps, SRE, IT operations, incident management, or related operational domains would be a big plus.
- Strong hands-on experience in Python-based AI/ML development is required.
- Proven experience owning production ML systems end to end, including data pipelines, training workflows, evaluation, deployment, inference, monitoring, retraining, and operational support.
- Experience building LLM-powered applications, AI agents, or multi-agent workflows for enterprise use cases is highly preferred.
- Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health prediction, capacity risk prediction, alert intelligence, or remediation recommendations.
- Familiarity with knowledge graphs, graph databases, and graph-based ML techniques for dependency-aware intelligence and operational reasoning.
- Experience using vector databases and retrieval frameworks for enterprise search, RAG systems, and agentic AI applications. Experience integrating AI services with platforms such as ServiceNow, Grafana, Prometheus, Splunk, AppDynamics, Datadog, Dynatrace, PagerDuty, Jira, or similar enterprise tools.
- Familiarity with MCP-based client or agent integrations is a plus.