
Search by job, company or skills
Job Description:
We are looking for an experienced AIOps Engineer with expertise in Google Cloud Platform (GCP), AIOps/MLOps/LLMOps, intelligent observability and monitoring, CI/CD, Python, Pub/Sub/Kafka, distributed computing, GitHub, data pipelines, and GCP services such as Vertex AI, BigQuery, Dataflow, GKE, and Cloud Operations Suite (Cloud Monitoring, Cloud Logging, Cloud Trace). This role will involve designing, implementing, and optimizing AI-driven IT operations solutions - including anomaly detection, event correlation, predictive incident management, and automated remediation - and ensuring smooth, reliable operation of platforms and ML workloads in production environments on GCP.
Responsibilities:
• Design, develop, and maintain AIOps pipelines on GCP for ingesting, processing, and analyzing telemetry data (logs, metrics, traces, events) to enable intelligent monitoring and observability.
• Build and deploy ML models for anomaly detection, event correlation, root-cause analysis, and predictive incident management using Vertex AI and BigQuery ML.
• Collaborate with data scientists, SREs, software engineers, and DevOps teams to embed AI-driven automation into IT operations using best practices in AIOps and MLOps.
• Automate end-to-end ML and operations workflows, including data preprocessing, model training, evaluation, deployment, and automated remediation, using tools like Vertex AI Pipelines, Kubeflow, or Cloud Composer (Apache Airflow).
• Implement CI/CD pipelines using Cloud Build, GitHub Actions, and Terraform (IaC) for automated deployment, testing, and monitoring of models and services on GCP.
• Utilize Pub/Sub, Dataflow, and Kafka for real-time telemetry ingestion, streaming analytics, and event-driven automation.
• Optimize GCP infrastructure (GKE, Compute Engine, BigQuery, Cloud Storage) for scalability, performance, reliability, and cost efficiency.
• Manage GitHub repositories for version control and collaboration on AIOps and machine learning projects.
• Integrate with ITSM/ITOM platforms (e.g., ServiceNow) and various data connectors to access and process operational data from different sources.
• Develop and maintain documentation for AIOps pipelines, GCP infrastructure, runbooks, and processes.
• Stay up to date on emerging technologies and best practices in AIOps, SRE, machine learning operations, and cloud engineering.
Qualifications:
• 5+ Years of prior experience in DevOps/SRE/Data Engineering, with strong exposure to AIOps and MLOps.
• 3+ Years of hands-on experience with GCP services (Vertex AI, BigQuery, Dataflow, Pub/Sub, GKE, Cloud Operations Suite) in production environments.
• Strong proficiency in Python programming language.
• Experience with observability and monitoring tools (Cloud Monitoring, Prometheus, Grafana, ELK/Splunk, Datadog, or similar) and distributed computing frameworks.
• Hands-on experience with CI/CD pipelines, Terraform/IaC, Kubernetes (GKE), and automation tools.
• Exposure in deploying a use case in production leveraging Generative AI involving prompt engineering and RAG Framework (e.g., LLM-assisted incident summarization or intelligent ticket triage).
• Familiarity with Pub/Sub, Kafka, or similar messaging systems.
• Strong problem-solving skills and the ability to iterate and experiment to optimize AI model behavior and operational automation.
• Excellent problem-solving skills and attention to detail.
• Ability to communicate effectively with diverse clients/stakeholders.
Education Background:
• Bachelor's or master's degree in computer science, Engineering, or a related field.
• Tier I/II candidates preferred.
Folks with shorter notice period to be preferred.
Job ID: 151780619