Search by job, company or skills

Site Reliability AI Engineer

Site Reliability AI Engineer

IntraEdge
7-9 Years
Not Disclosed
Early Applicant
  • Posted 2 months ago
  • Be among the first 10 applicants

Job Description

Observability/AIOps (5 to 8 yrs exp).

As an SRE with Observability focus you will:

Explore the complex IT estates of our clients to understand their observability/AIOps opportunities, identify the areas to improvise

Collaborate to architect unified observability and AIOps strategies which employ leading AI technology

Implement enterprise observability/AIOps technology and processes

Amplify observability/AIOps outcomes by accelerating adoption across technology and business organizations

Responsibilities include:

Architect observability solutions to address the gaps in order to reduce organizational MTTD and MTTR objectives.

Developing API-driven micro-services that combine into large and complex platforms

Planning and executing highly parallel distributed object storage transformations and migrations

Maintaining automated test suites using CI/CD tools

Participating in collaborative projects with small software engineering teams

Develop automation, processes, and tools designed to make our services simpler and more robust

Participate in troubleshooting, capacity planning and analysis, performance analysis activities

Advise management on service onboarding strategies and execution

Experience in architecting complex IT solutions

Understanding of observability dimensions(Metrics, logs, traces)

Excellent communication and stakeholder management skills

Development experience, comfortable working in multiple languages(Python, Java, Go and Ruby a plus)

Experience working in collaborative coding environments (peer review, continuous integration, etc)

7+ years of application development

Experience working in distributed remote teams across multiple time zones

Experience in large scale operations environments

7+ years of experience with Linux/Unix development or systems administration

3+ years of experience with networking systems and technologies

Deep understanding of network performance and security

Ability to identify tasks which require automation and implement required automation

Configuration Management tools experience with Puppet, Chef, SaltStack

Hands-on operational experience in a high-volume or critical production service environment - distributed systems, capacity planning, continuous deployment

Bachelors degree

We have opportunities to work with and learn:

Object Storage - Minio/S3/etc

Data Collection - OpenTelemetry/Grafana Alloy/etc

Message Bus - Kafka/NSQ/etc

Scaling Databases - Druid/Clickhouse/Cassandra/etc

Relational database technologies at large scale - Timescale/Vitess/Postgres/etc

Scheduling & Orchestration - Kubernetes/OpenShift/Docker

Cloud Platforms - AWS/Azure

Immediate- serving Notice preferred.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

NSQ

Vitess

Observability AIOps

Druid

OpenTelemetry

Clickhouse

CI CD tools

Networking systems

Minio

API-driven micro-services

Timescale

About Company

Similar Jobs

5-8 yrs
Hyderabad
Skills:
Ansible, Python, Kubernetes, site reliability, Slurm, GPU computing