Search Jobs

Search by job, company or skills

Fresher
Early Applicant
  • Posted 21 hours ago
  • Be among the first 10 applicants

Job Description

Summary:

We are seeking a highly skilled Software Engineer with a strong background in managing production incidents, debugging distributed systems, and implementing robust monitoring solutions. The ideal candidate will have extensive experience in programming and a keen interest in utilizing AI for enhancing system reliability.

Responsibilities:

  • Drive production incidents as the commander or primary responder with 24x7 on-call availability.
  • Debug distributed systems by correlating metrics, logs, and traces, and work comfortably with SLIs, SLOs, and error budgets.
  • Design and maintain Datadog dashboards, monitors, APM, tracing, and log pipelines.
  • Instrument services with Open Telemetry and design custom metrics, considering cardinality and cost.
  • Engage actively with AI engineering assistants to apply AI in reliability work, such as alert triage, log summarization, and RCA drafting.
  • Conduct capacity planning, load and performance testing, and ensure peak-event readiness.

Requirements:

  • Minimum of 8 years of experience in software engineering roles.
  • Experience in payments, fintech, or other high-availability regulated domains.

Required Skills:

  • Proficiency in Datadog for dashboards, monitors, APM, tracing, and log pipelines.
  • Experience with Open Telemetry for instrumenting services and designing custom metrics.
  • Strong programming skills in Java, Spring Boot, Node.js, and Python.

Preferred Skills:

  • Practical use of AI engineering assistants.


#AditiConsulting
# 26-05930

More Info

Key Skills

Similar Jobs

Bengaluru
Skills:
Continuous DeliverySqlSpring FrameworkSpring BootJavaContinuous IntegrationAgile MethodologiesKafkaReactAI-assisted software development tools
Bengaluru, India
Skills:
GolangGithubAgile MethodologiesArtifactoryGrafanaSoftware Development Life CycleTerraformSlackSystem DesignSonarqubeApplication DevelopmentPythonKubernetesTestingAWSInfrastructure as CodeEKSOperational stabilityGitHub ActionsTerraform CloudPublic Cloud services