Search by job, company or skills

Senior Site Reliability Engineer

8-13 Years
Early Applicant
Quick Apply
  • Posted 10 hours ago
  • Be among the first 30 applicants

Job Description

Key Job Responsibilities and Duties:

The core premise for the Booking SRE lies in treating operational and reliability problems of software systems as a software engineering problem. We code our way out of problems where operations are concerned addressing availability, scalability, latency, and efficiency challenges within the vast infrastructure here at Booking. We expect our SRE engineers to be software engineers that optimize systems rather than be system operators.

  • You will impact millions of people all over the globe with your creative solutions.
  • You work in one of the biggest e-commerce companies in the world.
  • You will solve exciting problems at scale by writing and deploying code across tens of thousands of servers.
  • Ensuring an everything as code mindset for yourself and your team.
  • You will have the opportunity to collaborate with many of the world's leading SREs.
  • You will be free to launch your own ideas and solutions within our sophisticated production environment.

Here are some of the tools and technologies we use to achieve this: Python, Go, Puppet, Kubernetes, Elasticsearch, Prometheus, HAProxy, Cassandra, Kafka, etc.

What you'll be doing:

  • Design, develop and implement software that improves the stability, scalability, availability and latency of the Booking.com products.
  • Take ownership of one or more services and have the freedom to do what is best for our business and customers.
  • Solve problems occurring with our highly available production systems and build solutions and automation to prevent them from happening again.
  • Build effective monitoring to supervise the health of your system, and jump in to handle outages.
  • Build and run capacity tests to manage the growth of your systems.
  • Plan for reliability by designing systems to work across our multinational data centers.
  • Develop tools to assist the product development teams with successfully deploying 1000s of change sets every day.
  • Be an advocate of engineering standard processes.
  • Share the on-call rotation and be an escalation contact for incidents.
  • Contribute to Booking.com's growth through interviewing, on-boarding, or other recruitment efforts.

What you'll bring:

  • 8+ years hands-on experience in software and site reliability engineering within the technology sector. Coupled with expertise with building, operating and maintaining sophisticated and scalable systems.
  • Solid experience in at least one programming language. We use Java, Python, Go, Ruby, Perl.
  • Experience with Infrastructure as Code technologies.
  • Knowledge of cloud computing fundamentals.
  • Solid foundation in Linux administration and troubleshooting.
  • Understanding of Service Level Agreements and Objectives.
  • Additional experience in OpenStack, Kubernetes, Networking, Security or Storage is desirable.
  • Supervising / observability technologies like Prometheus, Graphite, Grafana, Kibana, Elasticsearch are a plus.
  • Good interpersonal skills.
  • Proficient command of the English language, both written and spoken.

About Company

Booking Holdings is the world’s leading provider of online travel and related services, provided to consumers and local partners in more than 220 countries and territories through five primary consumer-facing brands: Booking.com, Priceline, Agoda, KAYAK and OpenTable, as well as through a network of subsidiary brands including Rocketmiles, Fareharbor, HotelsCombined, Cheapflights and Momondo, bringing greater value to both our customers and partners around the world. The mission of Booking Holdings is to make it easier for everyone to experience the world.

Job ID: 114282563

Similar Jobs

Bengaluru, India

Skills:

StormCassandraPrometheusKafkaTerraformDockerElasticsearchShell scriptingPostgresGitlabPythonAWSRustCloudformationRedisJenkinsCloudwatchGcpLinuxAnsibleSparkKubernetesGoFlinkArangoDBGitHub ActionsStackdriver

Bengaluru, India

Skills:

GithubElkPrometheusGrafanaDatadogShellTerraformDockerGitlabPythonAWSNew RelicJenkinsGitGcpPerlHelmAzureKubernetesGoHarnessGitHub ActionsLokiGitLab CIArgoCD

Bengaluru, India

Skills:

ApisPrometheusFluxGrafanaDatadogTerraformSplunkHelmKubernetesAI toolsLinux networking fundamentalsObservability toolsService architecturesArgoCD

Bengaluru

Skills:

GithubSamlPostgreSQLPythonAws

Bengaluru, India

Skills:

DatadogSqlKubernetesJenkinsPerl ScriptingGitLinuxWindows SysadminRelational Databases