
Search by job, company or skills
SRE Engineer is responsible for ensuring website uptime, optimizing performance, and maintaining security of the production application. This role involves monitoring site reliability, addressing technical issues, automating maintenance tasks, and collaborating with cross-functional teams to meet business objectives.
Responsibilities
Run the production environment by monitoring availability and taking a holistic view of system health
Build software and systems to manage platform infrastructure and applications
Improve reliability, quality, and time-to-market of our suite of software solutions
Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating for continual improvement
Provide primary operational support and engineering for multiple large-scale distributed software applications
Gather and analyze metrics from operating systems as well as applications to assist in performance tuning and fault finding
Partner with development teams to improve services through rigorous testing and release procedures
Participate in system design consulting, platform management, and capacity planning
Create sustainable systems and services through automation and uplifts
Balance feature development speed and reliability with well-defined service-level objectives
Required skills and qualifications
Bachelor's degree (or equivalent) in computer science or related discipline
Experience in SRE, DevOps, or similar roles.
Expertise in monitoring tools, infrastructure management, and automation.
Strong problem-solving skills and a collaborative mindset.
Proactive approach to identifying problems, performance bottlenecks, and areas for improvement
Job ID: 103729535
Skills:
Yaml, Bash, Json, Gcp, ECS, Azure, Kubernetes, Python, AWS, PingAM, GCE, server-less architectures, Fargate, ForgeRock, PingGateway, PingDS, PingIDM
Skills:
Python Scripting, Performance Tuning, Linux System Administration, Kubernetes cluster design, Docker OCI containers
Skills:
Unix, Elk, Prometheus, Bash, Grafana, Datadog, Gcp, Linux, Docker, Terraform, Ansible, Helm, Azure, Kubernetes, Python, AWS, GitOps, GitLab CI, GitHub Actions, OpenTelemetry, oci, ArgoCD
Skills:
Distributed Systems, Haskell, Polyglot Database Management, anomaly detection, Event Sourcing, Service Mesh Architecture, Real-time Data Streaming, Event-mesh Architecture, Chaos Testing, Client-side Distributed Data Caching
Skills:
PostgreSQL, Itil, Azure, Oracle, Kubernetes, Datadog, AWS, PagerDuty