

Search by job, company or skills

Work with product team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the mitigating critical customer incidents, or deployments or testing required to improve security, performance, availability, and scalability of service. Authority for end-to-end performance and operability. Partner with development teams in meeting SLA to unblock customers. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the effect of product architecture decisions on distributed systems. Professional curiosity and a desire to a develop deep understanding of services and technologies.
Responsible for the operation of production environments, including systems, supporting critical business operations. Will perform administration and analysis for multiple production environments and recommend new and novel solutions to improve availability, performance, and supportability. This is an opportunity to bring a combination of deep technical knowledge with administration/analysis knowledge of Oracle's Cloud Infrastructure to provide support to a wide range of complex production environment problems related to immense growth, scaling, leveraging the cloud, extremely high performance, and high availability requirements. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning.
We are looking for a Site Reliability Engineer to manage our cloud services. Responsibilities include:
Required Skills:
Career Level - IC3
Oracle Corporation is an American multinational computer technology corporation headquartered in Austin, Texas.In 2020, Oracle was the second-largest software company in the world by revenue and market capitalization.The company sells database software and technology (particularly its own brands), cloud engineered systems, and enterprise software products, such as enterprise resource planning (ERP) software, human capital management (HCM) software, customer relationship management (CRM) software (also known as customer experience), enterprise performance management (EPM) software, and supply chain management (SCM) software.
Job ID: 152068091
Skills:
AWS EKS, Golang, Elk, Shell Scripts, Grafana, Zabbix, Jenkins, Terraform, Ansible, Networking Protocols, Python, Monitoring stacks, GitLab CI, TICK, SIP, ArgoCD
Skills:
Java, Unix Administration, Monitoring Tools, Grafana, Terraform, Linux, Ansible, Python, Iscsi, Storage networking fundamentals, Go, cloud platforms, AI agents, deployment pipelines
Skills:
PostgreSQL, Prometheus, Kafka, Grafana, Terraform, Gitlab, Redis, New Relic, Load Balancers, Jenkins, Iam, Vault, Helm, Kubernetes, Cloud Logging, Cloud NAT, GitOps, Cloud DNS, Cloud SQL, Cloud Armor, Cloud Monitoring, Shared VPC, Private Service Connect, Artifact Registry, Secret Manager, ArgoCD
Skills:
Distributed Systems, Networking, Prometheus, Bash, Grafana, Devops, Linux, Kubernetes, Python, AWS, SRE, Production Engineering, Go, Containerized Environments, Observability, Platform Engineering, Monitoring
Skills:
AWS, S3, Vpc, Linux, Load Balancers, Datadog, Kubernetes, Bash, Python, Docker, Terraform, Vpn, Jenkins, Firewalls, GitHub Actions, Istio, subnets, GitLab CI, FluxCD, Go, Crossplane, EBS