
Senior Site Reliability Engineer
Oracle · Posted Oct 8
Cloud infrastructure, enterprise software, database systems, and AI-powered business applications
Get a personal compatibility score
Add a resume for personal matches
About the role
Oracle Cloud Infrastructure's Global Network Operations Center (GNOC) team is responsible for ensuring maximum possible service availability and performance across OCI's networks while delivering premier customer service. The team centrally manages OCI's networks and builds automated solutions to address common operational challenges across large-scale enterprise and cloud infrastructure. This Senior Site Reliability Engineer role covers fault and incident handling, root cause analysis, network changes, capacity and service lifecycle management, and scripting automation, with participation in operational rotations and on-call support.
What you will do
- Fault handling on incident tickets. provide break-fix support and escalation for event remediation, including leading root cause analysis (RCA) efforts
- Work closely with shift lead to ensure tickets are handled in a timely manner
- Using existing procedures and tooling, develop and safely complete network changes
- Mentor junior engineers as needed
- Participate in operational rotations providing break-fix support
- Possess analytical skills in resolving network issues with advanced troubleshooting and coordination with onsite support teams and vendors
- Identifying actionable incidents using a monitoring system, strong analytical problem-solving skills to mitigate network events/incidents, and following up on routine root cause analysis (RCA), coordinating with support teams and vendors
- Join major events/incidents calls, use technical and analytical skills to resolve network issues that impact Oracle customer/service, coordinate with SMEs, and provide RCA document
- Fault handling and escalation (identifying and responding to faults on OCI's systems and networks, collaborating closely with 3rd party suppliers, handling escalation through to resolution)
- Experience with Incident Response plans and strategies in cloud computing environments
- Have worked with large enterprise network infrastructure and cloud computing environments, supporting 24/7 and willing to work in rotational shifts in a network operations role
- Provide on-call support services as needed, job duties are varied and complex, needing independent judgment
- The role includes collaborating with networking automation services to integrate support tooling and frequently developing scripts to automate routine tasks
- You will use automation to complete work and develop scripts for routine tasks
- Takes proactive steps to design and architect infrastructure and/or service according to terms for reliability and functionality.
- Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads.
- Collaborates with the software development team to develop infrastructures and features that are reliable and scalable according to deployment requirements.
- Independently identifies opportunities for and drives prototyping (e.g., testing new applications or infrastructures, assisting in onboarding).
- Performs data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.
- Independently monitors services, maintains up-to-date knowledge of their performance, and documents their condition.
- Leverages comprehensive knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
- Provides health and performance reporting and takes appropriate actions based on trends in data.
- May independently perform provisioning to support infrastructure, applications, and services.
- May perform standard and non-standard decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
- Identifies opportunities for automation and assesses potential benefits.
- Develops automation tools or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.
- Independently conducts testing to ensure automation performs the task correctly and produces expected results.
- Communicates the scale, capacity, security, performance attributes, and requirements of services and technology within and sometimes beyond immediate team.
- Identifies and explains the potential impact of infrastructure, feature, and tool changes, considering their impact on team operations.
- Provides operational support for technology, escalating incidents and other standard and non-standard issues arising within Oracle services.
- Participates in on-call shifts to address issues.
- Resolves technical issues spanning various services, investigating and debugging products in order to reach SLOs (service level objectives).
- Documents incidents and performs root cause analyses according to standard reporting methods.
- Independently performs post-mortem procedures to prevent incident reoccurrence.
- Experiments with new tools and technologies to assess their potential impact on and improve infrastructure performance and reliability, ensuring adherence to security standards.
- Independently identifies and executes improvements for performance bottlenecks and deployments to ensure efficient resource usage, speed, and scalability.
- Develops knowledge of site reliability trends and shares new information with team members, management, and beyond to help others build, test, deploy and run services.
- Performs standard and non-standard analyses and provides clear data on production to contribute to business development decisions (e.g., design changes).
- Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
- Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Skills used in this role
What the employer is looking for
- Preference for individuals with experience in scripting and network automation - Python, Puppet, SQL, and/or Ansible
- Ability to act in a project lead role as needed
- Experience working in a large ISP or cloud provider environment
- Exposure to commodity Ethernet hardware (Broadcom/Mellanox)
- Protocol experience with BGP/OSPF/IS-IS, TCP, IPv4, IPv6, DNS, DHCP, MPLS
- Experience with networking protocols such as TCP/IP, VPN, DNS, DHCP, and SSL
- Experience supporting network technologies, especially Juniper, Cisco, Arista, firewalls, and switches
- Strong analytical skills and ability to collate and interpret data from various sources
- Ability to diagnose network alerts to assess and prioritize faults and respond or escalate accordingly
- Cisco and Juniper certifications are desired
- Highly motivated and self-starter
- Bachelor's degree is preferred with at least 3-5 years of network-related experience
- Strong oral and written communication skills
- Excellent time management and organization skills
- Comfortable and able to deal with a wide range of issues in a fast-paced environment
- Excellent organizational, verbal, and written communication skills
Benefits and support
- Compensation and benefits are detailed in the job posting
About Oracle
Oracle Corporation is a multinational computer technology company specializing in cloud-engineered hardware, enterprise software, and database management systems. It provides comprehensive cloud solutions such as Oracle Cloud Infrastructure (OCI) and Oracle Cloud Applications (OCA) to power digital transformations globally. Headquartered in Austin, Texas, Oracle is a major player in enterprise AI integration and database solutions.
- Industry
- Cloud Computing
- Company size
- 141000 employees
- Founded
- 1977
- Location
- Austin, Texas, USA
- Funding stage
- Public Company
Funding
Public Company · $56B raised
- 2025-10-24Post IPO Debt$18B
- 2025-02-03Conventional Debt$18B
Leadership
Lawrence J. Ellison
Chairman of the Board and Chief Technology Officer
Clay Magouyrk
Chief Executive Officer
Mike Sicilia
Chief Executive Officer
Hilary Maxson
Chief Financial Officer
Jae Evans
Global Chief Information Officer and Executive Vice President
Recent coverage
Oracle Newsroom
Oracle Announces $10 Million Partnership with the Nashville Symphony2026-10-04
Forbes
Oracle Is Absorbing A $300 Million Nuclear Bill That Wisconsin Ratepayers Were Stuck With2026-10-02
Oracle Newsroom
Oracle Ranks No. 1 in 2026 ISG Buyers Guide™ for Workforce Management Suites2026-10-02
Bitcoin.com
Oracle Gave Larry Ellison and His Co-CEOs Nearly $1 Billion in Stock Options2026-10-05