
Staff Reliability Engineer
Okta · Posted Sep 16
Cloud-based identity and access management solutions for workforce and customer security
Get a personal compatibility score
Add a resume for personal matches
About the role
Okta's TDI Network Engineering team is responsible for the global corporate network, building and supporting a high-performing, reliable network at scale, ensuring employees can work effectively from any location globally. This operations-focused Staff Reliability Engineer role is distinct from core Network Engineering and Network Security, centering on operational execution—responding to alerts, maintaining service availability, and ensuring system health across the global enterprise network. Reporting to the Network Engineering Manager, you will drive strategic reduction of systemic toil and technical debt, architect scalable platforms, and lead technical efforts to ensure an "Always Secure. Always On." environment.
What you will do
- Design and Own the resilience, health and availability of our entire global corporate network domain, managing operational responsibilities such as responding to alerts, monitoring health indicators, and executing reliability projects to ensure an "Always Secure. Always On." environment.
- Drive Strategic Reduction of systemic toil and technical debt across multiple teams by introducing process efficiencies, automating network operations, and building scalable self-service operational tooling.
- Collaborate and Influence closely with cross-functional stakeholders—including Business Technology, Workplace, Security, and executive leaders—challenging assumptions with grace and cascading relevant information to project teams.
- Make Critical Decisions and lead the resolution of complex network operations issues and alerts from a systems perspective, anticipating potential business challenges, monitoring leading indicators, and preventing future outages.
- Foster Learning and Talent by defining success for the whole team, cultivating an open and transparent environment, and actively mentoring team members through the P4 level to develop their skills and operational engineering best practices.
Skills used in this role
What the employer is looking for
- Typically requires 8+ years of related experience in a professional role with a Bachelor's degree; or 6+ years with a Master's degree; or 3+ years with a PhD; or equivalent experience.
- Deep expertise in AWS Networking and Palo Alto Networks solutions as core required technical competencies.
- Comprehensive operational experience in Distributed Systems & Networking fundamentals, including quick incident response to alerts, monitoring system health, and managing protocols such as WiFi, DNS, DHCP, VLANs, VPN, ACLs, Routing, and Firewall Policies.
- Strong proficiency in core technical skills: Cloud Platforms, IaC (e.g., Terraform/Ansible), Observability tools (e.g., Prometheus/Grafana), Programming (Python/Go), and Service Reliability Management (SLOs/SLIs) for large-scale enterprise environments.
- Proven track record of managing operational availability, delivering multi-quarter objectives, and executing technical projects within defined budgets and strategic VMTs (Vision, Mission, Targets).
- Demonstrated ability to navigate high levels of ambiguity, establish credibility with executive stakeholders, and model resilience during major system transitions or production incidents.
Preferred qualifications
- Desirable Technical Skills: Experience with Juniper/JUNOS switching/routing, Palo Alto Networks NGFWs, and enterprise office build and construction processes is highly desired and will help you hit the ground running.
Benefits and support
- Supporting Your Well-Being
- Driving Social Impact
- Developing Talent and Fostering Connection + Community
- Direct Customer Impact: You will have the opportunity to showcase your strong focus on customer and technology experiences, ensuring a secure, consistent end-user experience across a global enterprise network.
- Growth and Support: Join a collaborative environment that values systemic learning over personal errors, offering you the chance to eliminate manual toil through innovation, mentor peers, and occasionally travel to build impactful connections.
- We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate.
- Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.
About Okta
Okta is a leading independent provider of cloud-based identity and access management solutions that securely connect people and technologies to the right applications. The company powers both Workforce Identity and Customer Identity platforms, allowing organizations to securely manage user access and navigate modern challenges like agentic AI security.
- Industry
- Enterprise Software
- Company size
- 7500+ employees
- Founded
- 2009
- Location
- San Francisco, California, USA
- Funding stage
- Public Company
Funding
Public Company · $228M raised
- 2010-02-01Series A$10M
- 2011-08-01Series B$16.5M
- 2012-12-01Series C$25M
- 2013-09-01Series D$27M
- 2014-06-01Series E$75M
- 2015-09-01Series F$75M
- 2018-02-21Post IPO Debt$300M
Leadership
Co-Founder and Chief Executive Officer
Co-Founder and Executive Vice Chairman
Chief Marketing Officer
Chief Legal Officer
Recent coverage
StockTitan
Industry Leaders Form the Blueprint Alliance to Advance a Shared Architecture for Securing AI Agents2026-09-22
StockTitan
Okta Names Helen Riley to Board of Directors2026-09-18
ValueAdd VC
Okta Buys AI Identity Security Startup Permiso for ~$200M2026-07-30
StockTitan
Okta Announces Second Quarter Fiscal Year 2027 Financial Results2026-08-26