
Site Reliability Engineer
Akamai · Posted Sep 29
Cloud computing, cybersecurity, and content delivery network services
Get a personal compatibility score
Add a resume for personal matches
About the role
The Edge Reliability Engineering (ERE) team enhances performance, availability, and resilience for Akamai's media and web delivery products, optimizing core edge platform and web server technologies and acting as a technical escalation resource for global support. The team works across engineering, support, and operations to resolve critical, customer-impacting issues on Akamai's global edge, media delivery, and web delivery systems. As a Site Reliability Engineer, you will investigate reliability and performance problems across application, platform, network, and OS layers, use telemetry and SLOs to assess system health, and build automation and AI-powered workflows to improve diagnostics and incident response.
What you will do
- Taking responsibility for assigned reliability and performance investigations across Akamai's global edge platform, media delivery, and web delivery systems.
- Diagnosing problems across application, platform, network, and operating-system layers utilizing logs, metrics, traces, and diagnostic tools appropriately.
- Collaborating with Product, Engineering, Support, Network, and senior SRE teams to determine root causes and apply lasting solutions.
- Utilizing telemetry, SLIs, SLOs, KPIs, dashboards, and alerts to evaluate system health, determine customer impact, and enhance reliability improvements.
- Analyzing platform behavior, traffic patterns, and system bottlenecks to enhance performance, scalability, and resilience effectively across systems.
- Creating scripts, automation, internal tools, and workflows to minimize operational effort while enhancing diagnostics and incident response efficiency.
- Applying AI-powered analysis and creating workflows as needed, while ensuring precision, safety, and thorough human evaluation.
Skills used in this role
What the employer is looking for
- Demonstrate expertise in Computer Science, Engineering, or related fields, or possess substantial industry experience in large-scale SRE roles.
- Evaluate technical issues, assess system behavior, and derive insights from evidence to improve understanding and problem-solving abilities.
- Understand foundational networking and Internet technologies, such as caching, proxies, TLS, TCP/IP, DNS, and HTTP/HTTPS architectures.
- Demonstrate proficiency with Linux or Unix systems, command-line utilities, and foundational diagnostic methods.
- Demonstrate proficiency with structured data and telemetry or exhibit foundational SQL capabilities.
- Read and analyze C++ code to identify and resolve issues within core server infrastructure effectively.
Benefits and support
- FlexBase flexible work program: work from home, in an office, or a combination of both
About Akamai
Akamai is a global cloud computing, cybersecurity, and content delivery network company that powers and protects digital experiences online. Its distributed edge platform optimizes web and application performance while providing robust enterprise security solutions. The company serves major global enterprises across various industries, enabling scalable cloud infrastructure and advanced threat mitigation.
- Industry
- Cloud Computing
- Company size
- 11000+ employees
- Founded
- 1998
- Location
- Cambridge, Massachusetts, USA
- Funding stage
- Public Company
Funding
Public Company · $61.5M raised
- 2023-08-16Post Ipo Debt
Leadership
Chief Executive Officer & Co-Founder
EVP & Chief Financial Officer
EVP & Chief Technology Officer
Chief Operating Officer & GM, Cloud Technology Group
Recent coverage
Reuters
Akamai Signs $11.6 Billion Cloud Deal with Anthropic, Grants Warrant for Up to 5% Stake2026-09-24
Akamai Newsroom
Akamai Reports Second Quarter 2026 Financial Results2026-08-06
Akamai Newsroom
Akamai Completes Acquisition of Secure Enterprise Browser Provider LayerX2026-07-02