Akamai logo

Site Reliability Engineer

Akamai · Posted Sep 29

Cloud computing, cybersecurity, and content delivery network services

Bengaluru, IndiaFull-timeRemote · India-wideSenior Level3–6 years₹22.0L–₹35.0L yearly100+ applicants
Cloud ComputingCybersecurityContent Delivery NetworkInformation TechnologyPublic Company
Full time

Get a personal compatibility score

Add a resume for personal matches

About the role

The Edge Reliability Engineering (ERE) team enhances performance, availability, and resilience for Akamai's media and web delivery products, optimizing core edge platform and web server technologies and acting as a technical escalation resource for global support. The team works across engineering, support, and operations to resolve critical, customer-impacting issues on Akamai's global edge, media delivery, and web delivery systems. As a Site Reliability Engineer, you will investigate reliability and performance problems across application, platform, network, and OS layers, use telemetry and SLOs to assess system health, and build automation and AI-powered workflows to improve diagnostics and incident response.

What you will do

  • Taking responsibility for assigned reliability and performance investigations across Akamai's global edge platform, media delivery, and web delivery systems.
  • Diagnosing problems across application, platform, network, and operating-system layers utilizing logs, metrics, traces, and diagnostic tools appropriately.
  • Collaborating with Product, Engineering, Support, Network, and senior SRE teams to determine root causes and apply lasting solutions.
  • Utilizing telemetry, SLIs, SLOs, KPIs, dashboards, and alerts to evaluate system health, determine customer impact, and enhance reliability improvements.
  • Analyzing platform behavior, traffic patterns, and system bottlenecks to enhance performance, scalability, and resilience effectively across systems.
  • Creating scripts, automation, internal tools, and workflows to minimize operational effort while enhancing diagnostics and incident response efficiency.
  • Applying AI-powered analysis and creating workflows as needed, while ensuring precision, safety, and thorough human evaluation.

Skills used in this role

C++LinuxUnixSQLTCP/IPDNSTLSHTTP/HTTPSCachingProxiesNetworkingTelemetryDistributed SystemsAutomationProblem Solving

What the employer is looking for

  • Demonstrate expertise in Computer Science, Engineering, or related fields, or possess substantial industry experience in large-scale SRE roles.
  • Evaluate technical issues, assess system behavior, and derive insights from evidence to improve understanding and problem-solving abilities.
  • Understand foundational networking and Internet technologies, such as caching, proxies, TLS, TCP/IP, DNS, and HTTP/HTTPS architectures.
  • Demonstrate proficiency with Linux or Unix systems, command-line utilities, and foundational diagnostic methods.
  • Demonstrate proficiency with structured data and telemetry or exhibit foundational SQL capabilities.
  • Read and analyze C++ code to identify and resolve issues within core server infrastructure effectively.

Benefits and support

  • FlexBase flexible work program: work from home, in an office, or a combination of both

About Akamai

Akamai is a global cloud computing, cybersecurity, and content delivery network company that powers and protects digital experiences online. Its distributed edge platform optimizes web and application performance while providing robust enterprise security solutions. The company serves major global enterprises across various industries, enabling scalable cloud infrastructure and advanced threat mitigation.

Industry
Cloud Computing
Company size
11000+ employees
Founded
1998
Location
Cambridge, Massachusetts, USA
Funding stage
Public Company

Funding

Public Company · $61.5M raised

Apple
  • 2023-08-16Post Ipo Debt

Leadership

TL
Tom Leighton

Chief Executive Officer & Co-Founder

EM
Ed McGowan

EVP & Chief Financial Officer

RB
Robert Blumofe

EVP & Chief Technology Officer

AK
Adam Karon

Chief Operating Officer & GM, Cloud Technology Group