Qualcomm logo

Senior Site Reliability Engineer

Qualcomm · Posted Oct 9

Semiconductor, 5G wireless technology, mobile platform, and AI computing solutions

Chennai, Tamil Nādu, IndiaFull-timeOnsiteSenior Level5+ years₹30.0L–₹50.0L yearly74 applicants
SemiconductorsWireless TechnologyTelecommunicationsArtificial IntelligenceMobile ComputingPublic Company
Full time

Get a personal compatibility score

Add a resume for personal matches

About the role

Qualcomm's engineering group is hiring a Senior Site Reliability Engineer for a 24/7, follow-the-sun SRE team responsible for keeping critical production systems reliable, scalable, and secure. The team works across multiple time zones with on-call coverage that follows daylight hours, partnering closely with development teams to automate operations, strengthen observability, and improve system resilience. This role focuses on defining and reporting SLIs, SLOs, and error budgets, building automation and monitoring, leading incident response and post-mortems, tuning performance, and optimizing cloud cost on AWS, Docker, and Kubernetes.

What you will do

  • System Reliability: Ensure the reliability, availability, and performance of critical systems.
  • Service Level Objectives: Define, measure, and report on SLIs, SLOs, and error budgets, and use them to prioritize reliability work.
  • Automation: Develop and maintain automation scripts and tools to streamline operations.
  • Monitoring: Develop and maintain monitoring dashboards & alerts
  • Incident Management: Lead incident response efforts and post-mortem analysis to prevent future occurrences.
  • Performance Tuning: Optimize system performance and scalability.
  • Cost Optimization: Monitor and optimize cloud spend (e.g., right-sizing and autoscaling with Karpenter) to balance reliability with cost efficiency.
  • Security: Implement and maintain security best practices.
  • Documentation: Create and maintain comprehensive documentation for systems and processes.
  • Mentorship: Mentor junior engineers and champion SRE best practices across cross-functional teams.
  • On-Call Shifts: Own front-line 24/7 on-call rotations and incident response for critical production systems, acting as the reliability shield for the platform and observability engineering teams so they are not paged for production incidents.

Skills used in this role

PythonGoAWSTerraformAnsibleCloudFormationDockerKubernetesArgo CDLinkerdPrometheusGrafanaKarpenterELK StackLokiJenkinsGitLab CICircleCICC++JavaNetworkingProblem SolvingCommunicationLeadership

What the employer is looking for

  • Education: Bachelor’s degree in computer science, Engineering, or a related field. Advanced degrees are a plus.
  • Experience: 5+ years of experience in a similar role, with a strong background in software engineering and systems administration.
  • Programming Languages: Proficiency in one or more programming languages such as Python, Go.
  • Cloud Platforms: Extensive experience with AWS cloud platform.
  • Infrastructure as Code: Hands-on experience with tools like Terraform, Ansible, or CloudFormation.
  • Containerization and Orchestration: Expertise in Docker and Kubernetes. Kubernetes - Hands-on experience is a must.
  • Kubernetes technologies (with preferences): ArgoCD, Linkerd, Prometheus, Karpenter, etc…
  • Monitoring and Logging: Proficiency with monitoring tools like Prometheus, Grafana, and logging tools like ELK stack or Loki stack.
  • CI/CD Pipelines: Experience with continuous integration and continuous deployment tools such as Jenkins, GitLab CI, or CircleCI.
  • Networking: Strong understanding of networking concepts, protocols, and security.
  • Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience. OR Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience. OR PhD in Engineering, Information Systems, Computer Science, or related field.
  • 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
  • Problem-Solving: Excellent analytical and troubleshooting skills.
  • Communication: Strong verbal and written communication skills.
  • Collaboration: Ability to work effectively in a team environment and collaborate with cross-functional teams.
  • Leadership: Proven leadership skills and the ability to mentor junior engineers.
  • Remote Work: Comfortable working in a fully distributed, offshore setup and collaborating effectively with development teams across multiple locations.

Benefits and support

  • Compensation and benefits are detailed in the job posting

About Qualcomm

Qualcomm is a global leader in the development and commercialization of foundational technologies for the wireless industry, including mobile processors, 5G, and on-device artificial intelligence. The company designs and markets digital communications products, semiconductor solutions, and system software used globally across smartphones, automotive solutions, IoT devices, and computing systems. Through extensive R&D and intellectual property licensing, Qualcomm powers a highly interconnected and smarter world.

Industry
Semiconductors
Company size
50000+ employees
Founded
1985
Location
San Diego, California, USA
Funding stage
Public Company

Funding

Public Company · $53M raised

Western Technology Investment
  • 2018-11-01Grant

Leadership

CR
Cristiano R. Amon

President and Chief Executive Officer

AP
Akash Palkhiwala

Chief Financial Officer and Chief Operating Officer

JH

James H. Thompson

Chief Technology Officer

DM
Don McGuire

Chief Marketing Officer

AC

Ann Chaplin

General Counsel and Corporate Secretary