
Site Reliability Engineer
Autodesk · Posted Sep 18
3D design, engineering, architecture, and entertainment software solutions
Get a personal compatibility score
Add a resume for personal matches
About the role
Autodesk's Product Design and Manufacturing Solutions (PDMS) Platform Site Reliability Engineering (SRE) team is looking for a Site Reliability Engineer. In this role you will wear multiple hats, including first responder, performance analyst, system architect, capacity planner, and monitoring expert, helping build and operate reliable, scalable, secure, and high-performing cloud infrastructure that supports Autodesk products and customers. You will bring strong technical and communication skills, a passion for learning new technologies, and a problem-solving mindset.
What you will do
- Architect and implement hosting solutions for highly dynamic Software as a Service (SaaS) web applications, ensuring reliability, scalability, and performance
- Design, implement, and maintain Infrastructure as Code (IaC) solutions to support scalable, reliable, and secure global environments
- Develop and maintain well-documented engineering standards, processes, and best practices
- Implement infrastructure and application security best practices, including system hardening and the principle of least privilege
- Use modern infrastructure management tools such as Docker, Terraform, Amazon Web Services (AWS) CloudFormation, and AWS Cloud Development Kit (CDK) to manage and deploy containers and virtual machines
- Collaborate with Development, Quality Assurance, and Documentation teams throughout the product development lifecycle to ensure quality and reliability
- Automate operational processes and integrate new technologies to improve efficiency, reliability, and scalability
- Define and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) and manage error budgets to ensure reliability goals are achieved
- Partner with stakeholders to align technical strategies with business requirements
- Participate in on-call support and incident management, ensuring timely resolution and clear stakeholder communication
- Conduct blameless post-incident reviews to identify root causes, document learnings, and drive continuous improvement
- Take ownership of initiatives and contribute to a culture of continuous learning, operational excellence, and continuous improvement
Skills used in this role
What the employer is looking for
- 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related role supporting cloud-based applications
- Bachelor's degree in Computer Science or a related technical field
- Advanced hands-on experience with Linux administration, including monitoring, troubleshooting, reliability, performance, and security
- Experience managing large-scale cloud infrastructure, preferably on Amazon Web Services (AWS)
- Strong scripting skills using languages such as Bash, Python, or Perl
- Expert-level knowledge of AWS services, including Amazon Elastic Compute Cloud (EC2), Elastic Container Service (ECS), Elastic Kubernetes Service (EKS), AWS Lambda, Elastic Load Balancing (ELB), Amazon Simple Storage Service (S3), Identity and Access Management (IAM), Virtual Private Cloud (VPC), Amazon DynamoDB, and Amazon Relational Database Service (RDS)
- Hands-on experience with Docker, Kubernetes, and container technologies
- Proficiency with Infrastructure as Code (IaC) tools such as Terraform and AWS CloudFormation
- Experience with Continuous Integration and Continuous Deployment (CI/CD) tools and technologies such as Jenkins, JFrog Artifactory, and Git
- Experience with logging, monitoring, and observability tools such as Amazon CloudWatch, Splunk, Dynatrace, New Relic, and Grafana
- Experience with relational database technologies such as MySQL, PostgreSQL, and Microsoft SQL Server, along with Structured Query Language (SQL)
- Excellent analytical and problem-solving skills with the ability to work independently
- Excellent written and verbal communication skills
Preferred qualifications
- Experience using Artificial Intelligence (AI)-assisted engineering tools and development practices
- Experience designing and operating highly available, distributed cloud-native systems
- Experience with Site Reliability Engineering practices, including observability, capacity planning, incident management, and error budget management
- Experience automating infrastructure and operational processes at scale
- Knowledge of cloud security, infrastructure hardening, and compliance best practices
- Experience working in globally distributed engineering teams
- Demonstrates strong technical expertise in cloud infrastructure, Linux administration, containers, automation, and observability
- Applies Site Reliability Engineering principles to improve system availability, performance, scalability, and operational efficiency
- Uses analytical thinking and structured problem-solving to troubleshoot complex production issues and identify root causes
- Takes ownership of services and infrastructure while proactively identifying and addressing reliability risks
- Automates repetitive operational activities to improve engineering efficiency and reduce manual intervention
- Responds effectively to incidents while maintaining clear communication and driving timely resolution
- Learns from incidents and contributes to a blameless culture focused on continuous improvement
- Collaborates effectively with Software Engineering, Quality Assurance, and cross-functional teams to improve product and platform reliability
- Demonstrates curiosity and continuously develops expertise in emerging cloud, automation, observability, and Artificial Intelligence technologies
- Communicates technical concepts clearly and contributes to a collaborative culture focused on reliability, accountability, and engineering excellence
Benefits and support
- Compensation and benefits are detailed in the job posting
About Autodesk
Autodesk is a global leader in 3D design, engineering, and entertainment software, providing solutions for architecture, engineering, construction, product design, and media creation. Its flagship products, such as AutoCAD, Revit, and Fusion 360, empower professionals across industries to imagine, design, and make a better world. As a pioneer in the space, the company continues to drive innovation through cloud collaboration, digital twins, and AI-powered design workflows.
- Industry
- Software
- Company size
- 10001+ employees
- Founded
- 1982
- Location
- San Francisco, California, USA
- Funding stage
- Public Company
Leadership
President and Chief Executive Officer
Executive Vice President and Chief Financial Officer
Executive Vice President and Chief Technology Officer
Executive Vice President and Chief Operating Officer
Chief Marketing Officer
Recent coverage
Autodesk PR
Autodesk Advances Agentic AI in Its Three Industry Clouds2026-10-06
Autodesk PR
Autodesk names Diana Colella to lead architecture, engineering and construction business2026-10-06
Autodesk PR
MotionMaker in Maya named to Fast Company's 2026 Next Big Things in Tech list2026-10-06
Autodesk
Closing the gap between what we can imagine and what we can build: Autodesk University highlights Autodesk AI2026-09-15