
Software Engineer, Platform Infrastructure (Foundations)
Anyscale · Posted Sep 29
Distributed computing platform and managed cloud services for scaling artificial intelligence and machine learning workloads
Get a personal compatibility score
Add a resume for personal matches
About the role
Anyscale commercializes Ray, a popular open-source distributed computing framework used for scalable machine learning, and is building the best place to run Ray so developers can scale ML applications from laptop to cluster. The Infrastructure team builds the scalable, secure, and robust backbone for running distributed AI applications in the cloud, spanning both the control plane (cluster management, scheduling, user access) and the data plane (high-performance execution of distributed workloads). This role designs, implements, and optimizes that critical infrastructure, working on open-source Ray, the proprietary product, and their integration, with Kubernetes, cloud-native, and accelerator-focused projects.
What you will do
- Design, build, and scale services that orchestrate Ray clusters across cloud and on-prem environments, supporting both VM-based and Kubernetes-based deployments
- Optimize control plane components for large-scale, distributed AI/ML workloads
- Build intelligent scheduling and resource management systems for heterogeneous compute clusters
- Develop features to enhance the reliability, performance, scalability, and observability of Anyscale-managed Ray workloads
- Support and optimize accelerator integration (e.g., GPUs, TPUs).
- Handle container image management and dependency resolution for distributed workloads
- Participate in code reviews, design and architecture discussions
- Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues
- Collaborate with leading distributed systems and machine learning experts to push the boundaries of AI infrastructure
Skills used in this role
What the employer is looking for
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
- 3+ years of experience writing high-quality production code
- Hands-on experience in building and maintaining highly available, scalable, and performant distributed system
- Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments
- Deep understanding of networking, security, and authentication mechanisms in cloud environment
- Familiarity with observability stacks (Prometheus, Grafana etc)
- Proficiency in Go and Python
- Knowledge of low-level operating system foundations (Linux kernel, file systems, containers)
Benefits and support
- Compensation and benefits are detailed in the job posting
About Anyscale
Anyscale is the company behind Ray, an open-source distributed computing framework designed to scale artificial intelligence and machine learning workloads seamlessly from a laptop to the cloud. Its managed platform offers elastic GPU orchestration, distributed training, and batch inference capabilities to help development teams build production-grade AI applications without infrastructure complexities.
- Industry
- Artificial Intelligence
- Company size
- 200-500 employees
- Founded
- 2019
- Location
- San Francisco, California, USA
- Funding stage
- Series C
Funding
Series C · $281M raised
- 2021-10-15Series C$199M
- 2020-10-21Series B$40M
- 2019-08-02Series A$20M
- 2019-02-18Seed$22M
Leadership
Chief Executive Officer
Co-founder & Chief Strategy Officer
Co-founder & CTO
Co-founder & Executive Chairman
Recent coverage
Anyscale Press Release
Nscale Acquires Anyscale, Enhancing its Full Stack AI Cloud Platform2026-07-30
Anyscale Press Release
Anyscale Launches on Microsoft Azure as a Native Integration for Enterprises2026-06-02
Anyscale Press Release
Anyscale Names Christian Stano as Field CTO as Global Adoption of Ray Accelerates2026-03-19
Anyscale Press Release
Anyscale Cuts Multimodal AI Data Processing Costs by 80% with NVIDIA RTX2026-03-16