
Senior Forward Deployed Engineer I (AI Infra)
DigitalOcean · Posted Sep 30
Cloud infrastructure, virtual machines, managed database, and AI inference services
Get a personal compatibility score
Add a resume for personal matches
About the role
DigitalOcean's AI-Native Cloud team builds the simplest scalable cloud platform, and the Forward Deployed Engineering team sits at the intersection of deep systems engineering and customer architecture. The work spans heterogeneous GPU clusters (NVIDIA H100/H200/B200/B300 and AMD Instinct MI300X/MI325X), LLM inference and training platforms, Kubernetes orchestration, and infrastructure-as-code. This role embeds directly with strategic AI enterprises and high-growth startups to architect, implement, and fine-tune production AI infrastructure, debug low-level distributed stacks, and translate customer challenges into core platform features.
What you will do
- Embed & Execute: Act as the primary technical authority on heterogeneous AI infrastructure for high-value DigitalOcean customers, co-engineering custom GPU infrastructure solutions for their production workloads.
- Optimize Multi-Vendor AI Pipelines: Architect and fine-tune low-latency, high-throughput LLM serving platforms across NVIDIA CUDA Or AMD ROCm™ platforms using serving frameworks (e.g., vLLM, TensorRT-LLM, SGLang, TGI) and model execution techniques (quantization, KV caching, speculative decoding).
- Cluster Orchestration & SRE: Deploy, scale, and manage resilient Kubernetes clusters (DOKS/Bare Metal) tailored for compute-heavy AI workloads, utilizing tools like Ray, Slurm, and KubeFlow.
- Infrastructure as Code: Build scalable, repeatable blueprints using Terraform, Ansible, and Helm to automate multi-vendor GPU provisioning, high-speed networking, and storage stacks for customer deployments.
- Low-Level Heterogeneous Troubleshooting: Debug complex stack issues spanning host drivers (NVIDIA CUDA / AMD ROCm, HIP), container runtimes, inter-GPU communication libraries (NCCL / RCCL), high-speed interconnects (InfiniBand / RoCE / Infinity Fabric™), and distributed storage systems.
- Build for Scale: Translate common customer infrastructure challenges into core platform features, working directly with DigitalOcean’s product and core infrastructure teams to refine our cloud offering.
- Travel & Collaboration Requirements: Ability to travel up to 30% for customer engagements, strategic workshops, conferences, and internal collaboration. Ability to consistently overlap with North American business hours, including availability until at least noon Eastern Time, to collaborate effectively with customers, Product, Engineering, and go-to-market teams.
Skills used in this role
What the employer is looking for
- Cloud & Orchestration: Expertise with Linux systems engineering, Kubernetes, and Infrastructure as Code (Terraform, Helm).
- Heterogeneous GPU & Acceleration Stack: Hands-on experience managing NVIDIA Stack (CUDA, NCCL, NVLink, and Triton Inference Server ) Or AMD Stack ( ROCm™ RCCL, CDNA™)
- Inference & Distributed AI: Experience with modern LLM serving frameworks (vLLM, TensorRT-LLM, Ray Serve) running on both CUDA and ROCm backends.
- Networking & Storage: Deep understanding of high-performance interconnects (RDMA, InfiniBand, RoCE, AMD Infinity Fabric™) and high-throughput storage systems suited for massive datasets (e.g., Ceph, Lustre, NVMe-oF).
- Programming: Strong proficiency in Python and Go (C++, CUDA C/C++, or AMD HIP is a major plus).
Preferred qualifications
- AI Infrastructure & Forward Deployed Engineering Experience: 6+ years of experience working in Forward Deployed Engineering, AI Infrastructure, Technical Consulting roles supporting production AI systems.
- Customer Empathy & Technical Leadership: Ability to translate complex infrastructure concepts to engineering solutions and collaborate directly with client teams (CTOs, AI Leads).
- Builder Mentality: Preference for delivering production-ready code, low-latency container images, and deployment blueprints over slide decks.
- Agility: Comfortable navigating fast-moving environments and tuning model workloads for diverse accelerator architectures.
- Vendor & Strategic Partnership Collaboration: Experience collaborating with GPU vendors, infrastructure providers, model vendors, or ecosystem partners on benchmarking, optimization, technical validation, or launch readiness initiatives.
Benefits and support
- Reimbursement for relevant conferences, training, and education
- Access to LinkedIn Learning's 10,000+ courses
- Employee Assistance Program
- Local Employee Meetups
- Flexible time off policy
- Bonus in addition to base salary, determined based on company and individual performance
- Equity compensation to eligible employees, including equity grants upon hire and the option to participate in the Employee Stock Purchase Program
About DigitalOcean
DigitalOcean is an American cloud infrastructure provider and AI-native cloud platform built for developers, startups, and small-to-medium-sized businesses. The company provides simple, scalable virtual machines, managed databases, Kubernetes, and specialized inference engines designed to run modern application and agentic AI workloads.
- Industry
- Cloud Computing
- Company size
- 1001-5000 employees
- Founded
- 2011-06-24
- Location
- Broomfield, Colorado, USA
- Funding stage
- Public Company
Funding
Public Company · $174M raised
- 2013-07-01Seed$3.2M
- 2014-03-01Series A$37.2M
- 2014-12-01Debt Financing$50M
- 2015-07-01Series B$83M
- 2016-04-01Debt Financing$130M
- 2020-05-01Venture Round$50M
- 2021-03-24IPO$775.5M
Leadership
Chief Executive Officer
Chief Financial Officer
Chief Product and Technology Officer
Chief Marketing Officer
Recent coverage
Business Wire
DigitalOcean Introduces Agent Droplets: Everything an AI Agent Needs, One Simple Monthly Price2026-09-10
DigitalOcean IR
DigitalOcean Announces Second Quarter 2026 Financial Results2026-08-04
DigitalOcean Newsroom
Hippocratic AI Scales to 10 Million Patient Calls at 99.9% Clinical Safety on DigitalOcean's AI-Native Cloud2026-05-26
DigitalOcean Newsroom
DigitalOcean Acquires Katanemo Labs to Accelerate the Inference Cloud for the Agentic Era2026-04-01