
Staff Software Engineer-AI
DigitalOcean · Posted Sep 23
Cloud infrastructure, virtual machines, managed database, and AI inference services
Get a personal compatibility score
Add a resume for personal matches
About the role
DigitalOcean's Agentic AI organization provides a powerful inference cloud, Managed Agents, and robust Feedback systems that enable customers to run AI inference confidently at scale. The team is hiring a Staff Software Engineer to serve as a technical leader within the Feedback Systems team, driving the architecture for massive-scale infrastructure that simulates, tests, and evaluates AI agents. As an IC5 Staff Engineer, you will define the architectural vision for systems that simulate multi-agent, multi-turn deployments complete with tool integration, solving problems at the intersection of LLM orchestration, synthetic data generation, and behavioral simulation.
What you will do
- Simulation Architecture & Orchestration: Leading the end-to-end design and implementation of multi-agent, multi-turn simulation environments. You will architect systems where synthetic users and AI agents interact dynamically, leverage tools, and execute complex workflows to test boundary conditions.
- Persona & Scenario Generation: Designing ML pipelines that analyze historical user conversations to automatically extract, define, and synthesize realistic user personas and multi-turn simulation goals, mirroring real-world customer behavior.
- "What-If" Evaluation Frameworks: Building the core methodology and scoring infrastructure that allows customers to run alternative configuration scenarios, benchmark agent behavior, and safely evaluate non-deterministic agent outputs against defined success criteria.
- Architectural Leadership: Leading the end-to-end design and architecture of high-throughput, stateful workflow orchestration systems capable of managing complex, multi-turn AI agent simulations at massive scale.
- System Design & Integration: Defining robust, scalable API contracts and system boundaries bridging upstream telemetry data, asynchronous simulation engines, and secure remote execution environments.
- Technical Strategy: Driving the technical roadmap for the Feedback Systems team, balancing long-term scalability and resilience with iterative product delivery.
- Complex Problem Solving: Designing elegant solutions for hard distributed systems challenges, including rate limiting, backpressure, state management, and reliable execution of non-deterministic workflows.
- Mentorship & Elevation: Mentoring senior engineers, leading cross-organizational architectural reviews, and establishing engineering best practices for code quality, testing, and system observability.
- AI/ML Infrastructure Integration: Applying your practical experience with AI/ML platforms to design and implement the backend infrastructure that powers our evaluation engines, actively managing the complexities of integrating with LLMs, prompt routing, and non-deterministic agentic workflows.
- Cross-Functional Influence: Acting as the strategic technical bridge across the Agentic AI organization, partnering closely with Product Managers and peer engineering leaders to translate complex product requirements, evaluation methodologies, and experimental needs into a scalable, future-proof architectural roadmap.
Skills used in this role
What the employer is looking for
- Agentic Expertise: 5+ years of software engineering experience with deep proficiency in modern AI/ML frameworks, LLM orchestration (e.g., LangChain, AutoGen, CrewAI, or custom multi-agent frameworks), and production-grade Python and Go.
- Behavioral Modeling & Persona Synthesis: Background in processing natural language data (e.g., historical user chat logs, support tickets) to algorithmically extract user intent, synthesize realistic personas, and generate deterministic goals for simulation.
- Evaluation & "What-If" Benchmarking: Solid experience building evaluation frameworks for non-deterministic AI systems, including establishing metrics, guardrails, scoring rubrics, and regression testing methodologies for LLM configurations.
- Data Fluency & Orchestration: Strong understanding of managing complex state in asynchronous architectures, streaming LLM tokens, handling rate limits, and manipulating heavy data pipelines to feed simulation engines.
- Ownership & Pragmatism: A strong sense of technical ownership, a passion for balancing cutting-edge ML research with practical product delivery, and excellent communication skills to collaborate across a globally distributed team.
- Extensive Experience: 10+ years of software engineering experience, with a proven track record operating at a Staff, Principal, or Architect level designing mission-critical distributed systems.
- Distributed Systems Expertise: Expert-level understanding of designing highly concurrent, fault-tolerant, and globally scalable backend architectures.
- Advanced Orchestration: Deep architectural experience with stateful, durable workflow orchestration engines and managing complex asynchronous lifecycles at scale.
- API & Systems Integration: Extensive experience designing resilient, high-performance APIs (e.g., gRPC) and managing high-throughput message/event-driven architectures.
- AI/ML Engineering Experience: While you do not need to be an ML researcher, you have demonstrable experience building, scaling, or integrating backend infrastructure for AI/ML products. This includes hands-on experience working with LLMs, agentic architectures, and solving the unique infrastructure challenges of testing non-deterministic systems.
Benefits and support
- Hybrid work model (#LI-Hybrid)
- Why You'll Like Working for DigitalOcean
About DigitalOcean
DigitalOcean is an American cloud infrastructure provider and AI-native cloud platform built for developers, startups, and small-to-medium-sized businesses. The company provides simple, scalable virtual machines, managed databases, Kubernetes, and specialized inference engines designed to run modern application and agentic AI workloads.
- Industry
- Cloud Computing
- Company size
- 1001-5000 employees
- Founded
- 2011-06-24
- Location
- Broomfield, Colorado, USA
- Funding stage
- Public Company
Funding
Public Company · $174M raised
- 2013-07-01Seed$3.2M
- 2014-03-01Series A$37.2M
- 2014-12-01Debt Financing$50M
- 2015-07-01Series B$83M
- 2016-04-01Debt Financing$130M
- 2020-05-01Venture Round$50M
- 2021-03-24IPO$775.5M
Leadership
Chief Executive Officer
Chief Financial Officer
Chief Product and Technology Officer
Chief Marketing Officer
Recent coverage
Business Wire
DigitalOcean Introduces Agent Droplets: Everything an AI Agent Needs, One Simple Monthly Price2026-09-10
DigitalOcean IR
DigitalOcean Announces Second Quarter 2026 Financial Results2026-08-04
DigitalOcean Newsroom
Hippocratic AI Scales to 10 Million Patient Calls at 99.9% Clinical Safety on DigitalOcean's AI-Native Cloud2026-05-26
DigitalOcean Newsroom
DigitalOcean Acquires Katanemo Labs to Accelerate the Inference Cloud for the Agentic Era2026-04-01