
Senior Network Engineer – GPU Cluster Networking
AMD · Posted Oct 9
High-performance processors, graphics cards, data center solutions, and AI accelerators
Get a personal compatibility score
Add a resume for personal matches
About the role
AMD's IT Network Engineering team is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters and hyperscale AI and HPC environments. The team's primary focus is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters, spanning switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems. This role owns the network path from the GPU server and NIC through the data center switching fabric so that distributed AI training, large language model inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance, partnering with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams.
What you will do
- Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
- Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.
- Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.
- Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.
- Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.
- Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.
- Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
- Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.
- Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.
- Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.
- Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations
Skills used in this role
What the employer is looking for
- We are seeking a Senior Network Engineer with 12 to 20 yrs of experience to join the AMD IT Network Engineering team.
- The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.
- You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.
- You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.
- You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.
Preferred qualifications
- Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.
- Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.
- Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation
- Strong hands-on experience with RDMA and RoCEv2 in production environments.
- Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.
- Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.
- Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.
- Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.
- Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.
- Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
Benefits and support
- Compensation and benefits are detailed in the job posting
About AMD
AMD is a global semiconductor company that designs high-performance central processing units (CPUs), graphics processing units (GPUs), and adaptive processors for data centers, personal computers, gaming consoles, and embedded systems. The company specializes in advanced computing solutions and AI accelerators, positioning itself as a major leader in modern high-performance and artificial intelligence hardware. Through strategic developments and acquisitions, AMD continues to expand its ecosystem across cloud infrastructure, enterprise computing, and physical AI.
- Industry
- Semiconductors
- Company size
- 31000 employees
- Founded
- 1969-05-01
- Location
- Santa Clara, California, USA
- Funding stage
- Public Company
Funding
Public Company
- 2007-11-16Post Ipo
Leadership
Chair and Chief Executive Officer
Executive Vice President, Chief Financial Officer and Treasurer
Executive Vice President and Chief Technology Officer
Executive Vice President and Chief Commercial Officer
Executive Vice President and General Manager, Data Center Solutions Business Group
Recent coverage
Taipei Times
AMD to 'Substantially' Up Capacity: CEO2026-10-07
AMD Newsroom
AMD to Report Fiscal Third Quarter 2026 Financial Results2026-10-06
Venture Atlas
AMD Agrees to Buy World Labs for About $8.2 Billion2026-09-29