AMD logo

Senior Network Engineer – GPU Cluster Networking

AMD · Posted Oct 9

High-performance processors, graphics cards, data center solutions, and AI accelerators

Hyderabad, IndiaFull-timeOnsiteLead/Staff12–20 years₹45.0L–₹85.0L yearly50 applicants
SemiconductorsHardwareArtificial IntelligenceCloud ComputingGamingPublic Company
Full time

Get a personal compatibility score

Add a resume for personal matches

About the role

AMD's IT Network Engineering team is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters and hyperscale AI and HPC environments. The team's primary focus is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters, spanning switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems. This role owns the network path from the GPU server and NIC through the data center switching fabric so that distributed AI training, large language model inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance, partnering with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams.

What you will do

  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.
  • Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.
  • Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.
  • Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.
  • Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.
  • Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
  • Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.
  • Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.
  • Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.
  • Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations

Skills used in this role

RoCEv2RDMAEthernetBGPECMPEVPNVXLANVLANVRFQoSPFCECNDCQCNLinuxPCIeNUMAROCmRCCLSlurmKubernetesPrometheusGrafanagNMIJuniperJunos OS

What the employer is looking for

  • We are seeking a Senior Network Engineer with 12 to 20 yrs of experience to join the AMD IT Network Engineering team.
  • The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.
  • You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.
  • You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.
  • You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.

Preferred qualifications

  • Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.
  • Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.
  • Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation
  • Strong hands-on experience with RDMA and RoCEv2 in production environments.
  • Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.
  • Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.
  • Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.
  • Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.
  • Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.
  • Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting

Benefits and support

  • Compensation and benefits are detailed in the job posting

About AMD

AMD is a global semiconductor company that designs high-performance central processing units (CPUs), graphics processing units (GPUs), and adaptive processors for data centers, personal computers, gaming consoles, and embedded systems. The company specializes in advanced computing solutions and AI accelerators, positioning itself as a major leader in modern high-performance and artificial intelligence hardware. Through strategic developments and acquisitions, AMD continues to expand its ecosystem across cloud infrastructure, enterprise computing, and physical AI.

Industry
Semiconductors
Company size
31000 employees
Founded
1969-05-01
Location
Santa Clara, California, USA
Funding stage
Public Company

Funding

Public Company

  • 2007-11-16Post Ipo

Leadership

DL
Dr. Lisa Su

Chair and Chief Executive Officer

JH
Jean Hu

Executive Vice President, Chief Financial Officer and Treasurer

MP
Mark Papermaster

Executive Vice President and Chief Technology Officer

PG
Philip Guido

Executive Vice President and Chief Commercial Officer

FN
Forrest Norrod

Executive Vice President and General Manager, Data Center Solutions Business Group