CareerScanCareerScan
JobsCompanies
BlogContact
For Employers
Sign InRegister Free
CareerScanCareerScan

India's verified job platform connecting candidates directly with employers. 100% free applications with instant ATS resume scoring.

Chennai, Bengaluru & Hyderabad
Jobs by location
Jobs in ChennaiJobs in BengaluruJobs in HyderabadJobs in PuneJobs in Mumbai
Popular roles
AR Caller JobsHealthcare Medical CodingReact / Full Stack DeveloperData & Power BI AnalystCustomer Support Executive
Top companies
TCS CareersCognizant JobsInfosys OpeningsApollo HospitalsOmega Healthcare
Career services
Free ATS Resume CheckerAI Resume Builder (Free)AI Job MatcherSalary Guide & BenchmarksJob Alerts on WhatsApp
© 2026 CareerScan India. All rights reserved.256-bit SSL encrypted & verified
Back to all jobs
  1. Home
  2. Jobs
  3. Senior Network Engineer – GPU Cluster Networking
AM
Advanced Micro Devices, Inc

Senior Network Engineer – GPU Cluster Networking

US,CA,San Jose
Full-time
Posted 5d ago
2 views
Actively Hiring Urgent Opening Direct 1-Click Apply

Check Your Resume Match Score

Scan your resume against ATS criteria for this Senior Network Engineer – GPU Cluster Networking role at Advanced Micro Devices, Inc.

Apply for this position

Apply on Company Website
Notice a broken link or wrong info?

Job Description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.

  • THE ROLE:

We are seeking a Senior Network Engineer to join the AMD IT System Engineering team.

This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.

The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.

The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.

You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.

  • THE PERSON:

You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.

You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.

You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.

  • KEY RESPONSIBILITIES:
  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.

  • Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.

  • Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.

  • Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.

  • Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.

  • Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.

    • Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
  • Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.

  • Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.

  • Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.

  • Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations

  • PREFERRED EXPERIENCE:
  • Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.
  • Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.
  • Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation
  • Strong hands-on experience with RDMA and RoCEv2 in production environments.
  • Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.
  • Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.
  • Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.
  • Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.
  • Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.
  • Experience with

Juniper data center switching platforms and Junos OS

, including configuration and troubleshooting

  • Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.
  • Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.
  • Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.
  • Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.
  • ACADEMIC CREDITALS:
  • Bachelor’s or Master’s degree in Computer Engineering, or a related field, or equivalent practical experience.
  • LOCATION:

San Jose, CA OR Austin, TX

This role is not eligible for visa sponsorship.

#LI-BS1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Key Requirements & Skills

  • Experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or large-scale distributed computing environments
  • Experience with backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs or hyperscale compute environments
  • Deep knowledge of data center networking fundamentals: routing/switching, VLANs, subnetting, BGP, ECMP, QoS, MTU, switch buffering, network segmentation
  • Strong hands-on experience with RDMA and RoCEv2 in production environments
  • Experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet
  • Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures
  • Experience with BGP, ECMP, EVPN, VXLAN, and modern leaf-spine data center architectures
  • Strong understanding of GPU cluster topology: GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality
  • Experience with monitoring and observability using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, or sFlow
  • Experience with Juniper data center switching platforms and Junos OS
  • Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments
  • Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies
  • Experience designing backend networks for large language model training and communication-intensive distributed AI workloads
  • Bachelor's or Master's degree in Computer Engineering or related field, or equivalent practical experience

Benefits & Perks

medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.*

Frequently Asked Questions

How to apply for Senior Network Engineer – GPU Cluster Networking at Advanced Micro Devices, Inc?

Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.

What is the salary for this role?

Salary details will be discussed during the interview.

What experience is required?

This position is open to freshers and experienced candidates.

Is this position still open?

Yes, currently active and accepting applications.

ApplicationActively Hiring
Apply on Company Website
Broken link or expired?
AM

Advanced Micro Devices, Inc

More jobs at Advanced Micro Devices, Inc

Silicon Design Engineer 1

IN,Bangalore-Design Center

FPGA Prototyping Engineer

IN,Hyderabad-Sattva

SOFTWARE DEVELOPMENT ENGINEER (Python / C++)

IN,Hyderabad-Sattva

Share this Opening

Job Alerts for network_engineering

Receive email alerts whenever new network_engineering roles in US are posted.

Set Free Alert →

Similar Openings

Explore related active roles in network_engineering

View all
Actively Hiring
Zensar Technologies
Network Engineer L3
Zensar Technologies Verified
0-2 Yrs
Salary not disclosed
India
network_engineeringFull-time
Posted 5d ago
Apply Now
Actively Hiring
Apps Associates
Senior Network Engineer - Support
Apps Associates Verified
10+ years
Salary not disclosed
Hyderabad, Telangana, India
network_engineeringFull-time
Posted 5d ago
Apply Now
Actively Hiring
Digitide Solutions Limited
Senior Network Engineer
Digitide Solutions Limited Verified
0-2 Yrs
Salary not disclosed
India
network_engineeringFull-time
Posted 5d ago
Apply Now

Senior Network Engineer – GPU Cluster Networking

Advanced Micro Devices, Inc · US

Apply on Company Website