Scan your resume against ATS criteria for this Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure) role at Uvation.
Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
We are seeking a highly experienced
with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation
. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments. This is
. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in
. The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
platforms and large-scale infrastructure environments - Strong understanding of server hardware, including: - BIOS/UEFI - RAID controllers - Firmware management - iLO/iDRAC/IPMI - NICs and SmartNICs - HBA cards - Hardware diagnostics and troubleshooting - Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
for AI/ML workloads - Understanding of NVIDIA GPU technologies including: - A100, H100, H200, B200, or equivalent GPU platforms - NVIDIA DGX and OEM GPU servers - GPU provisioning and lifecycle management - GPU monitoring and performance optimization - Knowledge of AI Factory architecture and infrastructure requirements - Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads - Understanding of: - GPU resource allocation and scheduling - Multi-GPU systems - GPU networking requirements - High-bandwidth, low-latency infrastructure design - Familiarity with NVIDIA ecosystem technologies such as: - CUDA - NCCL - GPUDirect Storage - NVIDIA Fabric Manager - NVIDIA Base Command (preferred)
, including: - Cluster architecture - MON, OSD, MDS - RBD, CephFS, RGW - Capacity planning - Performance tuning - Failure recovery - Experience with high-performance AI storage platforms such as: - WEKA - VAST Data - Dell PowerScale - Pure Storage FlashBlade - NetApp - Understanding of: - NVMe-over-Fabrics (NVMe-oF) - RDMA - GPUDirect Storage - Parallel file systems - AI data pipelines
Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.
How to apply for Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure) at Uvation?
Click the "Apply on Company Website" button on this page to submit your application directly on the employer's official portal.
What is the salary for this role?
Salary details will be discussed during the interview.
What experience is required?
This position is open to freshers and experienced candidates.
Is this position still open?
Yes, currently active and accepting applications.
Explore related active roles in Manufacturing & Ops
Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
Uvation · India