uvation

Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

Malaysia full-time Senior Salary not listed
full-time Senior level Technology & IT Curated
Sign in to apply Free account — we bring you straight back to this role.

About the role

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms.

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)

Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management

Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments

Strong understanding of server hardware, including:
BIOS/UEFI

RAID controllers

Firmware management

iLO/iDRAC/IPMI

NICs and SmartNICs

HBA cards

Hardware diagnostics and troubleshooting

Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads

Understanding of NVIDIA GPU technologies including:
A100, H100, H200, B200, or equivalent GPU platforms

NVIDIA DGX and OEM GPU servers

GPU provisioning and lifecycle management

GPU monitoring and performance optimization

Knowledge of AI Factory architecture and infrastructure requirements

Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads

Understanding of:
GPU resource allocation and scheduling

Multi-GPU systems

GPU networking requirements

High-bandwidth, low-latency infrastructure design

Familiarity with NVIDIA ecosystem technologies such as:
CUDA

NCCL

GPUDirect Storage

NVIDIA Fabric Manager

NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

Advanced Linux storage administration:
LVM

XFS, EXT4

NFS

iSCSI

Fibre Channel SAN

Multipath I/O

Strong hands-on experience with Ceph, including:
Cluster architecture

MON, OSD, MDS

RBD, CephFS, RGW

Capacity planning

Performance tuning

Failure recovery

Experience with high-performance AI storage platforms such as:
WEKA

VAST Data

Dell PowerScale

Pure Storage FlashBlade

NetApp

Understanding of:
NVMe-over-Fabrics (NVMe-oF)

RDMA

GPUDirect Storage

Parallel file systems

AI data pipelines

Networking & Infrastructure

Strong networking knowledge:
Bonding

VLANs

Routing

MTU optimization

DNS

DHCP

Experience with high-performance data center networking:
100G/200G/400G Ethernet

RoCE

RDMA

Spine-Leaf architectures

Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies

Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

Experience with high availability, clustering, and disaster recovery

Strong troubleshooting skills across:
Linux operating systems

Hardware platforms

GPU infrastructure

Networking

Enterprise storage

Experience supporting mission-critical production environments

Bash and Python scripting for automation and operational efficiency

Experience creating operational documentation, runbooks, and infrastructure standards

Nice to Have

Kubernetes infrastructure (especially AI/ML and GPU integration)

KVM, VMware, OpenShift Virtualization, or similar virtualization platforms

Ansible automation

NVIDIA Base Command Manager

Slurm or HPC workload schedulers

Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)

Data Center Infrastructure Management (DCIM) tools

IPAM solutions

AWS, Azure, or hybrid cloud exposure

We Are Not Looking For

Candidates whose experience is primarily CI/CD pipeline engineering

Engineers focused mainly on Terraform, GitOps, or application delivery pipelines

Cloud-only administrators with limited bare metal, storage, or hardware experience

Professionals whose primary expertise is software development rather than infrastructure engineering

Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.

Originally posted on Himalayas

Interview prep

Walk in with sharper answers.

Use this as a quick practice sheet before you speak with the employer.

Senior
Technology & IT Administration API integration Operations Python Sales Senior level

Likely questions

  1. Tell us about work you have done that is close to the Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure) role.
  2. How would you approach your first 30 days at uvation?
  3. Which of Administration, API integration and Operations have you used recently, and what did it help you achieve?
  4. How have you led people, improved a process, or made a hard decision in a previous role?
  5. How do you handle busy days, changing priorities, or pressure at work?

Prepare before the call

  • A recent example that proves your experience with Administration, API integration and Operations.
  • One short story with a problem, your action, and the result.
  • Two examples that show the strengths listed on your CV.
  • A clear reason why this role and company interest you.
  • Your availability, preferred work style, and salary expectations.

Ask them

  • What would success look like in the first 90 days?
  • What are the main problems this hire should help solve?
  • How does the team give feedback and measure good work?
  • What does a normal working week look like for this role?
Practice line

I am interested in the Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure) role because I can bring practical experience in Administration, API integration and Operations, learn the team quickly, and contribute to the outcomes uvation needs from this hire.

Related jobs.

More roles from this company or category.