Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
Uvation · Romania
About The Role
Job Overview
We are seeking a highly experienced
Senior Linux Infrastructure Engineer
with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation
AI Factory / GPU infrastructure platforms
. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is
not a DevOps-focused role
. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in
Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms
.
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
- Deep expertise in
- bare metal server deployment, architecture, provisioning, and lifecycle management
- Experience operating
- Bare Metal as a Service (BMaaS)
- platforms and large-scale infrastructure environments
Strong understanding of server hardware, including
BIOS/UEFI
- RAID controllers
- Firmware management
- iLO/iDRAC/IPMI
NICs and SmartNICs
- HBA cards
- Hardware diagnostics and troubleshooting
- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
- Experience deploying and managing
- GPU-accelerated infrastructure
- for AI/ML workloads
Understanding of NVIDIA GPU technologies including
- A100, H100, H200, B200, or equivalent GPU platforms
- NVIDIA DGX and OEM GPU servers
- GPU provisioning and lifecycle management
- GPU monitoring and performance optimization
- Knowledge of AI Factory architecture and infrastructure requirements
- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
Understanding of
- GPU resource allocation and scheduling
- Multi-GPU systems
- GPU networking requirements
- High-bandwidth, low-latency infrastructure design
Familiarity with NVIDIA ecosystem technologies such as
CUDA
NCCL
GPUDirect Storage
NVIDIA Fabric Manager
NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
Advanced Linux storage administration
LVM
XFS, EXT4
NFS
iSCSI
Fibre Channel SAN
Multipath I/O
Strong hands-on experience with
Ceph
, including
Cluster architecture
MON, OSD, MDS
RBD, CephFS, RGW
- Capacity planning
- Performance tuning
- Failure recovery
Experience with high-performance AI storage platforms such as
WEKA
VAST Data
Dell PowerScale
Pure Storage FlashBlade
NetApp
Understanding of
NVMe-over-Fabrics (NVMe-oF)
RDMA
GPUDirect Storage
- Parallel file systems
- AI data pipelines
Networking & Infrastructure
Strong networking knowledge
Bonding
VLANs
Routing
MTU optimization
DNS
DHCP
Experience with high-performance data center networking
100G/200G/400G Ethernet
RoCE
RDMA
- Spine-Leaf architectures
- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
Experience with high availability, clustering, and disaster recovery
Strong troubleshooting skills across
- Linux operating systems
- Hardware platforms
- GPU infrastructure
Networking
- Enterprise storage
- Experience supporting mission-critical production environments
- Bash and Python scripting for automation and operational efficiency
- Experience creating operational documentation, runbooks, and infrastructure standards
Nice to Have
- Kubernetes infrastructure (especially AI/ML and GPU integration)
- KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
- Ansible automation
NVIDIA Base Command Manager
- Slurm or HPC workload schedulers
- Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
- Data Center Infrastructure Management (DCIM) tools
- IPAM solutions
- AWS, Azure, or hybrid cloud exposure
We Are Not Looking For
- Candidates whose experience is primarily CI/CD pipeline engineering
- Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
- Cloud-only administrators with limited bare metal, storage, or hardware experience
- Professionals whose primary expertise is software development rather than infrastructure engineering
Ideal Candidate
Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
