Job is active

Hardware Engineer

4 500 – 7 500 $
Full time1–3 yearsOn-siteSingapore

Location

FRASER STREET, DUO TOWER

Open in maps

About this job

We are seeking an experienced AI Hardware Engineer to support the design, deployment, validation, and troubleshooting of AI training clusters, GPU servers, networking, and storage infrastructure. The ideal candidate should have strong expertise in server hardware, GPU platforms, high-speed networking, and data center infrastructure to support large-scale AI/HPC environments.

Key Responsibilities

AI Server Hardware Management

  • Deploy, validate, and maintain AI GPU servers;
  • Perform hardware diagnostics and component replacement;
  • Analyze system logs, BMC logs, and hardware alerts;
  • Manage server hardware lifecycle.

GPU Platform Support

  • Deploy and validate NVIDIA GPU platforms;
  • Troubleshoot GPU-related;
  • Perform GPU benchmarking and stress testing;
  • Support CUDA, NCCL, and GPU fabric troubleshooting.

AI Cluster Deployment & Validation

  • Participate in AI/HPC cluster deployment;
  • Execute cluster hardware qualification testing;
  • Produce validation reports and documentation.

Network & Storage Support

  • Configure and maintain high-speed networking:
  • Support distributed storage systems:
  • Assist with performance analysis and troubleshooting.

Automation & Tool Development

  • Develop automation scripts for:
  • Hardware health checks

    Cluster validation

    Deployment automation

    Log collection

  • Build tools for testing and operations.
  • Good communication, teamwork, and ownership mindset.
  • Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.

Required Qualifications

Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.

Hardware

  • Strong knowledge of x86 server architecture;
  • Familiar with Intel, AMD, and NVIDIA Grace CPU platforms;
  • Experience with:
  • HGX
  • DGX
  • GB200 NVL72
  • GB300 NVL72
  • Knowledge of BMC/IPMI management.

GPU & AI Platform

  • Experience with NVIDIA GPU products H100, H200, B200, B300
  • Familiar with: CUDA ,NCCL ,NV Link ,NV Switch and GPU Direct RDMA

Linux

  • Strong Linux administration skills (Ubuntu, Rocky Linux);
  • Proficient in: Shell ,Python , Bash
  • Capable of independent troubleshooting.

Networking: Strong understanding of: TCP/IP , VLAN , BGP ,OSPF ,RDMA , InfiniBand and RoCE

Preferred Qualities

  • Experience operating AI training clusters; Kubernetes experience; Slurm administration
  • PXE deployment experience; GPU Fabric Manager expertise;
  • Experience with hyperscale AI datacenter deployments.

Market insight

25% above median
4 800 $

Based on 75 102 offers with salary for this country

Full salary breakdown

Similar jobs

RUNSUN SERVICE PTE. LTD. Singapore ·

Compliance Officer

5 000 – 5 500 $
Full time3–6 yearsOn-siteSingapore
RUNSUN SERVICE PTE. LTD. Singapore ·

Senior Compliance Officer

6 000 – 7 000 $
Full time3–6 yearsOn-siteSingapore
RUNSUN SERVICE PTE. LTD. Singapore ·

Network Operations Engineer

4 000 – 6 500 $
Full time1–3 yearsOn-siteSingapore
RUNSUN SERVICE PTE. LTD. Singapore ·

System Engineer

5 000 – 7 500 $
Full time1–3 yearsOn-siteSingapore

Frequently asked questions

What salary can I expect?

The employer lists 4 500 – 7 500 $ for this role at RUNSUN SERVICE PTE. LTD. in Singapore. For comparison, the local market median is about 4 800 $ based on 75 102 similar offers.

How do I apply for this job?

Open the original source page and contact the employer there. Finder never charges job seekers.

Are these jobs up to date?

Yes. Finder regularly refreshes vacancies from public sources and removes closed offers.

Where can I see employment type and work format?

Key conditions are shown above the description. You can also open related listings for RUNSUN SERVICE PTE. LTD. and Singapore.

Apply