YR23- AI Infrastructure Engineer |HPC/DevOps|CKA/CKAD Cert|Min 3 Years Experience
About this job
AI Infrastructure Engineer
5 days, Mon - Fri 8.30am to 5.30pm
Salary: $5,000 to $7,000
Location: Kaki Bukit Ave 1, Singapore 417938
Job scopes:
Compute & Cluster Management
- Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).
- Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
- Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
High-Performance Networking & Storage
- Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
- Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.
Automation & Infrastructure as Code (IaC)
- Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
- Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
Operations, Observability & Performance
- Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
- Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
- Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Requirements:
- Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
- Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
- Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
- High-Speed Networking: Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.
- Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
- Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
- Bachelor’s Degree in Computer Science, Information Technology, Computer
Engineering, or equivalent practical experience. - 3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
- Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified
Associate/Professional, AWS/Azure/GCP Solutions Architect).
WhatsApp me your resume for quicker processing! ☺️
WhatsApp: https://wa.me/6584067823 (Shiro)
SHEE YOKE RU|R26161758 | The Supreme Hr Advisory Pte Ltd EA No: 14C7279
Market insight
25% above medianExplore related jobs
Similar jobs
Telemarketer [5 DAYS / JURONG EAST] (KCKC)
AI Infrastructure Engineer - LCYL
Head of Security Operations & Delivery - LCYL
AI Security & Automation Lead - LCYL
Infrastructure Support Engineer (12-Month Contract) - LCYL
Associate Support Engineer (12-Month Contract) - LCYL
Frequently asked questions
What salary can I expect?
The employer lists 5 000 – 7 000 $ for this role at THE SUPREME HR ADVISORY PTE. LTD. in Singapore. For comparison, the local market median is about 4 800 $ based on 75 102 similar offers.
How do I apply for this job?
Open the original source page and contact the employer there. Finder never charges job seekers.
Are these jobs up to date?
Yes. Finder regularly refreshes vacancies from public sources and removes closed offers.
Where can I see employment type and work format?
Key conditions are shown above the description. You can also open related listings for THE SUPREME HR ADVISORY PTE. LTD. and Singapore.
