Back to D L Resources Pte Ltd jobs
D

HPC High Performance Computing IT Infra Engineer

Islandwide, Singapore
Contract, Full TimeEducation and Training

Job Description

Client: Research & Education Sector

Key Responsibilities Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments

  • Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu)
  • Administer physical and virtualized server environments (x86 architecture, VMware/KVM)
  • Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation
  • Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis)
  • Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration)
  • Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules
  • Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives
  • Maintain system documentation, operational procedures, and runbooks Required Skills & Experience 1–3 years of experience in system administration, infrastructure support, or IT operations Hands-on experience with Linux/Unix systems administration and command-line environments
  • Basic exposure to HPC, distributed systems, or parallel computing environments
  • Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts)
  • Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness)
  • Scripting and automation using Bash/Shell and/or Python
  • Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack)
  • Strong troubleshooting, analytical thinking, and incident resolution skills Preferred Skills Exposure to GPU computing environments (NVIDIA GPUs, CUDA – basic awareness) Familiarity with cloud platforms and HPC workloads on cloud (AWS, Azure)
  • Awareness of container technologies (Docker) and basic DevOps practices

Project / Environment Tech Stack. (Mostly On-Prem) The project environment is primarily on-premise, with approximately 80% of the infrastructure hosted on Dell and servers, and approximately 20% involving AWS-based HPC/GPU server environments. This provides exposure to both traditional data centre infrastructure and cloud-based GPU/HPC workloads. of: 80% On-Premise HPC Infrastructure

  • Dell servers, Huawei servers, physical data centre infrastructure
  • x86 server architecture, CPU-based compute infrastructure, bare-metal servers
  • Virtualized server environments using VMware / KVM where applicable
  • Linux/Unix-based systems including RHEL, CentOS, Ubuntu, AIX
  • HPC cluster components across compute, storage, and networking layers
  • High-performance storage / parallel file systems such as IBM Spectrum Scale (GPFS), Lustre, BeeGFS; with NAS / SAN / NFS exposure where applicable
  • Cluster networking and connectivity including TCP/IP, DNS, SSH, firewall concepts, high-throughput Ethernet; InfiniBand / RDMA exposure preferred
  • HPC workload scheduling and batch processing using Slurm, PBS, or LSF
  • Monitoring and logging tools such as Nagios, Zabbix, Prometheus, Grafana, ELK Stack, or Splunk

20% AWS HPC / GPU-Related Environment

  • AWS cloud infrastructure supporting HPC / GPU-related workloads
  • AWS GPU server exposure, including NVIDIA GPU-based compute instances
  • GPU computing environment exposure, including NVIDIA CUDA, NVIDIA drivers, GPU utilization, and GPU workload monitoring where applicable
  • Cloud HPC workload support involving compute, storage, networking, and security configurations
  • Exposure to AWS HPC services or related tools such as AWS ParallelCluster, AWS Batch, EC2 GPU instances, EBS / FSx / S3 storage, VPC, IAM, and security groups
  • Container and DevOps exposure where applicable, including Docker, Kubernetes, and basic CI/CD or automation practices
  • Hybrid HPC environment exposure involving on-premise infrastructure integrated with AWS-based compute or GPU resources

About D L Resources Pte Ltd

First seen: August 25, 2026
Last updated: September 19, 2026