Back to D L Resources Pte Ltd jobs
D
HPC High Performance Computing IT Infra Engineer
Islandwide, Singapore
Contract, Full TimeEducation and TrainingJob Description
Client: Research & Education Sector
Key Responsibilities Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments
- Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu)
- Administer physical and virtualized server environments (x86 architecture, VMware/KVM)
- Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation
- Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis)
- Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration)
- Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules
- Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives
- Maintain system documentation, operational procedures, and runbooks Required Skills & Experience 1–3 years of experience in system administration, infrastructure support, or IT operations Hands-on experience with Linux/Unix systems administration and command-line environments
- Basic exposure to HPC, distributed systems, or parallel computing environments
- Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts)
- Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness)
- Scripting and automation using Bash/Shell and/or Python
- Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack)
- Strong troubleshooting, analytical thinking, and incident resolution skills Preferred Skills Exposure to GPU computing environments (NVIDIA GPUs, CUDA – basic awareness) Familiarity with cloud platforms and HPC workloads on cloud (AWS, Azure)
- Awareness of container technologies (Docker) and basic DevOps practices
Project / Environment Tech Stack. (Mostly On-Prem) The project environment is primarily on-premise, with approximately 80% of the infrastructure hosted on Dell and servers, and approximately 20% involving AWS-based HPC/GPU server environments. This provides exposure to both traditional data centre infrastructure and cloud-based GPU/HPC workloads. of: 80% On-Premise HPC Infrastructure
- Dell servers, Huawei servers, physical data centre infrastructure
- x86 server architecture, CPU-based compute infrastructure, bare-metal servers
- Virtualized server environments using VMware / KVM where applicable
- Linux/Unix-based systems including RHEL, CentOS, Ubuntu, AIX
- HPC cluster components across compute, storage, and networking layers
- High-performance storage / parallel file systems such as IBM Spectrum Scale (GPFS), Lustre, BeeGFS; with NAS / SAN / NFS exposure where applicable
- Cluster networking and connectivity including TCP/IP, DNS, SSH, firewall concepts, high-throughput Ethernet; InfiniBand / RDMA exposure preferred
- HPC workload scheduling and batch processing using Slurm, PBS, or LSF
- Monitoring and logging tools such as Nagios, Zabbix, Prometheus, Grafana, ELK Stack, or Splunk
20% AWS HPC / GPU-Related Environment
- AWS cloud infrastructure supporting HPC / GPU-related workloads
- AWS GPU server exposure, including NVIDIA GPU-based compute instances
- GPU computing environment exposure, including NVIDIA CUDA, NVIDIA drivers, GPU utilization, and GPU workload monitoring where applicable
- Cloud HPC workload support involving compute, storage, networking, and security configurations
- Exposure to AWS HPC services or related tools such as AWS ParallelCluster, AWS Batch, EC2 GPU instances, EBS / FSx / S3 storage, VPC, IAM, and security groups
- Container and DevOps exposure where applicable, including Docker, Kubernetes, and basic CI/CD or automation practices
- Hybrid HPC environment exposure involving on-premise infrastructure integrated with AWS-based compute or GPU resources
About D L Resources Pte Ltd
First seen: August 25, 2026
Last updated: September 19, 2026