{bc}
linkedin

HPC System Engineer

EDGE
Abu Dhabi, UAE
Full-time
Mid-Senior
Onsite
Discovered 1 weeks ago
Linux administrationHPC cluster administrationHPC schedulersSlurmPBSLSF
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Linux administrationHPC cluster administrationHPC schedulers
Smart Apply

Full Job Posting

Role overview

Administer and optimize a high-performance computing environment supporting Computational Fluid Dynamics and Finite Element Analysis workloads.

Own the full HPC operations lifecycle, including infrastructure, schedulers, user support, performance optimization, and capacity planning.

Infrastructure and scheduling

  • Maintain compute nodes, login nodes, heterogeneous nodes, NAS storage servers, and InfiniBand interconnects.
  • Configure and tune Slurm, PBS, LSF, or Grid Engine schedulers.
  • Implement fair-share scheduling, job priorities, QoS limits, reservations, and preemption policies.
  • Recommend CPU and GPU architectures for engineering simulations.

Security and software

  • Manage accounts, permissions, storage quotas, secure SSH access, MFA, patches, and vulnerability fixes.
  • Maintain compliance with IT governance, data-security, and audit requirements.
  • Install and manage ANSYS, Siemens, NASTRAN, environment modules, compilers, MPI libraries, and drivers.

Performance and support

  • Monitor utilization, queue wait times, job failures, hardware health, and cluster performance.
  • Resolve network congestion, I/O bottlenecks, memory pressure, job crashes, solver errors, and environment issues.
  • Support batch scripting and job optimization, and provide user training and documentation.

Storage and capacity planning

  • Enforce quotas, purge policies, retention rules, backups, archival systems, and RAID configurations.
  • Maintain Lustre, GPFS, or BeeGFS parallel file systems for large CFD/FEA datasets.
  • Plan expansion for additional cores, GPUs, memory-heavy nodes, and faster interconnects.

Basic qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, Physics, or a related field.
  • 3–7 years of experience administering Linux-based HPC clusters.
  • Strong knowledge of HPC schedulers such as Slurm, PBS, or LSF.
  • Experience with MPI, OpenMP, and distributed computing.
  • Familiarity with CFD/FEA solvers and engineering workflows.
  • Proficiency in Bash or Python scripting.
  • Experience with high-speed networking and parallel file systems.

Preferred qualifications

  • Experience with GPU-accelerated workloads and CUDA-enabled solvers.
  • Knowledge of Singularity or Apptainer container technologies.
  • Experience with Grafana, Prometheus, or XDMoD monitoring tools.
  • Background supporting engineering or scientific computing environments.
  • Understanding of performance profiling and benchmarking.

Additional role requirement

  • English proficiency is required for all roles unless otherwise stated in the posting.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at EDGE