Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms . The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies. Key Responsibilities · HPC Infrastructure Management o Operate and maintain HPC clusters based on CentOS, RHEL , and hardware platforms like HPE and NVIDIA DGX . o Ensure optimal performance, scalability, and reliability of compute resources. · Storage Administration o Manage large-scale storage systems including Dell Isilon , VAST Storage , Lustre , and GPFS . o Implement data lifecycle management and optimize storage performance for HPC workloads. · Networking o Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication. o Troubleshoot network performance issues and ensure secure connectivity. · Cluster and Job Scheduling o Administer cluster management tools such as Bright Cluster Manager , Altair Grid Manager , and IBM LSF . o Optimize job scheduling and resource allocation for diverse workloads. · Monitoring and Automation o Implement monitoring solutions using Zabbix , Grafana , and ELK Stack . o Automate provisioning and configuration using Cobbler , Chef , Ansible , and AWS ParallelCluster . · Performance Tuning & Troubleshooting o Conduct performance benchmarking and tuning for HPC workloads. o Diagnose and resolve hardware/software issues across compute, storage, and network layers. · Security & Compliance o Ensure HPC environment adheres to security best practices and compliance standards. Required Skills & Qualifications · Technical Expertise o Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms ( HPE , NVIDIA DGX ). o Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions. o Proficiency in InfiniBand networking and high-speed interconnects. o Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager). · Automation & Scripting o Expertise in Ansible , Chef , Cobbler , and scripting languages (Bash, Python). o Experience with AWS ParallelCluster or similar cloud-based HPC solutions. · Monitoring & Logging o Practical experience with Zabbix , Grafana , and ELK Stack for system health and performance monitoring. · Soft Skills o Strong problem-solving and analytical skills. o Ability to work in a fast-paced environment and lead technical teams. o Excellent communication and documentation skills. Preferred Qualifications · Exposure to AI/ML workloads on HPC clusters. · Experience with containerization (Docker, Singularity) in HPC environments. · Knowledge of security hardening for HPC systems. Education · Bachelor’s or Master’s degree in Computer Science, Engineering, or related field. #LI-LK1
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, IND
Cognizant is seeking a Service Manager/Apps Manager to govern ITSM processes and deliver stable application services across sales and marketing functions. The role covers incident, problem and change management, service
, IND
, USA
, USA
, USA
, USA
, USA
, USA
, USA
Global professional services company providing technology and consulting services.
Visit company websiteJobs and hiring trendsFull-time
Senior
Onsite
Apply faster on company sites with our extension.