Support and administer production systems used by researchers and Research Centers.
Provide technical leadership/project management for system configuration, implementation, management, and user support for both new and existing systems.
Research and recommend new functionality for HPC management and administration tools by exploring system-wide impacts, working with functional users to define current and future processes.
Expertise with architecting, operating, and debugging large scale HPC network and storage infrastructure, including MPI, NCCL, RDMA, Infiniband, and parallel file systems
Works with scientific support specialists and assigns tasks and provides oversight as appropriate to HPC engineering team to support scientific researchers who use a broad spectrum of applications from diverse fields.
Analyze results of server monitoring and implement changes to improve performance, processing, and utilization.
Propose, maintain, and enforce policies, practices and security procedures.
Provide break/fix support, setup/installation support, escalation support, and solutions support.
Collaborate closely with a variety of stakeholders, both internal and external, on all aspects of projects.
Other duties as assigned.
In Addition to the Duties Described Above
Deploy, configure, and maintain large-scale Linux-based HPC clusters comprising CPU and GPU nodes, high-speed interconnects, and parallel file systems.
Implement and optimize workload schedulers (Slurm) and job submission policies to maximize system throughput and fair-share usage.
Administer and monitor distributed storage systems (GPFS, Lustre, WekaFS, Ceph, MinIO) to ensure reliability and performance across multi-petabyte environments.
Maintain high-speed fabric and network infrastructure (Infiniband, Ethernet) to support low-latency data transfer and MPI workloads.
Support research groups in deploying, testing, and optimizing scientific applications and AI/ML workflows on shared computing resources.
Develop and maintain automation and monitoring frameworks for system provisioning, metrics collection, and alerting (Prometheus, Grafana, ELK).
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
Participate in capacity planning, hardware lifecycle management, and evaluation of new technologies in collaboration with architects and management.
Ensure security and compliance through configuration hardening, patch management, and integration with campus identity and access control systems.
Document system designs, procedures, and troubleshooting guides to support knowledge transfer and team continuity.
Contribute to a collaborative engineering culture that emphasizes service quality, innovation, and continuous improvement in research computing operations.
Bachelor’s degree.
Six years of related experience.
Additional education may substitute for required experience and additional related experience may substitute for required education beyond a high school diploma/graduation equivalent, to the extent permitted by the JHU equivalency formula.
Eight plus years of experience in high-performance computing systems administration or engineering, including experience with cluster management, workload scheduling (e.g., Slurm), and distributed or parallel storage.
Deep proficiency in Linux systems administration, configuration management (Ansible, Puppet, or Salt), performance monitoring, and tuning for HPC workloads.
Experience with high-speed interconnects (Infiniband, 100/400 Gb Ethernet) and parallel file systems (e.g., GPFS, Lustre, BeeGFS, or WekaFS).
Working knowledge of containerization and orchestration (Singularity, Docker, Kubernetes for HPC).
Ability to automate deployments and routine operations through scripting (Bash, Python).
Familiarity with data-center operations, GPU acceleration, and research software environments (e.g., CUDA, MPI, AI/ML frameworks).
Strong analytical and troubleshooting skills, with proven ability to support complex research workloads in multi-user, multi-tenant environments.
Experience collaborating with faculty and research groups to translate scientific requirements into practical and performant computing solutions.
Classified Title: Sr. HPC Systems Engineer Role/Level/Range: ATP/04/PF Starting Salary Range: $85,500 - $149,800 Annually (Commensurate w/exp.) Employee group: Full Time Schedule: Mon-Fri, 8:30am-5pm FLSA Status: Exempt Location: Johns Hopkins Bayview Department name: IT@JH Research Computing Personnel area: University Administration
About Johns Hopkins University
Higher Education27300 employeesFounded 1876
Private research university and academic medical center.