{bc}
icims

HPC Infrastructure & Cluster Engineer

Bridge Core
Springfield, USA
Full-time
Senior · 5+ years experience
Onsite
USD 148000-179000 yearly / year
Discovered 1 weeks ago
LinuxRun:AIInfiniBandRed Hat OpenShiftKubernetesSLURM
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

LinuxRun:AIInfiniBand
Smart Apply

Full Job Posting

Key Responsibilities

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
  • Clearance Required: Active TS clearance (with SCI Eligibility) and eligibility to obtain CI Poly
  • Education/Experience:
  • Requires Bachelor's degree
  • 5+years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand)
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM)
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues
  • What is ideal?
  • Familiarity with parallel file systems and high-throughput storage architecture.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies
  • Intelligence Community Experience preferred
  • Recognizing great achievements do not go unnoticed by bcore through service anniversaries, spot awards, and employee referral bonuses
  • You’ll join a growing organization of passionate, top-shelf, IT engineering professionals with extensive experience in actively developing the technology revolution in the Intelligence community
  • The expected salary range within the Washington, DC metropolitan area is: $148,000 - $179,000. Final compensation is unique to each individual and will be determined based on factors such as experience, education, geographic location, and contractual requirements. This is not a guarantee.
  • Benefits include Health/Dental/Vision, 401(k), Paid Time Off, STD/LTD/Life Insurance/Voluntary Life Insurance, Stipends, Referral Bonuses, and more.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Bridge Core