Senior Infrastructure Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Sharon AI is building infrastructure for high-performance GPU compute and AI workloads.
The Senior Infrastructure Engineer will build, operate, and continuously improve the infrastructure powering the GPUaaS platform.
The role supports large-scale, multi-tenant infrastructure for AI training, fine-tuning, and real-time inference.
The position is remote across EMEA and reports to the Infrastructure Engineering Team Lead.
Key Skills for This Role
Full Job Posting
Role overview
Sharon AI is building infrastructure for high-performance GPU compute and AI workloads.
The Senior Infrastructure Engineer will build, operate, and continuously improve the infrastructure powering the GPUaaS platform.
The role supports large-scale, multi-tenant infrastructure for AI training, fine-tuning, and real-time inference.
The position is remote across EMEA and reports to the Infrastructure Engineering Team Lead.
What you will do
- Build and operate large-scale, multi-tenant GPUaaS infrastructure.
- Manage high-performance compute clusters with Kubernetes and/or HPC schedulers such as Slurm.
- Maintain GPU-optimized infrastructure, node lifecycles, and cluster scaling.
- Develop Infrastructure-as-Code with Terraform, Ansible, or similar tools.
- Automate provisioning, scaling, and configuration.
- Improve GPU utilization, scheduling efficiency, performance, reliability, and cost.
- Enhance monitoring, logging, alerting, security controls, and workload isolation.
- Support complex incidents, incident response, root cause analysis, capacity planning, and reliability improvements.
What we are looking for
- 6–10+ years of infrastructure engineering, platform engineering, or SRE experience.
- Deep cloud and/or bare-metal infrastructure expertise.
- Advanced distributed systems knowledge and experience operating infrastructure at scale.
- Strong hands-on Kubernetes and container orchestration experience.
- Experience managing or optimizing GPU-based systems and workloads.
- Strong Infrastructure-as-Code proficiency with Terraform, Ansible, or similar tools.
- Strong programming and scripting skills in Python, Go, and/or Bash.
- Experience with high-performance networking and storage systems.
- Observability, performance tuning, cost optimization, security hardening, and technical leadership capabilities.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
Preferred experience
- GPUaaS, IaaS, or neocloud platform experience.
- AI/ML workloads and frameworks such as PyTorch or TensorFlow.
- HPC environments and schedulers including Slurm or Ray.
- GPU technologies including CUDA, NCCL, or MIG.
- Kubeflow, Airflow, or similar orchestration tools.
- High-performance networking technologies such as RDMA or InfiniBand.
- Relevant cloud or Kubernetes certifications.
Why join Sharon AI
- Remote-first working environment across EMEA.
- Work with high-performance GPU infrastructure and AI/ML workloads.
- Work on large-scale distributed systems and production-grade platforms.
- Collaborate with platform, ML, and product teams on next-generation AI infrastructure.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Sharon AI, Inc
Senior Network Engineer
, UAE
Sharon AI is seeking a Senior Network Engineer to own network architecture, delivery, and operations across EMEA data centre and AI compute environments. The role requires advanced routing and security expertise, inciden
AI Platform Engineer
, UAE
Sharon AI is seeking a senior AI Platform Engineer to design, automate, and operate the platform powering its multi-site GPUaaS offering across EMEA. The role requires strong Kubernetes, GPU scheduling, MLOps, Infrastruc