{bc}
linkedin

Senior Infrastructure Engineer

Sharon AI, Inc
Abu Dhabi, UAE
Full-time
Mid-Senior
Onsite
Discovered 1 weeks ago
Kubernetes and container orchestrationInfrastructure-as-Code with Terraform, Ansible, or similar toolsGPU infrastructure and GPU-based workloadsDistributed systemsCloud and bare-metal infrastructurePython, Go, or Bash scripting
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Kubernetes and container orchestrationInfrastructure-as-Code with Terraform, Ansible, or similar toolsGPU infrastructure and GPU-based workloads
Smart Apply

Full Job Posting

Role overview

Sharon AI is building infrastructure for high-performance GPU compute and AI workloads.

The Senior Infrastructure Engineer will build, operate, and continuously improve the infrastructure powering the GPUaaS platform.

The role supports large-scale, multi-tenant infrastructure for AI training, fine-tuning, and real-time inference.

The position is remote across EMEA and reports to the Infrastructure Engineering Team Lead.

What you will do

  • Build and operate large-scale, multi-tenant GPUaaS infrastructure.
  • Manage high-performance compute clusters with Kubernetes and/or HPC schedulers such as Slurm.
  • Maintain GPU-optimized infrastructure, node lifecycles, and cluster scaling.
  • Develop Infrastructure-as-Code with Terraform, Ansible, or similar tools.
  • Automate provisioning, scaling, and configuration.
  • Improve GPU utilization, scheduling efficiency, performance, reliability, and cost.
  • Enhance monitoring, logging, alerting, security controls, and workload isolation.
  • Support complex incidents, incident response, root cause analysis, capacity planning, and reliability improvements.

What we are looking for

  • 6–10+ years of infrastructure engineering, platform engineering, or SRE experience.
  • Deep cloud and/or bare-metal infrastructure expertise.
  • Advanced distributed systems knowledge and experience operating infrastructure at scale.
  • Strong hands-on Kubernetes and container orchestration experience.
  • Experience managing or optimizing GPU-based systems and workloads.
  • Strong Infrastructure-as-Code proficiency with Terraform, Ansible, or similar tools.
  • Strong programming and scripting skills in Python, Go, and/or Bash.
  • Experience with high-performance networking and storage systems.
  • Observability, performance tuning, cost optimization, security hardening, and technical leadership capabilities.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.

Preferred experience

  • GPUaaS, IaaS, or neocloud platform experience.
  • AI/ML workloads and frameworks such as PyTorch or TensorFlow.
  • HPC environments and schedulers including Slurm or Ray.
  • GPU technologies including CUDA, NCCL, or MIG.
  • Kubeflow, Airflow, or similar orchestration tools.
  • High-performance networking technologies such as RDMA or InfiniBand.
  • Relevant cloud or Kubernetes certifications.

Why join Sharon AI

  • Remote-first working environment across EMEA.
  • Work with high-performance GPU infrastructure and AI/ML workloads.
  • Work on large-scale distributed systems and production-grade platforms.
  • Collaborate with platform, ML, and product teams on next-generation AI infrastructure.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Sharon AI, Inc