{bc}
linkedin

AI Platform Engineer

Sharon AI, Inc
Remote, UAE
Full-time
Entry
Remote
Discovered 2 weeks ago
KubernetesSlurmGPU schedulingMLOpsCI/CDTerraform
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

KubernetesSlurmGPU scheduling
Smart Apply

Full Job Posting

The Role

Design, build, and operate the platform layer powering Sharon AI's GPU-as-a-Service offering across EMEA.

Own architecture, automation, and reliability across Kubernetes, Slurm, GPU scheduling, MLOps, model serving, and observability.

This is a hands-on senior individual contributor role reporting to the Head of Operations.

Key Responsibilities

  • Design and own AI platform architecture across EMEA GPU clusters, including Kubernetes, Slurm, and container orchestration.
  • Lead CI/CD and MLOps tooling for training, fine-tuning, and inference workloads across multiple sites.
  • Implement multi-tenant GPU scheduling, quota management, and workload isolation at scale.
  • Design model-serving infrastructure for high availability, performance, and cost efficiency.
  • Build observability covering monitoring, logging, alerting, platform health, GPU utilization, and workload performance.
  • Drive Infrastructure-as-Code and automation with Terraform and Ansible.
  • Integrate the platform with the InfiniBand/RDMA network fabric.
  • Lead complex escalations, incident response, post-incident reviews, runbooks, capacity planning, and cost optimization.
  • Mentor platform engineers and support enterprise GPUaaS customer onboarding and technical escalations.

Skills & Experience

  • 6–10+ years in platform engineering, DevOps, MLOps, or SRE, ideally in HPC, cloud, or AI/ML infrastructure.
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field.
  • Production-scale Kubernetes and GPU scheduling experience is required.
  • Experience designing CI/CD and Infrastructure-as-Code practices for platform teams.
  • Advanced Python and Bash proficiency, plus Terraform and Ansible experience.
  • Strong Linux, networking, GPU infrastructure, distributed training, and production observability fundamentals.
  • Experience with PyTorch, TensorFlow, MLflow, Kubeflow, Ray, InfiniBand/RDMA, RoCEv2, Kubernetes certifications, or cloud certifications is advantageous.
  • Awareness of EMEA regulatory and data residency considerations, including GDPR, is required.

Location and Work Style

  • The role is EMEA-based and remote-first.
  • The position supports work across multiple EMEA sites and regions.

About Sharon AI

Sharon AI builds infrastructure for artificial intelligence, including high-performance compute, cloud platforms, and large-scale AI, ML, and HPC environments.

Why Join Sharon AI

  • Own the platform architecture for a growing GPUaaS and AI neocloud business.
  • Work hands-on with GPU infrastructure, Kubernetes, Slurm, MLOps, networking, and distributed AI workloads.
  • Influence reliability, automation, capacity, cost optimization, and customer workload operations across EMEA.
  • Provide technical leadership and mentorship while remaining a hands-on senior individual contributor.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Sharon AI, Inc