{bc}
linkedin

Senior AI Engineer

Gateworth Group
Dubai, UAE
Full-time
Mid-Senior
Onsite
Discovered 4 weeks ago
PythonPyTorchJAXTransformer architecturesLLM internalsDistributed systems
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

PythonPyTorchJAX
Smart Apply

Full Job Posting

Overview

Gateworth Group seeks a Senior AI Engineer to optimize LLM inference performance, work on transformer architectures, and contribute to next-generation AI systems.

The role requires strong Python skills, experience with modern LLMs, and a background in distributed systems and system-level optimization.

Main Responsibilities

  • Improve and optimise LLM inference performance across distributed, multi-chip and multi-node environments
  • Apply strong understanding of transformer architectures, including dense and Mixture-of-Experts (MoE) models
  • Benchmark leading LLMs (LLaMA, Mistral, Qwen, DeepSeek) across varied hardware stacks
  • Design and implement attention-level optimisations (Flash Attention, grouped-query, sliding-window)
  • Deliver model-level optimisation including quantisation (INT8/FP8), KV-cache strategies, batching and parallelism
  • Work closely with hardware, systems and compiler teams to co-design efficient inference pipelines
  • Build and maintain benchmarking frameworks to measure latency, throughput and scaling behaviour
  • Evaluate architectural trade-offs and contribute to deployment strategies for large-scale environments
  • Stay current with research across LLM architectures, inference optimisation and performance engineering

Qualifications

  • Strong understanding of transformer architectures, LLM internals and both dense/MoE models
  • Hands-on experience with modern LLMs (LLaMA, Mistral, Qwen, DeepSeek) and attention-level optimisation
  • Practical experience with inference optimisation: quantisation (INT8/FP8), KV-cache strategies, batching, pruning and parallelism
  • Strong Python skills with PyTorch or JAX, plus experience profiling and debugging performance bottlenecks
  • Background in distributed systems, large-scale inference workloads and system-level optimisation across hardware and runtime layers
  • Ideally 8+ years in deep learning, AI systems or performance engineering, with exposure to datacenter-scale inference (e.g., vLLM) and hardware-aware optimisation

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Gateworth Group