Base Career helps you apply smarter for this job.
Key skills for this role
AI systems Build, fine-tune, and evaluate LLM systems for domain-specific tasks (QLoRA / PEFT on open-weight models such as Llama-3 and Mistral).
Design reproducible evaluation harnesses and A/B test frameworks with tracked metrics: task success rate, safety rate, and latency distributions (p50/p95).
Architect multi-agent and RAG systems (LangGraph, FastAPI, vector databases) from prototype through production.
Implement safety guardrails — input/output validation, allowlist/denylist policies, and controls that reduce invalid or high-risk model actions.
Translate business use cases into deployable prototypes with measurable acceptance criteria, and demo them to stakeholders.
Platform & infrastructure Design and operate cloud infrastructure and MLOps workspaces (Azure, OCI, or GCP) for AI workloads on Kubernetes and containerized runtimes.
Build CI/CD pipelines and GitOps-based release promotion (Argo CD) across development, test, and production environments.
Implement end-to-end observability (Azure Monitor, Application Insights, ELK) with defined detection and response targets.
Apply network and perimeter security baselines (FW/WAF), automated code quality and SCA scanning (SonarQube, Black Duck), and gated pipelines.
Own disaster recovery design — automated backups, failover, and documented RTO/RPO commitments.
Engineering leadership Lead and mentor a cloud/AI operations team; define monitoring, incident response, and release governance practices with clear uptime and MTTR targets.
Standardize SDLC practices — branching strategy, PR governance, release management, delivery reporting — to improve lead time and deployment frequency.
Consolidate engineering tooling and workflows; drive migrations and platform standardization where fragmentation slows delivery.
Produce handover documentation and runbooks that make systems auditable and operationally transferable.
Support vendor and licensing negotiations for cloud enterprise agreements.
Job requirements Required Qualifications Bachelor's degree in Software Engineering, Computer Science, or a related field. 6–8+ years in software, DevOps, or platform engineering, including at least 2 years in an applied AI or ML engineering capacity.
Proven delivery of production AI/LLM systems — not only research or notebook-stage work.
Strong Python; comfortable with Bash and YAML.
Deep hands-on experience with Kubernetes, Docker/Podman, and Terraform.
Production experience with at least one major cloud (Azure preferred; OCI or GCP acceptable).
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
Riyadh, KSA
The employer is seeking a DevOps Engineer to build and operate reliable cloud and on-premises platforms. The role focuses on CI/CD, infrastructure automation, containers, monitoring, security, incident response, and cont
Riyadh, KSA
Saudi AZM is seeking a Senior IT Project Manager to lead technology projects from initiation through closure. The role requires 5+ years of IT project management experience, end-to-end delivery capability, strong stakeho
Riyadh, KSA
Riyadh, KSA
Riyadh, KSA
Riyadh, KSA
Riyadh, KSA
Riyadh, KSA
Riyadh, KSA
Riyadh, KSA
Demonstrated ownership of CI/CD at scale (Azure DevOps, GitHub Actions) and GitOps release models.
AI systems
Build, fine-tune, and evaluate LLM systems for domain-specific tasks (QLoRA / PEFT on open-weight models such as Llama-3 and Mistral).
Design reproducible evaluation harnesses and A/B test frameworks with tracked metrics: task success rate, safety rate, and latency distributions (p50/p95).
Architect multi-agent and RAG systems (LangGraph, FastAPI, vector databases) from prototype through production.
Implement safety guardrails — input/output validation, allowlist/denylist policies, and controls that reduce invalid or high-risk model actions.
Translate business use cases into deployable prototypes with measurable acceptance criteria, and demo them to stakeholders.
Platform & infrastructure
Design and operate cloud infrastructure and MLOps workspaces (Azure, OCI, or GCP) for AI workloads on Kubernetes and containerized runtimes.
Build CI/CD pipelines and GitOps-based release promotion (Argo CD) across development, test, and production environments.
Implement end-to-end observability (Azure Monitor, Application Insights, ELK) with defined detection and response targets.
Apply network and perimeter security baselines (FW/WAF), automated code quality and SCA scanning (SonarQube, Black Duck), and gated pipelines.
Own disaster recovery design — automated backups, failover, and documented RTO/RPO commitments.
Engineering leadership
Lead and mentor a cloud/AI operations team; define monitoring, incident response, and release governance practices with clear uptime and MTTR targets.
Standardize SDLC practices — branching strategy, PR governance, release management, delivery reporting — to improve lead time and deployment frequency.
Consolidate engineering tooling and workflows; drive migrations and platform standardization where fragmentation slows delivery.
Produce handover documentation and runbooks that make systems auditable and operationally transferable.
Support vendor and licensing negotiations for cloud enterprise agreements.
Saudi publicly listed digital technology company serving government and private organizations with business, software, and technology solutions.
Visit company websiteJobs and hiring trendsFull-time
Senior · 6+ years experience
Onsite
Apply faster on company sites with our extension.