Principal Evals Engineer - AI & Agentic Systems
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
The Principal Evals Engineer owns how the organisation measures whether AI systems work across assistants, retrieval systems, agentic workflows, voice applications, and document intelligence.
This is a hands-on individual contributor role focused on evaluation infrastructure, quality standards, and evidence for engineering decisions.
Key Skills for This Role
Full Job Posting
Organisation Context
The organisation is building government AI infrastructure, applications, and platforms for AI-native public services at scale.
Systems must operate reliably across Arabic and English and meet strict data sovereignty, privacy, security, and regulated-infrastructure requirements.
Role Overview
The Principal Evals Engineer owns how the organisation measures whether AI systems work across assistants, retrieval systems, agentic workflows, voice applications, and document intelligence.
This is a hands-on individual contributor role focused on evaluation infrastructure, quality standards, and evidence for engineering decisions.
Evaluation Strategy and Platform
- Define evaluation strategy, architecture, methodologies, and release decision criteria.
- Build evaluation harnesses, golden-set management, dataset versioning, automated grading, regression detection, and reporting infrastructure.
- Own judge-model selection, rubric design, and calibration against human labels.
AI and Multilingual Evaluation
- Measure grounding, citation correctness, tool use, multi-step reasoning, and failure recovery.
- Build Arabic evaluation datasets with dialect coverage, right-to-left validation, and calibrated judges.
- Evaluate non-deterministic systems where hallucination, poor grounding, behavioural drift, and prompt injection require more than conventional assertions.
Production and Release Quality
- Build online evaluation, sampling, human review, drift detection, alerting, and adversarial testing.
- Integrate AI evaluation with test automation and CI/CD quality gates.
- Use defensible evidence to support model swaps, prompt changes, framework migrations, and infrastructure decisions.
Evaluation Culture
- Enable engineering teams to run rigorous evaluations independently and design evaluation into systems from the beginning.
- Turn production quality escapes into evaluations that can detect the same failures before release.
Core Qualifications
- Staff- or Principal-level track record building evaluation, testing, or quality infrastructure at meaningful scale.
- Deep AI evaluation experience across output quality, retrieval, grounding, and regression detection.
- Experience with LLM-as-judge methods, rubric design, judge selection, human-label calibration, and automated-grading failure modes.
- Strong Python engineering skills for production evaluation harnesses, frameworks, and shared libraries.
- Statistical literacy covering sampling, confidence intervals, inter-rater agreement, and significance.
- Strong CI/CD engineering knowledge across API, web, and data surfaces.
Preferred Qualifications
- Experience with Ragas, DeepEval, Promptfoo, Braintrust, LangSmith, or Langfuse is strongly preferred.
- Arabic AI evaluation, adversarial AI security testing, voice evaluation, human evaluation programs, or regulated-environment delivery is preferred.
- Working knowledge of TypeScript or Java is valuable.
Technology Environment
The environment includes Python, TypeScript, Java, LangChain, LangGraph, Microsoft Agent Framework, pgvector, Qdrant, and evaluation platforms.
Testing and infrastructure technologies include Playwright, Cypress, Appium, Detox, Docker, Kubernetes, Azure, and CI/CD observability tools.
The stack also includes Great Expectations, Soda, k6, JMeter, Locust, GitHub Actions, GitLab CI, Grafana, and behavioural regression dashboards.
The technology stack reflects current engineering preferences rather than a rigid mandated standard.
Role Boundaries
This is not a traditional line-management role, a conventional QA leadership position, or a research-only role.
Success is measured by reliable evidence supporting engineering and product decisions and by identifying failures before users encounter them.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at oryxsearch.io
Principal Forward Deployed Engineer - AI & Agentic Systems
Abu Dhabi, UAE
The employer is seeking a principal forward deployed engineer to redesign government workflows and build agentic AI systems from discovery through production and adoption. The role requires hands-on production engineerin
Infrastructure Architect (Azure)
Abu Dhabi, UAE
A regulated Abu Dhabi organization is seeking an experienced Infrastructure Architect to design, build, govern, and operate enterprise Azure foundations. The hands-on role covers networking, security, AKS, observability,
AI Engineer / Forward Deployed Engineer
Abu Dhabi, UAE
The employer is seeking an AI Engineer / Forward Deployed Engineer to design, build, and deploy agentic systems that transform government services and operational workflows. The role combines AI engineering, product deve
Lead Data Scientist (Credit & Lending)
, UAE
The employer is seeking a lead data scientist to own credit decisioning and risk model development for fintech lending products. The role requires production credit-risk modelling, alternative-data expertise, advanced Py
Lead Applied AI Engineer (Research & Development)
, UAE
The employer is seeking a hands-on Lead Applied AI Engineer to drive research, prototyping, and production delivery for AI-powered financial infrastructure products. The role requires deep AI/ML expertise, strong Python
Lead Python Engineer (AI Automation)
Abu Dhabi, UAE
The employer is seeking a hands-on Lead Backend Engineer to own major platform areas while contributing approximately 70% engineering and 30% technical leadership. The role requires deep Python and backend expertise, dis
Lead Python Engineer (AI Automation)
, UAE
Oryxsearch.io is seeking a Lead Python Engineer (AI Automation) to take ownership of major areas of their platform and help scale both the technology and engineering team. The role involves designing and building scalabl
Principal Forward Deployed Engineer - AI & Agentic Systems
Abu Dhabi, UAE
Infrastructure Architect (Azure)
Abu Dhabi, UAE
AI Engineer / Forward Deployed Engineer
Abu Dhabi, UAE
Lead Data Scientist (Credit & Lending)
, UAE
Lead Applied AI Engineer (Research & Development)
, UAE
Lead Python Engineer (AI Automation)
Abu Dhabi, UAE
Lead Python Engineer (AI Automation)
, UAE