{bc}
linkedin

Principal Evals Engineer - AI & Agentic Systems

oryxsearch.io
Abu Dhabi, UAE
Full-time
Mid-Senior
Onsite
Discovered Yesterday
AI evaluationEvaluation architecturePythonLLM-as-judgeRubric design and judge calibrationGolden datasets
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

AI evaluationEvaluation architecturePython
Smart Apply

Full Job Posting

Organisation Context

The organisation is building government AI infrastructure, applications, and platforms for AI-native public services at scale.

Systems must operate reliably across Arabic and English and meet strict data sovereignty, privacy, security, and regulated-infrastructure requirements.

Role Overview

The Principal Evals Engineer owns how the organisation measures whether AI systems work across assistants, retrieval systems, agentic workflows, voice applications, and document intelligence.

This is a hands-on individual contributor role focused on evaluation infrastructure, quality standards, and evidence for engineering decisions.

Evaluation Strategy and Platform

  • Define evaluation strategy, architecture, methodologies, and release decision criteria.
  • Build evaluation harnesses, golden-set management, dataset versioning, automated grading, regression detection, and reporting infrastructure.
  • Own judge-model selection, rubric design, and calibration against human labels.

AI and Multilingual Evaluation

  • Measure grounding, citation correctness, tool use, multi-step reasoning, and failure recovery.
  • Build Arabic evaluation datasets with dialect coverage, right-to-left validation, and calibrated judges.
  • Evaluate non-deterministic systems where hallucination, poor grounding, behavioural drift, and prompt injection require more than conventional assertions.

Production and Release Quality

  • Build online evaluation, sampling, human review, drift detection, alerting, and adversarial testing.
  • Integrate AI evaluation with test automation and CI/CD quality gates.
  • Use defensible evidence to support model swaps, prompt changes, framework migrations, and infrastructure decisions.

Evaluation Culture

  • Enable engineering teams to run rigorous evaluations independently and design evaluation into systems from the beginning.
  • Turn production quality escapes into evaluations that can detect the same failures before release.

Core Qualifications

  • Staff- or Principal-level track record building evaluation, testing, or quality infrastructure at meaningful scale.
  • Deep AI evaluation experience across output quality, retrieval, grounding, and regression detection.
  • Experience with LLM-as-judge methods, rubric design, judge selection, human-label calibration, and automated-grading failure modes.
  • Strong Python engineering skills for production evaluation harnesses, frameworks, and shared libraries.
  • Statistical literacy covering sampling, confidence intervals, inter-rater agreement, and significance.
  • Strong CI/CD engineering knowledge across API, web, and data surfaces.

Preferred Qualifications

  • Experience with Ragas, DeepEval, Promptfoo, Braintrust, LangSmith, or Langfuse is strongly preferred.
  • Arabic AI evaluation, adversarial AI security testing, voice evaluation, human evaluation programs, or regulated-environment delivery is preferred.
  • Working knowledge of TypeScript or Java is valuable.

Technology Environment

The environment includes Python, TypeScript, Java, LangChain, LangGraph, Microsoft Agent Framework, pgvector, Qdrant, and evaluation platforms.

Testing and infrastructure technologies include Playwright, Cypress, Appium, Detox, Docker, Kubernetes, Azure, and CI/CD observability tools.

The stack also includes Great Expectations, Soda, k6, JMeter, Locust, GitHub Actions, GitLab CI, Grafana, and behavioural regression dashboards.

The technology stack reflects current engineering preferences rather than a rigid mandated standard.

Role Boundaries

This is not a traditional line-management role, a conventional QA leadership position, or a research-only role.

Success is measured by reliable evidence supporting engineering and product decisions and by identifying failures before users encounter them.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at oryxsearch.io