{bc}
linkedin

Principal Evaluation Engineer

REACH Energy
Abu Dhabi, UAE
Contract
Mid-Senior
Onsite
Discovered Today
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan
Smart Apply

Full Job Posting

Location: Abu Dhabi

Duration: Yearly renewable contract

Overview

We are looking for a Principal Evals Engineer to own how the AI Factory knows its AI systems work. You will set the technical direction for evaluation and quality engineering across the whole portfolio — GovGPT-class assistants, retrieval pipelines, agent workflows, voice and document systems, and the backend services and data pipelines underneath them — and you will build the evaluation infrastructure that turns "it seems better" into evidence a Director General can act on

Basic Qualifications

  • 12+ years of software, quality, or ML engineering experience, with a track record of operating at staff or principal level and of designing and owning evaluation or test infrastructure at scale — not writing tests or evals within frameworks built by others.
  • Deep, first-hand experience designing evaluation frameworks for AI / LLM-powered products in production: output quality assessment, retrieval and grounding evaluation, regression detection on non-deterministic behaviour, and the practices that close the loop between measurement and product improvement.
  • Hands-on depth with LLM-as-judge evaluation: rubric design, calibration against human labels, known limitations and failure modes, and the practices that make automated grading trustworthy
  • Strong programming foundation in Python — at the depth required to design evaluation harnesses, data tooling, and shared libraries that other engineers build on — with working ability in TypeScript / JavaScript or Java.
  • Experience evaluating RAG systems and agent-based systems specifically: retrieval quality, grounding correctness, citation accuracy, tool-use validation, multi-step reasoning, and failure recovery.
  • Demonstrated ability to set technical direction across multiple teams: evaluation architecture that holds up, harnesses that other engineers actually adopt, and standards that shape how an organisation ships.
  • Strong test automation and CI/CD engineering depth: framework design across API, web, and data surfaces; GitHub Actions, GitLab CI, Jenkins, or equivalent; test selection, parallelisation, and flake reduction at scale.
  • Experience with statistics for evaluation: sampling, confidence intervals, inter-rater agreement, significance for non-deterministic systems — enough to know when a measured improvement is real.
  • Practical use of AI coding agents — Codex, Claude Code, OpenCode, or equivalent — as a core part of your engineering workflow, and a view on how evaluation makes agent-built software safe to ship.
  • Hands-on technical depth that is current. You still write the harness, still run the experiment yourself when the answer cannot be delegated, and operate close enough to the systems to be credible to the engineers building them.
  • Strong written and verbal communication — you can author the evaluation strategy, defend a model decision in a design review, and explain quality evidence to executives and government entity stakeholders without losing the technical truth

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at REACH Energy