Principal Evals Engineer AI Agentic Systems Jobs Abu Dhabi UAE 2026

One of the region’s most ambitious Government AI programmes in Abu Dhabi — building the AI infrastructure, applications, and platforms required to operate AI-native public services at scale — is seeking a Principal Evals Engineer (AI & Agentic Systems) to own how the organisation determines whether its AI systems actually work. This is the most senior evaluation engineering role in a programme that must operate reliably across Arabic and English, within strict data sovereignty, privacy, security, and regulated infrastructure requirements. Relocation assistance is available for successful candidates and their families.

About the Opportunity — Principal AI Evals Engineer Abu Dhabi

Programme Scope: AI assistants, retrieval systems, agentic workflows, voice applications, and document intelligence — deployed across major public-sector organisations at significant scale

Evaluation Stack: Ragas, DeepEval, Promptfoo, Braintrust, LangSmith, Langfuse — LLM-as-judge patterns, golden datasets, automated grading, behavioural regression detection

Languages: Python (primary), TypeScript, Java — production-quality evaluation harnesses and shared libraries

Arabic Evaluation: Building dedicated Arabic evaluation datasets, dialect coverage, right-to-left validation, and judge calibration for bilingual AI systems

Relocation: Relocation assistance available for successful candidates and their families — international applications genuinely welcomed

 Why This Role — The Most Consequential Evals Role in the Region

Strategic Significance: This is not an internal AI team quality role — it is the principal evaluation engineering position for one of the region’s most ambitious government AI programmes, where the AI systems deployed support consequential workflows and senior decision-makers

Own the Quality Bar: You define the evidence required before model changes, prompt updates, framework migrations, or architectural changes reach production — in a programme where “it seems better” is not an acceptable quality standard

Build from Scratch: Establish the evaluation architecture, tooling, and engineering standards that teams across the entire programme will build against — a genuinely rare principal-level mandate

Technical Autonomy: Small, senior engineering teams with significant individual technical autonomy — your influence comes from the infrastructure you build and the evidence your evaluations enable

Position Overview

This Principal Evals Engineer — AI & Agentic Systems role in Abu Dhabi is a senior individual contributor position owning the evaluation strategy, architecture, platform, and standards for a government AI programme portfolio spanning AI assistants, retrieval systems, agentic workflows, voice applications, and document intelligence. You will define what gets measured and how — building evaluation harnesses, golden-set management, automated grading, LLM-as-judge infrastructure, behavioural regression detection, Arabic AI evaluation datasets, adversarial and production evaluation, CI/CD quality gates, and the evaluation culture across the programme’s engineering teams.

 Why This Role Matters: As Principal Evals Engineer for an Abu Dhabi government AI programme, you own the most fundamental quality question in applied AI — not whether the system was built as specified, but whether it actually works in the real world, across Arabic and English, under adversarial conditions, without hallucinating, without leaking data, and within the data sovereignty and regulated infrastructure requirements of a high-assurance public-sector environment, apply principal-level AI evaluation expertise — LLM-as-judge rubric design and calibration, golden dataset curation, retrieval grounding measurement, multi-step agentic reasoning evaluation, Arabic dialect coverage, and behavioural drift detection — in an environment where your evaluation evidence directly enables or blocks production deployment decisions for AI systems serving millions of people, build the shared evaluation platform and reference architecture that engineering teams across the programme use — converting the current state of “it seems better” into defensible, reproducible, statistically literate measurement evidence that supports every significant engineering decision in the portfolio, work at principal level with significant technical autonomy in a small, senior engineering team alongside Principal-level AI architects, product engineers, and forward-deployed engineers — a genuinely high-trust, high-accountability environment where your infrastructure and standards shape the programme’s quality discipline from the ground up, and relocate to Abu Dhabi with full relocation support for you and your family — joining one of the region’s most technically ambitious and consequential AI programmes at the precise moment when establishing rigorous evaluation practice is the highest-leverage investment the programme can make.

What You Own

Evaluation Strategy, Architecture & Shared Platform

  • Define what gets measured, at which layer, using which methodology, and how results feed into product and release decisions — establish the reference architecture engineering teams build against
  • Build the shared evaluation platform — evaluation harnesses, golden-set management, dataset versioning, automated grading, behavioural regression detection, and reporting infrastructure
  • Own judge-model selection, rubric design, and calibration against human labels — including understanding when automated grading cannot be trusted and human evaluation is required
  • Integrate AI evaluation with conventional test automation and CI/CD — turning production quality escapes into evaluations capable of detecting the same failure before release

Retrieval, Agents, Arabic & Multilingual Evaluation

  • Measure grounding, citation correctness, tool use, multi-step reasoning, and failure recovery across retrieval systems and agentic workflows
  • Build dedicated Arabic evaluation datasets — ensuring dialect coverage, right-to-left validation, and judge calibration that does not assume English evaluation methods transfer directly to Arabic
  • Evaluate AI assistants, document intelligence, voice applications, and conversational AI across Arabic and English in a bilingual government deployment environment
  • Ensure evaluation covers the failure modes specific to non-deterministic systems — hallucination, poor grounding, behavioural drift, prompt injection, and correlated failure patterns

Production, Adversarial & Release Evidence

  • Build online evaluation, sampling, human review, drift detection, and alerting infrastructure alongside adversarial testing for prompt injection, jailbreak resistance, and data leakage
  • Ensure model swaps, prompt changes, framework migrations, and infrastructure decisions are supported by defensible evaluation evidence before changing production behaviour
  • Define quality gates that block AI system changes from reaching production without sufficient evaluation evidence — making the quality bar measurable and enforceable
  • Lead production evaluation as an ongoing discipline — not a pre-launch exercise but a continuous monitoring and evidence-generation practice embedded in the programme

Evaluation Culture & Engineering Standards

  • Enable engineering teams to run rigorous evaluations independently — designing evaluation into systems from the beginning rather than adding it before launch
  • Build Python evaluation harnesses, frameworks, and shared libraries that other engineers across the programme depend upon and contribute to
  • Partner with Principal-level engineers responsible for AI architecture, product engineering, platform development, and forward deployment on evidence-based engineering decisions
  • Establish the statistical literacy, sampling methodology, confidence interval understanding, and inter-rater agreement standards that make observed AI improvements trustworthy

Qualifications & Requirements

Required Profile

  • Staff or Principal-level track record in evaluation, testing, or quality infrastructure used at meaningful scale (REQUIRED — years matter less than ownership)
  • Deep AI evaluation experience — output quality, retrieval, grounding, and regression detection across non-deterministic systems (REQUIRED)
  • LLM-as-judge practitioner — rubric design, judge selection, calibration against human labels, and understanding of automated grading failure modes (REQUIRED)
  • Strong Python engineer — able to build production-quality evaluation harnesses, frameworks, and shared libraries (REQUIRED)
  • Statistically literate — sampling, confidence intervals, inter-rater agreement, and statistical significance for AI improvement validation
  • Strong CI/CD engineering — framework design across API, web, and data surfaces including test selection, parallelisation, and flake management

Strongly Preferred

  • Experience with LLM evaluation platforms including Ragas, DeepEval, Promptfoo, Braintrust, LangSmith, or Langfuse
  • Arabic AI evaluation — golden datasets, dialect coverage, right-to-left validation, and judge calibration for bilingual systems
  • Adversarial and security testing for AI — structured red-teaming, prompt injection, and jailbreak resistance evaluation
  • Experience delivering technology in government, financial services, or another regulated/high-assurance environment
  • Voice and conversational AI evaluation; data pipeline and AI-service performance testing; human evaluation programme design

About This Government AI Programme Opportunity

The organisation building this government AI programme is creating the AI infrastructure, applications, and platforms required to operate AI-native public services at scale in Abu Dhabi — systems that must work reliably across Arabic and English, within strict data sovereignty, privacy, security, and regulated infrastructure requirements. Whether an AI system actually works cannot simply be a matter of opinion in this environment — it has to be measurable. This Principal Evals Engineer role owns how that measurement happens, and owns the evaluation architecture that the entire programme builds against. For a principal-level engineer who has already solved hard AI evaluation problems in production, this is one of the most consequential and technically demanding evaluation roles available in the global AI engineering community in 2026.

Career Excellence: Own AI evaluation for Abu Dhabi’s most ambitious government AI programme — where LLM-as-judge expertise, Arabic AI evaluation, agentic system measurement, and production quality infrastructure define a principal-level engineering career at the frontier of applied AI.

Who Should Apply?

  • Staff / Principal AI Evaluation Engineers: With production-scale evaluation infrastructure ownership — not just test case authoring but the harnesses, platforms, and standards that teams build against
  • LLM-as-Judge Practitioners: With rubric design, judge selection, calibration against human labels, and deep understanding of when automated grading fails and why
  • Arabic AI Evaluation Specialists: With experience building bilingual evaluation datasets, dialect-aware benchmarks, and calibrated judges for Arabic-English AI systems
  • Agentic & Retrieval System Evaluators: With grounding measurement, citation correctness, multi-step reasoning evaluation, and failure recovery detection experience for RAG and agentic workflows
  • Regulated/Government AI Engineers: Ready to relocate to Abu Dhabi (relocation fully supported) for a principal-level role in one of the region’s most ambitious and consequential government AI programmes

Recently Opening Job👇

Senior Tech Lead Jobs Abu Dhabi 2026

QA Automation Engineer Generative AI Tester Jobs Abu Dhabi UAE 2026

Leave a Comment

Select Your Degree:
Please select an option.
Select Your Experience:
Please select an option.
Select Currently Your Location:
Please select an option.
Please wait...
7
Aap ka agla page 7 second mein khulega...