
Forward Deployed Engineer - AI Assurance
Systems Limited
Visa & sponsorship
- Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.
Job description
ABOUT
:
Owns quality for AI-native applications โ functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).
KEY RESPONSIBILITIES
-
Build test plans and automation for AI-native application features (functional + AI-specific)
-
Design evaluation harnesses for model/agent outputs โ accuracy, consistency, hallucination rate
-
Run regression testing across model/prompt/config changes to catch silent quality drift
-
Red-team AI features for edge cases and adversarial inputs where relevant
-
Build automated eval pipelines integrated into CI/CD
-
Partner with AI Architects to define testability requirements before build starts
-
Own the quality gate before any AI feature ships to production
-
Communicate quality risk to delivery leadership in terms they can act on
-
Train delivery teams on AI-specific testing practices
-
Own the evals framework for the practice โ golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
-
Define eval acceptance thresholds per engagement and gate releases on them
-
Build eval engineering tooling โ dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
-
Instrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
-
5โ9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
-
Strong test automation skills (Python-based frameworks, CI/CD integration)
-
Understands AI-specific failure modes โ hallucination, bias, drift, non-determinism โ and designs tests for them
-
Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
-
Familiarity with red-teaming methodologies for AI systems
-
Clear, assertive communicator โ willing to block a release over a quality concern
-
Detail-oriented and methodical under delivery-timeline pressure
-
Collaborative but independent โ doesn't rubber-stamp under delivery pressure
-
Explains quality risk in business-impact terms, not just technical jargon
-
Hands-on evals engineering โ builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
-
Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
-
Understands RAG and agent eval metrics โ groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
-
Experience wiring evals and drift monitoring into CI/CD and production observability