Systems Limited logo

Forward Deployed Engineer - AI Assurance

Systems Limited

On-site๐Ÿ‡ธ๐Ÿ‡ฆSaudi ArabiaseniorPosted 2d ago

Visa & sponsorship

  • Employers in Saudi Arabia sponsor the residence visa by default, and nothing in the posting says otherwise.

Job description

ABOUT

:

Owns quality for AI-native applications โ€” functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).

KEY RESPONSIBILITIES

  • Build test plans and automation for AI-native application features (functional + AI-specific)

  • Design evaluation harnesses for model/agent outputs โ€” accuracy, consistency, hallucination rate

  • Run regression testing across model/prompt/config changes to catch silent quality drift

  • Red-team AI features for edge cases and adversarial inputs where relevant

  • Build automated eval pipelines integrated into CI/CD

  • Partner with AI Architects to define testability requirements before build starts

  • Own the quality gate before any AI feature ships to production

  • Communicate quality risk to delivery leadership in terms they can act on

  • Train delivery teams on AI-specific testing practices

  • Own the evals framework for the practice โ€” golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case

  • Define eval acceptance thresholds per engagement and gate releases on them

  • Build eval engineering tooling โ€” dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read

  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets

REQUIREMENTS & SKILLS

  • 5โ€“9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically

  • Strong test automation skills (Python-based frameworks, CI/CD integration)

  • Understands AI-specific failure modes โ€” hallucination, bias, drift, non-determinism โ€” and designs tests for them

  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results

  • Familiarity with red-teaming methodologies for AI systems

  • Clear, assertive communicator โ€” willing to block a release over a quality concern

  • Detail-oriented and methodical under delivery-timeline pressure

  • Collaborative but independent โ€” doesn't rubber-stamp under delivery pressure

  • Explains quality risk in business-impact terms, not just technical jargon

  • Hands-on evals engineering โ€” builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations

  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review

  • Understands RAG and agent eval metrics โ€” groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency

  • Experience wiring evals and drift monitoring into CI/CD and production observability