Leads customer implementations of LLM and generative-AI systems from architecture through production, building full-stack solutions and improving model quality.
About this role
This Quality Assurance Engineer joins the system quality team responsible for assuring IT service quality across business systems. The role focuses on establishing an EvalOps foundation so generative AI and AI agents can be used safely in business applications, while creating a repeatable model for evaluating and improving AI application quality.
The work covers designing, developing, and operating the EvalOps platform; evaluating Agent, RAG, and LLM applications; preparing Golden Datasets; building automated testing; analyzing and visualizing results; measuring the effects of prompt, RAG, and model-setting changes; and standardizing AI quality assurance processes. The role also includes defining evaluation metrics such as accuracy, usefulness, safety, response speed, and cost, creating dashboards and reports, and reviewing AI product release decisions.
Applicants must have experience in quality assurance for large-scale system development, or experience in system development with leadership of test design or reviews within a development process, or experience building a test-automation environment. Experience with AI, generative AI, LLMs, RAG, or AI quality evaluation is welcomed but is not a required condition stated for application.
This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.