AI Quality & Evaluation Engineer

Woven by Toyota

QA

GenAI

Flexible

About this role

This role is the first dedicated QA engineer for Woven by Toyota's AI products, focused on building sustainable quality practices for LLM-based systems. Responsibilities include defining quality dimensions such as accuracy, consistency, safety, fairness, and UX; conducting scenario-based and exploratory testing plus red teaming; analyzing LLM outputs to identify behavioral trends and failure patterns; and designing evaluation processes based on user feedback such as logs, ratings, and inquiries. The role also involves standardizing and automating evaluation workflows, collecting and aggregating evaluation data, establishing continuous quality checks, and building and operating quality feedback loops in collaboration with product development, MLOps, and data engineering teams. The engineer documents quality issues and risks and communicates them to relevant teams. Minimum qualifications include 3+ years of QA or testing experience with products that use LLMs, experience designing tests and test perspectives, and the ability to evaluate systems with ambiguous specifications. The position is based in Tokyo, reports to the function leader, offers a hybrid work arrangement, and involves business trips to Woven City several times a month.

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.