Leads customer implementations of LLM and generative-AI systems from architecture through production, building full-stack solutions and improving model quality.
About this role
This Service Reliability & Operations Engineer role supports SB OAI Japan’s delivery of AI products for enterprise customers. It sits at the intersection of customer operations, system reliability, development, security, legal, and governance, with responsibility for bringing AI-enabled systems into dependable production use.
The engineer leads production operation design and rollout, then improves ongoing service quality. Work includes operational architecture, monitoring and alert design, SLO and SLI management, incident response, change and release operations, performance and capacity management, and automation through observability, runbooks, and other repeatable mechanisms. The role also coordinates with customers, vendors, internal developers, security, legal, and audit teams.
Applicants must have at least two years of experience related to system operations or reliability improvement, experience designing or operating public-cloud environments such as AWS, GCP, or Azure, and an understanding of or interest in AI-product operations, including LLM or agent quality, safety, data handling, security, and evaluation.
This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.