Responsible for developing systems and services using generative AI/LLM, identifying business challenges, and promoting service construction.
About this role
This role leads reliability, availability, and scalability across multiple products and systems supporting client-facing generative-AI and LLM solutions and consumer entertainment and advertising platforms. The work spans infrastructure through application code in an AI-driven development environment, with a full-stack SRE remit rather than infrastructure operations alone.
Responsibilities include establishing the SRE team and leading member hiring and development. During incidents, the role handles application-layer fixes and backend-centered decisions. It introduces and advances SRE practices such as SLO and SLI design and error-budget operation, and leads monitoring and alert design and operational improvements. Incident response includes investigation, recovery, and the design and implementation of permanent corrective measures.
The role designs, builds, and operates infrastructure for the AI solutions and platforms. It designs highly available architectures in cloud environments such as AWS and GCP, codifies and standardizes infrastructure using infrastructure as code such as Terraform, and designs, builds, and optimizes CI/CD pipelines. The role also works closely with development teams to balance rapid development with stable operations.
In addition to handling individual incidents, the position leads the creation of organizational systems, standards, and culture for reliability. The stated activities include establishing SRE practices such as runbooks and postmortems. The role also meets with clients to hear their issues and requests and organize requirements.
The stated work location is in Toranomon, Minato-ku, Tokyo.
This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.