Cloud Engineer, Site Reliability Engineering

Square Enix

SRE

DevOps

Cloud

About this role

The role performs site reliability engineering for in-house backend systems supporting large-scale online RPGs and other services. The team defines SRE as engineering responsibility for the reliability of the entire cloud environment in which code is developed and executed, with broad responsibility and discretion.

Responsibilities include managing cloud infrastructure through infrastructure as code, building CI/CD pipelines, improving API performance, defining and visualizing service-level objectives, improving logging and monitoring platforms, and promoting the use of observability. The role also covers security management and improving work efficiency through AI, with the overall aim of increasing system reliability. The position involves monitoring the performance of backend systems supporting Square Enix games and contributing to end-user game experiences through SRE practices.

Required experience includes building and operating systems using public clouds such as AWS and Google Cloud; knowledge of and professional experience with container technologies such as Docker and Kubernetes; knowledge of and professional experience with Linux such as Ubuntu; and building and operating systems with infrastructure-as-code tools such as Terraform. The posting also requires communication skills for coordinating work involving multiple teams and a proactive approach to proposing improvements.

Preferred experience includes applying SRE practices such as SLOs and CI/CD; knowledge of and experience with RDBMS such as MySQL; building and operating monitoring platforms such as New Relic and Datadog; building and operating logging platforms such as Elastic Cloud and Cloud Logging; and using security monitoring platforms such as Sysdig and Security Command Center. High-load system operations, performance optimization, AI-based process improvement and automation, incident response, troubleshooting, and leading teams or projects are also listed as desirable. Coding skills in Go or Python for automating development and operations are preferred.

This is a permanent full-time position with no fixed contract period and a six-month probationary period. The work arrangement is flextime without core hours, with a standard working day of 7.5 hours excluding breaks. Overtime may occur. The workplace is the Tokyo headquarters in Shibuya, Tokyo, or a home or equivalent location when approved; a work-from-home system is available. Benefits include transportation expenses, health insurance association resort facilities, health examinations, paid leave, and social insurance. Salary is determined according to company regulations based on experience and ability.

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.