Site Reliability Engineer (SRE)

TIER IV

SRE

DevOps

Cloud

About this role

This SRE improves the reliability of Web.Auto and the infrastructure supporting autonomous-driving development and operations. Assigned areas depend on experience rather than requiring ownership of every listed system.

Work can include deployment and rollback systems, ECS/EKS/Lambda services, GPU training infrastructure, monitoring, operational automation, security and cloud-cost controls. The platform also involves HPC technologies such as Slurm, Lustre and InfiniBand, with regional deployment and data-management requirements.

Applicants need at least five years in DevOps/SRE; cloud, container, networking and security fundamentals; Linux/Unix development; and software or web development in a general-purpose language such as Python or Go. Required experience includes infrastructure as code, orchestration, CI/CD and Git. Japanese speaking, reading and writing, plus technical-English reading and writing, are required.

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.