SRE / Cloud Infrastructure Engineer

atama plus

SRE

Cloud

TypeScript

Python

About this role

The SRE / Cloud Infrastructure Engineer designs, builds, and operates scalable, highly available cloud infrastructure for products that deliver individualized learning. atama+ is an AI learning product provided to more than 4,500 cram schools and preparatory-school classrooms in Japan. The organization also operates atama+ learning schools and offers programs for universities and, from March 2026, employee education programs for companies.

The role covers infrastructure design and construction, automation of operational systems, and improvements to overall performance, reliability, and scalability so students can focus on learning. Responsibilities include designing AWS cloud infrastructure and CI/CD pipelines with availability, maintainability, and security in mind; building infrastructure with Terraform; creating and optimizing CircleCI deployment pipelines; designing and operating an internal developer platform; implementing internal CDK libraries; and preparing guidelines that help application developers build and operate infrastructure.

Day-to-day operations include analyzing performance, finding bottlenecks, implementing infrastructure-side improvements, capacity planning, resource management, and scaling strategies that reflect business requirements. The engineer builds and runs Datadog monitoring, handles alerts, analyzes and responds to major incidents with relevant teams, performs post-incident root-cause analysis, and plans measures to prevent recurrence. The role also implements security measures, handles regular updates, and investigates security alerts. Repetitive tasks are identified for automation, while continuous improvement processes, design and operations documentation, and recurring knowledge-sharing sessions are established and maintained.

The SRE team consists of one team leader and three cloud infrastructure engineers. It works closely with product development and data platform teams on company-wide infrastructure strategy and execution. The wider engineering organization includes application engineers, QA, and product managers, with cross-functional collaboration on end-to-end system design and implementation. New projects may involve leading an infrastructure team while designing and building new infrastructure from the beginning.

Named technologies include AWS, Google Cloud, Docker, ECS, Terraform, AWS CDK, CircleCI, GitHub Actions, Atlantis, Datadog, Amazon Aurora PostgreSQL, AWS Lambda, AWS Step Functions, Amazon Bedrock, Python, TypeScript, Cursor, GitHub Copilot, Claude Code, Devin, and Takumi. The listed work location is 2-1-2 Koraku, Bunkyo-ku, Tokyo, at Sumitomo Fudosan Iidabashi Building No. 5, first floor.

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.