About this role

ANDPAD is hiring a Site Reliability Engineer to balance stable service operation with developer experience for its construction-DX platform (265,000+ companies, 690,000+ users) built as many microservices. Work includes: building, operating and updating shared platform systems and supporting per-product infrastructure; resolving performance issues and technical debt with dev teams; coordinating system maintenance; managing SaaS/accounts; infrastructure cost optimization; infra-security measures; and platform engineering. Environment: AWS, Google Cloud, Elastic Cloud; Kubernetes (EKS), ECS on Fargate, Linkerd, Cloud Run; Ruby, Go, Rails, Sidekiq; CI/CD (CircleCI, GitHub Actions, CodeBuild/CodePipeline, etc.); observability (Datadog, Bugsnag, Sentry); IaC (Terraform, CloudFormation, SAM). Required: basic network/security knowledge, design and production operation of cloud systems, ability to find and solve bottlenecks, and hands-on Docker/Kubernetes experience. Welcome: RDBMS (MySQL/RDS) operation, monitoring/observability design, programming to improve performance, and English technical communication (mainly reading/writing). Flextime with core hours; remote-friendly. Business-level Japanese is expected. Full-time, based in Tokyo (Mita).

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.