Build AI/ML solutions for patient and healthcare operations, covering technical validation, model design, product integration, evaluation, and continuous quality improvement.
About this role
The SRE will work across the company’s online and offline healthcare services to improve platform reliability, safety, scalability, and developer productivity. Responsibilities include leading gradual re-architecture and operational improvements for a highly available medical platform, designing, building, and operating infrastructure with AWS and GCP, and managing infrastructure as code with Terraform.
The role includes defining and operating SLOs and SLIs, improving observability with tools such as Datadog, and defining and improving reliability metrics across services. The SRE will also improve and stabilize CI/CD pipelines used by multiple services and teams, including CircleCI and GitHub Actions, and improve deployment flows using ECS and Fargate.
Security and incident responsibilities include planning and implementing measures such as vulnerability assessments and WAF configuration in response to healthcare-domain requirements. The role participates directly in incident response, leads responses when service failures occur, and creates recurrence-prevention measures through postmortem analysis. It also involves standardizing incident-response criteria, reducing dependencies on legacy configurations and application specifications, and continuously strengthening security based on the 3省2ガイドライン and the AWS Security Maturity Model.
The position leads technical decisions involving technology selection and architecture design, and works with development teams to promote and communicate SRE practices. The team uses regular meetings, daily stand-ups, Slack, Slack Huddles, Google Meet, and Gather.town for asynchronous and synchronous communication.
The stated development environment includes Ruby, Ruby on Rails, TypeScript, React, Next.js, Dart, Flutter, GraphQL, AWS services including ECS Fargate, CloudFront, and Lambda, GCP with BigQuery for a data analysis platform, Datadog, PostgreSQL, CircleCI, Autify, AWS CodePipeline, GitHub Actions, Terraform, Metabase, Amplitude, Figma, GitHub, Slack, and Notion.
The posting states that the position is fully remote, with the listed office in Hamamatsucho, Minato-ku, Tokyo also available, and offers flextime.
This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.