Lead platform engineering for AI-native software development and operations, building automation, quality, monitoring, and secure systems for sensitive medical data.
About this role
This Senior SRE works on the platform that supports Ubie's healthcare products using generative AI. The role is not limited to incident response: it develops and operates infrastructure that keeps services secure and available while enabling product teams to deliver quickly. A small platform team owns the company's product infrastructure, from the edge layer, including CDN and WAF, through a Kubernetes and Istio (Envoy) service mesh to managed databases. The role can expand its area of ownership and is expected to investigate and fix bottlenecks at any layer.
A major focus is defining infrastructure that improves development speed without trading away security or availability. The platform supports products deployed in Japan and the United States and must accommodate different medical regulations in each country. The role also plans capacity where model pricing, TPM rate limits and latency affect product design, and where LLM usage directly affects cost. Reliability, product capability and cost are considered together.
Responsibilities include developing and operating infrastructure compliant with country and product regulations; introducing and spreading Infrastructure as Code; providing performance visibility and observability; improving service levels through SRE practices; optimizing cost; and continuously reducing toil. Cost work includes designing committed-use discounts, moving environments toward spot capacity, removing unused resources, connecting allocated cost to product economics, and leading vendor negotiations and competitive quotes.
The team is also incorporating an internal Infra Agent into operations. It already handles some investigation, inquiry response and cost analysis. The SRE will help determine which work can be delegated to the agent and which requires human attention, including designing an on-call model that is still being established. The position covers automation, agent boundaries, SLOs and postmortems so that the organization learns from incidents and reduces their recurrence.
The expected profile can define a future infrastructure direction, remain hands-on, and work with other teams to deliver company-wide systems. The person should be able to design a platform in which reliability enables speed, rather than treating reliability, cost and speed as an unavoidable trade-off, and help product teams adopt SRE practices so the platform is used. The posting does not state an education requirement or a years-of-experience threshold.
The platform team uses Terraform, GitHub Actions and Design Doc reviews. Employees are expected to share information actively through channels such as note, X and LinkedIn, and are encouraged to speak at or attend talks, events and study sessions. The stated work location is Tokyo, Japan. The scope of duties may change to any duties designated by the company.
This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.