SRE / Infrastructure Engineer, ML Platform

Money Forward

SRE

Python

About this role

This full-time SRE and infrastructure role owns the infrastructure for the ML Ops platform supporting digital banking and group Fintech products. The work focuses on serving interfaces and production model endpoints, with responsibility for stable operation, monitoring, security, and high-availability infrastructure for the planned 2027 digital bank launch.

Responsibilities include designing, building, coding, and continuously improving infrastructure for serving interfaces, data pipelines, and ML development environments in AWS using Terraform. The role designs and maintains financial-grade network security, including mutual TLS authentication, private network connections such as AWS PrivateLink and VPC Peering, and ongoing security monitoring.

The position establishes observability using CloudWatch and related monitoring of logs, metrics, and traces. It includes defining SLOs and SLAs, launching an on-call structure, and standardizing incident-response processes. It also designs and builds the interface between the Databricks data platform and the ML Ops platform, including data processing and pipeline integration, and provides infrastructure support for ML Ops.

The engineer collaborates with the SRE responsible for the shared ML Ops infrastructure standard, as well as data-platform and product-platform teams in other organizations and group companies. The current delivery structure includes a product manager, engineering manager, data scientists, and ML engineers working in a scrum team. The role is dedicated to the digital banking and Fintech area while coordinating with the broader platform.

Required experience includes approximately five or more years working as an SRE or infrastructure engineer; designing, building, and operating AWS infrastructure such as ECS, API Gateway, VPC, and IAM; and practical Terraform experience covering infrastructure codification, module design, and automated operation through CI/CD. Candidates should have experience designing secure network architectures using VPC Peering, AWS PrivateLink, and TLS or mutual TLS; operating cloud monitoring and alerting systems and analyzing performance logs; and building or optimizing build and deployment pipelines with GitHub Actions and Docker.

Experience supporting infrastructure for MLOps or data platforms using Amazon SageMaker, Databricks, or Airflow is welcome, as are experience establishing on-call and incident-management operations, financial or payments security and audit work, cross-organizational platform coordination, and AI-assisted development. The listed technology stack includes AWS ECS Fargate, API Gateway, VPC, PrivateLink, SageMaker, S3, Glue, CloudWatch, Databricks, Terraform, GitHub Actions, Docker, Claude Code, Python, Shell, SQL, Slack, and Notion. Datadog and PagerDuty are not currently adopted, leaving technology selection open.

Business-level Japanese sufficient for client communication and basic business-level English equivalent to TOEIC 700 or higher are required. The position is based in Minato, Tokyo, under a discretionary work system where applicable, with a standard schedule of 9:30–18:30 and a 60-minute break. It uses a hybrid work style with two days of office attendance generally required and three or more days recommended. Annual compensation is JPY 7,008,000–11,004,000, including fixed allowances for specified overtime and late-night work. The probationary period is three months. Benefits include social insurance, housing-related allowances, paid leave, seasonal leave, a defined-contribution pension, an employee stock ownership plan, book-purchase support, health examinations, and conference support.

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.