Senior–Lead Infrastructure/SRE Engineer

Safie

Python

Shell

AWS

About this role

This role sits in the infrastructure group responsible for the cloud foundation of Safie’s video platform. The environment runs primarily on AWS, with more than 2,000 servers and very large volumes of video data. The team works closely with server-side groups to keep the platform reliable as connected cameras, traffic, and services expand.

The work covers designing and operating highly available cloud architecture, optimizing server costs, monitoring and tuning services, reducing operational toil, and evaluating database technologies as transaction volumes grow. The listed environment includes Linux, Python, ShellScript, AWS, Docker, Terraform, Ansible, Packer, MySQL, Redis, PostgreSQL, TiDB, Prometheus, Grafana, GitHub Actions, Git, and related monitoring and delivery tools. A leader-track version also includes developing a team of around five to six members and promoting architecture, operations, and process improvements.

Applicants need experience designing, building, and operating systems of at least several dozen Linux servers; practical experience with container technologies and infrastructure as code such as Docker and Terraform; at least five years of AWS operations; experience operating and maintaining ten or more AWS services; CI/CD implementation experience; and a record of leading efficiency or productivity improvements as business scale increased.

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.