Leads customer implementations of LLM and generative-AI systems from architecture through production, building full-stack solutions and improving model quality.
About this role
This role sits in SoftBank’s IT and Solutions Operations organization, supporting private cloud infrastructure and computing platforms used for generative AI development. The environment includes large-scale on-premises infrastructure spanning GPU servers, dedicated storage, job management systems, virtualization, Active Directory, DNS, and mail servers.
The engineer investigates incidents and high-load performance issues, coordinates root-cause resolution, develops preventive and recurrence-prevention measures, and improves operational efficiency through programming and internal tools. The role also includes creating operational procedures and documentation and coordinating outsourced members. The posting requires at least three years of experience operating, maintaining, and troubleshooting Linux and Windows operating systems, experience investigating infrastructure failures and controlling vendors, and hands-on understanding of operating systems and middleware behavior.
This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.