Researcher, Natural Language Processing (NLP) for Multimodal Machine Learning

Sony Group

ML

Research

NLP

About this role

This research role sits in an R&D organization focused on large-scale generative AI for entertainment content creation and production across music, film, and games. The work connects fundamental research with Sony business groups, universities, product teams, and studio workflows.

The researcher will investigate multimodal learning, multimodal LLMs, music and video understanding, agents, reasoning, controllable generation, discrete-data generative models, image and audio captioning, text-to-image and text-to-audio systems, vision-language pre-training, and commonsense knowledge graphs. Responsibilities include developing large-scale data, submitting papers to conferences such as ACL, EMNLP, NeurIPS, ICLR, and CVPR, collaborating on research, and deploying technologies in products and AI-assisted content-creation tools. Applicants need a master's degree or equivalent practical experience, at least three years using Python, C/C++, and Linux/Unix, at least two years in machine learning and NLP with PyTorch or TensorFlow, and evidence of research ability through papers, open-source software, or other scientific work.

This is an AI-generated summary of the employer's original posting — details can be incomplete, out of date or simply wrong. Always confirm everything on the official posting before applying.