Senior Applied Scientist
ZillowAbout the role
About the team
Are you interested in innovating new multimodal technologies to power experiences that change how millions of people using Zillow experience and find their next home? Zillow’s AI Team plays an important part in delivering unique AI-powered experiences for the hundreds of millions of customers that visit Zillow websites each month. The AI Media Insights team focuses on multimodal foundational models (CV, NLP, 3D) to develop next-generation indoor and outdoor understanding and reasoning capabilities that enable Zillow's customers to make sense of homes in order to find their dream home. Our team includes research, product, and engineering. We empower the next generation of home shopping by crafting new innovative technologies that deliver unrivaled insights and interactive AI-fueled experiences to our customers.About the role
As a Senior Applied Scientist on the AI Media Insights team, you’ll be researching and developing state-of-the-art multimodal foundational technologies. This role also requires an entrepreneurial approach to bridge the gap between research and production in the constantly evolving field of Multimodal Foundation Models. You will build algorithms and deep learning models that power a wide range of experiences for buyers, agents and renters. We are seeking an innovative, collaborative and customer-focused applied scientist with deep expertise in Multimodal Machine Learning and Computer Vision.
This role is encouraged to:
Challenge established assumptions and expectations by coming up with new ideas that fuel innovation and creativity.
Apply a growth-mindset, proficiency with modern frameworks, and first principles to ambiguous customer problems to rapidly iterate on novel solutions and ways of working in this space.
Collaborate closely with other applied scientists, engineering, design and user research to understand, scope, design, prototype, implement and iterate on internal- and external-facing systems supporting and implementing next generation AI applications.
Cultivate connections with other teams for critical dependencies and infrastructure.
Contribute to carrying and growing our team culture of rapid innovation and creative frugality.
Who you are
You are a roll-up-the-sleeves and get-it-done researcher with a deep understanding of metrics, applied research, and supporting and iterating on models at scale. You have:
3-5 years of experience in computer vision, with a focus on vision and language, multimodal generation, image, video. A Ph.D. in Computer Science or Engineering, Machine Learning, Computer Vision, or related domain. (Or an MSc in a relevant field with 5+ years of R&D experience)
A basic understanding of geometry and 3D.
Solid background in Multi-task, Multimodal (Vision + Language) research, building models & algorithms with a proven track record of publication at top tier Computer Vision and/or Machine Learning venues (e.g., CVPR, ICCV, ECCV, ICML and NeurIPS).
Proficiency with a high-level programming language (Python is preferred)
Experience training machine learning models in frameworks like Tensorflow and PyTorch
Familiarity with the latest advancements in Multimodal Foundational Models
The communication skills to influence, collaborate with, and educate others (whom you may need to educate on methods and requirements in experimentation and statistics).
Experience prototyping, developing, and implementing algorithmic solutions and new technologies with diverse analytics and data.
Tenacity to embrace and solve comp
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s