Senior Applied Data Scientist
MicrosoftAbout the role
Document Understanding Team plays a pivotal and central role in understanding various public web information needs of users and building next-generation models to stay ahead of the curve and push the platform capabilities to go hand in hand with modeling improvements. Work in our team is unique as you'll be able to work with industry-leading scales of data (raw and training data), computing resources and best-quality large language models like GPT4, multimodal large language model, etc.
In this role, you will be leading projects from idea creation through implementation, experimentation and delivering improvements to real world scenarios, working closely with various partners. We are looking for a passionate and motivated team member who has solid, hands-on experience in transforming business problems into ML problems, collecting high-quality labels, and developing state-of-the-art models to address product challenges and drive value for end users. You will also leverage your software engineering skills and ML expertise in fields such as natural language processing and information retrieval to help create the next generation of text representation and understanding techniques at Microsoft.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
- Lead innovation and development of deep learning models for document understanding and their usage in downstream tasks, e.g., Copilot, Generative Search, Search, Spam, QA and recommendation, etc.
- Identify opportunities in Web Data space to solve using ML at scale of 100s of Billions of documents.
- Push the state-of-the-art in those areas through multiple aspects, for example:
- Defining the problem space
- Gathering training data at scale
- Exploring model design and architecture
- Exploring learning objectives and tasks
- Build Feature Generation Algorithms used at Index Generation time.
- Build Automated Document Understanding Training Pipelines.
- Guide team members to develop new technologies that lead to solutions that impact real production scenarios in Microsoft.
- Work closely with various teams in WebXT to understand common needs and build technical roadmap for addressing them.
Qualifications
Required Qualifications
- Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 4+ years related experience (e.g., statistics predictive analytics, research)
- OR Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 3+ years related experience (e.g., statistics, predictive analytics, research)
- OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 1+ year(s) related experience (e.g., statistics, predictive analytics, research)
- OR equivalent experience.
- 4+ years of experience in product development in the areas of Software Engineering and Machine Learning/Deep Learning
- Hands-on experience in developing algorithms and models using deep learning frameworks such as TensorFlow, PyTorch, etc.
Other Requirements
Candidates must be able to meet Microsoft, customer and/or government security screening requirements that are required for this role. These requirements include, but are not limited to the following specialized security screenings:
- Microsoft Cloud Background Ch
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s