Jobs and Careers
SI
Sr Principal Software Engineer, Quantization (AI2324)
SiMa.aiSan Jose, United Statesfull_timeVerifiedPosted 27 Oct 2024
💰 $305,000/yr($240,000/yr – $305,000/yr)
About the role
Job Title: Sr Principal Software Engineer, Quantization Job Location: San Jose, CA
Job Number: AI2324 Job Description: SiMa.ai is seeking an outstanding researcher working on efficient deep learning to join the MLSoC Platform Architecture team. We are passionate about pushing the boundaries of Edge AI with power efficient inferencing. We are particularly interested in Post Training Quantization and Pruning techniques applied to quantization of CNN and Transformer based Neural Networks for inference primarily on int8 Machine Learning Accelerator (MLA) and on mixed precision MLA. You will work with an amazing team of engineers that pushes the boundaries and your contributions will have a chance to create a real impact in our products. Sr. Principal Engineer Key Responsibilities (including but not limited to):
Job Number: AI2324 Job Description: SiMa.ai is seeking an outstanding researcher working on efficient deep learning to join the MLSoC Platform Architecture team. We are passionate about pushing the boundaries of Edge AI with power efficient inferencing. We are particularly interested in Post Training Quantization and Pruning techniques applied to quantization of CNN and Transformer based Neural Networks for inference primarily on int8 Machine Learning Accelerator (MLA) and on mixed precision MLA. You will work with an amazing team of engineers that pushes the boundaries and your contributions will have a chance to create a real impact in our products. Sr. Principal Engineer Key Responsibilities (including but not limited to):
- Research, design and implement novel methods to improve PTQ techniques for both int8 and mixed-precision (int8 + bf16) quantization.
- Collaborate with other team members to understand the limitations of our Machine Learning Accelerator and adapt your strategy based on their input.
- Prototype PTQ techniques using Fake Quantization in PyTorch, as well as modify internal tools to implement quantized operators to verify accuracy.
- Understand state-of-the-art research in PTQ and apply it to CNN and Transformer based Neural Networks.
- Help define timeline and deliverables and be accountable for them.
- PhD in electrical engineering or computer science with 6+ years research numerical methods and tools in efficient Neural Network inferencing.
- Proficient in techniques like HAWQ2, and RL based methods for Mixed-precision quantization.
- Proficient in state-of-the-art PTQ techniques like Optimum Brain Compression for LLMs.
- Proficient with PyTorch or other Quantization exploration frameworks like Model Compression Toolkit.
- Excellent programming skills in C++, Python.
- Co-authored internal technical presentations, research papers and disclosures/patents on key technical topics
- Noteworthy technical contributions, which were multi-disciplinary and in collaboration with other cross-functional teams.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s