Jobs and Careers
FI

Staff AI Inference and Acceleration Engineer

Figure
San Jose, USAfull_timePosted 26 Jun 2026

About the role

<p>Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.</p> <p>We are looking for a Staff AI Inference &amp; Acceleration Engineer to join the Platform Software team and own the on-board inference architecture for Figure’s humanoid robots. You will be the technical authority on how AI workloads are mapped, optimized, and executed across the robot’s compute hardware — driving down power consumption and cost while meeting the strict latency and reliability demands of a real-time autonomous system.</p> <p><strong>Responsibilities:</strong></p> <ul> <li>Own the on-board inference architecture — mapping models to available accelerators (NPU, GPU, DSP, CPU) based on latency, power, and memory budgets.</li> <li>Partition inference workloads across heterogeneous compute resources, balancing real-time performance with power and thermal constraints.</li> <li>Define and maintain a system-level compute budget across all inference tasks running on the robot.</li> <li>Evaluate next-generation acceleration hardware and contribute to the definition of future compute platform requirements.</li> <li>Optimize inference toolchains end-to-end — from model export through runtime execution — for target hardware.</li> <li>Apply quantization (INT8, INT4, mixed-precision), pruning, operator fusion, and other compression techniques to reduce compute, memory, and power footprint.</li> <li>Profile inference pipelines to identify and eliminate bottlenecks in latency, memory bandwidth, and power consumption.</li> <li>Optimize kernel scheduling, memory layout, and data movement across the compute hierarchy.</li> <li>Partner closely with the AI/ML team to define model architecture constraints that are hardware-friendly from the outset.</li> <li>Work with the Platform Software team on runtime integration, scheduling, and power management.</li> <li>Engage with silicon vendors and research teams to track the accelerator landscape and influence hardware roadmaps.</li> </ul> <p><strong>Requirements:</strong></p> <ul> <li>M.S. or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or a related field — or equivalent industry experience.</li> <li>At least 8 years of industry experience in hardware acceleration, ML systems, or compute architecture.</li> <li>Deep understanding of AI/ML inference — model formats (ONNX, TFLite, etc.), inference runtimes, and deployment pipelines.</li> <li>Hands-on experience optimizing models for edge or embedded hardware using

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Figure

View company profile →