Jobs and Careers
GR

Senior Software Engineer, Perception (Robotics)

Grab
Chinafull_timeVerifiedPosted 20 Aug 2026

About the role

Company Description

About Grab and Our Workplace

Grab is Southeast Asia's leading superapp. From getting your favourite meals delivered to helping you manage your finances and getting around town hassle-free, we've got your back with everything. In Grab, purpose gives us joy and habits build excellence, while harnessing the power of Technology and AI to deliver the mission of driving Southeast Asia forward by economically empowering everyone, with heart, hunger, honour, and humility.

Job Description

About the Team:

The Robotics Technology team is a core part of Grab's long-term vision to build urban embodied AI. Our engineers take full ownership of the product lifecycle: designing and manufacturing hardware in-house, developing control and machine-learning systems, and rigorously testing in real-world conditions and production fleet operations. We are executing an ambitious growth plan to expand our robotics fleet across cities over the coming years, and we are focused on delivering highly productive, safe and efficient robot delivery services that help address current delivery labour shortages.

Based in Singapore and China, we offer opportunities to work on the latest autonomy, deploy solutions in complex environments, and directly influence the future of last-mile logistics. If you're excited by tangible impact, large-scale systems and cross-functional engineering, you'll find meaningful challenges and rapid career growth here.

Get to know the role

As a Senior Perception & Prediction Engineer, you will build and ship systems that turn raw multi-sensor data into a reliable, real-time and predictive representation of the world. You will work across a robust modular perception stack and modern learning-based approaches, contributing through model development, evaluation, debugging, integration and on-robot validation.

On top of a shipped multi-sensor detection baseline, you will deepen three connected capabilities: multi-object tracking, generic / open-set world understanding, and motion prediction. You will also help evaluate and productionise end-to-end temporal perception-prediction models and VLA / embodied foundation models for open-vocabulary understanding, long-tail reasoning and task-conditioned robot intelligence.

We are pragmatic about new research: a model earns its place through measurable closed-loop value, reliable grounding, real-time performance and safety. Classical geometry, filtering and modular components remain important as interpretable baselines, safety fallbacks and guardrails. This role offers the opportunity to take promising research from prototype to fleet data, embedded deployment and real-world robot behaviour.

You will report to the Senior Principal Perception & Prediction Engineer and work onsite at a Grab office.

The critical tasks you will perform

  • You will develop and improve multi-object tracking — data association, Bayesian state estimation (Kalman / EKF / UKF and motion models), track lifecycle, and ID stability through detector gaps, occlusions and crowded scenes — and evolve it toward learning-based tracking, including learned association, joint detection-and-tracking and query / transformer-based trackers.
  • You will build and harden generic / open-set world understanding — class-agnostic obstacle detection, occupancy grids / occupancy networks / BEV occupancy, clustering and learned generic-object branches — so the robot safely reacts to rare or unknown objects.
  • You will build motion prediction for pedestrians, cyclists, vehicles and other agents, including multi-modal trajectory forecasting, interaction-aware prediction, occupancy flow, uncertainty estimation and well-defined interfaces into behaviour and planning.
  • You will develop and evaluate end-to-end temporal perception-prediction approaches, such as joint detection-tracking-forecasting, streaming BEV representations, agent / trajectory queries and learned world models, while maintaining strong modular baselines and production fallbacks.
  • You will explore VLM / VLA and embodied foundation models for open-vocabulary perception, semantic scene reasoning, task-conditioned understanding, long-tail discovery, auto-labeling and teacher-student supervision. You will test grounding, hallucination, temporal consistency, latency and robustness, then distil, quantise or adapt useful capabilities for on-robot deployment rather than stopping at demos.
  • You will strengthen multi-sensor and temporal integration across camera, LiDAR, radar and IMU, ensuring geometry, timing, identities, predictions and semantic context remain consistent over time

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Grab

View company profile →