Site Reliability Engineer
GeotabAbout the role
Who we are:
Geotab ® is a global leader in IoT and connected transportation and certified “Great Place to Work™.” We are a company of diverse and talented individuals who work together to help businesses grow and succeed, and increase the safety and sustainability of our communities. Geotab is advancing security, connecting commercial vehicles to the internet and providing web-based analytics to help customers better manage their fleets. Geotab’s open platform and Geotab Marketplace ®, offering hundreds of third-party solution options, allows both small and large businesses to automate operations by integrating vehicle data with their other data assets. Processing billions of data points a day, Geotab leverages data analytics and machine learning to improve productivity, optimize fleets through the reduction of fuel consumption, enhance driver safety and achieve strong compliance to regulatory changes. Our team is growing and we’re looking for people who follow their passion, think differently and want to make an impact. Ours is a fast paced, ever changing environment. Geotabbers accept that challenge and are willing to take on new tasks and activities - ones that may not always be described in the initial job description. Join us for a fulfilling career with opportunities to innovate, great benefits, and our fun and inclusive work culture. Reach your full potential with Geotab. To see what it’s like to be a Geotabber, check out our blog and follow us @InsideGeotab on Instagram. Join our talent network to learn more about job opportunities and company news.Who you are:
We are always looking for amazing talent who can contribute to our growth and deliver results! Geotab is seeking Site Reliability Engineer professionals who with training, will be able to quickly contribute to the Site Reliability team. If you love technology, are passionate about engineering support, and are keen to join an industry leader — we would love to hear from you!
What you'll do:
As a part of Site Reliability Engineering team, your key area of responsibility is to ensure the availability, reliability, and performance of Geotab's core products for our customers. This role acts as a primary escalation point, diagnosing and resolving complex application issues impacting service availability and performance of multiple large scale applications that support thousands of customers globally. SRE supports production applications and infrastructure, focusing on restoring normal service operations efficiently and contributing to long-term system stability.
How you'll make an impact
-
Act as a primary escalation point for critical production application/product issues.
-
Rapidly troubleshoot complex problems across the application stack, utilizing observability tools to identify root causes.
-
Coordinate effectively with development, infrastructure, and other technical teams during incidents to implement fixes and restore service swiftly.
-
Clearly communicate incident status, impact, and resolution steps to internal stakeholders.
-
Collaborate with team members to improve monitoring tools, dashboards, and alerting mechanisms for proactive detection of issues impacting Critical User Journeys (CUJs) within the application/product and computing architecture. Our complex environment encompasses monolithic applications, microservices, and a vast ecosystem of millions of hardware units.
-
Monitor application/product and system health proactively using a combination of tools to ensure high availability and adherence to Service Level Objectives (SLOs) / Service Level Agreements (SLAs).
-
Identify opportunities and implement automation tools/scripts to streamline routine operational tasks, reduce manual effort (toil), and improve response times.
-
Conduct system tests to validate performance, reliability, and successful remediation of issues.
-
Recommend design and process enhancements based on operational experience to improve overall application reliability and maintainability.
-
Participate in post major incident reviews (PMIRs) to analyze disruptions, document findings, track corrective actions to prevent recurrence, and identify areas of improvement for incident response processes.
-
Contribute to building a culture of learning from incidents.
-
Participate in a 24x7 on-call rotation to
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s