About the role
<p><strong>About Us</strong></p> <p>Sage is on a mission to improve care and quality of life for older adults, starting with those residing in senior living facilities. Falls are the leading cause of injury-related death among adults over 65. And yet, fall prevention and emergency response systems for older adults are archaic and ineffective. At Sage we've built a more modern way of understanding when older adults need help, including methods for residents to alert caregivers when in need of help, and corresponding software for caregivers to triage response. Our company mission is to create a product that our client counterparts love, and this role is a key part of that objective.</p> <p>Sage is a small, tight team of ambitious, multi-disciplinary entrepreneurs. We are a software-enabled, mission-driven company, and are focused only on the problems that are central to achieving that mission. At Sage, we work hard and fast but also know that to build a truly important company, we need to treat our work as a marathon, and not a sprint. The journey matters.</p> <p><strong>About this Role</strong></p> <p>Sage provides life-saving functionality that improves the lives of our older population. This role is critical to ensure Sage can live up to its mission to be a 24x7, highly available platform for elder care. As a Site Reliability Engineer, you’ll partner with engineering teams across the organization to achieve four 9s of uptime for our platform.</p> <p><strong>Responsibilities</strong></p> <ul> <li><strong>Design and evolve highly reliable system architectures</strong>, ensuring high availability, fault tolerance, and scalability across Sage’s production infrastructure.</li> <li><strong>Lead complex incident response efforts</strong>, coordinating across engineering teams to quickly diagnose and resolve production issues while driving thorough post-incident reviews and long-term reliability improvements.</li> <li><strong>Define and implement organization-wide observability practices</strong>, including metrics, logging, tracing, and actionable alerting to ensure strong visibility into system health.</li> <li><strong>Establish and maintain reliability standards</strong>, including defining SLIs, SLOs, and error budgets, and partnering with engineering teams to integrate these practices into the software development lifecycle.</li> <li><strong>Drive automation and infrastructure improvements</strong> that reduce operational toil and improve the efficiency and reliability of deployments, monitoring, and operational workflows.</li> <li><strong>Partner with engineering teams on system design and architecture reviews</strong>, ensuring reliability, scalability, and operational best practices are conside