Jobs and Careers
XE
Lead Engineer, Site Reliability Engineering - Observability
XeroSan Mateo, United Statesfull_timeVerifiedPosted 7 Mar 2025
About the role
Our Purpose At Xero, we’re here to help you supercharge your business. We do this by automating routine tasks, surfacing actionable insights and connecting businesses with the right data, advisors and apps. When that happens, we’re not only making life better for small business, we’ll be building a stronger economy that can change the world.
As a Lead Engineers within the Site Reliability Engineering, (SRE), Observability team you’ll have a thorough knowledge of industry-leading observability practices and extensive hands-on experience. You’ll have a proven ability to provide strong technical mentorship, guiding engineers to upskill and enabling a focus on continual improvement. Working closely with the Product Manager, Team Lead, Principal and Lead Engineers you’ll have a product mindset and contribute your technical expertise and leadership to align team deliverables with the wider SRE and Xero initiatives.
You’ll help to adapt and grow observability at Xero, informed by a strong understanding of systems and reliability engineering and modern SRE principles. You’ll drive uplift in observability at Xero, paving the way for engineering teams to adopt Open Telemetry. You’ll be a strong advocate for the customer while contributing to the technical direction and roadmap for SRE products. You’ll model a growth mindset and help improve our services by identifying gaps, promoting capability growth, discovering technical solutions to business problems and championing modern practices.
Research has shown that women and underrepresented groups are less likely to apply to jobs unless they meet every single com
As a Lead Engineers within the Site Reliability Engineering, (SRE), Observability team you’ll have a thorough knowledge of industry-leading observability practices and extensive hands-on experience. You’ll have a proven ability to provide strong technical mentorship, guiding engineers to upskill and enabling a focus on continual improvement. Working closely with the Product Manager, Team Lead, Principal and Lead Engineers you’ll have a product mindset and contribute your technical expertise and leadership to align team deliverables with the wider SRE and Xero initiatives.
You’ll help to adapt and grow observability at Xero, informed by a strong understanding of systems and reliability engineering and modern SRE principles. You’ll drive uplift in observability at Xero, paving the way for engineering teams to adopt Open Telemetry. You’ll be a strong advocate for the customer while contributing to the technical direction and roadmap for SRE products. You’ll model a growth mindset and help improve our services by identifying gaps, promoting capability growth, discovering technical solutions to business problems and championing modern practices.
What You'll Do:
- Design systems to improve adoption of Xero's observability tools with a strong focus on reducing toil in managing our monitoring and logging platforms.
- A strong focus on developing and growing engineers through technical mentoring and coaching. Provide leadership around observability standards and practices.
- Create systems that support and enable our product teams to uplift their observability practices.Improve the implementation of system instrumentation as and when required.
- Be a key member of the pod leadership, contributing to technical strategy, feasibility, backlog management and enabling delivery.
- Participate with on-call roster.
- Empower other engineering teams at Xero to achieve a high standard of system awareness so they can create efficient, scalable and reliable applications for Xero's customers.
What You'll Bring With You:
- Experience with agile software development methodology including continuous integration and delivery.
- An understanding of how solutions architecture or architecture design works in a large software delivery organization.
- Experience building and implementing observability with large distributed cloud environments (ideally AWS).Excellent knowledge of reliability and observability concepts and practices.
- An understanding of Open Telemetry and how it works.
- Experience being on call and helping to resolve production incidents in a complex environment.
- Experience in instrumenting applications and integrating with monitoring solutions like New Relic, Datadog, Dynatrace, SignalFX, Scalyr, Sumo Logic or Splunk (ideally New Relic).Proficiency in one or more object-oriented programming languages such as C#, JavaScript, Golang, Python etc. Experience with DevOps tooling, eg. Linux, Docker, Kubernetes, IaC, CICD tools.
- The ability to help structure work to make optimal use of the team’s resources.
- The ability to set quarterly and annual objectives for the team in collaboration with the Product Manager and Team LeadProven ability to engage, influence and build relationships with internal stakeholders.
- Experience in managing and maintaining healthy observability platforms for a large user base.
Research has shown that women and underrepresented groups are less likely to apply to jobs unless they meet every single com
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s