Senior Site Reliability Engineer
MicrosoftAbout the role
Microsoft is looking for a Senior Site Reliability Engineer (SRE) to support and expand Viva Engage. Viva Engage (formerly Yammer) is the industry-defining social network for the enterprise. We provide a platform for millions of employees, including those from 85% of Fortune 500 companies, to build community and culture, share knowledge, and connect with their leaders and each other.
The user base for Viva Engage is growing quickly. The Site Reliability team is responsible for keeping the services reliable as we scale and modernize our tech stack. We need a SRE who knows how to manage the conflicting priorities of keeping things running today while making sure we have the architecture we need for the future.
Acquired by Microsoft in 2012, Viva Engage combines the benefits of a startup - rapid innovation, cutting-edge technology, outsized individual impact - with the advantages of working for one of the most successful software companies in the world. We believe in mission-driven work and in this post-Covid world, our platform has become more indispensable than ever as it fosters connection and a sense of belonging among remote teams.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
- Participate in on-call rotations and incident responses throughout product development and operations cycles. On-call will require responding to support requests after normal business hours to include the weekends and/or holidays in a designated Microsoft office.
- Monitor system performance and proactively identify and resolve issues to ensure high availability and performance.
- Develop and maintain automation tools and processes for deployment, monitoring, and configuration management.
- Apply troubleshooting skills, debugging tools, and examines logs, telemetry, and other methods to verify assumptions and customer impact. Proactively and reactively address findings with customer and/or service engineering efficiently via written and verbal communications.
- Lead blameless postmortems for root cause and production resiliency.
- Consult with developers to design services that scale in Azure.
- Mentor team members and contribute to the overall growth and development of the SRE team.
- Stay current with industry trends, emerging technologies, and best practices in site reliability engineering and cloud computing.
Qualifications
Required Qualifications:
- 6+ years technical experience in software engineering, network engineering, or systems administration
- OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration.
Other Requirements:
Ability to meet Microsoft, customer and/or government security screening requirements are
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s