Senior Site Reliability Engineer
MicrosoftAbout the role
Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.
Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform for the age of AI, powering a new class of data-first applications and driving a data culture.
Within Azure Data, the databases team builds and maintains Microsoft's operational Database systems. We store and manage data in a structured way to enable multitude of applications across various industries.
Running software as a service means more than just developing and releasing features. Ensuring reliability and serviceability is critical part of software cycle. This is where you come into the picture. As a Senior Site Reliability Engineer, you will ensure the service of Azure SQL Database or Managed Instance runs smoothly with required reliability and availability. You will design and implement software to automatically resolve issues. You will work closely with feature teams to design, implement and release features that are reliable and serviceable. You will be a cross-domain expertise who has a holistic view of our cloud service.
We do not just value differences or different perspectives. We seek them out and invite them in so we can tap into the collective power of everyone in the company. As a result, our customers are better served.
Responsibilities
• Act as subject matter expert for configuring, troubleshooting and monitoring Azure Database/Managed Instance services.
• Identify opportunities and implement automation to resolve and reduce live-site incidents.
• Design and implement solutions to improve service health, manageability, reliability and telemetry.
• Own, triage, investigate, and resolve service issues with emphasis on broad communications, learning, and teaching throughout the process.
• Author and maintain functional and technical documentation. Define and maintain process and procedures to run enterprise service.
• Interact with customers as result of escalation from support for issues including performance and availability.
• Ability to meet on-call responsibilities periodically.
Qualifications
Required/Minimum Qualifications
6+ years technical experience in software engineering, network engineering, or systems administration
o OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration
o OR Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administratio
Experience troubleshooting SQL Server query plan issue
Other Requirements
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check:
- This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
Preferred/Additional Qualifications
- 7+ years of software development, managing/supporting SQL product(s) such as SQL Server, Azure SQL Database or SQL Managed instance.
- Demonstrated troubleshooting skills in SQL Server/Azure SQL Database/Managed instance with deep understanding in one or more of the following areas in Query Processing, Storage Engine, SQL Operating System (SQL OS) layer, High Availability, Replication, Connectivity
- Deep understanding of Windows Operating System level concepts such as processes, threading, memory allocation, and the network stack; understanding of how applications are affected by the above, and ability to debug same.
- Proficient programming skills using managed code such as C#/Java. Ability to read native C/C++ code to debug issues and find answers no
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s