Senior Software Engineer
MicrosoftAbout the role
Microsoft is on a mission to empower every person and every organization on the planet to achieve more. Our culture is centered on embracing a growth mindset, a theme of inspiring excellence, and encouraging teams and leaders to bring their best each day. In doing so, we create life-changing innovations that impact billions of lives around the world. You can help us achieve our mission.
Cloud Operations + Innovation (CO+I) is the engine that powers Microsoft’s core cloud platforms and services that millions of people use every day. With more than 95% of Fortune 500 business on Azure, 180 million using Office 365, and millions using other services – all running on Microsoft's cloud infrastructure – CO+I builds and operates the foundation upon which Microsoft’s mission to empower every person and organization comes to life.
Are you passionate about cloud computing? Do you get excited about taking a hands-on approach to transforming Microsoft’s most critical business through investigation, data analysis, and automation? If so, come and help us build the most reliable & efficient datacenter infrastructure on the planet. The CO+I Critical Environment Systems Intelligence (CESI) team is responsible for designing and delivering solutions to support global datacenter operations and to improve availability. CESI is helping to drive CO+I’s transition to a customer centric, data driven, live service culture. As a Senior Software Engineer, you will be a key player in this transition.
As a Senior Software Engineer on the CO+I Critical Environment Service Intelligence (CESI) team, you will be responsible for analyzing live and historical telemetry to create and maintain critical environment anomaly detections. You will partner with CO+I teams to build off and leverage each other to ensure that we are driving continuous improvements across the fleet for critical environment anomaly detections and detection onboarding. You will work with massive amounts of data with low latency requirements across cutting edge technologies, with the potential for significant potential impact to both internal partners and external customers.
In alignment with our Microsoft values, we are committed to cultivating an inclusive work environment for all employees to positively impact our culture every day.
Responsibilities
- Create innovative solutions to complex and unconventional problems.
- Leads discussions for the architecture of products/solutions and creates proposals for architecture by testing design hypotheses and helping to refine code plans
- Contribute to the development and design of software services, applications, and tools that are secure, highly available, scalable, and reliable to meet the needs of the business.
- Leads by example within the team by producing extensible and maintainable code.
- Optimizes, debugs, refactors, and reuses code to improve performance and maintainability, effectiveness, and return on investment (ROI).
- Applies metrics to drive the quality and stability of code, as well as appropriate coding patterns and best practices.
- Triage live site incidents to Identify patterns and create correlations between anomalies and anomaly detections via live and historical data.
- Analyze manually detected critical environment anomalies and automate detections.
- Design and implement monitoring and alerting solutions to ensure the health of datacenter critical environments.
- Develop code, scripts, systems, tools, or platforms that automate moderately complex but repetitive operations processes (e.g., monitoring, alerting, deploying products and updates, debugging) at scale; reviews existing automation code and scripts to evaluate reusability, extendibility, and scalability within an organization.
- Collaborate with software development and data engineering teams to ensure that software systems are designed with security, privacy, and reliability in mind.
- Leverage live and historical EPMS/BAS and server/network telemetry to research, Investigate, analyze, design, and deploy solutions for data center critical environment anomaly detections.
- Participate in incident response and post-mortem analysis to identify and address root causes of issues (DRI).
Embody our Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds. Free account required — sign up in 30sApply for this role