Senior Software Engineer - CTJ - Poly
MicrosoftAbout the role
As stewards of a rapidly growing and evolving cloud ecosystem, Azure Cloud+AI builds the foundation of the Azure cloud platform on commodity datacenter hardware to deliver infrastructure and services for hosting customer applications and data. Within this space, the Capacity Infrastructure Service (CIS) team provides an automation platform for capacity provisioning functions and global buildout to meet increasing and changing capacity demands in a consistent, efficient, safe, and secure way.
As a Senior Software Engineer for CIS, you will be responsible for growing and maintaining the capabilities of our platform through design, deployment, and operation of related services within an airgapped environment. This opportunity enables you to collaborate with stakeholders across the infrastructural layer of Azure, learn what it takes to build and manage a cloud at Azure-scale, and work on something highly strategic to Microsoft and extremely relevant in the industry.
Microsoft’s mission is to empower every person and organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
- Creates, implements, optimizes, debugs, refactors, and reuses code to establish and improve performance and maintainability, effectiveness, and return on investment (ROI).
- Applies debugging tools and examines logs, telemetry, and other methods to verify assumptions through writing and developing code proactively before issues occur and reactively as issues occur for products. Conducts retrospective debugging of solutions to identify root causes of problems.
- Acts as a Designated Responsible Individual (DRI) and guides other engineers by developing and following the playbook, working on call to monitor system/product/service for degradation, downtime, or interruptions, alerting stakeholders about status and initiates actions to restore system/product/service for simple and complex problems when appropriate.
- Maintains operations of live service as issues arise on a rotational, on-call basis. Implements solutions and mitigations to more complex issues impacting performance or functionality of Live Site service and escalates as necessary. Reviews and writes issues postmortem and shares insights with the team.
- Drives identification of dependencies and the development of design documents for a product, application, service, or platform.
- Leverages subject-matter expertise of product features and partners with appropriate stakeholders (e.g., project managers) to drive a workgroup's project plans, release plans, and work items.
- Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of products while also driving consistency in monitoring and operations at scale.
- Embody our
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s