Senior Site Reliability Engineer
GovXAbout the role
GOVX is seeking an experienced Senior Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our production systems through automation, observability, and operational excellence. This position is remote but must be located in one of the following states: California, Washington, Texas, Tennessee, Florida, Colorado, or New York.
The Senior Site Reliability Engineer (SRE) plays a key role in maintaining resilient infrastructure, monitoring critical services, and improving deployment and recovery processes across environments. The Senior Site Reliability Engineer works under the direction of the Director of Engineering and collaborates closely with Site Reliability Engineers, Automation Engineers, and other members of the engineering organization.
This position will report to the Director of Engineering.
Responsibilities
- Maintain scalable, secure, and reliable cloud services ensuring reliable system operations within Service Level Objectives.
- Implement and manage monitoring, alerting, and observability systems using Prometheus, Grafana, and Azure Monitor to proactively identify and resolve issues.
- Develop and maintain automation scripts and tools in PowerShell, Bash, and C# to improve deployment efficiency, system reliability, and developer productivity.
- Create, refine, and maintain detailed runbooks for production systems to ensure consistent operational procedures and effective incident response.
- Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to measure and maintain system reliability.
- Collaborate with software engineers and automation engineers to integrate reliability practices into CI/CD pipelines using Azure DevOps.
- Design and implement intelligent alerting strategies that ensure high signal-to-noise ratios and enable rapid triage of critical issues.
- Participate in incident response, post-incident reviews, and blameless root cause analysis to drive continuous improvement of system reliability and uptime.
- Contribute to deployment strategy evolution, including blue-green and canary deployments, to minimize downtime and release risk.
- Collaborate closely with Automation Engineers to enhance automated validation and testing of production environments.
- Monitor system health, capacity, and performance, providing data-driven insights and recommendations for optimization.
- Conduct chaos engineering experiments and resilience testing to proactively identify and address system weaknesses.
- Develop and maintain disaster recovery and business continuity plans, including regular failover testing.
- Participate in the on-call rotation for platform services, ensuring high availability and rapid incident resolution.
- Proactively monitor and respond to production support tickets and alerts within established SLA timeframes, delivering first-level diagnosis, troubleshooting, and escalation as needed to maintain system reliability
- Continuously improve incident response playbooks and reduce Mean Time to Recovery (MTTR).
- Participate in sprint planning, stand-ups, and retrospectives to ensure alignment with development and operational objectives.
- Identify opportunities to improve resiliency, reduce toil, and strengthen the reliability culture across the engineering organization.
- Collaborate with security and compliance teams to ensure infrastructure meets regulatory and security standards.
- Support cost optimization efforts by monitoring cloud resource usage and recommending efficiency improvements.
- Explore and integrate AI/ML-based observability tools for predictive monitoring and anomaly detection.
Requirements
- 8+ years of professional experience in site reliability, infrastructure, or systems engineering roles.
- Proficiency with Azure cloud infrastructure, services, and resource management
- Experience in operating systems, network concepts, protocols, and architecture. Microsoft/Linux operating systems, active directory, OSI.
- Technical ability in Node JS, .NET/C# and knowledge of both current and legacy architecture, software development practices, and conventions.
- Strong experience with Rest APIs
- Hands-on experience with containerization and orchestration using Kubernetes and microservices architecture.
- Strong automation and scripting skills in PowerShell, Bash.
- Experience with Infrastructure as Code tools for provisioning and configuration management.
- Deep understandin
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s
Similar roles
Senior Mechanical Engineer
General Dynamics Mission Systems
$125,275/yr
Product Deliverable Lead (PDL)- Trainer Documentation (Senior Advanced Integrated Logistics Support Specialist)
General Dynamics Mission Systems
$128,856/yr
Senior Director, Transitional Housing
Samaritan Daytop Village
$123,000/yr