Sr. Site Reliability Engineer (SRE)
SteampunkAbout the role
Overview
Design. Disrupt. Repeat. Be an agent of change on a team committed to achieving client-focused, mission-driven excellence. Steampunk is looking for an experienced Site Reliability Engineer with an appetite for taking on new challenges. Who We Are Steampunk is the explosive collision of human-centered design and traditional government contracting. An employee-owned company with a startup mindset and time-tested approaches tailored for the federal government, we’re passionate about creating solutions that are impactful, practical, scalable, and most importantly, that meet our clients’ ever-changing needs. At Steampunk, we believe in disrupting the status quo and setting the pace in the ecosystem of government contractors, while repurposing tried-and-true methodologies. We believe in empowering our people to find creative solutions to intractable problems. We believe the best environment in which to grow and thrive is outside our comfort zone. While good design makes for a good product, we believe human-centered design makes for an excellent one. We also believe effective teams are powered by diverse perspectives, backgrounds, and experiences. To that end, Steampunk is an equal opportunity employer committed to promoting diversity of race, gender, sexual orientation, religion, ethnicity, national origin, disability status, and protected veteran status, amongst our ranks. Additionally, we participate in the E-Verify program. Why Steampunk? Our people are the very core of what we do; their expertise and hunger for new and exciting challenges fuel our relentless pursuit of mission success. As part of our team of “Punks,” you’ll test the status quo, explore new boundaries, and set the bar high for how government clients expect to engage with contractors. Because we value our employees’ work/life balance (and believe those who work hard deserve to play hard), we offer a very competitive benefits package, including telework/flex scheduling, health/dental with orthodontics/vision insurance upon hire, paid time off with a sell-back benefit and carryover option, 11 Federal Holidays, 100% paid military leave, 100% 401(k) plan match upon hire, professional development/education reimbursement, all flexible spending accounts, and more.
Contributions
As a Sr. Steampunk Site Reliability Engineer (SRE), you will be responsible for working with program development teams, infrastructure and platform services teams, and traditional operations and maintenance teams to embrace and embody a shared responsibility for the reliability of an organizations’ applications and infrastructure. As an SRE, your primary responsibility is to combine aspects of software engineering with traditional operations to maintain and improve the reliability, availability, and performance of cloud, infrastructure, and large-scale software systems and services while minimizing downtime and mitigating potential failures.
There are a wide variety of responsibilities you will be delivering in this role:
- Infrastructure Optimization: Conduct in-depth analyses of infrastructure, identifying areas for improvement in terms of performance, scalability, and resource utilization. Collaborate with development and operations teams to implement enhancements, utilizing software engineering and/or infrastructure-as-code principles to streamline deployment processes and ensure consistency across environments.
- Reliability Metrics and Reporting: Define and implement key reliability metrics, service-level objectives (SLOs), and service-level indicators (SLIs) to measure and report on the health of our systems. Establish monitoring and alerting mechanisms to proactively identify potential issues before they impact users.
- Automation and Tooling: Design and implement automation tools to reduce manual toil, streamline repetitive tasks, and enhance overall operational efficiency. Leverage software development techniques to create robust, scalable tooling that supports our reliability goals, and collaborate with development teams to integrate reliability features into the development lifecycle.
- Performance Optimization using Software Development Techniques: Collaborate with software development teams to optimize the performance and resilience of services through code improvements, architectural enhancements, and performance tuning. Integrate automated testing and profiling into the development pipeline to identify and address performance bottlenecks early in the development lifecycle.
- Capacity Planning and Scaling: Collaborate with infrastructure teams to forecast capacity requirements, ensuring our systems can seamlessly scale to meet growing user demands. Implement strategies for auto-scaling and load balancing to optimize resource
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s