Advanced Site Reliability Engineer (SRE)
Federal Reserve SystemAbout the role
Company
Federal Reserve Bank of RichmondWhen you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future.Bring your passion and expertise, and we’ll provide the opportunities that will challenge you and propel your growth—along with a wide range of benefits and perks that support your health, wealth, and life. In addition to competitive compensation, we offer a comprehensive benefits package that includes tuition assistance, generous paid time off, top-notch health care benefits, child and family care leave, professional development opportunities, a 401(k) match, pension, and more. All brought together in a flexible work environment where you can truly find balance.
About the Opportunity
As an Advanced SRE, you will be accountable for the development, maturation, and iterative implementation of SRE practices for the cloud foundational product line in the Federal Reserve.
The SRE team is part of the Cloud Operations department and has the overall responsibility for the operability and reliability of the numerous cloud foundational environments in the FRS. The team is responsible to define and implement best practices for observability, establish and maintain service level indicators (SLIs) and service level objectives (SLO); error budgets and error budget policies; tracking and addressing toil, conducting blameless post-mortems and incorporating preventative and proactive SRE practices in the SDLC for continual improvement.
The SRE team interfaces with internal System IT Application Delivery/Development Teams (ADTs) and other stakeholders for planning, delivery, and service management. Team owns the implementation and driving of continuous operational improvement initiatives working closely with Architects, Engineers, peer product teams, as well as marketing and cloud program teams.
What Will Be Expected of You
- As an Advanced SRE and primary SRE Engineer, work closely with leaders to establish and iteratively implement the SRE practice. Gain insights into product operations, build relationships and influence SRE ways of working in product teams
- Lead the establishment of SLIs, SLOs, Error budgets, policies and work with respective engineers to instrument, visualize and offer a means for peer engineers and developers to gain greater insight into operational performance (Observability)
- Develop and Mentor Junior SREs
- Creating and maintaining automation, scripts and code associated with improved operations.
- Automate all facets of operations in compliance with security standards.
- The ideal candidate is someone who loves building and maintaining reliable and scalable systems, is passionate about continual improvement.
- The candidate has a understanding of ITSM operations and has proven experience being a key player in the transformation of traditional operations to cloud, SRE and DevOps.
- Identify, track and address Toil,
- Conduct Post-Mortems
- Identify and implement continuous improvement in various facets of production operations.
- Offer advanced technical support for cross product issues/incidents
- Leveraging SRE tooling to develop, implement and deliver on SRE mission.
- Conduct Chaos Testing
- Stay current with industry trends and source new ways for our business to improve
- Other duties assigned as necessary
Qualifications
- 3-5+ years of Site Reliability Engineering experience. DevOps a plus
- 3-5+ years of supporting production cloud environments
- Bachelor’s degree in computer science, Information Systems, or equivalent background or equivalent experience.
- Strong analytic and problem-solving skills.
- Team player and influencer
- Self-motivated individual with the ability to prioritize and manage changing priorities.
- Strong customer service and communication skills.
- Independent critical thinking and decision-making abilities.
- Excellent written and oral communication abilities.
Expertise you will bring:
- Extensive knowledge and understanding of working in AWS & Azure environments & services
- Experience with Gitlab, Service Now, Dynatrace
- Proficiency in scripting/programming languages (GOLANG, Python, Perl, R)
- Dashboarding and visualization
- Experience as an advanced SRE in support of a cloud environment
- Experience supporting infrastructure for large multi-services applications.
- Experience working with continuous deployment in micro-services a
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s