Jobs and Careers
T.

Lead Site Reliability Engineer (SRE)

T. Rowe Price
Maryland, United States, United Statesfull_timeVerifiedPosted 22 Apr 2024

About the role

Department: CDO Technology Group 

Summary: 

We are seeking a highly motivated and experienced Lead Site Reliability Engineer (SRE) to join our CDO Technology Group. As an SRE, you will play a crucial role in ensuring the availability, latency, performance, efficiency, and stability of our critical infrastructure, which supports a range of data platforms, applications, and services. You will collaborate closely with development teams to implement and maintain reliable and scalable systems while adhering to industry best practices and security standards. 

Responsibilities: 

Availability

  • Proactively monitor and proactively identify potential issues that could impact the availability of our systems. 

  • Implement and maintain automated alerting mechanisms to notify the appropriate parties of potential outages or performance degradation. 

  • Collaborate with development teams to design and implement solutions that enhance system resilience and reduce downtime. 

Latency: 

  • Analyze performance metrics to identify and resolve latency bottlenecks in our infrastructure. 

  • Implement performance optimization techniques and tools to improve the overall responsiveness of our systems. 

  • Work with development teams to ensure that new features and code changes do not introduce performance regressions.

Performance:

  • Develop and maintain metrics dashboards to track key performance indicators (KPIs) for our critical systems. 

  • Identify performance trends and anomalies that may indicate potential issues or areas for improvement. 

  • Recommend and implement performance optimization strategies to enhance the overall efficiency of our systems. 

Efficiency: 

  • Optimize resource utilization and minimize unnecessary expenditure on IT infrastructure. 

  • Collaborate with development teams to optimize resource allocation for new applications and services. 

Release Management:

  • Participate in the release planning process to ensure that software releases are conducted smoothly and without disruptions. 

  • Develop and implement automated deployment and rollback procedures to mitigate risks associated with software updates. 

  • Monitor the performance of new releases and address any issues that arise promptly. 

Monitoring:

  • Design, implement, and maintain a comprehensive monitoring infrastructure to track the health and performance of our systems. 

  • Analyze monitoring data to identify potential issues and proactively troubleshoot problems before they impact users. 

  • Develop and implement alerts and notifications for critical events to ensure timely intervention. 

Emergency Response:

  • Respond promptly to incidents and work collaboratively to resolve them in a timely manner. 

  • Analyze root causes of incidents to identify and implement preventive measures to minimize their recurrence. 

  • Document incident responses and lessons learned to enhance our incident handling processes.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

T. Rowe Price

View company profile →