Jobs and Careers
GU

Lead Site Reliability Engineer

Guild Mortgage
San Diego, United StatesRemotefull_timeVerifiedPosted 22 Feb 2025
💰 $173,000/yr($127,000/yr$173,000/yr)

About the role

Guild Mortgage Company, closing loans and opening doors since 1960. As a mortgage banking firm, we are dedicated to serving the homeowner/buyer. Our goal is to provide affordable home financing for our customers, utilizing the best terms available while providing a level of professionalism and service unsurpassed in the lending industry.

Position Summary

The Lead Site Reliability Engineer is responsible for driving the organizational reliability strategy and conducting resiliency design reviews to ensure the reliability, scalability, and performance of our company's software systems and applications meet organizational service level objectives (SLOs) and error budgets. The role is responsible for leading a team of Site Reliability Engineers in designing, implementing, and maintaining the infrastructure and tools necessary to support our platforms, as well as improving our monitoring, automation, and deployment processes. This role involves strategic planning, technical leadership, and collaboration with various stakeholders including Guild’s Product Delivery, Data Services, DevOps, DataOps, Governance, and Infrastructure teams to support organizational goals.

Essential Functions

  • Lead, mentor, and develop a team of Site Reliability engineers, fostering a collaborative and innovative work environment.
  • Oversee an SRE team and drive the reliability strategy for the organization.
  • Conduct resiliency design reviews and lead complex problem-solving efforts. 
  • Design, implement, and maintain monitoring systems to track the performance, availability, and reliability of services.
  • Respond to incidents promptly, investigate root causes, and coordinate efforts to mitigate and resolve them.
  • Analyze performance data, and plan for scalability and capacity requirements.
  • Identify and optimize performance bottlenecks, both at the infrastructure and application levels.
  • Automate repetitive tasks and processes to improve efficiency and reduce manual intervention.
  • Implement and enforce change management practices to ensure safe and controlled changes to the production environment.
  • Design and implement fault-tolerant systems and practices to minimize downtime and ensure service availability.
  • Collaborate with the GRC team on developing and maintaining disaster recovery plans and procedures relevant to the software supported to minimize the impact of catastrophic failures.
  • Work with the Incident Management and other teams to conduct a thorough analysis of incidents, document postmortem reports, and implement improvements based on lessons learned.
  • Work closely with development, operations, and other teams to foster a culture of reliability, and provide feedback on system design and architecture for improved reliability.
  • Perform other duties as assigned.

    Qualifications

    • Bachelor's Degree directly related to the position or equivalent, preferred. Bachelor's degree or 8+ years demonstrated work experience or an equivalent combination of related training and experience and at least three of those years spent in a leadership level role(s) required.
    • Minimum of eight years demonstrated work experience.
    • Minimum three years supervisory or leadership experience. Proven leadership experience and ability to manage a team, required.
    • Ability to create DR strategies and execute DR drills.
    • Collaborate with stakeholders to define RPO / RTO for Guild’s system footprint.
    • Expert in Cloud-based redundancy, high availability, and reliability strategies.
    • Expert in reliability, scalability, and performance optimization.
    • Expert at maintaining Linux / Unix and Windows systems administration, provisioning, configuration, monitoring, and troubleshooting Web Servers in a 7x24 customer facing environment.
    • Strong Linux and Windows Administration & scripting.
    • Solid Database Administration skills (MySQL, MariaDB, RDS, Sql Server, and Azure Storage services).
    • Deep knowledge of current methodologies in high performance operations and scalable multi-site implementations.
    • Proven Experience with large-scale software implementation (high transaction volume, high- availability concepts).
    • Deep knowledge of software deployment, versioning (GIT) and release management processes.
    • Deep knowledge with infrastructure design, implementation, and support.
    • Proficient at automated provisioning, automated configuration management, and containerization solutions and tools.
    • Experienced in cloud-based hosting solutions (AWS, Azure, GCP).
    • Experienced with Cloud server environments (AWS, Google Cloud, or Azure).
    • Experienced in Agile software development best practices uti

    Apply for this role

    Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

    Apply Now →Generate Application Kit

    Free account required — sign up in 30s

    Company

    Guild Mortgage

    View company profile →