Jobs and Careers
GE

Senior Director, Reliability Tools and Practices

GEICO
MD Chevy Chase (Office) - JPS, United Statesfull_timeVerifiedPosted 3 Jan 2024
💰 $315,000/yr($195,000/yr$315,000/yr)

About the role

GEICO Introduction:

GEICO (Government Employees Insurance Company) was founded in 1936 and insures more than 27 million vehicles in all 50 states and the District of Columbia. A member of the Berkshire Hathaway family of companies, GEICO constantly strives to make lives better by protecting people against unexpected events while saving them money. GEICO is one of the nation's largest and fastest-growing auto insurers, known for our low rates, outstanding service, and clever marketing, but we are so much more. GEICO Tech is constantly looking for new ways to anticipate our customers’ needs by asking big questions and challenging ourselves to think outside the box. Join our team as we continue to build and innovate! 

Position Overview:

GEICO is seeking an experienced and visionary technical Director / Senior Director of Reliability Tools and Practices Engineering within the Site Reliability Engineering (SRE) organization. You will play a critical role in ensuring the reliability, availability, and performance of the company’s systems and services by defining, developing, and delivering Reliability Tooling and Platforms to Geico engineering organization. You will lead a team responsible for developing, implementing, and maintaining tools, practices, and processes that enable their organization to achieve world-class reliability and operational excellence. This role combines technical expertise, leadership, and strategic thinking to drive continuous improvement in reliability and scalability.

Key Responsibilities:

Tooling Strategy:

  • Develop and execute a comprehensive tooling strategy to enhance the reliability and scalability of our systems.
  • Identify, evaluate, and implement cutting-edge tools and technologies that align with industry best practices.

Tooling Development:

  • Lead the development and maintenance of custom tools and automation solutions tailored to the specific needs of our SRE teams.
  • Collaborate with cross-functional teams to ensure tooling meets business requirements.

Monitoring and Alerting:

  • Define and implement robust monitoring and alerting practices, ensuring that SRE teams have timely and actionable insights into system performance and issues.

Incident Response:

  • Establish and maintain incident response processes, including incident escalation procedures, post-incident reviews, and incident management tooling.

Capacity Planning:

  • Work closely with Capacity Planning teams to ensure adequate resources are provisioned to meet growing demands and develop tools and practices for proactive capacity management.

Reliability Best Practices:

  • Define and promote reliability best practices across the organization. Lead efforts to improve service-level objectives (SLOs) and error budgeting.

Documentation:

  • Ensure comprehensive documentation of reliability tools and practices, making them accessible to SRE teams and promoting knowledge sharing.

Team Leadership:

  • Build and lead a high-performing team of reliability engineers and tooling specialists.
  • Provide mentorship, guidance, and professional development opportunities.

Vendor Relations:

  • Manage relationships with third-party tooling vendors, negotiate contracts, and stay informed about emerging trends and innovations in the field.

Compliance and Security:

  • Ensure that all reliability tools and practices adhere to security and compliance standards, and drive efforts to continuously enhance security.

Qualifications:

  • Bachelor's or Master's degree in Computer Science, Information Technology, or related field, or equivalent practical experience. Advanced degree preferred. 
  • Proven experience leading full-stack SW development teams preferably within a Site Reliability Engineering, Engineering Productivity, Observability or Workflow Automation domains using agile SW development methodologies and DevOps practices.
  • Strong expertise in defining, developing, and managing reliability tools and automation solutions as product owner.
  • Knowledge of chaos engineering principles and tools.
  • Experience with continuous integration and continuous delivery (CI/CD) pipelines.
  • Proficiency in monitoring, alerting, and incident response tools such as Prometheus, Grafana, ELK, PagerDuty, etc.
  • Experience with cloud platforms (e.g., AWS, Azure, GCP) and container orchestration (e.g., Kubernetes).
  • Familiarity with industry best practices in reliability engineering, including SLOs, error budgets, and incident management.
  • Expertise in incident management processes, including creating incident response playbooks, incident triaging strategies, a

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

GEICO

View company profile →