Jobs and Careers
GE
Staff Engineer – SRE (Business Continuity and Disaster Recovery)
GEICOChevy Chase, United Statesfull_timeVerifiedPosted 22 Aug 2024
💰 $260,000/yr($115,000/yr – $260,000/yr)
About the role
Position Responsibilities
As a Staff Engineer, you will:
- Develop and drive the overall strategy for Business Continuity and Disaster Recovery (BCDR), aligning it with the organization's business goals and objectives
- Provide thought leadership in BCDR, staying ahead of industry trends and emerging technologies to enhance our resilience posture
- Conduct comprehensive risk assessments to identify potential threats and vulnerabilities
- Design and implement robust risk mitigation strategies and plans to ensure continuous business operations
- Lead the design and architecture of resilient and scalable systems, considering both on-premises and cloud-based solutions.
- Collaborate with cross-functional teams to integrate BCDR considerations into the development and deployment processes
- Develop and maintain comprehensive incident response plans to address various disaster scenarios
- Conduct regular simulations and drills to ensure the readiness of the organization in the event
of a disaster
- Hands-on software engineering and SDLC best practices (Technical Review Documents, Architecture, Software Development, Software Reviews, Testing, Production Readiness Reviews, among others)
- Evaluate, select, and implement cutting-edge technologies and tools to enhance our BCDR capabilities including but not limited to processes, compliance, and visibility
- Stay current with industry best practices and emerging technologies to continuously improve our BCDR capabilities
- Work closely with executive leadership, IT teams, and other stakeholders to communicate the importance of BCDR and foster a culture of resilience.
- Act as a trusted advisor, providing guidance on BCDR matters to technical and non-technical stakeholders.
- Be a role model and mentor, helping to coach and strengthen the technical expertise and knowledge of our engineering and product community. Influence and educate executives
- Analyze cost and forecast, incorporating them into business plans
- Determine and support resource requirements, evaluate operational processes, measure outcomes to ensure desired results, and demonstrate adaptability and sponsoring continuous learning
Qualifications
- Fluency and specialization in software development and best practices using modern programming languages such as Go, Java, and Python
- Understanding of SQL and NoSQL databases, including stateful services management and storage
- Understanding of networking, caches, key/value stores, load balancing, global load balancing, queues, DNS and CDN.
- Deep knowledge of SRE practices, methodologies, and principles, along with a solid understanding of on prem and public cloud-based network, compute, and storage technologies
- In-depth knowledge of hybrid cloud architecture, IaaS and PaaS technologies, container orchestration platforms (e.g., Kubernetes), cloud efficiency and observability etc.
- Strong background in incident management
- Ability to create incident response playbooks, runbooks, incident triaging strategies, and post-incident analysis to drive continuous improvement in system reliability and availability
- Experience with open-source management and monitoring tools
- Experience with infrastructure automation, tooling, and configuration management frameworks (e.g., Puppet, Chef, Ansible, Terraform, Pulumi, etc.)
- Familiarity with cloud security best practices and compliance standards
- Excellent leadership skills with a passion for mentoring and fostering professional growth
- Strong problem-solving and analytical abilities, with a keen eye for detail and a passion for driving operational excellence
- Visionary thinker with the ability to anticipate future challenges and opportunities
- Exceptional leadership and communication skills
- Strong analytical and problem-solving capabilities
- Proven record of accomplishment of successfully leading and building software in large and complex organizations
- One or more of the following or relevant certifications are highly desired:
- Certified Information Systems Security Professional (CISSP)
- Certified Business Continuity Professional (CBCP)
- AWS Certified DevOps Engineer
- AWS Cloud Practitioner
- Google Professional Cloud DevOps Engineer
Experience
- 10+ years of professional experience in software engineering
- 8+ years of experience with architecture and design
- 6+ years of experience in open-source frameworks
- 4+ years of experience with AWS, GCP, Azure, or another cloud service
Education
- Bachelor's degree in computer science, Information Systems, or equivalent education or work experience
#L
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s