Jobs and Careers
SY

Stability and Resilience Engineer with Python Dev Experience

Synechron
Alpharetta, United Statesfull_timeVerifiedPosted 8 Oct 2024
💰 $110,000/yr($100,000/yr$110,000/yr)

About the role

We are

At Synechron, we believe in the power of digital to transform businesses for the better. Our global consulting firm combines creativity and innovative technology to deliver industry-leading digital solutions. Synechron’s progressive technologies and optimization strategies span end-to-end Artificial Intelligence, Consulting, Digital, Cloud & DevOps, Data, and Software Engineering, servicing an array of noteworthy financial services and technology firms. Through research and development initiatives in our FinLabs we develop solutions for modernization, from Artificial Intelligence and Blockchain to Data Science models, Digital Underwriting, mobile-first applications and more. Over the last 20+ years, our company has been honored with multiple employer awards, recognizing our commitment to our talented teams. With top clients to boast about, Synechron has a global workforce of 14,000+, and has 55 offices in 20 countries within key global markets.

Our challenge

We are seeking a highly skilled Stability and Resilience Engineer to join our dynamic team. This role is crucial for ensuring the stability, reliability, and resilience of our systems and applications, with a focus on proactive measures to prevent downtime and mitigate risks. The ideal candidate will have a strong technical background, a passion for problem-solving, and a commitment to excellence in system performance.

Additional Information* 

The base salary for this position will vary based on geography and other factors. In accordance with law, the base salary for this role if filled within Alpharetta, GA is $100K - $110K/year & benefits (see below).

The Role

Responsibilities:

1.       System Monitoring and Analysis:

  • Implement and manage monitoring solutions to track system performance and health.

  • Analyze system metrics to identify patterns and irregularities that may indicate potential failures.

2.       Incident Management:

  • Develop and execute incident response plans to address system outages and failures.

  • Conduct root cause analysis to understand incidents and implement corrective actions.

3.       Collaboration:

  • Work closely with development, operations, and other engineering teams to ensure systems are designed with stability and resilience in mind.

  • Participate in architectural reviews and design discussions to provide insights on resilience best practices.

4.       Documentation and Reporting:

  • Maintain thorough documentation of system configurations, incident reports, and resilience testing results.

  • Prepare reports and presentations to communicate system performance, risks, and improvement strategies to stakeholders.

5.       Continuous Improvement:

  • Stay current with industry trends, tools, and techniques in system stability and resilience.

  • Propose and implement innovations to enhance system performance and reliability.

Requirements:

You are:

  • Bachelor’s degree in Computer Science, Engineering, or related field; Master’s degree preferred.

  • Proven experience in systems engineering, reliability engineering, or a similar role.

  • Strong understanding of system architecture, cloud infrastructure, and networking concepts.

  • Proficiency in monitoring tools and techniques (e.g., Prometheus, Grafana, Dynatrace, etc.).

  • Experience with incident management and response frameworks (ITIL, SRE principles).

  • Knowledge of scripting and programming languages (e.g., Python, Bash, Go).

  • Familiarity with DevOps practices and CI/CD pipelines.

  • Excellent problem-solving skills and the ability to work under pressure.

  • Strong communication and collaboration skills, with an ability to work effectively in a team environment.

It would be great if you also had:

  • Relevant certifications (e.g., AWS Certified Solutions Architect, Google Cloud Professional DevOps Engineer, Certified Kubernetes Administrator).

  • Experience with container orchestration and microservices architecture.

  • Previous involvement in disaster recovery planning and execution.

We can offer you:

  • A highly competitive compensation and benefits package

  • A multinational organization with 55 offices in 20 countries and the possibility to work abroad

  • Lapt

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Synechron

View company profile →