Jobs and Careers
AM

Senior Service Reliability Engineer

Amadeus
United Statesfull_timeVerifiedPosted 30 Jun 2024

About the role

Job Title

Senior Service Reliability Engineer

Summary of the role:

As part of the Amadeus Global Operations Americas organization, the Senior Service Reliability Engineer is responsible to support revenue generating systems in production environments. The engineer works closely with other SREs in the monitoring, maintenance, and support related to incident and problem resolution of Amadeus Hospitality production applications. The position is responsible to support the solution full stack including all technology layers comprised in the proprietary software hospitality solutions and platform. The Senior Service Reliability Engineer reports to the Director of Platform and SRE and partners with the Architect, System, Network, Cloud and Platform Engineering teams while leading work on projects or day to day operational activities aimed to ensure the reliability of the Hospitality production applications. The candidate must be able to provide expert and prompt technology operations support in a high energy, fast paced environment. He or she will serve as senior resource and mentor to other Service Reliability Engineers. The role requires to work autonomously within defined processes and procedures. The role contributes to build an healthy and collaborative environment, leading by example.

In this role you'll:

  • Provide support related to production systems availability incidents and problems.
  • Provide support related to production systems latency incidents and problems.
  • Provide support related to production systems performance incidents and problems.
  • Provide support related to production operations efficiency issues.
  • Provide emergency response to production systems incidents.
  • Demonstrates advanced knowledge of job-relevant issues, products, systems, and processes.
  • Demonstrates advanced knowledge of function-specific procedures.
  • Applies knowledge/judgment to achieve business goals.
  • Foresees, identifies, and resolves problems.
  • Keeps up-to-date technically and applies new knowledge to job.
  • Participate in On-Call production support
  • Performs other reasonable duties as required for this position.
  • Manage to help ensure application incident metrics trend in the right direction (e.g., minimal response times, incident backlogs, and reliance on other teams while resolving incidents).
  • Help drive feedback loops from Incident and problem Management back into Delivery teams, so that quality improves over time.
  • Ensure all application monitoring are implemented and updated appropriately.
  • Determine need for new Problem tickets or association with Known Errors or other existing Problems, to help advance continual operational improvement.
  • Automate repetitive tasks
  • Drive technical projects from conception to completion
  • Develops specific goals and plans to prioritize, organize, and accomplish work.
  • Provides direction and assistance to other teams regarding projects.
  • Understands and meets the needs of key stakeholders.
  • Communicates concepts in a clear and persuasive manner that is easy to understand.
  • Demonstrates an understanding of business priorities.
  • Supports achievement of performance goals, budget goals, team goals, etc.
  • Analyzes information and evaluates results to choose the best solution and solve problems.
  • Generates and provides accurate and timely results in the form of reports, presentations, etc.
  • Plans, develops, implements, and evaluates the quality of operations.


About the Ideal Candidate:

  • Bachelor's degree educated or equivalent
  • 8+ years automated operational deployment and support
  • English; French and/or Spanish as plus
  • Expert Linux operating systems and platforms such as RedHat/CentOS.
  • Strong Windows operating systems experience.
  • Extensive experience with troubleshooting web server technologies.
  • Extensive experience with any middleware such as Tomcat, Jboss or other application server.
  • Extensive experience with any monitoring, alerting, or pipeline analysis tool such as Datadog, Splunk, Prometheus, Zabbix and/or SCOM.
  • Strong understanding of network technology concepts and usage.
  • Expert level understanding of Linux shell scripting.
  • Outstanding ability to troubleshoot issues throughout the application and infrastructure stack.
  • Strong experience with version control and repositories such as Git and Bitbucket.
  • Strong experience developing scripts such as Bash, Python and Ruby.
  • Experience with Automation solutions like Ansible and/or Chef.
  • Expert understanding of Java application servers.
  • Expert experience with Linux RPM repositories such as YUM, PULP,and  RedHa

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Amadeus

View company profile →