Senior Service Reliability Engineer
AmadeusAbout the role
Job Title
Senior Service Reliability EngineerWorking on cloud infrastructure and CI/CD pipelines to ensure reliability, security and strategical evolution, as well as delivery of product features.
As part of Amadeus Hospitality, the main objective of the Media division is to drive demand and boost client market shares by advertising their products in all the possible mediums and channels (web searches, web, social, GDS, mobile). Our customers are hotels who aim to attract more guests, cities or regions who want to attract more visitors as well as airlines who strive to attract more travellers.
Along the years, Amadeus has created a leading data-driven travel advertising platform complemented by strategic partnerships with advertising giants (Google, Facebook, Microsoft, … to name a few), that can reach billions of travellers globally every day. The Media division is in fast expansion aiming to reach a billion in revenue in the next couple of years. Our engaged teams are key to achieve this goal and we want YOU to be part of the adventure.
Unlike many media agencies which only rely on third party software to run advertising campaigns, Amadeus has created its own advertising platform praised by our customers, for handling unique use cases and allowing a more targeted approach. This platform was built with a microservices architecture, using some of the most recent technologies like Golang, Scala, Python, gRPC, GraphQL, Kafka, Apache Airflow, Apache Beam, Aerospike, Redis, PostgreSQL. It reaches 160k transactions per second, relying on a PB-size Google BigQuery hosted data warehouse, an innovative AI running up to 250k predictions per second and is hosted on GCP (Google Cloud Platform). Our machine learning models are built with Python, Tensorflow, Keras, LightGBM and more.
Our infrastructure also follows the recent standards of the industry (Terraform, Kubernetes, Helm, FluxCD, GitHub Actions, CircleCI, GCP [Google Cloud Platform], GKE [Google Kubernetes Engine], Thanos, Prometheus, Grafana, Apache Airflow / Google Cloud Composer, Apache Beam / Google Cloud Dataflow, Google BigQuery, Aerospike, Redis, PostgreSQL, Google BigTable, ...) and is built with a GitOps mindset. Our SRE team is in charge of our CI/CD pipelines, too, in a healthy and close collaboration with the developer teams.
The next step the evolution of our platform is increasing our geographical footprint by deploying our stack to other continents, as well as gradually bringing parts of the Hospitality advertising business to this platform.
As Senior Service Reliability Engineer, your contribution will be paramount in all the future major transformations, on the stability and on the security of the platform.
The responsibility of the SRE team covers the following:
On-going evolution of the infrastructure and implementation of new standards
Support projects of geographical expansion
Monitoring and support of the infrastructure
Support and mentor the R&D teams
Ensure the stability of the platform
Ensure security of the platform
Maintain and evolve CI/CD
In this role you will:
Maintain and evolve our production and staging infrastructure.
Manage the full application stack on Google Cloud.
Participate in the analysis of new requirements and develop core services to support and accelerate our IT teams productivity.
Bring SRE expertise to evaluate and productionise new services and data processing solutions.
Provide technical guidance and expertise to the organisation.
Participate to the construction of multi-year roadmaps for infrastructure evolution and industrialisation of processes.
Focus on developing systems automation and provisioning frameworks to increase delivery teams autonomy and ownership.
Mentor, support and coach IT team members regarding tools, concepts and cloud best practices. Effectively communicate with team members, peers, and management.
About the ideal candidate:
Has excellent written and oral communication skills (English).
Demonstrates strong analytical thinking.
Possesses strong troubleshooting skills – we are usually the last line of support for all infrastructure issues.
Has hands-on experience with Kubernetes (or OpenShift) and good understanding of Kubernetes concepts & the Kubernetes ecosystem.
Has prior experience working with one of the major cloud providers: AWS, GCP, Azure.
Is acquainted with concepts and practical aspects of the Prometheus based monitoring stack (PromQL, Prometheus, Thanos, Grafana).
Is familiar with CI/CD concepts & to
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s