Jobs and Careers
TH
Site Reliability Engineer (S3NS, an alliance between Thales and Google Cloud)
ThalesRomaniafull_timeVerifiedPosted 5 Dec 2025
About the role
Location: Bucharest, Romania<p></p><p></p><p><span>Thales is a global technology leader trusted by governments, institutions, and enterprises to tackle their most demanding challenges. From quantum applications and artificial intelligence to cybersecurity and 6G innovation, our solutions empower critical decisions rooted in human intelligence. Operating at the forefront of defence and security, aerospace and space, cybersecurity and digital identity, we’re driven by a mission to build a future we can all trust.</span></p><p><span> </span></p><p><span>In Romania, we are advancing innovation through software engineering, research and development, delivering solutions in key markets in which Thales Group operates. Our engineers design, develop and integrate solutions that impact global industries – from fully operational systems and subsystems for naval warfare and maritime security operations, to air traffic management systems, satellite-based solutions, tactical indoor simulations, identity and biometric technologies and more.</span></p><p></p><p><b>About the Role:</b></p><p>S3NS was born from an industrial partnership between Thales, a global leader in cybersecurity, and Google Cloud, a global leader in cloud solutions. Our ambition is to offer the best of both worlds to all organizations concerned with protecting their sensitive data (public institutions, OIVs, OSEs, etc.).</p><p></p><p><b>Your Day-to-Day:</b></p><p>As an SRE Engineer, your mission will be at the heart of operating our GCP universe:</p><ul><li>You will be in charge of the execution and management of the entirety of a GCP universe.</li><li>You will manage 24/7 production incidents and develop a deep understanding of the technical services that form the backbone of the GCP services used by millions of customers.</li><li>You will benefit from an accelerated and intensive training program on GCP technologies to be fully trained on the Google technical stack (e.g., Borg, Colossus, Spanner, Andromeda, etc.). <b>You will also be able to train with experts from Google.</b></li><li>You will be part of an SRE team responsible for operating sovereign GCP services, some of which are directly customer-facing and others which constitute technical infrastructures with availability requirements of 99.99% or more.</li><li>Within this SRE organization, you will have the opportunity to take on the complex challenges associated with the unique scale of a cloud platform, while leveraging your expertise in incident resolution, complexity analysis, and understanding large-scale system design.</li></ul><p></p><p><b>Key Responsibilities:</b></p><ul><li><b>SLI/SLO Monitoring:</b> Monitor the availability, scalability, latency, and efficiency of sovereign GCP services by handling production incidents.</li><li><b>Incident Resolution:</b> Fix problems and ensure system reliability; perform periodic on-call duties according to a <i>follow-the-sun</i> model.</li><li><b>Team Collaboration:</b> Collaborate with GCP service experts worldwide to help mitigate and resolve incidents.</li><li><b>Automation & Knowledge:</b> Document knowledge to ensure all S3NS SREs work with the same information, standardize resolution flows, and improve operational <i>playbooks</i>.</li><li><b>Post-Incident Reviews:</b> After an incident, gather teams to perform a post-mortem, understand the causes, draw lessons, and encourage continuous improvement.</li></ul><p></p><p><i>In accordance with the SecNumCloud qualification requirements for our services, delivered by ANSSI, this position is subject to reinforced security requirements. The successful candidate will have to undergo a security investigation conducted by our services, in accordance with our personnel security policy.</i></p><p></p><p><b>Your Profile</b></p><p><b>What motivates you:</b></p><ul><li><b>Technological innovation</b>, the <b>Cloud</b>, and the operation of services and infrastructure in <b>"as code"</b> mode.</li><li>The operation and management of <b>critical large-scale systems</b> with high availability.</li><li>The curiosity to explore the <b>technical DNA of Google</b>: discovering how a Cloud works at an <i>hyperscaler</i> and mastering the technologies developed over more than 20 years.</li><li>The opportunity to join specialized teams in key areas (Compute, Storage, Data, <span>Observability/Tooling).</span></li></ul><p></p><p><b>Your profile and experience:</b></p><ul><li><b>Education:</b> Graduate of an engineering school or holder of a Master's degree.</li><li><b>SRE & Automation Experience:</b> Minimum <b>three (3) years of proven experience</b> in Site Reliability Engineering and operations automation (Security, Compliance, Problem Resolution).</li><li><b>Regulated Context:</b> Significant experience in <b>highly regulated markets</b> (Banking, Insurance, Medical, etc.).</li><li><b>International Exposure:</b> Exposure to an international environment with an <b>excellent level of English required</b>.</
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s