Site Reliability Engineer Jobs in Paris, EEA

28 verified site reliability engineer openings in Paris

  • Senior Site Reliability Engineer (x/f/m)

    Doctolib · Paris, France

    On-site
    about 10 hours agoApply →

    <h2><strong>Your Impact</strong></h2> <p>We are looking for a <strong>Senior Site Reliability Engineer</strong> to join our SRE team dedicated to platform reliability within Platform Engineering.</p> <p>Your mission will be to ensure Doctolib's platform remains reliable, scalable, and resilient at a European scale across infrastructure, observability, and cross-cutting reliability initiatives. You will work within a team driving reliability standards across 170+ applications, contributing directly to supporting 520,000 health professionals and 90 million patients in their daily healthcare journey.</p> <p>Working in the tech team at Doctolib means taking ownership of critical systems, driving reliability improvements end-to-end, and partnering closely with product and engineering teams to enable fast, safe delivery.</p> <h2><strong>What you'll do</strong></h2> <p>Your responsibilities include but are not limited to:</p> <ul> <li>Build and maintain infrastructure automation and infrastructure-as-code at scale, with a strong focus on consistency, reliability, and developer experience across 170+ applications</li> <li>Identify and lead large-scale cross-cutting reliability initiatives, including improvements to incident detection, response, and postmortem analysis capabilities</li> <li>Design, build, and improve infrastructure components that support reliability, scalability, and observability across the platform</li> <li>Define and drive SLOs, error budgets, and alerting standards across multiple product teams</li> <li>Take part in the on-call rotation, and actively contribute to improving our on-call experience by reducing noise and ensuring actionable telemetry</li> <li>Partner with software engineering teams to embed reliability practices early in the development lifecycle</li> </ul> <h2><strong>Who you are</strong></h2> <p>Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.</p> <p><strong>You'll be a great fit if you:</strong></p> <ul> <li>Have solid hands-on experience (5y+) in a Site Reliability Engineering role within a large-scale, multi-team production environment</li> <li>Have proven experience with cloud platforms such as AWS, GCP, or Azure</li> <li>Have strong experience with containerization and orchestration technologies, Kubernetes is a must, its deployment and scaling strategies ecosystem</li> <li>Have experience working closely with software engineering teams to co-design reliable systems</li> <li>Have hands-on experience with infrastructure as code, particularly Terraform, applied in real production

  • Senior Site Reliability Engineer

    Photoroom · Paris, France

    On-site
    9 days agoApply →

    Senior Site Reliability Engineer at Photoroom. Apply via Ashby.

  • Senior Site Reliability Engineer

    Alice bob · Paris, Paris

    On-site
    10 days agoApply →

    Alice & Bob is developing the first universal, fault-tolerant quantum computer to solve the world’s hardest problems. The quantum computer we envision building is based on a new kind of superconducting qubit: the Schrödinger cat qubit 🐈‍⬛. In comparison to other superconducting platforms, cat qubits have the astonishing ability to implement quantum error correction autonomously!   We're a diverse team of 250+ brilliant minds from over 35 countries united by a single goal: to revolutionise computing with a practical fault-tolerant quantum machine. Are you ready to take on unprecedented challenges and contribute to revolutionising technology? Join us, and let's shape the future of quantum computing together! About the role We're a fast-moving scale-up and our infrastructure needs to keep pace. That’s why we're looking for a Senior Site Reliability Engineer to own the reliability, automation, and security of our GCP platform, and to set the bar for how we run production as we grow. You'll spend most of your time building and operating systems, but you'll also mentor more junior engineers and shape the practices the rest of the team works by. If you like having real ownership, moving quickly without breaking things, and turning manual toil into automation, you'll fit right in.

  • Software Engineer - Site Reliability Engineering

    Qube Research & Technologies · Paris, France

    On-site
    19 days agoApply →

    <p> </p> <p>Qube Research & Technologies (QRT) is a global quantitative and systematic investment manager, operating in all liquid asset classes across the world. We are a technology and data driven group implementing a scientific approach to investing. Combining data, research, technology, and trading expertise has shaped our collaborative mindset, which enables us to solve the most complex challenges. QRT’s culture of innovation continuously drives our ambition to deliver high quality returns for our investors.</p> <p><span data-teams="true">QRT is seeking a Software Engineer with strong Python skills and a deep understanding of systems design and reliability principles to help evolve and scale our high-performance trading platform. You'll build tools, automation, and design systems that support the speed and reliability our business depends on.</span></p> <p><strong>Your future role within QRT</strong></p> <ul> <li>Collaborate with developers, researchers, traders, and platform teams to build and improve business-critical systems</li> <li>Build automation for operations, deployment, monitoring, and incident response</li> <li>Evaluate and implement tools that balance performance, complexity, and maintainability</li> <li>Own and evolve our observability stack: metrics, logs, tracing, and alerts</li> <li>Participate in the team oncall rota, leading incident response and driving continuous improvement through postmortems and root cause analysis</li> </ul> <p><strong>Your Present Skill Set</strong></p> <ul> <li>Bachelor's degree in Computer Science or a strongly related field</li> <li>Strong Python programming skills</li> <li>Strong problem-solving skills and deep understanding of computer science fundamentals</li> <li>Commitment to code clarity, documentation, and reducing technical debt</li> <li>Hands-on experience with Linux/Unix systems in production environments</li> <li>Proven track record designing and delivering scalable systems in production environments</li> </ul> <p><em>Preferred Qualifications</em></p> <ul> <li>Experience with infrastructure-as-code tools (Terraform, Ansible)</li> <li>Experience with monitoring and observability tools (Prometheus, Grafana or similar)</li> <li>Experience with container orchestration (Kubernetes, Docker, ECS)</li> <li>Knowledge of low-latency systems, trading environments, market data, or exchange protocols</li> </ul> <p>QRT is an equal opportunity employer. We welcome diversity as essential to our success. QRT empowers employees to work openly and respectfully to achieve collective success. In addition to professional achievement, we

  • Senior Site Reliability Engineer

    Algolia · Paris, France

    On-site
    about 1 month agoApply →

    <div class="content-intro"><p>At Algolia, we’re proud to be a pioneer and market leader in AI Search, empowering 17,000+ businesses to deliver blazing-fast, predictive search and browse experiences at internet scale. Every week, we power over 30 billion search requests — four times more than Microsoft Bing, Yahoo, Baidu, Yandex, and DuckDuckGo combined.</p> <p>In 2021, we raised $150 million in Series D funding, quadrupling our valuation to $2.25 billion. This strong foundation enables us to keep investing in our market-leading platform and serving incredible customers like Under Armour, PetSmart, Stripe, Gymshark, and Walgreens.</p></div><p>Algolia is set to enable every company to create world-class Search and Discovery experiences with an API-first approach. Performance and Scalability is at the heart of our mission: we power 1.5 trillion searches a year, for 10K+ customers all over the world. </p> <p>If you're a problem solver, able to think outside the box and eager to nurture others and learn from them, then this is your challenge!</p> <h3><strong>The Team</strong></h3> <p>The Fleet team is a Site Reliability Engineering team focusing on one goal: the Search products should always be available. To make this possible, the Fleet team creates pragmatic solutions to optimize the Search products availability and costs at scale, taking into account the needs of customers, the product teams, and the many engineering teams involved in delivering a unique Search Experience to our customers.</p> <h3><strong>The Opportunity</strong></h3> <p>The team is looking for an individual who has a first experience of building and operating scalable architectures. You will contribute to the delivery of solutions that support other engineering teams and will have a direct impact on the success of Algolia's Search products. </p> <p>In this role, you'll help design and implement systems focused on reliability, scalability, and cost efficiency, while also having opportunities to grow your skills and collaborate with team members.</p> <p><strong>Your role will include</strong></p> <ul> <li>Operating the Search products, building self-healing and automated incident response mechanisms</li> <li>Building components that improve reliability and performance</li> <li>Monitoring and computing the SLO and the error budget of the product you operate</li> <li>Reducing the toil and the technical debt by automating tasks and increasing the quality of existing components</li> <li>Managing Incidents and Customer Requests</li> </ul> <p><strong>You might be a good fit if you have</strong></p> <ul> <li>5 years experience in a scalable environment</li> <li>Knowle

  • Site Reliability Engineer - Observability

    Proton · Geneva; Paris, France

    On-site
    about 1 month agoApply →

    <h3>Join Proton and build a better internet where privacy is the default</h3> <p>Proton was founded in 2014 by scientists from CERN on a simple truth: <strong>privacy is a fundamental human right</strong>. Since then, we've built the world's largest encrypted email service (Proton Mail) and expanded into Proton VPN, Proton Drive, Proton Pass, and Proton Calendar — tools used by millions globally to protect their freedom, fight censorship, and keep their data safe. In some situations, Proton has literally helped save lives.</p> <p>We are profitable, independent (no VC control), and selectively hire from the top ~1% of applicants. Our 700+ team members across 50+ countries come from leading organizations and elite academic backgrounds. We move fast, keep hierarchy light, and prioritize impact over optics. If you want to do meaningful work with exceptionally high-caliber people, this is it. Check our open-source projects <a href="https://proton.me/community/open-source">here</a>.</p> <h3><strong><strong class="Lexical__textBold">Purpose of the role:</strong></strong>  </h3> <p>We're a small, tool-agnostic team that owns the observability infrastructure behind Proton's services — the logs, metrics, traces, and alerts that keep systems running smoothly for the millions of users who trust us with their privacy. We run on open-source stacks across Proton's on-premise data centers, and we dogfood heavily: we're our own first customers. We favor simple, solid solutions over large engineering efforts, and we believe good systems emerge iteratively. You'll join a group that values frank, open communication and a problem-solving mentality — if you want narrow scope and a fixed backlog, this isn't the right fit.</p> <h3> </h3> <h3>Tech Stack and Tools</h3> <ul> <li>Languages: Python, Go</li> <li>Observability: open-source stacks for logs, metrics, traces, alerting; OpenTelemetry</li> <li>Orchestration: Kubernetes</li> <li>GitOps: ArgoCD</li> <li>Infrastructure-as-code: Terraform, Ansible, Puppet</li> <li>Storage at scale: ClickHouse</li> <li>Platform: Linux, on-premise data centers</li> </ul> <h3>What You'll Do</h3> <ul> <li>Design, deploy, and operate observability pipelines for logs, metrics, traces, and alerts across Proton's services using open-source technologies.</li> <li>Partner with development and platform teams to ship practical alerting, dashboarding, and integration solutions that engineers actually rely on.</li> <li>Build reusable templates and tooling that streamline onboarding, incident response, and analysis.</li> <li>Champion observability best practices across teams and raise the ba

  • Lead Site Reliability Engineer (SRE)

    Amo · Paris, France

    On-site
    about 2 months agoApply →

    Lead Site Reliability Engineer (SRE) at Amo. Apply via Ashby.

  • Site Reliability Engineer - Secteur Public - Paris - CDI

    Theodo · Paris, Paris

    On-site
    about 2 months agoApply →

    L’histoire du groupe Theodo et son succès Theodo accompagne depuis 2009 les entreprises innovantes dans la conception, le développement et le déploiement de produits digitaux ingénieux - tels que la plateforme TF1+, l'application LCL Pro ou l'application France Identité Numérique - en tirant parti du meilleur de la technologie et de l’approche Lean.   Theodo connait une croissance exceptionnelle depuis 15 ans : nos équipes rassemblent plus de 700 Theodoers, principalement des Software Engineers, passionnés de technologie et d’amélioration continue. En 2025, le groupe Theodo génère 100M€ de CA.   Chez Theodo GovTech, nous accompagnons nos clients - ministères, collectivités territoriales et grands opérateurs publics - dans la conception, le développement et la maintenance de produits numériques accessibles, durables et utiles pour la société. Rejoindre notre aventure, c'est partager une ambition commune : utiliser le numérique pour augmenter l’efficacité de l'action publique et améliorer le quotidien des agents et des usagers. Pourquoi on recrute ? Nos équipes interviennent au cœur des organisations publiques, sur des projets complexes, contraints, souvent souverains, où l’infrastructure n’est pas un détail… mais un enjeu central du produit. Aujourd’hui, on renforce nos équipes avec un·e Ops / SRE Build, capable de concevoir et builder l’infrastructure, dans des environnements on-premise ou cloud privé.   L’équipe que tu rejoins : Sur chaque projet, tu seras intégré(e) au sein d’une équipe projet avec 1 Tech lead en charge de ton encadrement, 1 à 2 Dévelopeur, 1 Associate Product Manager et 1 Lead Product Manager. Tu seras managé(e) par Matthieu, notre CTO qui a plus de 14 ans d’expérience dans la tech. Tu travailleras aussi avec Maxime, notre CEO, expert du secteur public.   Tes missions principales sont : - Concevoir des architectures d’infrastructure adaptées à des contextes contraints (sécurité, souveraineté, performance) - Builder et déployer des infrastructures from scratch ou en refonte - Diffuser l’approche DevOps au cœur de l’équipe agile (rituels, cadrage, documentation) - Challenger les choix techniques pour optimiser coûts, performance et maintenabilité - Analyser les défauts et incidents pour renforcer la robustesse du système - Expliquer, documenter et rendre lisibles des décisions techniques complexes - Former l’équipe “en situation de travail” aux bonnes pratiques DevOps   Profil recherché Hard skills - 5 ans minimum d’expérience en SRE sur des environnements on-premise ou cloud privé. - Excellente maîtrise des concepts DevOps et web complexes - Très bonnes bases en conteneurisation et orchestration - Solides compétences en réseaux et systèmes Linux - Capacité à concevoir et expliquer une architecture claire - Exigence forte sur la qualité, la lisibilité et la maintenabilité   Soft skills - Pragmatisme : créer un maximum de valeur avec les contraintes existantes - Esprit d’équipe : la réussite collective avant l’ego - Curiosité et envie de progresser (et de faire progresser les autres) - Humilité : savoir demander de l’aide et challenger avec bienveillance   Process de recrutement L’entretien se déroule en 5 étapes sous une dizaine de jours : Un call de qualification (30 min) Un entretien de culture fit (1 heure) Un entretien théorique de connaissance infra et archi (1 heure) Un entretien technique (1 heure) Un final (1 heure)

  • Site Reliability Engineer - Pôle Run - CDI Paris - Theodo Cloud

    Theodo · Paris, Paris

    On-site
    about 2 months agoApply →

    L’histoire du groupe Theodo et son succès Theodo accompagne depuis 2009 les entreprises innovantes dans la conception, le développement et le déploiement de produits digitaux ingénieux - tels que la plateforme TF1+, l'application LCL Pro ou l'application France Identité Numérique - en tirant parti du meilleur de la technologie et de l’approche Lean.   Theodo connait une croissance exceptionnelle depuis 15 ans : nos équipes rassemblent plus de 700 Theodoers, principalement des Software Engineers, passionnés de technologie et d’amélioration continue. En 2025, le groupe Theodo génère 100M€ de CA.   La practice Run de Theodo a été lancé en 2021 et propose des offres d’infogérance modulables selon trois blocs : React (réaction aux incidents), Maintain (maintien en conditions opérationnelles de l’infrastructure), et Change (évolutions de l’infrastructure).  Ton équipe : L’équipe technique de Theodo Cloud est constituée de deux pôles : - le pôle Build, dédié à nos projets de création, migration et optimisation d’infrastructures - le pôle Run, dédié à l’infogérance des infrastructures de nos clients.   Notre pôle Run connaît une forte croissance et recherche un Site Reliability Engineer pour renforcer l’équipe et ainsi pouvoir développer le nombre de clients infogérés par Theodo.   L’équipe Run est composée de 7 personnes, intervenant sur de la réaction à incident, du maintien en conditions opérationnelles des infrastructures et des optimisations.   Tes missions : En tant que SRE Run, ton rôle est d’assurer la stabilité et la disponibilité des infrastructures de nos clients : - Réagir aux incidents en production et mener des investigations en cas de problèmes techniques - Réaliser les post-mortem et assurer la communication avec le client - Répondre aux interrogations de nos clients et leur fournir des recommandations pour améliorer la qualité de leur infrastructure.   La réalisation d’astreintes est également possible, elle se fait toujours sur la base du volontariat.   Ce poste est idéal pour les personnes qui sont curieuses et souhaitent progresser vite : il offre l’opportunité de travailler sur un large panel d’infrastructures et d’outils. C’est l’assurance de ne jamais s’ennuyer et d’avoir à résoudre des bugs complexes en production.   Les opportunités d’évolution : La première étape d’évolution chez Theodo Cloud est de devenir Lead sur le pôle Run, et de devenir ainsi garant de la montée en compétences de ton équipe. Nous proposons ensuite deux tracks d’évolution possibles : - une track Management, qui t’emmènera vers un rôle d’Engineering Manager - une track Expertise, qui te permettra d’accéder à un rôle de Staff Engineer   Profil recherché : Nous recherchons une personne ayant :   - une solide formation académique en école d’ingénieur - une 1ère expérience professionnelle avec les technologies infra / Cloud, et notamment Kubernetes - de la rigueur et une bonne gestion du stress, afin d’être capable de réagir vite et intervenir sur des environnements en production   Tu ne coches pas 100% des critères mais tu penses tout de même correspondre à notre recherche ? Ne t’auto-censure pas et envoie-nous ta candidature, on se fera un plaisir de l’étudier !   Pour information : - Le poste est proposé en CDI, basé depuis nos locaux à Paris. - Nous permettons à nos équipes de télé-travailler jusqu’à 3 jours par semaine. - Le salaire proposé dépend du niveau d’expérience de chacun : nous en parlons toujours dès le premier appel téléphonique.   Déroulement des entretiens : Nous répondons à toutes les candidatures dans un délai de 3 jours. Si ton profil nous plait, nous te proposons le process suivant :   - un appel téléphonique de 30 minutes avec un membre de l’équipe recrutement, pour comprendre tes critères de recherche - un entretien d’une heure avec le recruteur ou la recruteuse avec qui tu as échangé à la première étape, pour parler plus en détail de la culture Theodo - un entretien de logique - un entretien technique de débug sur Docker de 45 minutes - un échange de 30 minutes avec un membre de l'équipe dirigeante   Tu seras accompagné(e) par une personne de l’équipe recrutement, qui sera là pour te coacher tout au long du process.

  • Senior Site Reliability Engineer (SRE)

    Swile · Paris, France

    On-site
    about 2 months agoApply →

    At Swile, we believe that good products can help reduce friction in daily professional life and boost employee satisfaction. Today, we provide innovative solutions in various areas such as Fintech, Travel, HR, and Employee Benefits to more than 5.5 million users in 85,000 companies in France and Brazil.   Your role as a Senior Site Reliability Engineer (SRE) centers around creatively solving problems, ensuring a balance between speed, reliability, pragmatism and excellence. It's less about the specific technologies used and more about crafting innovative solutions to business and engineering challenges.  

  • Site Reliability Engineer - SRE

    Scaleway · Paris, Paris

    On-site
    about 2 months agoApply →

    OUR STORY:   🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow ! Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.   Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.   With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.   Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.   📍 Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.   WHY WE NEED YOU?   Our growth is driving us to strengthen our SRE team to support and scale our production environments. Your mission will be to build and maintain reliable, observable, and secure infrastructure in order to ensure optimal service availability for our customers around the world.   YOUR FUTURE TEAM   We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together.   You will be part of a team of experienced Site Reliability Engineers. The team is responsible for maintaining and evolving core infrastructure and observability tools, supporting product teams, and improving the reliability of Scaleway’s services.   YOUR DAILY ROUTINE   - Build and optimize tooling to automate monitoring, diagnosis, and remediation of production incidents - Troubleshoot high-impact production issues in collaboration with other engineering teams - Participate in an on-call rotation to handle incidents and ensure service continuity - Implement and maintain observability solutions to monitor infrastructure and application health - Contribute to infrastructure lifecycle management across different environments - Promote and apply best practices in terms of stability, resiliency, scalability, and security - Maintain clear technical documentation for tools and procedures - Contribute to system and tool evolution based on production feedback - Collaborate closely with development teams to ensure infrastructure readiness - Participate in team rituals and knowledge-sharing initiatives   ABOUT YOU   SOFTSKILLS : - Proactive and solution-oriented mindset - Passion for automation and continuous improvement - Strong collaboration and communication skills - Ability to work independently and in a team - Willingness to mentor and share knowledge   💻 HARDSKILLS :    - Experience with Go, Python or Rust - Strong scripting skills (Bash, Python) - Hands-on experience with Linux systems (Ubuntu/Debian) - Knowledge of networking (TCP/IP, DNS, BGP, load-balancing, IPv6, etc.) - Experience in cloud environments and infrastructure (bare metal, VMs, containers, orchestrators) - Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.) - Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.) - Experience managing relational databases (PostgreSQL) - Understanding of CI/CD pipelines (GitLab) - Comfortable with English (written and spoken)   WHAT YOU WILL FIND AT SCALEWAY ++++   - Hybrid work: We offer up to 3 days of remote work per week. - Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities. - Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches. - Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life. International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.  - Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.   🚀 Why join the Scaleway adventure ?  ✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI. ✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges. ✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.   🔜 THE NEXT STEPS … - Discovery call with a recruiter (30 min) - Interview with the manager to understand your technical skills and approach to the role (45 min) - Technical interview to validate your expertise (1h) - Interview with the Head of the Tribe to deepen your discussions and assess your fit with the team (45 min) - HR interview to tour our offices and meet your future colleagues          

  • Site Reliability Engineer (SRE) - AI GPU Clusters

    Scaleway · Paris, Paris

    On-site
    about 2 months agoApply →

    OUR STORY:   🇪🇺 Join Scaleway and shape the sovereign cloud of tomorrow ! Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.   Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.   With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.   Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.   📍 Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.   WHY WE NEED YOU? Our growth is driving us to strengthen our SRE team to support and scale our production environments. Your mission will be to build and maintain reliable, observable, and secure infrastructure in order to ensure optimal service availability for our customers around the world.   #HPC #AI #GPU #CLUSTERS   YOUR FUTURE TEAM We work in a collaborative and international environment where the diversity of Scalers, combined with a spirit of sharing, helps bring new projects to life every day, advancing our ambitions together. You will join a newly formed team dedicated to building and operating Scaleway’s future AI infrastructure. As part of this group, you will design, maintain, and scale core systems and observability tools, partner with product teams, and ensure the reliability and performance of AI services across Scaleway.   YOUR DAILY ROUTINE - Build a large AI infrastructure with monitoring, diagnosis, and remediation of production incidents- Troubleshoot high-impact production issues in collaboration with other engineering teams - Participate in an on-call rotation to handle incidents and ensure service continuity - Implement and maintain observability solutions to monitor AI infrastructure and application health - Contribute to AI infrastructure lifecycle management across different environments and countries - Promote and apply best practices in terms of stability, resiliency, scalability, and security - Maintain clear technical documentation for tools and procedures - Contribute to system and tool evolution based on production feedback - Collaborate closely with development teams to ensure infrastructure readiness- Participate in team rituals and knowledge-sharing initiatives   ABOUT YOU   🎯 SOFTSKILLS :  - Proactive and solution-oriented mindset - Passion for automation and continuous improvement - Strong collaboration and communication skills - Ability to work independently and in a team - Willingness to mentor and share knowledge   💻 HARDSKILLS :  - Experience with Python, Go or C++  - Strong scripting skills (Bash, Python) - Hands-on experience with Linux systems (Ubuntu/Debian) - Preferred hands-on experience with GPU & HPC infrastructure  - Knowledge of networking (TCP/IP, DNS, BGP, load-balancing, IPv6, etc.) - Familiarity with monitoring and logging tools (Prometheus, Grafana, Elastic, etc.) - Comfortable with Infrastructure-as-Code (Ansible, Salt, AWX, etc.) - Experience managing relational databases (MariaDB) - Understanding of CI/CD pipelines (GitLab) - Comfortable with English (written and spoken)   WHAT YOU WILL FIND AT SCALEWAY ++++ Hybrid work: We offer up to 3 days of remote work per week. Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities. Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches. Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life. International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French. Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.     🚀 Why join the Scaleway adventure ?  ✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI. ✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges. ✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.   🔜 THE NEXT STEPS … - Discovery call with a recruiter (30 min) - Technical interview to validate your expertise (1h) - Interview with the manager to understand your approach to the role (45 min) - Interview with the Head of the Tribe to deepen your discussions and assess your fit with the team (45 min) - HR interview to tour our offices and meet your future colleagues

  • Site Reliability Engineer Cloud&Infra (H/F)

    Safran ai · Paris, Paris

    On-site
    about 2 months agoApply →

    Safran.AI - De l'IA qui opère dans le réel Filiale de Safran Electronics & Defense, Safran.AI conçoit des solutions d'IA appliquées à des données complexes telles que les images satellite haute résolution, les flux vidéo FMV et les signaux acoustiques. Ces solutions s'appuient sur des algorithmes capables de détecter et d'identifier automatiquement des objets d'intérêt dans les secteurs du renseignement, de la défense et de l'aéronautique.  Depuis son intégration à Safran en septembre 2024, Safran.AI contribue également à la transformation du groupe, en appliquant les solutions d’IA aux domaines de l’industrie 4.0.  À titre d’exemple, l’analyse d’images automatisée par l’IA peut assister les contrôleurs en charge de l’inspection de pièces critiques en les aidant à détecter les anomalies éventuelles à partir de clichés numériques.    Fort de nos 250 collaborateurs, vous évoluerez au sein d’équipes passionnées et pluridisciplinaires, réunissant des talents parmi les plus reconnus du secteur, tous animés par une même exigence d’excellence et d’innovation technologique.    Les enjeux du poste En tant qu'ingénieur Cloud & Infrastructure vous serez amené à :   → Construire et améliorer le socle technique pour la gestion des cloud et de l’orchestrateur Kubernetes → Gérer et augmenter la couverture de supervision → Gérer les outils de forge logiciel → Fournir un support à l'utilisation du socle par les SRE présent dans les équipes de développement (embedded SRE) → Fournir des modules industrialisés et sécurisés pour instancier de l'infrastructure cloud dans un premier temps et de l'infrastructure “on-premise” et “air gap” dans un second temps → Assister l'équipe sécurité dans la mise en place des best practice sécurité → Réaliser la MCO du socle et de ces composants Coté stack : L'équipe Cloud & Infrastructure côtoie les technologies suivantes : Cloud : AWS, OVH, ... Forge :Github, Sonar, Artifactory, … Monitoring : Grafana, Mimir, Loki, OpenTelemetry, … IAC : Terraform, Ansible, ArgoCD, ... Réseau : Cloudflare, WARP, répartition de charge (LB), routage et filtrage ... Sécurité : Hashicorp Vault, Hardening (ANSSI/CIS/STIG), Prisma, IAM, ... Attention : la capacité à obtenir une habilitation Défense est obligatoire pour ce poste.  A propos de vous → Vous avez entre 1 et 2 ans d'expérience sur un poste similaire ou aux fortes composantes DevOps. → Vous avez de bonnes connaissances en réseaux et en sécurité → Vous avez une bonne capacité à expliquer des sujets techniques à différents publics. → Vous avez un bon sens de l'organisation, autonomie et capacité à établir des rapports → Expérience de travail sur des infrastructures cloud (AWS) et on-premise. → Expérience sur de l'installation on-premise “air gap” serait un plus. → Expérience de travail sur Docker avec un système d'orchestration. → Expérience dans la mise en œuvre de pipelines CI / CD et des défis qui y sont liés. → Vous avez un large bagage technique, de l'infrastructure au développement en passant par le réseau et la sécurité. → Vous êtes un ambassadeur de l'automatisation (terraform, ansible) et de la qualité. Si vous ne remplissez pas 100% des critères ci-dessus, pas de panique, vous pouvez nous indiquer les raisons pour lesquelles vous pensez tout de même être un bon candidat pour ce rôle ! Ce que vous trouverez chez Safran.AI   Safran.AI réuni des sujets techniques de fond, des équipes qui maîtrisent leur domaine et une culture où la rigueur et l'entraide coexistent.   Structure à taille humaine adossée à un groupe industriel de premier plan, elle permet de travailler sur des enjeux réels sans renoncer à la proximité.  Nos bénéfices d’entreprise   📈 Participation & intéressement  💼 Plans d’épargne groupe avec abondement jusqu’à 3 000 € / an  ⏳ Compte Épargne Temps après 1 an d’ancienneté  🏥 Mutuelle familiale prise en charge à 55 %  🍽️ Carte Swile de 11 € / jour travaillé pris en charge par SAFRAN AI à hauteur de 59,1%  🚆 Jusqu’à 75 % des frais de transport pris en charge ou forfait mobilité durable  🏡 Jusqu’à 3 jours de télétravail par semaine selon les postes   👶 10 semaines de congé second parent & maintien de salaire lié à la grossesse  📚 Offre de formation et développement des compétences  🧘 Avantages bien-être : Moka Care, conférences & accès sport via Wellpass  Votre parcours de recrutement Un échange de 45 minutes avec un recruteur pour en apprendre plus sur vous, vos attentes et vous donner plus de détails sur la vie chez Safran.AI  

  • Applied AI Engineer, Site Reliability Engineer - EMEA

    Mistral · Paris, Paris

    On-site
    about 2 months agoApply →

    About Mistral  At Mistral AI, we believe in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to integrate seamlessly into daily working life.   We democratize AI through high-performance, optimized, open-source and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet enterprise needs, whether on-premises or in cloud environments. Our offerings include le Chat, the AI assistant for life and work.   We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited. Join us to be part of a pioneering company shaping the future of AI. Together, we can make a meaningful impact. See more about our culture on https://mistral.ai/careers. About the team The Applied AI team is Mistral's customer-facing technical organization. We work directly with enterprise clients from pre-sales through implementation to deploy cutting-edge AI solutions that deliver measurable business impact. Our team combines deep ML expertise with strong customer engagement skills, operating like startup CTOs who own end-to-end project execution. Our SRE team works transversally across customer engagements, enabling value creation through Mistral tech - at scale. By joining the team you will bridge the gap between cutting-edge AI research and real-world enterprise applications, ensuring our solutions are robust, scalable, and aligned with both customer needs and Mistral's technological vision. About The Job You will be one of the founding engineers of the Applied AI SRE sub-team. Your mission, alongside the team, is to build and operate the framework to ensure Mistral’s solution delivery is reliable and sustainable - and applied uniformly across all our accounts, both Mistral-hosted and customer-hosted. You should already have a strong understanding of what operational excellence looks like, and you’re ready to scale your impact. You will operate in four concurrent modes: - BUILD - Design for a fleet of Mistral platforms and apps. Build proactivity to reduce reactivity. Productize reliability, author runbooks, create SLO templates, implement observability. - RUN - Operate the Tier-1 customer environments that Mistral are contracted to operate. Ensure SLO compliance, own on-call and incident response, manage drift, partner with Technical Support as L3 escalation, champion high signal post-mortems. - ENABLE - Productize how Mistral deploy, secure, and scale our Applied AI solutions. Engineer on-demand provisioning, author security baseline packages, embed security guardrails, automate everything. - SECURE - Own the security operations layer for our customer-side deployments. Lead CVE response across the fleet, ship supply-chain integrity controls (SBOM, signed images, provenance), co-page with InfoSec on security incidents, enforce secure-config baselines. This is a framework-first, fleet management role at heart. If you're excited by the difference between solving one customer's problem and structurally solving the class of problem for every customer, this is the role. How We Work in Applied AI • We care about people and outputs. • What matters is what you ship, not the time you spend on it • Bureaucracy is where urgency goes to vanish. You talk to whoever you need to talk to. The best idea wins, whether it comes from a principal engineer or someone in their first week. • Always ask why. The best solutions come from deep understanding, not from copying what worked before • We say what we mean. Feedback is direct, timely, and given because we care. • No politics. Low ego, high standards. • We embrace an unstructured environment and find joy in it. About you • Fluent in English. • 5+ years in SRE, Production Engineering, or DevOps, with a record of shipping tooling. • Strong multi-tenant Kubernetes fluency, namespace segmentation, network policy, RBAC, admission control, operations at scale. • On-call discipline: incident response, blameless post-mortem culture, runbook-first mindset. • Observability stack in production: Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Signoz. • Infrastructure as code: Terraform, Ansible (or close equivalents). • Proficient in Python and/or Golang for tooling and automation. • Security mindset: you treat secure-SDLC, CVE response, and supply-chain integrity as reliability properties of the shipped artifact, not as someone else's job. • Strong written communication skills: runbooks, post-mortems, and customer-facing incident comms are core deliverables of this role. • Comfortable operating with high autonomy in an ambiguous, fast-paced environment — and disciplined enough to defend the team's scope when work tries to spill in. • Solid Linux internals, networking debug, and distributed-systems fundamentals. Strong plus • Cloud or application security background (AppSec, K8s security, supply chain — SBOM, cosign, SLSA). At least one of our early hires must bring this; if it's you, flag it. • Experience operating LLM / model-serving stacks in production • Experience with multi-cloud or on-prem hybrid customer environments (AWS, GCP, Azure, sovereign clouds). • Open-source contributions, particularly in SRE, observability, or security tooling.     By applying, you agree to our Applicant Privacy Policy.

  • Site Reliability Engineer

    Mistral · Paris, Paris

    On-site
    about 2 months agoApply →

    About Mistral    At Mistral AI, we believe in the power of AI to simplify tasks, save time, and enhance learning and creativity. Our technology is designed to integrate seamlessly into daily working life.   We democratize AI through high-performance, optimized, open-source and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet enterprise needs, whether on-premises or in cloud environments. Our offerings include le Chat, the AI assistant for life and work.   We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited.   Join us to be part of a pioneering company shaping the future of AI. Together, we can make a meaningful impact. See more about our culture on https://mistral.ai/careers.   Role Summary    We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and exceed our internal and external customers' expectations.   Location: Remote - Europe Reporting line: Team Lead, Site Reliability Engineer   What you will do   As a Site Reliability Engineer, you balance the day-to-day operations on production systems with long-term software engineering improvements to reduce operational toil and foster the reliability, availability, and performance of these systems.   Operations • Design, build, and maintain scalable, highly available and fault-tolerant infrastructures to support our web services and ML workloads • Make sure our platform, inference and model training environments are always highly available and enable seamless replication of work environments across several HPC clusters • Operate systems and troubleshoot issues in production environments (interrupts, on-call responses, users admin, data extraction, infrastructure scaling, etc.) • Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime • Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our client-facing APIs and large training runs • Participate occasionally in on-call rotations to respond to incidents and perform root cause analysis to prevent future occurrences   Development • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform • Collaborate with AI/ML researchers to develop and implement solutions that enable safe and reproducible model-training experiments • Build a cloud-agnostic platform offering an abstraction layer between science and infrastructure • Design and develop new workflows and tooling to improve to the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.) • Collaborate with the security team to ensure infrastructure adheres to best security practices and compliance requirements • Document processes and procedures to ensure consistency and knowledge sharing across the team • Contribute to open-source projects, research publications, blog articles and conferences   About you   • Master’s degree in Computer Science, Engineering or a related field • 7+ years of experience in a DevOps/SRE role • Strong experience with cloud computing and highly available distributed systems • Exposure to site reliability issues in critical environments (issue root cause analysis, in-production troubleshooting, on-call rotations...) • Experience working against reliability KPIs (observability, alerting, SLAs) • Hands-on experience with CI/CD, containerization and orchestration tools (Docker, Kubernetes...) • Knowledge of monitoring, logging, alerting and observability tools (Prometheus, Grafana, ELK Stack, Datadog...) • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation • Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices • Strong understanding of networking, security, and system administration concepts • Excellent problem-solving and communication skills • Self-motivated and able to work well in a fast-paced startup environment   Your application will be all the more interesting if you also have: • experience in an AI/ML environment • experience of high-performance computing (HPC) systems and workload managers (Slurm) • worked with modern AI-oriented solutions (Fluidstack, Coreweave, Vast...)   By applying, you agree to our Applicant Privacy Policy.  

  • Mistral Cloud - Site Reliability Engineer

    Mistral · Paris, Paris

    On-site
    about 2 months agoApply →

    About Mistral  At Mistral AI, we believe in the power of AI to simplify tasks, save time and enhance learning and creativity. Our technology is designed to integrate seamlessly into daily working life.  We democratize AI through high-performance, optimized, open-source and cutting-edge models, products and solutions. Our comprehensive AI platform is designed to meet enterprise needs, whether on-premises or in cloud environments. Our offerings include le Chat, the AI assistant for life and work. We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between France, USA, UK, Germany and Singapore. We are creative, low-ego and team-spirited. Join us to be part of a pioneering company shaping the future of AI. Together, we can make a meaningful impact. See more about our culture on https://mistral.ai/careers. Role summary We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our Cloud platform and customer facing applications. You will work closely with our software engineers and product teams to ensure our systems meet and exceed our internal and external customers' expectations.  More information on Mistral Cloud here: https://mistral.ai/products/compute What you will do Operations • Design, build, and maintain scalable, highly available and fault-tolerant infrastructures • Operate systems and troubleshoot issues in production environments (interrupts, on-call responses, users admin, data extraction, infrastructure scaling, etc.) • Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime • Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging and alerting systems) for both our customer-facing APIs and large training runs • Participate occasionally in on-call rotations to respond to incidents and perform root cause analysis to prevent future occurrences   Development • Drive continuous improvement in infrastructure automation, deployment, and orchestration • Collaborate with software engineers to develop and implement solutions that enable safe and reproducible model-training experiments • Help build a cloud platform offering an abstraction layer between science, engineering and infrastructure • Design and develop new workflows and tooling to improve the reliability, availability and performance of our systems (automation scripts, refactoring, new API-based features, web apps, dashboards, etc.) • Collaborate with the security team to ensure infrastructure adheres to best security practices and compliance requirements • Document processes and procedures to ensure consistency and knowledge sharing across the team • Contribute to open-source projects, research publications, blog articles and conferences About you • Master’s degree in Computer Science, Engineering or a related field • 5+ years of experience in a DevOps/SRE role • Strong experience with bare metal infrastructure and highly available distributed systems • Exposure to site reliability issues in critical environments (issue root cause analysis, in-production troubleshooting, on-call rotations...) • Experience working against reliability KPIs (observability, alerting, SLAs) • Hands-on experience with CI/CD, containerization and orchestration tools (Docker, Kubernetes...) • Knowledge of monitoring, logging, alerting and observability tools (Prometheus, Grafana, ELK Stack, Datadog...) • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation • Proficiency in scripting languages (Python, Go, Bash...) and knowledge of software development best practices • Strong understanding of networking, security, and system administration concepts • Excellent problem-solving and communication skills • Self-motivated and able to work well in a fast-paced startup environment   Your application will be all the more interesting if you also have: • experience in an AI/ML environment • experience of high-performance computing (HPC) systems and workload managers (Slurm) • worked with modern AI-oriented solutions (Fluidstack, Coreweave, Vast...) Hiring Process • Introduction Call - 30 min • Manager Interview - 30 min • Technical Interview / System Design - 45 min • Technical Interview / Deep Dive - 60 min • Culture-fit Discussion - 30 min • References Our Culture We're driven to build a strong company culture and are looking for individuals with solid alignment with the following: • Reason with rigor • Are you audacious enough? • Make our customers succeed • Ship early and accelerate • Leave your ego aside  Engineering blog Our first Engineering blog post is live, you can check it out here !

  • Site Reliability Engineer

    Blablacar · Paris, France

    On-site
    about 2 months agoApply →

    About BlaBlaCar BlaBlaCar is the world’s leading community-based travel app enabling 27 million members a year to carpool or travel by bus in 21 countries. Our team of 800 employees counts over 50 nationalities and is spread across our 5 global offices, 30% working fully remotely. Your Mission By joining our Foundations department, you will be working alongside talented individuals grouped in small agile teams that each have strong ownership on their piece of these goals. Foundations is composed of seven teams which “provide consistent, easy to use, infrastructures, services, and expertise to support BlaBlaCar’s growth and evolution”.  The Site Reliability Engineering team (SRE) is responsible to provide best in class Observability, Alerting and Incident management tools and processes to service teams. As an enabling team, we help BlaBlacar engineers to efficiently improve their service reliability. Empowering developers and bringing them our reliability expertise are at the core of our daily work.  Technical stack: Core Infrastructure: Kubernetes, Google Cloud Platform GitOps/Delivery: GitHub, Terraform, Flux, Helm, Jenkins Observability/Incident Management: Datadog, Opentelemetry, Grafana IRM, In house Synthetic Tests platform: Playwright, Qualcium, SauceLabs Languages: Go / Python for Tooling, Typescripts/JS for the testing platform Your responsibilities  Support software engineers by creating, maintaining, and improving observability and alerting tools and frameworks. You embrace the use of AI, leveraging agentic to eliminate toil and streamline your daily tasks Own the Service Level Objectives (SLOs) framework, assist in the design and maintenance of indicators (SLI) and objectives to ensure service reliability. Owning the incident management process by defining best practices, standards, and ensuring continuous improvement through post-mortems and chaos engineering. While developers handle incidents within their scope, you could step in as Incident Commander during high-severity incidents, leading coordination efforts . Develop and maintain tools, such as Terraform modules or Go apps, to help automate and enhance reliability across services. Build and promote reporting on operational metrics and incidents to drive distributed and continuous improvement. Your qualifications  1 to 5 years of experience in SRE, DevOps, or Software Engineering roles Working in a multidisciplinary environment will request strong communication skills : you'll need to adapt your communication level to other teams expertise and be able to understand their needs Strong knowledge of observability tools (e.g., Datadog) and understanding of metrics, logging, and tracing. Troubleshooting/oncall experience in production environments, diagnosing and resolving technical issues effectively (experience with Kubernetes is a plus). Full working proficiency in English  Fit with our BlaBlaPrinciples Thriving in a collaborative, fast-growing and innovative environment Ability to take ownership, aligned with business priorities and navigating in different contexts   Nice to have: Familiarity with incident management platforms (e.g., Grafana IRM) is a bonus Experience working with Service Level Objectives (SLOs) and Service Level Indicators (SLIs) Exposure to programming in Go or a strong interest in learning it. Experience in integrating Opentelemetry  Backend services are built using multiple programming languages: while development skills aren't required, familiarity with object-oriented programming and scripting languages is an advantage. Familiarity with web/mobile testing tools or a strong curiosity to understand how software is tested at scale.   What we have to offer Hybrid status for this role : 2-3 days at the Office 4 additional weeks on top of legal maternity/paternity leaves 50% healthcare coverage (Alan) Financial support for home office equipment Minimum 25 days holiday per year  Local meal plan policy (Swile card) 50% transportation paid (Forfait Mobilité Durable) Free unlimited carpooling & bus rides Personal growth via trainings, mentorship, and internal mobility opportunities Employee Stock ownership plan Regular team building events 1 day off per year to test our product   Interested in joining the ride? a 45-min video-call with Maxime, Talent Acquisition Manager,  to get to know you, understand your career expectations and answer your questions a 60-min video-call with Damien Bertau, Hiring Manager, to discuss your experience and share more details about the team a 90-min system design interview with 2 team members to discuss about your technical expertise a 45-min video-call with Maxime Fouilleul, Head of Foundations, to get a wider vision of the department and its strategy   Our hiring process lasts on average 25-30 days, offers usually come within 48 hours. Please note that one of these interviews will be onsite. 

  • Site Reliability Engineer (EngX)

    Blablacar · Paris, France

    On-site
    about 2 months agoApply →

    About BlaBlaCar BlaBlaCar is the world’s leading community-based travel app enabling 27 million members a year to carpool or travel by bus in 21 countries. Our team of 800 employees counts over 50 nationalities and is spread across our 5 global offices, 30% working fully remotely. Your Mission SRE in the Engineering Experience Team, part of the Foundations Department, are responsible for designing, building and maintaining the Software Delivery platform, tools and standards that enable teams to confidently release changes up to production. We aim to accelerate delivery, simplify the Engineering experience, guarantee reliable workflow and satisfy Engineering needs at scale. By joining our Foundations Department, you will be working alongside talented individuals grouped in small agile teams that each have strong ownership on their stack and roadmaps. Foundations is composed of four teams (Engineering Experience, Cloud Infrastructure, Site Reliability Engineering & Quality Assurance) which “provide consistent, easy to use, secured infrastructure, services, and expertise to support BlaBlaCar’s growth and evolution”. The Engineering Experience Team has four main objectives, driving its roadmap: Reduce BlaBlaCar product’s time to market by designing, building and maintaining state of the art CI/CD and associated tooling to streamline day-to-day delivery from development teams Improve developers efficiency in providing AI tooling and infrastructure, ensuring the compliance with internal policies while keeping enough flexibility for experimentation Drive development teams towards autonomy through the provision of comprehensive training and support, clear guidelines, and effective tooling Leverage our existing FinOps framework to enable precise cost control for the Software-as-a-Service (SaaS) we use and manage, strategically balancing this with the need to support innovation and the adoption of new functionalities The role requires a global vision of the Engineering perimeter.You will champion the adoption and sharing of best practices among Engineering. Your approach should be that of an enabler, not a gatekeeper. You embrace the use of AI, leveraging code generators and assistants to eliminate toil and streamline your daily tasks and make development teams life easier. Crucially, your strong communication skills will be essential for ensuring a clear understanding of our users' needs. To fulfill the mission, you will be working with several stakeholders :  The Product & Engineering teams, working with service team to ensure best understanding and usage of our Software Factory components. Developer Experience Engineers, to ensure that the best-in-class user experience is prioritized from the start. External SaaS providers, to deliver cutting-edge support for our internal users, as well as analyzing and recommending subscription adjustments to maximize the value BlaBlaCar derives from these services. Technical stack: Core Infrastructure: Google Cloud Platform, Kubernetes GitOps/Delivery: GitHub, Github Copilot, Github Actions & Jenkins, Terraform, Flux, Helm Datastores: Postgres, Cassandra, Elasticsearch, Kafka Observability: Datadog, Grafana Languages: Go for Infra/Security Tooling, Java for backend services, Python for data Your ResponsibilitiesIn cooperation with your Engineering Manager and the Engineering Experience team: Design, build and improve parts of our Software Factory, specifically but not only Continuous Integration and Continuous Delivery, to address scaling and resiliency needs on our cloud platforms; Implement tools and services to ease the work of developers and automate problem resolution; Collaborate with engineers and help them improve software development lifecycle and processes. Investigate and fix service issues; Your Qualifications You can demonstrate a strong experience with large scale continuous integration/delivery systems (e.g. GitHub Actions  or Jenkins); You can demonstrate an experience with Cloud platforms, container and process isolation technologies, especially Docker and Kubernetes; You can demonstrate an experience with an SRE/DevOps oriented language (e.g. Go); You can demonstrate a good knowledge of Linux/Unix fundamentals; You embrace change, prioritize high-value tasks, and are results-driven and impact-oriented; You are a humble, collaborative, and communicative team player, focused on enabling developer empowerment and autonomy, eager to share knowledge and learn from others; You are at ease with English speaking. If you don’t meet 100% of the qualifications outlined above, tell us why you’d still be a great fit for this role in your application! What we have to offer Hybrid status for this role : 2-3 days at the Office 4 additional weeks on top of legal maternity/paternity leaves 50% healthcare coverage (Alan) Financial support for home office equipment Minimum 25 days holiday per year  Local meal plan policy (Swile card) 50% transportation paid (Forfait Mobilité Durable) Free unlimited carpooling & bus rides Personal growth via trainings, mentorship, and internal mobility programs Employee Stock ownership plan Regular team building events 1 day off per year to test our product Interested in joining the ride? Here’s what your hiring journey will look like: a 45-min video-call with Maxime, Talent Acquisition Managers to get to know you, understand your career expectations, and answer your first questions a 60-min video-call with your future manager, Jean-Baptiste Favre, Engineering Manager, to get to know you, present you the team, and discuss your technical fit for the role. a technical assignment to evaluate your technical skills followed by a 60-min video-call with two Engineers. a 30-min video-call with Maxime Fouilleul, Head of Engineering, for vision fit and rounding off the process 

  • Site Reliability Engineer - Infrastructure Systems

    Proton · Geneva; Paris, France

    On-site
    about 2 months agoApply →

    <div id="content" class="highlighter-context page view" data-inline-comments-target="true" data-testid="page-content-only"> <div class="_19pkidpf _2hwx1wug _otyridpf _18u01wug _1bsb1osq"> <div> <div id="main-content" class="wiki-content css-k6juse e5xcnr80" data-testid="pageContentRendererTestId" data-vc="pageContentRendererTestId" data-test-appearance="full-page"> <div class="renderer-overrides"> <div class="cc-11exwsr"> <div class="ak-renderer-wrapper is-full-page cc-1jke4yk"> <div class="cc-7j7jco"> <div class="ak-renderer-document"> <p data-renderer-start-pos="18" data-local-id="384391c2485b"><strong data-renderer-mark="true">Join Proton and build a better internet where privacy is the default</strong></p> <p data-renderer-start-pos="88" data-local-id="70873c53c148">Proton was founded in 2014 by scientists from CERN on a simple truth: <strong data-renderer-mark="true">privacy is a fundamental human right</strong>. Since then, we’ve built the world’s largest encrypted email service (Proton Mail) and expanded into Proton VPN, Proton Drive, Proton Pass, and Proton Calendar—tools used by millions globally to protect their freedom, fight censorship, and keep their data safe. In some situations, Proton has literally helped save lives!</p> <p data-renderer-start-pos="518" data-local-id="bb6a25ba35fe">We are profitable, independent (no VC control), and selectively hire from the top ~1% of applicants. Our 500+ team members across 50+ countries come from leading organizations and elite academic backgrounds. We move fast, keep hierarchy light, and prioritize impact over optics. If you want to do meaningful work with exceptionally high-caliber people, this is it. Join us and do work you can truly be proud of. Check our open-source projects <a class="_ymio1r31 _ypr0glyw _zcxs1o36 _mizu1v1w _1ah3dkaa _ra3xnqa1 _128mdkaa _1cvmnqa1 _4davt94y _4bfu1r31 _1hms8stv _ajmmnqa1 _vchhusvi _kqswh2mm _ect4ttxp _syaz13af _1a3b1r31 _4fpr8stv _5goinqa1 _f8pj13af _9oik1r31 _1bnxglyw _jf4cnqa1 _30l313af _1nrm1r31 _c2waglyw _1iohnqa1 _9h8h12zz _10531ra0 _1ien1ra0 _n0fx1ra0 _1vhv17z1" href="https://proton.me/community/open-source" data-renderer-mark="true">here</a>!</p> <p class="Lexical__paragraph"><strong><strong class="Lexical__textBold">Purpose of the role:</strong></strong>  </p> <p class="Lexical__paragraph">You will be part of the Infrastructure Systems team which provides all base platforms in Proton (Kubernetes, VM orchestration, Bare metal provisioning) as well as all the critical services for them (DNS, DH

  • Site Reliability Engineer - Application Edge

    Proton · Geneva; Paris, France

    On-site
    about 2 months agoApply →

    <div id="content" class="highlighter-context page view" data-inline-comments-target="true" data-testid="page-content-only"> <div class="_19pkidpf _2hwx1wug _otyridpf _18u01wug _1bsb1osq"> <div> <div id="main-content" class="wiki-content css-k6juse e5xcnr80" data-testid="pageContentRendererTestId" data-vc="pageContentRendererTestId" data-test-appearance="full-page"> <div class="renderer-overrides"> <div class="cc-11exwsr"> <div class="ak-renderer-wrapper is-full-page cc-1jke4yk"> <div class="cc-7j7jco"> <div class="ak-renderer-document"> <p data-renderer-start-pos="53"><strong data-renderer-mark="true">Join Proton and build a better internet where privacy is the default</strong></p> <p data-renderer-start-pos="123">At Proton, we believe that privacy is a fundamental human right and the cornerstone of democracy. Since our inception in 2014, founded by a team of scientists from CERN, we have dedicated ourselves to providing free and open-source technology to millions worldwide, ensuring access to privacy, security, and freedom online.</p> <p data-renderer-start-pos="448">Our journey began with Proton Mail, the largest secure email service globally, and has since expanded to include Proton VPN, Proton Calendar, Proton Drive, and Proton Pass. These tools empower individuals and organizations to take control of their personal data, break away from Big Tech’s invasive practices, and defeat censorship. Our work impacts hundreds of millions of lives, from activists on the front lines defending freedom to leaders in governments protecting sensitive information. In some cases, Proton’s services have even been instrumental in saving lives by enabling secure and private communications in high-risk situations.</p> <p data-renderer-start-pos="1090">Proton is a profitable company that does not rely upon VC funding, supporting over 100 million user accounts with a growing team of over 500 people from over 50 different countries, from the world's top companies and universities. We value intelligence, learning potential, and ambition in our hiring process. Adaptability is key as we navigate uncharted territories and redefine how business is conducted online.</p> <p data-renderer-start-pos="1505">Hiring at Proton is highly selective, with less than 1% of candidates hired. We believe smaller teams of exceptional talent will always prevail over larger teams with lower talent density. You will have the opportunity work with many of the world's top minds in their fields, ranging from former international math and science olympiad winners to chess champions.</p> <p data-renderer-start-pos="1870">We have a global minds