Jobs and Careers
IT

Senior Site Reliability Engineer

Iterable
PortugalRemotefull_timeVerifiedPosted 16 Dec 2025

About the role

<p>Iterable is the leading AI-powered customer engagement platform that helps leading brands like Redfin, SeatGeek, Priceline, Calm, and Box create dynamic, individualized experiences at scale. Our platform empowers organizations to activate customer data, design seamless cross-channel interactions, and optimize engagement—all with enterprise-grade security and compliance. Today, nearly 1,200 brands across 50+ countries rely on Iterable to drive growth, deepen customer relationships, and deliver joyful customer experiences.</p> <p>Our success is powered by extraordinary people who bring our core values—Trust, Growth Mindset, Balance, and Humility—to life. We foster a culture of innovation, collaboration, and inclusion, where ideas are valued and individuals are empowered to do their best work. That’s why we’ve been recognized as one of <a href="https://www.inc.com/best-workplaces/2022">Inc’s Best Workplaces</a> and <a href="https://iterable.com/blog/inc-names-iterable-one-of-americas-fastest-growing-companies/">Fastest Growing Companies</a>, and were recognized on Forbes’ list of America’s Best Startup Employers in 2022. Notably, Iterable has also been listed on<a href="https://blog.wealthfront.com/announcing-2021-career-launching-companies/"> Wealthfront’s Career Launching Companies List</a> and has held a top 10 ranking on the<a href="https://wearegirlsclub.com/top-25-companies-where-women-want-to-work/"> Top 25 Companies Where Women Want to Work</a>.</p> <p>With a global presence—including offices in San Francisco, New York, Denver, London, and Lisbon, plus remote employees worldwide—we are committed to building a diverse and inclusive workplace. We welcome candidates from all backgrounds and encourage you to apply. Learn more about our story and mission on our<a href="https://iterable.com/culture/"> Culture</a> and<a href="https://iterable.com/company/"> About Us</a> pages. Let’s shape the future of customer engagement together!</p><p><strong>How you will make an impact:</strong></p> <p>As a Senior Engineer on the Cloud Platform team, your impact will be measured by the continuous improvement of our platform’s reliability, scalability, and security posture.</p> <ul> <li>SLO Ownership &amp; Error Budget Management: Take direct ownership of the established Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for core platform services (e.g., latency, availability, error rate). You will manage and use the Error Budget as the primary drivers to prioritize reliability work</li> <li>Scale and HArden the Core Platform: Apply deep technical expertise in Kubernetes, AWS, traffic management, and Infrastructure-as-Code to scale and harden the foundational platform that powers Iterable’s product workloads. </li> <li>Drive Systemic Improvements: This role centers on <strong>hands-on engineering skill</strong>, <strong>technical leadership</strong>, and <strong>systemic reliability improvements</strong> within our complex, distributed multi-region platform.</li> </ul> <h4><strong>What you’ll do</strong></h4> <ul> <li><strong>Kubernetes Platform Engineering</strong><strong><br/></strong>Use your Kubernetes and AWS expertise to evolve EKS lifecycle, multi-tenant isolation, and regional consistency, ensuring clusters remain secure, performant, and predictable as we scale.</li> <li><strong>Traffic &amp; Ingress Reliability</strong><strong><br/></strong>Apply advanced knowledge of cloud-native traffic management, and API gateways to strengthen routing, authentication, rate-limiting, and secure communication protocols (like mTLS). This focus will dramatically improve both the reliability and security posture of the platform’s public and internal service access points.</li> <li><strong>Infrastructure-as-Code at Scale</strong><strong><br/></strong>Demonstrate mastery in IaC to manage complex, multi-region architecture. Use tools like Terraform Cloud to build reusable modules, validate changes through policy-as-code, and establish safe multi-account patterns our teams can rely on.</li> <li><strong>Security &amp; Access Control</strong><strong><br/></strong>Drive a zero-trust posture by establishing service guardrails and access controls across the platform: This includes: implementing policy-as-code solutions,brokering least-privilege access for platform using cloud Identity and Access Management (IAM) best practices, and Integrating and managing identity providers to define Role-Based Access Control (RBAC) across environments.</li> <li><strong>Reliability Engineering &amp; Incident Leadership</strong><strong><br/></strong>Demonstrate strong diagnostic and incident-response leadership to rapidly isolate issues across clusters, networks, and workloads. Your ultimate responsibility will be to lead and drive systemic, long-term fixes and root-cause investigations, ensuring all necessary actions are taken to eliminate repeat failures .</li> <li><strong>Collaboration, Influence &amp; Mentorship</strong><strong><br/>

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Iterable

View company profile →