Senior Site Reliability Engineer
GorgiasAbout the role
Meet Gorgias, the customer service platform designed for ecommerce merchants, and built to provide amazing experience to shoppers at scale on Shopify, BigCommerce, and Magento. Our product empowers merchants to manage all their customer service in one place over email, live chat, voice, Facebook, Instagram, Twitter, and SMS.
Everything we do is for our customers, and we’re currently serving over 12,000+ ecommerce merchants, including : Steve Madden, Timbuk2, Decathlon, and Sports Illustrated. They love us for our innovative product, our focus on their ecommerce needs, and, of course, our lightning-fast customer service response time.
We raised $25 million in our Series B round in December 2020 and $30 million in our Series C round in 2022. We more than doubled in size in every meaningful way: annual recurring revenue, the size of our customer base, and the size of our Gorgias team, for starters.
We’re still growing fast and looking for new teammates who want to grow with us.
About The SRE Team
✍️ We are seeking a highly skilled and experienced Senior Site Reliability Engineer (SRE) to join our team. As an SRE at Gorgias, you will play a crucial role in ensuring the reliability, scalability, and performance of our systems, enabling the seamless delivery of our products and services.
🙋 The SRE team at Gorgias maintains the core infrastructure and services that make up the heart of our product. We have the privilege of building solutions for our high throughput systems and TB-scale data stores serving billions of queries per day. Due to our efforts many core workloads see an average of sub-millisecond response times.
What You Will Do:
Work closely with other development teams to design, implement, and maintain scalable and reliable infrastructure solutions.
Implement and manage infrastructure as code (IaC) using tools such as Terraform to automate deployment and configuration processes. Provide modules for other teams to leverage in their projects.
Write scripts and programs in Python, Bash, Go, etc. for in-house tools and automation.
Utilize Kubernetes to orchestrate containerized applications, ensuring efficient scaling, load balancing, and fault tolerance.
Work closely with Cloud Providers (e.g., AWS, GCP, Azure) to optimize infrastructure, leverage cloud-native services, and ensure high availability as well as cost-optimizations.
Manage and optimize PostgreSQL databases, including performance tuning, backup, and recovery strategies.
Develop and maintain monitoring, alerting, and logging solutions to proactively identify and address potential issues.
Implement security best practices, conduct regular security audits, and participate in incident response activities.
Participate in an on-call rotation to respond to and resolve production incidents promptly.
What You Should Have:
Bachelor's degree in Computer Science or equivalent work experience.
5+ years experience as a Site Reliability Engineer or similar role, with a focus on maintaining high-performance, scalable, and reliable high-throughput web systems.
Proficiency in using Kubernetes for container orchestration and management.
3+ years experience with Cloud Providers ( GCP, AWS) and a deep understanding of cloud services and architectures.
Proficient in scripting and programming languages such as Python, Bash, Go, or NodeJS.
Comfortable and confident in Linux systems and the command line.
Solid understanding of infrastructure as code (IaC) principles and experience with tools like Terraform.
Experience with continuous integration and deployment (CI/CD) pipelines.
Excellent problem-solving and troubleshooting skills.
Strong communication and collaboration skills with the ability to work effectively in a team environment.
Bonus Points If You Have:
Certification in Kubernetes (e.g., Certified Kubernetes Administrator - CKA).
Certification in a Cloud Provider platform (e.g., AWS Certified Solutions Architect, Google Cloud Professional Cloud Architect).
Experience in managing and optimizing PostgreSQL databases.
Perks and Benefits
🏖️ 5-week vacation plus 2 weeks RTT (We follow each country's appropriate PTO Laws)
🤕 Paid sick leave
🧸 Paid parental leave (16 weeks)
💻 MacBook Pro
🍽️ Personal credit card to buy lunches (we use Swile)
🏥 We provide private health insurance (we use
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s