Jobs and Careers
SP
Senior Site Reliability Engineer
SpreedlyRemote, United States, United StatesRemotefull_timeVerifiedPosted 13 Jul 2025
About the role
About Us:
Spreedly is the world's leading Open Payments Platform, sitting at the center of a network processing more than $50b of GMV annually. Spreedly's Payments Orchestration platform enables and optimizes digital transactions with the world’s most complete payment services marketplace. Built on Spreedly’s PCI-compliant architecture, our Advanced Vault solution combines a modern feature-set with rule-based configurations to optimize the vaulting experience for all stored payment methods. Global enterprises and hyper-growth companies grow their digital business faster by relying on our payments platform. Hundreds of customers worldwide secure card data in our PCI-compliant vault and use tokenized card data to enable and optimize over $45 billion of annual transaction volumes with any payment service.
Our vision is that the world is better with a diversified, inclusive payment ecosystem. Our mission is to accelerate commerce with an open, secure, and flexible payment platform that welcomes all payment participants. Our employees help us execute our vision by building a culture focused on autonomy, transparency, and collaboration in a dynamic, high-growth organization.
Product Offering:
Spreedly provides an open payments platform. The platform’s connectivity provides payments performance. Key products and services include:
Payment Gateway Integration: Connects merchants, platforms, and marketplaces to multiple payment gateways and payment services.Tokenization: Securely stores and manages payment data with a universal tokenization service.Transaction Routing: Enables intelligent routing of transactions to optimize success rates and costs.Payment Vault: A secure storage solution for sensitive payment information.Fraud Tools Integration: Integrates with various fraud prevention tools to enhance transaction security.
About the Role:
As a Senior Site Reliability Engineer (SRE) at Spreedly, you will focus on ensuring the reliability, observability, and scalability of our globally distributed payments platform. You will lead efforts to stabilize and optimize our infrastructure, build platform services, and champion best practices that enhance system performance and resilience. A strong candidate for this role brings deep experience in designing and operating highly available, scalable cloud architectures while fostering a culture of reliability across the organization.
In this role, you will leverage your expertise in software development, infrastructure, and operations to ensure our applications and systems are reliable, scalable, and efficient. You will work across the entire application stack, using a diverse range of tools and technologies to support our mission-critical system.
Spreedly is the world's leading Open Payments Platform, sitting at the center of a network processing more than $50b of GMV annually. Spreedly's Payments Orchestration platform enables and optimizes digital transactions with the world’s most complete payment services marketplace. Built on Spreedly’s PCI-compliant architecture, our Advanced Vault solution combines a modern feature-set with rule-based configurations to optimize the vaulting experience for all stored payment methods. Global enterprises and hyper-growth companies grow their digital business faster by relying on our payments platform. Hundreds of customers worldwide secure card data in our PCI-compliant vault and use tokenized card data to enable and optimize over $45 billion of annual transaction volumes with any payment service.
Our vision is that the world is better with a diversified, inclusive payment ecosystem. Our mission is to accelerate commerce with an open, secure, and flexible payment platform that welcomes all payment participants. Our employees help us execute our vision by building a culture focused on autonomy, transparency, and collaboration in a dynamic, high-growth organization.
Product Offering:
Spreedly provides an open payments platform. The platform’s connectivity provides payments performance. Key products and services include:
Payment Gateway Integration: Connects merchants, platforms, and marketplaces to multiple payment gateways and payment services.Tokenization: Securely stores and manages payment data with a universal tokenization service.Transaction Routing: Enables intelligent routing of transactions to optimize success rates and costs.Payment Vault: A secure storage solution for sensitive payment information.Fraud Tools Integration: Integrates with various fraud prevention tools to enhance transaction security.
About the Role:
As a Senior Site Reliability Engineer (SRE) at Spreedly, you will focus on ensuring the reliability, observability, and scalability of our globally distributed payments platform. You will lead efforts to stabilize and optimize our infrastructure, build platform services, and champion best practices that enhance system performance and resilience. A strong candidate for this role brings deep experience in designing and operating highly available, scalable cloud architectures while fostering a culture of reliability across the organization.
In this role, you will leverage your expertise in software development, infrastructure, and operations to ensure our applications and systems are reliable, scalable, and efficient. You will work across the entire application stack, using a diverse range of tools and technologies to support our mission-critical system.
Responsibilities:
- System Reliability & Performance: Ensure the reliability, availability, and performance of Spreedly’s globally distributed payments platform, processing $4B monthly production systems through monitoring, automation, and continuous improvement.
- Application Development Support: Collaborate with development teams to improve the reliability and performance of Ruby on Rails and Elixir applications.
- Observability & Monitoring: Implement and maintain robust observability solutions using Datadog and OpenTelemetry, enabling proactive identification alerting, and resolution of issues.
- Incident Management: Lead incident response efforts by participating in a shared on-call rotation to maintain 24/7 system reliability, including root cause analysis, resolution, and implementing measures to prevent recurrence.
- Automation & Tooling: Develop and maintain automation tools to reduce manual intervention, streamline operations, and enhance developer productivity.
- Database Performance Tuning: Monitor, analyze, and optimize the performance of relational databases, identifying and resolving bottlenecks to maintain data integrity and efficiency.
- Thought Leadership: Lead by example, infusing modern SRE best practices and fostering a culture of reliability and performance within the engineering organization.
- Mentorship: Provide technical guidance and mentorship to team members, fostering a culture of learning and collaboration.
Requirements:
- Observability Tools: Hands-on experience with Datadog, OpenTelemetry, Sentry, and Sumo Logic or similar monitoring and observability platforms, with a focus on actionable metrics and alerts.
- Programming Expertise: Proficiency in a modern programming language, with a proven ability to write clea
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s