Senior Site Reliability Engineer, Performance
FleetioAbout the role
Description
Our Platform Engineering team is looking for a Senior Site Reliability Engineer to help run, maintain and improve the performance of our Ruby on Rails Stack and Infrastructure. You will help scale our application using best in class architecture and software design. This includes training, software engineering, system design and operational practices that support the needs of our engineers and customers while accounting for future growth. You will be entrusted with proactively identifying and owning initiatives that will help improve the performance, reliability, and scalability of our application stack and databases.
At Fleetio, we foster a culture of 'Product Engineers', where we value and prioritize engineers who enjoy being part of the product discovery process. Our Engineering and Product teams are structured as autonomous PODs that execute within one focal area toward a defined product vision. We strive to deliver easy-to-use software, and our goal as engineers is to quickly and continuously deliver meaningful value to our customers. We've optimized our CI/CD tools and processes to easily get code into our production environments, resulting in an average of 40 deploys per week.
Fleetio is a modern software platform that helps thousands of organizations around the world manage their fleet operations. Transportation technology is a hot market and we’re leading the charge, with raving fans and new customers signing up every day. We raised $144M in Series C in June of 2023 and are on an exciting trajectory as a company. Fleetio is also a proud founding member of the Rails Foundation!
More About Our Team and Company
- Watch our culture videos: https://fleet.io/culture
- Engineering culture, interview process and videos: https://www.fleetio.com/careers/engineering
- Fleetio overview video: https://www.youtube.com/watch?v=IlvIbwZT3oU
- Fleetio Go overview video: https://www.fleetio.com/go
- More about the Fleetio platform: https://www.fleetio.com/features
- API docs: developer.fleetio.com
- Test drive Fleetio to get an even better feel for what we're building: https://www.fleetio.com/register
This is a remote opportunity and is open to candidates in the United States, Canada, or Mexico.
Who You Are
Our ideal candidate is an Infrastructure Engineer experienced in scaling Ruby on Rails applications, with a passion for optimization and performance improvements. You bring a strong background in Site Reliability and Infrastructure Engineering for Rails applications. You follow Agile and DevOps principles, can effectively influence teams to achieve goals, and demonstrate excellent problem-solving skills in our fast-paced environment.
Your Impact
As a Site Reliability Engineer on Fleetio’s Platform Engineering team, you will:
- Proactively identify, triage, and resolve performance issues
- Enhance system observability by monitoring performance metrics across Ruby, Rails, and database systems, including SLOs and SLIs
- Guide product engineers on Ruby/Rails performance and database best practices through code reviews and pair programming
- Optimize performance through instance configuration and monitoring
- Collaborate with other SREs to proactively identify and address performance bottlenecks
- Lead database capacity planning and upgrade initiatives
- Manage the database-specific components of disaster recovery planning and execution
- Oversee backup systems and pre-production databases
- Create and maintain infrastructure and operations documentation
- Participate in the on-call rotation
Your Experience
- 5+ years of Ruby/Rail Experience
- 3+ years of AWS Experience
- Kubernetes experience
- Experience with profiling and benchmarking source code
- Effective at code review, and identifying potential performance problems before they reach production
- Experience with Datadog or other APM tools
- Excellent written and verbal communication skills
Considered a Plus
- Infrastructure as Code tools (Terraform)
- Deep understanding of cloud network fundamentals (routing, firewalls, load balancers, CDNs, VPCs, etc.)
- Experience with distributed event and data stores, such as Kafka, Redis, Elasticsearch, Memcached, and TimescaleDB
- You know a thing or two about the fleet management industry
Be
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s