Site Reliabilty Engineer
FloatAbout the role
Who We Are
Float is the leading resource management software for professional services teams. Since 2012, we’ve grown every year—independently, self-funded, and profitably. We’re rated #1 for resource management on G2 and trusted by 4,500+ customers worldwide.
As a certified B Corporation, we’re committed to making a positive impact on our team, customers, the environment, and the remote community. Our 50+ person team works 100% remotely across the globe, with perks and benefits designed to support us in living our Best Work Life. You'll collaborate with teammates across Australia, Mexico, the UK, Nigeria, Canada, and the US. Learn more about our data security practices for employment or service contracts here. Browse our blog to get a glimpse of life at Float and check out our Glassdoor employer reviews. See why our customers love Float on G2 .
We’re on a scale-up journey, and we’re seeking people who thrive in this stage. We want Float to be the place where you have the autonomy and opportunity to do the best work of your career.
Why We’re Hiring For This Role
Float’s infrastructure has grown rapidly, meaning more customers, more complex systems, and more opportunities to build for scale. As the scale of our systems increases, we’re growing our SRE team to match. You’ll be the third site reliability engineer, and will be working alongside our QA team. This role is about stepping into a high-impact space: helping us automate smarter, improve visibility across engineering, and ensure reliability as we scale. You’ll join a team that’s laying the groundwork for stronger SLAs and an even better experience for our customers.
This role will report into Chris, our Team Lead for SRE & QA. Check out this video where he explains the important role you will play within our SRE team. Watch this video!
You’ll be working asynchronously with a bright, dedicated team from across the globe, with a strong focus on taking complex problems and creating solutions that feel simple and intuitive for our customers.
What You’ll Be Responsible For
Early on, you’ll jump right into:
- Upgrade paths: Maintain and validate the processes that keep our Kubernetes infrastructure up-to-date, ensuring upgrades happen smoothly, safely, and regularly.
- Service hygiene: Remove noisy, unused, or misfiring boot alerts and improve the team's ability to trust alerts as meaningful signals.
- Service integration: Partner with engineers to configure services within our clusters and support service migrations where possible.
- Kubernetes optimisation: Review and optimise usage across Kubernetes services, including right-sizing scale node specifications.
Once you are a bit more settled, we expect that you will jump into the following projects:
- Service mesh & ingress security: Lead our exploration and implementation of service mesh options and harden ingress layers to defend against spam and abuse.
- Incident response playbooks: Define and roll out standardised playbooks to improve clarity and speed during production incidents.
- CDC layer support: Build deep familiarity with our next-gen data layer (CDC) to support new teams building on top of it.
- SLO coaching & support: Help teams define, measure, and meet reliability goals—enabling engineering to own quality into production and drive better outcomes for customers.
What You’ll Need To Be Successful
We want you to love your work and believe that these skills will allow you to succeed in the role. Applying these skills requires:
- Bash + programming language: Confident writing scripts in Bash and proficient in at least one go-to language (ideally PHP, NodeJS, or Python).
- Kubernetes: Strong production experience
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s