Site Reliability Engineer
AlloyAbout the role
Alloy is where you belong!
Alloy helps solve the identity risk problem for companies that offer financial products by enabling them to outpace fraud and confidently serve more people around the world. Over 600 of the world’s largest financial institutions and fintechs turn to Alloy to take control of fraud, credit, and compliance risk, and grow with the clearest picture of their customers.
Through our values: Be Bold, Get Scrappy, Collaborate, and Celebrate Our Differences, we are creating a workplace where you can grow, thrive, and belong. See how we’ve been continuously recognized and named one of Inc. Magazine’s Best Workplaces, Forbes America’s Best Startup Employers, Best Fintech to Work for by American Banker, year after year.
Check out our investors and read more about us here.
About the team
Alloy’s Infrastructure Team is composed of 6 engineers who are responsible for implementing, improving, and maintaining Alloy’s infrastructure.
Alloy operates in a hybrid-work environment. We look to foster collaboration and community by having our local employees onsite twice a week.
What you'll be doing
Our services are used by leading fintechs and top tier banks. Our clients rely on Alloy’s uptime to serve their clients. Reporting into the Engineering Manager of Infrastructure, you'll help architect and build infrastructure solutions to improve our uptime and exceed our SLOs. We are looking for someone who can:
- Provision and manage a variety of AWS resources, using Terraform.
- Implement solutions for deploying applications to Kubernetes in production.
- Help architect and build secure and reliable systems and deployment pipelines.
- Write and review code comfortable.
- Think pragmatically; can justify when to build and when to buy. It's important to recognize that what we’re creating is part of something bigger, and always look to eliminate constraints and keep releases flowing.
- Implement and use tools like Datadog, Splunk, or Newrelic to find where the latency in distributed systems is coming from, then propose and implement solutions to reduce that latency.
- Participate in on-call rotations, but spend your day building systems that are resilient and self-healing, so alerts don’t happen.
Writing infrastructure as code
- Our infrastructure is represented in Terraform. You’ll be working with the team to make improvements to our IAC, as well as provision new infrastructure in support of our broader engineering goals.
Automating
- We use a collection of AWS Tools, Github Actions, third-party solutions, and custom scripts to automate releases, deployments, scaling events, and other actions. You’ll help improve existing automated processes as well as create new ones to improve resiliency, performance, and predictability.
Supporting application developers
- Our shared goal of shipping quality code means you’ll be looking for and eliminating constraints in the deployment pipeline, helping to maintain dev-prod parity, and looking for ways to ensure applications get to production quickly, predictably, and drama-free.
Continuously improving
You’ll be continuously looking for the opportunities to improve uptime, autoscaling and autorecovery, while reducing RTO and RPO of our products. Suggesting and implementing new cloud services and taking advantage of the latest AWS releases.
Offering cost optimizations, improving our logging, monitoring, and alerting solutions.
You’ll also participate in our on-call rotation. Fortunately, because of our focus on continuous improvement, resiliency, and automation, incidents are rare, recoverable, and quickly followed up by blameless postmortems.
Who we’re looking for
- 4+ years of experience working on a DevOps, SRE, or Infrastructure team
- Bonus points for software engineering experience
- Experience writing Infrastructure as Code
- Experience running and troubleshooting applications in Docker
- Solid experience with CI/CD with tools like GitHub Actions, CircleCI, Travis, etc
- Experience configuring and using tools like Datadog, Cloudwatch, ELK/EFK, etc for troubleshooting
- Experience being in an on-call rotation
- Programming experience with one or more of python, javascript, golang, etc
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s