Senior Site Reliability Engineer - Payward Services
KrakenAbout the role
Building the Future of Crypto
Our Krakenites are a world-class team with crypto conviction, united by our desire to discover and unlock the potential of crypto and blockchain technology.
What makes us different?
Kraken is a mission-focused company rooted in crypto values. As a Krakenite, you’ll join us on our mission to accelerate the global adoption of crypto, so that everyone can achieve financial freedom and inclusion. For over a decade, Kraken’s focus on our mission and crypto ethos has attracted many of the most talented crypto experts in the world.
Before you apply, please read the Kraken Culture page to learn more about our internal culture, values, and mission. We also expect candidates to familiarize themselves with the Kraken app. Learn how to create a Kraken account here.
As a fully remote company, we have Krakenites in 70+ countries who speak over 50 languages. Krakenites are industry pioneers who develop premium crypto products for experienced traders, institutions, and newcomers to the space. Kraken is committed to industry-leading security, crypto education, and world-class client support through our products like Kraken Pro, Desktop, Wallet, and Kraken Futures.
Become a Krakenite and build the future of crypto!
Proof of work
The team
This role is fully remote, with a strong preference for candidates in EU timezones. The Payward Services (PWS) business unit powers Kraken's B2B and institutional product suite, serving external partners and institutional clients under contractual SLAs.
As a Senior SRE, you will partner with PWS development and operations teams to manage infrastructure, improve CI/CD pipelines, and support operational excellence. You will bring expertise in infrastructure, monitoring, and automation to ensure performant, resilient, and continuously improving services.
The opportunity
Manage and support infrastructure for Payward Services, including Nomad, Kubernetes, databases, and 3rd party system integration
Provide operational support across multiple teams, helping debug issues in staging and production environments
Participate in incident response and post-incident reviews to improve system resilience
Consult with teams on performance, monitoring, and alerting best practices — with awareness of partner-facing SLA commitments
Build tooling, automation, and dashboards to improve observability and empower development teams
Maintain and troubleshoot CI pipelines, ensuring reliable and fast build, test, and deployment cycles
Collaborate with developers, QA, and product managers to streamline development and release cycles
Support a fully distributed team operating across multiple timezones
Skills you should HODL
5+ years in DevOps or SRE role
Proficiency with hybrid-cloud infrastructure environments
Git source version-control and CI/CD configuration proficiency
Deep understanding of monitoring and alerting systems, preferably Prometheus and Grafana
Ability to debug complex distributed systems, networks, and Linux operating systems issues
Containerization and orchestration experience (Docker, Nomad, Kubernetes a plus)
Strong scripting skills (Bash, Python, or Go)
Self-starter capable of thriving independently and remotely in fast-paced environments
Nice to haves
Background working with distributed systems and technologies (Kafka, gRPC, Redis, etc.)
Experience operating services with external SLAs or in a B2B/enterprise context
Experience with benchmarking, performance tuning, and identi
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s