Director of Site Reliability Engineering (SRE) - New York
GalaxyAbout the role
Who We Are:
Galaxy is a global leader in digital assets and data center infrastructure, delivering solutions that accelerate progress in finance and artificial intelligence. We believe that blockchain and digital asset innovation will transform how value moves through the world – and we’re building the products and services to make that future a reality.
Our institutional digital assets platform spans trading, investment banking, asset management, staking, self-custody, and tokenization technology. We also invest in and operate cutting-edge data center infrastructure to power AI and high-performance computing, addressing the growing demand for scalable energy and compute in the U.S.
We work at the intersection of finance and technology, helping institutions, startups, and developers navigate a digitally native economy. Led by CEO and Founder Michael Novogratz, our team blends deep crypto expertise with institutional experience and a shared commitment to shaping the future of Web3 and AI.
Galaxy is headquartered in New York City, with offices across North America, Europe, the Middle East, and Asia.
To learn more about our businesses and products, visit www.galaxy.com.
What We Value:
We are a diverse team of free thinkers, and fast movers united to help investors and creators energize the global economy. We are looking for individuals who thrive in a culture of builders and overachievers and embrace high performance, transparent feedback, and a mission-first approach. Our culture shapes our way of working and gets us where we want to be.
- Seek Excellence.
- Be Selective To Be Effective.
- Be Highly Aligned, Loosely Coupled.
- Disagree Transparently.
- Encourage Independent Decision-Making.
- Build Dream Teams.
Who You Are:
We are seeking a Director of Site Reliability Engineering (SRE) to lead our reliability, automation, and performance initiatives. You will drive the development of scalable, secure, and resilient infrastructure while working closely with engineering and security teams. A deep understanding of crypto/blockchain ecosystems, Kubernetes, containers, Terraform, and observability is required.
What You’ll Do:
- Strategic Leadership & Reliability Engineering
- Define and execute the SRE strategy to enhance system reliability, scalability, and security
- Lead a team of SREs, fostering a culture of automation, observability, and continuous improvement
- Implement and refine SLOs, SLAs, and error budgets to balance innovation with system stability
- Collaborate with security teams to enforce best practices for crypto-related workloads
- Infrastructure & Automation
- Architect and manage containerized environments using Kubernetes and related cloud-native technologies
- Oversee Infrastructure as Code (IaC) using Terraform, ensuring repeatability and compliance
- Develop CI/CD pipelines to accelerate software delivery with security and compliance in mind
- Improve auto-scaling, failover, and disaster recovery strategies for highly available systems
- Observability & Incident Response
- Build and enhance observability capabilities using logging, tracing, and metrics tools (e.g., Datadog, Prometheus, OpenTelemetry)
- Define monitoring and alerting strategies to proactively detect and mitigate failures
- Lead post-mortem reviews, driving improvements in system design and operational playbooks
- Develop automated solutions to minimize MTTR (Mean Time to Recovery) and improve incident response
- Blockchain & Crypto-Specific Operations
- Optimize infrastructure for crypto/blockchain nodes, validators, and smart contracts
- Implement security best practices for key management, RPC endpoints, and decentralized protocols
- Ensure regulatory compliance and high availability for blockchain-based applications
What We’re Looking For:
- 10+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering
- Deep expertise in Kubernetes, containers, and service mesh architectures
- Strong proficiency in Terraform and other Infrastructure as Code (IaC) tools
- Extensive experience with AWS and on-prem environments
- Hands-on experience with observability stacks (e.g., Datadog, Prometheus, Grafana, OpenTelemetry)
- Experience securing and optimizing crypto/blockchain infrastructure (e.g.,Ethereum, Solana, Bitcoin, L2s)
- Proven experience leading high-performing SRE teams
- Ability to work cross-functionally with engineering, security, and product teams
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s