Senior Manager, Production Engineering
GeminiAbout the role
Empower the Individual Through Crypto
Gemini is a crypto exchange and custodian that allows customers to buy, sell, store, and earn more than 30 cryptocurrencies like bitcoin, bitcoin cash, ether, litecoin, and Zcash. Gemini is a New York trust company that is subject to the capital reserve requirements, cybersecurity requirements, and banking compliance standards set forth by the New York State Department of Financial Services and the New York Banking Law. Gemini was founded in 2014 by twin brothers Cameron and Tyler Winklevoss to empower the individual through crypto.
Crypto is about giving you greater choice, independence, and opportunity. We are here to help you on your journey. We build crypto products that are simple, elegant, and secure. Whether you are an individual or an institution, we want to help you buy, sell, and store your bitcoin and cryptocurrency. Crypto is not just a technology, it's a movement.
At Gemini, our mission is to empower the individual and that includes giving our employees flexibility of choice — our Office Optional Policy allows employees to choose to work from one of our physical locations or from home.
The Department: Platform
Our Platform organization’s purpose is to enable Gemini to scale effectively and empower our engineering teams to focus on building innovative financial products and experiences for individuals around the world. Within Platform, the Site Reliability Engineering team is responsible for partnering with Gemini’s other engineering teams to ensure all our systems are architected, engineered and deployed to be resilient, reliable and performant.
The Embedded SRE team is a part of Site Reliability Engineering with a focus on engaging directly with our other engineering teams to onboard them onto our platform systems, reviewing and recommending design and architectural decisions, and guiding our engineering teams on how to implement the tooling provided by the larger Platform organization required to ensure systems can scale and react to changing conditions, with continuous improvement loops.
The Role: Senior Manager, Production Engineering
In this position, you will lead a team of skilled Site Reliability Engineers responsible for the design, deployment, and maintenance of our production systems. You will play a crucial role in ensuring the reliability, scalability, and performance of our infrastructure, as well as driving continuous improvement initiatives. Your expertise in SRE practices and experience with the listed technologies will enable you to effectively guide the team towards achieving operational excellence.
Responsibilities:
- Lead and manage a team of Site Reliability Engineers, fostering a culture of collaboration, innovation, and operational excellence.
- Develop and execute the SRE team's strategic goals, objectives, and roadmap in alignment with the overall business objectives.
- Oversee the design, implementation, and maintenance of highly available and scalable production systems.
- Drive continuous improvement initiatives by identifying areas for enhancement and implementing best practices, automation, and process improvements.
- Collaborate with cross-functional teams and Departments to ensure smooth integration of applications and systems.
- Define and enforce Service Level Objectives (SLOs) and Service Level Agreements (SLAs) to ensure system reliability and uptime.
- Monitor system performance, troubleshoot issues, and ensure timely incident response, root cause analysis, and problem resolution.
- Implement effective monitoring, logging, and alerting systems to proactively identify and mitigate potential issues.
- Stay up-to-date with industry trends, emerging technologies, and best practices related to SRE and DevOps, and apply them to improve operational efficiency.
Minimum Qualifications:
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
- Proven experience as a Site Reliability Engineer or similar role, with at least 5 years of hands-on experience in managing production systems.
- Strong expertise in the listed technologies: Ansible, Concourse CI, Jenkins, Github Actions, EKS (Kubernetes), Linux Administration.
- Demonstrated experience in leading and managing a team of technical professionals.
- Solid understanding of SRE principles, including reliability, scalability, availability, and performance.
- Proficient in scripting and automation (e.g., Python, Bash, or similar).
- Experience with infrastructure-as-code (IaC) tools, configuration management, and CI/CD pipelines.
- Knowledge of cloud platforms (e.g., AWS, Azure, or Google Cloud) and containerization technolo
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s