Site Reliability Engineering Team Lead
OANDAAbout the role
Everyone at OANDA is focused on our vision to transform how our customers can meet all their currency needs. We are revolutionising the world of currency trading by providing innovative trading experiences, currency data and analytics solutions. Dare to be open, bold, focused - own it and apply! The future is now!
OANDA is looking for a passionate Team Lead of Site Reliability Engineering to lead a team, organize and monitor work processes, and apply practices to solve difficult problems.
How do we work?
As an SRE Team Leader, you will be responsible for leading the relationship with our development teams, acting as the champion for reliability best practices including observability, automation, high-availability, fault tolerance, and full-lifecycle ownership.
The perfect candidate for this role has a strong data-driven approach to improving the performance of our products in on-premise and cloud environments.
In this role, you will:
Deploy a team of production engineers to development teams to champion SRE and DevOps best practices and to ensure our products are designed for reliability and high availability
Provide technical guidance regarding system architecture (analyzing data, making suitable recommendations)
Ensure security principles and respond to deficiencies
Make data-driven decisions by pushing monitoring, instrumentation, and observability as core tenets of our development practice
Demonstrate ability to manage large cross-functional projects, scope, and strategy
Build partnerships and work collaboratively with others to meet shared objectives
Be responsible for performance management and budget plans for a department
What skillset you need, to be successful in this role:
Excellent communication (English) and presentation skills
Strong leadership skills, including mentoring, team-building, and conflict resolution
Organizational and strategic management skills, including budget and business planning and forecasting
Experience with Git based code hosting and collaboration tools (BitBucket, GitHub, etc.)
Familiarity with monitoring tools and platforms (e.g. Grafana, Prometheus, Zabbix)
Understanding of SRE principles, including monitoring, alerting, error budgets, fault analysis, and other common reliability engineering concepts
Strong experience working in cloud-native and on-premise environments, in bare metal, virtualized (VM), and containerized/orchestrated deployments (GCP, Docker, Kubernetes, Cloudflare, CircleCI, ArgoCD, GoCD, etc.)
Experience working with Infrastructure as a code and configuration management tool (Ansible, Terraform, Helm, etc.)
Deep knowledge of Linux
CI/CD pipeline management
OANDA Global Corporation is a diverse and global team with offices around the world. We value the unique skills and experiences each individual brings to OANDA. We are committed to creating and sustaining a collegial work environment in which all individuals are treated with dignity and respect and one which reflects the diversity of the community in which we operate. We provide an inclusive and accessible environment for everyone. Candidates selected for an interview will be contacted directly. If you require accommodation during the recruitment and selection process, please let us know. We will work with you to provide as seamless a recruitment experience as possible.
Learn more about our culture here.
Review OANDA Privacy Policy and learn more about how we treat your personal data and protect your privacy.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s