Site Reliability Engineer
Bosonai · Toronto, Toronto
On-site11 days agoApply →About The Role We're seeking an experienced Network Engineer to design, build, and optimize the high-performance networking infrastructure powering our AI/ML operations in Toronto. You'll work at the cutting edge of network technology—managing InfiniBand and ultra-high-speed Ethernet fabrics that connect NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, and hundreds of servers. You'll be hands-on with the full lifecycle of our network infrastructure: planning, building, testing, deploying, and keeping everything running at peak performance. That means troubleshooting issues as they arise, monitoring network performance and throughput, developing automation to streamline operations, and working closely with HPC and ML teams to ensure they have the bandwidth they need. You'll also help us plan for future capacity and evaluate emerging network technologies as we scale to meet increasingly demanding workloads.
Cloud Performance Engineering - Site Reliability Engineer ( Remote Canada)
Smiledigitalhealth · Toronto, Ontario
Remote15 days agoApply →Working for a company like Smile Digital Health means supporting our mandate for #BetterGlobalHealth. We strive towards this goal every day, and the results can be seen in the impact of our innovative health data platform and data management solutions, which are used in over 20 countries. We were #19 on Deloitte's Technology Fast 50 Ranking for 2024! Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform. At its heart, the Smile platform enables people and organizations to better manage healthcare data. We help generate and liberate structured healthcare data to ensure effective delivery across care teams and health systems bringing #BetterGlobalHealth to patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners. This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.
Site Reliability Engineer, Inference Infrastructure
cohere · Toronto, USA
On-site16 days agoApply →Senior Site Reliability Engineer
highlightta · Toronto, Ontario, USA
On-site16 days agoApply →Manager, Site Reliability Engineering
docebo · Toronto, Ontario, USA
On-site17 days agoApply →Senior Site Reliability Engineer (SRE)
acquird · Toronto - Hybrid, USA
Hybrid21 days agoApply →Site Reliability Engineer (Senior or Staff), Deployments
mongodb · Toronto, Toronto
On-siteabout 1 month agoApply →Site Reliability Engineer - Canada Wide - Remote
Newton · Toronto, Ontario
Remoteabout 2 months agoApply →Say hello to Newton! We're changing how Canadians trade crypto. Our goal? To make financial freedom something everyone can achieve. We give our customers the tools and knowledge they need to navigate the crypto world. At Newton, you'll work with a remote team spread across Canada, but you'll never feel distant. Ready to be part of something meaningful? Join a team that’s all about pushing boundaries and getting things done. Some of our values: ● Customer first mindset - Commitment to integrity and transparency to our users! ● A dynamic team fueled by collaboration uniting our strengths to overcome any obstacles. Together we build success. We persevere, adapt, and come back stronger, turning obstacles into opportunities. ● We strive for continuous improvement and embrace creativity and encourage experimentation. We push the boundaries of what’s possible and continuously explore new ideas, technologies, and solutions. Role Overview: We're looking for a Site Reliability Engineer to improve the reliability, resilience, and operational readiness of our services. You’ll work closely with engineering teams to improve system design and operational excellence. You’ll help prevent incidents, lead response efforts, and drive improvements through post-mortems. Your mission: ensure our systems are reliable, scalable, and resilient. Responsibilities will include: • Implementing the improvements to the reliability, fault tolerance, scalability, and performance of our infrastructure • Managing incidents using your technical know-how to involve the appropriate teams and automate away manual practices • Providing support to our critical services by responding to automated alerts through our on-call rotation • Define and maintain SLIs, SLOs,SLA, and error budgets to guide reliability decisions • Improve observability across our systems (metrics, logs, tracing) to reduce time to detection and resolution • Make production issues easier to detect, troubleshoot, and resolve • Improving monitoring, alerting, dashboards, tracing and runbooks for critical services • Leading postmortems and follow-up actions to reduce repeat incidents Who you are: • You have experience designing and operating scalable, reliable systems in AWS or a similar cloud environment • You have handled on-call shifts for critical systems • You are experienced with chaos engineering (i.e. Gremlin) • You are able to dive in and debug live production systems • You enjoy working in a growing system, and writing and deploying code without any downtime • You have experience scripting and/or development (i.e. Linux Shell, Python, Javascript, Java) • You are a self-starter, taking initiative in an ambiguous space preferably within a start-up environment
Site Reliability Engineer (.Net)
tipaltisolutions · Toronto, Canada
On-siteabout 2 months agoApply →Senior Site Reliability Engineer (Auth0)
okta · Toronto, Canada
On-siteabout 2 months agoApply →Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
mongodb · Montreal; Toronto, Montreal; Toronto
On-siteabout 2 months agoApply →Staff Site Reliability Engineer, Fabric
mongodb · Toronto; Vancouver, Toronto; Vancouver
On-siteabout 2 months agoApply →Site Reliability Engineer
momentumfinancialservicesgroup · Toronto, Canada
On-siteabout 2 months agoApply →Site Reliability Engineer
maintainx · Montreal & Toronto, Montreal & Toronto
On-siteabout 2 months agoApply →Senior Site Reliability Engineer
fundedclub · Toronto Ontario Canada, Toronto Ontario Canada
On-siteabout 2 months agoApply →Site Reliability Engineer - Ops & Automation
cerebrassystems · Sunnyvale CA or Toronto Canada, Sunnyvale CA or Toronto Canada
On-siteabout 2 months agoApply →Senior DevOps & Site Reliability Engineer - Americas
appspace · Toronto, Canada
On-siteabout 2 months agoApply →