Principal Site Reliability Engineer (Northridge, CA)
MedtronicAbout the role
At Medtronic you can begin a life-long career of exploration and innovation, while helping champion healthcare access and equity for all. You’ll lead with purpose, breaking down barriers to innovation in a more connected, compassionate world.
A Day in the Life
Medtronic is hiring a Principal Site Reliability Engineer (SRE). The Principal Site Reliability Engineer is responsible for the overall health, performance, and reliability of mission-critical applications and systems, including SAP, Salesforce, Jira, and Confluence. This role combines software engineering and IT operations to ensure seamless deployment, stability, and scalability of systems and applications. You will work closely with cross-functional teams to design and implement infrastructure solutions, develop automation strategies, and drive continuous improvement initiatives. The ideal candidate will be a proactive problem solver with a deep understanding of software development, cloud infrastructure, and DevOps practices. This role also includes a focus on security, compliance, and business continuity, ensuring that applications and infrastructure are resilient, secure, and aligned with industry standards.Responsibilities may include the following and other duties may be assigned.
Design, implement, and maintain scalable, highly available systems architecture for critical business applications like SAP, Salesforce, Jira, and Confluence.
Proactively monitor and manage performance, availability, and capacity while implementing automation strategies to streamline processes and reduce manual efforts.
Collaborate with development and IT teams for seamless deployments, configuration changes, and infrastructure upgrades, ensuring alignment with business needs.
Conduct root cause analysis, resolve complex technical issues, and implement SRE best practices, focusing on service-level objectives (SLOs) and indicators (SLIs).
Manage and optimize cloud infrastructure (including hybrid and multi-cloud environments) to ensure cost efficiency, performance, and scalability.
Lead disaster recovery planning, business continuity initiatives, and capacity planning to maintain system resilience and availability.
Implement security and compliance best practices, monitoring for vulnerabilities, and ensuring adherence to industry standards (e.g., SOX, GDPR, HIPAA).
Drive change and release management processes, planning and coordinating production deployments, and optimizing workflows.
Serve as a technical expert in enterprise applications, developing comprehensive documentation, and mentoring junior team members.
Collaborate with stakeholders to prioritize projects, align SRE initiatives with business goals, and provide technical leadership across the organization.
Knowledge, Skills, and/or Abilities
In-depth knowledge of enterprise application systems, including SAP, Salesforce, Jira, and Confluence.
Advanced experience with cloud platforms (AWS, Azure, or GCP) and cloud-native services, such as container orchestration (Kubernetes) and serverless computing.
Strong proficiency in scripting languages such as Python, Bash, or PowerShell for automating infrastructure and application management tasks.
Solid understanding of CI/CD pipelines, GitOps, and Infrastructure as Code using tools like Jenkins, GitHub Actions, Terraform, or CloudFormation.
Experience with monitoring, logging, and observability tools such as Prometheus, Grafana, ELK Stack, or Splunk.
Familiarity with IT service management (ITSM) practices and tools like ServiceNow.
Excelle
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s