Jobs and Careers
PR

Senior Software Engineer IS, Cloud Infrastructure *Virtual*

Providence
United Statesfull_timeVerifiedPosted 21 May 2025

About the role

Providence is one of the nation's leading non-profit healthcare systems with 119,000+ caregivers/employees serving more than 5 million unique patients across 51 hospitals and 800+ clinics. Our locations range from metropolitan centers to rural settings across seven states: Alaska, California, Montana, New Mexico, Oregon, Texas, and Washington. As a mission-based, not-for-profit healthcare provider, our commitment to providing compassionate care to all lives on through our five core values: Compassion, Dignity, Justice, Excellence, and Integrity.

 

Providence caregivers are not simply valued – they’re invaluable. Join our team at Enterprise Information Services and thrive in our culture of patient-focused, whole-person care built on understanding, commitment, and mutual respect. Your voice matters here, because we know that to inspire and retain the best people, we must empower them.

 

We are seeking a skilled Sr Software Engineer – Infrastructure Telemetry and Site Reliability Engineer (SRE) to join our dynamic platform team. The ideal candidate will be responsible for ensuring the reliability, availability, and performance of our systems while leveraging telemetry data to enhance monitoring and observability. This role is critical in maintaining our high service standards and continuously improving our infrastructure. 

 Providence welcomes virtual work for applicants who reside in one of the following states: 

  • Oregon

Key Responsibilities 

  • Lead the design, develop, and implement monitoring, logging, and alerting solutions to ensure system reliability and performance
  • Utilize telemetry data to identify and troubleshoot issues, optimize system performance, and enhance overall observability
  • Collaborate with development and operations teams to ensure seamless integration of monitoring and alerting tools
  • Write and maintain scripts for infrastructure management and automation (e.g., Python, PowerShell, Bash)
  • Automate repetitive tasks to improve efficiency and reduce manual intervention
  • Participate in on-call rotations and incident response, providing timely resolution to system outages and performance issues
  • Develop and maintain documentation for system architecture, processes, and procedures related to telemetry and site reliability
  • Design and implementation of cloud infrastructure using Infrastructure as Code (IaC) tools such as Terraform, AWS CloudFormation, or Azure Resource Manager
  • Automate deployment pipelines using CI/CD tools such as Jenkins, GitHub Actions, or Azure DevOps
  • Collaborate with cross-functional teams to design and implement scalable and resilient infrastructure solutions
  • Conduct root cause analysis of incidents and implement corrective actions to prevent recurrence
  • Drive the adoption of best practices in site reliability engineering and telemetry within the organization

Required qualifications:

  • Bachelor’s degree in computer engineering, computer science, mathematics, engineering -OR- a combination of equivalent education and experience
  • 5 or more years of related experience
  • Extensive experience with object-oriented programming in C#, Java, Python or equivalent
  • Experience with source code control systems such as Git
  • SQL integration development experience with SQL/NoSQL
  • Experience with Agile software development methodologies and tools such as Azure Devops, TFS, and Jira
  • Proven track record of working both independently and collaboratively as part of a multi-disciplined team
  • Experience designing and successfully implementing a large project

 

Preferred qualifications:

  • Experience in a healthcare setting
  • Software development
  • 5 or more years of experience in software engineering with a focus on site reliability engineering, DevOps, IaC and Cloud Infrastructure or another related field 
  • Strong knowledge of monitoring, logging, and alerting tools (e.g., Datadog, Prometheus, Grafana, ELK stack, Splunk, New Relic)
  • Proficiency in programming and scripting languages (e.g., Python, Go, Bash)
  • Experience with cloud platforms (e.g., AWS, Azure, Google Cloud) and containerization technologies (e.g., Docker, Kubernetes)
  • Strong understanding of Linux/Unix systems and networking concepts
  • Experience with configuration management and automation tools (e.g., Terraform, Ansible, Puppet, Chef)
  • Familiarity with CI/CD pipelines and tools (e.g., Jenkins, GitLab CI, CircleCI) is a plus
  • Experience with site reliability engineering practices and principles, such as error budgets and service level objectives (SLOs)
  • Knowledge of data analy

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Providence

View company profile →