Jobs and Careers
RE

Senior Site Reliability Engineer - Data Platform

Red Hat
Remote, Ireland, IrelandRemotefull_timeVerifiedPosted 6 Mar 2024

About the role

<h2>About the job</h2> <p>The Data Development, Insights &amp; Strategy Team (DDIS) is a highly focused effort to lead digital-first execution and transformation at Red Hat leveraging data strategically for our customers, partners, and associates.DDIS team is seeking Site Reliability Engineers (SRE) focused on Data as-a-service Infrastructure to join our team. We are looking for strong site reliability engineers to define, lead SRE practice for the next generation SaaS applications for the Red Hat data products at Cloud scale - warehouses, analytical and machine learning workloads. As an SRE, you will contribute to development and operations of our GitOps based infrastructure as code automation platforms to manage data services environments with a deep focus on aligning the product, processes, and policies.</p> <p>In this role, you will have an opportunity to influence the complex challenges of scale &amp; security to develop, operate Red Hat data platforms. Also you will be partnering with Sales, Finance, Marketing focused domain teams across Red Hat and compliance and security stakeholders to deliver a secure SaaS platform for Red Hat and partner teams.</p> <h2>What you will do</h2> <ul> <li>Manage, deploy, and operate cloud solutions at scale using the principles of Site Reliability Engineering</li> <li>Participate in the design and development of new features to enable Data  'as-a-service'</li> <li>Design and write automation software to provision, upgrade, monitor, and heal Data  'as-a-service'</li> <li>Identify single points of failure and other high-risk architecture issues; propose and implement more resilient resolutions</li> <li>Define Service level Objectives, implement them along with runbooks</li> <li>Participate in product release cycles, deploying code to integration, staging and production environments, integrating with CI/CD tooling, monitoring and change management</li> <li>Interact with automated monitoring and healing infrastructure to ensure healthy environments</li> <li>Help and develop peers through knowledge sharing, mentoring and collaboration</li> <li>Create and maintain standard operating procedures (SOPs) for performing maintenance tasks, applying configuration changes and remediating problems in our environment</li> <li>Participate in a follow-the-sun on-call rotation</li> <li>Contribute software tests and participate in peer review to increase the quality of our codebase</li> <li>Work as part of a globally distributed team</li> </ul> <p>#LI-REMOTE</p> <p>#LI-AM4</p> <h2>What you will bring</h2> <ul> <li>Software engineering experience using object-oriented languages; Golang/Python are preferred</li> <li>Experience developing, managing Infrastructure as code automation platforms - Terraform is preferred</li> <li>Experience in troubleshooting as-a-service offerings (SaaS, PaaS, etc.)</li> <li>Experience with developing, deploying and managing applications on any of the public cloud services - AWS is preferred</li> <li>Experience with Kubernetes or OpenShift</li> <li>Prior experience with Snowflake, Fivetran is a plus</li> <li>Superior communications skills and experience working directly with and presenting to stakeholders</li> <li>Ability to quickly learn new technologies and follow industry trends</li> <li>Excellent communication, presentation, and writing skills</li> </ul>

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Red Hat

View company profile →