Senior Software Engineer
Red HatAbout the role
Architect and implement observability solutions for Red Hat's product release pipelines running on OpenShift, leveraging Prometheus, SignalFx, Grafana, and Honeycomb to provide real-time monitoring, logging, and distributed tracing.
*Telecommuting role to be performed anywhere in the U.S.
What You Will Do:
Optimize Resource Requests/Limits, Horizontal Pod Autoscaling (HPA), and Vertical Pod Autoscaling (VPA) to ensure efficient utilization of OpenShift workloads.
Develop energy consumption monitoring and reporting solutions using Red Hat’s Power Monitoring Operator and Kepler to track carbon footprint and sustainability metrics.
Integrate FinOps reporting by leveraging Red Hat Cost Management Operator, providing insights into cloud expenditures and optimizing resource allocations including the design and implementation of a tagging strategy to enable granularity on the cost consumption on the product / release level.
Define and implement SLIs, SLOs, and KPIs to measure and enhance service reliability and pipeline availability.
Write infrastructure code in Groovy, Python, and Go to provide, configure, and manage the lifecycle for developer infrastructure using Ansible.
Develop internal tools to automate the creation of dashboards, SLI charts, alerts, and integrate with other monitoring tools.
Optimize alerting and incident management workflows using PagerDuty, ensuring proactive response strategies based on SLIs, SLOs, and KPIs for automated pipelines.
Implement metrics, logging, and distributed tracing using industry standard tools like Prometheus, Grafana, OpenTelemetry, Jaeger, and SignalFx.
Monitor and optimize Tekton-based CI/CD pipelines, ensuring visibility into build, test, and deployment performance.
Enhance logging and tracing within Tekton tasks and pipelines, leveraging OpenTelemetry for better debugging and troubleshooting.
Lead and mentor a team of engineers, fostering a culture of technical excellence in observability, cloud cost management, and automation, and supporting their professional growth and career progress by identifying and addressing their skill gaps and providing them with assignments that represent growth opportunities.
Advise on security best practices, performance tuning, and compliance strategies for observability and monitoring solutions.
Ensure reliability in Red Hat’s product build and release pipeline by working with Brew, Pungi, rhpkg, Errata. Collaborate with SRE, DevOps, FinOps, and security teams to improve observability and cost-efficiency across OpenShift-based deployments.
Provide E2E project management for projects with impact on software engineering that require the coordination of cross-organizational efforts, e.g. data center migrations.
Participate in the talent acquisition process by planning and conducting coding challenges and technical interviews with candidates.
Optimize deployment strategies using ArgoCD ApplicationSets ensuring efficient multi-environment management.
What You Will Bring:
Master’s degree (U.S. or foreign equivalent) in Computer Science or related field and eight (8) years of experience in the job offered or related role.
Must have five (5) years of experience with: Linux operating system; writing infrastructure code in Groovy, Python, and Go to provide, configure, and manage the lifecycle for developer infrastructure using Ansible; and Containers, GitLab, Jenkins, and JSON.
Must have five (5) years of coding experience, including code reviews.
Must have four (4) years of experience working with Red Hat’s product build, compose, and release tools including BREW, PUNGI, RHPKG, and ERRATA.
Must have four (4) years of experience with: observability, monitoring, and cloud service optimization, including experience with implementing solutions using Prometheus, SignalFX, Grafana and Honeycomb, to provide real-time monitoring, logging and distributed tracing; OpenShift, Kubernetes, and containerized environments; CI/CD automation using Tekton, GitLab CI, Jenkins, and Ansible; debugging issues across multiple software layers; working with multi-tenant OpenShift clusters in enterprise environments; optimizing Resource Requests/Limits, using Horizontal Pod Autoscaling (HPA), and Vertical Pod Autoscaling (VPA) to ensure efficient utilization of OpenShift workloads; defining, developing, and implementing SLIs, SLOs, and KPIs to measure and enhance service reliability and pipeline availability; developing tools to automate the creation of dashboards, SLI charts, alerts, and integrate with other monitoring tools; monitoring and optimizing Tekton-based CI/CD pipelines, ensuring visibility into b
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s