Site Reliability Engineer
ManTechAbout the role
General information
Description & Requirements
***This is for a future opportunity***
MANTECH seeks motivated, career, and customer-oriented Site Reliability Engineer (SRE) for a new initiative. This effort supports the rapid design, deployment, operation, and sustainment of enterprise-scale AI, data, and mission platform capabilities across cloud, edge, and classified operational environment
This role supports the operational reliability, scalability, monitoring, and incident response for the enterprise AI systems. You will focus on operational outcomes and optimizing system performance.
Responsibilities include but are not limited to:
Apply core reliability engineering principles to ensure high availability and stability of production systems.
Manage incident response, root cause analysis, and post-mortem processes for the AI platform.
Implement and optimize observability operations using OpenTelemetry, Prometheus, Grafana, Loki, or Tempo.
Oversee capacity planning, performance optimization, and FinOps practices.
Define and continuously monitor Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
Minimum Qualifications:
Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
5 or more years of experience in Site Reliability Engineering (SRE), DevOps, or production operations.
Extensive experience with cloud-native infrastructure, particularly Kubernetes.
Deep knowledge of monitoring, alerting, and logging systems.
Proven ability to automate operational tasks and reduce toil.
Preferred Qualifications:
Hands-on experience with the full observability stack: OpenTelemetry, Prometheus, Grafana, Loki, and Tempo.
Experience with FinOps and
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s