Senior System Engineering - Engineering Operations
AT&TAbout the role
Job Description:
This position requires office presence of a minimum of 5 days per week and is only located in the location(s) posted. No relocation is offered.
Join AT&T and reimagine the communications and technologies that connect the world. Our Consumer Technology experience team is delivering innovative and reliable technology solutions to power differentiated, simplified customer experiences. Bring your bold ideas and fearless risk-taking to redefine connectivity and transform how the world shares stories and experiences that matter. When you step into a career with AT&T, you won’t just imagine the future-you’ll create it.
As a Senior System Engineer with the Engineering Operations team, the candidate will have an SRE mindset, have a strong sense of ownership, a relentless drive and passion to detect and resolve problems. As functionality gets operationalized and adopted across the enterprise, he/she will be entrusted to help maintain a healthy production ecosystem that is responsive, user friendly and reliable to our user community.
The role will partner with various stakeholders including product development teams, business leaders, end-user communities to ensure production applications are designed for Non-Functional Requirements. This is an exciting, hands-on, technical System Engineering position responsible for Site Reliability Engineering aspects such as developing a deep functional and technical knowledge-base of the application, creation of run books, developing observability features of the application in terms of alerts, monitoring and dashboards that enable proactive incident and problem detection, triaging of the incidents and conducting blameless post-mortems (after action reviews).
Additionally, the person is responsible for Problem Management Engineering functions such as proactive problem detection, intake of problems from end users, functional triage, customer impact assessments, preliminary technical analysis, logging the issues as defects and working with product/dev/quality teams in getting resolution.
Seeking a candidate with 5+ years of supporting large scale applications in production with an Engineering approach (SRE) – such as Java EE apps, ERP or CRM apps in an operations capacity. You will be responsible for operations support aspects of mission critical applications that are used by front office and retail agents for selling, service and support functions.
Candidates possess a strong Engineering background and are well-experienced in Observability tool-sets, and leverage them effectively to create alerts, dashboards. Additionally, the candidates are experienced in constantly looking out for automation opportunities and eliminate manual work.
Key Roles and Responsibilities:
• Incident Management: Lead the response to production issues, ranging from identifying and troubleshooting problems to implementing immediate fixes. Ensure minimal downtime and adherence to service level agreements (SLAs).
• Observability: Build alerting, monitoring and dashboards that identify problems proactively;
• Problem Solving: Utilize strong analytical, technical and functional skills to diagnose and resolve complex issues within production environments with a focus on immediate impact mitigation; work with dev teams to implement long-term solutions to prevent recurrence of incidents.
• RunBooks and Documentation: Create and maintain comprehensive documentation for system architecture, configuration, deployment procedures, and troubleshooting guides
• Automation: Develop and maintain scripts and automation tools to streamline operations, deployment processes, and repetitive tasks. Focus on automating recovery processes and routine maintenance tasks to improve system reliability and efficiency
• Non Functional Requirements : Working with development teams, identify and provide the non-functional requirements and acceptance criteria during design and development, and ensure that these are met prior to moving the features to production
• Performance Optimization: Monitor application performance using APM (Application Performance Management) tools such as Dynatrace, App Dynamics and ELK. Identify bottlenecks and work with dev teams to optimize the performance of applications through code improvements, configuration tuning, and resource optimization.
• Adopt SRE best practices: Work with dev teams to define Non-Functional Requirements such as reliability, performance, scalability, application logging for observability, etc. Define SLI/SLOs, Error Budgets, Automation focus
• Conduct Blameless Postmortems / After Action Reviews: Work with dev/architect/quality engineering teams to identify and document patterns
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s