Staff Engineer II
Western Alliance BankAbout the role
Job Title:
Staff Engineer IILocation:
Block 23What you'll do:
The Staff Engineer II – Monitoring & Performance Engineering is responsible for the implementation, configuration, administration, and support of enterprise monitoring, observability, and performance management platforms. This role supports critical applications, infrastructure, cloud services, and network technologies to ensure system reliability, availability, and operational excellence.The position requires hands-on expertise with Datadog, Tripwire, Application Performance Monitoring (APM), and infrastructure troubleshooting. The engineer partners with Infrastructure, Network, Security, Cloud, Architecture, Application Development, and Operations teams to implement monitoring solutions, resolve complex technical issues, and improve service reliability across the enterprise.
Design, implement, configure, administer, and support enterprise monitoring and observability platforms, including Datadog and Tripwire.
Configure and maintain APM, infrastructure monitoring, log management, distributed tracing, alerting, dashboards, and reporting capabilities.
Onboard applications, infrastructure, cloud services, middleware platforms, and network technologies into enterprise monitoring solutions.
Develop and maintain monitoring standards, alerting strategies, dashboards, and operational reporting.
Troubleshoot complex application, infrastructure, cloud, and network issues and perform root cause analysis.
Collaborate across Infrastructure, Network, Security, Cloud, Architecture, Application Development, and Operations teams to resolve cross-functional technical issues.
Create and maintain technical documentation, operational procedures, and support runbooks.
Evaluate and recommend improvements to monitoring, observability, and performance management capabilities.
Participate in technology modernization and continuous improvement initiatives.
Participate in a 24x7 on-call rotation supporting enterprise monitoring platforms and production services.
Respond to production incidents, support major incident management activities, and participate in after-hours maintenance and operational support as required.
What you'll need:
Bachelor’s degree in computer science, Engineering, Information Technology, or equivalent experience.
7+ years of experience in Systems Engineering, Infrastructure Engineering, Monitoring, Observability, Performance Engineering, or related technical disciplines.
Hands-on experience designing, implementing, configuring, administering, and supporting enterprise monitoring and observability platforms, including Datadog and Tripwire.
Experience onboarding applications, infrastructure, cloud services, middleware platforms, and network technologies into monitoring solutions. Experience supporting: Application Performance Monitoring (APM), Infrastructure Monitoring, Distributed Tracing, Log Management, Alerting and Notifications, Dashboards and Operational Reporting
Strong troubleshooting and root cause analysis skills across infrastructure, applications, middleware, cloud platforms, and networks.
Strong understanding of TCP/IP networking fundamentals, including DNS, HTTP/HTTPS, TLS/SSL, routing, switching, network connectivity, and latency analysis.
Experience supporting enterprise productio
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s