SRE Insights Senior Engineer
LPL FinancialAbout the role
Are you a team player? Are you curious to learn? Are you interested in working in meaningful projects? Do you want to work with cutting-edge technology? Are you interested in being part of a team that is working to transform and do things differently? If so, LPL Financial is the place for you!
Ā
Job Overview:
We are seeking a dynamic, motivated and experienced Engineer for our Site Reliability Engineering Insights Team. In this role, you will leverage observability platforms and application logs for proactive insights and complex root cause analysis. This role is pivotal to improve observability, derive actionable insights, integrate monitoring practices into infrastructure and application pipelines using code-driven automation, ensuring system reliability, scalability, and performance. You will collaborate with cross-functional teams to identify problem patterns, hotspots and provide constant feedback to improve monitoring and stability.
This is an exciting opportunity to drive meaningful change and enhancing the advisor experience. If you are passionate about SRE and Observability and have a track record of success, we invite you to apply and be part of our journey toward greater resilience and efficiency.
Responsibilities:
Incident Management and Root Cause Analysis: Perform triage, detection, debugging, and resolution of complex multi stack production incidents. Respond rapidly and effectively to minimize the impact on advisor workflows while maintaining high service delivery standards.
Performance Optimization: Proactively identify opportunities to optimize application performance, database queries, and infrastructure utilization. Provide regular health and performance reports using analytics and reporting tools
Monitoring and Observability: Provide feedback to improve E2E Full-Stack Monitoring, Real User Monitoring (RUM) and Synthetic Monitoring across applications, services, and infrastructure.
Analytics: Review Incident, Problem, Error Data, DB Performance Stats and identify hotspots and proactive opportunities for performance and non-functional improvements
Training and Development: Mentor and develop other team members, providing training on observability tools and processes. Stay current with industry best practices and technologies, fostering a culture of continuous learning and professional growth.
Incident Management and Root Cause Analysis: Perform triage, detection, debugging, and resolution of complex multi stack production incidents. Respond rapidly and effectively to minimize the impact on advisor workflows while maintaining high service delivery standards.
Performance Optimization: Proactively identify opportunities to optimize application performance, database queries, and infrastructure utilization. Provide regular system health and performance reports using analytics and reporting tools
Monitoring and Observability: Provide feedback to improve E2E Full-Stack Monitoring, Real User Monitoring (RUM) and Synthetic Monitoring across applications, services, and infrastructure.
Analytics: Review Incident, Problem, Error Data, DB Performance Stats and identify hotspots and proactive opportunities for performance and non-functional improvements
Training and Development: Mentor and develop other team members, providing training on observability tools and processes. Stay current with industry best practices and technologies, fostering a culture of continuous learning and professional growth.
What are we looking for?
We want strong collaborators who can deliver a world-class client experience. We are looking for people who thrive in a fast-paced environment, are client-focused, team oriented, and are able to execute in a way that encourages creativity and continuous improvement.
Requirements:
Experience: 8+ years in SRE, DevOps, or related fields
Expertise in .NET stack
Observability Tools: Hands-on experience with Dynatrace, ELK or equivalent logging platform, Davis AI, and its APIs. Familiarity and hands on experience with complementary tools such as Prometheus, Grafana
Analytical Abilities: Able to analyze complex issues, data patterns, find hotspots, derive actionable insights
Cloud Platforms: Strong knowledge of AWS, Azure, or Google Cloud, and integrating Dynatrace with their services.
Preferences:
Containerization: Familiarity with containerized environments (e.g., Docker, Kubernetes)
Proven experience with Agile development processes.
Strong understanding of incident management frameworks (e.g., ITIL)
#LI-Hybrid
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights ā in under 60 seconds.
Apply Now āGenerate Application KitFree account required ā sign up in 30s