Jobs and Careers
LU
Staff Engineer, IT Infrastructure Observability
Lucid MotorsUnited Statesfull_timeVerifiedPosted 9 Mar 2024
About the role
Leading the future in luxury electric and mobility
At Lucid, we set out to introduce the most captivating, luxury electric vehicles that elevate the human experience and transcend the perceived limitations of space, performance, and intelligence. Vehicles that are intuitive, liberating, and designed for the future of mobility.
We plan to lead in this new era of luxury electric by returning to the fundamentals of great design – where every decision we make is in service of the individual and environment. Because when you are no longer bound by convention, you are free to define your own experience.
Come work alongside some of the most accomplished minds in the industry. Beyond providing competitive salaries, we’re providing a community for innovators who want to make an immediate and significant impact. If you are driven to create a better, more sustainable future, then this is the right place for you.Staff Engineer, IT Infrastructure Observability
We are currently seeking a Staff Engineer, IT Infrastructure Observability for our Newark location. This position requires an experienced technical expert in observability in network, physical systems, virtual systems, backup and storage systems.
Our ideal candidate exhibits a can-do attitude and approaches his or her work with vigor and determination. Candidates will be expected to demonstrate excellence in their respective fields, to possess the ability to learn quickly and to strive for perfection within a fast-paced environment.
Role:
Minimum Qualifications:
Preferred Qualifications:
We are currently seeking a Staff Engineer, IT Infrastructure Observability for our Newark location. This position requires an experienced technical expert in observability in network, physical systems, virtual systems, backup and storage systems.
Our ideal candidate exhibits a can-do attitude and approaches his or her work with vigor and determination. Candidates will be expected to demonstrate excellence in their respective fields, to possess the ability to learn quickly and to strive for perfection within a fast-paced environment.
Role:
- Highly skilled in operational efficiency, optimal utilization, and system resiliency for a real-time streaming analytics platform
- Utilizing knowledge and experience in monitoring systems and applications, conduct initial troubleshooting and root cause analysis
- Supporting the integration of new technologies with monitoring systems to include deployment and decommissioning of monitoring agents
- Implement service level metrics and service level objectives that act as service-level health indicators
- Producing metrics and capacity planning reports as needed to support the monitored environments
- Expert in collecting metrics for performance related monitoring
- Design processes that help improve observability and system resiliency
- Work with other system administrators, engineers, and vendors to resolve hardware and software issues
- Maintain monitoring documentation, diagrams and standard operating procedures as required
- Coordinate and participate in key process improvements as they relate to operations monitoring
- Experience defining, creating, and supporting monitoring dashboards
- Identify potential trends in performance gaps and recommend modifications to the standardized work documents by using process improvement principles
- Triage site availability incidents and proactively work towards reducing MTTR for customer-impacting incidents
- Coordinate with Project Management as a subject matter expert (SME) on various projects, for infrastructure and IT operations. Automate as required
- Capable of working and collaborating on multiple projects or tasks with high attention to detail
- Keep abreast of emerging technologies to identify, research, evaluate and present concepts and solutions to management for implementation considerations
- Advocate standards, best practices, policies, and procedures
Minimum Qualifications:
- Bachelor’s degree in Information Technology, Computer Science, or related field
- 8+ years of Site Reliability Engineer or production Engineer
- 5+ years of Windows, Linux server environment
- 2+ years of experience working with network
- 5+ years of coding experience in Python, Perl, Bash
- Experience with Systems Observability & Monitoring experience
- Experience with Influx DB Grafana, Prometheus, Splunk, SolarWinds ELK and APM
Preferred Qualifications:
- Knowledge of orchestration engines and package management including Kubernetes and Helm
- Good understanding of modern application, version control and development flow
- Customer service oriented with strong interpersonal and leadership skills
- Knowledge of containers and cloud platforms (AWS, Azure and/or GCP)
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s