Jobs and Careers
IN
Senior Site Reliability Engineer(Mandarin preferred)
IntelliPro Group Inc.San Diego, United Statesfull_timeVerifiedPosted 2 Jul 2025
About the role
Job Responsibilities
• Work closely with cross-functional teams to ensure the company has the right set of tools to
generate, collect, analyze, visualize and alert on operational data.
• Participate in an on-call rotation to ensure 24/7/365 availability of company's production
system. Own & operate critical open-source services like Elasticsearch,
• Kafka, RabbitMQ, Redis. Build tools and design processes that help improve observability and
system resiliency of the platform.
• Triage Site Availability Incidents and proactively work towards reducing MTTR for customer
impacting incidents.
• Partner with Service owners to implement Service Level Metrics & Service Level Objectives.
Establish design patterns for monitoring, benchmarking and deploying new features for the
backend services.
• Develop and maintain technical documentation, network diagrams, runbooks, and procedures.
Increase efficiency, respond to production incidents and prevent repeatable issues, improve the
reliability and performance of the infrastructure.
Job Requirements
• Bachelor’s degree or a foreign equivalent in Computer Science, Information Systems or a related
field, plus 3 years of experience in the job offered or as a Computer Systems Engineer, Software
Engineer or related job titles.
• Applicable experience must include at least 3 years of experience with: (1) supporting missioncritical, real-time, high-traffic applications in cloud environments; (2) Knowledge of Cloud
systems, continuous integration/build systems, Java, SQL and NoSQL databases; (3) observability
tools such as Grafana, Prometheus, Zabbix; (4) scripting/programming languages (Python or
GoLang); (5) one or more OSS technologies (Elasticsearch, Kafka or Redis); (6) container
technology like Docker, Kubernetes, Mesos.
Total Salary:200-270k
• Work closely with cross-functional teams to ensure the company has the right set of tools to
generate, collect, analyze, visualize and alert on operational data.
• Participate in an on-call rotation to ensure 24/7/365 availability of company's production
system. Own & operate critical open-source services like Elasticsearch,
• Kafka, RabbitMQ, Redis. Build tools and design processes that help improve observability and
system resiliency of the platform.
• Triage Site Availability Incidents and proactively work towards reducing MTTR for customer
impacting incidents.
• Partner with Service owners to implement Service Level Metrics & Service Level Objectives.
Establish design patterns for monitoring, benchmarking and deploying new features for the
backend services.
• Develop and maintain technical documentation, network diagrams, runbooks, and procedures.
Increase efficiency, respond to production incidents and prevent repeatable issues, improve the
reliability and performance of the infrastructure.
Job Requirements
• Bachelor’s degree or a foreign equivalent in Computer Science, Information Systems or a related
field, plus 3 years of experience in the job offered or as a Computer Systems Engineer, Software
Engineer or related job titles.
• Applicable experience must include at least 3 years of experience with: (1) supporting missioncritical, real-time, high-traffic applications in cloud environments; (2) Knowledge of Cloud
systems, continuous integration/build systems, Java, SQL and NoSQL databases; (3) observability
tools such as Grafana, Prometheus, Zabbix; (4) scripting/programming languages (Python or
GoLang); (5) one or more OSS technologies (Elasticsearch, Kafka or Redis); (6) container
technology like Docker, Kubernetes, Mesos.
Total Salary:200-270k
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s