Senior Site Reliability Engineer
PlayStation GlobalAbout the role
Why PlayStation?
PlayStation isn’t just the Best Place to Play — it’s also the Best Place to Work. Today, we’re recognized as a global leader in entertainment producing The PlayStation family of products and services including PlayStation®5, PlayStation®4, PlayStation®VR, PlayStation®Plus, acclaimed PlayStation software titles from PlayStation Studios, and more.
PlayStation also strives to create an inclusive environment that empowers employees and embraces diversity. We welcome and encourage everyone who has a passion and curiosity for innovation, technology, and play to explore our open positions and join our growing global team.
The PlayStation brand falls under Sony Interactive Entertainment, a wholly-owned subsidiary of Sony Corporation.
Senior Site Reliability Engineer
Los Angeles, CA
As a member of the technical operations SRE team within the partner engineering group, you will carry the responsibility of keeping services and applications on the platform available, resilient, and secure while continually enabling our feature teams to deliver exciting features and experiences to our content creators and operators worldwide. You will drive, and lead technical initiatives, identify and contribute towards process and technology improvements amplifying experiences for content creators, operators, and players.
Responsibilities:
- Support application delivery and operations of internal and public-facing services within AWS cloud environment, ensuring availability, resiliency, scalability, and performance.
- Facilitate delivery and releases of new services and features to customers while ensuring operational readiness.
- Pursue operational improvements and toil reduction thru automation and tooling.
- Improve observability on our platform by implementing robust monitoring and alerting patterns. Develop rich, informative dashboards / reports on applications and services that provide relevant insights and meaningful alerting to reduce MTTD and MTTR.
- Collaborate with development, platform and security to inspire, implement, and deliver end-to-end system performance, resiliency and security across all services and tools within the platform.
- Evaluate hosting resource usage and spend and optimize by applying standard methodologies, patterns and technologies such as spot instances and auto-scaling.
- Participate in rotational on-call support to triage, resolve production incidents and conduct root cause analysis to identify and drive improvements.
Key Qualifications:
- Build, deploy, operate and support services at scale
- PASSIONATE(!) desire to automate and improve everything including process improvements, standardizing tools and technologies
- Excellent problem solving skills that span user experience, system, infrastructure, and networking (TCP/IP).
- Drive operational and infrastructure requirements that promote availability, reliability, performance, and security at global scale
- Customer and peer relationship focused with strong interpersonal and communication skills, inspire change across teams and mentor others.
Required Skills:
- Operating distributed, critical customer-facing services or applications in production at a global scale.
- In depth understanding of Unix/Linux systems internals and networking
- Source code management tools (GitHub preferred)
- Software development experience in one or more of following: Python, Go, Node.js, Java.
- Building and deploying Infrastructure as Code: CloudFormation/Terraform
- Building continuous integration and continuous delivery (CICD) pipelines in Jenkins or similar
- Operating and running Web Application/APIs in AWS cloud infrastructure including managed services such as Lambda, RDS, DynamoDB and Elasticache.
- AWS systems and network protocols (ie: ALB, R53, API-Gateway, TCP/IP, HTTP/HTTPS, DNS)
- Delivering production content using CDN technologies (Akamai, Cloudfront)
- Container technologies and orchestration (ie: Docker, Kubernetes, EKS)
- Application monitoring tools: DataDog, CloudWatch, Splunk, Grafana
- Data Reporting & Analytics: SQL, MySQL, Oracle.
Preferred Experience:
- BS degree in Computer Science, Software Engineering, or related technical area
- 7+ years professional experience
- 2+ years AWS Cloud - deploying, tuning and operating Web/API services at scale.
#LI-KS1
Please refer to our Candidate Privacy Notice for more information about how we process your personal information, and your da
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s