Principal Site Reliability Engineer - Exadata Cloud Service
OracleAbout the role
Job Description
Are you interested in solving the complex challenges involved in building and operating large-scale, distributed cloud infrastructure?
Oracle Cloud Infrastructure is developing the next generation of cloud technologies operating in highly available, scalable, secure, distributed, and multi-tenant environments. Our mission is to provide customers with an enterprise-grade cloud infrastructure platform that delivers exceptional reliability, scalability, security, and performance for mission-critical databases, applications, and workloads.
Learn more about Oracle Cloud Infrastructure at:
https://cloud.oracle.com/cloud-infrastructure
About the Role
Oracle is building and expanding its next-generation Platform as a Service cloud offering and the support experience required to operate them at global scale. As our cloud services grow, we are expanding our team of highly skilled, customer-focused Site Reliability Engineers.
As a Principal Site Reliability Engineer, you will help operate, support, and improve Oracle Exadata Cloud Service within Oracle Cloud Infrastructure. You will work at the intersection of software engineering, database engineering, systems administration, cloud infrastructure, automation, artificial intelligence, and production operations. Oracle Exadata is a full-stack engineered system designed to improve the performance, scalability, security, and availability of Oracle Database workloads. It includes specialized capabilities engineered with Oracle Database to accelerate online transaction processing, analytics, consolidation, and machine learning workloads.
In this role, you will support highly complex Exadata environments, lead the resolution of critical production issues, develop software and automation, improve service observability, and influence the architecture and operational readiness of new service capabilities. You will also serve as a technical leader within the organization. You will mentor engineers, establish engineering standards, lead cross-functional initiatives, and represent the voice of the customer to product development, service engineering, and architecture teams.
This position is integral to the reliability of the Exadata Cloud Service platform and to the success of Oracle’s customer relationships.
Responsibilities
- Design, develop, test, and deliver software and automation that improve the availability, scalability, latency, security, operability, and efficiency of Oracle Database as a Service offering.
- Lead the investigation and resolution of complex technical issues spanning Exadata Cloud Service, Autonomous Database, Oracle Database, operating systems, virtualization, storage, networking, and cloud infrastructure.
- Coordinate response to high-severity incidents, including technical diagnosis, mitigation, stakeholder communication, recovery, and post-incident review.
- Perform detailed root-cause analysis and develop corrective and preventive solutions that reduce the likelihood and impact of recurrence.
- Build automation to eliminate repetitive operational work, reduce human error, accelerate incident response, and improve fleet-management efficiency.
- Apply AI-assisted engineering and operations techniques to improve anomaly detection, incident correlation, troubleshooting, knowledge retrieval, capacity forecasting, and operational decision-making.
- Evaluate and integrate generative AI, machine learning, and large language model capabilities into appropriate SRE workflows while maintaining security, privacy, accuracy, and human oversight.
- Develop tools that use operational telemetry, logs, metrics, traces, events, and historical incident data to identify patterns and provide actionable insights.
- Define, implement, and continuously improve service-level indicators, service-level objectives, error budgets, alerts, dashboards, and operational health metrics.
- Improve monitoring and observability across distributed database and infrastructure services.
- Participate in the architecture, design, implementation, and operational-readiness review of large-scale distributed DBaaS features.
- Conduct research, prototyping, and proof-of-concept development for new service capabilities, reliability improvements, automation frameworks, and AI-enabled operational tools.
- Act as a trusted technical advisor to customers and internal stakeholders, helping solve complex database, infrastructure, cloud, and DevOps challenges.
- Create and deliver best-practice recommendations, sample code, runbooks, troubleshooting guides, technical documentation, and operational procedures.
- Contribute to making Oracle
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s