Posted:2 days ago|
Platform:
Hybrid
Full Time
At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for Enterprises as a diverse team of fellow creators and inventors. We act with the speed and attitude of a start-up, with the scale and customer-focus of the leading enterprise software company in the world. Compute is one of the core organisations within OCI. We are responsible for providing Compute power i.e. VMs and BMs. Cloud pretty much cannot exists without our org. The Compute org comprises of a family of critical foundational infrastructure services that drive OCIs hardware lifecycle activities.
Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the design and delivery of the mission critical stack, with focus on security, resiliency, scale, and performance. Authority for end-to-end performance and operability. Partner with development teams in defining and implementing improvements in service architecture. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the affect of product architecture decisions on distributed systems. Professional curiosity and a desire to a develop deep understanding of services and technologies.
Incident Management Support and troubleshooting of Staging/Production environments Response and Resolve incidents as per SLA's Organise, Anticipate, Plan and work as On-Call in shifts for multiple services (Open to work in shifts & shows flexibility) Maintain Service High Availability Release Management Test and Deploy solutions and automate to replace manual processes Build and maintain deployment tools/procedures Zero downtime deployments and a high availability mindset Define and build innovative solution methodologies and assets around infrastructure, cloud migration and deployment operations at scale. Work with service teams to resolve complex issues that require troubleshooting and knowledge of code. Keep documentation up to date and resolving similar tickets with lower turnaround time and within SLA Ensure production security posture Ensure monitoring is robust and effective Change Management Perform Root Cause Analysis
Oracle
Upload Resume
Drag or click to upload
Your data is secure with us, protected by advanced encryption.
Browse through a variety of job opportunities tailored to your skills and preferences. Filter by location, experience, salary, and more to find your perfect fit.
We have sent an OTP to your contact. Please enter it below to verify.
Practice Python coding challenges to boost your skills
Start Practicing Python Now
hyderabad, bengaluru
20.0 - 35.0 Lacs P.A.
noida, uttar pradesh, india
Salary: Not disclosed
30.0 - 35.0 Lacs P.A.
Salary: Not disclosed
hyderabad, telangana, india
Salary: Not disclosed
noida, uttar pradesh, india
Salary: Not disclosed
ahmedabad, gujarat, india
Salary: Not disclosed
chennai, tamil nadu, india
Salary: Not disclosed
trivandrum, kerala, india
Salary: Not disclosed
45.0 - 50.0 Lacs P.A.