Join Oracle Cloud Infrastructure (OCI) as a Senior Core Infrastructure Engineer and play a pivotal role in shaping the future of cloud computing. In this role, you will lead the design, development, and operation of compute operability solutions, ensuring the reliability, scalability, and performance of OCI’s compute infrastructure. You’ll work with a team of innovative engineers to build and operate massive-scale, integrated cloud services that power businesses and organizations worldwide.
As a software engineer on the OCI Compute team, you will focus on enhancing the operability of our compute services, driving automation, and optimizing system reliability. Your work will directly impact the performance of mission-critical workloads for Oracle’s global customers, solving complex challenges in distributed systems, high-availability computing, and operational excellence.
Major focus areas of software maintained by the team include:
-
Designing and implementing highly scalable systems capable of functioning across numerous regions in a worldwide cloud footprint.
-
Building systems for orchestrating large-scale fleet management actions
-
Working on a strategic system to manage power in our new and upcoming AI datacenters
-
Contributing to our Compute usage systems which are critical to OCI billing
The successful hire will have passion for automation, preventing customer impact, operational excellence, and solving large-scale, complex problems.
Note that this role is onsite in our downtown Nashville, TN offices and Oracle is 5 days in the office.
Internal Responsibilities
• Design and implement scalable, reliable, and high-performance compute operability solutions for OCI.
• Develop tools, frameworks, and automation to enhance the operational efficiency of compute infrastructure.
• Troubleshoot and resolve technical issues in large-scale distributed systems.
• Drive the adoption of best practices for system reliability, performance tuning, and operational excellence.
• Mentor junior engineers
• Innovate to improve system availability, reduce latency, and optimize resource utilization.
• Participate in on-call rotations to ensure 24/7 service reliability.
• Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
• 3+ years of experience in software engineering
• Expertise in compute operability, including virtualization, containerization, or orchestration technologies (e.g., KVM, Docker, Kubernetes).
• Strong programming skills in languages such as Go, Python, Java, or C++.
• Strong data analysis experience and proficiency in SQL
• Experience with large-scale system design, automation, and operational tools (e.g., Grafana, Terraform, Prometheus).
• Familiarity with cloud computing concepts, including IaaS, PaaS, or serverless architectures.
• Excellent problem-solving skills and a track record of resolving technical challenges.
• Strong communication and collaboration skills to work effectively in a globally distributed team.
• Experience with OCI, AWS, Azure, or Google Cloud Platform.
• Contributions to open-source projects or a strong portfolio of technical innovation.