Department Description
Oracle Cloud Infrastructure (OCI) delivers cloud infrastructure services for a wide range of customer workloads. OCI Kubernetes Engine (OKE) is OCI’s managed Kubernetes service, helping customers deploy and operate containerized applications at scale.
Our team develops and operates software that supports Kubernetes worker nodes, including Linux systems integration and container networking. We build new capabilities, improve performance and reliability, and maintain the automation needed to test, release, and operate a globally available service.
Position Overview
We’re looking for hands-on engineers with expertise and passion for solving difficult problems in distributed systems, virtualized infrastructure, and highly available services. If this is you, at Oracle, you can design and build innovative new systems from the ground up.
The ideal candidate for this team is a strong software engineer with broad technical knowledge and experience, including areas such as networking, storage, internet protocols, and operating systems. We design, build, and operate distributed, highly available systems to build, update, and deploy Kubernetes, plus automation and tooling for testing, deployments, and other needs.
Required Qualifications
- Bachelor’s degree in Computer Science or a related technical field, or equivalent practical experience.
- Five or more years of software engineering experience building, delivering, and operating large-scale, highly available distributed systems.
- Experience developing and maintaining large Java or Go codebases, along with experience building automation and tooling using languages such as Python.
- Strong knowledge of data structures, algorithms, operating systems, and distributed systems.
- A systematic approach to debugging and problem-solving, with the ability to work independently and take responsibility for outcomes.
- Clear written and verbal communication skills and the ability to collaborate effectively across engineering teams.
Preferred Qualifications
- Experience developing or managing containerized workloads using Kubernetes.
- Experience with overlay networking or container networking on Linux.
- Experience with Linux system and service management.
- Experience building multi-tenant, virtualized infrastructure.
Internal Responsibilities
Own components within OKE platform services, from design and implementation through testing, deployment, and ongoing operation.
Investigate and resolve complex issues across related distributed services, delivering fixes that protect dependent services and preserve safe upgrade paths.
Analyze usage, performance, and error budgets to identify and implement targeted improvements in resilience and capacity.
Lead team-level improvements in automated testing, deployment tooling, and operational workflows to reduce manual work and improve release reliability.
Collaborate with engineering teams across OCI and with external partners to resolve dependencies and deliver integrated solutions.
Create and maintain technical documentation, examples, and adoption guidance that help other teams use and operate the components you own.
Participate in an on-call rotation to investigate service incidents, address customer escalations, and restore service.
Perform operational maintenance and follow through on improvements that prevent recurring issues.
External Responsibilities
Own components within OKE platform services, from design and implementation through testing, deployment, and ongoing operation.
Investigate and resolve complex issues across related distributed services, delivering fixes that protect dependent services and preserve safe upgrade paths.
Analyze usage, performance, and error budgets to identify and implement targeted improvements in resilience and capacity.
Lead team-level improvements in automated testing, deployment tooling, and operational workflows to reduce manual work and improve release reliability.
Collaborate with engineering teams across OCI and with external partners to resolve dependencies and deliver integrated solutions.
Create and maintain technical documentation, examples, and adoption guidance that help other teams use and operate the components you own.
Participate in an on-call rotation to investigate service incidents, address customer escalations, and restore service.
Perform operational maintenance and follow through on improvements that prevent recurring issues.