Job Description
Oracle Cloud Infrastructure, or OCI, is building and operating a suite of massive-scale, highly available cloud services in a broadly distributed, multi-tenant cloud environment. OCI’s mission is to provide secure, reliable, and performant cloud infrastructure for customers running some of the world’s most business-critical workloads.
The OCI Exec Service team owns a foundational execution platform used by OCI lifecycle and deployment systems to run workloads such as deployment validation, canaries, build-related executions, artifact workflows, and operational automation across commercial and government regions.
We are looking for a strong Senior software engineer who can design, build, and operate distributed systems in production. You will work on core components of Exec Service, including APIs, orchestration workflows, execution scheduling, Kubernetes-based execution infrastructure, service reliability, security hardening, observability, and automation. This role requires hands-on engineering depth, strong ownership, and the ability to work independently on major features while collaborating across OCI service teams.
Who are we looking for?
We are looking for engineers with strong distributed systems fundamentals and practical production experience. You should have experience designing major features, taking them through design review, implementation, testing, launch, and operational support. You understand how to build services that are resilient, observable, secure, and maintainable.
The ideal candidate is a strong coder and a distributed systems generalist who can dive deep into service APIs, orchestration logic, Kubernetes execution environments, persistence layers, networking behavior, and production operations. You should value simplicity, correctness, automation, and operational excellence. You should be comfortable working in ambiguous technical areas, decomposing complex problems, and delivering durable solutions.
What are the biggest challenges for the team?
Exec Service sits on the critical path for OCI lifecycle operations, region bootstrap, deployment validation, canaries, and internal execution workflows. The team is evolving the platform to support higher scale, stronger isolation, improved security posture, better self-service debugging, and broader adoption across OCI.
Key challenges include:
Internal Responsibilities
Responsibilities
As a Software Developer on the Exec Service team, you will:
Design, implement, test, and operate major features for OCI Exec Service.
Build and improve distributed service components such as frontend APIs, orchestration workflows, execution scheduling, and execution lifecycle management.
Contribute to Kubernetes-based execution infrastructure, including workload scheduling, cluster interaction, pod execution, and operational tooling.
Improve service reliability through better monitoring, alerting, dashboards, runbooks, and automated remediation.
Participate in production operations, including on-call support, incident response, root-cause analysis, and corrective action delivery.
Partner with OCI lifecycle, deployment, build, test, canary, security, and region-build teams to deliver platform capabilities.
Drive design reviews, write technical documentation, and communicate tradeoffs clearly.
Contribute to security hardening efforts related to workload isolation, credentials, network access, logs, and compliance requirements.
Identify opportunities to use AI and automation to improve operational triage, service debugging, and developer productivity.
Required Qualifications
4+ years of software engineering experience building and operating distributed services.
Strong development experience in Java, C++, C#, Go, or a similar object-oriented or systems programming language.
Strong understanding of data structures, algorithms, operating systems, networking, and distributed systems fundamentals.
Experience designing, implementing, and launching major features in production services.
Experience participating in design reviews, writing technical documentation, and collaborating with cross-functional stakeholders.
Experience operating production services, including monitoring, alerting, incident response, and root-cause analysis.
Good understanding of databases, persistence layers, NoSQL systems, or distributed storage technologies.
Ability to work independently on well-scoped projects and drive them to completion.
Strong communication skills and ability to work effectively in a collaborative engineering environment.
Preferred Qualifications
Hands-on experience developing services on a public cloud platform such as Oracle Cloud, AWS, or Azure.
Experience developing, operating, or debugging Kubernetes-based applications or clusters.
Experience with service orchestration, workflow engines, schedulers, execution platforms, or job-running systems.
Experience with high-scale, multi-tenant infrastructure services.
Experience with cloud networking, service-to-service authentication, mTLS, IAM, resource principals, or similar security models.
Familiarity with threat modeling, security reviews, CVE remediation, SAST/DAST, CIS benchmarks, or compliance-driven engineering.
Interest in applying AI to production operations, on-call triage, service debugging, or developer productivity.
External Responsibilities
Responsibilities
As a Software Developer on the Exec Service team, you will:
Design, implement, test, and operate major features for OCI Exec Service.
Build and improve distributed service components such as frontend APIs, orchestration workflows, execution scheduling, and execution lifecycle management.
Contribute to Kubernetes-based execution infrastructure, including workload scheduling, cluster interaction, pod execution, and operational tooling.
Improve service reliability through better monitoring, alerting, dashboards, runbooks, and automated remediation.
Participate in production operations, including on-call support, incident response, root-cause analysis, and corrective action delivery.
Partner with OCI lifecycle, deployment, build, test, canary, security, and region-build teams to deliver platform capabilities.
Drive design reviews, write technical documentation, and communicate tradeoffs clearly.
Contribute to security hardening efforts related to workload isolation, credentials, network access, logs, and compliance requirements.
Identify opportunities to use AI and automation to improve operational triage, service debugging, and developer productivity.
Required Qualifications
4+ years of software engineering experience building and operating distributed services.
Strong development experience in Java, C++, C#, Go, or a similar object-oriented or systems programming language.
Strong understanding of data structures, algorithms, operating systems, networking, and distributed systems fundamentals.
Experience designing, implementing, and launching major features in production services.
Experience participating in design reviews, writing technical documentation, and collaborating with cross-functional stakeholders.
Experience operating production services, including monitoring, alerting, incident response, and root-cause analysis.
Good understanding of databases, persistence layers, NoSQL systems, or distributed storage technologies.
Ability to work independently on well-scoped projects and drive them to completion.
Strong communication skills and ability to work effectively in a collaborative engineering environment.
Preferred Qualifications
Hands-on experience developing services on a public cloud platform such as Oracle Cloud, AWS, or Azure.
Experience developing, operating, or debugging Kubernetes-based applications or clusters.
Experience with service orchestration, workflow engines, schedulers, execution platforms, or job-running systems.
Experience with high-scale, multi-tenant infrastructure services.
Experience with cloud networking, service-to-service authentication, mTLS, IAM, resource principals, or similar security models.
Familiarity with threat modeling, security reviews, CVE remediation, SAST/DAST, CIS benchmarks, or compliance-driven engineering.
Interest in applying AI to production operations, on-call triage, service debugging, or developer productivity.