We are looking for seasoned and talented developers to join the Observability & Supervision team within Oracle Cloud Infrastructure (OCI). Our team is part of Autonomous OCI, a new service platform powering OCI services that encompasses low-level execution runtime, application management and high-level workflows for change management. Our mission is to help OCI developers concentrate on building their services while we ensure they run efficiently and reliably at cloud scale.
Observability & Supervision is how AOCI sees and heals itself. We build the systems that collect and query metrics and logs across the fleet, evaluate the health of every workload in real time, and decide what to do when something goes wrong. Our goal is an autonomous platform: one that detects, diagnoses and recovers from failures on its own, scales itself ahead of demand, and escalates to humans only when it truly needs to. This means building low-latency, highly available distributed systems where milliseconds and correctness matter, and where the platform is only as reliable as the component watching over it.
These are exciting times, and our team is still new and growing. This is your opportunity to build ambitious new initiatives with broad impact across OCI. We need you to challenge existing engineering assumptions and boundaries, bring your expertise in highly performant, reliable software engineering and help us bring OCI to the next level. We are looking for engineers who are self-motivated, passionate about solving complex software challenges, and able to dive deep into systems to understand and improve them. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.
What You Will Do
• Design and build the supervision control loop of the Kilt platform: health evaluation, automated recovery, escalation policies and self-healing workflows.
• Build scalable telemetry pipelines for metrics and logs, including ingestion, storage, querying and alerting, operating at OCI fleet scale.
• Develop platform triggers that drive autonomous actions such as recovery and autoscaling from live platform state, with tight end-to-end latency budgets.
• Write performance-critical, memory-safe services in Rust alongside Java and Go components.
• Engineer for high availability: replication, partitioning, failover and graceful degradation, validated through load, performance and chaos testing.
• Partner with runtime, compute and service teams across OCI to make observability and recovery a built-in property of every service running on Kilt.
Internal Responsibilities
Basic Qualifications
• BS or MS degree in Computer Science or relevant technical field involving coding, or equivalent practical experience
• 4+ years of full-time professional experience in software development
• Demonstrated ability to write great code using Rust, Java, GoLang, or similar languages
• Proven ability to deliver products and experience with the full software development lifecycle
• Experience working on large-scale, highly distributed services infrastructure
• Experience working in an operational environment with mission-critical tier-one livesite servicing
• Systematic problem-solving approach, strong communication skills, a sense of ownership, and drive
• Experience designing architectures that demonstrate deep technical depth in one area, or span many products, to enable high availability, scalability, market-leading features and flexibility to meet future business demands
Preferred Qualifications
• Experience as technical lead on a large-scale cloud service
• Production experience with Rust, particularly for low-latency or high-throughput systems
• Experience building observability systems: metrics, logging, tracing, time-series databases, query engines or alerting platforms
• Experience building automated remediation, self-healing or autoscaling systems, and designing control loops that act safely without human intervention
• Knowledge of distributed systems fundamentals: consensus, replication, gossip protocols, partitioning and failure detection
• Hands-on experience developing and maintaining services on a public cloud platform (e.g., AWS, Azure, Oracle)
• Experience working on Kubernetes
• Knowledge of Infrastructure as Code (IaC) languages, preferably Terraform
• Strong knowledge of databases (SQL and NoSQL), including embedded storage engines
• Strong knowledge of Computer Networking (OSI layers, HTTP, DNS, TCP/IP, DHCP, Routers, Gateways, Subnets, etc.)
• Knowledge of Linux internals, Linux/Unix troubleshooting skills
• Familiarity with host virtualization technologies (KVM, Containers, Docker, etc.)
• Able to effectively communicate technical ideas verbally and in writing (technical proposals, design specs, architecture diagrams and presentations)
- Experience with hiring, mentorship and raising the talent bar
External Responsibilities
Basic Qualifications
• BS or MS degree in Computer Science or relevant technical field involving coding, or equivalent practical experience
• 4+ years of full-time professional experience in software development
• Demonstrated ability to write great code using Rust, Java, GoLang, or similar languages
• Proven ability to deliver products and experience with the full software development lifecycle
• Experience working on large-scale, highly distributed services infrastructure
• Experience working in an operational environment with mission-critical tier-one livesite servicing
• Systematic problem-solving approach, strong communication skills, a sense of ownership, and drive
• Experience designing architectures that demonstrate deep technical depth in one area, or span many products, to enable high availability, scalability, market-leading features and flexibility to meet future business demands
Preferred Qualifications
• Experience as technical lead on a large-scale cloud service
• Production experience with Rust, particularly for low-latency or high-throughput systems
• Experience building observability systems: metrics, logging, tracing, time-series databases, query engines or alerting platforms
• Experience building automated remediation, self-healing or autoscaling systems, and designing control loops that act safely without human intervention
• Knowledge of distributed systems fundamentals: consensus, replication, gossip protocols, partitioning and failure detection
• Hands-on experience developing and maintaining services on a public cloud platform (e.g., AWS, Azure, Oracle)
• Experience working on Kubernetes
• Knowledge of Infrastructure as Code (IaC) languages, preferably Terraform
• Strong knowledge of databases (SQL and NoSQL), including embedded storage engines
• Strong knowledge of Computer Networking (OSI layers, HTTP, DNS, TCP/IP, DHCP, Routers, Gateways, Subnets, etc.)
• Knowledge of Linux internals, Linux/Unix troubleshooting skills
• Familiarity with host virtualization technologies (KVM, Containers, Docker, etc.)
• Able to effectively communicate technical ideas verbally and in writing (technical proposals, design specs, architecture diagrams and presentations)
- Experience with hiring, mentorship and raising the talent bar