In this high-impact role, you will build distributed network observability software that transforms measurements and telemetry into trusted, actionable network intelligence. You will work with active probes, topology and routing information, flow signals, and switch telemetry to understand network behavior and identify where performance problems originate.
The vision for Next Gen Network Observability extends beyond helping network operators diagnose problems. The platform will provide the intelligence layer for self-healing networks, enabling automated control systems to steer traffic, mitigate congestion, isolate failures, and restore healthy network behavior. Achieving this vision requires accurate diagnoses, confidence-aware decisions, explainable actions, and safe closed-loop integration between observability and network control.
This position requires strong systems knowledge and hands-on experience with networking, algorithms, telemetry, and distributed computing. You will contribute to topology-aware analytics, congestion detection, queue and path reconstruction, fault localization, and production services operating at cloud scale.
Internal Responsibilities
Key Responsibilities
- Design and develop new service features, driving functionality from
concept to production. - Write high-quality, production-grade code, delivering reliable and
scalable solutions. - Own service operations and support, including troubleshooting,
incident response, and customer issue resolution. - Collaborate closely with cloud platform teams, including compute,
networking, search, and storage. - Contribute to a fast-growing, early-stage team, helping shape direction
and execute on ambitious new initiatives. - Develop automated tests and validation frameworks for detection accuracy, localization quality, false-positive rates, control-action safety, and system performance.
- Write robust, maintainable, and well-tested production code in languages such as Java, Go, C++, Rust, or Python.
- Use AI-assisted software development tools and practices across design, implementation, testing, debugging, code review, documentation, and operational analysis.
- Investigate and resolve complex issues involving distributed services, telemetry pipelines, real-time network state, and production network behavior.
- Participate in architecture, and reviews, helping maintain high standards for correctness, scalability, reliability, and operational readiness.
- Partner with network controller, SRE, hardware, network operations, security, AI/ML, and product teams to deliver closed-loop capabilities end to end.
Preferred Qualifications
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent professional experience.
- 3-5+ years of experience building distributed systems, network software, network measurement platforms, or telemetry and observability systems.
- Strong knowledge of networking fundamentals, including L2/L3 forwarding, routing, switching, BGP, ECMP, data center fabrics, queues and buffers, congestion, and packet loss.
- Advanced programming experience in at least one language such as Java, Go, C++, Rust, or Python.
- Strong understanding of algorithms, data structures, distributed systems concepts, system reliability, and performance engineering.
- Experience with one or more of the following: active network measurement, streaming telemetry, time-series or stream processing, graph algorithms, statistical inference, or multi-source data correlation.
- Familiarity with ML pipelines and models, including data preparation, model evaluation, deployment, versioning, monitoring, and integration into production services.
- Experience using AI-assisted development tools and workflows across the software development lifecycle.
- Experience building systems that process high-volume, time-sensitive data in a distributed environment.
- Demonstrated experience delivering resilient and performant software in production, cloud, or large enterprise environments.
- Strong ownership, problem-solving, communication, and cross-functional collaboration skills.
- Experience contributing to technical reviews, collaborating with architects and other senior engineers, and delivering complex projects.
Join the OCI Networking – Next Gen Network Observability team to build the intelligence layer for self-healing networks—transforming trusted network measurements and diagnoses into safe, automated control actions that improve availability, resilience, and customer experience—helping Oracle’s cloud networks operate with greater visibility, resilience, and efficiency.
External Responsibilities
Key Responsibilities
- Design and develop new service features, driving functionality from
concept to production. - Write high-quality, production-grade code, delivering reliable and
scalable solutions. - Own service operations and support, including troubleshooting,
incident response, and customer issue resolution. - Collaborate closely with cloud platform teams, including compute,
networking, search, and storage. - Contribute to a fast-growing, early-stage team, helping shape direction
and execute on ambitious new initiatives. - Develop automated tests and validation frameworks for detection accuracy, localization quality, false-positive rates, control-action safety, and system performance.
- Write robust, maintainable, and well-tested production code in languages such as Java, Go, C++, Rust, or Python.
- Use AI-assisted software development tools and practices across design, implementation, testing, debugging, code review, documentation, and operational analysis.
- Investigate and resolve complex issues involving distributed services, telemetry pipelines, real-time network state, and production network behavior.
- Participate in architecture, and reviews, helping maintain high standards for correctness, scalability, reliability, and operational readiness.
- Partner with network controller, SRE, hardware, network operations, security, AI/ML, and product teams to deliver closed-loop capabilities end to end.
Preferred Qualifications
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent professional experience.
- 3-5+ years of experience building distributed systems, network software, network measurement platforms, or telemetry and observability systems.
- Strong knowledge of networking fundamentals, including L2/L3 forwarding, routing, switching, BGP, ECMP, data center fabrics, queues and buffers, congestion, and packet loss.
- Advanced programming experience in at least one language such as Java, Go, C++, Rust, or Python.
- Strong understanding of algorithms, data structures, distributed systems concepts, system reliability, and performance engineering.
- Experience with one or more of the following: active network measurement, streaming telemetry, time-series or stream processing, graph algorithms, statistical inference, or multi-source data correlation.
- Familiarity with ML pipelines and models, including data preparation, model evaluation, deployment, versioning, monitoring, and integration into production services.
- Experience using AI-assisted development tools and workflows across the software development lifecycle.
- Experience building systems that process high-volume, time-sensitive data in a distributed environment.
- Demonstrated experience delivering resilient and performant software in production, cloud, or large enterprise environments.
- Strong ownership, problem-solving, communication, and cross-functional collaboration skills.
- Experience contributing to technical reviews, collaborating with architects and other senior engineers, and delivering complex projects.
Join the OCI Networking – Next Gen Network Observability team to build the intelligence layer for self-healing networks—transforming trusted network measurements and diagnoses into safe, automated control actions that improve availability, resilience, and customer experience—helping Oracle’s cloud networks operate with greater visibility, resilience, and efficiency.