Oracle Cloud Infrastructure is seeking Senior Infrastructure Engineers to design, build, and operate the high-performance virtual networking dataplane that underpins OCI gateway services. In this role, you will develop major C and C++ features for packet processing, routing, network address translation, gateway functionality, and security-policy enforcement. You will optimize CPU and memory efficiency, concurrency, latency, and throughput for production systems operating at cloud scale.
You will partner with networking, control-plane, infrastructure, security, and operations teams to build fault-tolerant distributed systems with strong tenant isolation and safe deployment practices. The role includes end-to-end feature ownership: architecture, implementation, testing, staged rollout, observability, on-call support, incident response, and continuous operational improvement. Experience with fast dataplane technologies such as DPDK, VirtIO, eBPF, XDP, SR-IOV, SmartNICs, or userspace packet-processing frameworks is especially valuable.
Internal Responsibilities
Key Responsibilities:
Dataplane Scalability and Performance
- Implement and enhance C and C++ components for OCI’s virtual networking dataplane, including routing, network address translation, network gateways, and security-policy enforcement.
- Build packet-processing features that support scalable, multi-tenant cloud networking with guidance from senior engineers.
- Profile and optimize CPU utilization, memory usage, concurrency, latency, and throughput in large-scale production environments.
- Contribute to fast dataplane development using technologies such as DPDK, VirtIO, Linux networking, SR-IOV, SmartNICs, or userspace packet-processing frameworks.
- Develop and maintain unit, integration, system, performance, and scale tests.
Participate in design and code reviews, applying established patterns for performance, correctness, and maintainability.
Reliability and Availability
- Implement fault-tolerant networking components that handle failures, maintenance events, and software updates with minimal customer impact.
- Contribute to resilient designs using health checks, retries, timeouts, graceful draining, failover, and recovery mechanisms.
- Develop fault-injection, brownout, and recovery tests to validate service behavior.
Support staged deployments, automated validation, monitoring, rollback, and recovery procedures.
Observability and Operational Excellence
- Build and maintain service metrics, alarms, dashboards, telemetry, canaries, and operational health checks.
- Diagnose Linux, networking, performance, and production issues using logs, metrics, packet captures, traces, profilers, and debugging tools.
- Create and improve runbooks, diagnostic tools, and automation for common operational issues.
Participate in the team’s on-call rotation, incident response, root-cause analysis, and corrective-action follow-through.
Security and Multi-Tenant Isolation
Automation and Change Management
- Maintain automation and infrastructure-as-code used to build, test, deploy, and operate the service.
- Implement automated checks and safeguards for configuration changes, software releases, and rollback procedures.
- Own assigned features through implementation, testing, deployment, production validation, and support.
Core Responsibilities:
Planning and Execution
- Plan and deliver assigned work with limited supervision, identifying dependencies and risks early.
- Manage priorities and adjust execution as requirements, timelines, or production needs change.
- Communicate progress, blockers, and technical decisions clearly.
Collaboration and Partnership
- Work closely with networking, control-plane, infrastructure, security, and operations teams.
- Participate in design reviews, code reviews, incident reviews, and knowledge-sharing activities.
- Document technical decisions and contribute to team standards and best practices.
Problem Solving
- Independently troubleshoot standard and moderately complex issues using established engineering practices.
- Analyze data from logs, metrics, traces, packet captures, and tests to identify root causes.
- Escalate complex cross-service issues appropriately while contributing to their resolution.
Continuous Learning and Improvement
- Build expertise in C/C++, systems programming, cloud networking, Linux, and high-performance dataplane technologies.
- Seek feedback and apply it to improve technical and operational skills.
- Identify practical opportunities to improve testing, tooling, deployment, observability, and team workflows.
External Responsibilities
Key Responsibilities:
Dataplane Scalability and Performance
- Implement and enhance C and C++ components for OCI’s virtual networking dataplane, including routing, network address translation, network gateways, and security-policy enforcement.
- Build packet-processing features that support scalable, multi-tenant cloud networking with guidance from senior engineers.
- Profile and optimize CPU utilization, memory usage, concurrency, latency, and throughput in large-scale production environments.
- Contribute to fast dataplane development using technologies such as DPDK, VirtIO, Linux networking, SR-IOV, SmartNICs, or userspace packet-processing frameworks.
- Develop and maintain unit, integration, system, performance, and scale tests.
Participate in design and code reviews, applying established patterns for performance, correctness, and maintainability.
Reliability and Availability
- Implement fault-tolerant networking components that handle failures, maintenance events, and software updates with minimal customer impact.
- Contribute to resilient designs using health checks, retries, timeouts, graceful draining, failover, and recovery mechanisms.
- Develop fault-injection, brownout, and recovery tests to validate service behavior.
Support staged deployments, automated validation, monitoring, rollback, and recovery procedures.
Observability and Operational Excellence
- Build and maintain service metrics, alarms, dashboards, telemetry, canaries, and operational health checks.
- Diagnose Linux, networking, performance, and production issues using logs, metrics, packet captures, traces, profilers, and debugging tools.
- Create and improve runbooks, diagnostic tools, and automation for common operational issues.
Participate in the team’s on-call rotation, incident response, root-cause analysis, and corrective-action follow-through.
Security and Multi-Tenant Isolation
Automation and Change Management
- Maintain automation and infrastructure-as-code used to build, test, deploy, and operate the service.
- Implement automated checks and safeguards for configuration changes, software releases, and rollback procedures.
- Own assigned features through implementation, testing, deployment, production validation, and support.
Core Responsibilities:
Planning and Execution
- Plan and deliver assigned work with limited supervision, identifying dependencies and risks early.
- Manage priorities and adjust execution as requirements, timelines, or production needs change.
- Communicate progress, blockers, and technical decisions clearly.
Collaboration and Partnership
- Work closely with networking, control-plane, infrastructure, security, and operations teams.
- Participate in design reviews, code reviews, incident reviews, and knowledge-sharing activities.
- Document technical decisions and contribute to team standards and best practices.
Problem Solving
- Independently troubleshoot standard and moderately complex issues using established engineering practices.
- Analyze data from logs, metrics, traces, packet captures, and tests to identify root causes.
- Escalate complex cross-service issues appropriately while contributing to their resolution.
Continuous Learning and Improvement
- Build expertise in C/C++, systems programming, cloud networking, Linux, and high-performance dataplane technologies.
- Seek feedback and apply it to improve technical and operational skills.
- Identify practical opportunities to improve testing, tooling, deployment, observability, and team workflows.