The Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for top tier enterprises around the world. Our cloud offers unmatched hyper-scale, multi-tenant services deployed in more than 40 regions worldwide.
The mission of our Network Development Engineering team is to provide services that allow our customers to drive operational excellence in OCI networks at scale. Our customers want auto-remediation of incidents, touchless and automated operations such as upgrading 1000s of devices and pushing configuration changes safely, adding network capacity seamlessly. We help Oracle build best-in-class network automation services to meet our customers' needs.
As a Sr Network Development Engineer, you will be responsible for helping design, build, test, deploy and operate highly scalable solutions to support Oracle’s next-generation Cloud Infrastructure. You should be comfortable with building within complex distributed systems involving the management of 1000s of devices. You should be a rock-solid coder with a strong understanding of networking and distributed systems. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn. You should also be able to manage multiple projects and priorities in a fast-paced environment, be results-driven, and thrive even when faced with ambiguity and uncertainty. Most importantly, you should be a team player, someone who is eager to learn, and willing to do whatever it takes to help the team succeed. You will participate in on-call rotation for engineers.
Preferred Qualifications:
- 3+ years of experience in network and software development
- Proficiency with Python or similar programming language
- experience with Linux
- Proficiency with network technologies and protocols (TCP/IP, BGP, OSPF, MPLS, IS-IS, RSVP-TE, EVPN, VxLAN etc.)
- Excellent knowledge of data structures and algorithms
- Excellent organizational, verbal, and written communication skills
Additional Qualifications:
- Knowledge of cloud computing & networking technologies
- Experience in large Enterprises, ISP, or Cloud providers.
- Experience developing service-oriented systems, REST, python
- Experience developing network automation or device management solutions
- Experience working in a network support role
- Experience working in a large ISP or cloud provider environment
- Experience with network modeling and programming – YANG, OpenConfig, NETCONF
- Experience with Ticket systems like Jira, and Version control systems like Git.
- Knowledge of Scrum & Agile Methodologies
- Strong technical knowledge in cloud networking, high performance computing, and GPU systems.
Internal Responsibilities
Key Responsibilities
Network Design, Development, and Validation:
-Designs and develops advanced network fabric and systems to support large-scale enterprise, data center, or cloud-based networking environments.
-Validates production networks and develops scaling strategies for varied use cases.
-Participates in the network solution and architecture design/validation process, providing expertise and guidance.
-Identifies and assesses complex risks in network design before deployment, recommending mitigations.
-Collaborates with vendors and internal groups (e.g., quality assurance) to ensure hardware, firmware, software, and cloud network code.
-Designs playbooks for investigations to resolve common issues related to network areas.
-Analyzes network workflows and proposes optimizations to enhance performance and efficiency.
Automation and Scripting:
-Designs and builds moderately complex modules within moderately complex automation frameworks for testing and operations to enhance efficiency.
-Automates critical network tasks, including production and lab environments and optimizes repetitive support functions to improve processes.
-Builds and customizes moderately complex dashboards, telemetry systems, and alerting mechanisms to proactively monitor network health.
-Writes and enhances advanced scripts to automate routine tasks across teams and business units.
-Uses and adjusts infrastructure pipelines, configuration items, using established playbooks and resources.
-Independently configures tools to meet requirements of new products or services.
Testing and Quality Assurance:
-Develops standalone, moderately complex test cases and test strategies to ensure and enhance network integrity.
-Serves as technical lead for test practices and strategies, providing guidance to the team and collaborating with other internal stakeholders in the development of robust testing protocols.
-Implements new approaches to improve break-fix validation and post-event root cause analysis (RCA), ensuring issues are resolved and lessons are integrated within the team.
-Conducts L1 network change reviews and creates and presents network changes to change management review boards.
-Collaborates with stakeholders to validate systems and environments through pre-production testing, ensuring compliance with internal standards prior to deployment.
-Provides audit documentation and evidence as part of Governance, Risk, and Compliance (GRC) to ensure alignment with governance and security requirements.
Monitoring and Reliability:
-Works with others (e.g. monitoring teams) to maintain advanced dashboards, telemetry systems, and alerting mechanisms to monitor network health, including defining metrics, trends, and thresholds to ensure service level objectives (SLOs).
-Uses tools from the monitoring team to monitor alerting systems and provides feedback to enhance the detection and diagnosis of network anomalies.
-Participates in operational support rotations, providing guidance in incident response and root cause investigations, and collaborating with stakeholders across team to resolve escalated issues.
-Develops and enhances tools and solutions for frontline support teams to drive efficiencies in responding to network failure conditions.
Cross-Team Collaboration and Leadership:
-Works alongside project program managers and uses templated project plan to ensure achievement of critical test and deployment milestones and project alignment.
-Mentors junior engineers and provides subject matter expert (SME) guidance on technical issues.
-Engages with customers to clarify technical issues, and contributes to the delivery of root cause analysis (RCA) documentation.
-Contributes to roadmap and software architecture discussions, ensuring ideas are shared within team and implemented where necessary.
-Collaborates with vendors and internal teams to ensure adherence to networking standards, identifying areas of misalignment and identifying resolutions.
Performance and Capacity Management:
-Analyzes network performance and service metrics (e.g., latency, throughput, packet loss) to identify network efficiency issues and takes next steps to scale network.
-Uses capacity information through security and performance metrics to ensure the infrastructure can support anticipated growth.
Core Responsibilities
Planning & Execution:
-Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
Collaboration & Partnership:
-Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Problem Solving:
-Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.
Continuous Learning:
-Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools, and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement:
-Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.
External Responsibilities
Key Responsibilities
Network Design, Development, and Validation:
-Designs and develops advanced network fabric and systems to support large-scale enterprise, data center, or cloud-based networking environments.
-Validates production networks and develops scaling strategies for varied use cases.
-Participates in the network solution and architecture design/validation process, providing expertise and guidance.
-Identifies and assesses complex risks in network design before deployment, recommending mitigations.
-Collaborates with vendors and internal groups (e.g., quality assurance) to ensure hardware, firmware, software, and cloud network code.
-Designs playbooks for investigations to resolve common issues related to network areas.
-Analyzes network workflows and proposes optimizations to enhance performance and efficiency.
Automation and Scripting:
-Designs and builds moderately complex modules within moderately complex automation frameworks for testing and operations to enhance efficiency.
-Automates critical network tasks, including production and lab environments and optimizes repetitive support functions to improve processes.
-Builds and customizes moderately complex dashboards, telemetry systems, and alerting mechanisms to proactively monitor network health.
-Writes and enhances advanced scripts to automate routine tasks across teams and business units.
-Uses and adjusts infrastructure pipelines, configuration items, using established playbooks and resources.
-Independently configures tools to meet requirements of new products or services.
Testing and Quality Assurance:
-Develops standalone, moderately complex test cases and test strategies to ensure and enhance network integrity.
-Serves as technical lead for test practices and strategies, providing guidance to the team and collaborating with other internal stakeholders in the development of robust testing protocols.
-Implements new approaches to improve break-fix validation and post-event root cause analysis (RCA), ensuring issues are resolved and lessons are integrated within the team.
-Conducts L1 network change reviews and creates and presents network changes to change management review boards.
-Collaborates with stakeholders to validate systems and environments through pre-production testing, ensuring compliance with internal standards prior to deployment.
-Provides audit documentation and evidence as part of Governance, Risk, and Compliance (GRC) to ensure alignment with governance and security requirements.
Monitoring and Reliability:
-Works with others (e.g. monitoring teams) to maintain advanced dashboards, telemetry systems, and alerting mechanisms to monitor network health, including defining metrics, trends, and thresholds to ensure service level objectives (SLOs).
-Uses tools from the monitoring team to monitor alerting systems and provides feedback to enhance the detection and diagnosis of network anomalies.
-Participates in operational support rotations, providing guidance in incident response and root cause investigations, and collaborating with stakeholders across team to resolve escalated issues.
-Develops and enhances tools and solutions for frontline support teams to drive efficiencies in responding to network failure conditions.
Cross-Team Collaboration and Leadership:
-Works alongside project program managers and uses templated project plan to ensure achievement of critical test and deployment milestones and project alignment.
-Mentors junior engineers and provides subject matter expert (SME) guidance on technical issues.
-Engages with customers to clarify technical issues, and contributes to the delivery of root cause analysis (RCA) documentation.
-Contributes to roadmap and software architecture discussions, ensuring ideas are shared within team and implemented where necessary.
-Collaborates with vendors and internal teams to ensure adherence to networking standards, identifying areas of misalignment and identifying resolutions.
Performance and Capacity Management:
-Analyzes network performance and service metrics (e.g., latency, throughput, packet loss) to identify network efficiency issues and takes next steps to scale network.
-Uses capacity information through security and performance metrics to ensure the infrastructure can support anticipated growth.
Core Responsibilities
Planning & Execution:
-Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
Collaboration & Partnership:
-Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Problem Solving:
-Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.
Continuous Learning:
-Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools, and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement:
-Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.