Designs and develops advanced network systems for scalable enterprise, data center, or cloud-based environments, ensuring alignment with existing infrastructure patterns and collaborating with vendors and internal stakeholders for comprehensive validation. Designs and builds advanced modules within automation frameworks to support testing, operations, and service reliability, automating high-impact network tasks and developing proactive monitoring systems. Designs advanced test strategies and reusable test cases, leading cross-functional collaboration, refining validation methods, and ensuring systems meet compliance and deployment standards. Designs advanced systems for network health monitoring, partners with teams to refine anomaly detection, leads incident response and root cause analysis, and builds efficient tools for frontline support teams. Drives milestone achievements with project managers, mentors junior engineers, leads customer engagements and root cause analysis (RCA) documentation, and ensures integration meets technical standards. Leads network performance analysis to drive scalable improvements and forecasts infrastructure needs for growth.
Location: on-site Nashville TN
Internal Responsibilities
Network Design, Development, and Validation:
- NW engineering background - experience in designing and configuring CLOS NW for cloud services
- Lead firmware upgrades, configuration refreshes, feature rollouts, and controlled network changes in partnership with operations and deployment teams.
- Deep experience designing large-scale leaf-spine and multi-fabric data center networks aka CLOS architecture.
- Hands-on expertise with Ethernet switching, BGP, EVPN/VXLAN, IP addressing, routing scale, ECMP, QoS, and network segmentation.
- Strong knowledge of management, OOB, console-server, and production data-center network architectures.
- Experience with network operating systems and hardware platforms such as Arista EOS, Cisco Nexus/NX-OS, NVIDIA Spectrum/Cumulus, or comparable platforms.
- Experience designing or operating AI/HPC/GPU cluster networking, including RDMA/RoCE, lossless Ethernet, and high-bandwidth fabrics.
- Experience leading firmware lifecycle management, configuration standardization, vulnerability remediation, and network change governance.
- Participates in solution architecture discussions, providing technical guidance on network requirements, tradeoffs, and design decisions.
- Identifies and assesses complex risks in network design and recommends mitigation strategies before deployment.
- Collaborates with vendors and internal stakeholders to align on hardware, firmware, software, and cloud network code requirements.
- Designs playbooks for resolving common and uncommon network issues involving moderately complex systems.
Analyzes network workflows to identify inefficiencies and proposes scalable solutions to improve performance.
Automation and Scripting:
- Experience with network automation using Python, Ansible, Terraform, APIs, Git, and CI/CD.
- Experience with capacity planning, hardware-resource/TCAM analysis, and large-scale routing-table design.
- Can partner with NW automation teams to design and develop NW modeling and configuration rendering solutions, to be able to review configuration diffs and mitigate configuration issues on NW switches.
- Designs and builds advanced modules within automation frameworks to support testing, operations, and service reliability across multiple environments.
- Automates high-impact network tasks for production and lab environments, and optimizes repetitive, cross-team workflows to improve operational efficiency.
- Develops and maintains advanced dashboards, telemetry tools, and alerting systems to enable proactive monitoring and faster issue resolution.
- Writes, enhances, and documents reusable scripts that streamline routine network operations across teams and product domains.
- Modifies and extends infrastructure pipelines and configuration management tools to meet evolving requirements using existing playbooks and custom logic.
Independently configures and extends tools to support new products or services, ensuring alignment with platform standards and integration best practices.
Testing and Quality Assurance:
- Designs and develops advanced test strategies and reusable test cases that validate complex network systems and enhance long-term network integrity and reliability, including testing network features in lab environment and rolling it out in production.
- Leads team-level test practices and cross-functional collaboration, helping define robust testing protocols and ensuring consistency in execution across environments.
- Implements and refines post-incident validation methods, integrating break-fix outcomes and lessons learned into continuous test improvements.
- Reviews and approves L1 and L2 network changes, and presents high-impact changes to cross-team change management boards, ensuring risks are understood and mitigated.
- Partners with cross-functional teams to drive pre-production validation efforts, ensuring systems and environments meet compliance and deployment-readiness criteria.
Prepares and delivers audit documentation and test evidence to support internal Governance, Risk, and Compliance (GRC) processes, maintaining alignment with security and policy requirements.
Monitoring and Reliability:
- Works with others (e.g. monitoring teams) to customize dashboards, telemetry pipelines, and alerting systems, defining service level objectives (SLOs) based thresholds and metrics to monitor network health.
- Partners with monitoring and Site Reliability Engineering (SRE) teams to refine alerting systems and enhance early detection of network anomalies and degradation.
- Leads incident response during support rotations, driving root cause analysis and coordinating cross-team resolution for escalated issues.
Builds and improves internal tools that enable frontline support teams to efficiently respond to network failures and recurring operational issues.
Cross-Team Collaboration and Leadership:
- Drives test and deployment milestones in collaboration with project/program managers, adapting plans to address risks and ensure cross-team alignment.
- Mentors junior engineers and acts as a technical SME, providing hands-on guidance in resolving complex issues.
- Leads customer engagements on technical escalations and delivers clear, actionable root cause analysis (RCA) documentation.
- Influences team-level roadmap and architecture decisions, contributing solutions that align with broader engineering goals.
Coordinates with vendors and internal teams to resolve standards misalignment, ensuring integration meets technical and operational expectations.
Performance and Capacity Management:
- Leads analysis of network performance metrics (e.g., latency, throughput, packet loss) to identify systemic inefficiencies and drive scalable improvements.
Forecasts infrastructure needs using performance and capacity data, ensuring systems are prepared for anticipated traffic and service growth.
Planning & Execution:
Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately-sized project or initiative. Efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
Collaboration & Partnership:
Collaborates across the organization to align on expectations and achieve shared objectives. Leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs. Supports inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.
Problem Solving:
Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices. Proactively escalates unresolved or critical issues with a thorough assessment and suggests potential solutions. Reviews, contributes to, and documents problem-solving strategies.
Continuous Learning:
Pursues learning opportunities to expand knowledge and skills and/or tools in new areas and stays abreast of the latest industry trends and best practices. Proactively seeks and leverages ongoing feedback and training to improve skills. Coaches and mentors junior team members, fostering continuous learning and knowledge sharing within and across teams.
Continuous Improvement:
Develops ideas, recommends updates, and/or collaborates on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams, and evaluates the impact on key stakeholders. Solicits feedback from others on ideas for alternative approaches and methods for continued improvement.
Performance and Development:
Contributes to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Qualifications:
- Deep experience designing large-scale leaf-spine and multi-fabric data center networks.
- Hands-on expertise with Ethernet switching, BGP, EVPN/VXLAN, IP addressing, routing scale, ECMP, QoS, and network segmentation.
- Strong knowledge of management, OOB, console-server, and production data-center network architectures.
- Experience with network operating systems and hardware platforms such as Arista EOS, Cisco Nexus/NX-OS, NVIDIA Spectrum/Cumulus, or comparable platforms.
- Experience designing or operating AI/HPC/GPU cluster networking, including RDMA/RoCE, lossless Ethernet, and high-bandwidth fabrics.
- Experience with network telemetry and observability: gNMI, streaming telemetry, SNMP, syslog, inventory/CMDB integration, and monitoring systems.
- Experience leading firmware lifecycle management, configuration standardization, vulnerability remediation, and network change governance.
- Strong incident-management and root-cause-analysis skills.
- Ability to communicate technical architecture and tradeoffs to engineering, operations, leadership, and customers.
#LI-CG2
External Responsibilities
Network Design, Development, and Validation:
- NW engineering background - experience in designing and configuring CLOS NW for cloud services
- Lead firmware upgrades, configuration refreshes, feature rollouts, and controlled network changes in partnership with operations and deployment teams.
- Deep experience designing large-scale leaf-spine and multi-fabric data center networks aka CLOS architecture.
- Hands-on expertise with Ethernet switching, BGP, EVPN/VXLAN, IP addressing, routing scale, ECMP, QoS, and network segmentation.
- Strong knowledge of management, OOB, console-server, and production data-center network architectures.
- Experience with network operating systems and hardware platforms such as Arista EOS, Cisco Nexus/NX-OS, NVIDIA Spectrum/Cumulus, or comparable platforms.
- Experience designing or operating AI/HPC/GPU cluster networking, including RDMA/RoCE, lossless Ethernet, and high-bandwidth fabrics.
- Experience leading firmware lifecycle management, configuration standardization, vulnerability remediation, and network change governance.
- Participates in solution architecture discussions, providing technical guidance on network requirements, tradeoffs, and design decisions.
- Identifies and assesses complex risks in network design and recommends mitigation strategies before deployment.
- Collaborates with vendors and internal stakeholders to align on hardware, firmware, software, and cloud network code requirements.
- Designs playbooks for resolving common and uncommon network issues involving moderately complex systems.
Analyzes network workflows to identify inefficiencies and proposes scalable solutions to improve performance.
Automation and Scripting:
- Experience with network automation using Python, Ansible, Terraform, APIs, Git, and CI/CD.
- Experience with capacity planning, hardware-resource/TCAM analysis, and large-scale routing-table design.
- Can partner with NW automation teams to design and develop NW modeling and configuration rendering solutions, to be able to review configuration diffs and mitigate configuration issues on NW switches.
- Designs and builds advanced modules within automation frameworks to support testing, operations, and service reliability across multiple environments.
- Automates high-impact network tasks for production and lab environments, and optimizes repetitive, cross-team workflows to improve operational efficiency.
- Develops and maintains advanced dashboards, telemetry tools, and alerting systems to enable proactive monitoring and faster issue resolution.
- Writes, enhances, and documents reusable scripts that streamline routine network operations across teams and product domains.
- Modifies and extends infrastructure pipelines and configuration management tools to meet evolving requirements using existing playbooks and custom logic.
Independently configures and extends tools to support new products or services, ensuring alignment with platform standards and integration best practices.
Testing and Quality Assurance:
- Designs and develops advanced test strategies and reusable test cases that validate complex network systems and enhance long-term network integrity and reliability, including testing network features in lab environment and rolling it out in production.
- Leads team-level test practices and cross-functional collaboration, helping define robust testing protocols and ensuring consistency in execution across environments.
- Implements and refines post-incident validation methods, integrating break-fix outcomes and lessons learned into continuous test improvements.
- Reviews and approves L1 and L2 network changes, and presents high-impact changes to cross-team change management boards, ensuring risks are understood and mitigated.
- Partners with cross-functional teams to drive pre-production validation efforts, ensuring systems and environments meet compliance and deployment-readiness criteria.
Prepares and delivers audit documentation and test evidence to support internal Governance, Risk, and Compliance (GRC) processes, maintaining alignment with security and policy requirements.
Monitoring and Reliability:
- Works with others (e.g. monitoring teams) to customize dashboards, telemetry pipelines, and alerting systems, defining service level objectives (SLOs) based thresholds and metrics to monitor network health.
- Partners with monitoring and Site Reliability Engineering (SRE) teams to refine alerting systems and enhance early detection of network anomalies and degradation.
- Leads incident response during support rotations, driving root cause analysis and coordinating cross-team resolution for escalated issues.
Builds and improves internal tools that enable frontline support teams to efficiently respond to network failures and recurring operational issues.
Cross-Team Collaboration and Leadership:
- Drives test and deployment milestones in collaboration with project/program managers, adapting plans to address risks and ensure cross-team alignment.
- Mentors junior engineers and acts as a technical SME, providing hands-on guidance in resolving complex issues.
- Leads customer engagements on technical escalations and delivers clear, actionable root cause analysis (RCA) documentation.
- Influences team-level roadmap and architecture decisions, contributing solutions that align with broader engineering goals.
Coordinates with vendors and internal teams to resolve standards misalignment, ensuring integration meets technical and operational expectations.
Performance and Capacity Management:
- Leads analysis of network performance metrics (e.g., latency, throughput, packet loss) to identify systemic inefficiencies and drive scalable improvements.
Forecasts infrastructure needs using performance and capacity data, ensuring systems are prepared for anticipated traffic and service growth.
Planning & Execution:
Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately-sized project or initiative. Efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
Collaboration & Partnership:
Collaborates across the organization to align on expectations and achieve shared objectives. Leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs. Supports inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.
Problem Solving:
Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices. Proactively escalates unresolved or critical issues with a thorough assessment and suggests potential solutions. Reviews, contributes to, and documents problem-solving strategies.
Continuous Learning:
Pursues learning opportunities to expand knowledge and skills and/or tools in new areas and stays abreast of the latest industry trends and best practices. Proactively seeks and leverages ongoing feedback and training to improve skills. Coaches and mentors junior team members, fostering continuous learning and knowledge sharing within and across teams.
Continuous Improvement:
Develops ideas, recommends updates, and/or collaborates on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams, and evaluates the impact on key stakeholders. Solicits feedback from others on ideas for alternative approaches and methods for continued improvement.
Performance and Development:
Contributes to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Qualifications:
- Deep experience designing large-scale leaf-spine and multi-fabric data center networks.
- Hands-on expertise with Ethernet switching, BGP, EVPN/VXLAN, IP addressing, routing scale, ECMP, QoS, and network segmentation.
- Strong knowledge of management, OOB, console-server, and production data-center network architectures.
- Experience with network operating systems and hardware platforms such as Arista EOS, Cisco Nexus/NX-OS, NVIDIA Spectrum/Cumulus, or comparable platforms.
- Experience designing or operating AI/HPC/GPU cluster networking, including RDMA/RoCE, lossless Ethernet, and high-bandwidth fabrics.
- Experience with network telemetry and observability: gNMI, streaming telemetry, SNMP, syslog, inventory/CMDB integration, and monitoring systems.
- Experience leading firmware lifecycle management, configuration standardization, vulnerability remediation, and network change governance.
- Strong incident-management and root-cause-analysis skills.
- Ability to communicate technical architecture and tradeoffs to engineering, operations, leadership, and customers.
#LI-CG2