Manage a team that designs, develops, troubleshoots and debugs software programs for databases, applications, tools, networks etc.
Internal Responsibilities
- Lead the end-to-end NPI lifecycle for current and next-generation high-performance NICs supporting OCI GPU and AI infrastructure, including architecture, qualification, performance, reliability, and production readiness.
- Partner with NIC, GPU, silicon, and systems vendors and internal OCI teams to bring new networking technologies from early development through large-scale fleet deployment.
- Build and lead the engineering organization responsible for OCI’s Collective Communication Library (CCL), defining its architecture, technical strategy, and performance roadmap.
- Drive optimization of collective communication, RDMA, GPU networking, and topology-aware algorithms across servers, racks, clusters, and AI superclusters.
- Own GPU cluster networking performance and develop benchmarking, modeling, and performance-analysis capabilities to identify and eliminate hardware and software bottlenecks.
- Drive hardware/software co-design across NICs, firmware, drivers, network fabrics, runtimes, and collective communication software.
- Establish automated qualification, regression, performance, reliability, and fault-testing processes to ensure cloud-scale production readiness.
- Build and develop a high-performing engineering organization spanning NIC NPI, networking software, collective communications, systems, performance engineering, and validation.
- Lead complex cross-functional programs, establish technical roadmaps and engineering metrics, manage external technology partnerships, and communicate technical strategy, risks, and results to senior OCI leadership.
Qualifications
- BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
- 15+ years of experience in systems, networking, distributed computing, HPC, GPU infrastructure, or related areas.
- Significant experience leading engineering teams responsible for complex hardware/software systems.
- Deep expertise in high-performance networking, including Ethernet, RDMA, RoCE, PCIe, DMA, and modern NIC architectures.
- Strong knowledge of GPU systems, distributed GPU communication, collective communication algorithms, and large-scale distributed training or HPC systems.
- Experience taking new hardware technologies from early engineering stages through qualification and large-scale production deployment.
- Strong systems performance-analysis skills across hardware, firmware, operating systems, networking, runtimes, and applications.
- Experience working with silicon vendors and leading complex cross-functional engineering programs.
- Excellent technical and executive communication skills.
- Experience with technologies such as NCCL, MPI, UCX, RDMA, RoCEv2, GPUDirect RDMA, CUDA, SmartNICs/DPUs, NIC firmware/drivers, congestion control, and large-scale AI cluster networking is highly desirable. Experience with GPU clusters containing thousands to tens of thousands of accelerators is strongly preferred.
External Responsibilities
- Lead the end-to-end NPI lifecycle for current and next-generation high-performance NICs supporting OCI GPU and AI infrastructure, including architecture, qualification, performance, reliability, and production readiness.
- Partner with NIC, GPU, silicon, and systems vendors and internal OCI teams to bring new networking technologies from early development through large-scale fleet deployment.
- Build and lead the engineering organization responsible for OCI’s Collective Communication Library (CCL), defining its architecture, technical strategy, and performance roadmap.
- Drive optimization of collective communication, RDMA, GPU networking, and topology-aware algorithms across servers, racks, clusters, and AI superclusters.
- Own GPU cluster networking performance and develop benchmarking, modeling, and performance-analysis capabilities to identify and eliminate hardware and software bottlenecks.
- Drive hardware/software co-design across NICs, firmware, drivers, network fabrics, runtimes, and collective communication software.
- Establish automated qualification, regression, performance, reliability, and fault-testing processes to ensure cloud-scale production readiness.
- Build and develop a high-performing engineering organization spanning NIC NPI, networking software, collective communications, systems, performance engineering, and validation.
- Lead complex cross-functional programs, establish technical roadmaps and engineering metrics, manage external technology partnerships, and communicate technical strategy, risks, and results to senior OCI leadership.
Qualifications
- BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
- 15+ years of experience in systems, networking, distributed computing, HPC, GPU infrastructure, or related areas.
- Significant experience leading engineering teams responsible for complex hardware/software systems.
- Deep expertise in high-performance networking, including Ethernet, RDMA, RoCE, PCIe, DMA, and modern NIC architectures.
- Strong knowledge of GPU systems, distributed GPU communication, collective communication algorithms, and large-scale distributed training or HPC systems.
- Experience taking new hardware technologies from early engineering stages through qualification and large-scale production deployment.
- Strong systems performance-analysis skills across hardware, firmware, operating systems, networking, runtimes, and applications.
- Experience working with silicon vendors and leading complex cross-functional engineering programs.
- Excellent technical and executive communication skills.
- Experience with technologies such as NCCL, MPI, UCX, RDMA, RoCEv2, GPUDirect RDMA, CUDA, SmartNICs/DPUs, NIC firmware/drivers, congestion control, and large-scale AI cluster networking is highly desirable. Experience with GPU clusters containing thousands to tens of thousands of accelerators is strongly preferred.