Mentors teams and leads the architecture of highly scalable, interdependent distributed systems. Identifies and removes performance/scalability bottlenecks for hyper‑scale workloads; defines scalability requirements with stakeholders; and designs elastic, high‑impact systems while advancing innovation in data plane platforms. Engineers and oversees fault‑tolerant, in‑service‑upgradable designs; optimizes resilience mechanisms (load‑shedding, throttling, rate‑limiting); and sets SLO‑aligned durability and availability standards across dependent services. Establishes KPIs and advanced telemetry; applies formal verification for complex features; and develops robust replication/synchronization strategies. Advises and leads resolution of complex production issues, sets operational readiness and SOP standards, and directs incident response and RCAs. Architects advanced security controls, drives remediation and compliance, and delivers enterprise‑level automation (IaC) and change strategies enabling safe, automated patching, updates, and rollbacks.
Internal Responsibilities
Key Responsibilities
System Design & Architecture - System Scalability:
–Mentor the team in the architecture and design of highly scalable, interdependent distributed systems, ensuring horizontal and vertical scalability and overall performance, including leveraging distributed state management tools.
–Lead the identification of performance and scalability bottlenecks and recommend solutions to optimize code and/or systems for large-scale data processing and high-throughput requirements to improve performance for hyper-scale systems.
–Lead collaboration with stakeholders to define system scalability requirements, ensuring the defined requirements meet customer expectations.
–Leverage deep expertise to design high-impact, interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).
–Drive innovation in the use of data plane platforms.
–Evaluate whether systems are meeting nonfunctional scalability requirements, and proactively anticipate growing business needs within the business unit.
System Design & Architecture - System Reliability Design:
–Design and oversee the implementation of fault-tolerant, interdependent systems capable of withstanding in-service updates by implementing sophisticated redundancy, replication, and automatic failover capabilities.
–Lead the design and implementation of systems that effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
–Guide the optimization of advanced mechanisms to handle network unreliability, including load-shedding, throttling, and rate-limiting.
–Design interdependent systems that are durable and adhere to service level objectives (SLOs), driving standards for availability and durability of other computing services within the organization
System Design & Architecture - System Reliability Performance:
–Define key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running, interdependent systems.
–Drive the creation and customization of highly complex dashboards, telemetry systems, and alerting mechanisms, proactively ensuring system health and reliability.
System Design & Architecture - Correctness / Availability:
–Maintain expertise in industry standards for verifying correctness and apply existing techniques to interdependent systems.
–Formally verify complex features (e.g., via TLA+) to ensure system design correctness for various interdependent systems.
–Develop advanced strategies for data replication and synchronization, ensuring robust data integrity and availability
Compliance & Security:
–Architect advanced security measures to protect data and applications in multi-tenant environments, and lead initiatives to enhance data and application protection.
–Guide the execution of comprehensive remediation plans to address identified security vulnerabilities.
–Ensure cloud infrastructure is in compliance with industry standards and regulations, and guide documentation efforts across projects.
Automation & Change Management:
–Develop enterprise-level automation tools and strategies (e.g., Infrastructure as Code (IaC)) and oversee their implementation.
–Drive alignment of change management plans and organizational initiatives for patching, updating, and rolling back applications, and design interdependent systems to allow for automation of these processes.
Minimum Qualifications:
• Bachelor’s degree in Computer Science or equivalent proven experience
• 10+ years of experience building and operating large scale, highly available, cloud based distributed systems
• Specialist skill in a modern programming language such as Java, C, C++, C#, Go, or Python, with proficiency in additional languages preferred
• Validated understanding of operating system fundamentals
• Strong understanding of data models and distributed persistence technologies
• Thorough understanding of the latest security principles, techniques, and protocols
• Strong troubleshooting and performance tuning skills
• Proficiency in network, distributed, asynchronous, and concurrent programming
• Knowledge of professional software engineering standard methodologies for the full software development process
• Experience building and operating scalable infrastructure software or distributed systems
• Proven track record to achieve stretch goals in a highly innovative and fast-paced environment
• Passion for technical leadership and mentoring
• Strong verbal and written communication skills
• Strong analytical skills, with excellent problem-solving abilities
Preferred Qualifications
• Experience in Agile/SCRUM enterprise-scale software development
• Experience building automated network and security solutions
• Knowledge of Machine Learning fundamentals
• Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures
• Working familiarity with storage principles, protocols and practices
• Working familiarity with building secure software using modern security principles
External Responsibilities
Key Responsibilities
System Design & Architecture - System Scalability:
–Mentor the team in the architecture and design of highly scalable, interdependent distributed systems, ensuring horizontal and vertical scalability and overall performance, including leveraging distributed state management tools.
–Lead the identification of performance and scalability bottlenecks and recommend solutions to optimize code and/or systems for large-scale data processing and high-throughput requirements to improve performance for hyper-scale systems.
–Lead collaboration with stakeholders to define system scalability requirements, ensuring the defined requirements meet customer expectations.
–Leverage deep expertise to design high-impact, interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).
–Drive innovation in the use of data plane platforms.
–Evaluate whether systems are meeting nonfunctional scalability requirements, and proactively anticipate growing business needs within the business unit.
System Design & Architecture - System Reliability Design:
–Design and oversee the implementation of fault-tolerant, interdependent systems capable of withstanding in-service updates by implementing sophisticated redundancy, replication, and automatic failover capabilities.
–Lead the design and implementation of systems that effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
–Guide the optimization of advanced mechanisms to handle network unreliability, including load-shedding, throttling, and rate-limiting.
–Design interdependent systems that are durable and adhere to service level objectives (SLOs), driving standards for availability and durability of other computing services within the organization
System Design & Architecture - System Reliability Performance:
–Define key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running, interdependent systems.
–Drive the creation and customization of highly complex dashboards, telemetry systems, and alerting mechanisms, proactively ensuring system health and reliability.
System Design & Architecture - Correctness / Availability:
–Maintain expertise in industry standards for verifying correctness and apply existing techniques to interdependent systems.
–Formally verify complex features (e.g., via TLA+) to ensure system design correctness for various interdependent systems.
–Develop advanced strategies for data replication and synchronization, ensuring robust data integrity and availability
Compliance & Security:
–Architect advanced security measures to protect data and applications in multi-tenant environments, and lead initiatives to enhance data and application protection.
–Guide the execution of comprehensive remediation plans to address identified security vulnerabilities.
–Ensure cloud infrastructure is in compliance with industry standards and regulations, and guide documentation efforts across projects.
Automation & Change Management:
–Develop enterprise-level automation tools and strategies (e.g., Infrastructure as Code (IaC)) and oversee their implementation.
–Drive alignment of change management plans and organizational initiatives for patching, updating, and rolling back applications, and design interdependent systems to allow for automation of these processes.
Minimum Qualifications:
• Bachelor’s degree in Computer Science or equivalent proven experience
• 10+ years of experience building and operating large scale, highly available, cloud based distributed systems
• Specialist skill in a modern programming language such as Java, C, C++, C#, Go, or Python, with proficiency in additional languages preferred
• Validated understanding of operating system fundamentals
• Strong understanding of data models and distributed persistence technologies
• Thorough understanding of the latest security principles, techniques, and protocols
• Strong troubleshooting and performance tuning skills
• Proficiency in network, distributed, asynchronous, and concurrent programming
• Knowledge of professional software engineering standard methodologies for the full software development process
• Experience building and operating scalable infrastructure software or distributed systems
• Proven track record to achieve stretch goals in a highly innovative and fast-paced environment
• Passion for technical leadership and mentoring
• Strong verbal and written communication skills
• Strong analytical skills, with excellent problem-solving abilities
Preferred Qualifications
• Experience in Agile/SCRUM enterprise-scale software development
• Experience building automated network and security solutions
• Knowledge of Machine Learning fundamentals
• Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures
• Working familiarity with storage principles, protocols and practices
• Working familiarity with building secure software using modern security principles