Lead the development of scalable, elastic, and highly available distributed systems for hyperscale cloud environments. Architect critical components and drive performance, reliability, operational excellence, security, and automation.
System Design & Architecture
- Design and develop scalable, elastic distributed systems supporting horizontal and vertical scaling.
- Define scalability requirements and optimize code and data paths for high-throughput, hyperscale workloads.
- Leverage distributed state management and data-plane platforms for large-scale retrieval, storage, and processing.
- Build fault-tolerant, highly available systems using redundancy, replication, failover, and appropriate consistency and availability tradeoffs.
- Implement load shedding, throttling, and rate limiting while meeting defined SLOs.
- Define KPIs, telemetry, dashboards, and alerts to proactively monitor system health and performance.
- Design performance, load, fault-injection, and brownout testing to validate scalability, correctness, and resilience.
- Implement replication and synchronization mechanisms to ensure data integrity, durability, and availability.
Operational Excellence & Incident Management
- Proactively diagnose and resolve complex production issues and ensure operational readiness.
- Support incident response, root cause investigations, and operational support rotations.
- Design systems for in-service maintenance and upgrades with minimal customer impact.
- Mentor engineers on troubleshooting and operational best practices.
Security & Compliance
- Implement robust security controls for multi-tenant cloud environments, including encryption and access controls.
- Remediate security gaps and ensure compliance with applicable standards and requirements.
- Maintain accurate security and compliance documentation.
Automation & Change Management
- Develop Infrastructure as Code (IaC) and automation for managing cloud infrastructure.
- Enable safe, repeatable patching, upgrades, deployments, and rollbacks through effective change-management practices.
Internal Responsibilities
Below are the core responsibilities: Planning & Execution
- Plan and coordinate moderately complex engineering initiatives, managing priorities, timelines, dependencies, risks, and deliverables.
- Provide technical oversight and adjust execution plans as priorities or requirements evolve.
Collaboration & Partnership
- Collaborate across teams and organizations to align technical direction and achieve shared objectives.
- Understand customer and stakeholder needs and communicate technical decisions, risks, and tradeoffs effectively.
- Foster an inclusive and collaborative engineering environment.
Problem Solving
- Analyze and resolve complex technical issues using data, telemetry, and sound engineering judgment.
- Escalate critical issues with clear assessments and recommended solutions.
- Document key technical decisions, solutions, and lessons learned.
Continuous Learning
- Stay current with distributed systems, cloud infrastructure, reliability, security, and automation technologies and practices.
- Seek feedback, continuously develop technical expertise, and promote knowledge sharing.
- Coach and mentor junior engineers.
Continuous Improvement
- Identify and implement improvements to engineering processes, architectures, tools, and operational workflows.
- Drive improvements in efficiency, scalability, reliability, security, and engineering effectiveness.
Performance & Development
- Support team growth through mentoring and knowledge sharing.
- Participate in candidate interviews, technical assessments, and hiring recommendations.
External Responsibilities
Below are the core responsibilities: Planning & Execution
- Plan and coordinate moderately complex engineering initiatives, managing priorities, timelines, dependencies, risks, and deliverables.
- Provide technical oversight and adjust execution plans as priorities or requirements evolve.
Collaboration & Partnership
- Collaborate across teams and organizations to align technical direction and achieve shared objectives.
- Understand customer and stakeholder needs and communicate technical decisions, risks, and tradeoffs effectively.
- Foster an inclusive and collaborative engineering environment.
Problem Solving
- Analyze and resolve complex technical issues using data, telemetry, and sound engineering judgment.
- Escalate critical issues with clear assessments and recommended solutions.
- Document key technical decisions, solutions, and lessons learned.
Continuous Learning
- Stay current with distributed systems, cloud infrastructure, reliability, security, and automation technologies and practices.
- Seek feedback, continuously develop technical expertise, and promote knowledge sharing.
- Coach and mentor junior engineers.
Continuous Improvement
- Identify and implement improvements to engineering processes, architectures, tools, and operational workflows.
- Drive improvements in efficiency, scalability, reliability, security, and engineering effectiveness.
Performance & Development
- Support team growth through mentoring and knowledge sharing.
- Participate in candidate interviews, technical assessments, and hiring recommendations.