Team owns the layer between clients and backend workloads: proxy and L7 services, API Gateway, and inference gateway for AI traffic. Every request to backend services passes through this layer, so latency, availability, and correctness here directly affect everything downstream. You'll build and grow the team, set its technical direction, and own these services end to end, from design through production operations.
Internal Responsibilities
Problems the team works on:
Proxy and L7: request routing, TLS termination, connection management, config propagation across a large fleet, and keeping tail latency predictable under load.
API Gateway: authentication and authorization, rate limiting and quotas, traffic policies, and giving service teams a safe, self-serve way to expose APIs.
Inference gateway: routing AI inference traffic, which behaves differently from typical web traffic. Requests are long-running and often streamed, cost varies widely by request, and backend GPU capacity is expensive and limited. That makes routing, queuing, fairness, and capacity efficiency hard problems.
What you'll do:
Build and grow the team: hire engineers and tech leads, set expectations, give direct feedback, and develop people over time.
Define the team's roadmap and priorities with leadership and partner teams.
Stay close enough to the technical work to review designs, weigh tradeoffs, and make calls when needed.
Own the operational health of the services: on-call, incident response, postmortems, and follow-through on reliability work.
Work closely with networking, security, platform, and ML infrastructure teams on problems that cross team boundaries.
What we're looking for:
Experience managing engineers who build and operate distributed systems in production.
A solid technical background in at least one of: proxies and load balancing, API gateways, service mesh, networking, or high-throughput backend services.
Experience running latency-sensitive services, including on-call and incident management.
Experience building something new, such as starting a team, launching a new service, or taking a product from early design to production.
A track record of hiring well and helping engineers grow.
Nice to have:
Hands-on experience with Envoy, NGINX or similar.
Familiarity with LLM inference serving (batching, streaming, GPU utilization).
Experience building internal platforms used by many engineering teams.
External Responsibilities
Problems the team works on:
Proxy and L7: request routing, TLS termination, connection management, config propagation across a large fleet, and keeping tail latency predictable under load.
API Gateway: authentication and authorization, rate limiting and quotas, traffic policies, and giving service teams a safe, self-serve way to expose APIs.
Inference gateway: routing AI inference traffic, which behaves differently from typical web traffic. Requests are long-running and often streamed, cost varies widely by request, and backend GPU capacity is expensive and limited. That makes routing, queuing, fairness, and capacity efficiency hard problems.
What you'll do:
Build and grow the team: hire engineers and tech leads, set expectations, give direct feedback, and develop people over time.
Define the team's roadmap and priorities with leadership and partner teams.
Stay close enough to the technical work to review designs, weigh tradeoffs, and make calls when needed.
Own the operational health of the services: on-call, incident response, postmortems, and follow-through on reliability work.
Work closely with networking, security, platform, and ML infrastructure teams on problems that cross team boundaries.
What we're looking for:
Experience managing engineers who build and operate distributed systems in production.
A solid technical background in at least one of: proxies and load balancing, API gateways, service mesh, networking, or high-throughput backend services.
Experience running latency-sensitive services, including on-call and incident management.
Experience building something new, such as starting a team, launching a new service, or taking a product from early design to production.
A track record of hiring well and helping engineers grow.
Nice to have:
Hands-on experience with Envoy, NGINX or similar.
Familiarity with LLM inference serving (batching, streaming, GPU utilization).
Experience building internal platforms used by many engineering teams.