Operated large-scale, high-performance synchronous distributed systems in the cloud, working directly with load balancers, reverse proxies, TLS/mTLS and routing, and using metrics, logs and traces to debug issues, optimize performance and improve overall reliability
Worked hands-on with Kubernetes-based infrastructure and Infrastructure-as-Code tooling — defining and evolving manifests, routing and deployment strategies, and managing DNS configuration — with a solid understanding of how these pieces fit together to deliver resilient traffic paths in production
Uses AI as part of their daily engineering workflow, leveraging AI tools and agents to investigate issues, write and review code, explore solutions and automate repetitive work, while consistently validating outputs and holding quality and safety as non-negotiable standards
Led end-to-end infrastructure platform projects from design to production go-live, breaking work into clear milestones, coordinating across stakeholders, managing risks and trade-offs, and following through post-rollout to stabilize and evolve the solution
Experience with Istio or other service-mesh technologies such as Linkerd or Envoy-based solutions — or legacy stacks like Finagle — designing and evolving traffic policies, resilience features and observability for service-to-service communication
Worked with AWS networking and compute primitives (ALB/NLB, security groups, VPC, Route 53, IAM) in production environments, with an understanding of how these layers interact with application traffic
Hands-on experience using Infrastructure-as-Code tools such as Pulumi or Terraform to design, version and roll out changes to cloud and Kubernetes infrastructure in a safe, repeatable way
Previously built platform capabilities or shared libraries for other engineers — abstractions, SDKs or internal tooling — that make it easier and safer for product teams to consume traffic and networking capabilities at scale