Experience operating SaaS platforms serving large-scale customer workloads
Experience working within Kubernetes-based microservices environments
Experience supporting globally distributed production environments
Experience with GitOps and ArgoCD
Experience implementing AI-assisted operational tooling or automation workflows
P22415_3261422
Strong experience operating large-scale production services in AWS and/or GCP
Deep expertise with Linux and Kubernetes in production environments
Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues
Extensive experience with Infrastructure as Code technologies such as Terraform and Helm
Strong software engineering skills in Golang and/or Python
Experience building automation and internal engineering platforms
Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies
Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, TLS, service networking, and traffic management
Experience with observability platforms, monitoring strategies, and production telemetry
Experience with or strong interest in AI-assisted engineering and operational automation
Strong expertise operating customer-facing production systems subject to SLA
Experience leading incident response and driving operational improvements
Deep understanding of reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning
Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices
Proven ability to balance reliability, scalability, security, and engineering velocity
This role supports US FedRAMP projects and requires the employee to be a US Person (US Citizen or Green Card Holder) to meet FedRAMP compliance and security clearance standards
Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design
Experience implementing operational controls, compliance standards (e.g., FedRAMP, SOC2, HIPAA), and best practices in highly regulated or security-sensitive government cloud environments is highly preferred
Demonstrated success contributing to complex engineering initiatives
Strong collaboration and communication skills
Experience working effectively within globally distributed engineering organizations spanning multiple timezones and cultures
Experience collaborating with engineers and contributing to technical capabilities within an organization
Ability to influence technical direction through expertise, partnership, and execution