The infrastructure team owns shared platforms and capabilities, while adjacent Data Platform, Data Engineering, and application teams own their applications and data products. You will partner with those teams to provide smooth integrations, operational guidance, and dependable infrastructure that helps them move quickly and safely
Please note we expect the Senior Software Engineering to work East Coast hours
Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, Cloudflare, and related cloud-native technologies
Deliver SRE and DevOps initiatives that improve reliability, scalability, observability, deployment safety, and operational readiness
Build reusable infrastructure modules, automation, and self-service workflows that reduce manual work and improve the developer experience
Help define and implement service-level indicators, service-level objectives, monitoring, alerting, and error-budget practices for critical systems
Participate in incident response and improve operational outcomes through clear runbooks, effective post-incident reviews, and durable corrective actions
Strengthen disaster-recovery readiness through recovery planning, automation, testing, and remediation of identified gaps
Improve CI/CD workflows and infrastructure delivery so engineering teams receive faster feedback and can deploy confidently
Partner with application, Data Platform, and Data Engineering teams to understand infrastructure needs and help teams operate their workloads effectively
Improve cloud efficiency through thoughtful architecture, capacity planning, Kubernetes resource optimization, cost visibility, and automation
Contribute to technical standards, architecture decisions, documentation, and the evolution of the team’s sustainable 24/7 operating model