Дополнительно
At Crusoe, our Production Engineering team plays a pivotal role in ensuring the reliability and performance of our infrastructure. Production Engineer at Crusoe is dedicated to detecting, analyzing, and preventing issues to maintain high Service Level Agreements (SLAs) through Service Level Indicators (SLIs) and Service Level Objectives (SLOs). Through automation and proactive remediation, our Production Engineers resolve common errors automatically, then advise various engineering teams how to build resilient code. We anticipate and resolve issues before they impact our customers, conduct thorough post-mortems, and drive continuous improvement. Our customer-centric approach ensures that clients always have access to the virtual machines they depend on. This role is crucial for maintaining the "gold standard" reliability and performance of Crusoe's AI platform. The ideal candidate has experience in Production Engineering practices, understanding of distributed systems, networking, Linux, and a passion for automation and problem-solving. This is a full-time position
Building Kubernetes Platform: Focus on scaling tooling and features dedicated to Crusoe's Managed Kubernetes and Managed VM platforms for external customers
Collaboration and Planning: Collaborate with the team in morning stand-up meetings to discuss ongoing projects, recent incidents, and priorities for the day. Collaborate on action plans for deploying new data centers or retrofitting existing ones. Work closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment
System Monitoring and Alerting: Review overnight alerts and system performance metrics to ensure everything is running smoothly. Analyze system logs and develop tools to enhance our monitoring capabilities
Incident Response and Problem Solving: Engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones. Resolve common errors automatically through automation and proactive remediation
Performance Monitoring and Optimization: Stay focused on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers
Documentation and Knowledge Sharing: Document work, share insights with the team, and plan for the next day's challenges, always with a customer-centric mindset