Please note, this team is hiring across all levels and candidates are individually assessed and appropriately leveled based upon their skills and experience
We're looking for an engineer to ensure the reliable operation of production environments for our Data Infrastructure and products, running at scale on large-volume distributed cloud systems. You'll focus on maximizing system reliability, automating routine tasks, and improving production efficiency
This role offers hands-on exposure to modern cloud technologies — Docker, Kubernetes, networking, and platforms like AWS and GCP — while you contribute directly to system uptime and user experience for a large-scale distributed application
You'll lead production monitoring and incident response, drive automation to reduce manual work, and collaborate with Engineering on root cause analysis, CI/CD support, and capacity planning
Monitoring: Use observability dashboards to monitor system performance, error rates, and resource utilization. Define new dashboards and alerts as and when required and write technical runbooks for Incident response
Incident Response: Act as the initial point of contact for production alerts, mitigate/solve the ongoing issues, escalate highly complex problems to the Engineering team, and conduct thorough root cause analysis for incidents
Automation: Identify repetitive manual tasks and streamline them through automation
Support CI/CD: Help create, test, and maintain automated application deployment pipelines
Capacity Management: Assist in tracking system resource usage (CPU, memory, storage) and traffic volume to help forecast and scale infrastructure needs