Bachelor’s degree in Computer Science, Engineering, or a related field
5+ years of hands-on experience in Site Reliability Engineering or related roles
Proven experience in any cloud (AWS/GCP/Azure)
Experience with implementing SRE practices such as SLO/SLI, Error budgets, Postmortems, Reducing Toil, capacity planning, and Incident Management
Python or other scripting/programming language
Strong background in monitoring tools
Proficiency in CI/CD tools, infrastructure as code, and configuration management
Solid knowledge of container orchestration technologies (Kubernetes, Docker)
English language proficiency at an Upper-Intermediate level (B2) or higher
Certification in Kubernetes, AWS/GCP/Azure, or similar technologies
Proven experience in DevOps
Expertise in deployment and management of LLMs, including technologies like RAG
Knowledge of managing and optimizing AI/ML models in production environments, including basic deployment, monitoring, and maintenance
Native AI cloud services: AWS Bedrock, Google Vertex AI, Azure AI
Experience in designing, building, and operating AI agents and agentic frameworks
Coding Agents: Claude Code, OpenCode, Cursor, Trae, Antigravity