3+ years in a DevOps, SRE, or MLOps role with a focus on cloud infrastructure and a background in cloud services (AWS, GCP, Azure)
Skills in building and managing CI/CD pipelines (Jenkins, GitLab CI, or cloud-native services) and proficiency in at least one scripting language (e.g., Python, Bash)
Familiarity with IaC tools (e.g., AWS CDK, CloudFormation, Terraform) and containerization/orchestration (Docker, Kubernetes)
Track record of deploying and operating LLM inference (e.g., vLLM, Triton, TGI, Ray Serve, KServe/Seldon)
Hands-on experience with LLM/app tracing and metrics (e.g., OpenTelemetry + Langfuse, Arize Phoenix, WhyLabs) and in building evaluation pipelines (offline/online, regression suites)
Skills in operating retrieval pipelines: embedding generation, indexing/refresh strategies, vector DBs (Pinecone, Weaviate, Milvus, FAISS), and relevance monitoring
Experience in running multi-agent workflows (LangGraph, CrewAI, AutoGen-like), including state management, retries, rate limits, tool-failure handling, and step-level auditing
Experience in implementing guardrails: secrets isolation, tool/API permissions, prompt-injection defenses, data leakage prevention, PII redaction, and policy enforcement
Background integrating agents with external tools using MCP (or similar tool-calling standards) and operating tool registries is a plus
Fluent in English (B2+ level)
Master's degree or PhD in Computer Science, AI, Machine Learning, or a related field
Experience with cloud-native GenAI services like AWS Bedrock, Azure AI Foundry, or Google Vertex AI
Familiarity with the architecture and operational challenges of Large Language Models (LLMs)
Experience designing or managing multi-agent systems or complex, orchestrated workflows
Knowledge of monitoring and observability tools like Prometheus, Grafana, or Datadog
Relevant cloud or DevOps certifications
Strong problem-solving skills and the ability to work effectively in a fast-paced, collaborative environment