Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent experience
Proven experience in architecture roles involving cloud, data, or production Generative AI/LLM systems
Strong knowledge of cloud platforms (either Azure, AWS or GCP) and Infrastructure as Code tools like Terraform, Bicep, CDK, CloudFormation
Proficiency with containerization, orchestration, and API management (Docker, Kubernetes, API gateways/service meshes)
Experience with CI/CD and release management for ML/LLM workloads (Jenkins, GitHub Actions, GitLab CI, Azure DevOps)
Comprehensive expertise with LLM-based solutions, including deploying and operating LLM inference (e.g., vLLM, Triton, TGI, Ray Serve, KServe/Seldon), LLM/app tracing and metrics (e.g., OpenTelemetry, Langfuse, Arize Phoenix, WhyLabs), and building evaluation pipelines (offline/online, regression suites)
Ability to design and oversee data and retrieval pipelines (embedding generation, indexing/refresh strategies, vector DBs such as Pinecone, Weaviate, Milvus, FAISS, and relevance monitoring)
Experience architecting secure, scalable agentic systems, including multi-agent workflows (LangGraph, CrewAI, AutoGen-like), state management, retries, rate limits, tool-failure handling, step-level auditing, and integration with external tools via Model Context Protocol (MCP) or similar standards
Proven track record implementing security and compliance guardrails: secrets isolation, tool/API permissions, prompt-injection defenses, data leakage prevention, PII redaction, policy enforcement, and operating tool registries
Advanced proficiency in English (B2+/C1)