5+ years of experience building production backend systems, distributed systems, or infrastructure platforms
Strong systems design skills and experience owning significant systems from design through production
Depth in at least one of the following
AI agent systems, orchestration, tool use, evaluation, or grounding
Knowledge graphs or graph data modeling
Search, retrieval, ranking, RAG, or semantic search systems
Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems
Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms
Comfortable working across languages such as Go, TypeScript, Python, or Rust
Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack
Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve
We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure
This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure
Responsible for delivering the software but also for operating and supporting it in production
Why this Role
You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective
You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure
Remote based in India