Design, build, and maintain production-grade LLM integration pipelines — including retrieval-augmented generation (RAG), prompt engineering, output parsing, and chain orchestration
Develop and operate AI features within Jeeves's core financial products: spend categorization, document extraction, anomaly detection, financial Q&A, and automated reconciliation
Implement structured output validation, fallback handling, and confidence scoring to ensure AI decisions meet reliability standards for financial use cases
Evaluate and integrate AI frameworks and tools (LangChain, LlamaIndex, OpenAI API, Anthropic API, HuggingFace, vector databases) and advocate for the right tool for the job
Establish prompt versioning and evaluation practices to ensure AI outputs remain accurate and consistent as models and data evolve
Design and maintain vector search pipelines using databases such as Pinecone, Weaviate, or pgvector to power semantic search and RAG-based features
Build document ingestion and chunking pipelines for Jeeves's financial data — processing invoices, receipts, policy documents, and transaction records
Optimize retrieval quality through embedding model selection, chunk strategy, metadata filtering, and re-ranking techniques
Collaborate with data scientists to take trained ML models from experimental notebooks to production serving infrastructure
Build and maintain model serving endpoints with appropriate latency SLOs, input validation, and output monitoring
Implement model performance monitoring and data drift detection to ensure production models remain accurate over time
Support model retraining workflows by designing clean data pipelines and feature engineering that can be continuously updated
Integrate AI services cleanly with Jeeves's backend microservices — designing clear API contracts, circuit breakers, and graceful degradation patterns
Write high-quality, testable backend code in Python or Go/Node.js to power AI-integrated features
Instrument AI components with structured logging, distributed tracing, latency dashboards, and alerting to ensure operational visibility
Build human-in-the-loop review workflows for AI decisions that require oversight — particularly for high-value financial actions