How we test prompts and retrieval quality, detect regressions, monitor cost and latency, evaluate user-facing quality, and safely evolve models, prompts, datasets, and providers over time
The right person is a strong software engineer with practical LLM application experience and a genuine quality mindset. You care not only that an AI feature works in a demo, but that it performs consistently for customers, degrades gracefully, provides traceable results, and improves through feedback loops. And you have real opinions about what makes an evaluation trustworthy — not just that one exists, but whether it measures the right thing
Partner with product engineers to design, build, and improve AI features using LLMs, RAG, semantic search, and agentic workflows
Contribute directly to production codebases in Python, C#/.NET, or related technologies
Improve prompt, retrieval, context assembly, ranking, grounding, and response-generation patterns
Help teams make practical architecture tradeoffs across quality, latency, cost, privacy, and maintainability
Support model and provider evaluations, migrations, fallback strategies, and rollout plans
Make AI quality measurable — define what "good enough" means
This is a defining pillar of the role, not an afterthought. It is not enough to have an evaluation; you will be responsible for whether our evaluations are adequate
Design and implement evaluation pipelines for LLM-powered features, and build scoring methodologies from first principles rather than reaching for the nearest metric
Build and maintain versioned golden datasets covering real-world use cases, edge cases, failure modes, and customer-critical workflows
Implement LLM-as-judge, heuristic, human-feedback, and task-specific quality scoring approaches
Establish the criteria that determine whether an existing evaluation is sufficient for a given feature and risk profile — and identify gaps before they become production issues
Establish prompt and retrieval regression testing as part of the development lifecycle
Define quality gates and thresholds that help teams know when an AI feature is ready to ship
Instrument LLM interactions, RAG pipelines, tool calls, and agent workflows using observability platforms (OnBoard currently uses Arize; comparable tools include Langfuse, LangSmith, W&B, and OpenTelemetry-based stacks)
Track latency, token usage, cost, retrieval quality, groundedness, failure modes, safety signals, and user feedback
Build dashboards and alerts that surface meaningful product and engineering signals, not just raw telemetry
Analyze production traces to identify quality issues, cost spikes, regressions, and improvement opportunities
Create runbooks and response patterns for common LLM and AI-product failure modes
Ensure AI systems handle sensitive data appropriately and align with security, privacy, SOC 2, ISO 27001, and data-residency requirements
Monitor guardrails, policy enforcement, content safety signals, and safety-related anomalies as a distinct observability concern
Support auditability and traceability of AI interactions where required
Accountability
Adaptability
AI Curiosity / Innovation
Applied Learning
Business Acumen
Collaboration
Customer Focus
Dealing with Ambiguity
Decision Making
Driving for Results
Initiating Action
Planning and Organizing
Technical / Professional Knowledge