RAG Retrieval Evaluation Engineer

Builds and interprets evaluation frameworks for RAG retrieval quality, including recall, precision, groundedness, and answer relevance metrics, to systematically improve system accuracy.

A RAG Retrieval Evaluation Engineer helps teams answer a deceptively hard question: is our retrieval-augmented generation system actually retrieving the right information, and is the model actually using it correctly? Many RAG systems are built and deployed without any systematic way to measure this, leaving teams guessing based on anecdotal spot checks or user complaints. This role fills that gap by helping design and interpret evaluation frameworks that quantify retrieval quality and answer faithfulness in a structured, repeatable way. The assistant works by first understanding your current RAG setup and what evaluation, if any, already exists, then helping you define a test set of representative queries with known correct answers or ground-truth documents. From there, it guides you through selecting and calculating relevant metrics, such as recall@k and precision@k for retrieval quality, mean reciprocal rank for ranking quality, and groundedness or faithfulness scores that check whether the generated answer is actually supported by retrieved context rather than hallucinated. It also explains how to evaluate answer relevance, distinguishing between a technically grounded response and one that actually addresses the user's question well. Expect practical guidance on building evaluation datasets when you do not already have one, including techniques like using historical query logs, synthetic query generation, or domain-expert annotation. The assistant can also explain how to set up automated evaluation loops using LLM-as-judge approaches, describing their strengths and known biases so you interpret results correctly rather than trusting them blindly. Results from working with this role typically include a concrete evaluation plan, a set of metrics tailored to your specific failure modes, and a clear methodology for tracking whether changes to chunking, retrieval, or prompting actually improve outcomes over time rather than just feeling better anecdotally. This role is especially valuable for teams moving a RAG prototype toward production, teams experiencing inconsistent or hallucinated answers who need to diagnose whether the problem is retrieval or generation, and organizations that need defensible, measurable quality benchmarks to report to stakeholders or auditors. It turns retrieval quality from a vague impression into a measurable, improvable engineering discipline.

🔒 Unlock the AI System Prompt

Sign in with Google to access expert-crafted prompts. New users get 10 free credits.

Sign in to unlock