In order to skill up for this new AI world I have been working on a RAG project. Let me start by saying, yes, I understand that the hot frontier right now is harnesses, but I want an understanding of how to feed information to an LLM before we add in the complexities of tool calling (building an agent itself), and then also the enforcement of that tool calling/what to do with the results.
Last year, I started with the question “How do you have an LLM respond in the voice of a character”. There are a lot of grammar/word choice quirks that read like the output is AI and there is a lot of debate on whether that can be prompted out of the output. I do see a lot of clearly generated blog and social media posts. Some people claim that this is on account of the poster being lazy/inexperienced with guiding the output. So I set out to see if it could be different, and I was interested in building a character either way.
When I looked at the question of the voice of a character, I realized that I needed to understand who this character was in order to understand what their voice might be like. What I’ve learned over the course of building with an LLM is that your output is only as good as your input. I felt like that maxim also applied in this case. So, over the course of a couple weeks, I, along with my new sidekick, Claude, started generating a backstory, world, and voice around a character. This post though, isn’t about the generation of the character data, I want to explore a specific aspect of building a RAG, namely evaluations.
For now, I will gloss over the building of the RAG itself. Suffice it to say, I had this corpus of story and character building that I fed into a RAG to generate character responses. The important part for this post is that I wanted a mechanism to evaluate how the generated responses changed when I modified the prompt, added more data, changed to a different model, etc. I’ve always been big on testing so this is really a natural extension of that inclination.
I want to share what I’ve learned along the way although I don’t claim by any means to be an expert on this topic. (You’ll also see that I made lots of mistakes along the way.)
Preparation
The first obstacle I came up against was how to do the evaluations. The LLM’s inclination was to just build an evaluation framework from scratch, but that seemed like overkill to me so I evaluated some existing frameworks. This is not a comprehensive overview and as we probably all know, the field is changing very quickly so new frameworks are being released all the time.
Frameworks Considered:
- Braintrust – rejected because it’s a hosted platform
- LangSmith – same issue as above
- DeepEval – has built in metrics that are used often in industry to score rag conversations
I ended up going with DeepEval because I didn’t want to rely on a hosted platform. The fact that it had some built in metrics that I could use also seemed like it might come in handy so that I could see what people with more experience than me were using to evaluate their RAGs.
One thing to understand first is DeepEval’s format for evaluating a case
{
input: <this is the user input>,
context: <system prompt, other information>,
retrieval_context: <what is pulled from storage and passed to the LLM to generate a response>,
actual_output: <the response that the LLM generated>
}
Trial 1: Trying out the framework
So to start, I needed test user input in order to evaluate responses.
The test cases I was looking at were:
- a general greeting
- responding to a question about the character’s family
- responding to a question about the character’s hobbies.
Then, I chose the faithfulness metric since on the surface it seemed like being faithful to the source material would be something which I would want to test. This metric is actually four separate calls to the judge model as given below for the version of DeepEval that I was using.
Faithfulness metric prompts
# 1. Extract facts from the retrieval context
Based on the given text, please generate a comprehensive list of FACTUAL, undisputed truths, that can inferred from the provided text.
These truths, MUST BE COHERENT. They must NOT be taken out of context.
Example:
[same Einstein / photoelectric paragraph as step 2, with the answer under a "truths" key]
===== END OF EXAMPLE ======
**
IMPORTANT: Please make sure to only return in JSON format, with the "truths" key as a list of strings. No words or explanation is needed.
Only include truths that are factual, BUT IT DOESN'T MATTER IF THEY ARE FACTUALLY CORRECT.
**
Text:
{{ retrieval_context }}
JSON:
The words "FACTUAL, undisputed truths" come from the truths_extraction_limit parameter. If you set it to N, that phrase becomes "the N most important FACTUAL, undisputed truths per document".
# 2. Extract claims from actual_output
Based on the given text, please extract a comprehensive list of FACTUAL, undisputed truths, that can inferred from the provided actual AI output.
These truths, MUST BE COHERENT, and CANNOT be taken out of context.
Example:
Example Text:
"Albert Einstein, the genius often associated with wild hair and mind-bending theories, famously won the Nobel Prize in Physics—though not for his groundbreaking work on relativity, as many assume. Instead, in 1968, he was honored for his discovery of the photoelectric effect, a phenomenon that laid the foundation for quantum mechanics."
Example JSON:
{
"claims": [
"Einstein won the noble prize for his discovery of the photoelectric effect in 1968.",
"The photoelectric effect is a phenomenon that laid the foundation for quantum mechanics."
]
}
===== END OF EXAMPLE ======
**
IMPORTANT: Please make sure to only return in JSON format, with the "claims" key as a list of strings. No words or explanation is needed.
Only include claims that are factual, BUT IT DOESN'T MATTER IF THEY ARE FACTUALLY CORRECT. The claims you extract should include the full context it was presented in, NOT cherry picked facts.
You should NOT include any prior knowledge, and take the text at face value when extracting claims.
You should be aware that it is an AI that is outputting these claims.
**
AI Output:
{{ actual_output }}
JSON:
# 3. Judge each claim
This is the rendered text version. <RETRIEVAL_CONTEXT> and <CLAIMS> are placeholders I passed in.
Based on the given claims, which is a list of strings, generate a list of JSON objects to indicate whether EACH claim contradicts any facts in the retrieval context. The JSON will have 2 fields: 'verdict' and 'reason'.
The 'verdict' key should STRICTLY be either 'yes', 'no', or 'idk', which states whether the given claim agrees with the context.
Provide a 'reason' ONLY if the answer is 'no' or 'idk'.
The provided claim is drawn from the actual output. Try to provide a correction in the reason using the facts in the retrieval context.
Expected JSON format:
{{
"verdicts": [
{{
"verdict": "yes"
}},
{{
"reason": <explanation_for_contradiction>,
"verdict": "no"
}},
{{
"reason": <explanation_for_uncertainty>,
"verdict": "idk"
}}
]
}}
**
IMPORTANT: Please make sure to only return in JSON format, with the 'verdicts' key as a list of JSON objects.
Generate ONE verdict per claim - length of 'verdicts' MUST equal number of claims.
No 'reason' needed for 'yes' verdicts.
Only use 'no' if retrieval context DIRECTLY CONTRADICTS the claim - never use prior knowledge.
Use 'idk' for claims not backed up by context OR factually incorrect but non-contradictory - do not assume your knowledge.
Vague/speculative language in claims (e.g. 'may have', 'possibility') does NOT count as contradiction.
**
Retrieval Contexts:
<RETRIEVAL_CONTEXT>
Claims:
['<CLAIMS>']
JSON:
The score for the faithfulness metric is supposed to work as follows:
- Pull claims from the actual_output. The judge LLM breaks actual_output into separate factual claims.
- Pull truths from the retrieval_context. It extracts factual statements from retrieval_context.
- Give each claim from (1) a verdict against those truths from (2):
- – yes: the claim agrees with the retrieval_context
- – no: the claim directly contradicts the retrieval_context
- – idk: the retrieval_context doesn’t mention it
The default setting for pass/fail is
I ran into an issue almost immediately with one of the judges giving me the following reasoning for failing a response.
“The score is 0.00 because the actual output fundamentally misrepresents the retrieval context by attributing human activities to an AI. Specifically, the output incorrectly claims that ‘the AI’ engages in motorcycle riding and pickup basketball games, when the retrieval context clearly describes these as hobbies of a human individual, not an artificial intelligence.”
Turns out, in the prompt there is the line “You should be aware that it is an AI that is outputting these claims.” which made the judge reject responses for not acknowledging AI generation and speaking in first person. The responses are supposed to read as if the character is speaking to the user, making this criteria incorrect for my use case.
Another issue is that some of the things retrieved from storage aren’t facts. Some of them are directives for how the character should act and so framing things as facts doesn’t work for all my cases. For the family case though, facts is mostly fine, but I still had to deal with the issue of the judge wanting acknowledgement of the response being from an AI.
Trial 2: Trying out custom criteria
The DeepEval framework provides a way to use custom written criteria when the default ones don’t quite fit. So, since I still wanted to test whether the LLM was using the material correctly, I tried to write my own criteria to run across the evals.
As I began thinking about how to structure the criteria, I realized that one challenge with the family case was that the character should be reluctant to share all the details of his family, but the hobbies case didn’t have any such reluctance. This meant that the criteria I wrote had to be flexible with regards to how much of the retrieved_context was revealed in the response.
I tried to resolve this by using an injected attitude that each case could set which would determine the level of deflectiveness the response should have with regards to the retrieved_context. But if you remember from the structure of DeepEval’s cases, there’s no field called attitude so instead I appropriated the context field and directed the judge to look there for attitude.
My own criteria: Coloured By Memory
1. The expected attitudes for this input are given in the context
2. Take the input and come up with answer shapes based on the expected attitudes
3. Read the retrieval context and note whether any of it is relevant to the input
4. Check whether the reply uses any relevant retrieval context to answer the input
5. Penalize using retrieval context that is not relevant to the input
Different levels of deflectiveness
_LEVEL_STEPS = {
"not deflective": (
"Penalize using the retrieval context in a reply that does not try to answer the input",
"Penalize not using relevant retrieval context",
),
"lightly deflective": (
"Penalize using the retrieval context only as evasion, with none of it reaching the answer",
"Penalize not using relevant retrieval context. A reply that gives part of it and withholds part is using it",
),
}
Astute readers might note that I added complexity by having multiple attitudes without even trying just a single attitude to begin with. Those of us who adhered to TDD might call this taking too big of a step, and quite frankly, you would be correct. The pitfall of starting out using an LLM to help code is that it feels like I can take bigger steps than is prudent (especially when I”m still just trying to understand the problem space). I offer up only that I am human and make mistakes (may my machine overlords look kindly upon my foibles)
The end result of using this criteria was that the judge was unable to determine if the responses were evasive and penalized not using all of the retrieval context instead. It also had a hard time determining if the retrieval_context was relevant in the first place. I suspect that trying to use the context field for something it wasn’t originally designed for may have gotten in the way. Fighting your framework mostly leads to more headaches is something that I am learning anew.
Next time, we’ll look at taking a smaller step with just the family case.














