Skip to content

fix(ragas): join list contexts in Faithfulness before judging - #246

Open
Yuriy H (jurayh) wants to merge 1 commit into
braintrustdata:mainfrom
jurayh:fix/faithfulness-list-context
Open

Yuriy H (jurayh) wants to merge 1 commit into
braintrustdata:mainfrom
jurayh:fix/faithfulness-list-context

Conversation

@jurayh

Copy link
Copy Markdown

Faithfulness is the only ragas scorer that does not join a list context before sending it to the judge. The prompt ends up containing the Python list repr:

context: ['Paris is in France.', 'It is the capital.']
statements: ['Paris is the capital of France.']

instead of the documents themselves. Every other scorer in the same file (ContextRelevancy, ContextRecall, ContextPrecision, ContextEntityRecall) joins list contexts with newlines first, and the TypeScript scorer flattens them. The judge still gets the gist, so nothing looks broken, but it scores against a degraded prompt.

This joins list contexts in both the sync and async paths, matching the sibling scorers. Regression tests capture the context the judge receives for list and string inputs. String contexts are unchanged.

Found while auditing the ragas scorers. Related to #245, which fixed the same kind of silent degradation in AnswerCorrectness.

Faithfulness passed the context to the judge unjoined, so a list context
reached the prompt as a Python list repr (['Paris is in France.', 'It is
the capital.']) instead of the documents themselves. Every other ragas
scorer joins list contexts with newlines, and the TypeScript scorer
flattens them. Join in both the sync and async paths and add regression
tests that capture the context the judge receives.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant