Oliver Wakeford
All projects
Applied research
Applied researchShipped2026

Does Retrieval Make Generated Copy Sound Like the Client?

In daily use since July 2026. Voice score vs blind human ratings: Spearman −0.04 (n = 48)

Retrieval-Augmented GenerationEvaluationLLMsStatistics

Publication

Master's internship report, ACL format

Submitted August 2026

1,091 items · 8 clients · in daily use by the content team

Built with

PythonSentence TransformersFAISSBM25SciPyClaude SonnetClaude Code

Client work

The corpus is eight clients' published material and the repository is private. The method and the numbers are described here, but the data can't be shared.

At Chiron, a marketing consultancy, I built the system the content team uses to draft posts in each client's voice. It finds a client's closest past posts and hands them to Claude as examples. The live version is a Claude Code skill that searches the client's published posts for a few keywords, then has Claude draft and check its own work.

My M1 internship was the offline study behind that design, on 1,091 published items from eight clients, with 20 captions per client held out as a 160-item test set. I compared BM25, dense search (bge-small with FAISS) and a hybrid of the two. Dense scored highest on relevance (0.765, against 0.754 hybrid and 0.730 BM25, with 95% bootstrap intervals), but relevance was measured in the same embedding space dense search ranks by, so that comparison flatters it. BM25 gave the most varied examples for a small relevance cost, and it keeps everything in files the team can read. That's why production runs on keywords.

Over 24 briefs and eight client accounts, every retrieval variant raised brand-voice similarity over a no-retrieval baseline, from 0.700 to between 0.737 and 0.759, and all five held after Holm correction (p < 0.001).

Then I checked that score against people. 48 generated drafts were rated blind at the agency in one consensus session, and the score didn't track the ratings (Spearman −0.04 (n = 48)). The report concludes the automatic scores shouldn't be relied on alone. I tested the LLM judge too. Reversing the order of each pair showed no position bias I could detect, and its agreement with a second, embedding-based judge was only fair (Cohen's kappa 0.39).

Two of my own mistakes are in the report. When I scored the live system's drafts against the clients' published posts, the published posts were sitting in the pool they were scored against, so each one matched itself. Excluding those self-matches brought their score from about 0.81 down to about 0.77, a gap about the size of the retrieval effect. And a finding in my June draft, that retrieval made house-style errors worse, didn't survive Holm correction when I re-ran it, so the final version withdraws it.

The content team has used the system every day since the end of July 2026.

What this does not show

  • Generation results rest on 24 briefs, one sample per condition. The direction is the finding; the magnitudes are rough.
  • The human ratings are one consensus session over 48 drafts, which is thin. The claim is only that the automatic score didn't track them.
  • The comparison between the live system and the clients' own posts is observational. Briefs, retriever and post-processing all differ, and the report names each.
  • Client work. The corpus is proprietary and the repository is private.