Results through July 21, 2026 at 6:30 p.m. ET

Which Odyssey translation wins a blind reading test?

We hid the names and asked readers to choose between eight English translations. This report covers the first 685 completed blind tests across 8,043 head-to-head choices.

685
total ballots
8
translations
8,043
head-to-heads
1,989
lowest exposure

Which translation won?

Finding: Fagles led this test. Robert Fagles’s Penguin translation ranked first among the eight tested translations in the 674-ballot published analytical cohort. It led the raw results, the regularized pairwise model, the scene-adjusted model, and all 50,000 whole-ballot bootstrap resamples. This is strong evidence about reader preference within this study but not a universal verdict on the best Odyssey translation.

Reader’s choice

Fagles

Underrated

Lattimore

Noteworthy

Mendelsohn

All-reader ranking

Pairwise strength estimates the chance of beating an average translation in this field after accounting for opponent strength. It is not the percentage of readers who chose that translation. The 95% ranges are shown by default.

This experiment measures relative reader preference among eight translations, across a study panel that ultimately included eight short scenes, in a self-selected and mostly American social-media sample. It does not measure fidelity to the Ancient Greek, the experience of reading an entire translation, or the preferences of all Odyssey readers.

Which translations were compared?

Answer: the eight translations with the highest estimated named search attention among 17 candidates.

  • Alexander PopeOUP
  • Samuel ButlerOpen Road
  • Robert FitzgeraldPicador
  • Richmond LattimoreHarper
  • Robert FaglesPenguin
  • Stanley LombardoHackett
  • Emily WilsonNorton
  • Daniel MendelsohnChicago

We evaluated 17 candidates using trailing 12-month U.S. search-volume estimates for exact full-name-plus-Odyssey queries and capped the test at eight to keep the comparison task manageable. This is a reproducible measure of named search attention—not sales, readership, classroom assignment, or literary importance. The report ranks these eight translations, not every English Odyssey.

Search attention across all 17 candidates

Linear scale · trailing 12-month average

4.3× Wilson over Fagles80× Wilson over Lombardo

Who took the test?

Sample: a mostly U.S., mobile, social-media convenience sample recruited during attention around Nolan’s Odyssey.

Device

Geography

Paid-ad response gender

First-touch source

The largest traffic source was our Instagram post. TikTok, Facebook, YouTube, and readers sharing the test also contributed.

The convenience sample came from the cultural moment that inspired the test: Christopher Nolan’s Odyssey. Social media reached people already interested in the film and Homer. That audience interest is the study’s primary selection bias.

Device, geography, and source figures come from campaign and visitor analytics. Gender comes from paid campaign-delivery data and was not collected on the ballot. These figures describe the audience reached by the campaign; they cannot be mapped precisely to either the 685 counted completions or the 674-ballot published analytical cohort. We did not retain raw IP addresses or direct identifiers; one-way IP hashes were used to detect retakes.

From movie attention to translation choice

Each lane is scaled to its own peak so the timing remains visible.

Does popularity predict preference?

Finding: named search attention was a weak proxy for preference in this eight-translation comparison. Emily Wilson had the highest named search attention but did not lead the blind test.

Popularity vs. pairwise strength

Eight tested translations · US Google search attention

Moving right means more named search attention; moving up means stronger pairwise strength. The lines at 1,000 average monthly searches and 50% pairwise strength are arbitrary visual guides. Moving the guides changes the quadrant labels, not the points or underlying data.

Leader

Fagles

Underrated

FitzgeraldLattimoreMendelsohnLombardo

Popular

Wilson

Buried

ButlerPope

Popularity uses DataForSEO's trailing 12-month estimate of average monthly US Google searches for each exact full-name-plus-Odyssey query. It is named search attention—not sales, readership, or literary quality.

Do readers prefer archaic language?

Emily Wilson argues that readers may mistake archaic language for greater authority or authenticity. Here is her full point:

“Mild stylistic archaism is often accepted without question in translations of ancient texts and can be presented as if it were a mark of authenticity. But of course, the English of the nineteenth or early twentieth century is no closer to Homeric Greek than the language of today. The use of a noncolloquial or archaizing linguistic register can blind readers to the real, inevitable, and vast gap between the Greek original and any modern translation.”

Below, we compare publication date with blind reader preference. Newer translations look stronger across the full 300-year span. Remove Pope and the pattern weakens. Among translations since 1960, it mostly disappears. This does not refute Wilson’s point that readers may be biased toward “mild stylistic archaism,” but it suggests they were not reflexively choosing the most archaic English in this test.

With Pope, the line slopes upward across all eight translations (correlation 0.71). Among translations published since 1960, it becomes -0.16. In plain terms: the pattern mostly separates older verse styles from the modern field; it does not steadily reward the newest translation.

Do Google AI mentions reflect preference?

Across the eight translations, Google AI mentions had a weak descriptive relationship with blind preference (correlation 0.22). Wilson appears most often in DataForSEO’s observed Google AI corpus but ranks sixth in pairwise strength.

Google AI mentions vs blind preference

Do ChatGPT results reflect preference?

This relationship is also weak (correlation 0.32). ChatGPT surfaces some leaders, but Mendelsohn ranks third in the blind test and appears rarely.

ChatGPT mentions vs blind preference

Did people just skim?

Answer: very few, and it did not matter. We excluded 11 timed completions under 90 seconds. Fagles still leads when they are restored.

How long did readers take?

DiscardedSkimSlowMedian 4:01

The 90-second floor removes 11 completions that required processing 24 excerpts and making 12 decisions at several hundred words per minute before allowing any decision time. Skim covers 1:30–2:14; slow covers 2:15 or more.

Does skimming change preference?

The all-reader view contains the full published cohort of 674 ballots. Fagles leads overall.

What if you discarded zero ballots?

Finding: Fagles still leads. Restoring all 11 completions under 90 seconds leaves the pairwise ranking unchanged. Raw match shares also move very little; Fitzgerald and Wilson tie at 79 top matches.

All 685 counted ballots

Pairwise strength estimates the chance of beating an average translation in this field after accounting for opponent strength.

This is just for auditing purposes. The original ranking is more robust because it applies the report’s consistent 90-second quality threshold. The zero-discard view shows that the winner was not created by that cutoff.

Did people cheat?

We found no coordinated bot or voting campaign. The site counted one ballot per anonymous browser and blocked retakes and later submissions from browsers that had already cast a counted ballot.

To check for coordinated activity on shared networks, we converted each IP address into a private one-way hash: a nonreversible fingerprint that lets matching addresses be compared without keeping the raw address in ballot records. Retakes never entered the public totals.

How were the scenes selected?

To select the eight-scene panel, we used GPT-5.6 Sol to construct a fixed pool of 25 candidate scenes using only Samuel Butler’s public-domain Project Gutenberg translation. Candidate scenes were ranked by their recurrence in nine study guides under the assumption that well-known plot points would increase test completion rate. GPT-5.6 Sol then narrowed that fixed pool to eight scenes selected for narrative and semantic breadth across the poem.

The panel-selection input did not contain any of the seven other tested translations or any reader results. It therefore could not directly optimize for wording that favored any competing translator. We evaluated the selected panel against all 1,081,575 possible eight-scene subsets of the same 25 candidates. Just 5,059 subsets, or 0.47%, matched or exceeded it on both recognition and semantic breadth. This is a benchmark to defend against cherry-picking. It doesn’t mean scenes were randomly sampled.

Recognition and breadth

Our panel compared with every eight-scene combination from the same 25 candidates

Middle 50% of panelsChosen panel

Fagles leads overall, but other translators win individual scenes. Pick one to compare its first- and eighth-place passages.

Opening

1,197 blind choices

#1Robert Fagles · Penguin

69% pairwise

Sing to me of the man, Muse, the man of twists and turns driven time and again off course, once he had plundered the hallowed heights of Troy.

#8Alexander Pope · OUP

26% pairwise

The man for wisdom’s various arts renown’d, Long exercised in woes, O Muse! resound; Who, when his arms had wrought the destined fall Of sacred Troy, and razed her heaven-built wall.

Global scene-adjusted ranking

This model gives equal weight to scenes, opponents, and recruitment waves and controls for first/second position. The top three remain unchanged, only Fitzgerald and Lombardo swap fourth/fifth.

What happens if you segment readers by taste?

Finding: not much. We searched for two to five recurring reader groups—such as one camp favoring literal translations and another favoring more literary ones—but found no division that was both stable and meaningfully distinct.

The two-group solution was highly repeatable, reappearing in 97.6% of resamples, but its measured separation was 16.4%, below our stated 20% qualification threshold. We therefore treat it as evidence of modest recurring taste variation rather than two sharply distinct camps.

No coherent reader camps emerged

674 ballots · candidate solutions from two to five groups

0 qualifying solutions

A reader type had to clear both gates. Some splits repeated reliably, but every split fell short on separation.

What this means: the aggregate winner does not appear to conceal a large, coherent minority camp. It does not mean every reader wants the same thing.

That does not mean readers agreed on every passage. It means the overall result does not appear to be an artificial average produced by several large, opposing taste camps.

What's the most accurate translation?

We are not classicists nor do we know Ancient Greek, so we’ve provided some essays below. From what we’ve learned, the question isn’t so simple. Every translation necessarily interprets and transforms the original. Each choice preserves some features while changing, narrowing, or sacrificing others. Here is a simple three-part model that begins to explain some of the difficulty.

Textual fidelity

What did Homer say?

Homer calls Odysseus πολύτροπον (polytropon), or “many-turning.” It can mean traveled, clever, adaptable, evasive, or deceptive.

  • Lattimore: “the man of many ways.” Broad, but vague.
  • Fagles: “the man of twists and turns.” Vivid and idiomatic; it foregrounds turning, change, and complexity.
  • Wilson: “a complicated man.” Clear, but psychological.
  • Mendelsohn: “so many roundabout ways / To wander.” Broad, but longer.

Emily Wilson argues that Fagles amplifies the drama and cumulative motion of Homer’s compact opening. Her comparison is illuminating, and it is also a first-person defense of her own translational method.

Poetic fidelity

What did Homer's poetry do?

Literal facts are not the whole poem. In Book 18, Homer uses πετάννυμι (petannymi), “to open or spread out,” when desire wakes in Penelope’s suitors.

  • Fagles: Athena will “fan” and “inflame” their hearts.
  • Lombardo: Penelope will “make their blood pound.”
  • Wilson: desire opens inside them “like a sail.”

Corinne Pache uses this example to show how a less literal sentence can keep the Greek image. Meter constrains the translator’s available choices: line length, rhythm, syntax, and imagery compete for limited space. Wilson uses short pentameter. Mendelsohn uses a longer six-beat line. Kim Montpelier explains Mendelsohn’s method. Gregory Nagy explains why he values several translations.

Interpretive fidelity

What did Homer leave undecided?

Every translation picks a meaning. πολύμητις (polymetis) means “rich in metis.” That can mean intelligence, planning, cunning, trickery, or deception. Wilson’s “lord of lies” picks one side. Pache thinks it picks too hard.

Social language has the same problem. “Maid” can hide enslavement. “Sluts” or “whores” can add a judgment Homer did not make. Wilson writes about inherited bias, and Pache reviews her choices. Pasquale Toscano argues that a new translation should reveal new possibilities, not try to replace its predecessors.

“My goal is not to rank them, nor to criticize my esteemed predecessors and colleagues.” - Emily Wilson

Wilson’s essay instead compares how particular translations echo and depart from the Greek. This is not our expertise. If it is yours, follow the sources above and judge the arguments for yourself.

What bias should readers know about?

As of July 22, 2026, Scrivium has no financial or editorial relationship with any tested translator or publisher. We do have a commercial interest in attracting readers to our interactive mythology course, but we do not think that shaped the result.

How was the test designed?

We hid translator names, dates, and publishers. Readers were shown up to 12 comparisons. For 668 of the 674 published ballots, the opening eight formed a balanced cycle: every available translator appeared exactly twice. Before each of the final four comparisons, the test refit a regularized Bradley–Terry model to that reader’s previous non-skipped choices. It first prioritized translators with the least evidence and, when skips had fragmented the comparison network, pairs that connected those fragments. Among the eligible pairs, it favored strong contenders with similar scores, greater uncertainty, and no previous head-to-head meeting. Scene assignment followed the phase-specific rules detailed below. Left and right positions were independently randomized for every comparison. Skips changed neither translator’s preference score, although the pair and scene still counted as having been shown. Detailed implementation facts:

  • Initialization: Every translator began at ability 0, equivalent to a neutral 50% predicted win probability.

  • Updating scores: After each non-skipped choice, the browser refit the model from all previous choices. It ran 80 gradient iterations with L2 regularization of 0.7, shrinking sparse estimates toward zero. Skips were omitted from the ability fit.

  • Eligible pairs: The starting universe was every unordered pair among the translators present in that ballot—28 pairs for eight translators.

  • Coverage and connectivity filters: If non-skipped choices left disconnected groups, only pairs joining two groups were considered. It then retained pairs containing the largest possible number of translators tied for the fewest decisions.

  • Repeated pairs: Previously shown pairs were excluded whenever an unseen pair survived those filters. Repeats were permitted only when necessary and were penalized according to 1 / (1 + previous meetings). A skipped pair still counted as previously shown.

  • Closeness: For abilities aᵢ and aⱼ, closeness was 1 / (1 + |aᵢ − aⱼ|). This was only one part of the composite score, alongside contender strength, uncertainty, connectivity, and novelty.

  • Final selection: The algorithm kept the six highest-scoring eligible pairs and made a weighted selection among them. Higher scores were more likely, but the highest score was not automatically chosen.

  • Scene assignment: The method changed prospectively. Published ballots 1–6 used the earlier four-scene cycle. Ballots 7–447 balanced scene exposure while also considering editorial strength, lexical contrast, and pair freshness. For published ballots 448–674, the selector removed the immediately previous scene when alternatives existed, retained the least-shown scenes, sorted them by ID, and selected one with a ballot-specific deterministic hash. Only this final phase was fully blind to candidate identity, wording, and the reader’s selections.

  • Unavailable material: Pair selection only saw translators in the session’s candidate set; scene assignment only saw passages present in that session. Nothing unavailable was imputed or backfilled.

  • Left/right order: The logical pair was independently flipped with a 50/50 cryptographic random draw before every comparison.

  • Randomness and seeds: There was no fixed study-wide assignment seed. The server cryptographically shuffled translator and passage order and generated a random ballot UUID. That UUID became a per-ballot seed for reproducible adaptive-pair and scene assignment. Left/right flips remained separately cryptographically random and unseeded.

Also note, a JSON validation issue was corrected prospectively. Daniel Mendelsohn was unavailable for the first 6 counted ballots; all six were retained, and he entered on ballot 7. Harmonious household and The Lotus-Eaters were unavailable for the first 457 counted ballots, of which 447 remain in the published cohort; they entered on counted ballot 458, or published ballot 448. No ballots were rewritten or excluded because of these corrections. The report separately excludes 11 ballots for finishing under 90 seconds; 10 of those ballots were collected during the scene-affected period. Opening rounds balanced the options then available, while the scene-adjusted model controls for translator-by-scene and recruitment-wave effects.

The published ranking fits a regularized Bradley–Terry model:

P(i ≻ j) = σ(βᵢ − βⱼ) = 1 / (1 + e−(βᵢ−βⱼ))
β is a translation’s estimated ability. We report pairwise strength as its average probability of beating each of the other seven translations. A unit L2 penalty pulls weakly observed estimates toward the field average instead of allowing extreme scores from sparse matchups.

The compensated scene model adds partially pooled scene effects, recruitment wave effects, and a first-position term:

logit P(i ≻ j) = (αᵢ − αⱼ) + (sₖᵢ − sₖⱼ) + (wₜᵢ − wₜⱼ) + δ
k identifies the scene, t the recruitment wave, and δ the first-option effect. Partial pooling keeps small scene samples from overreacting. In five-fold grouped cross-validation, mean held-out log loss improved from 0.669 to 0.633; the final sampler had R̂ = 1.00 and no divergences.

The 95% ranges come from 50,000 whole-ballot bootstrap resamples, preserving the dependence among one reader’s choices. Published charts use 674 ballots and 7,919 non-skipped choices. Retakes and later writes from the same anonymous browser never entered the public totals.

50,000 / 50,000

whole-ballot reruns ranked Fagles first

5.5% lower

held-out prediction error after scene adjustment

R̂ 1.00 · 0

maximum R̂ and sampler divergences; minimum effective sample size 1,230

Audit our data

The frozen snapshot, analysis settings, downloads, and reconciliation checks are below.

Aggregate downloads

Download the rankings and model diagnostics used in this report. These files contain aggregate results only: no ballot IDs, IP addresses, or individual ballot records.

Frozen source and analysis settings

The source snapshot and analysis choices used for every reported figure.

Snapshot cutoff
July 21, 2026 at 6:30 p.m. ET
Redis reconciliation
Exact
Counted ballots
685
Published cohort
674
Bootstrap runs
50,000
Model
Regularized and scene-adjusted Bradley-Terry

Media Inquiries

For interviews, data and methodology questions, publication-ready charts, or factual corrections, contact Kyle Cureau at kyle@scrivium.com or text Kyle at 415-390-6475 and include your outlet and deadline.