What our data tells us about building more reliable AI translation
Every AI translation company talks about accuracy the same way. Pick a benchmark, run a test, publish a score. We do this too, our model comparison testing is some of our most-cited content. But a benchmark only tells you how a model performs on a sentence someone picked for the test. It doesn't tell you what your actual users are doing, or how differently invested they are depending on why they showed up.
So we looked at our own usage data: whether people uploaded a file or pasted text, and how long they stayed. The gap between paying and free behavior was bigger than we expected, and it's a more honest signal than anything a survey would have told us.
What does user behavior reveal that surveys can't?
None of what follows came from asking anyone what they were doing. It came from watching what they did, which is a more honest signal than any survey question. Two behaviors separate paying users from free users cleanly: whether they upload a file rather than paste text, and how long they stay once they do.
People staying 38 minutes and uploading a file are investing something a two-minute pasted sentence never requires. That's the clearest difference between sustained, invested use and a quick test you'll find in this dataset, and it didn't take a single survey question to see it.
Why does this matter for AI translation reliability?
Time and effort are the only signals here, and they're consistent enough to build an argument on. The people uploading a file rather than pasting text, and staying on the platform for far longer once they do, are showing something a label like "paying" or "free" can't capture on its own: how much they're investing in the outcome. Someone spending 38 minutes on a translation is depending on the result being right in a way a two-minute test never is.
That's the exact gap our SMART consensus approach was built to close, and the population depending on it most is exactly the one this data just described: people who show, through how long they stay and what they choose to upload, that walking away with a wrong answer costs them something. Running the same text through multiple independent models, and showing where they land together versus where they don't, gives them something a single fluent-sounding output can't: a way to tell a reliable translation from a risky one before it costs them anything. We've shown this pattern directly in our own head-to-head model testing, and again as we've expanded the model pool in later rounds.
"Session length is the signal I pay closest attention to. Something that holds someone's attention for 38 minutes, when a test submission gets abandoned in under two, tells us where a wrong translation would actually cost someone. That's the population we test SMART's consensus threshold against before we ship any change to it." - Rachelle Garcia, AI Lead at Tomedes / MachineTranslation.com
1. Behavior reveals investment better than any label could. File uploads and session length separate sustained, invested use from a quick test more reliably than account status alone.
2. Effort is a signal reliability has to answer to. The people investing the most time in a submission are the ones with the most riding on the result being right.
What this data really tells us is who's depending on SMART to be right, and it's a bigger, more specific group than "anyone translating something important." That's the group product and content decisions here are actually supposed to serve. It's also why the next thing we test won't be another leaderboard score. It'll be whether the mechanism holds up for the users who are clearly staking the most on it.
Frequently asked questions
1. What's the biggest behavioral difference between paying and free users on MachineTranslation.com?
File uploads and session length. Roughly three in five paying sessions involve uploading a document, versus about one in ten free sessions, and paying users' sessions run many times longer. Neither of these was asked in a survey. Both were simply observed in usage data.
2. Why does session length matter for AI translation reliability?
Someone spending significantly longer on a single submission, and choosing to upload a full document rather than paste a sentence, is investing effort that a quick test never requires. That level of investment is a signal that the result matters to them.
3. Does this data change how MachineTranslation.com's SMART consensus feature is built?
It does not change the mechanism, but it clarifies who depends on it most. SMART runs every translation through multiple independent AI models and surfaces the version the majority agree on, which matters most for the users most invested in getting a translation right.
4. Is AI translation reliable enough for professional or high-stakes work?
Single-model AI translation carries real risk on nuanced or high-stakes language, since errors often go undetected by the reader. This is why high-stakes translations on MachineTranslation.com are paired with consensus verification and, where needed, human review.
William drives content strategy and growth across Tomedes and MachineTranslation.com, with a focus on user behaviour, SEO, and what makes people choose one translation solution over another. He writes about the decisions behind the marketing, not just the outcomes.