What kind of evidence would you need?
As readers of Biblonia know, I have a stubborn interest, to say the least, in etymology and the origins of language. And origins in general. No wonder I spent 4 years writing a PhD, although it was in political and textual history rather than linguistics. Origins tickle me, and it’s a good tickle, more like gentle stroking, full of endorphins and dopamine.
When I look up a word in a dictionary, I tend to go straight to the etymon, and based on my searches, I end up with a Proto-Indo-European (PIE) root most of the time, a reconstruction of one of the oldest documented language families on Earth.
But how further back can we travel, before we hit a wall? If I had time to write another PhD, it would surely be on this topic. But since other commitments come first, I am content to dip my dilletant toes in the water.
But seriously, how far back can we go in putting the language puzzle together?
Starting with PIE, the reconstruction is done via the comparative method: systematic sound correspondences across attested daughter languages, such as Sanskrit, Greek, Latin, Gothic, Hittite, Tocharian, etc, formalised by 19th century scholars and refined into something called the Neogrammarian regularity principle. That’s pretty uncontroversial among linguists today.
Going further back, that’s where things get more contentious.
The comparative method relies on systematic sound correspondences across many lexical items, but the signal-to-noise ratio collapses as languages diverge.
The rough consensus “wall” is that reliable reconstruction via standard comparative method tops out around 8,000–10,000 years before our common era. That’s roughly the point where chance resemblance, borrowing, and universal onomatopoeia or nursery-word effects like mama/papa patterns become statistically indistinguishable from genuine cognates.
Beyond Proto-Indo-European, there has been a proposal to link families of languages together into something called ‘Nostratic’, from ‘nostras’ meaning our own, a coinage going back to the first years of the 20th century and the work of a linguist by the name of Holger Pederson.
Indo-European is maybe 6,000 years old and still just barely reconstructable. Nostratic would need to be 10,000–15,000 years old, double the depth, at a point where regular sound change has had far more time to obscure or erase the original signal. It’s less like reading a faded photocopy and more like reading a photocopy of a photocopy of a photocopy. Annoying, isn’t it?!
And beyond that, there’s less consensus and far more speculation, like something called The ‘Proto-World theory’. That’s the hypothesis that all human languages, everywhere, descend from a single common ancestor spoken by the first anatomically modern humans, tens of thousands of years ago. Some proponents suggest 50,000–150,000 years, tracking roughly with the out of Africa dispersal of Homo sapiens.
Walls, walls and more walls. We simply cannot currently demonstrate Proto-World or similar claims scientifically with existing comparative methods, and possibly never will be able to, because the evidence has genuinely decayed past recoverability. It pays to know where to stop. Not asking questions, but expecting good answers.



