Proto-Indo-European homeland south of the Caucasus?

User Camulogène Rix at Anthrogenica posted an interesting excerpt of Reich’s new book in a thread on ancient DNA studies in the news (emphasis mine):

Ancient DNA available from this time in Anatolia shows no evidence of steppe ancestry similar to that in the Yamnaya (although the evidence here is circumstantial as no ancient DNA from the Hittites themselves has yet been published). This suggests to me that the most likely location of the population that first spoke an Indo-European language was south of the Caucasus Mountains, perhaps in present-day Iran or Armenia, because ancient DNA from people who lived there matches what we would expect for a source population both for the Yamnaya and for ancient Anatolians. If this scenario is right the population sent one branch up into the steppe-mixing with steppe hunter-gatherers in a one-to-one ratio to become the Yamnaya as described earlier- and another to Anatolia to found the ancestors of people there who spoke languages such as Hittite.

The thread has since logically become a trolling hell, and it seems not to be working right for hours now.

Reich’s proposal based on ancestral components to explain the formation of a people and language is a continuation of their emphasis on ancestry to explain cultures and languages. It seems quite interesting to see this happen again, given their current trend to surreptitiously modify their previous ‘Yamnaya ancestry’ concept and Yamnaya millennia-long R1a-R1b community (that supposedly explains a Yamna -> Corded Ware -> Bell Beaker migration) to a more general ‘steppe people’ sharing a ‘steppe ancestry’ who spoke a ‘steppe language’.

steppe-ancestry
Interesting arrows of dispersal of steppe ancestry, from Yamna -> Corded Ware -> Bell Beaker, from David Reich’s new book (yes, from 2018, number one bestseller in Amazon.com).

This new idea based on ancestral components suffers thus from the same essential methodological problems, which equate it – yet again – to pure speculation:

  1. It is a conclusion based on the genomic analysis of few individuals from distant regions and different periods, and – maybe more disturbingly – on the lack of steppe ancestry in the few samples at hand.
  2. Wait, what? Steppe ancestry? So they are trying to derive potential genetic connections among specific prehistoric cultures with a poorly depicted genetic sketch, based on previous flawed concepts (instead of on anthropological disciplines), which seems a rather long stretch for any scientist, whether they are content with seeing themselves as barbaric scientific conquerors of academic disciplines or not. In other words, statistics is also science (in fact, the main one to assert anything in almost any scientific field), and you cannot overcome essential errors (design, sampling, hypothesis testing) merely by using a priori correct statistical methods. Results obtained this way constitute a statistical fallacy.

  3. Even if the sampling and hypothesis testing were fine, to derive anthropological models from genomic investigation is completely wrong. Ancestral component ≠ population.
  4. To include not only potential migrations, but also languages spoken by these potential migrants? It’s sad that we have a need to repeat it, but if ancestral component ≠ population, how could ancestral component = language?

The Proto-Indo-European-speaking community

This is what we know about the formation of a Proto-Indo-European community (i.e. a community speaking a reconstructible Proto-Indo-European language) in the Pontic-Caspian steppe, which is based on linguistic reconstruction and guesstimates, tracing archaeological cultures backwards from cultures known to have spoken ancient (proto-)languages, and helping both disciplines with anthropological models (for which ancient genomics is only helping select certain details) of migration or – rarely – cultural diffusion:

NOTE. The following dates are obviously simplified. Read here a more detailed linguistic assessment based on phonology.

neolithic_steppe-anatolian-migrations
Most likely Pre-Proto-Anatolian migration with Suvorovo-Novodanilovka chiefs in the North Pontic steppe and the Balkans.
  • ca. 5000 BC. Early Proto-Indo-European (or Indo-Uralic) spoken probably during the formation and development of a loose Early Khvalynsk – Sredni Stog I cultural-historical community over the Pontic-Caspian steppe region, whose indigenous population probably had mainly Caucasus hunter-gatherer ancestry.
  • ca. 4500 BC. Khvalynsk probably speaking Middle Proto-Indo-European expands, most likely including Suvorovo-Novodanilovka chiefs into the North Pontic steppe, and probably expanding R1b-M269 lineages for the first time.
  • ca. 4000 BC. Separated communities develop, including North Pontic cultures probably gradually dominated by R1a-Z645 (potentially speaking Proto-Uralic); and Khvalynsk (and Repin) cultures probably dominated by R1b-L23 lineages, most likely developing a Late Proto-Indo-European already separated from Proto-Anatolian.
  • ca. 3500 BC. A Proto-Corded Ware population dominated by R1a-Z645 expands to the north, and slightly later an early Yamna community develops from Late Khvalynsk and Repin, expanding to the west of the Don River, and to the east into Afanasevo. This is most likely the period of reduction of variability and expansion of subclades of R1a-Z645 and R1b-L23 that we expect to see with more samples.
  • ca. 3000 BC. Expansion of Corded Ware migrants in northern Europe, and Yamna migrants along the Danube and into the Balkans, with further reduction and expansion of certain subclades.
  • ca. 2500 BC. Expansion of Bell Beaker migrants dominated by R1b-L51 subclades in Europe, and late Corded Ware migrants in east Yamna expanding R1a-Z93 subclades.

All these events are compatible with language reconstruction in mainstream European schools since at least the 1980s, supported by traditional archaeological research of the past 20 years, and is being confirmed with Genomics.

For those willingly lost in a myriad of new dreams boosted by the shallow comment contained in David Reich’s paragraph on CHG ancestry, even he does not doubt that the origin of Late Proto-Indo-European lies in Yamna, to the north of the Caucasus, based on Anthony’s (2007) account:

yamnaya-migrations-reich
Both images from the book, posted by Twitter user Jasper at https://twitter.com/jaspergregory.

NOTE: By the way, David Anthony, one of the main sources of information for Reich’s group, never considered Corded Ware to have received Yamna migrants, and althought he changed his model due to the conclusions of the 2015 papers, he has recently changed his model again to adapt it to the inconsistencies found in phylogeography.

CHG ancestry and PIE homeland south of the Caucasus

As for the potential origins of CHG ancestry in early Proto-Indo-European speakers, I already stated clearly my opinion quite recently. They may be attributed to:

Just to be clear, an expansion of Proto-Anatolian to the south, through the Caucasus, cannot be discarded today. It will remain a possibility until Maykop and more Balkan Chalcolithic and Anatolian-speaking samples are published.

However, an original Early Proto-Indo-European community south of the Caucasus seems to me highly unlikely, based on anthropological data, which should drive any conclusion. From what I could read, here are the rather simplistic arguments used:

  • Gimbutas and Maykop: Maykop was thought to be (in Gimbutas’ times) a rather late archaeological culture, directly connected to a Transcaucasian Copper Age culture ca. 2400-2300 BC. It has been demonstrated in recent years that this culture is substantially older, and even then language guesstimates for a Late PIE / Proto-Anatolian would not fit a migration to the north. While our ignorance may certainly be used to derive far-fetched conclusions about potential migrations from and to it, using Gimbutas (or any archaeological theory until the 1990s) today does not make any sense. Still less if we think that she favoured a steppe homeland.

NOTE. It seems that the Reich Lab may have already access to Maykop samples, so this suggested Proto-Indo-European – Maykop connection may have some real foundation. Regardless, we already know that intense contacts happened, so there will be no surprise (unless Y-DNA shows some sort of direct continuity from one to the other).

  • Gamkrelidze & Ivanov: they argued for an Armenian homeland (and are thus at the origin of yet another autochthonous continuity theory), but they did so to support their glottalic theory, i.e. merely to support what they saw as favouring their linguistic model (with Armenian being the most archaic dialect). The glottalic theory is supported today – as far as I know – mainly by Kortlandt, Jagodziński, or (Nostraticist) Bomhard, but even they most likely would not need to argue for an Armenian homeland. In fact, their support of a Graeco-Aryan group (also supported by Gamkrelidze & Ivanov) would be against this, at least in archaeological terms.
  • Colin Renfrew and the Anatolian homeland: This conceptual umbrella of language spreading with farming everywhere has changed so much and so many times in the past 20 years, with so many glottochronological and archaeological estimates circulating, that you can support anything by now using them. Mostly used today for abstract models of long-lasting language contacts, cultural diffusion, and constellation analogies. Anyway, he strives to keep up-to-date information to revise the model, that much is certain:
  • Glottochronology, phylogenetic trees, Swadesh list analysis, statistical estimates, psychics, pyramid power, and healing crystals: no, please, no.
Science Magazine
“A first line of evidence comes from linguistic analysis based on quantitative lexical data, which returned a tree compatible with the Anatolian hypothesis

In principle, unlike many other recent autochthonous continuity theories, I doubt there can be much racial-based opposition anywhere in the world to an origin of Proto-Indo-European in the Middle East, where the oldest civilizations appeared – apart, obviously, from modern Northeast and Northwest Caucasian, Kartvelian, or Semitic speakers, who may in turn have to revisit their autochthonous continuity theories radically…

Nevertheless, it is obvious that prehistoric (and many historic) migrations are signalled by the reduction in variability and expansion of certain Y-DNA haplogroups, and not just by ancestral components. That is generally accepted, although the reasons for this almost universal phenomenon are not always clear.

In fact, Proto-Anatolian and Common Anatolian speakers need not share any ancestral component, PCA cluster, or any other statistical parameter related to steppe populations, not even the same Y-DNA haplogroups, given that approximately three thousand years might have passed between their split from an Indo-Hittite community and the first attested Anatolian-speaking communities…We must carefully follow their tracks from Anatolia ca. 1500 BC to the steppe ca. 4500 BC, otherwise we risk creating another mess like the Corded Ware one.

In my opinion, the substantial contribution of EHG ancestry and R1a-M417 lineages to the Pontic-Caspian steppe (probably ca. 6500 BC) from Central or East Eurasia is the most recent sizeable genomic event in the region, and thus the best candidate for the community that expanded a language ancestral to Proto-Indo-European – whether you call it Pre-Proto-Indo-European, Pre-Indo-Uralic, or Eurasiatic, depending on your preferences.

An early (and substantial) contribution of CHG ancestry in Khvalynsk relative to North Pontic cultures, if it is found with new samples, may actually be a further proof of the Caucasian substrate of Proto-Indo-European proposed by Kortlandt (or Bomhard) as contributing to the differentiation of Middle PIE from Uralic. Genomics could thus help support, again, traditional disciplines in accepting or rejecting academic controversial theories.

Conclusion

In the case of an Early PIE (or Indo-Uralic) homeland, genomic data is scarce. But all traditional anthropological disciplines point to the Pontic-Caspian steppe, so we should stick to it, regardless of the informal suggestion written by a renown geneticist in one paragraph of a book conceived as an introduction to the field.

It seems we are not learning much from the hundreds of peer-reviewed, statistically (superficially, at least) sound genetic papers whose anthropological conclusions have been proven wrong by now. A lot of people should be spending their time learning about the complex, endless methods at hand in this kind of research – not just bioinformatics – , instead of fruitlessly speculating about wild unsubstantiated proposals.

As a final note, I would like to remind some in the discussion, who seem to dismiss the identification of CHG with Proto-Indo-European by supporting a “R1a-R1b” community for PIE, of their previous commitment to ancestral components in identifying peoples and languages, and thus their support to Reich’s (and his group’s) fundamental premises.

You cannot have it both ways. At least David Reich is being consistent.

Related:

The uneasy relationship between Archaeology and Ancient Genomics

Allentoft Corded Ware

News feature Divided by DNA: The uneasy relationship between archaeology and ancient genomics, Two fields in the midst of a technological revolution are struggling to reconcile their views of the past, by Ewen Callaway, Nature (2018) 555:573-576.

Interesting excerpts (emphasis mine):

In duelling 2015 Nature papers6,7the teams arrived at broadly similar conclusions: an influx of herders from the grassland steppes of present-day Russia and Ukraine — linked to Yamnaya cultural artefacts and practices such as pit burial mounds — had replaced much of the gene pool of central and Western Europe around 4,500–5,000 years ago. This was coincident with the disappearance of Neolithic pottery, burial styles and other cultural expressions and the emergence of Corded Ware cultural artefacts, which are distributed throughout northern and central Europe. “These results were a shock to the archaeological community,” Kristiansen says.

(…)

Still, not everyone was satisfied. In an essay8 titled ‘Kossinna’s Smile’, archaeologist Volker Heyd at the University of Bristol, UK, disagreed, not with the conclusion that people moved west from the steppe, but with how their genetic signatures were conflated with complex cultural expressions. Corded Ware and Yamnaya burials are more different than they are similar, and there is evidence of cultural exchange, at least, between the Russian steppe and regions west that predate Yamnaya culture, he says. None of these facts negates the conclusions of the genetics papers, but they underscore the insufficiency of the articles in addressing the questions that archaeologists are interested in, he argued. “While I have no doubt they are basically right, it is the complexity of the past that is not reflected,” Heyd wrote, before issuing a call to arms. “Instead of letting geneticists determine the agenda and set the message, we should teach them about complexity in past human actions.”

Many archaeologists are also trying to understand and engage with the inconvenient findings from genetics. (…)
[Carlin:] “I would characterize a lot of these papers as ‘map and describe’. They’re looking at the movement of genetic signatures, but in terms of how or why that’s happening, those things aren’t being explored,” says Carlin, who is no longer disturbed by the disconnect. “I am increasingly reconciling myself to the view that archaeology and ancient DNA are telling different stories.” The changes in cultural and social practices that he studies might coincide with the population shifts that Reich and his team are uncovering, but they don’t necessarily have to. And such biological insights will never fully explain the human experiences captured in the archaeological record.

Reich agrees that his field is in a “map-making phase”, and that genetics is only sketching out the rough contours of the past. Sweeping conclusions, such as those put forth in the 2015 steppe migration papers, will give way to regionally focused studies with more subtlety.

This is already starting to happen. Although the Bell Beaker study found a profound shift in the genetic make-up of Britain, it rejected the notion that the cultural phenomenon was associated with a single population. In Iberia, individuals buried with Bell Beaker goods were closely related to earlier local populations and shared little ancestry with Beaker-associated individuals from northern Europe (who were related to steppe groups such as the Yamnaya). The pots did the moving, not the people.

This final paragraph apparently sums up a view that Reich has of this field, since he repeats it:

Reich concedes that his field hasn’t always handled the past with the nuance or accuracy that archaeologists and historians would like. But he hopes they will eventually be swayed by the insights his field can bring. “We’re barbarians coming late to the study of the human past,” Reich says. “But it’s dangerous to ignore barbarians.”

I would say that the true barbarians didn’t have a habit or possibility to learn from the higher civilizations they attacked or invaded. Geneticists, on the other hand, only have to do what they expect archaeologists to do: study.

EDIT (30 MAR 2018): A new interesting editorial of Nature, On the use and abuse of ancient DNA.

See also:

Y-DNA haplogroup R1b-Z2103 in Proto-Indo-Iranians?

chalcolithic_early-asia

We already know that the Sintashta -> Andronovo migrants will probably be dominated by Y-DNA R1a-Z93 lineages. However, I doubt it will be the only Y-DNA haplogroup found.

I said in my predictions for this year that there could not be much new genetic data to ascertain how Pre-Indo-Iranian survived the invasion, gradual replacement and founder effects that happened in terms of male haplogroups after the arrival of late Corded Ware migrants, and that we should probably have to rely on anthropological explanations for language continuity despite genetic replacement, as in the Basque case.

Nevertheless, since we have very few samples, I think we could still see a clear genetic contribution from Yamna to Corded Ware immigrants in the North Caspian region (from Abashevo, in turn a mix of Fatyanovo/Balanovo and Catacomb/Poltavka cultures) in terms of:

  • Ancestral components and PCA in new Sintashta-Petrovka, Andronovo, and/or later samples – similar the ‘steppe’ drift seen in Potapovka relative to Sintashta samples, both formed by incoming Corded Ware migrants – ; and
  • R1b-L23 subclades, either appearing scattered during the Sintashta melting pot (of Abashevo/R1a-Z645 and East Yamna-Poltavka/R1b-Z2103 peoples), or resurging after this period, as we have seen in Pre-Balto-Slavic territory.

This contribution could better explain the obvious language continuity in the region, beautifully complementing the complex anthropological model we have now of archaeological continuity of Sintashta and Potapovka with the previous Poltavka, seen in a similar material and symbolic culture that survived the arrival of newcomers.

A lot of people seem to be looking like crazy since O&M 2018 for some sort of connection between Corded Ware and Yamna migrants in Eastern and Central Europe (wheter in SNP calls of samples published, or among almost forgotten academic papers), either to support the ideas of the 2015 papers – for those who relied on their conclusions and built (even if only mentally) far-fetched migration models around it – , or just because of some sort of absurd continuity theory involving modern R1a-Z645 subclades:

NOTE. The situation we have seen with the hundreds of samples from O&M 2018, and with the recent additional Eastern European samples, depict an unexpected absolutely clear-cut distinction in Y-DNA haplogroups between Corded Ware and Yamna/Bell Beaker: I really can’t see how the situation could be more obvious for everyone, so I doubt any further samples will make certain people change their minds. Their hope is, I guess, that just one sample may give some more oxygen to infinite pet theories, as we are still surprisingly seeing even with reactionary R1b autochthonous continuists in Western Europe…

However, looking into the most likely future for the field, what we should be expecting right now is continuity of Yamna ancestry and lineages in early Proto-Indo-Iranian territory. Since we only have a few samples from Sintashta-Petrovka, Potapovka, and Andronovo, I think there might be a sizeable number of R1b-Z2103 subclades in the territory inhabited by those who – no doubt – spread the language into Central Asia.

Haplogroup_R1b_(Y-DNA)
Modern Y-DNA haplogroup R1b distribution, by Maulucioni at Wikipedia

While full population replacement by R1a-Z93 lineages in the North Caspian region ca. 2000 BC is not impossible, I don’t think it is very likely, since we already know that there are R1b-Z2103 lineages widely distributed in Indo-Iranian-speaking territory, and Z93 is now known to be an older subclade than YFull’s mean formation date suggested (due to the Ukraine_Eneolithic I6561 sample‘s SNP call), so what we can infer now that actually happened in Sintashta -> Andronovo is not exactly the spread of haplogroup Z93 during its formation, but rather a regional reduction in its variability coupled with the expansion of some of its subclades.

The main question, after the South Asia paper is finally published, will then be:

  1. Given that Yamna peoples were an elite group of patrilineally-related families mainly of R1b-L23 subclades:
  2. Accepting that PCA, ADMIXTURE, and other statistical methods are not relevant (alone) for ethnolinguistic identification: e.g. Yamna ‘outliers’ and East Bell Beaker migrants of R1b-L23 lineages without steppe ancestry; N1c1a1a-L392 lineages and Siberian ancestry unrelated to Uralic speakers; R1a-Z645 and steppe ancestry in North-East Europe related to Uralic-speaking cultures
  3. If we find now, as I expect, genetic continuity of east Yamna in Sintashta -> Andronovo (relative to other late Corded Ware peoples), probably including haplogroup R1b-Z2103 mixed with R1a-Z93 before its further reduction of subclades (e.g. to L657) and expansion during its subsequent spread southward…

bronze_age_early_Asia-andronovo
Diachronic map of migrations in Asia ca. 2250-1750 BC

Why exactly do we need Corded Ware to explain migrations of Late Indo-European speakers?

In other words: if we had the data we have today in 2015, would we have a need for Corded Ware to explain Indo-European migrations from the steppe? Are some people so blinded by their will to (appear to) be right in their past interpretations that they can’t just let go?

NOTE. On a side note, wouldn’t it be nice for this paper to publish some other R1b-L23 (x2103) sample – maybe even R1b-L51 – in Yamna, Andronovo, or Afanasevo territory, to end both autochthonous continuity theories (of North-Eastern and Western Europe) at the same time?

I really hope someone in David Reich’s team understands this matter, or else they will still identify Corded Ware as the (now probably ‘a’ instead) vector of expansion of Indo-European languages, and some of us will still have fun for another 2 or 3 years with such conclusions, until someone in the lab realizes that ancestry ≠ population ≠ ethnic identification ≠ language.

NOTE. It seems rather dull to read how people are discussing in the Twitterverse conventional constructs like ‘human race‘ as found in Reich’s op-ed in The New York Times, as if such grandiose semantic discussions had any practical meaning, when basic anthropological questions actually relevant for Genomics, like the essential ancestral component ≠ people tenet seem not to be of interest for anyone in the field….

Since our Indo-European demic difusion model (and its consequences for our reconstruction of North-West Indo-European) and this blog are becoming more and more popular each day – judging by the constant growth in visits in the past 6 months or so – , I guess the simplemindedness and predictability of certain geneticists is benefitting traditional anthropology directly, driving more and more amateur geneticists to look for sound academic models to answer the growing inconsistencies of genetic research.

NOTE. I am not saying the rejection of Corded Ware as spreading Indo-European is definitive. Maybe more samples within some years will depict a clear ancient expansion of Early or Middle Proto-Indo-Europeans from Khvalynsk to the forest-steppe and forest zone, and later with certain Corded Ware migrants into Central Europe, over whose territory a Late Indo-European dialect from Bell Beakers became the superstrate, as some have proposed in the past – e.g. to explain Krahe’s Old European hydronymy. I really doubt you could demonstrate such an old ethnolinguistic identification with a clear, unbroken archaeological trail, though, and we know now that this old hydronymy is probably of Late Indo-European nature (possibly even more recent).

What I am saying is: with the data we have now, it does not make any sense to keep the anthropological models invented by geneticists ex nihilo in 2015, and the hundred different alternative Late Indo-European migration models that arebornwitheachnewpaper.

These Yamna -> Corded Ware migration models didn’t have any sense for me since early 2016, but now after O&M 2017, and especially O&M 2018, I don’t think any geneticist with a little knowledge in Linguistics or Archaeology (if they are decent about their quest for truth in describing ancient European migrations) would buy them, if not for some sort of created ‘tradition’. So let’s ditch Corded Ware as Late Indo-European-speaking, let’s accept that late Corded Ware migrants should most likely be identified as early Uralic speakers, and then future data will tell if we are – again – wrong.

Please, don’t let Genomics become another pseudoscience based solely on Bioinformatics like glottochronology: let anthropologists (preferably mainstream archaeologists, but also the true Indo-Europeanists, linguists) help you interpret your raw data. Don’t deceive yourselves thinking that you have read enough about the Indo-European question, or that you know enough Indo-Europeanists (say what?) to derive your own conclusions.

Use the South Asia paper to begin expressly retracting the Corded Ware mess.

Please pretty please with sugar on top?

Related:

For commenters: this post concerns an anthropological question, and deals with the expansion of Late Proto-Indo-European speakers from Yamna, and the controversy surrounding the role of Corded Ware migrants that a handful of academics propose spread from it, based on a renewed model of Gimbutas’ outdated Kurgan theory and on the so-called ‘Yamnaya’ ancestry.

It happens so that the discussion has turned lately mainly to ancient Y-DNA haplogroups, because they help confirm previous mainstream anthropological models of cultural diffusion and migration. It is obviously not reasonable to judge prehistoric ethnolinguistic migrations from ca. 5,000 years ago based on historical nation-states and ethnic or religious concepts invented since the Middle Ages, coupled with “your” people’s main modern (or your own) paternal lineage.

EDIT (27 MAR 2018): Minor corrections and post made shorter.

Oldest N1c1a1a-L392 samples and Siberian ancestry in Bronze Age Fennoscandia

Open access preprint at bioRxiv, Ancient Fennoscandian genomes reveal origin and spread of Siberian ancestry in Europe, by Lamnidis et al. (2018).

Abstract (emphasis mine):

European history has been shaped by migrations of people, and their subsequent admixture. Recently, evidence from ancient DNA has brought new insights into migration events that could be linked to the advent of agriculture, and possibly to the spread of Indo-European languages. However, little is known so far about the ancient population history of north-eastern Europe, in particular about populations speaking Uralic languages, such as Finns and Saami. Here we analyse ancient genomic data from 11 individuals from Finland and Northwest Russia. We show that the specific genetic makeup of northern Europe traces back to migrations from Siberia that began at least 3,500 years ago. This ancestry was subsequently admixed into many modern populations in the region, in particular populations speaking Uralic languages today. In addition, we show that ancestors of modern Saami inhabited a larger territory during the Iron Age than today, which adds to historical and linguistic evidence for the population history of Finland.

Interesting excerpts (edited):

While the Siberian genetic component described here was previously described in modern-day populations from the region, we gain further insights into its temporal depth. Our data suggest that this fourth genetic component found in modern-day north-eastern Europeans arrived in the area around 4,000 years ago at the latest, as illustrated by ALDER dating using the ancient genome-wide data from Bolshoy Oleni Ostrov. The upper bound for the introduction of this component is harder to estimate. The component is absent in the Karelian hunter-gatherers (EHG) 3 dated to 8,300-7,200 yBP as well as Mesolithic and Neolithic populations from the Baltics from 8,300 yBP and 7,100-5,000 yBP respectively. While this suggests an upper bound of 5,000 yBP for the arrival of Siberian ancestry, we cannot exclude the possibility of its presence even earlier, yet restricted to more northern regions, as suggested by its absence in populations in the Baltic during the Bronze Age. Our study also presents the earliest occurrence of the Y-chromosomal haplogroup N1c in Fennoscandia. N1c is common among modern Uralic speakers, and has also been detected in Hungarian individuals dating to the 10th century, yet it is absent in all published Mesolithic genomes from Karelia and the Baltics.

The large Siberian component in the Bolshoy individuals from the Kola Peninsula provides the earliest direct genetic evidence for an eastern migration into this region. Such contact is well documented in archaeology, with the introduction of asbestos-mixed Lovozero ceramics during the second millenium BC, and the spread of even-based arrowheads in Lapland from 1,900 BCE. Additionally, the nearest counterparts of Vardøy ceramics, appearing in the area around 1,600-1,300 BCE, can be found on the Taymyr peninsula, much further to the east. Finally, the Imiyakhtakhskaya culture from Yakutia spread to the Kola Peninsula during the same period. Contacts between Siberia and Europe are also recognised in linguistics. The fact that the Siberian genetic component is consistently shared among Uralic-speaking populations, with the exceptions of Hungarians and the non-Uralic speaking Russians, would make it tempting to equate this component with the spread of Uralic languages in the area. However, such a model may be overly simplistic. First, the presence of the Siberian component on the Kola Peninsula at ca. 4000 yBP predates most linguistic estimates of the spread of Uralic languages to the area. Second, as shown in our analyses, the admixture patterns found in historic and modern Uralic speakers are complex and in fact inconsistent with a single admixture event. Therefore, even if the Siberian genetic component partly spread alongside Uralic languages, it likely presented only an addition to populations carrying this component from earlier.

admixture-uralic
Plot of ADMIXTURE (K=3) results containing West Eurasian populations and the Nganasan. Ancient individuals from this study are represented by thicker bars.

The novel genome-wide data here presented from ancient individuals from Finland opens new insights into Finnish population history. Two of the three higher coverage individuals and all six low coverage individuals from Levänluhta showed low genetic affinity to modern-day Finnish speakers of the area. Instead, an increased affinity was observed to modern-day Saami speakers, now mostly residing in the north of the Scandinavian Peninsula. These results suggest that the geographic range of the Saami extended further south in the past, and hints at a genetic shift at least in the western Finnish region during the Iron Age. The findings are in concordance with the noted linguistic shift from Saami languages to early Finnish. Further ancient DNA from Finland is needed to conclude to what extent these signals of migration and admixture are representative of Finland as a whole.

fennoscandia-pca
PCA plot of 113 Modern Eurasian populations, with individuals from this study projected on the principal components. Uralic speakers are highlighted in light purple.

The two samples of haplogroup N1c1a1a-L392/L1026, dated ca. 1500 BC, come from the site Bolshoy Oleniy Ostrov, in the Kola Peninsula.

Bolshoy Oleniy Ostrov (Great Reindeer Island), situated in the Kola Bay of the Barents Sea and separated from the mainland by Yekarerininsky Island and two straits, harbors the ancient cemetery of an unknown Early Metal Age culture. The preservation of artifacts made from bone and antler, wooden structures, as well as human remains is remarkable for the location and age this site represents. Altogether 19 skeletons of adults and children have been recognized from both single and collective burials of the site, together with more than 250 artifacts. (…) Apart from these excavations, approximately 25 burials were revealed in 1934 during the construction of fortifications. (…) Radiocarbon dates are provided by Moiseyev and Khartanovich in their 2012 study, placing the site in middle to the late 2nd millennium BC (…)

After seing how Late Indo-European languages spread with Yamna and (mainly) R1b-L23 lineages, we are now obtaining proof of how Siberian ancestry – likely accompanying N1c-L392 lineages – was probably related to an early archaeological Siberian influence in the easternmost region of North-East Europe, seen also probably in linguistics.

NOTE. Whereas I proposed – based mainly on common guesstimates – that R1a-M417 and EHG ancestry might have signaled the arrival of an early Yukaghir substratum to NE Europe, later acquired by Uralic spreading over this territory, while N1c1a1a lineages with the Seima-Turbino phenomenon might have given Uralic its later Altaic traits, it is indeed possible – and more likely with the findings in this paper – that N1c1a1a lineages may have in fact spread Yukaghir languages, especially if (like the Leiden school) one supports an Indo-Uralic community.

The linguistic effect of this migration may depend on one’s preferred model for Proto-Uralic and its strata, and especially on one’s position in the Proto-Uralic vs. Proto-Uralo-Yukaghir controversy. Although I really didn’t have a strong opinion on this matter, it is clear from my texts that (unlike Kortlandt) I didn’t consider Yukaghir to share a common ancestor with Uralic languages. What genomics is showing right now seems to me directly translatable to a linguistic model, and we should therefore reject an original Proto-Uralo-Yukaghir community.

Also, it seems that the Finnish population peak which expanded today’s prevalent N1c-L392 lineages – after the Iron Age bottleneck which likely reduced its haplogroup diversity – may have been associated with the event that displaced the Saami population from Finland after ca. 1000 AD.

I think it is becoming still clearer where Uralic languages came from.

Related:

Uralic as a Corded Ware substrate of Indo-Iranian, and loanwords in Finno-Ugric

Asko Parpola has recently published a new paper, Finnish vatsa ~ Sanskrit vatsá and the formation of Indo-Iranian and Uralic languages.

Abstract:

Finnish vatsa ‘stomach’ < PFU *vaćća < Proto-Indo-Aryan *vatsá- ‘calf’ < PIE *vet-(e)s-ó- ‘yearling’ contrasts with Finnish vasa- ‘calf’ < Proto-Iranian *vasa- ‘calf’. Indo-Aryan -ts- versus Iranian -s- refl ects the divergent development of PIE *-tst- in the Iranian branch (> *-st-, with Greek and Balto-Slavic) and in the Indo-Aryan branch ( > *-tt-, probably due to Uralic substratum). The split of Indo-Iranian can be traced in the archaeological record to the differentiation of the Yamnaya culture in the North Pontic and Volga steppes respectively during the third millennium BCE, due to the use of separate sources of metal: the Iranian branch was dependent on the North Caucasus, while the Indo-Aryan branch was oriented towards the Urals. It is argued that the Abashevo culture of the Mid-Volga-Kama-Belaya basins and the Sejma-Turbino trade network (2200–1900 BCE) were bilingual in Proto-Indo-Aryan and PFU, and introduced the PFU as the basis of West Uralic (Volga-Finnic) into the Netted Ware Culture of the Upper Volga-Oka (1900–200 BCE).

He updates thus his quite recent model from On the emergence, contacts and dispersal of Proto-Indo-European, Proto-Uralic and Proto-Aryan in an archaeological perspective (2017).

In it he supported a North-West Indo-European expansion with Corded Ware, and a Neolithic Proto-Uralic community in East Europe (associated with the Comb Ware culture), as I did before the famous 2015 papers.

In fact, he supports that the satemization trend of Proto-Indo-Iranian is due to a Proto-Finno-Ugric substratum in its population in the Volga-Ural region, similar to the model I propose (with the Corded Ware substratum hypothesis).

NOTE. While for Parpola the ‘satemizing’ substratum of Balto-Slavic (a NWIE dialect) may not come exactly from the same Finno-Ugric population as for Indo-Iranian, but from a different Uralic dialect (as I explain in my hypothesis), for the few extant supporters of an Indo-Slavonic group there should not be any problem identifying the same ancient substrate as for the Proto-Indo-Iranian population…

Now that North-West Indo-European is clearly associated with the Yamna -> Bell Beaker expansion, I understand that his previous model is obsolete and needs a revision.

I find it especially difficult to understand (in light of his previous theory) why he compares Indo-Aryan *vatsa– and Iranian *vasa– to assert that the former is the origin of the loanword in Finno-Ugric, when the Proto-Indo-Iranian form is essentially the same as the Indo-Aryan one, with respect to the *w– evolution into *v– in both PII and late FU dialects…

NOTE: I wrote him yesterday asking for this issue, I will post here his answer.

EDIT (20 MAR 2018): The summary of his answer regarding his selection of Indo-Aryan *vatsa– vs. Iranian *vasa– (instead of just PII *watsa-/vatsa-) is one based on Archaeology (and likley guesstimates), since he understands the split into Iranian and Indo-Aryan to have happened early within the Yamna culture, so that the cultural admixture of Abashevo must have happened after the separation.

netted-ware-parpola
Potential spread of Finnic. “Distribution of the Netted Ware according to Carpelan (2002: 198). A: Emergence of the Netted Ware on the Upper Volga c. 1900 calBC. B: Spread of Netted Ware by c. 1800 calBC. C: Early Iron Age spread of Netted Ware. (After Carpelan 2002: 198 > Parpola 2012a: 151.)

His effort to link the actual expansion of Finno-Ugric to Corded Ware territory, linking it also partially to population movements from the Seima-Turbino phenomenon – probably associated with the initial expansion of N1c lineages – is another good example of convergence of the different anthropological theories thanks to recent Genomic studies.

Related:

Reactionary views on new Yamna and Bell Beaker data, and the newest IECWT model

You might expect some rambling about bad journalism here, but I don’t have time to read so much garbage to analyze them all. We have seen already what they did with the “blackness” or “whiteness” of the Cheddar Man: no paper published, just some informal data, but too much sensationalism already.

Some people who supported far-fetched theories on Indo-European migrations or common European haplogroups are today sharing some weeping and gnashing of teeth around forums and blogs – although, to be fair, neither Olalde et al. (2018) nor Mathieson et al. (2018) actually gave any surprising new data that you couldn’t infer before… People are nevertheless in the middle of the five stages of grief (for whatever expectations they had for new samples), and acceptance will surely take some time.

They will be confronted with two options:

  1. Keep fighting for what they believed, however wrong it turns out to be – after all, we still see all kinds of autochthonous continuists out there, no matter how much data there is against their views. People want to be supporters of a West European origin of R1b-M269 linked to Vasconic languages, fans of R1b-M269 continuity in Central Europe, Uralic speakers who believe in hidden N1c communities in Mesolithic or Neolithic Eastern Europe, fans of the OIT and Indian origins of R1a-M417…
  2. Just accept what seems now clear, change their model, and go on.
wiik-modified
Modified from Wiik for the current autochthonous continuity fans: Vasconic-Uralic distribution and Indo-European folk distribution

For me, the second option sounds quite simple, since whatever happens – markers of Indo-European migration being R1a or R1b, Corded Ware or Bell Beaker, or bothour group’s aim for the past 15 years or so is to support a North-West Indo-European proto-language, so any of the most reasonable anthropological models are a priori compatible with that. My model of Indo-European demic diffusion fits best the most recent proto-language guesstimates, though.

However, I understand that if I had been buying or selling dreams – and I mean literally buying or selling fantasies of whiteness and Europeanness (hidden behind an idealized concept of “Indo-European”, and ancestral components disguised as populations), beginning with the ‘R1a-M417/CWC’ and ‘Yamnaya ancestry’ craze of the 2015 papers – , and I realized data didn’t support that money exchange, I would be frustrated, too.

There is a funny mental process going on here for some of these people, as far as I could read today. Let me review some history of the Indo-European question here before getting to the point:

  1. Firstly linguists reconstructed (and are still doing it) Proto-Indo-European and other ancestral Indo-European proto-languages.
  2. Then archaeologists tried to identify certain ancestral cultures with these actual communities with help from linguistic guesstimates and dialectal classifications,
  3. using anthropological models of migration or cultural diffusion.
  4. Then genetic data came to support one of these alternative anthropological models, if possible.

Now some (amateur) geneticists are apparently disregarding what “Indo-European” means, and why Yamna was considered the best candidate for the expansion of Late Indo-European languages, and question the very sciences of Linguistics and Archaeology as unreliable, instead of questioning their own false assumptions and wrong interpretations from genetic papers.

Really? Genomics (especially ancestral components) now defines what an Indo-European population, culture, and language is? If that is not a fallacy of circular reasoning, I wonder what is.

The modified IECWT model

The surprise today came from the quick reaction of one member of the IECWT workgroup, Guus Kroonen, in his draft Comments to Olalde et al. 2018., The Beaker phenomenon and the genomic, transformation of northwest Europe, Nature.

Allentoft Corded Ware
The IECWT workgroup’s so-called “Steppe model” until today, as presented in Haak et al. (2015).

He and – I can only guess – the whole IECWT workgroup finally rejected their characteristic Corded Ware -> Bell Beaker migration model – which they defended as “The steppe model” of Indo-European migration in Haak et al. 2015. They now defend a proposal similar to Anthony (2007).

Fan fact: Anthony changed his mind recently to partially support what Heyd said in 2007. While I did not dislike Anthony’s new model, I consider it wrong.

The Danish group – unsurprisingly – sticks nevertheless to the hypothesis of some kind of autochthonous Germanic in Scandinavia being defined by Corded Ware migrants and haplogroup R1a, and being somehow special and older among Proto-Indo-European dialects because of its non-Indo-European substrate – although in fact Kroonen’s original linguistic paper didn’t imply so.

While this new change of the workgroup’s model brings it closer to Heyd (2007), and parallels in that sense the adaptation process of Anthony (but always one step behind), what they are proposing right now seems not anymore a modified Kurgan model, as I described it: it is essentially The Kurgan model of Marija Gimbutas (1963), with Bell Beakers spreading a language ancestral to Italo-Celtic, and Corded Ware spreading some kind of mythical Germano-Balto-Slavic

I find it odd that he would not cite Gimbutas, Heyd – as Anthony recently did – , or the most recent paper of Mallory on the language expanded with Bell Beakers, but just the workgroup’s papers and other old ones, to present this “new” theory.

However simple and (obviously) rapidly drafted it was, following the publications in Nature, it does not seem right: They were first, they were right, acknowledge them. Period.

It is interesting how the wrong interpretations of the ‘Yamnaya ancestral component’ (you know, that bulletproof “Yamna R1a-R1b community” and Yamna->Corded Ware migration that never happened) is affecting everyone involved in Indo-European studies.

Related:

Spatio-temporal deixis and cognitive models in early Indo-European

helios-rising

Interesting article, Spatio-temporal deixis and cognitive models in early Indo-European, by Annamaria Bartolotta, Cognitive Linguistics (2018); 29(1):1-44.

Abstract (emphasis mine):

This paper is a comparative study based on the linguistic evidence in Vedic Sanskrit and Homeric Greek, aimed at reconstructing the space-time cognitive models used in the Proto-Indo-European language in a diachronic perspective. While it has been widely recognized that ancient Indo-European languages construed earlier (and past) events as in front of later ones, as predicted in the Time-Reference-Point mapping, it is less clear how in the same languages the passage took place from this ‘archaic’ Time-RP model or non-deictic sequence, in which future events are behind or follow the past ones in a temporal sequence, to the more recent ‘post-archaic’ Ego-RP model that is found only from the classical period onwards, in which the future is located in front and the past in back of a deictic observer. Data from the Rigveda and the Homeric poems show that an Ego-RP mapping with an ego-perspective frame of reference (FoR) could not have existed yet at an early Indo-European stage. In particular, spatial terms of front and behind turn out to be used with reference not only to temporal events, but also to east and west respectively, thus presupposing the existence of an absolute field-based FoR which the temporal sequence is metaphorically related to. Specifically, sequence is relative position on a path appears to be motivated by what has been called day orientation frame, in which the different positions of the sun during the day motivate the mapping of front onto ‘earlier’ and behind onto ‘later’, without involving ego’s ‘now’. These findings suggest that early Indo-European still had not made use of spatio-temporal deixis based on the tense-related ego-perspective FoR found in modern languages.

Featured image, from the article: Helios rising from the sea (blacas red-figured calyx-krater, fifth century B.C., British Museum). Related quote from the article:
“Interestingly, the archeological evidence supports that time could be spatialized along the lateral axis. In ancient Greek art the sun is represented as moving from right to left. Such orientation can be observed, for instance, on the Blacas red-figured calyx-krater of the fifth century B.C. (London, British Museum), where Helios is found at the extreme right of the scene and proceeds to the viewer’s left, following Eos, i.e., the dawn.”

See also:

The myth of mixed language, the concepts of culture core and package, and the invention of ‘Steppe folk’

neolithic-steppe

I recently read some papers which, albeit apparently unrelated, should be of interest for many today.

Mixed language

The myth of the mixed languages, by Kees Versteeg, in Advances in Maltese linguistics, ed. by Benjamin Saade and Mauro Tosco, 217-238. Berlin and New York: Mouton de Gruyter, 2017 [uncorrected proofs]

This paper focuses on the usefulness of the label ‘mixed languages’ as an analytical tool. Section 1 sketches the emergence of the biological paradigm in linguistics and its effect on the contemporary debate about mixed languages. Sections 2 and 3 discuss two processes that have been held responsible for the emergence of mixed languages, code switching and extreme borrowing. Section 4 compares these two mechanisms with the categories of change in Thomason & Kaufman (1988), while Section 5 offers some conclusions about the status of mixed languages as a special category.

Although the paper is a must read for language contact and language change (code-switching, borrowing, shifting), a good summary may save you some time if you are not interested in linguistics:

Speakers may either shift to a new language while retaining traces of their old language, or they may stick to their original language while borrowing from another language with which they come in touch (…)

[Bakker] distinguishes two types of communities in which mixing is found: isolated mixed marriage communities with (asymmetrical) bilingualism; and nomadic communities that shift to a dominant language, but retain a substantial part of their lexicon as a private or secret register, closely connected with the community’s identity. The history of these communities provides us with plausible scenarios to explain the idiosyncrasies of the speech pattern of the speakers belonging to them.

In what way, then, does it help to put both categories under one label of mixed languages? I believe Backus (2003: 263) is right when he suggests that the question of whether a certain set of features constitutes a mixed language is perhaps not a very interesting one. The question should not be whether given certain features a language may be categorized as mixed, but what the linguistic effects of different kinds of contact (trade, work, conquest, mixed marriages, colonization, marginalization, etc.) are, and to what extent these effects correlate with the type of contact. At no point is it necessary to posit a category of mixed languages. In fact, the myth of the mixed languages may have been perpetuated because of the relative weirdness of the initial cases, notably that of Michif, which represent phenomena so unique that it is understandable that some scholars came to believe that they could only be explained by special mechanisms. The position taken here is that if we focus on the speakers’ behavior, the phenomena in question become much more understandable. The crucial point is that languages do not mix, people do.

Culture core and package

The article Isolation-by-distance, homophily, and “core” vs. “package” cultural evolution models in Neolithic Europe, by Shennana, Crema, and Kerig (2015):

Recently there has been growing interest in characterising population structure in cultural data in the context of ongoing debates about the potential of cultural group selection as an evolutionary process. Here we use archaeological data for this purpose, which brings in a temporal as well as spatial dimension. We analyse two distinct material cultures (pottery and personal ornaments) from Neolithic Europe, in order to: a) determine whether archaeologically defined “cultures” exhibit marked discontinuities in space and time, supporting the existence of a population structure, or merely isolation-by-distance; and b) investigate the extent to which cultures can be conceived as structuring “cores” or as multiple and historically independent “packages”. Our results support the existence of a robust population structure comparable to previous studies on human culture, and show how the two material cultures exhibit profound differences in their spatial and temporal structuring, signalling different evolutionary trajectories.

Our results suggest distinct evolutionary histories in the spatial and temporal variation of personal ornament and pottery, with different rates of innovation, patterns of descent, and dynamics of diffusion. Ornament data do show statistically significant values of ΦST using pottery-defined population structures, but the magnitude is extremely small, and partial Mantel tests suggest that much of this pattern is explained by isolation by distance. These results are in line with a model
of culture represented by independent “packages” of multiple coherent units rather than one characterised by a distinct and fairly isolated “core” surrounded by a “periphery” of elements prone to crosscultural transmission. The alternative hypothesis is that one element was part of the “core” tradition, whilst the other was peripheral. This scenario is however less likely given that both elements are generally regarded as expression of local lines of transmission and/or signalling.

The robust support for a population structure in the pottery data shows that some degree of homophily must have biased the transmission process, but this bias was confined within the single “package”, rather than affecting other aspects of the material culture. In other words similarity (or dissimilarity) of pottery style was not influencing the transmission process of personal ornaments and vice-versa. If this was the case, we should have observed a stronger agreement between in the spatio-temporal distribution of the two datasets, a pattern we failed to observe. Personal ornaments are often seen as group-identity markers, but the fact that our study appears to indicate a stronger role for isolation by distance in accounting for variation in ornaments suggests that this assumption may not be valid, or alternatively that these groups cross-cut the archaeological cultures traditionally recognised. Thus,while our study has provided strong evidence of population structure affecting patterns of cultural interaction, in this case at least the distinct patterns observed point to a modular, ‘package’ model. It has also shown that we can identify population structuring from the evidence of the archaeological record without continuing to attempt the fruitless task of correlating its patterns with past ethnolinguistic units.

pottery-ornament-neolithic-europe
Location of material culture (ornament and pottery) data of 361 Neolithic sites in central
Europe

A situation more interesting than Neolithic Europe for those following this blog will probably be the one on West Europe during the Bell Beaker expansion.

Regarding Iberia, I have already talked about the possibilities of “resurgence” after the arrival of Bell Beaker migrants, in this blog and in the Indo-European demic diffusion model.

About the British Isles, you can read e.g. the stepped and different cultural changes that happened after the arrival of East Bell Beakers in What was and what would never be: changing patterns of interaction and archaeological visibility across north-west Europe from 2,500 to 1,500 cal BC, by Wilkin, N., Vander Linden, M. 2015. In : Anderson-Whymark, H., Garrow, D., Sturt, F. (eds)Continental connections. Exploring cross -Channel relationships from the Mesolithic to the Iron Age. Oxford: Oxbow: 99-121.

britain-ireland-bell-beaker
The changing cultural geography of north-west Europe c. 3000-1500 BC. Solid line: architecture; dashed line: material culture;, dash-dot line: funerary practices

This logical description different cultural changes brings up a question obvious to many, and can be summed up by “does the arrival of (North-West Indo-European-speaking) East Bell Beakers mean the end of non-Indo-European cultures in Western and Northern Europe?” The answer is obviously – as in the rest of Europe – quite simply No. Many non-Indo-European groups must have survived the initial expansion of East Bell Beakers in many regions, as the pre-Roman situation (already quite simplified after the expansion of Celtic and Germanic) testifies.

“Resurgence” of local groups as seen in genomic data is the most direct connection with survival of previous non-IE cultures (and thus languages), but obviously not the only mechanism of language survival, since we can see founder effects – such as those seen in modern Basque speakers (mainly of R1b subclades), and in Ugro-Finnic speakers in north-east Europe (mainly of N subclades).

Steppe folk

I am recently stumbling more and more often upon the concept of ‘Steppe people‘ by amateur geneticists, whether in Anthrogenica, in blogs and blog comments, or even in research papers, where people used to talk about ‘Yamnaya people’ or ‘Yamnaya folk’. As I said, I expected this since I questioned the concept of the ‘Yamnaya ancestral component‘, so I feel vindicated by this change.

However, whereas most will be using these names simply following the redeeming term “steppe admixture”, and thus refer to a Neolithic steppe population that shared a common admixture component (probably ca. 5000 BC), some will obviously go further and identify this component with Proto-Indo-European (like that, in general terms), since it is the logical sequence for those who consider the term “Indo-Europeans” as an umbrella for a certain ethnic proportion and a link with modern populations in stupid autochthonous continuity theories, where prehistoric language and culture are irrelevant.

NOTE. Evidently, those who supported quite strongly the fact that R1a-M417 subclades associated with Corded Ware migrants (i.e. mainly R1a-Z645) stemmed directly from Yamna migrants are shifting terminology to “steppe people” in light of recent data, so that they can support an older steppe community in the likely case that no R1a is found in Yamna, while keeping open the possibility to revert to a more direct support of the Yamna -> Corded Ware model in case just one R1a sample is found… So no anthropological model at all here, just a personal desire to be fullfilled in any possible way.

Those thinking naively about an imaginary ‘Steppe folk’, living in loosely connected steppe cultures, speaking a mixed ‘Steppe language’, may well keep inventing potentially popular peoples and languages based on admixture, such as West Hunter-Gatherer folk (Vasconic?), East Hunter-Gatherer folk (Uralic? – or maybe today they can invent a Siberian folk bringing Finno-Ugric after 500 BC), Caucasus Hunter-Gatherer folk (??), etc. So welcome back to the 1930s!… Or was it the 2000s?

wiik-modified
Image modified by me from Wiik’s original, for those of you nostalgic fans of autochthonous continuity theories: You can now again support a native Vasconic-Uralic-Indo-European “folk” distribution in Neolithic Europe.

Some people like to talk about how “Science” wins against Academia, especially when they try to defend this pseudo-subfield they are inventing and venerating on the go to characterize ancient populations, where new genomic methods are king, and the other fields involved are just noise they easily use or dismiss to support their own desires and preconceptions.

NOTE. I felt like many fans of Genomics are very well represented in this recent article at ArXiv: “23andMe confirms: I’m super white” — Analyzing Twitter Discourse On Genetic Testing

Too bad for them. The misuse of this new field might be popular today among certain amateur geneticists, but it will not stand the test of time, similar to how the initial hype around radiocarbon analysis (for Archaeology) or glottochronology (for Linguistics) eventually faded, and they became just another tool among traditional methods. In Science, time puts everything where it belongs.

Whether you like it or not, Indo-European (and Uralic) questions will be solved – as they have been for a long time now, there is nothing new under the sun – with Historical Linguistics first, then Prehistoric Archaeology, then Anthropology (models of migration, cultural diffusion, founder effects, etc.), and only then Genomics, which may (or may not) help solve certain controversial aspects, by supporting one or other anthropological model. Period.

See also:

The Indo-European demic diffusion model, and the “R1b – Indo-European” association

yamna_bell_beaker_cut

Beginning with the new year, I wanted to commit myself to some predictions, as I did last year, even though they constantly change with new data.

I recently read Proto-Indo-European homelands – ancient genetic clues at last?, by Edward Pegler, which is a good summary of the current state of the art in the Indo-European question for many geneticists – and thus a great example of how well Genetics can influence Indo-European studies, and how badly it can be used to interpret actual cultural events – although more time is necessary for some to realize it. Notice for example the distribution of ‘Yamnaya’ in 3000 BC, all the way to Latvia (based on the initial findings of Mathieson et al. 2017), and the map of 2000 BC with ‘Corded Ware’, both suggesting communities linked by admixture and unrelated to actual cultures.

Some people – especially those interested in keeping a simplistic picture of Europe, either divided into admixture groups or simplistic R1b-Vasconic / R1a-Indo-European / N1c-Uralic (or any combination thereof) – want (others) to believe that I am linking ‘Indo-Europeans’ with haplogroup R1b. That is simply not true. In fact, my model dismisses such simplistic identifications of the reconstructible proto-languages with any modern peoples, admixtures, or haplogroups.

vasconic-uralic
Simplistic Vasconic/R1b-Uralic/N1c distribution, and intruding Indo-European/R1a, according to Wiik.

The beauty of the model lies, therefore, precisely in that if you take any modern group speaking Indo-European languages, none can trace back their combination of language, admixture, and/or haplogroup to a common Indo-European-speaking people. All our ancestral lines have no doubt changed language families (and indeed cultures), they have admixed, and our European regions’ paternal lines have changed, so that any dreams of ‘purity’ or linguistic/cultural/regional continuity become absurd.

That conclusion, which should be obvious to all, has been denied for a long time in blogs and forums alike, and is behind the effort of many of those involved in amateur genetics.

Main linguistic aim

The main consequence of the model, as the title of the paper suggests, is that reconstructible Indo-European proto-languages expanded with people, i.e. with actual communities, which is what we can assert with the help of Genomics. From a personal (or ethnic, or political) point of view genomics is useless, but from an anthropological (and thus linguistic) point of view, genomics can be a very useful tool to decide between alternative models of language diffusion, which has given lots of headaches to those of us involved in Indo-European studies.

The demic diffusion theory for the three main stages of the proto-language expansion was originally, therefore, a dismissal of impossible-to-prove cultural diffusion models for the proto-language – e.g. the adoption of Late Proto-Indo-European by Corded Ware groups due to a patron-client relationship (as proposed by Anthony), or a long-lasting connection between cultures (as proposed by Kristiansen, and favoured by “constellation analogy” proponents like Clackson, who negated the existence of common proto-languages). It also means the acceptance of the easiest anthropological model for language change: migration and – consequently – replacement.

By the time of the famous 2015 papers, I had been dealing for some time with the idea that the shared features between Indo-Iranian and Balto-Slavic may have been due to a common substrate, and must have therefore had some reflection in genomic finds. The data on these papers, and the addition of a weak connection between Pre-Germanic and Balto-Slavic communities, together with their clearest genetic link – R1a-M417 subclades (especially European Z283) – made it still easier to propose a Corded Ware substrate, partially common to the three.

Allentoft Corded Ware
Allentoft et al. “Arrows indicate migrations — those from the Corded Ware reflect the evidence that people of this archaeological culture (or their relatives) were responsible for the spreading of Indo-European languages. All coloured boundaries are approximate.”

Before the famous 2015 papers (and even after them, if we followed their interpretation), we were left to wonder why the supposed vector of expansion of Indo-European languages, Corded Ware migrants – represented by R1a-Z645 subclades, and supposedly continued unchanged into modern populations in its ‘original’ ancestral territories, Balto-Slavic and Indo-Iranian – , were precisely the (phonetically) most divergent Indo-European languages – relative to the parent Late Indo-European proto-language.

My paper implied therefore the dismissal of an unlikely Indo-Slavonic group, as proposed by Kortlandt, and of a still less factible Germano-Slavonic, or Germano-Indo-Slavonic (?) group, as loosely implied by some in the past, and maybe supported in certain archaeological models (viz. Kristiansen or partially Anthony), and presently by some geneticists since their simplistic 2015 papers on “massive migrations from the steppe“, and amateur genetic fans with infinite pet theories, indeed.

A common Corded Ware substrate to Balto-Slavic and Indo-Iranian, and common also partially between Balto-Slavic and Germanic (as supported by Kortlandt, too, albeit with different linguistic connotations), would explain their common features. The Corded Ware culture (and Uralic, tentatively proposed by me as the group’s main language family) is a strong potential connection between them, further supported by phylogeography, too.

Other consequences

Interpretations in my paper help thus dismiss the simplistic Yamna -> Corded Ware -> Bell Beaker migration model implied with phylogeography in the 2000s, and revived again by geneticists and Kristiansen’s workgroup based on the famous 2015 papers, whereby – due to the “Yamnaya ancestral component” – the Yamna culture would have been composed of communities of R1a-M417 and R1b-M269 lineages which remained against all odds ‘related but separated’ for more than two thousand years, sharing a common unitary language (why? and how?), and which expanded from Yamna (mainly R1b-L23) into Corded Ware (mainly R1a-M417) and then into Bell Beaker (mainly R1b-L51), in imaginary migration waves whose traces Archaeology has not found, or Anthropology described, before.

While phylogeography (especially the distribution of ancient samples of certain R1b and R1a subclades) was the main genetic aspect I used in combination with Archaeology and Anthropology to challenge the reliability of the “Yamnaya ancestral component” in assessing migrations – and thus Kristiansen’s now-popular-again modified Kurgan model – , my main aim was to prove a recent expansion of Late Proto-Indo-European from the steppe, and a still more recent expansion of a common group of speakers of North-West Indo-European, the language ancestral to Italo-Celtic, Germanic, and probably Balto-Slavic (or ‘Temematic’, the NWIE substrate of Balto-Slavic, according to some linguists).

My arguments serve for this purpose, and modern distributions of haplogroups or admixture are fully irrelevant: I am ready to change my view at any time, regarding the role of any haplogroup, or ancestral component, archaeological data, or anthropological migration model, to the extent that it supports the soundest linguistic model.

proto-indo-european-stages
Stages of Proto-Indo-European evolution. IU: Indo-Uralic; PU: Proto-Uralic; PAn: Pre-Anatolian; PToch: Pre-Tocharian; Fin-Ugr: Finno-Ugric. The period between Balkan IE and Proto-Greek could be divided in two periods: an older one, called Proto-Greek (close to the time when NWIE was spoken), probably including Macedonian, and spoken somewhere in the Balkans; and a more recent one, called Mello-Greek, coinciding with the classically reconstructed Proto-Greek, already spoken in the Greek peninsula (West 2007). Similarly, the period between Northern Indo-European and North-West Indo-European could be divided, after the split of Pre-Tocharian, into a North-West Indo-European proper, during the expansion of Yamna to the west, and an Old European period, coinciding with the formation and expansion of the East Bell Beaker group.

Gimbutas’ old theory of sudden and recent expansion served well to support a real community of Proto-Indo-European speakers, as did later the Yamna -> Corded Ware -> Bell Beaker theory that circulated in the 2000s based on modern phylogeography, and as did later partially Anthony’s updated steppe theory (2007). On the other hand, Kristiansen’s long-lasting connections among north-west Pontic steppe cultures and Globular Amphorae and Trypillian cultures, did not fit well with a close community expanding rapidly – although recent genetic data on Trypillia and Globular Amphorae might be compelling him to improve his migration theory.

So, if data turns out to be not as I expect now, I will reflect that in future versions of the paper. I have no problem saying I am wrong. I have been wrong many times before, and something I am certain is that I am wrong now in many details, and I am going to be in the future.

If, for example, R1b-L23(xZ2105) is demonstrated to come from Hungary and not the steppe (as supported by Balanovsky) or R1a-M417 samples are proved to have expanded with West Yamna settlers (as recently proposed by Anthony, see below the Balto-Slavic question), I would support the same model from a linguistic point of view, but modified to reflect these facts. Or if a direct migration link is found in Archaeology from Yamna to Corded Ware, and from Corded Ware to Bell Beaker (as proposed in the 2015 papers), I will revise that too (again, see the image below). Or, if – as Lazaridis et al. (2017) paper on Minoans and Mycenaeans suggested – the Anatolian hypothesis (that is, one of the multiple ones proposed) turns out to be somehow right, I will support it.

calcolithic-expansion
My map of Late Proto-Indo-European expansion (A Grammar of Modern Indo-European, 2006), following Gimbutas and Mallory.

Haplogroups are the least important aspect of the whole model, they are just another data that has to be taken into account for a throrough explanation of migrations. It has become essential today because of the apparent lack of vision on the part of geneticists, who failed to use them to adjust their findings of admixture with findings of haplogroup expansions, favouring thus a marginal theory of long-lasting steppe expansion instead of the mainstream anthropological models.

Since many of these alternative scenarios seem less and less likely with each new paper, it is probably more efficient to talk about which developments are most likely to challenge my model.

Main points

My main predictions – based mostly on language guesstimates, archaeological cultures, and anthropological models of migration -, even with the scarce genomic data we had, have been proven right until know with new samples from Mathieson et al. (2017) and Olalde et al. (2017), among other papers of this past year. These were my original assumptions:

(1) A Middle Proto-Indo-European expansion defined by the appearance of steppe ancestry + reduction in haplogroup diversity and expansion of (mainly) R1b-M269 and R1b-L23 lineages;

(2) A Late Proto-Indo-European expansion defined by steppe ancestry + reduction in haplogroup diversity and expansion of (mainly) R1b-L23 subclades; and

(3) A North-West Indo-European expansion defined by steppe ancestry + reduction in haplogroup diversity and expansion of (mainly) R1b-L51 subclades.

The expansion of Corded Ware peoples, associated with steppe ancestry + reduction in haplogroup diversity and expansion of (mainly) R1a-Z645 subclades, represents thus a different migration, which is compatible with the different nature of the Corded Ware culture, unrelated to Yamna and without migration waves from one to the other (although there were certainly contacts in neighbouring regions).

As you can see, neither of the 3+1 expansion models imply that no other haplogroup can be found in the culture or regions involved (others have in fact been found, and still the models remain valid): these migrations imply a reduction of haplogroup diversity, and the expansion of certain subclades as is common in population expansions throughout history. While we all accept this general idea, some people have difficulties accepting just those cases not compatible with their dreams of autochthonous continuity.

Nevertheless, there are still voids in genetic investigation.

Controversial aspects

In my humble opinion, these are potential conflict periods and the most likely areas of change for the future of the theory:

1. When and how did R1b-M269 lineages become “chiefs” in the steppe?

Based on scarce data from Khvalynsk, it seems that during the Neolithic there were many haplogroups in the North Pontic and North Caspian steppes. A reduction to R1b-M269 subclades must have happened either just before or (as I support) during (the migrations that caused) the Suvorovo-Novodanilovka expansion among Sredni Stog, probably coinciding also with the expansion (or one of the expansions) of CHG ancestry (and thus the appearance of ‘Steppe component’ in the steppe). My theory was based initially on Anthony’s account and TMRCA of haplogroups of modern populations (both ca. 4200-4000 BC), but recent samples of the Balkans (R1b-M269 and steppe ancestry) seem to trace the population expansion some centuries back.

If my assessment is correct, then modern populations of haplogroup R1b-M269* and R1b-L23* in the Balkans probably reflect that ancient expansion, and samples related to Proto-Anatolian cultures in the Balkans will most likely be of R1b-M269 subclades and R1b-L23*. After admixture in the Balkans, posterior migrations of Anatolian languages into Anatolia might be associated with a different admixture component and haplogroups, we don’t have enough data yet.

If the haplogroup reduction and expansion in Khvalynsk happened later than the Suvorovo-Novodanilovka expansion, then we might find the expansion of Pre- or Proto-Anatolian associated with many different haplogroups, such as R1b (xM269), R1a, I, J, or G2, and more or less associated with steppe ancestry in the Balkans.

Another reason for finding such variety of haplogroups in ancient samples from the Balkans would be that this Khvalynsk group of “chiefs” traversed – and mixed with – the Sredni Stog population. Nevertheless, if we suppose homogeneity in haplogroups in Khvalynsk during the expansion, a high proportion of different haplogroups explained by admixture with the local population of Sredni Stog would challenge the whole “chief domination” explanation by Anthony, and we would have to return to the “different culture” theory by Rassamakin and potentially an older migration from Khvalynsk. In any case, both researchers show clear links of the Suvorovo-Novodanilovka phenomenon to Khvalynsk, and a differentiation with the surrounding Sredni Stog culture.

A less likely model would support the identification of the whole Eneolithic Pontic-Caspian steppe as a loose Indo-Hittite-speaking community, which would be in my opinion too big a territory and too loose a cultural bond to justify such a long-lasting close linguistic connection. This will probably be the refuge of certain people looking desperately for R1a-IE connections. However, the nature of the western steppe will remain distinct from Late Proto-Indo-European, which must have developed in the Yamna culture, so autochthonous continuity is not on the table anymore, in any case…

suvorovo-novodanilovka-region
Coexistence of the Varna-Gumelniţa culture and the Suvorovo phase of the sceptre-bearer communities. 1 — Fălciu; 2 — Fundeni-Lungoţi; 3 — Novoselskaja; 4 — Suvorovo; 5 — Casimcea; 6 — Kjulevča; 7 — Reka Devnja; 8 — Drama; 9 — Gonova mogila; 10 — Reževo; 11 — geographically separate Decea variant of the sceptre bearer group (after Govedarica, Manzura 2011: Abb. 5, adapted).

2. How did R1a-M417 (and especially R1a-Z645) haplogroups came to dominate over the Corded Ware cultures?

If I am right (again, based on TMRCA of modern populations), then it is precisely at the time of the potential expansion of Proto-Corded Ware from the Dnieper-Dniester forest, forest-steppe, and steppe regions, ca 3300-3000. Furholt’s recent radiocarbon analysis and suggestions of a Lesser Poland origin of the third or A-horizon, on which disparate archaeologists such as Anthony or Klejn rely now, seem to suggest also that Corded Ware was a cultural complex rather than a compact culture reflecting a migration of peoples – similar thus to the Bell Beaker complex.

This cultural complex interpretation of Corded Ware contrasts with the quite homogeneous late samples we have, suggesting clear migration waves in northern Europe, at least at some point in time, so Genomics will be a great tool to ascertain when and from where approximately did Corded Ware peoples expand. Right now, it seems that Eneolithic Ukraine populations are the closest to its origin, so the traditional interpretation of its regional origin by Kristiansen or Anthony remains valid.

3. How was Indo-Iranian adopted by Corded Ware invaders?

This is rather an anthropological question. We need reasonable models of founder effect/cultural diffusion necessary for that to happen – similar to the ones necessary to explain the arrival of N1c subclades into north-east Europe, or the arrival of R1b subclades in Basque/Iberian-speaking regions in south-west Europe. My description of potential events in the eastern steppe – based partially on Anthony – is merely a short sketch. Genomic data is unlikely to offer more than it does today (replacement of haplogroups, and gradually of some steppe component, by late Corded Ware groups in the steppe), but let’s see what new samples can contribute.

As for what some Indians – and other people willing to confront them – are looking for, regarding R1a-M417 and/or Indo-European origins in India, I don’t see the point, we already know a) that the origin of the expansion is in the steppe and b) that Hindu nationalist biggots will not accept results from research that oppose their views. I don’t expect huge surprises there, just more fruitless discussions (fomented by those who live from trolling or conspiracies)…

4. Yamna settlers from Hungary

Anthony’s new theory – and the nature of Balto-Slavic – hinges on the presence of R1a-M417 subclades (associated with later Corded Ware samples) in Yamna settlers of Hungary, potentially originally from the North Pontic area, where the oldest sample has been found.

My ‘modified’ version of Anthony’s new model (the only I deem just remotely factible) includes the expansion of a Proto-Corded Ware from Lesser Poland, but (given the overwhelming R1b found in East Bell Beaker), with R1a-M417 being associated with the region. How to explain this language change with objective data? Well, we have Bell Beaker expanding to these areas at a later time, so we would need to find R1b-L23 settlers in Lesser Poland, and then a resurge of R1a-M417 haplogroup. If not, resorting yet again to cultural diffusion Yamna “patrons” to Corded Ware “clients” of Lesser Poland would bring us to square one, now with the ‘steppe ancestry’ controversy included…

Since some Eastern Europeans are (for no obvious reason whatsoever) putting their hopes on that IE-R1a-CWC association, let’s hope some samples of R1a-M417 in Yamna or Hungary give them a break, so that they can begin accepting something closer to mainstream anthropological models. We could then work from there a Yamna-> Bell Beaker / North-West Indo-European association truce, and from there keep accepting that no single haplogroup from Yamna settlers is linked with modern languages, cultures or ethnic groups.

yamna-region
localization of Central-European funerary monuments with elements of the Pit Grave culture (after Bátora 2006);

5. How and when was Balto-Slavic associated with haplogroup R1a?

If we accept the Southern or Graeco-Aryan nature of Balto-Slavic with influence from an absorbed North-West Indo-European dialect, “Temematic” (as Kortlandt does), then Indo-Slavonic adopted in the steppe from Potapovka by Sintashta and Poltavka populations divided ca. 2000 BC into Indo-Iranian (migrating to the east with Andronovo), and Balto-Slavic (migrating westward with the Srubna culture). History from there is not straightforward, and it should follow Srubna, Thraco-Cimmerian, or other late expansions from cultures of the steppe.

On the other hand, if it is a Northern dialect related closely to Germanic and Italo-Celtic (in a North-West Indo-European group), then its origin has to be found in the initial expansion of East Bell Beakers, and its development into either the Únětice culture (of Balkan and thus potentially “Southern IE” influence), or the Mierzanowice-Nitra culture (of Corded Ware and thus potentially Uralic influence), or maybe from both, given the intermediate substrate found in Germanic and Balto-Slavic.

It is my opinion that the association of Balto-Slavic with haplogroup R1a is quite early after the East Bell Beaker expansion, probably initially with the subclade typically associated with West Slavic, R1a-M458. I have not much data to support this (apart from the most common linguistic model), just modern haplogroup distribution maps and common TMRCA, and highly hypothetical archaeological-anthropological models. Genetics will hopefully bring more data.

Let’s see also what information on ancient haplogroups we can obtain from the Tollense valley (already showing a close cluster with modern West Slavic populations) and steppe regions.

6. How did Germanic, Celtic, and Italic expand?

Germanic is probably the most interesting one. Following the expansion of R1b-L51 subclades (especially R1b-U106) and steppe ancestry (a confounding factor, with the previous expansion of R1a-Z284 subclades) in Scandinavia is going to be fascinating. Anthropological models already point to a linguistic and archaeological expansion of Pre-Germanic with Bell Beaker peoples.

The expansion of Celtic seems to be associated with chiefdoms, untraceable today in terms of haplogroups, and it seems thus different from previous expansions. New studies might tell how that happened, if it was actually in successive ways, as proposed, or maybe we don’t have enough data yet to reach conclusions.

We don’t know either how Italic expanded into the Italian Peninsula, or whether Latin expanded with peoples from Italy, if at all, or it was mostly a cultural diffusion event, as it seems.

Regarding Etruscan, while I think it is a controversy initiated based on fantastic accounts, and ignited with few finds of Middle Eastern ancestry (that seem logical from the point of view of regional contacts), it will be important for Italian linguists and archaeologists, also to accept the most likely scenario.

As for Palaeo-Hispanic languages, while steppe ancestry is found quite reduced in R1b-L51 subclades (after so many different expansions and admixture events since the departure from the steppe), their distribution from the Chalcolithic onwards and the resurgence of native haplogroups may serve to ascertain which Pre-Roman tribes were associated with the oldest regions where these subclades dominated. For that aim, a closer look at the developments in Aquitania and other pre-Roman Vasconic- and Iberian-speaking regions may shed some light on how founder effects might develop to leave the native language intact (in a case similar to the adoption of Indo-Iranian by post-Corded Ware Sinthastha and Potapovka in the eastern Pontic-Caspian steppe).

NOTE: Although mostly unrelated, linguistic questions may also be somehow altered with a change of migration models. For example, our current Corded Ware Substrate Hypothesis – strongly contested by Kortlandt and others – implies that Uralic was potentially the language spoken by Eneolithic Ukraine / Proto-Corded Ware peoples, therefore early Uralic languages were spoken by Corded Ware peoples, as a substrate for Germanic and Balto-Slavic, and Balto-Slavic and Indo-Iranian. If an Indo-Hittite branch different from Late PIE is accepted for Eneolithic Ukraine (thus suggesting a millennia-long cultural-historical community in the steppe), then the model still stands (e.g. Ger. and BSl. *-mos/-mus, as stated by Kortlandt, would correspond to the oldest morphological IE layer). As you can read in the different versions of our model, the different possibilities for the common substrate are stated, and the most likely one selected. But the most likely a priori option sometimes turns out to be wrong…

NOTE 2: You can comment whatever you want here, but I opened a specific thread in our forum if you want serious comments on the model to stuck and be further discussed.

Featured images: from the book Interactions, changes and meanings. Essays in honour of Igor Manzura on the occasion of his 60th birthday. Țerna S., Govedarica B. (eds.). 2016. Kishinev: Stratum Plus.

See also:

Prehistoric loan relations: Foreign elements in the Proto-Indo-European vocabulary

ancient-indo-european-world-fantasy

An interesting ongoing web project, Prehistoric loan relations, on potential loans of Proto-Indo-European words, from Uralic-Yukaghir, Caucasian, and Middle Eastern influence.

Based on a Ph.D. thesis by Bjørn (2017) Foreign elements in the Proto-Indo-European vocabulary (PDF).

From the website (emphasis mine):

This page allows historical linguists to compare and scrutinize proposed prehistoric lexical borrowings from the perspective of Proto-Indo-European. The first entries are all (135 in total) extracted from my master’s thesis “Foreign elements in the Proto-Indo-European vocabulary” (Bjørn 2017). Comments are encouraged at the bottom of each entry. New entries will be added, also on request.

Take this not as the conclusion, but an invitation to join the conversation.

So, we welcome the invitation, and hope that this new project thrives.

Also, I loved his fantasy-like map of the central Eurasian region (featured image on this post).

Related: