An interesting aspect of the paper, hidden among so many relevant details, is a clearer picture of how the so-called Yamnaya or steppe ancestry evolved from Samara hunter-gatherers to Yamna nomadic pastoralists, and how this ancestry appeared among Proto-Corded Ware populations.
Please note: arrows of “ancestry movement” in the following PCAs do not necessarily represent physical population movements, or even ethnolinguistic change. To avoid misinterpretations, I have depicted arrows with Y-DNA haplogroup migrations to represent the most likely true ethnolinguistic movements. Admixture graphics shown are from Wang et al. (2018), and also (the K12) from Mathieson et al. (2018).
1. Samara to Early Khvalynsk
The so-called steppe ancestry was born during the Khvalynsk expansion through the steppes, probably through exogamy of expanding elite clans (eventually all R1b-M269 lineages) originally of Samara_HG ancestry. The nearest group to the ANE-like ghost population with which Samara hunter-gatherers admixed is represented by the Steppe_Eneolithic / Steppe_Maykop cluster (from the Northern Caucasus Piedmont).
Steppe_Eneolithic samples, of R1b1 lineages, are probably expanded Khvalynsk peoples, showing thus a proximate ancestry of an Early Eneolithic ghost population of the Northern Caucasus. Steppe_Maykop samples represent a later replacement of this Steppe_Eneolithic population – and/or a similar population with further contribution of ANE-like ancestry – in the area some 1,000 years later.
This is what Steppe_Maykop looks like, different from Steppe_Eneolithic:
NOTE. This admixture shows how different Steppe_Maykop is from Steppe_Eneolithic, but in the different supervised ADMIXTURE graphics below Maykop_Eneolithic is roughly equivalent to Eneolithic_Steppe (see orange arrow in ADMIXTURE graphic above). This is useful for a simplified analysis, but actual differences between Khvalynsk, Sredni Stog, Afanasevo, Yamna and Corded Ware are probably underestimated in the analyses below, and will become clearer in the future when more ancestral hunter-gatherer populations are added to the analysis.
2. Early Khvalynsk expansion
We have direct data of Khvalynsk-Novodanilovka-like populations thanks to Khvalynsk and Steppe_Eneolithic samples (although I’ve used the latter above to represent the ghost Caucasus population with which Samara_HG admixed).
We also have indirect data. First, there is the PCA with outliers:
Second, we have data from north Pontic Ukraine_Eneolithic samples (see next section).
Third, there is the continuity of late Repin / Afanasevo with Steppe_Eneolithic (see below).
3. Proto-Corded Ware expansion
It is unclear if R1a-M459 subclades were continuously in the steppe and resurged after the Khvalynsk expansion, or (the most likely option) they came from the forested region of the Upper Dnieper area, possibly from previous expansions there with hunter-gatherer pottery.
Supporting the latter is the millennia-long continuity of R1b-V88 and I2a2 subclades in the north Pontic Mesolithic, Neolithic, and Early Eneolithic Sredni Stog culture, until ca. 4500 BC (and even later, during the second half).
Only at the end of the Early Eneolithic with the disappearance of Novodanilovka (and beginning of the steppe ‘hiatus’ of Rassamakin) is R1a to be found in Ukraine again (after disappearing from the record some 2,000 years earlier), related to complex population movements in the north Pontic area.
NOTE. In the PCA, a tentative position of Novodanilovka closer to Anatolia_Neolithic / Dzudzuana ancestry is selected, based on the apparent cline formed by Ukraine_Eneolithic samples, and on the position and ancestry of Sredni Stog, Yamna, and Corded Ware later. A good alternative would be to place Novodanilovka still closer to the Balkan outliers (i.e. Suvorovo), and a source closer to EHG as the ancestry driven by the migration of R1a-M417.
The first sample with steppe ancestry appears only after 4250 BC in the forest-steppe, centuries after the samples with steppe ancestry from the Northern Caucasus and the Balkans, which points to exogamy of expanding R1a-M417 lineages with the remnants of the Novodanilovka population.
4. Repin / Early Yamna expansion
We don’t have direct data on early Repin settlers. But we do have a very close representative: Afanasevo, a population we know comes directly from the Repin/late Khvalynsk expansion ca. 3500/3300 BC (just before the emergence of Early Yamna), and which shows fully Steppe_Eneolithic-like ancestry.
Compared to this eastern Repin expansion that gave Afanasevo, the late Repin expansion to the west ca. 3300 BC that gave rise to the Yamna culture was one of colonization, evidenced by the admixture with north Pontic (Sredni Stog-like) populations, no doubt through exogamy:
This admixture is also found (in lesser proportion) in east Yamna groups, which supports the high mobility and exogamy practices among western and eastern Yamna clans, not only with locals:
We don’t have a comparison with Ukraine_Eneolithic or Corded Ware samples in Wang et al. (2018), but we do have proximate sources for Abashevo, when compared to the Poltavka population (with which it admixed in the Volga-Ural steppes): Sintashta, Potapovka, Srubna (with further Abashevo contribution), and Andronovo:
The two CWC outliers from the Baltic show what I thought was an admixture with Yamna. However, given the previous mixture of Eneolithic_Steppe in north Pontic steppe-forest populations, this elevated “steppe ancestry” found in Baltic_LN (similar to west Yamna) seems rather an admixture of Baltic sub-Neolithic peoples with a north Pontic Eneolithic_Steppe-like population. Late Repin settlers also admixed with a similar population during its colonization of the north Pontic area, hence the Baltic_LN – west Yamna similarities.
NOTE. A direct admixture with west Yamna populations through exogamy by the ancestors of this Baltic population cannot be ruled out yet (without direct access to more samples), though, because of the contacts of Corded Ware with west Yamna settlers in the forest-steppe regions.
A similar case is found in the Yamna outlier from Mednikarovo south of the Danube. It would be absurd to think that Yamna from the Balkans comes from Corded Ware (or vice versa), just because the former is closer in the PCA to the latter than other Yamna samples. The same error is also found e.g. in the Corded Ware → Bell Beaker theory, because of their proximity in the PCA and their shared “steppe ancestry”. All those theories have been proven already wrong.
NOTE. A similar fallacy is found in potential Sintashta→Mycenaean connections, where we should distinguish statistically that result from an East/West Yamna + Balkans_BA admixture. In fact, genetic links of Mycenaeans with west Yamna settlers prove this (there are some related analyses in Anthrogenica, but the site is down at this moment). To try to relate these two populations (separated more than 1,000 years before Sintashta) is like comparing ancient populations to modern ones, without the intermediate samples to trace the real anthropological trail of what is found…Pure numbers and wishful thinking.
Marital structure. The intensity of interethnic marriages puts the existence of the Ulchi population at risk. The colorful ethnic composition of the Ulchi settlements is reflected in the marriage structure [see featured image]. We found that the proportion of single-ethnic marriages of the Ulchi is on average 51%. The greatest number of such marriages takes place in the village of Bulava. Marriages of Ulchi with Russians are in second place. Marriages with indigenous peoples of the Far East, Nanais, Nivkhs, Evenks, and others, are in third place. Thus, almost half of the Ulchi marriages are with representatives of other nationalities. Such a significant level of interethnic mixing makes it possible to talk about intense processes of assimilation of this indigenous people and puts to the forefront the problem of loss of the unique gene pool of the Ulchi.
Haplogroup C (its branch M48) was genotyped for its five subbranches with markers M86, B470, F13686, B93, and the marker at position 16645386 (GRCh37), which was found by our team for the first time. Variant B93 is rare in the Ulchi, and 14 samples (that is, more than a quarter of the entire gene pool of the Ulchi, Fig. 2) belong to M86 and its subvariants. Therefore, we genotyped STR markers of C-M86 carriers for the Ulchi and neighboring Amur populations and analyzed the relationships of detected haplotypes on the phylogenetic network (Fig. 3, STR haplotypes are available from authors upon request).
(…) On the network, different clusters are associated with different populations: most Mongols belong to F13686, all Evenks of the Amur River region with this haplogroup form a subcluster within F13686, and part of Upper Nanais is the basis of cluster B470.
An estimate of the age of the entire haplogroup C-F12355 obtained from the data of genome-wide sequencing of seven specimens is 2400 ± 500 years (O.P. Balanovsky, unpublished data). That is, the common ancestor of all the studied representatives of various peoples with this haplogroup lived not so long ago, the first millennium BC. The formation time of cluster F13686 is somewhat later: 1990 ± 600 years.
(…) obvious traces of the interaction of the gene pool of the Ulchi with neighboring and remote peoples of the Far East and Central Asia in the time range of the last one to three thousand years were revealed. This shows that the results of work  on the similarity of the gene pool of the ancient (age of 7500 years) Neolithic genomes of the Amur River region to the Ulchi probably indicate not the uniqueness of the Ulchi, but the fact that this ancient gene pool was preserved in a vast circle of populations of the Far East interwoven with gene flows both with each other and, to a lesser extent, with populations of Central Asia.
The expansion of C2b1a2a-M86 (among many basal C2-M217 samples) is thus possibly associated with the spread of Tungusic, which puts C2b1a at the root of the Micro-Altaic expansion, with a formation date ca. 12700 BC, TMRCA 12500 BC (and not only Mongolian). This shows that Micro-Altaic is connected with a local population which shows a clear continuity since at least 3500 BC. This, however, tells us little about the origin of the language.
That leaves the ancestral N lineages found among Far East Asians as Palaeo-Siberian in origin, and their late expansions to the west not particularly linked with any of the known Palaeo-Siberian ethnolinguistic groups, let alone a supposed “Uralo-Altaic” language…
It has been known for a long time that the Caucasus must have hosted many (at least partially) isolated populations, probably helped by geographical boundaries, setting it apart from open Eurasian areas.
David Reich writes in his book the following about India:
The genetic data told a clear story. Around a third of Indian groups experienced population bottlenecks as strong or stronger than the ones that occurred among Finns or Ashkenazi Jews. We later confirmed this finding in an even larger dataset that we collected working with Thangaraj: genetic data from more than 250 jati groups spread throughout India (…)
Rather than an invention of colonialism as Dirks suggested, long-term endogamy as embodied in India today in the institution of caste has been overwhelmingly important for millennia. (…)
The Han Chinese are truly a large population. They have been mixing freely for thousands of years. In contrast, there are few if any Indian groups that are demographically very large, and the degree of genetic differentiation among Indian jati groups living side by side in the same village is typically two to three times higher than the genetic differentiation between northern and southern Europeans. The truth is that India is composed of a large number of small populations.
There is little doubt now, based on findings spanning thousands of years, that the Mesolithic and Neolithic Caucasus hosted various very small populations, even if the ancestral components may be reduced to the few known to date (such as ANE, EHG, AME*, ENA, CHG, and other “deep” ancestral components).
NOTE. I will call the ancestral component of Dzudzuana/Anatolian hunter-gatherers Ancient Middle Easterner (AME), to give a clear idea of its likely extension during the Late Upper Palaeolithic, and to avoid using the more simplistic Dzudzuana, unless it is useful to mention these specific local samples.
Genetic labs have a strong fixation with ancestry. I guess the use of complex statistical methods gives professionals and laymen alike the feeling of dealing with “Science”, as opposed to academic fields where you have to interpret data. I think language reveals a lot about the way people think, and the fact that ancestral components are called ‘lineages’ – while not wrong per se – is a clear symptom of the lack of interest in the true lineages: Y-DNA haplogroups.
It has become quite clear that male-biased migrations are often the ones which can be confidently followed for actual population movements and ethnolinguistic identification, at least until the Iron Age. The frequently used Palaeolithic clusters offer a clear example of why ancestry does not represent what some people believe: They merely give a basic idea of sizeable population replacements by distant peoples.
Both concepts are important: sizeable and distant peoples. For example, during the Upper Palaeolithic in Europe there was a sizeable population replacement of the Aurignacian Goyet cluster by the Gravettian Vestonice cluster (probably from populations of far eastern Russia) coupled with the arrival of haplogroup I, although during the thousands of years that this material culture lasted, the previously expanded C1a2 lineages did not disappear, and there were probably different resurgence and admixture events.
Haplogroup I certainly expanded with the Gravettian culture to Iberia, where the Goyet ancestry did not change much – probably because of male-driven migrations -, to the extent that during the Magdalenian expansions haplogroup I expanded with an ancestry closer to Goyet, in what is called a ‘resurge’ of the Goyet cluster – even though there is a clear replacement of male lines.
The Villabruna (WHG) cluster is another good example. It probably spread with haplogroup R1b-L754, which – based on the extra ‘East Asian’ affinity of some samples and on modern samples from the Middle East – came probably from the east through a southern route, and not too long before the expansion of WHG likely from around the Black Sea, although this is still unclear. The finding of haplogroup I in samples of mostly WHG ancestry could confuse people that do not care about timing, sub-structured populations, and gene flow.
NOTE. If you don’t understand why ‘clusters’ that span thousands of years don’t really matter for the many Palaeolithic population expansions that certainly happened among hunter-gatherers in Europe, just take a look at what happened with Bell Beakers expanding from Yamna into western Europe within 500 years.
If we don’t thread carefully when talking about population migrations, these terms are bound to confuse people. Just as the fixation on “steppe ancestry” – which marks the arrival in Chalcolithic Europe of peoples from the Pontic-Caspian region – has confused a lot of researchers to this day.
When I began to write about the Indo-European demic diffusion model, my concern was to find a single spot where a North-West Indo-European proto-language could have expanded from ca. 2000 BC (our most common guesstimate). Based on the 2015 papers, and in spite of their conclusions, I thought it had become clear that Corded Ware was not it, and it was rather Bell Beakers. I assumed that Uralic was spoken to the north (as was the traditional belief), and thus Corded Ware expanded from the forest zone, hence steppe ancestry would also be found there with other R1a lineages.
With the publication of Mathieson et al. (2017) and Olalde et al. (2017), I changed my mind, seeing how “steppe ancestry” did in fact appear quite late, hence it was likely to be the result of very specific population movements, probably directly from the Caucasus. Later, Mathieson published in a revision the sample from Alexandria of hg R1a-M417 (probably R1a-Z645, possibly Z93+), which further supported the idea that the migration of Corded Ware peoples started near the North Pontic forest-steppe (as I included in a the next revision).
The question remains the same I repeated recently, though: where do the extra Caucasus components (i.e. beyond EHG) of Eneolithic Ukraine/Corded Ware and Khvalynsk/Yamna come from?
Considering 2-way mixtures, we can model Karelia_HG as deriving 34 ± 2.8% of its ancestry from a Villabruna-related source, with the remainder mainly from ANE represented by the AfontovaGora3 (AG3) sample from Lake Baikal ~17kya.
AG3 was likely of haplogroup Q1a (as reported by YFull, see Genetiker), and probably the ANE ancestry found in Eastern Europe accompanied a Palaeolithic migration of Q1a2-M25 (formed ca. 22600 BC, TMRCA ca. 14300 BC).
Combined with what we know about the Eneolithic Steppe and Caucasus populations – it is likely that ANE ancestry remained the most important component of some of the small ghost populations of the Caucasus until their emergence with the Lola culture.
The first sample we have now attributed to the EHG cluster is Sidelkino, from the Samara region (ca. 9300 BC), mtDNA U5a2. In Damgaard et al. (Science 2018), Yamnaya could be modelled as a CHG population related to Kotias Klde (54%) and the remaining from ANE population related to Sidelkino (>46%), with the following split events:
A split event, where the CHG component of Yamnaya splits from KK1. The model inferred this time at 27 kya (though we note the larger models in Sections S2.12.4 and S2.12.5 inferred a more recent split time).
A split event, where the ANE component of Yamnaya splits from Sidelkino. This was inferred at about about 11 kya.
A split event, where the ANE component of Yamnaya splits from Botai. We inferred this to occur 17 kya. Note that this is above the Sidelkino split time, so our model infers Yamnaya to be more closely related to the EHG Sidelkino, as expected.
An ancestral split event between the CHG and ANE ancestral populations. This was inferred to occur around 40 kya.
Other samples classified as of the EHG cluster:
Popovo2 (ca. 6250 BC) of hg J1, mtDNA U4d – Po2 and Po4 from the same site (ca. 6550 BC) show continuity of mtDNA.
Karelia_HG, from Juzhnii Oleni Ostrov (ca. 6300 BC): I0211/UzOO40 (ca. 6300 BC) of hg J1(xJ1a), mtDNA U4a; and I0061/UzOO74 of hg R1a1(xR1a1a), mtDNA C1
UzOO77 and UzOO76 from Juzhnii Oleni Ostrov (ca. 5250 BC) of mtDNA R1b.
Samara_HG from Lebyanzhinka (ca. 5600 BC) of hg R1b1a, mtDNA U5a1d.
About the enigmatic Anatolia_Neolithic-related ancestry found in Pontic-Caspian steppe samples, this is what Wang et al. (2018) had to say:
We focused on model of mixture of proximal sources such as CHG and Anatolian Chalcolithic for all six groups of the Caucasus cluster (Eneolithic Caucasus, Maykop and Late Makyop, Maykop-Novosvobodnaya, Kura-Araxes, and Dolmen LBA), with admixture proportions on a genetic cline of 40-72% Anatolian Chalcolithic related and 28-60% CHG related (Supplementary Table 7). When we explored Romania_EN and Greece_Neolithic individuals as alternative southeast European sources (30-46% and 36-49%), the CHG proportions increased to 54-70% and 51-64%, respectively. We hypothesize that alternative models, replacing the Anatolian Chalcolithic individual with yet unsampled populations from eastern Anatolia, South Caucasus or northern Mesopotamia, would probably also provide a fit to the data from some of the tested Caucasus groups.
The first appearance of ‘Near Eastern farmer related ancestry’ in the steppe zone is evident in Steppe Maykop outliers. However, PCA results also suggest that Yamnaya and later groups of the West Eurasian steppe carry some farmer related ancestry as they are slightly shifted towards ‘European Neolithic groups’ in PC2 (Fig. 2D) compared to Eneolithic steppe. This is not the case for the preceding Eneolithic steppe individuals. The tilting cline is also confirmed by admixture f3-statistics, which provide statistically negative values for AG3 as one source and any Anatolian Neolithic related group as a second source
Detailed exploration via D-statistics in the form of D(EHG, steppe group; X, Mbuti) and D(Samara_Eneolithic, steppe group; X, Mbuti) show significantly negative D values for most of the steppe groups when X is a member of the Caucasus cluster or one of the Levant/Anatolia farmer-related groups (Supplementary Figs. 5 and 6). In addition, we used f- and D-statistics to explore the shared ancestry with Anatolian Neolithic as well as the reciprocal relationship between Anatolian- and Iranian farmer-related ancestry for all groups of our two main clusters and relevant adjacent regions (Supplementary Fig. 4). Here, we observe an increase in farmer-related ancestry (both Anatolian and Iranian) in our Steppe cluster, ranging from Eneolithic steppe to later groups. In Middle/Late Bronze Age groups especially to the north and east we observe a further increase of Anatolian farmer related ancestry consistent with previous studies of the Poltavka, Andronovo, Srubnaya and Sintashta groups and reflecting a different process not especially related to events in the Caucasus.
(…) Surprisingly, we found that a minimum of four streams of ancestry is needed to explain all eleven steppe ancestry groups tested, including previously published ones (Fig. 2; Supplementary Table 12). Importantly, our results show a subtle contribution of both Anatolian farmer-related ancestry and WHG-related ancestry (Fig.4; Supplementary Tables 13 and 14), which was likely contributed through Middle and Late Neolithic farming groups from adjacent regions in the West. The discovery of a quite old AME ancestry has rendered this probably unnecessary, because this admixture from an Anatolian-like ghost population could be driven even by small populations from the Caucasus.
While it is not yet fully clear, the increased Anatolian_Neolithic-like ancestry in Ukraine_Eneolithic samples (see below) makes it unlikely that all such ancestry in Corded Ware groups comes from a GAC-related contribution. It is likely that at least part of it represents contributions from populations of the Caucasus, based on the mostly westward population movements in the steppe from ca. 4600 BC on, including the Suvorovo-Novodanilovka expansion, and especially the Kuban-Maykop expansion during the final Eneolithic into the North Pontic area.
NOTE. Since CHG-like groups from the Caucasus may have combinations of AME and ANE ancestry similar to Yamna (which may thus appear as ‘steppe ancestry’ in the North Pontic area), it is impossible to interpret with precision the following ADMIXTURE graphic:
The East Asian contribution to samples from the WHG samples (like Loschbour or La Braña), as specified in Fu et al. (2016), does not seem to be related to Baikal_EN, and appears possibly (in the ADMIXTURE analysis) integrated into he Villabruna component. I guess this implies that the shared alleles with East Asians are quite early, and potentially due to the expansion of R1b-L754 from the East.
It would be interesting to know the specific material culture Sidelkino belonged to – i.e. if it was related to the expansion of the North-Eastern Technocomplex – , and its Y-DNA. The Post-Swiderian expansion into eastern Europe, probably associated with the expansion of R1b-P297 lineages (including R1b-M73, found later in Botai and in Baltic HG) is supposed to have begun during the 11th millennium BC, but migrations to the Urals and beyond are probably concentrated in the 9th millennium, so this sample is possibly slightly early for R1b.
NOTE. User Rozenfeld at Anthrogenica posted this, which I think is interesting (in case anyone wants to try a Y-SNP call):
there is something strange with Sidelkino EHG: first, its archaeological context is not described in the supplementary. Second, its sex is not listed in the supplementary tables. Third, after looking for info about this sample, I found that: “Сиделькино-3. Для снятия вопроса о половой принадлежности индивида была проведена генетическая экспертиза, выявившая принадлежность останков мужчине.”(translation: Sidelkino-3. To resolve the question about sex of the remains, the genetic analysis was conducted, which showed that remains belonged to male), source: http://static.iea.ras.ru/books/7487_Traditsii.pdf
So either they haven’t mentioned his Y-DNA in the paper for some reason, or there are more than one Sidelkino sample and the male one has not yet been published. The coverage of the Sidelkino sample from the paper is 2.9, more than enough to tell Y-DNA haplogroup.
My speculative guess right now about specific population movements in far eastern Europe, based on the few data we have:
The expansion of the North-Eastern Technocomplex first around the 9th millennium BC, most likely expanded R1b-P279 ca. 11300 BC, judging by its TMRCA, with both R1b-M73 (TMRCA 5300) and R1b-M269 (TMRCA 4400 BC) info (with extra El Mirón ancestry) back, and thus Eurasiatic.
The expansion of haplogroup J1 to the north may have happened before or after the R1b-P279 expansion. Judging by the increase in AG3-related ancestry near Karelia compared to Baltic_HG, it is possible that it expanded just after R1b-P279 (hence possibly J1-Y6304? TMRCA 9700 BC). Its long-lasting presence in the Caucasus is supported by the Satsurblia (ca. 11300 BC) and the Dolmen BA (ca. 1300 BC) samples.
The expansion of R1a-M17 ca. 6600 BC is still likely to have happened from the east, based on the R1a-M17 samples found in Baikalic cultures slightly later (ca. 5300 BC). The presence of elevated Baikal_EN ancestry in Karelia HG and in Samara HG, and the finding of R1a-M417 samples in the Forest Zone after the Mesolithic suggests a connection with the expansion of Hunter-Gatherer pottery, from the Elshanka culture in the Samara region northward into the Forset Zone and westward into the North Pontic area.
The expansion of R1b-M73 ca. 5300 BC is likely to be associated with the emergence of a group east of the Urals (related to the later Botai culture, and potentially Pre-Yukaghir). Its presence in a Narva sample from Donkalnis (ca. 5200 BC) suggest either an early split and spread of both R1b-P297 lineages (M73 and M269) through Eastern Europe, or maybe a back-migration with hunter-gatherer pottery.
R1b-M269 spread successfully ca. 4400 BC (and R1b-L23 ca. 4100 BC, both based on TMRCA), and this successful expansion is probably to be associated with the Khvalynsk-Novodanilovka expansion. We already know that Samara_HG ca. 5600 was R1b1a, so it is likely that R1b-M269 appeared (or ‘resurged’) in the Volga-Ural region shortly after the expansion of R1a-M17, whose expansion through the region may be inferred by the additional AG3 and Baikal_EN ancestry. Interesting from Samara_HG compared to the previous Sidelkino sample is the introduction of more El Mirón-related ancestry, typical of WHG populations (and thus proper of Baltic groups).
NOTE. The TMRCA dates are obviously gross approximations, because a) the actual rate of mutation is unknown and b) TMRCA estimates are based on the convergence of lineages that survived. The potential finding of R1a-Z645 (possibly Z93+) in Ukraine Eneolithic (ca. 4000 BC), and the potential finding of R1b-L23 in Khvalynsk ca. 4250 BC complicates things further, in terms of dates and origins of any subclade.
The question thus remains as it was long ago: did R1b-M269 lineages expand (‘return’) from the east, near the Urals, or directly from the north? Were they already near Samara at the same time as the expansion of hunter-gatherer pottery, and were not much affected by it? Or did they ‘resurge’ from populations admixed with Caucasus-related ancestry after the expansion of R1a-M17 with this pottery (since there are different stepped expansions from the Samara region)? We could even ask, did R1a-M17 really expand from the east, i.e. are the dates on Baikalic subclades from Moussa et al. (2016) reliable? Or did R1a-M17 expand from some pockets in the Pontic-Caspian steppe, taking over the expansion of HG pottery at some point?
The most interesting aspect from the new paper (regarding Indo-Uralic migrations) is that Ancestral Middle Easterner ancestry will probably be a better proxy for the Anatolia_Neolithic component found in Ukraine Mesolithic to Eneolithic, and possibly also for some of the “more CHG-like” component found among Pontic-Caspian steppe populations, all likely derived from different admixture events with groups from the Caucasus.
NOTE. Even the supposed gene flow of Neolithic Iranian ancestry into the Caucasus can be put into question, since that means possibly a Dzudzuana-like population with greater “deep ancestry” proportion than the one found in CHG, which may still be found within the Caucasus.
If it was not clear already that following ‘steppe ancestry’ wherever it appears is a rather lame way of following Indo-European migrations, every single sample from the Caucasus and their admixture with Pontic-Caspian steppe populations will probably show that “steppe ancestry” is in fact formed by a variety of steppe-related ancestral components, impossible to follow coherently with a single population. Exactly what is happening already with the Siberian ancestry.
If the paper on the Dzudzuana samples has shown something, is that the expansion of an ANE-like population shook the entire Caucasus area up to the Zagros Mountains, creating this ANE – AME cline that are CHG and Iran_N, with further contributions of “deep ancestries” (probably from the south) complicating the picture further.
If this happens with few known samples, and we know of an ANE-like ghost population in the Caucasus (appearing later in the Lola culture), we can already guess that the often repeated “CHG component” found in Ukraine_Eneolithic and Khvalynsk will not be the same (except the part mediated by the Novodanilovka expansion).
This ANE-like expansion happened probably in the Late Upper Palaeolithic, and reached Northern Europe probably after the expansion of the Villabruna cluster (ca. 12000 BC), judging by the advance of AG3-like and ENA-like ancestry in later WHG samples.
The population movements during the Mesolithic and Early Neolithic in the North Pontic area are quite complicated: the extra AME ancestry is probably connected to the admixture with populations from the Caucasus, while the close similarity of Ukraine populations with Scandinavian ones (with an increase in Villabruna ancestry from Mesolithic to Neolithic samples), probably reveal population movements related to the expansion of Maglemose-related groups.
These Maglemose-related groups were probably migrants from the north-west, originally from the Northern European Plains, who occupied the previous Swiderian territory, and then expanded into the North Pontic area. The overwhelming presence of I2a (likely all I2a2a1b1b) lineages in Ukraine Neolithic supports this migration.
The likely picture of Mesolithic-Neolithic migrations in the North Pontic area right now is then:
Expansion of R1a-M459 from the east ca. 12000 BC – probably coupled with AG3 and also some Baikal_EN ancestry. First sample is I1819 from Vasilievka (ca. 8700 BC), another is from Dereivka ca. 6900 BC.
Expansion of R1b-V88 from the Balkans in the west ca. 9700 BC, based on its TMRCA and also the Balkan hunter-gatherer population overwhemingly of this haplogroup from the 10th millennium until the Neolithic. First sample is I1734 from Vasilievka (ca. 7252 BC), which suggests that it replaced the male population there, based on their similar EHG-like adxmixture (and lack of sizeable WHG increase), and shared mtDNA U5b2, U5a2.
Expansion of I2a-Y5606 probably ca. 6800 based on its TMRCA with Janislawice culture. Supporting this is the increase in WHG contribution to Neolithic samples, including the spread of U4 subclades compared to the previous period.
Expansion of R1a-M17 starting probably ca. 6600 BC in the east (see above).
NOTE. The first sample of haplogroup I appears in the Mesolithic: I1763 (ca. 8100 BC) of haplogroup I2a1, probably related to an older Upper Palaeolithic expansion.
It is becoming more and more clear with each new paper that – unless the number of very ancient samples increases – the use of Y-chromosome haplogroups remains one of the most important tools for academics; this is especially so in the steppes, in light of the diversity found in populations from the Caucasus. A clear example comes from the Yamna – Corded Ware similarities:
The presence of haplogroups Q and R1a-M459 (xM17) in Khvalynsk along with a R1b1a sample, which some interpreted as being akin to modern ‘mixed’ populations in the past, is likely to point instead to a period of Khvalynsk-Novodanilovka expansion with R1b-M269, where different small populations from the steppe were being integrated into the common Khvalynsk stock, but where differences are seen in material culture surrounding their burials, as supported by the finding of R1b1 in the Kuban area already in the first half of the 5th millennium. The case would be similar to the early ‘mixed’ Icelandic population.
Only after the emergence of the Samara culture (in the second half of the 6th millennium BC), with a sample of haplogroup R1b1a, starts then the obvious connection with Early Proto-Indo-Europeans; and only after the appearance of late Sredni Stog and haplogroup R1a-M417 (ca. 4000 BC) is its connection with Uralic also clear. In previous population movements, I think more haplogroups were involved in migrations of small groups, and only some communities among them were eventually successful, expanding to be dominant, creating ever growing cultures during their expansions.
Indeed, if you think in terms of Uralic and Indo-European just as converging languages, and forget their potential genetic connection, then the genetic + linguistic picture becomes simplified, and the upper frontier of the 6th millennium BC with a division North Pontic (Mariupol) vs. Volga-Ural (Samara) is enough. However, tracing their movements backwards – with cultural expansions from west to east (with the expansion of farming), and earlier east to west (with hunter-gatherer pottery), and still earlier west to east (with the north-eastern technocomplex), offers an interesting way to prove their potential connection to macrofamilies, at least in terms of population movements.
I am quite convinced right now that it would be possible to connect the expansion of R1b-L754 subclades with a speculative Nostratic (given the R1b-V88 connection with Afroasiatic, and the obvious connection of R1b-L297 with Eurasiatic). Paradoxically, the connection of an Indo-Uralic community in the steppes (after the separation of Yukaghir) with any lineage expansion (R1a-M17, R1b-M269, or even Q, I or J1) seems somehow blurrier than one year ago, possibly just because there are too many open possibilities.
David Reich says about the admixture with Neanderthals, which he helped discover:
At the conclusion of the Neanderthal genome project, I am still amazed by the surprises we encountered. Having found the first evidence of interbreeding between Neanderthals and modern humans, I continue to have nightmares that the finding is some kind of mistake. But the data are sternly consistent: the evidence for Neanderthal interbreeding turns out to be everywhere. As we continue to do genetic work, we keep encountering more and more patterns that reflect the extraordinary impact this interbreeding has had on the genomes of people living today.
I think this is a shared feeling among many of us who have made proposals about anything, to fear that we have made a gross, evident mistake, and constantly look for flaws. However, it seems to me that geneticists are more preoccupied with being wrong in their developed statistical methods, in the theoretical models they are creating, and not so much about errors in the true ancient ethnolinguistic picture human population genetics is (at least in theory) concerned about. Their publications are, after all, constantly associating genetic finds with cultures and (whenever possible) languages, so this aspect of their research should not be taken lightly.
Seeing how David Anthony or Razib Khan (among many others) have changed their previously preferred migration models as new data was published, and they continue to be respected in their own fields, I guess we can be confident that professionals with integrity are going to accept whatever new picture appears. While I don’t think that genetic finds can change what we can reconstruct with comparative grammar, I am also ready to revise guesstimates and routes of expansion of certain dialects if R1a-Z645 is shown to have accompanied Late Proto-Indo-Europeans during their expansion with Yamna, and later integrated somehow with Corded Ware.
However, taking into account the obsession of some with an ancestral, uninterrupted R1a—Indo-European association, and the lack of actual political repercussion of Neanderthal admixture, I think the most common nightmare that all genetic researchers should be worried about is to keep inflating this “Yamnaya ancestry”-based hornet’s nest, which has been constantly stirred up for the past two years, by rejecting it – or, rather, specifying it into its true complex nature.
This succession of corrections and redefinitions, coupled with the distinct Y-DNA bottleneck of each steppe population, will eventually lead to a completely different ethnolinguistic picture of the Pontic-Caspian region during the Eneolithic, which is likely to eventually piss off not only reasonable academics stubbornly attached to the CWC-IE idea, but also a part of those interested in daydreaming about their patrilineal ancestors.
Sometimes it’s better to just rip off the band-aid once and for all…
I was reading The Bronze Age Landscape in the Russian Steppes: The Samara Valley Project (2016), and I was really surprised to find the following excerpt by David W. Anthony:
The Samara Valley links the central steppes with the western steppes and is a north-south ecotone between the pastoral steppes to the south and the forest-steppe zone to the north [see figure below]. The economic contrast between pastoral steppe subsistence, with its associated social organizations, and forest-zone hunting and fishing economies probably explains the shifting but persistent linguistic border between forest-zone Uralic languages to the north (today largely displaced by Russian) and a sequence of steppe languages to the south, recently Turkic, before that Iranian, and before that probably an eastern dialect of Proto-Indo-European (Anthony 2007). The Samara Valley represents several kinds of borders, linguistic, cultural, and ecological, and it is centrally located in the Eurasian steppes, making it a critical place to examine the development of Eurasian steppe pastoralism.
Khokhlov (translated by Anthony) further insists on the racial and ethnic divide between both populations, Abashevo to the north, and Poltavka to the south, during the formation of the Abashevo – Sintashta-Potapovka community that gave rise to Proto-Indo-Iranians:
Among all cranial series in the Volga-Ural region, the Potapovka population represents the clearest example of race mixing and probably ethnic mixing as well. The cultural advancements seen in this period might perhaps have been the result of the mixing of heterogeneous groups. Such a craniometric observation is to some extent consistent with the view of some archaeologists that the Sintashta monuments represent a combination of various cultures (principally Abashevo and Poltavka, but with other influences) and therefore do not correspond to the basic concept of an archaeological culture (Kuzmina 2003:76). Under this option, the Potapovka-Sintashta burial rite may be considered, first, a combination of traits to guarantee the afterlife of a selected part of a heterogeneous population. Second, it reflected a kind of social “caste” rather than a single population. In our view, the decisive element in shaping the ethnic structure of the Potapovka-Sintashta monuments was their extensive mobility over a fairly large geographic area. They obtained knowledge of various cultures from the populations with whom they interacted.
Interesting is also this excerpt about the predominant population in the Abashevo – Sintashta-Potapovka admixture (which supports what Chetan said recently, although this does not seemed backed by Y-DNA haplogroups found in the richest burials), coupled with the sign of incoming “Uraloid” peoples from the east, found in both Sintashta and eastern Abashevo:
The socially dominant anthropological component was Europeoid, possibly the descendants of Yamnaya. The association of craniofacial types with archaeological cultures in this period is difficult, primarily because of the small amount of published anthropological material of the cultures of steppe and forest belt (Balanbash, Vol’sko-Lbishche) and the eastern and southern steppes (Botai-Tersek). The crania associated with late MBA western Abashevo groups in the Don-Volga forest zone were different from eastern Abashevo in the Urals, where the expression of the Old Uraloid craniological complex was increased. Old Uraloid is found also on a single skull of Vol’sko-Lbishche culture (Tamar Utkul VII, Kurgan 4). Potentially related variants, including Mongoloid features, could be found among the Seima-Turbino tribes of the forest-steppe zone, who mixed with Sintashta and Abashevo. In the Sintashta Bulanova cemetery from the western Urals, some individuals were buried with implements of Seima-Turbino type (Khalyapin 2001; Khokhlov 2009; Khokhlov and Kitov 2009). Previously, similarities were noted between some individual skulls from Potapovka I and burials of the much older Botai culture in northern Kazakhstan (Khokhlov 2000a). Botai-Tersek is, in fact, a growing contender for the source of some “eastern” cranial features.
The wave of peoples associated with “eastern” features can be seen in genetics in the Sintashta outliers from Narasimhan et al. (2018), and it probably will be eventually seen in Abashevo, too. These may be related to the Seima-Turbino international network – but most likely it is directly connected to Sintashta through the starting Andronovo and Seima-Turbino horizons, by admixing of prospective groups and small-scale back-migrations.
Corded Ware – Yamna similarities?
So, if peoples of north-eastern Europe have been assumed for a long time to be Uralic speakers, what is happening with the Corded Ware = IE obsession? Is it Gimbutas’ ghost possessing old archaeologists? Probably not.
It is about certain cultural similarities evident at first sight, which have been traditionally interpreted as a sign of cultural diffusion or migration. Not dissimilar to the many Bell Beaker models available, where each archaeologist is pushing certain differences, mixing what seemed reasonable, what still might seem reasonable, and what certainly isn’t anymore after the latest ancient DNA data.
The initial models of Gimbutas, Kristiansen, or Anthony – which are known to many today – were enunciated in the infancy of archaeological studies in the regions, during and just after the fall of the USSR, and before many radiocarbon dates that we have today were published (with radiocarbon dating being still today in need of refinement), so it is only logical that gross mistakes were made.
We have similar gross mistakes related to the origins of Bell Beakers, and studying them was certainly easier than studying eastern data.
Gimbutas believed – based mainly on Kurgan-like burials – that Bell Beaker formed from a combination of Yamna settlers with the Vučedol culture, so she was not that far from the truth.
The expansion of Corded Ware from peoples of the North Pontic forest-steppe area, proposed by Gimbutas and later supported also by Kristiansen (1989) as the main Indo-European expansion – , is probably also right about the approximate origins of the culture. Only its ‘Indo-European’ nature is in question, given the differences with Khvalynsk and Yamna evolution.
Anthony only claimed that Yamna migrants settled in the Balkans and along the Danube into the Hungarian steppes. He never said that Corded Ware was a Yamna offshoot until after the first genetic papers of 2015 (read about his newest proposal). He initially claimed that only certain neighbouring Corded Ware groups “adopted” Indo-European (through cultural diffusion) because of ‘patron-client’ relationships, and was never preoccupied with the fate of Corded Ware and related cultures in the east European forest zone and Finland.
So none of them was really that far from the true picture; we might say a lot people are more way off the real picture today than the picture these three researchers helped create in the 1990s and 2000s. Genetics is just putting the last nail in the coffin of Corded Ware as a Yamna offshoot, instead of – as we believed in the 2000s – to Vučedol and Bell Beaker.
So let’s revise some of these traditional links between Corded Ware and Yamna with today’s data:
Even more than genetics – at least until we have an adequate regional and temporary sampling – , archaeological findings lead what we have to know about both cultures.
It is essential to remember that Corded Ware, starting ca. 3000/2900 BC in east-central Europe, has been proposed to be derived from Early Yamna, which appeared suddenly in the Pontic-Caspian steppes ca. 3300 BC (probably from the late Repin expansion), and expanded to the west ca. 3000.
The question at hand, therefore, is if Corded Ware can be considered an offshoot of the Late PIE community, and thus whether the CWC ethnolinguistic community – proven in genetics to be quite homogeneous – spoke a Late PIE dialect, or if – alternatively – it is derived from other neighbouring cultures of the North Pontic region.
NOTE. The interpretation of an Indo-Slavonic group represented by a previous branching off of the group is untenable with today’s data, since Indo-Slavonic – for those who support it – would itself be a branch of Graeco-Aryan, and Palaeo-Balkan languages expanded most likely with West Yamna (i.e. R1b-L23, mainly R1b-Z2103) to the south.
The convoluted alternative explanation would be that Corded Ware represents an earlier, Middle PIE branch (somehow carrying R1a??) which influences expanding Late PIE dialects; this has been recently supported by Kortlandt, although this simplistic picture also fails to explain the Uralic problem.
❔ Kurgans: The Yamna tradition was inherited from late Repin, in turn inherited from Khvalynsk-Novodanilovka proto-Kurgans. As for the CWC tradition, it is unclear if the tumuli were built as a tradition inherited from North and West Pontic cultures (in turn inherited or copied from Khvalynsk-Novodanilovka), such as late Trypillia, late Kvityana, late Dereivka, late Sredni Stog; or if they were built because of the spread of the ‘Transformation of Europe’, set in motion by the Early Yamna expansion ca. 3300-3000 BC (as found in east-central European cultures like Coţofeni, Lizevile, Șoimuș, or the Adriatic Vučedol). My guess is that it inherits an older tradition than Yamna, with an origin in east-central Europe, because of the mound-building distribution in the North Pontic area before the Yamna expansion, but we may never really know.
❌ Burial rite: Yamna features (with regional differences) single burials with body on its back, flexed upright knees, poor grave goods, common orientation east-west (heads to the west) inherited from Repin, in turn inherited from Khvalynsk-Novodanilovka. CWC tradition – partially connected to Złota and surrounding east-central European territories (in turn from the Khvalynsk-Novodanilovka expansion) – features single graves, body in fetal position, strict gender differentiation – men on the right, women on the left -, looking to the south, graves with standardized assemblages (objects representing affirmation of battle, hunting, and feasting). The burial rites clearly represent different ideologies.
❌ Corded decoration: Corded ware decoration appears in the Balkans during the 5th millennium, and represents a simple technique whereby a cord is twisted, or wrapped around a stick, and then pressed directly onto the fresh surface of a vessel leaving a characteristic decoration. It appears in many groups of the 5th and 4th millennium BC, but it was Globular Amphorae the culture which popularized the drinking vessels and their corded ornamentation. It appears thus in some regional groups of Yamna, but it becomes the standard pottery only in Corded Ware (especially with the A-horizon), which shows continuity with GAC pottery.
❌ Economy: Yamna expands from Repin (and Repin from Khvalynsk-Novodanilovka) as a nomadic or semi-nomadic purely pastoralist society (with occasional gathering of wild seeds), which naturally thrives in the grasslands of the Pontic-Caspian, lower Danube and Hungarian steppes. Corded Ware shows agropastoralism (as late Eneolithic forest-steppe and steppe groups of eastern Europe, such as late Trypillian, TRB, and GAC groups), inhabits territories north of the loess line, with heavy reliance of hunter-gathering depending on the specific region.
❌ Cattle herding: Interestingly, both west Yamna and Corded Ware show more reliance on cattle herding than other pastoralist groups, which – contrasted with the previous Eneolithic herding traditions of the Pontic-Caspian steppe, where sheep-goats predominate – make them look alike. However, the cattle-herding economy of Yamna is essential for its development from late Repin and its expansion through the steppes (over western territories practising more hunter-gathering and sheep-goat herding economy), and it does not reach equally the Volga-Ural region, whose groups keep some of the old subsistence economy (read more about the late Repin expansion). Corded Ware, on the other hand, inherits its economic strategy from east European groups like TRB, GAC, and especially late Trypillian communities, showing a predominance of cattle herding within an agropastoral community in the forest-steppe and forest zones of Volhynia, Podolia, and surrounding forest-steppe and forest regions.
❔ Horse riding: Horse riding and horse transport is proven in Yamna (and succeeding Bell Beaker and Sintashta), assumed for late Repin (essential for cattle herding in the seas of grasslands that are the steppes, without nearby water sources), quite likely during the Khvalynsk expansion (read more here), and potentially also for Samara, where the predominant horse symbolism of early Khvalynsk starts. Corded Ware – like the north Pontic forest-steppe and forest areas during the Eneolithic – , on the other hand, does not show a strong reliance on horse riding. The high mobility and short-term settlements characteristic of Corded Ware, that are often associated with horse riding by association with Yamna, may or may not be correct, but there is no need for horses to explain their herding economy or their mobility, and the north-eastern European areas – the one which survived after Bell Beaker expansion – did certainly not rely on horses as an essential part of their economy.
NOTE: I cannot think of more supposed similarities right now. If you have more ideas, please share in the comments and I will add them here.
✅ EHG: This is the clearest link between both communities. We thought it was related to the expansion of ANE-related ancestry to the west into WHG territory, but now it seems that it will be rather WHG expanding into ANE territory from the Pontic-Caspian region to the east (read more on recent Caucasus Neolithic, on , and on Caucasus HG).
NOTE. Given how much each paper changes what we know about the Palaeolithic, the origin and expansion of the (always developing) known ancestral components and specific subclades (see below) is not clear at all.
❔ CHG: This is the key link between both cultures, which will delimit their interaction in terms of time and space. CHG is intermediate between EHG and Iran N (ca. 8000 BC). The ancestry is thus linked to the Caucasus south of the steppe before the emergence of North Pontic (western) and Don-Volga-Ural (eastern) communities during the Mesolithic. The real question is: when we have more samples from the steppe and the Caucasus during the Neolithic, how many CHG groups are we going to find? Will the new specific ancestral components (say CHG1, CHG2, CHG3, etc.) found in Yamna (from Khvalynsk, in the east) and Corded Ware (probably from the North Pontic forest-steppe) be the same? My guess is, most likely not, unless they are mediated by the Khvalynsk-Novodanilovka expansion (read more on CHG in the Caucasus).
❌ WHG/EEF: This is the obvious major difference – known today – in the formation of both communities in the steppe, and shows the different contacts that both groups had at least since the Eneolithic, i.e. since the expansion of Repin with its renewed Y-DNA bottleneck, and probably since before the early Khvalynsk expansion (read more on Yamna-Corded Ware differences contrasting with Yamna-Afanasevo, Yamna-Bell Beaker, and Yamna-Sintashta similarities).
NOTE 1. Some similarities between groups can be seen depending on the sampled region; e.g. Baltic groups show more similarities with southern Pontic-Caspian steppe populations, probably due to exogamy.
NOTE 2. We have this information on the differences in “steppe ancestry” between Yamna and Corded Ware, compared to previous studies, because now we have more samples of neighbouring, roughly contemporaneous Eneolithic groups, to analyse the real admixture processes. This kind of fine scale studies is what is going to show more and more differences between Khvalynsk-Yamna and Sredni Stog-Corded Ware as more data pours in. The evolution of both communities in archaeology and in PCA (see below) is probably witness to those differences yet to be published.
❌ R1: Even though some people try very hard to think in terms of “R1” vs. (Caucasus) J or G or any other upper clade, this is plainly wrong. It is possible, given what we know now, that Q1a2-M242 expanded ANE ancestry to the west ca. 13000 BC, while R1b-P279 expanded WHG ancestry to the east with the expansion of post-Swiderian cultures, creating EHG as a WHG:ANE cline. The role of R1a-M459 is unknown, but it might be related to any of these migrations, or others (plural) along northern Eurasia (read more on the expansion of R1b-P279, on Palaeolithic Q1a2, and on R1a-M417).
NOTE. I am inclined to believe in a speculative Mesolithic-Early Neolithic community involving Eurasiatic movements accross North Eurasia, and Indo-Uralic movements in its western part, with the last intense early Uralic-PIE contacts represented by the forming west (Mariupol culture) and east (Don-Volga-Ural cultures, including Samara) communities developing side by side. Before their known Eneolithic expansions, no large-scale Y-DNA bottleneck is going to be seen in the Pontic-Caspian steppe, with different (especially R1a and R1b subclades) mixed among them, as shown in North Pontic Neolithic, Samara HG, and Khvalynsk samples.
Corded Ware and ‘steppe ancestry’
If we take a look at the evolution of Corded Ware cultures, the expansion of Bell Beakers – dominated over most previous European cultures from west to east Europe – influenced the development of the whole European Bronze Age, up to Mierzanowice and Trzciniec in the east.
The only relevant unscathed CWC-derived groups, after the expansion of Sintashta-Potapovka as the Srubna-Andronovo horizon in the Eurasian steppes, were those of the north-eastern European forest zone: between Belarus to the west, Finland to the north, the Urals to the east, and the forest-steppe region to the south. That is, precisely the region supposed to represent Uralic speakers during the Bronze Age.
This inconsistency of steppe ancestry and its relation with Uralic (and Balto-Slavic) peoples was observed shortly after the publication of the first famous 2015 papers by Paul Heggarty, of the Max-Planck Institute for Evolutionary Anthropology (read more):
Haak et al. (2015) make much of the high Yamnaya ancestry scores for (only some!) Indo-European languages. What they do not mention is that those same results also include speakers of other languages among those with the highest of all scores for Yamnaya ancestry. Only these are languages of the Uralic family, not Indo-European at all; and their Yamnaya-ancestry signals are far higher than in many branches of Indo-European in (southern) Europe. Estonian ranks very high, while speakers of the very closely related Finnish are curiously not shown, and nor are the Saami. Hungarian is relevant less directly since this language arrived only c. 900 AD, but also high.
These data imply that Uralic-speakers too would have been part of the Yamnaya > Corded Ware movement, which was thus not exclusively Indo-European in any case. And as well as the genetics, the geography, chronology and language contact evidence also all fit with a Yamnaya > Corded Ware movement including Uralic as well as Balto-Slavic.
Both papers fail to address properly the question of the Uralic languages. And this despite — or because? — the only Uralic speakers they report rank so high among modern populations with Yamnaya ancestry. Their linguistic ancestors also have a good claim to have been involved in the Corded Ware and Yamnaya cultures, and of course the other members of the Uralic family are scattered across European Russia up to the Urals.
NOTE. Although the author was trying to support the Anatolian hypothesis – proper of glottochronological studies often published from the Max Planck Institute – , the question remains equally valid: “if Proto-Indo-European expands with Corded Ware and steppe ancestry, what is happening with Uralic peoples?”
For my part, I claimed in my draft that ancestral components were not the only relevant data to take into account, and that Y-DNA haplogroups R1a and R1b (appearing separately in CWC and Yamna-Bell Beaker-Afanasevo), together with their calculated timeframes of formation – and therefore likely expansion – did not fit with the archaeological and linguistic description of the spread of Proto-Indo-European and its dialects.
In fact, it seemed that only one haplogroup (R1b-M269) was constantly and consistenly associated with the proposed routes of Late PIE dialectal expansions – like Anthony’s second (Afanasevo) and third (Lower Danube, Balkan) waves. What genetics shows fits seamlessly with Mallory’s association of the North-West Indo-European expansion with Bell Beakers (read here how archaeologists were right).
More precise inconsistencies were observed after the publication of Olalde et al. (2017) and Mathieson et al. (2017), by Volker Heyd in Kossinna’s smile (2017). Letting aside the many details enumerated (you can read a summary in my latest draft), this interesting excerpt is from the conclusion:
Simple solutions to complex problems are never the best choice, even when favoured by politicians and the media. Kossinna also offered a simple solution to a complex prehistoric problem, and failed therein. Prehistoric archaeology has been aware of this for a century, and has responded by becoming more differentiated and nuanced, working anthropologically, scientifically and across disciplines (cf. Müller 2013; Kristiansen 2014), and rejecting monocausal explanations. The two aDNA papers in Nature, powerful and promising as they are for our future understanding, also offer rather straightforward messages, heavily pulled by culture-history and the equation of people with culture. This admittedly is due partly to the restrictions of the medium that conveys them (and despite the often relevant additional detail given as supplementary information, which is unfortunately not always given full consideration).
While I have no doubt that both papers are essentially right, they do not reflect the complexity of the past. It is here that archaeology and archaeologists contributing to aDNA studies find their role; rather than simply handing over samples and advising on chronology, and instead of letting the geneticists determine the agenda and set the messages, we should teach them about complexity in past human actions and interactions. If accepted, this could be the beginning of a marriage made in heaven, with the blessing smile of Gustaf Kossinna, and no doubt Vere Gordon Childe, were they still alive, in a reconciliation of twentieth- and twenty-first-century approaches. For us as archaeologists, it could also be the starting point for the next level of a new archaeology.
The question was made painfully clear with the publication of Olalde et al. (2018) & Mathieson et al. (2018), where the real route of Yamna expansion into Europe was now clearly set through the steppes into the Carpathian basin, later expanded as Bell Beakers.
After 568 AD the nomadic Avars settled in the Carpathian Basin and founded their empire, which was an important force in Central Europe until the beginning of the 9th century AD. The Avar elite was probably of Inner Asian origin; its identification with the Rourans (who ruled the region of today’s Mongolia and North China in the 4th-6th centuries AD) is widely accepted in the historical research.
Here, we study the whole mitochondrial genomes of twenty-three 7th century and two 8th century AD individuals from a well-characterised Avar elite group of burials excavated in Hungary. Most of them were buried with high value prestige artefacts and their skulls showed Mongoloid morphological traits.
The majority (64%) of the studied samples’ mitochondrial DNA variability belongs to Asian haplogroups (C, D, F, M, R, Y and Z). This Avar elite group shows affinities to several ancient and modern Inner Asian populations.
The genetic results verify the historical thesis on the Inner Asian origin of the Avar elite, as not only a military retinue consisting of armed men, but an endogamous group of families migrated. This correlates well with records on historical nomadic societies where maternal lineages were as important as paternal descent.
The mitochondrial genome sequences can be assigned to a wide range of the Eurasian haplogroups with dominance of the Asian lineages, which represent 64% of the variability: four samples belong to Asian macrohaplogroup C (two C4a1a4, one C4a1a4a and one C4b6); five samples to macrohaplogroup D (one by one D4i2, D4j, D4j12, D4j5a, D5b1), and three individuals to F (two F1b1b and one F1b1f). Each haplogroup M7c1b2b, R2, Y1a1 and Z1a1 is represented by one individual. One further haplogroup, M7 (probably M7c1b2b), was detected (sample AC20); however, the poor quality of its sequence data (2.19x average coverage) did not allow further analysis of this sample.
European lineages (occurring mainly among females) are represented by the following haplogroups: H (one H5a2 and one H8a1), one J1b1a1, three T1a (two T1a1 and one T1a1b), one U5a1 and one U5b1b (Table S1).
We detected two identical F1b1f haplotypes (AC11 female and AC12 male) and two identical C4a1a4 haplotypes (AC13 and AC15 males) from the same cemetery of Kunszállás; these matches indicate the maternal kinship of these individuals. There is no chronological difference between the female and the male from Grave 30 and 32 (AC11 and AC12), but the two males buried in Grave 28 and 52 (AC13 and AC15) are not contemporaries; they lived at least 2-3 generations apart.
The Avar period elite shows the lowest and non-significant genetic distances to ancient Central Asian populations dated to the Late Iron Age (Hunnic) and to the Medieval period, which is displayed on the ancient MDS plot (Fig. 4); these connections are also reflected on the haplogroup based Ward-type clustering tree (Fig. 3). Building of these large Central Asian sample pools is enabled by the small number of samples per cultural/ethnic group. Further mitogenomic data from Inner Asia are needed to specify the ancient genetic connections; however, genomic analyses are also set back by the state of archaeological research, i.e. the lack of human remains from the 4th-5th century Mongolia, which would be a particularly important region in the study of the Avar elite’s origin.
The investigated elite group from the Avar period elite also shows low genetic distances and phylogenetic connections to several Central and Inner Asian modern populations. Our results indicate that the source population of the elite group of the Avar Qaganate might have existed in Inner Asia (region of today’s Mongolia and North China) and the studied stratum of the Avars moved from there westwards towards Europe. Further genetic connections of the Avars to modern populations living to East and North of Inner Asia (Yakuts, Buryats, Tungus) probably indicate common source populations.
Sadly, no Y-DNA is available from this paper, although haplogroups Q, C2, or R1b (xM269) are probably to be expected, given the reported mtDNA. A replacement of the male population with subsequent migrations is obvious from the current distribution of Y-DNA haplogroups in the Carpathian Basin.
Hungarians and Corded Ware
Ancient Hungarians are important to understand the evolution, not only of Ugric, but also of Finno-Ugric peoples and their origin, since they show a genetic picture before more recent population expansions, genetic drift, and bottlenecks in eastern Europe.
In Ob-Ugric peoples, from the scarce data found in Pimenoff et al. (2018), we can see how Siberian N subclades expanded further after the separation of Magyars, evidenced by the inverted proportion of haplogroups R1a and N in modern Khantys and Mansis compared to Hungarians, and the diversity of N subclades compared to modern Fennic peoples.
Similarly to Hungarians, the situation of modern Estonians (where R1a and N subclades show approximately the same proportion, ca. 33%) is probably closer to Fennic peoples in Antiquity, not having undergone the latest strong founder effect evident in modern Finns after their expansion to the north.
In Semino et al. (2001) they found among 45 Palóc from Budapest and northern Hungary: 60% R1a, 13% R1b, 11% I, 9% E, 2% G, 2% J2.
In Csányi et al. (2008) Among 100 Hungarian men, 90 of whom from the Great Hungarian Plain: 30% R1a, 15% R1b, 13% I2a1, 13% J2, 9% E1b1b1a, 8% I1, 3% G2, 3% J1, 3% I*, 1% E*, 1% F*, 1% K*. Among 97 Székelys, in Romania: 20% R1b, 19% R1a, 17% I1, 11% J2, 10% J1, 8% E1b1b1a, 5% I2a1, 5% G2, 3% P*, 1% E*, 1% N.
In Pamjav et al. (2011), among 230 samples expected to include 6-8% Gypsy peoples: 26% R1a, 20% I2a, 19% R1b, 7% I, 6% J2, 5% H, 5% G2a, 5% E1b1b1a1, 3% J1, <1% N, <1% R2.
In Pamjav et al. (2017), from the Bodrogköz population: R1a-M458 (20.4%), I2a1-P37 (19%), R1b-M343 (15%), R1a-Z280 (14.3%), E1b-M78 (10.2%), and N1c-Tat (6.2%).
NOTE. The N1c-Tat found in Bodrogköz belongs to the N1c-VL29 subgroup, more frequent among Balto-Slavic peoples, which may suggest (yet again) an initial stage of the expansion of N subclades among Finno-Ugric peoples by the time of the Hungarian migration.
3.2% N (1.4% Z9136, 0.5% M2019/VL67, 0.5% Y7310, 0.9% Z16981)- note: only unrelated males are sampled
2.3% Q (1.2% YP789, 0.9% M346, 0.2% M242)
R1a-Z280 stands out in FDNA (which we have to assume has no geographic preference among modern Hungarians), while R1a-M458 is prevalent in the north, which probably points to its relationship with (at least West) Slavic populations.
NOTE. For more on the analysis of probability of the actual subclade, see here.
Bronze Age R1a-Z93 samples of central-east Europe – like the Balkans BA sample (ca. 1750-1625 BC) from Merichleri, of R1a1a1b2 subclade – correspond most likely to the expansion of Iranian-speaking peoples in the early 2nd millennium BC, probably to the westward expansion of the Srubna culture.
The specific subclade of King Béla III, on the other hand, probably corresponds to the more recent expansion of Magyar tribes settled in the region during the 9th century AD, so the specific subclade must have separated from those found in central-east Europe and in Andronovo during the Corded Ware expansion.
The study by Csányi et al. (2008), where the Tat C allele was found in 2 of 4 ancient samples, showed thus a potential 50:50 relationship of N1c in ancient Magyars, which is striking given the modern 1-3% a mere 1,000 years later, without any relevant population movement in between. This result remains to be reproduced with the current technology.
In fact, recent studies of ancient Magyars, from the 10th to the 12th century, have not shown any N1c sample, and have confirmed instead the ancient presence of R1a (two other samples, interred near Béla III), R1b (four samples), I2a (two samples) J1, and E1b, a mixed genetic picture which is more in line with what is expected.
So the question that I recently posed about east Corded Ware groups remains open: were Proto-Ugric peoples mainly of R1a-Z282 or R1a-Z93 subclades? Without ancient DNA from Middle Dnieper, Fatyanovo, Afanasevo, and the succeeding cultures (like Netted Ware) in north-eastern Europe, it is difficult to say.
It is very likely that they are going to show mainly a mixture of both R1a-Z282 and R1a-Z93 lineages, with later populations showing a higher proportion of R1a-Z280 subclades. Whether this mixture happened already during the Corded Ware period, or is the result of later developments, is still unknown. What is certain is that Hungarian N1a1a1a-L708 subclades belong to more recent additions of Siberian haplogroups to the Ugric stock, probably during the Iron Age, just centuries before the Magyar expansion.
The Keriyan, Lopnur and Dolan peoples are isolated populations with sparse numbers living in the western border desert of our country. By sequencing and typing the complete Y-chromosome of 179 individuals in these three isolated populations, all mutations and SNPs in the Y-chromosome and their corresponding haplotypes were obtained. Types and frequencies of each haplotype were analyzed to investigate genetic diversity and genetic structure in the three isolated populations. The results showed that 12 haplogroups were detected in the Keriyan with high frequencies of the J2a1b1 (25.64%), R1a1a1b2a (20.51%), R2a (17.95%) and R1a1a1b2a2 (15.38%) groups. Sixteen haplogroups were noted in the Lopnur with the following frequencies: J2a1 (43.75%), J2a2 (14.06%), R2 (9.38%) and L1c (7.81%). Forty haplogroups were found in the Dolan, noting the following frequencies: R1b1a1a1 (9.21%), R1a1a1b2a1a (7.89%), R1a1a1b2a2b (6.58%) and C3c1 (6.58%). These data show that these three isolated populations have a closer genetic relationship with the Uygur, Mongolian and Sala peoples. In particular, there are no significant differences in haplotype and frequency between the three isolated populations and Uygur (f=0.833, p=0.367). In addition, the genetic haplotypes and frequencies in the three isolated populations showed marked Eurasian mixing illustrating typical characteristics of Central Asian populations.
My knowledge of written Chinese is almost zero, so here are some excerpts with the help of Google Translate:
The source of 179 blood samples used in the study is shown in Figure 1. The Keriyan blood samples were collected from Dali Yabuyi Township, Yutian County (39 samples). The blood samples of the Lopnur people were collected from Kaerqu Township, Yuli County (64 cases); the blood samples of the Dolan people were collected from the town of Uluru, Awati County (76).
The composition and frequency of the Keriyan people’s haplogroup are closest to those of the Uighurs, and both Principal Component Analysis and Phylogenetic Tree Analysis show that their kinship is recent. We initially infer that the Keriyan are local desert indigenous people. They have a connection with the source of the Uighurs. Chen et al.  studied the patriarchal and maternal genetic analysis of the Keriyan people and found that they are not descendants of the Tibetan ethnic group in the West. The Keriyan people are a mixed group of Eastern and Western Europeans, which may originate from the local Vil group. Duan Ranhui  and other studies have shown that the nucleotide variability and average nucleotide differences in the Keriyan population are between the reported Eastern and Western populations. The phylogenetic tree also shows that the populations in Central Asia are between the continental lineage of the eastern population and the European lineage of the western population, and the genetic distance between the Keriyan and the Uighurs is the closest, indicating that they have a close relationship.
Regarding the origin of the Lopnur people, Purzhevski judged that it was a mixture of Mongolians and Aryans according to the physical characteristics of the Lopnur people. In 1934, the Sino-Swiss delegation discovered the famous burials of the ancient tombs in the Peacock River. After research, they were the indigenous people before the Loulan period; the researcher Yang Lan, a researcher at the Institute of Cultural Relics of the Chinese Academy of Social Sciences, said that the Lopnur people were descendants of the ancient “Landan survivors”. However, the Loulan people speaking an Indo-European language, and the Lopnur people speaking Uyghur languages contradict this; the historical materials of the Western Regions, “The Geography of the Western Regions” and “The Western Regions of the Ming Dynasty” record the Uighurs who lived in Cao Cao in the late 17th and early 18th centuries. Because of the occupation of the land by the Junggar nobles and their oppression, they fled. Some of them were forced to move to the Lop Nur area. There are many similar archaeological discoveries and historical records. We have no way to determine their accuracy, but they are at different times, and there is a great difference in what is heard in the same region. (…) The genetic characteristics of modern Lopnur people are the result of the long-term ethnic integration of Uyghurs, Mongols, and Europeans. This is also consistent with the similarity of the genetic structure of the Y chromosome of Lopnur in this study with the Uighurs and Mongolians. For example, the frequency of J haplogroup is as high as 59.37%, while J and its downstream sub-haplogroup are mainly distributed in western Europe, West Asia and Central Asia; the frequency of O, R haplogroup is close to that of Mongolians.
According to Ming History·Western Biography, the Mongolians originated from the Mobei Plateau and later ruled Asia and Eastern Europe. Mongolia was established, and large areas of southern Xinjiang and Central Asia were included. Later, due to the Mongolian king’s struggle for power, it fell into a long-term conflict. People of the land fled to avoid the war, and the uninhabited plain of the lower reaches of the Yarkant River naturally became a good place to live. People from all over the world gathered together and called themselves “Dura” and changed to “Dang Lang”. The long-term local Uyghur exchanges that entered the southern Mongolian monks and “Dura” were gradually assimilated . According to the report, locals wore Mongolian clothes, especially women who still maintained a Mongolian face . In 1976, the robes and waistbands found in the ancient time of the Daolang people in Awati County were very similar to those of the ancients. Dalang Muqam is an important part of Daolang culture. It is also a part of the Uyghur Twelve Muqam, and it retains the ancient Western culture, but it also contains a larger Mongolian culture and relics. The above historical records show that the Daolang people should appear in the Chagatai Khanate and be formed by the integration of Mongolian and Uighur ethnic groups. Through our research, we also found that the paternal haplotype of the Daolang people is contained in both Uygur and Mongolian, and the main haplogroups are the same, whereas the frequencies are different (see Table 3). The principal component analysis and the NJ analysis are also the same. It is very close to the Uyghur and the Mongolian people, which establishes new evidence for the “mixed theory” in molecular genetics.
If the nomenclature follows a recent ISOGG standard, it appears that:
The presence of exclusively R1a-Z93 subclades and the lack of R1b-M269 samples is compatible with the expansion of R1a-Z93 into the area with Proto-Tocharians, at the turn of the 3rd-2nd millennium BC, as suggested by the Xiaohe samples, supposedly R1a(xZ93).
Lacking proper assessment of ancient DNA from Proto-Tocharians, this potential early Y-DNA replacement is still speculative*. However, if that is the case, I wonder what the Copenhagen group will say when supporting this, but rejecting at the same time the more obvious Y-DNA replacement in East Yamna / Poltavka in the mid-3rd millennium with incoming Corded Ware-related peoples. I guess the invention of an Indo-Tocharian group may be near…
*NOTE. The presence of R1b-M269 among Proto-Tocharians, as well as the presence of R1b-M269 among Tarim Basin peoples in modern and ancient times is not yet fully discarded. The prevalence of R1a-Z93 may also be the sign of a more recent replacement by Iranian peoples, before the Mongolian and Turkic expansions that probably brought R1b(xM269).
Also, the presence of R1b (xM269) samples in east Asia strengthens the hypothesis of a back-migration of R1b-P297 subclades, from Northern Europe to the east, into the Lake Baikal area, during the Early Mesolithic, as found in the Botai samples and later also in Turkic populations – which are the most likely source of these subclades (and probably also of Q1a2 and N1c) in the region.
Because, if YFull‘s (and Iain McDonald‘s) estimation of the split of R1b-L23 in L51 and Z2103 (ca. 4100 BC, TMRCA ca. 3700 BC) was wrong, by as much as the R1a-Z645 estimates proved wrong, and both subclades were older than expected, then maybe R1b-L51 was not part of the Yamna expansion, but rather part of an earlier expansion with Suvorovo-Novodanilovka into central Europe.
That is, R1b-L51 and R1b-Z2103 would have expanded wih Khvalynsk-Novodanilovka migrants, and they would have either disappeared among local populations, or settled and expanded with successful lineages in certain regions. I think this may give rise to two potential models.
A hidden group in the European east-central steppes?
Here is what Heyd (2011), for example, has to say about the effect of the Khvalynsk-Novodanilovka expansion in the 4th millennium BC, with the first Kurgan wave that shuttered the social, economic, and cultural foundations of south-eastern Europe (before the expansion of west Yamna migrants in the region):
As the Boleraz and Baden tumuli cases in Serbia and Hungary demonstrate, there are earlier, 4th millennium cal. B.C. round tumuli in the Carpathian basin. There are also earlier north-Pontic steppe populations who infiltrated similar environments west of the Black Sea prior to the rise of the Yamnaya culture. This situation can be traced back to the 2nd half of the 5th millennium cal. B.C. to a group of distinct burials, zoomorphic maceheads, long flint blades, triangular flint points, etc., summarized under the term Suvurovo-Novodanilovka (Govedarica 2004; Rassamakin 2004; Anthony 2007; Heyd forthcoming 2011). They also erected round personalized tumuli, though smaller in size and height, above inhumations of single individuals. Suvorovo and Casimcea are the key examples in the lower Danube region of Romania. In northeast Bulgaria, the primary grave of Polska Kosovo (ochre-stained supine extended body position: information communicated by S. Alexandrov) can also be seen as such, as should the Targovishte-“Gonova mogila” primary grave 1 in the Thracian plain with a burial arranged in a supine position with flexed legs, southeast-northwest orientated, and strewed with ochre (Kanchev 1991 , p. 56- 57; Ivanova Gaydarska 2007). In addition to the many copper and shell beads, the 17.4cm long obsidian blade is exceptional, which links this grave to the Csongrád-“Kettoshalom” grave in the south Hungarian plain (Ecsedy 1979). It also yielded an obsidian blade ( 13.2cm long) and copper, shell and limestone beads.
However, no traces of a tumulus have been recorded above the Kettoshalom tomb. Conventionally, it is dated to the Bodrogkeresztur-period in east Hungary, shortly after 4000 cal. B.C., which would correspond very well with the suggested Cernavodă I (or its less known cultural equivalent in the Thracian plain) attribution for the “Gonova mogila” grave, a cultural background to which the Csongrád grave should have also belonged. Bodrogkeresztur and Cernavodă I periods are not the only examples of 4th millennium cal. B.C. tumuli and burials displaying this steppe connection. Indeed we can find this early steppe impact throughout the 4th millennium cal. B.C. These include adscriptions to the Horodiștea II (Corlateni-Dealul Stadole, grave I: Burtanescu l 998, p. 37; Holbocai, grave 34: Coma 1998, p. 16); to Gordinești-Cernavodă 11 (Liești-Movila Arbănașu, grave 22: Brudiu 2000); to Gorodsk-Usatovo (Corlăteni Dealul Cetăţii, grave I: Comșa 1998, p. 17- 18, in Romania; Durankulak, grave 982: Vajsov 2002, in Bulgaria); and to Cernavodă III(Golyama Detelina, tum. 4: Leshtakov, Borisov 1995), and early (end of 4th millennium cal. B.C.) Ezero in Ovchartsi, primary grave (Kalchev 1994, p. 134-138) and Golyama Detelina, tum. 2 (Kanchev 1991) in Bulgaria. Also the Boleráz and Baden tumuli of Banjevac-Tolisavac and Mokrin in the south Carpathian basin account for this, since one should perhaps take into account primary grave 12 of the Sárrédtudavari-Orhalom tumulus in the Hungarian Alfold: a left-sided crouched juvenile ( 15- 17 y) individual in an oval, NW-SE orientated grave pit 14C dated to 3350-3100 cal. B.C. at 2 sigma (Dani, Ncpper 2006). Neither the burial custom (no ochre strewing or depositing a lump of ochre has been recorded), nor date account for its ascription to the Yamnaya!
All of these tumuli and burials demonstrate, though, that there is already a constant but perhaps low-level 4th millennium cal. B.C. steppe interaction, linking the regions of the north of the Black Sea with those of the west, and reaching deep into the Carpathian basin. This has to be acknowledged. even if these populations remain small, bounded to their steppe habitat with an economy adapted to this special environment, and are not always visible in the record. Indirect hints may help in seeing them, such as the frequent occurrence of horse bones, regarded as deriving from domesticated horses, in Hungarian Baden settlements (Bokonyi 1978; Benecke 1998), and in those of the south German Cham Culture (Matuschik 1999, p. 80-82) and the east German Bernburg Culture (Becker 1999; Benecke 1999). These occur, however, always in low numbers, perhaps not enough to maintain and regenerate a herd. Does this point us towards otherwise archaeologically hidden horsebreeders in the Carpathian basin, before the Yamnaya? In any case, I hope to make one case clear: these are by no means Yamnaya burials in the strict definition! Attribution to the Yamnaya in its strict definition applies.
Also, about the expansion of Yamna settlers along the steppes:
However, it should have been made clear by the distribution map of the Western Yamnaya that they were confining themselves solely to their own, well-known, steppe habitat and therefore not occupying, or pushing away and expelling, the locally settled farming societies. Also, living solely in the steppes requires another lifestyle, and quite different economic and social bases, most likely very different to the established farming societies. Although surely regarded as incoming strangers, they may therefore not have been seen as direct competitors. This argument can be further enforced when remembering that the lowlands and the steppes in the southeast of Europe had already been populated throughout the 4th millennium cal. B.C., as demonstrated above, by societies with a similar north-Pontic steppe origin and tradition, albeit in lower numbers. It is only for these groups that the Yamnaya may have become a threat, but their common origin and perhaps a similar economic/ social background with comparable lifestyles would surely have assisted to allow rapid assimilation. More important, though, is that farming societies in this region may therefore have been accustomed to dealing and interacting with different people and ethnic strangers for a long time. (…)
When assessing farming and steppe societies’ interaction from a general point of view, attitudes can diverge in three main directions:
the violent one; with raids, fights, struggles, warfare, suppression and finally the superiority and exploitation of the one over the other;
the peaceful one; with a continuous exchange of gifts, goods, work, information and genes in a balanced reciprocal system, leading eventually to the merging of the two societies and creation of a new identity;
the neutral one; with the two societies ignoring each other for a long time.
What we see from trying to understand the record of the Yamnaya, based on their tumuli and burials, and the local and neighbouring contemporary societies, based on their settlements, hoards, and graves, is likely a mixture of all three scenarios, with the balance perhaps more towards exchange in a highly dynamic system with alterations over time. However, violence and raids cannot be ruled out; they would be difficult to see in the archaeological record; or only indirectly, such as the building of hill forts, particularly the defence-like chain of Vucedol hillforts along the south shore of the Danube on the Serbian/Croatian border zone (Tasic 1995a), and the retreat of people into them (Falkenstein 1998, p. 261-262), with other interpretations also possible. And finally, we are dealing here with very different local and neighbouring societies, as well as with more distant contemporary ones, looking, in reality, rather like a chequer board of societies and archaeological cultures (see Parzinger 1993 for the overview). These display different regional backgrounds and traditions leading to different social and settlement organizations, different economic bases and material cultures in the wide areas between Prut and Maritza rivers, and Black Sea and Tisza river. They surely found their individual way of responding to the incoming and settling Yamnaya people.
The best data we have about this potential non-Yamna origin of R1b-L51 – and thus in favour of its admixture in the Carpathian basin – lies in:
The majority of R1a-Z2103 subclades found to date among Yamna samples.
The limited presence of (ancient and modern) R1b-L51 in eastern Europe and India, whose isolated finds are commonly (and simplistically) attributed to ‘late migrations’.
The presence of R1b-L51 (xZ2103) in cultures related to the ‘Yamna package’, but supposedly not to Yamna settlers. So for example I7043, of haplogroup R1b-L151(xU106,xP312), ca. 2500-2200 BC from Szigetszentmiklós-Üdülősor, probably from the Bell Beaker (Csepel group), but maybe from the early Nagýrev culture.
The expansion of its subclades apparently only from a single region, around the Carpathian basin, in contrast to R1b-Z2103.
The already ‘diluted’ steppe admixture found in the earliest samples with respect to Yamna, which points to the appearance after the Yamna admixture with the local population.
Ukrainian archaeologists (in contrast to their Russian colleagues) point to the relevance of North Pontic cultures like Kvitjana and Lower Mikhailovka in the development of Early Yamna in the west, and some eastern European researchers also believe in this similarity.
If R1b-Z2103 and R1b-L51 had expanded with Suvorovo-Novodanilovka migrants to the west, and had admixed later as Hungary_LCA-LBA-like peoples with Yamna migrants during the long-term contacts with other ‘kurganized cultures’ ca. 2900-2500 BC in the Great Hungarian Plains, it could explain some peculiar linguistic traits of North-West Indo-European, and also why R1b-Z2103 appears in cultures associated with this earlier ‘steppe influence’ (i.e. not directly related to Yamna) such as Vučedol (with a R1b-Z2103 sample, see below). That could also explain the presence of R1b-L151(xP312, xU106) in similar Balkan cultures, possibly not directly related to Yamna.
A hidden group among north or west Pontic Eneolithic steppe cultures?
The expansion of Khvalynsk as Novodanilovka into the North Pontic area happened through the south across the steppe, near the coast, with the forest-steppe region working as a clear natural border for this culture of likely horse-riding chieftains, whose economy was probably based on some rudimentary form of mobile pastoralism.
Although archaeologists are divided as to the origin of each individual Middle Eneolithic group near the Black Sea after the end of the Khvalynsk-Novodanilovka period, it seems more or less clear that steppe cultures like Cernavodă, Lower Mikhailovka, or Kvitjana are closer (or “more archaic”) in their steppe features, which connects them to Volga–Ural and Northern Caucasus cultures, like Northern Caucasus, Repin or Khvalynsk.
On the other hand, forest-steppe cultures like Dereivka (including Alexandria) show innovative traits and contacts with para- or sub-Neolithic cultures to the north, like Comb-Pit Ware groups, apart from corded decoration influenced by Trypillian groups to the west, especially in their later (‘Proto-Corded Ware‘) stage after ca. 3500 BC.
If Ukrainian researchers like Rassamakin are right, Early Yamna expanded not only from Repin settlers, but also from local steppe cultures adopting Repin traits to develop an Early Yamna culture, similar to how eastern (Volga–Ural groups) seem to have synchronously adopted Early Yamna without massive affluence of Repin settlements.
Furthermore, local traits develop in southern groups, like anthropomorphic stelae (shared with Kemi-Oba, direct heir of Lower Mikhailovka), and rich burials featuring wagons. These traits are seen in west Yamna settlers.
Problems of this model include:
On the North Pontic area – in contrast to the Volga–Ural region – , there was a clear “colonization” wave of Repin settlers, also supported by Ukrainian researchers, based on the number of new settlements and burials, and on the progressive retreat of Dereivka, Kvitjana, as well as (more recent) Maykop- and Trypillia-related groups from the North Pontic area ca. 3350/3300 BC. It seems unlikely that these expansionist, semi-nomadic, cattle-breeding, patrilineally-related steppe clans that were driving all native populations out of their territories suddenly decided, at some point during their spread into the North Pontic area ca. 3300-3100 BC, to join forces with some foreign male lineages from the area, and then continue their expansion to the west…
Similar to the fate of R1b-P297 subclades in the Baltic after the expansion of Corded Ware migrants, previous haplogropus of the North Pontic region – such as R1a, R1b-V88, and I2 subclades basically disappeared from the ancient DNA record after the expansion of Khvalynsk-Novodanilovka, and then after the expansion of Yamna, as is clear from Yamna, Afanasevo, and Bell Beaker samples obtained to date. This, in combination with what we know about Y-chromosome bottlenecks in post-Neolithic expansions, leaves little space to think that a big enough territorial group with a majority of “native” haplogroups could survive later expansions (be it R1b-L51 or R1a-Z645).
Supporting an expansion of the same male (and partly female) population, the Yamna admixture from east to west is quite homogeneous, with the only difference found in (non-significant) EEF-like proportion which becomes elevated in distant areas [apart from significant ‘southern’ contribution to certain outlier samples]. Based on the also homogeneous Y-DNA picture, the heterogeneity must come, in general, from the female exogamy practiced by expanding groups.
There is a short period, spanning some centuries (approximately 3300-2700 BC), in which the North Pontic area – especially the forest-steppe territories to the west of the Dnieper, i.e. the Upper Dniester, Boh, and Prut-Siret areas – are a chaos of incoming and emigrating, expanding and shrinking groups of different cultures, such as late Trypillian groups, Maykop-related traits, TRB, GAC, (Proto-)Corded Ware, and Early Yamna settlements. No natural geographic frontier can be delimited between these groups, which probably interacted in different ways. Nevertheless, based on their cultural traits, admixture, and especially on their Y-DNA, it seems that they never incorporated foreign male lineages, beyond those they probably had during their initial expansion trends.
The further expansionist waves of Early Yamna seen ca. 3100 BC, from the Danube Delta to the west, give an overall image of continuously expanding patrilineal clans of R1b-M269 subclades since the Khvalynsk-Novodanilovka migration, in different periodic steps, mostly from eastern Pontic-Caspian nuclei, usually overriding all encountered cultures and (especially male) populations, rather than showing long-term collaboration and interaction. Such interaction is seen only in exceptional cases, e.g. the long-term admixture between Abashevo and Poltavka, as seen in Proto-Indo-Iranian peoples and their language.
We are living right now an exemplary ego-, (ethno-)nationalism-, and/or supremacy-deflating moment, for some individuals of eastern and northern European descent who believed that R1a or ‘steppe ancestry proportions’ meant something special. The same can be said about those who had interiorized some social or ethnolinguistic meaning for the origin of R1b in western Europe, N1c in north-eastern Europe, as well as Greeks, Iranians, Armenians, or Mediterranean peoples in general of ‘Near Eastern’ ancestry or haplogroups, or peoples of Near Eastern origin and/or language.
These people had linked their haplogroups or ancestry with some fantasy continuity of ‘their’ ancestral populations to ‘their’ territories or languages (or both), and all are being proven wrong.
Apart from teaching such people a lesson about what simplistic views are useful for – whether it is based on ABO or RH group, white skin, blond hair, blue eyes, lactase persistence, or on the own ancestry or Y-DNA haplogroup -, it teaches the rest of us what can happen in the near future among western Europeans. Because, until recently, most western Europeans were comfortably settled thinking that our ancestors were some remnant population from an older, Palaeolithic or Mesolithic population, who acquired Indo-European languages by way of cultural diffusion in different periods, including only minor migrations.
Judging by what we can see now among some individuals of Northern and Eastern European descent, the only thing that can worsen the air of superiority among western Europeans is when they realize (within a few years, when all these stupid battles to control the narrative fade) that not only are they the cultural ‘heirs’ of the Graeco-Roman tradition that began with the Roman Empire, but that most of them are the direct patrilineal descendants of Khvalynsk, Yamna, Bell Beaker, and European Bronze Age peoples, and thus direct descendants of Middle PIE, Late PIE, and NWIE speakers.
The finding of R1b-L51 and R1b-Z2103 among expanding Suvorovo-Novodanilovka chieftains, with pockets of R1b-L51 remaining in steppe-like societies of the Balkans and the Carpathian Basin, would have beautifully complemented what we know about the East Yamna admixture with R1a-Z93 subclades (Uralic speakers) ca. 2600-2100 BC to form Proto-Indo-Iranian, and about the regional admixtures seen in the Balkans, e.g. in Proto-Greeks, with the prevalent J subclades of the region.
It would have meant an end to any modern culture or nation identifying themselves with the ‘true’ Late PIE and Yamna heirs, because these would be exclusively associated with the expansion of R1b-Z2103 subclades with late Repin, and later as the full-fledged Late PIE with Yamna settlers to south-east and central Europe, and to the southern Urals. The language would have had then obviously undergone different language changes in all these territories through long-lasting admixture with other populations. In that sense, it would have ended with the ideas of supremacy in western Europe before they even begin.
The most likely future
However limited the evidence, it seems that R1b-L51 expanded with Yamna, though, based on the estimates for the haplogroups involved, and on marginal hints at the variability of L23 subclades within Yamna and neighbouring populations. If R1b-L51 expanded with West Repin / Early Yamna settlers, this is why they have not yet been found among Yamna samples:
The subclade division of Yamna settlers needs not be 50:50 for L51:Z2103, either in time or in space. I think this is the simplistic view underlying many thoughts on this matter. Many different expanding patrilineal clans of L23 subclades may have been more or less successful in different areas, and non-Z2103 may have been on the minority, or more isolated relative to Z2103-clans among expanding peoples on the steppe, especially on the east. In fact, we usually talk in terms of “Z2103 vs. L51” as if
these two were the only L23 subclades; and
both had split and succeeded (expanding) synchronously;
that is, as if there had not been multiple subclades of both haplogroups, and as if there had not been different expansion waves for hundreds of years stemming from different evolving nuclei, involving each time only limited (successful) clans. Many different subclades of haplogroups L23 (xZ2103, xL51), Z2103, and L51 must have been unsuccessful during the ca. 1,500 years of late Khvalynsk and late Repin-Early Yamna expansions in which they must have participated (for approximately 60-75 generations, based on a mean 20-25 years).
If we want to imagine a pocket of ‘hidden’ L51 for some region of the North Pontic or Carpathian region, the same can be imagined – and much more likely – for any unsampled territory of expanding late Repin/Early Yamna settlers from the Lower Don – Lower Volga region (probably already a mixed society of L51 and Z2103 subclades since their beginning, as the early Repin culture, ca. 3800 BC), with L51 clans being probably successful to the west.
The Repin culture expanded only in small, mobile settlements from the Lower Don – Lower Volga to the north, east, and south, starting ca. 3500/3400 BC, in the waves that eventually gave a rather early distant offshoot in the Altai region, i.e. Afanasevo. Starting ca. 3300 BC in the archaeological record, the majority of R1b-Z2103 subclades found to date in Afanasevo also supports either
a mixed Repin society, with Z2103-clans predominating among eastern settlers; or
a Repin society marked by haplogroup L51, and thus a cultural diffusion of late Repin/Early Yamna traits among neighbouring (Khvalynsk, Samara, etc.) groups of essentially the same (early Khvalynsk-Novodanilovka) genetic stock in the Volga–Ural region.
Both options could justify a majority of Z2103 in the Lower Volga–Ural region, with the latter being supported by the scattered archaeological remains of late Repin in the region before the synchronous emergence of Early Yamna findings in the whole Pontic-Caspian steppe.
Most Z2103 from Yamna samples to date are from around 3100 BC (in average) onward, and from the right bank of the Lower Don to the east, particularly from the Lower Volga–Ural area (especially the Samara region), which – based on the center of expansion of late Repin settlers – may be depicting an artificially high Z2103-distribution of the whole Yamna community.
Yamna sample I0443, R1b-L23 (Y410+, L51-), ca. 3300-2700 BCE from Lopatino II, points to an intermediate subclade between L23 and L51, near one of the supposed late Repin sites (based on kurgan burials with late Repin cultural traits) in the Samara region.
Other Balkan cultures potentially unrelated to the Yamna expansion also show Z2103 (and not only L51) subclades, like I3499 (ca. 2884-2666 calBC), of the Vučedol culture, from Beli Manastir-Popova zemlja, which points to the infiltration of Yamna peoples in other cultures. In any case, the appearance of R1b-L23 subclades in the region happens only after the Yamna expansion ca. 3100 BC, probably through intrusions into different neighbouring regions, if these Balkan cultures are not directly derived from Yamna settlements (which is probably the case of the Csepel Bell Beaker or early Nagýrev sample, see above).
The diversity of haplogroups found in or around the Carpathian Basin in Late Chalcolithic / Early Bronze Age samples, including L151(xP312, xU106), P312, U106, Z2103, makes it the most likely sink of Yamna settlers, who spread thus with expanding family clans of different R1b-L23 subclades.
Even though some Yamna vanguard groups are known to have expanded up to Saxony-Anhalt before ca. 2700 BC, haplogroup Z2103 seems to be restricted to more eastern regions, which suggests that R1b-L51 was already successful among expanding West Yamna clans in Hungary, which gave rise only later to expanding East Bell Beakers (overwhelmingly of L151 subclades). The source of R1b-L51 and L151 expansion over Z2103 must lie therefore in the West Yamna period, and not in the Bell Beaker expansion.
The R1b-Z2103 found in Poltavka, Catacomb, and to the south point to a late migration displacing the western R1b-L51, only after the late Repin expansion. This is also seen in the steppe ancestry and R1b-Z2103 south of the Caucasus, in Hajji Firuz, which points to this route as a potential source of the supposed “Earliest Proto-Indo-Iranian” (the mariannu term) of the Near East. A similar replacement event happened some centuries later with expanding R1a-Z93 subclades from the east wiping out haplogroup R1b-Z2103 from the Pontic-Caspian steppe.
Many ancient samples from Khvalynsk, Northern Caucasus, Yamna, or later ones are reported simply as R1b-M269 or L23, without a clear subclade, so the simplistic ‘Yamna–Z2103’ picture is not real: if one takes into account that Z2103 might have been successful quite early in the eastern region, it is more likely to obtain a successful Y-SNP call of a Z2103 subclade in the Volga–Ural region than a xZ2103 one.
‘Western’ features described by archaeologists for West Yamna settlers, associated with Kemi Oba and southern Yamna groups in the North Pontic area – like rich burials with anthropomorphic stelae and wagons – are actually absent in burials from settlers beyond Bulgaria, which does not support their affiliation with these local steppe groups of the Black Sea. Also, a mix with local traditions is seen accross all Early Yamna groups of the Pontic-Caspian steppe, and still genetics and common cultural traits point to their homogeneization under the same patrilineal clans expanding continuously for centuries. The maintenance of local traditions (as evidenced by East Bell Beakers in Iberia related to Iberian Proto-Beakers) is often not a useful argument in genetics, especially when the female population is not replaced.
Middle Proto-Indo-European expanded with Khvalynsk-Novodanilovka after ca. 4800 BC, with the first Suvorovo settlements dated ca. 4600 BC.
Archaic Late Proto-Indo-European expanded with late Repin (or Volga–Ural settlers related to Khvalynsk, influenced by the Repin expansion) into Afanasevo ca. 3500/3400 BC.
Late Proto-Indo-European expanded with Early Yamna settlers to the west into central Europe and the Balkans ca. 3100 BC; and also to the east (as Pre-Proto-Indo-Iranian) into the southern Urals ca. 2600 BC.
North-West Indo-European expanded with Yamna Hungary -> East Bell Beakers, from ca. 2500 BC.
Proto-Indo-Iranian expanded with Sintashta, Potapovka, and later Andronovo and Srubna from ca. 2100 BC.
It seems that the subclades from Khvalynsk ca. 4250-4000 BC were wrongly reported – like those of Narasimhan et al. (2018). However, even if they are real and YFull estimates have to be revised, and even if the split had happened before the expansion of Suvorovo-Novodanilovka, the most likely origin of R1b-L51 among Bell Beakers will still be the expansion of late Repin / Early Yamna settlers, and that is what ancient DNA samples will most likely show, whatever the social or political consequences.
The only relevance of the finding of R1b-L51 in one place or another – especially if it is found to be a remnant of a Middle PIE expansion coupled with centuries of admixture and interaction in the Carpathian Basin – is the potential influence of an archaic PIE (or non-IE) layer on the development of North-West Indo-European in Yamna Hungary -> East Bell Beaker. That is, more or less like the Uralic influence related to the appearance of R1a-Z93 among Proto-Indo-Iranians, of R1a-Z284 among Pre-Germanic peoples, and of R1a-Z282 among Balto-Slavic peoples.
I think there is little that ancient DNA samples from West Yamna could add to what we know in general terms of archaeology or linguistics at this point regarding Late PIE migrations, beyond many interesting details. I am sure that those who have not attributed some random 6,000-year-old paternal ancestor any magical (ethnic or nationalist) meaning are just having fun, enjoying more and more the precise data we have now on European prehistoric populations.
As for those who believe in magical consequences of genetic studies, I don’t think there is anything for them to this quest beyond the artificially created grand-daddy issues. And, funnily enough, those who played (and play) the ‘neutrality’ card to feel superior in front of others – the “I only care about the truth”-type of lie, while secretly longing for grandpa’s ethnolinguistic continuity – are suffering the hardest fall.
This study included 321 samples from men throughout Corsica; samples from Provence and Tuscany were added to the cohort. All samples were typed for 92 Y-SNPs, and Y-STRs were also analyzed.
Haplogroup R represented approximately half of the lineages in both Corsican and Tuscan samples (respectively 51.8% and 45.3%) whereas it reached 90% in Provence. Sub-clade R1b1a1a2a1a2b-U152 predominated in North Corsica whereas R1b1a1a2a1a1-U106 was present in South Corsica. Both SNPs display clinal distributions of frequency variation in Europe, the U152 branch being most frequent in Switzerland, Italy, France and Western Poland. Calibrated branch lengths from whole Y chromosome sequencing [44,45] and ancient DNA studies  both indicated that R1a and R1b diversification began relatively recently, about 5 Kya, consistent with Bronze Age and Copper Age demographic expansion. TMRCA estimations are concordant with such expansion in Corsica.
Haplogroup G reached 21.7% in Corsica and 13.3% in Tuscany. Sub-clade G2a2a1a2-L91 accounted for 11.3% of all haplogroups in Corsica yet was not present in Provence or in Tuscany. Thirty-four out of the 37 G2a2a1a2-L91 displayed a unique Y-STR profile, illustrated by the star-like profile of STR networks (Fig 1). G2a2a1a2-L91 and G2a2a-PF3147(xL91xM286) show their highest frequency in present day Sardinia and southern Corsica compared to low levels from Caucasus to Southern Europe, encompassing the Near and Middle East [21,47–50]. Ancient DNA results from Early and Middle Neolithic samples reported the presence of haplogroup G2a-P15 [51–53], consistent with gene flow from the Mediterranean region during the Neolithic transition. Td expansion time estimated by STR for P15-affiliated chromosomes was estimated to be 15,082+/-2217 years ago . Ötzi, the 5,300-year-old Alpine mummy, was derived for the L91 SNP . A genetic relationship between G haplogroups from Corsica and Sardinia is further supported by DYS19 duplication, reported in North Sardinia , and observed in the southern part of the Corsica in 9 out of 37 G2a2a1a2-L91 chromosomes and in 4 out of 5 G2a2a-PF3147(xL91xM286) chromosomes, 3 of which displayed an identical STR profile (S4 Table).
This lineage has a reported coalescent age estimated by whole sequencing in Sardinian samples of about 9,000 years ago. This could reflect common ancestors coming from the Caucasus and moving westward during the Neolithic period , whereas their continental counterparts would have been replaced by rapidly expanding populations associated with the Bronze Age [46,54,55]. Estimated TMRCA for L91 lineage in Corsica is 4529 +/- 853 years. G-L497 showed high frequencies in Corsica compared to Provence and Tuscany, and this haplogroup was common in Europe, but rare in Greece, Anatolia and the Middle East. Fifteen out of the 17 Corsican G2a2b2a1a1b-L497 displayed a unique Y-STR profile (S4 Table) with an estimated TMRCA of 6867 +/- 1294 years. Haplogroup G2a2b1-M406, associated with Impressed Ware Neolithic markers, along with J2a1-DYS445 = 6 and J2a1b1-M92 [22,49], had very low levels in Corsica. Conversely, G2a2b2a-P303was highly represented and seemed to be independent of the G2a2b1-M406 marker. The 7 G2a2b2a-P303(xL497xM527) Corsican chromosomes displayed a unique Y-STR profile (S4 Table).
Haplogroup J, mainly represented by J2a1b-M67(xM92), displayed intermediate frequencies in Corsica compared to Tuscany and Provence. J2a1b-M67(xM92) derived STR network analysis displayed a quite homogeneous profile across the island with an estimated TMRCA of 2381 +/- 449 years (Fig 1) and individuals displaying M67 were peripheral compared to Northwestern Italians (S2 Fig). The haplogroup J2a1-Page55(xM67xM530), characteristic of non-Greek Anatolia , was found in the north-west of Corsica. Haplogroup J2a1-DYS445 = 6 was found in the north-west with DYS391 = 10 repeats, and in the far south with DYS391 = 9 repeats, the former was associated with Anatolian Greek samples, whereas the second was found in central Anatolia . The 7 J2b2a-M241 displayed a unique Y-STR profile (S4 Table), they were only detected in the Cap Corse region, this sub-haplogroup shows frequency peaks in both the southern Balkans and northern-central Italy  and is associated with expansion from the Near East to the Balkans during Neolithic period .
Haplogroup E, mainly represented by E1b1b1a1b1a-V13, displayed intermediate frequencies in Corsica compared to Tuscany and Provence. E1b1b1a1b1a-V13 was thought to have initiated a pan-Mediterranean expansion 7,000 years ago starting from the Balkans  and its dispersal to the northern shore of the Mediterranean basin is consistent with the Greek Anatolian expansion to the western Mediterranean , characteristic of the region surrounding Alaria, and consistent with the TMRCA estimated in Corsica for this haplogroup. A few E1b1a-V38 chromosomes are also observed in the same regions as V13.