Bell Beakers and Mycenaeans from Yamnaya; Corded Ware from the forest steppe


I have recently written about the spread of Pre-Yamnaya or Yamnaya ancestry and Corded Ware-related ancestry throughout Eurasia, using exclusively analyses published by professional geneticists, and filling in the gaps and contradictory data with the most reasonable interpretations. I did so consciously, to avoid any suspicion that I was interspersing my own data or cherry picking results.

Now I’m finished recapitulating the known public data, and the only way forward is the assessment of these populations using the available datasets and free tools.

Understanding the complexities of qpAdm is fairly difficult without a proper genetic and statistical background, which I won’t pretend to have, so its tweaking to get strictly correct results would require an unending game of trial and error. I have sadly little time for this, even taking my tendency to procrastination into account… so I have used a simple model akin to those published before – in particular, the outgroup selection by Ning, Wang et al. (2019), who seem to be part of the only group interested in distinguishing Yamnaya-related from Corded Ware-related ancestry, probably the most relevant question discussed today in population genomics regarding the Proto-Indo-European and Proto-Uralic homelands.

Supplementary Table 13. P values of rank=2 and admixture proportions in modelling Steppe ancestry populations as a three-way admixture of Eneolithic steppe Anatolian_Neolithic and WHG using 14 outgroups.
Left populations: Test, Eneolithic_steppe, Anatolian_Neolithic, WHG.
Right populations: Mbuti.DG, Ust_Ishim.DG, Kostenki14, MA1, Han.DG, Papuan.DG, Onge.DG, Villabruna, Vestonice16, ElMiron, Ethiopia_4500BP.SG, Karitiana.DG, Natufian, Iran_Ganj_Dareh_Neolithic.

I have used for all analyses below a merged dataset including the curated one of the Reich Lab, the latest on Central and South Asia by Narasimhan, Patterson et al. (2019), on Iberia by Olalde et al. (2019), and on the East Baltic by Saag et al. (2019), as well as datasets including samples from Wang et al. (2019) and Lamnidis et al. (2018). I used (and intend to use) the same merged dataset in all cases, despite its huge size, to avoid adding one more uncontrolled variable to the analyses, so that all results obtained can be compared.

I try to prepare in advance a bunch of relevant files with left pops and right pops for each model:

  1. It seems a priori more reasonable to use geographically and chronologically closer proxy populations (say, Trypillia or GAC for Steppe-related peoples) than hypothetic combinations of ancestral ones (viz. Anatolian farmer, WHG, and EHG).
  2. This also means using subgroups closer to the most likely source population, such as (Don-Volga interfluve) Yamnaya_Kalmykia rather than (Middle Volga) Yamnaya_Samara for the western expansion of late Repin/early Yamnaya, or the early Germany_Corded_Ware.SG or Czech_Corded Ware for the group closest to the Proto-Corded Ware population (see below), likely neighbouring the Upper Vistula region.
  3. I usually test two source populations for different targets, which seems like a much more efficient way of using computer resources, whenever I know what I want to test, since I need my PC back for its normal use; whenever I don’t know exactly what to test, I use three-way admixture models and look for subsets to try and improve the results.

I have probably left out some more complex models by individualizing the most relevant groups, but for the time being this would have to do. Also, no other formal stats have been used in any case, which is an evident shortcoming, ruling out an interpretation drawn directly and only from the results below.

Full qpAdm results for each batch of samples are presented in a Google Spreadsheet, with each tab (bottom of the page) showing a different combination of sources, usually in order of formally ‘best’ (first to the left) to ‘worst’ (last to the right) fits, although the order is difficult to select in highly heterogeneous target groups, as will be readily visible.

Disintegration, migration, and imports of the Azov–Black Sea region. First migration event (solid arrows): Gordineşti–Maikop expansion (groups: I – Bursuchensk; II – Zhyvotylivka; III – Vovchans’k; IV – Crimean; V – Lower Don; VI – pre-Kuban). Second migration event (hollow arrows): Repin expansion. After Rassamakin (1999), Demchenko (2016).

Corded Ware origins

The latest publications on the Yampil barrow complex have not improved much our understanding of the complexity of Corded Ware origins from an archaeological point of view, involving multiple cultural (hence likely population) influences. This bit is from Ivanova et al., Baltic-Pontic Studies (2015) 20:1, and most hypotheses of the paper remain unanswered (except maybe for the relevance of the Złota group):

In the light of the above outline therefore one should argue that the ‘architecture of barrows’ associated in the ‘Yampil landscape’ of the Middle Dniester Area with the Eneolithic (specifically, mainly with the TC), precedes the development of a similar phenomenon that can be observed from 2900/2800 BC in the Upper Dniester Area and drainage basin of the Upper Vistula, associated with the CWC [Goslar et al. 2015; Włodarczak 2006; 2007; 2008; Jarosz, Włodarczak 2007]. The most consuming research question therefore is whether ritual customs making use of Eneolithic (Tripolye) ‘barrow architecture’ could have penetrated northwards along the Dniester route, where GAC communities functioned. One could also ask what role the rituals played among the autochthons [Kośko 2000; Włodarczak 2008; 2014: 335; Ivanova, Toshchev 2015b].

This issue has already been discussed with a resulting tentative systemic taxonomy in the studies of Włodarczak, arguing for the Złota culture (ZC) in the Vistula region as an illustration of one of the (Małopolska) reception centres of civilization inspirations from the oldest Pontic ‘barrow culture’ circle associated with the Eneolithic and Early Bronze Age [Włodarczak 2008]. Notably, it is in the ZC that one can notice a set of cultural traits (catacomb grave construction, burial details, forms and decoration of vessels) analogous to those shared by the north-western Black Sea Coast groups of the forest-steppe Eneolithic (chiefly Zhyvotilovka-Volchansk) and the Late Tripolye circle (chiefly Usatovo-Gordinești-Horodiștea-Kasperovtsy).

Globular Amphorae culture „exodus” to the Danube Delta: a – Globular Amphorae culture; b – GAC (1), Gorodsk (2), Vykhvatintsy (3) and Usatovo (4) groups of Trypillia culture; c – Coţofeni culture; d – northern border of the late phase of Baden culture;red arrows – direction of Globular Amphora culture expansion; blue arrow – direction of „reflux” of Globular Amphora culture (apud Włodarczak, 2008, with changes).

Taking into account that I6561 might be wrongly dated, we cannot include the Corded Ware-like sample of the end-5th millennium BC in the analysis of Corded Ware origins. That uncertainty in the chronology of the appearance of “Steppe ancestry” in Proto-Corded Ware peoples complicates the selection of any potential source population from the CHG cline.

Nevertheless, the lack of hg. R1a-M417 and sizeable Pre-Yamnaya-related ancestry in the sampled Pontic forest-steppe Eneolithic populations (represented exclusively by two samples from Dereivka ca. 3600-3400 BC) would leave open the interesting possibility that a similar ancestry got to the forest-steppe region between modern Poland and Ukraine during the known complex population movements of the Late Eneolithic.

It is known that Corded Ware-derived groups and Steppe Maykop show bad fits for Pre-Yamnaya/Yamnaya ancestry, and also that Steppe Maykop is a potential source of “Steppe-related ancestry” within the Eneolithic CHG mating network of the Pontic-Caspian steppes and forest-steppes. Testing Corded Ware for recent Trypillia and Maykop influences, proper of Late Trypillia and Late Maykop groups in the North Pontic area (such as Zhyvotylivka–Vovchans’k and Gordineşti) side by side with potential Pre-Yamnaya and Yamnaya sources makes thus sense:

Now, the main obvious difference between Khvalynsk-Yamnaya and Corded Ware is the long-lasting, pervasive Y-chromosome bottlenecks under R1b lineages in the former, compared to the haplogroup variability and late bottleneck under R1a-M417 in the latter, which speaks in favour – on top of everything else – of a different community of sub-Neolithic hunter-gatherers including hg. R1a-M417 hijacking the expansion of Steppe_Maykop-related ancestry around the Volhynian-Podolian Upland.

Akin to how Yamnaya patrilineal descendants hijacked regional EEF (±CWC) ancestry components mainly through exogamy, dragging them into the different expanding Bell Beaker groups (see below), but kept their Indo-European languages, these hunter-gatherers that admixed with peoples of “Steppe ancestry” were the most likely vector of expansion of Uralic languages in Eastern Europe.

PCA of ancient Eurasian samples. Marked likely Proto-Corded Ware samples and potential origin of its PCA cluster based on qpAdm results. See full PCA and more related files.

Baltic Corded Ware

One of the most interesting aspects of the results above is the surprising heterogeneity of the different regional groups, which is also reflected in the Y-DNA variability of early Corded Ware samples.

Seeing how Baltic CWC groups, especially the early Latvia_LN sample, show particularly bad fits with the models above, it seems necessary to test how this population might have come to be. My first impression in 2017 was that they could represent early Corded Ware groups admixed with Yamnaya settlers through their interactions along the Dnieper-Dniester corridor.

However, I recently predicted that the most likely admixture leading to their ancestry and PCA cluster would involve a Corded Ware-like group and a group related to sub-Neolithic cultures of eastern Europe, whose best proxy to date are EHG-like Khvalynsk samples (i.e. excluding the outlier with Pre-Yamnaya ancestry, I0434):

Detail of the PCA of the Corded Ware expansion. See full PCA and more related files.

Late Corded Ware + Yamnaya vanguard

Relevant are also the mixtures of Corded Ware from Esperstedt, and particularly those of the sample I0104, which I have repeated many times in this blog I suspected to be influenced by vanguard Yamnaya settlers:

The infeasible models of CWC + Yamnaya_Kalmykia ± Hungary_Baden (see below for Bell Beakers) and the potential cluster formed with other samples from the Baltic suggest that it could represent a more complex set of mixtures with sub-Neolithic populations. On the other hand, its location in Germany, late date (ca. 2500 BC or later), and position in the PCA, together with the good fits obtained for Germany_Beaker as a source, suggest that the increase in Steppe-related ancestry + EEF makes it impossible for the model (as I set it) to directly include Yamnaya_Kalmykia, despite this excess Steppe-related ancestry actually coming from Yamnaya vanguard groups.

I think it is very likely that the future publication of EEF-admixed Yamnaya_Hungary samples (or maybe even Yamnaya vanguard samples) will improve the fits of this model.

These results confirm at least the need to distrust the common interpretation of mixtures including late Corded Ware samples from Esperstedt (giving rise to the “up to 75% Yamnaya ancestry of CWC” in the 2015 papers) as representative of the Corded Ware culture as a whole, and to keep always in mind that an admixture of European BA groups including Corded Ware Esperstedt as a source also includes East BBC-like ancestry, unless proven otherwise.

Yamnaya vanguard groups in Corded Ware territory before the expansion of Bell Beakers (ca. 2500 BC). See full map.

Bell Beaker expansion

A hotly (re)debated topic in the past 6 months or so, and for all the wrong reasons, is the origin of the Bell Beaker folk. Archaeology, linguistics, and different Y-chromosome bottlenecks clearly indicate that Bell Beakers were at the origin of the North-West Indo-European expansion in Europe, while the survival of Corded Ware-related groups in north-eastern Europe is clearly related to the expansion of Uralic languages.

NOTE. For the interesting case of Proto-Indo-Iranians expanding with Corded Ware-like ancestry, see more on the formation of Sintashta-Potapovka-Filatovka from East Uralic-speaking Abashevo and Pre-Proto-Indo-Iranian-speaking Poltavka herders. See also more on R1a in Indo-Iranians and on the social complexity of Sintashta.

Nevertheless, every single discarded theory out there seems to keep coming back to life from time to time, and a new wave of interest in “Bell Beaker from the Single Grave culture” somehow got revived in the process, too, because this obsession – unlike the “Bell Beakers from Iberia Chalcolithic” – is apparently acceptable in certain circles, for some reason.

We know that Iberian Beakers, British Beakers, or Sicilian EBA – representing the most likely closest source population of speakers of Proto-Galaico-Lusitanian, Pre-Celtic Indo-European, and Proto-Elymian, respectively – have already been successfully tested for a direct origin among Western European Beakers in Olalde et al. (2018), Olalde et al. (2019), and Fernandes et al. (2019).

This success in ascertaining a closer Beaker source is probably due to the physical isolation of the specific groups (related to Germany_Beaker, Netherlands_Beaker, and NE_Mediterranean_Beaker samples, respectively) after their migration into regions dominated by peoples without Steppe-related ancestry. Furthermore, Celtic-speaking populations expanding with Urnfield south of the Pyrenees also show a good fit with a source close to France_Beaker.

So I decided to test sampled Bell Beaker populations, to see if it could shed light to the most likely source population of individual Beaker groups and the direction of migration within Central Europe, i.e. roughly eastwards or westwards. As it was to be expected for closely related populations (see the relevant discussion here), an attempt to offer a simplistic analysis of direction based on formal stats does not make any sense, because most of the alternative hypotheses cannot be rejected:

Not only because of the similar values obtained, but because it is absurd to take p-values as a measure of anything, especially when most of these conflicting groups with slightly ‘better’ or ‘worse’ p-values represent multiple different mixtures of the type (Yamnaya + EEF) + (Corded Ware + EEF ± Yamnaya), impossible to distinguish without selecting proper, direct ancestral populations…

A further example of how explosive the Bell Beaker expansion was into different territories, and of their extensive local admixture, is shown by the unsuccessful attempt by Olalde et al. (2018) to obtain an origin of the EEF source for all Beaker groups (excluding Iberian Beakers):

Investigating the genetic makeup of Beaker-complex-associated individuals. Testing different populations as a source for the Neolithic ancestry component in Beaker-complex-associated individuals. The table shows P values (* indicates values > 0.05) for the fit of the model: ‘Steppe_EBA + Neolithic/Copper Age’ source population.
Map of attested Yamnaya pit-grave burials in the Hungarian plains; superimposed in shades of blue are common areas covered by floods before the extensive controls imposed in the 19th century; in orange, cumulative thickness of sand, unfavourable loamy sand layer. Marked are settlements/findings of Boleráz (ca. 3500 BC on), Baden (until ca. 2800 BC), Kostolac (precise dates unknown), and Yamna kurgans (from ca. 3100/3000 BC on).

Now, there is a simpler way to understand what kind of Steppe-related ancestry is proper of Bell Beakers. I tested two simple models for some Beaker groups: Yamnaya + Hungary Baden vs. Corded Ware + GAC Poland. After all, the Bell Beaker folk should prefer a source more closely related to either Yamnaya Hungary or Central European Corded Ware:

Interestingly, models including Yamnaya + Baden show good fits for the most important groups related to North-West Indo-Europeans, including Bell Beakers from Germany, the Netherlands, Italy, and Poland, representing the most likely closest source populations of speakers of Pre-Proto-Celtic, Pre-Proto-Germanic, Proto-Italo-Venetic, and Pre-Proto-Balto-Slavic, respectively.

The admixed Yamnaya samples from Hungary that will hopefully be published soon by the Jena Lab will most likely further improve these fits, especially in combination with intermediate Chalcolithic populations of the Middle and Upper Danube and its tributaries, to a point where there will be an absolute chronological and geographical genomic trail from the fully Yamnaya-like Yamnaya settlers from Hungary to all North-West Indo-European-speaking groups of the Early Bronze Age.

The only difference between groups will be the gradual admixture events of their source Beaker group with local populations on their expansion paths, including peoples of mainly EEF, CWC+EEF, or CWC+EEF+Yamnaya related ancestry. There is ample evidence beyond ancestry models to support this, in particular continued Y-DNA bottlenecks under typical Yamnaya paternal lineages, mainly represented by R1b-L51 subclades.

Distribution of the Bell Beaker East Group, with its regional provinces, as of c. 2400 cal BC (after Heyd et al. 2004, modified). See full maps.

European Early Bronze Age

European EBA groups that might show conflicting results due to multiple admixture events with Corded Ware-related populations are the Únětice culture and the Nordic Late Neolithic.

The results for Únětice groups seem to be in line with what is expected of a Central European EBA population derived from Bell Beakers admixed with surrounding poulations of East Bell Beaker and/or late (Epi-)Corded Ware descent.

Potential models of mixture for Nordic Late Neolithic samples – despite the bad fits due to the lack of direct ancestral CWC and BBC groups from Denmark – seem to be impossible to justify as derived exclusively from Single Grave or (even less) from Battle Axe peoples, supporting immigration waves of Bell Beakers from the south and further admixture events with local groups through maritime domination.

PCA of ancient European samples. Marked are Bronze Age clusters. See full PCAs.

Balkans Bronze Age

The potential origin of the typical Corded Ware Steppe-related ancestry in the social upheaval and population movements of the Dnieper-Dniester forest-steppe corridor during the 4th millennium BC raises the question: how much do Balkan Bronze Age groups owe their ancestry to a population different than the spread of Pre-Yamnaya-like Suvorovo-Novodanilovka chieftains? Furthermore, which Bronze Age groups seem to be more likely derived exclusively from Pre-Yamnaya groups, and which are more likely to be derived from a mixture of Yamnaya and Pre-Yamnaya? Do the formal stats obtained correspond to the expected results for each group?

Since the expansion of hg. I2a-L699 (TMRCA ca. 5500 BC) need not be associated with Yamnaya, some of these values – together with the assessment of each individual archaeological culture – may question their origin in a Yamnaya-related expansion rather than in a Khvalynsk-related one.

NOTE. These are the last ones I was able to test yesterday, and I have not thought these models through, so feel free to propose other source and target groups. In particular, complex movements through the North Pontic area during the Late Eneolithic would suggest that there might have been different Steppe-ancestry-related vs. EEF-related interactions in the north-west and west Pontic area before and during the expansion of Yamnaya.


One of the key Indo-European populations that should be derived from Yamnaya to confirm the Steppe hypothesis, together with North-West Indo-Europeans, are Proto-Greeks, who will in turn improve our understanding of the preceding Palaeo-Balkan community. Unfortunately, we only have Mycenaean samples from the Aegean, with slight contributions of Steppe-related ancestry.

Still, analyses with potential source populations for this Steppe ancestry show that the Yamnaya outlier from Bulgaria is a good fit:

The comparison of all results makes it quite evident the why of the good fits from (Srubnaya-related) Bulgaria_MLBA I2163 or of Sintashta_MLBA relative to the only a priori reasonable Yamnaya and Catacomb sources: it is not about some hypothetical shared ancestor in Graeco-Aryan-speaking East Yamnaya– or even Catacomb-Poltavka-related groups, because all available Yamnaya-related peoples are almost indistinguishable from each other (at least with the sampling available today). These results reflect a sizeable contribution of similar EEF-related populations from around the Carpathians in both Steppe-related groups: Corded Ware and Yamnaya settlers from the Balkans.

Cultural groups in and around the Balkans during the Early Bronze Age. See full maps.

qpAdm magic

In hobby ancestry magic, as in magic in general, it is not about getting dubious results out of thin air: misdirection is the key. A magician needs to draw the audience attention to ‘remarkable’ ancestry percentages coupled with ‘great’ (?) p-values that purportedly “prove” what the audience expects to see, distracting everyone from the true interesting aspects, like statistical design, the data used (and its shortcomings), other opposing models, a comparison of values, a proper interpretation…you name it.

I reckon – based on the examples above – that the following problems lie at the core of bad uses of qpAdm:

  1. In the formal aspect, the poor understanding of what p-values and other formal stats obtained actually mean, and – more importantly – what they don’t mean. The simplistic trend to accept results of a few analyses at face value is necessarily wrong, in so far as there is often no proper reasoning of what is being assessed and how, and there is never a previous opinion about what could be expected if the alternative hypotheses were true.
  2. In the interpretation aspect, the poor judgement of accompanying any results with simplistic, superficial, irrelevant, and often plainly wrong archaeological or linguistic data selected a posteriori; the inclusion of some racial or sociopolitical overtones in the mixture to set a propitious mood in the target audience; and a sort of ritualistic theatrics with the main theme of ‘winning’, that is best completed with ad hominems.

If you get rid of all this, the most reasonable interpretation of the output of a model proposed and tested should be similar to Nick Patterson’s words in his explanation of qpWave and qpAdm use:

Here we see that, at least in this analysis there are reasonable models with CordedWareNeolithic is a mix of either WHG or LBKNeolithic and YamnayaEBA. (…) The point of this note is not to give a serious phylogenetic analysis but the results here certainly support a major Steppe contribution to the Corded Ware population, which is entirely concordant with the archaeology [?].

Very far, as you can see, from the childish “Eureka! I proved the source!”-kind of thinking common among hobbyists.

The Mycenaean case is an illustrative example: if the Yamnaya outlier from Bulgaria were not available, and if one were not careful when designing and assessing those mixture models, the interpretation would range from erroneous (viz. a Graeco-Aryan substrate, as I initially thought) to impossible (say, inventing migration waves of Sintashta or Srubnaya peoples into Crete). The models presented above show that a contribution of Yamnaya to Mycenaeans couldn’t be rejected, and this alone should have been enough to accept Yamnaya as the most likely source population of “Steppe ancestry” in Proto-Greeks, pending intermediate samples from the Balkans. In other words, one could actually find that ‘the best’ p-values for source populations of Mycenaeans is a combination of modern Poles + Turks, despite the impracticality of such a model…

I haven’t been able to reproduce results which supposedly showed that Corded Ware is more likely to be derived from (Pre-)Yamnaya than other source population, or that Corded Ware is better suited as the ancestral population of Bell Beakers. The analyses above show values in line with what has been published in recent scientific papers, and what should be expected based on linguistics and archaeology. So I’ll go out on a limb here and say that it’s only through a careful selection of outgroups and samples tested, and of as few compared models as possible, that you could eventually get this kind of results and interpretation, if at all.

Whether that kind of special care for outgroups and samples is about (a) an acceptable fine-tuning of the analyses, (b) a simplistic selection dragged from the first papers published and applied indiscriminately to all models, or (c) cherry picking analyses until results fit the expected outcome, is a question that will become mostly irrelevant when future publications continue to support an origin of the expansion of ancient Indo-European languages in Khvalynsk- and Yamnaya-related migrations.

Feel free to suggest (reasonable) modifications to correct some of these models in the comments. Also, be sure to check out other values such as proportions, SD or SNPs of the different results that I might have not taken into account when assessing ‘good’ or ‘bad’ fits.


On the Ukraine Eneolithic outlier I6561 from Alexandria


Over the past week or so, since the publication of new Corded Ware samples in Narasimhan, Patterson et al. (2019) and after finding out that the R1a-M417 star-like phylogeny may have started ca. 3000 BC, I have been ruminating the relevance of contradictory data about the Ukraine_Eneolithic_o sample from Alexandria, its potential wrong radiocarbon date, and its implications for the Indo-European question.

How many other similar ‘controversial’ samples are there which we haven’t even considered? And what mechanisms are in place to control that the case of Hajji_Firuz_CA I2327 is not repeated?

Ukraine Eneolithic outlier I6561

It was not the first time that I (or many others) have alternatively questioned its subclade or its date, but the contradictory data seem to keep piling up. We can still explain all these discrepancies by assuming that the radiocarbon date is correct – seeing how it is a direct and newly reported lab analysis – because it is an isolated individual from a poorly sampled region, so he may actually be the first one to show features proper of later Corded Ware-related samples.

PCA of ancient Eurasian samples. An interpretation of the evolution of the Pontic-Caspian steppe populations in the Eneolithic. See full PCA.

The individual seems to be especially relevant for the Indo-European and Uralic homeland question. The last one to mention this sample in a publication was Anthony (2019), who considered it in common with two other Eneolithic samples from Dereivka to show how Anatolian farmer-related ancestry first appeared in the recently opened CHG mating network of the Pontic-Caspian steppes and forest-steppes during the Middle Eneolithic, after the expansion of Khvalynsk:

The currently oldest sample with Anatolian Farmer ancestry in the steppes in an individual at Aleksandriya, a Sredni Stog cemetery on the Donets in eastern Ukraine. Sredni Stog has often been discussed as a possible Yamnaya ancestor in Ukraine (Anthony 2007: 239- 254). The single published grave is dated about 4000 BC (4045–3974 calBC/ 5215±20 BP/ PSUAMS-2832) and shows 20% Anatolian Farmer ancestry and 80% Khvalynsk-type steppe ancestry (CHG&EHG). His Y-chromosome haplogroup was R1a-Z93, similar to the later Sintashta culture and to South Asian Indo-Aryans, and he is the earliest known sample to show the genetic adaptation to lactase persistence (I3910-T). Another pre-Yamnaya grave with Anatolian Farmer ancestry was analyzed from the Dnieper valley at Dereivka, dated 3600-3400 BC (grave 73, 3634–3377 calBC/ 4725±25 BP/ UCIAMS-186349). She also had 20% Anatolian Farmer ancestry, but she showed less CHG than Aleksandriya and more Dereivka-1 ancestry, not surprising for a Dnieper valley sample, but also showing that the old fifth-millennium-type EHG/WHG Dnieper ancestry survived into the fourth millennium BC in the Dnieper valley (Mathieson et al. 2018).

The main problem is that this sample has more than one inconsistent, anachronistic data compared to its reported precise radiocarbon date ca. 4045–3974 calBCE (5215±20BP, PSUAMS-2832). I summarized them on Twitter:

  • First known R1a-M417 sample, with subclade R1a-Y26 (Y2-), with formation date and TMRCA ca. 2750 BC (CI 95% ca. 3750–1950 BC), and proper of much later Steppe_MLBA bottlenecks. The closest available sample would be the Poltavka outlier of hg. R1a-Z94 (ca. 2700 BC), from a mixed cemetery that could belong to a later (likely Abashevo) layer; the closest related subclade is probably found in sample I12450 of Butkara_IA (ca. 800 BC).
  • NOTE. The formation date of upper clade R1a-Z93 is estimated ca. 3000 BC, with a CI 95% ca. 3550–2550 BC, suggesting that the actual TMRCA range for the subclade has most likely a lower maximum formation date than estimated with the available samples under Y3.

  • Ancestry and PCA cluster like Steppe_MLBA (see PCA below), different from neighbouring Sredni Stog samples of the roughly coetaneous Dereivka site (ca. 3600-3400 BC), and from a later Yamnaya sample from Dereivka (ca. 2800 BC), even more shifted toward WHG-related ancestry.
  • Allele for lactase persistence (I3910-T), found only much later among Bell Beakers, and still later in Sintashta and Steppe_MLBA samples. This suggests a strong selection in northern Europe and South Asia stemming from steppe-related (and not forest-steppe-related) peoples, postdating the age of massive Indo-European migrations.
  • Hajji Firuz Chalcolithic outlier

    My impression is that the Hajji_Firuz Chalcolithic outlier, initially dated ca. 5900-5500 BC, had much less reason to be questioned than this sample, since Pre-Yamnaya ancestry was (and apparently is still) believed by members of the Reich Lab to have come from south of the Caucasus, and to have arrived around that time or earlier to the North Caspian steppe, i.e. before the 5th millennium BC.

    The formation date of its initially reported haplogroup, R1b-Z2103, is ca. 4100 BC (CI 95% 4800-3500 BC), which seems also roughly compatible with that date and site – at least as compatible as R1a-Y3(xY2) is for ca. 4000 BC -, so it could have been interpreted as a migrant from the South Caspian region, potentially related to Proto-Anatolians, especially before the description of the Caucasus genetic barrier in Wang et al (2018). For some reason, though, the Hajji_Firuz sample was questioned, but this one didn’t even merited an interrogation mark.

    There was already a similar situation with two samples (RISE568 and RISE569) initially reported as belonging to Czech Corded Ware groups, that turned out to be Early Slavs ca. 3,000 years younger, in turn more closely related to Bell Beaker-derived cultures of Central-East Europe. It seems little has changed since that case.

    All in all, my guess is that genomic data of I6561 would have been a priori more compatible with a later period, during the expansion of East Corded Ware groups: at least Middle Dnieper culture, potentially Multi-Cordoned Ware culture, but most likely a Srubnaya-related one, given the most likely SNP mutation and TMRCA date, and the haplogroup variability found in the few samples available from that culture.

    PCA of ancient Eurasian samples. Marked I6561 sample within the cluster formed by Srubnaya samples. See full PCA.

    Compatibility checks

    I tried to start a thread on the possibility that the radiocarbon date was wrong, and IF it were, how likely it would be that formal stats could actually show this, or how could we automatically prevent ancestry magic fiascos.

    In other words: if this guy were a Srubnaya-related individual actually dated e.g. ca. 1700 BC, and someone would try to ‘prove’ – based on the current open source tools alone – that he was the ancestor of expanding peoples of the 4th and 3rd millennium BC (i.e. Balkan outliers, Yamnaya, Corded Ware, you name it), could these results be formally challenged?

    I was hoping for some original brainstorming where people would propose crazy, essentially impossible to understand statistical models, say plotting dozens of well-studied mutations of different geographically related ancient samples with their reported dates, to visually highlight samples that don’t exactly fit with such a feature-based time series analysis; I mean, the kind of theoretical models I wouldn’t even be able to follow after the first two tweets or so. I didn’t receive an answer like that, but still:

    I have nothing to add to these answers, because I agree that all contradictory data are circumstancial.

    The current absolute lack of this kind of validity checks for ancestry models is disappointing, though, and leaves the so-called outliers in a dangerous limbo between “potentially very interesting samples” and “potentially wrongly dated samples”. Radiocarbon date is thus – together with compatibility of population source in terms of archaeological cultures and their potential relationship – a necessary variable to take into account in any statistical design: an error in one of these variables means a catastrophic error in the whole model.

    Formal stats

    For example, in these qpAdm models, I assumed Srubnaya, Ukraine_Eneolithic_outlier, and Bulgaria_MLBA samples were roughly coetaneous and potentially related to the Srubnaya-SabatinovkaNoua cultural horizon, hence stemming from a source close to:

    1. Abashevo-like individuals (whose best proxy to date should be Poltavka_outlier I0432) potentially admixed with Poltavka-like herders; or
    2. Potapovka-like individuals potentially admixed with Catacomb-like peoples (whose best proxy until recently were probably Yamnaya_Kalmykia*).

    *To avoid adding more potential errors by merging different datasets, I have used only proxy samples available in the Reich Lab’s curated dataset of published ancient DNA.

    Srubnaya and Noua-Sabatinovka cultural horizon during the MLBA. See full maps.

    Apart from the lack of more models for comparison (I’m not going to dedicate more time to this), the results can’t be interpreted without a proper sampling and context, either, because (1) Poltavka_o may actually be from a much later group closely related to Srubnaya; (2) Bulgaria_MLBA is only one sample; and (3) there are only two samples from Potapovka; so the models here presented are basically useless, as many similar models that have been tested looking just for a formal “best fit”.

    So feel free to chime in and contribute with ideas as to how to detect in the future whether a sample is ancestral to or derived from others. I will post here informative answers from Twitter, too, if there are any. I don’t think a discussion about the potentially wrong date in this specific sample is very useful, because this seems impossible to prove or disprove at this point. Just what tools or data would you use to at least try and assess whether samples are compatible with its reported date or not – preferably in some kind of automated sieve that takes dozens or hundreds of samples into account.

    On the bright side, there is so much more than formal stats to arrive to relevant inferences about prehistoric populations, their movements and languages. That’s why I6561 didn’t matter for the conclusion by Anthony (2019) that it was the R1b-rich Eneolithic Don-Volga-Caucasus region the most likely Indo-Anatolian and Late Proto-Indo-European homeland, due to the creation of a wide Eneolithic mating network with extended exogamy practices, where Y-chromosome bottlenecks seem to be one of the main genomic data to take into account from the Neolithic to the Middle Bronze Age.

    And that is the same reason why it doesn’t matter that much for the Proto-Indo-European or Uralic question for me, either.


Proto-Tocharians: From Afanasievo to the Tarim Basin through the Tian Shan


A reader commented recently that there is little information about Indo-Europeans from Central and East Asia in this blog. Regardless of the scarce archaeological data compared to European prehistory, I think it is premature to write anything detailed about population movements of Indo-Iranians in Asia, especially now that we are awaiting the updates of Narasimhan et al (2018).

Furthermore, there was little hope that Tocharians would be different than neighbouring Andronovo-like populations (see a recent post on my predicted varied admixture of Common Tocharians), so the history of both unrelated Late PIE languages would have had to be explained by the admixture of Afanasievo-related groups with peoples of Andronovo descent and their acculturation.

However, data reported recently by Ning, Wang et al. Current Biology (2019) confirmed that peoples of mainly Afanasievo ancestry – as opposed to those of Corded Ware-related ancestry expanding with the Srubna-Andronovo horizon – spread the Tocharian branch of Proto-Indo-European from the Altai into the Tian Shan area, surviving essentially unadmixed into the Early Iron Age.

This genetic continuity of Tocharians will no doubt help us disentangle a great part the ethnolinguistic history of speakers of the Tocharian branch of Proto-Indo-European, from Pre-Proto-Tocharians of Afanasievo to Common Tocharians of the Late Bronze Age/Iron Age eastern Tian Shan.

NOTE. Tocharian’s isolation from the rest of Late PIE dialects and its early and intense language contacts have always been the key to support an early migration and physical separation of the group, hence the traditional association with Afanasievo, a late Repin/early Yamna offshoot. Even with the current incomplete archaeological and genetic picture, there is no other option left for the expansion of Tocharian.

It is not possible to use the currently available ancestry data to map the evolution of Afanasievo ancestry, lacking a proper geographical and temporal transect of Central and East Asian groups. In spite of this, Ning, Wang, et al. (2019) is a huge leap forward, discarding some archaeological models, and leaving only a few potential routes by which Tocharians may have spread southward from the Altai.

NOTE. I have updated the maps of prehistoric cultures accordingly, with colours – as always – reflecting the language/ancestry evolution of the different groups, even though the archaeological data of some groups of Xinjiang remains scarce, so their ethnolinguistic attribution – and the colours picked for them – remain tentative.

A rough timeline of related archaeological sites from North Eurasia. Image modified from Yang (2019).


The recent book Ancient China and its Eurasian Neighbors. Artifacts, Identity and Death in the Frontier, 3000–700 BCE, by Linduff, Sun, Cao, and Liu, Cambridge University Press (2017) offers an interesting summary of the introduction of metalworking into western China.

Here are some relevant excerpts (emphasis mine):

Although [the Xinjiang] route is not uniformly agreed upon (Shelach-Lavi 2009: 134–46), this western transmission has been thought to have passed through eastern Kazakhstan, especially as it is manifest in Semireiche, with Yamnaya, Afanasievo (copper) and Andronovo (tin bronze) peoples (Mei 2000: Fig. 3). From Xinjiang this knowledge has been thought to have traveled through the Gansu Corridor via the Qijia peoples (Bagley 1999) and then into territories controlled by dynastic China. The dating of this process is still a problem, as the sites and their contents in Xinjiang are consistently later than those in Gansu, suggesting that the point of contact was in Gansu and that the knowledge then spread from there westward.

1. Eneolithic Altai

Afanasievo expansion ca. 3300-2600 BC. See full culture and ancient DNA maps.

The Afanasievo sites, as they are identified in Mongolia, for instance, make up an Eneolithic culture analogous to that of southern Siberia (3100/2500–2000 BCE) in the Upper Yenissei Valley that is characterized by copper tools and an economy reliant on horse, sheep and cattle breeding as well as hunting. (…) The Afanasievo is best known through study of its burials, which typically include groups of round barrows (kurgans), each up to 12 m in diameter with a stone kerb and covering a central pit grave containing multiple inhumations. In their Siberian context, burial pottery types and styles have suggested contacts with the slightly earlier Kelteminar culture of the Aral and Caspian Sea area.

The Afanasievo culture monuments, located in the northern Altai and in the Minusinsk Basin (the western Sayan), have been seen as analogous evidence for cross-Eurasian exchange. These complexes contain small collections of metal, and many of the items are made of brass, although golden, silver and iron ornaments were also identified. A mere one-fourth of these objects are tools and ornaments, while the rest consist of unshaped remains and semi-manufactured objects. Its metallurgical tradition has recently been dated by Chernykh to as early as 3100 to 2700 BCE (1992),making it more compatible chronologically with the early brass-using sites in Shaanxi mentioned above. Kovalev and Erdenebaatar have excavated barrows in Bayan-Ulgii, Mongolia, that have been carbon-dated to the first half of the third millennium BCE and associated by ceramic types and styles and burial patterns with the Afanasievo (Kovalev and Erdenebaatar 2009: 357–58). These mounded kurgans were covered with stone and housed rectangular, wooden-faced tombs that included Afanasievo-type bronze awls, plates and small “leaf-shaped” knife blades (Kovalev and Erdenebaatar 2009: Figs. 6 and 7).

They also excavated sites belonging to the more recently identified Chemurchek archaeological culture, located in the foothills of the Mongolian Altai (Kovalev 2014, 2015) (Fig. 2.6). These sites are carbon-dated to the same period as the Afanasievo burials or to c. 3100/2500–1800 BCE (six barrows in Khovd aimag and four in Bayan-Ulgo aimag). In the rectangular stone kerbed Chemurchek slab burials (Ulaaanhus sum, Bayan-ul’gi aimag and so forth), bronze items included awls; and at Khovd aimag, Bulgan sum, in addition to stone sculptures, three lead and one bronze ring were excavated (Kovalev and Erdenebaatar 2009: Figs. 2 and 3; Fig. 2.6). Although we will not know if they were produced locally until much further investigation is undertaken, these discoveries do document knowledge of various uses and types of metal objects in western and south central Mongolia. The types of metal items thus far recovered are simple tools (awls) and rings (ornamental?) not unlike those associated with Andronovo archaeological cultures as well.

This is a complex circumstance where archaeological evidence is not complete, but raises very important questions about transmission of metallurgical knowledge to and from areas in present-day China. In the 1970s some Afanasievo mounds were excavated in Central Mongolia by a Soviet–Mongolian expedition led by V. V. Volkov and E. A. Novgorodova (Novgorodova 1989: 81–85). Unfortunately, these mounds did not yield metal objects, only ceramics, but they show that the Afanasievo culture with the Eneolithic metallurgical tradition of manufacturing pure copper items had already moved east at least far as central Mongolia. In 2004, Kovalev and Erdenebaatar investigated a large Afanasievo mound, Kulala ula, in the extreme northwest of Mongolia, near the Russian border (Kovalev and Erdenebaatar 2009). There they found a copper knife and awl (Fig. 2.5). There are five C14 dates on wood, coal and human bones from this mound, which belong to the period 2890–2570 BCE. This shows that the Afanasievo culture were carriers of technology and produced artifacts in the first half of the third millennium BCE and that they also moved south along the foothills of the Mongolian Altai. Afanasievo culture in Altai and the Minusinsk basin is dated by C14 to 3600–2500 BCE (Svyatko et al. 2009; Polyakov 2010). In the north of Xinjiang in the Altai district, several typical egg-shaped vessels and two censers of Afanasievo types were found. Some of these have been obtained from the stone boxes (chambers of megalithic graves of the Chemurchek culture) (Kovalev 2011). Thus, the Afanasievo tradition of pure copper metallurgy must have spread to the northern foothills of the Tienshan Mountains no later than the mid-third millennium BCE. The links with Afanasievo and local cultures adjacent to and south of the mountains into present-day China can now be assumed.

Afanasievo – Chemurchek evolution ca. 2600-2200 BC. See full culture and ancient DNA maps.

2. Bronze Age Altai

Kovalev and Erdenebaatar (2014a) and later Tishkin, Grushin, Kovalev and Munkhbayar (2015) in Western Mongolia conducted large-scale excavations of megalithic barrows of the Chemurchek culture (dated about 2600–1800 BCE). This peculiar culture appeared in Dzungaria and the Mongolian Altai in the second quarter of the third millennium BCE and for some time existed together with the late Afanasievo culture, as evidenced by the findings of Afanasievo ceramics in Chemurchek graves, in the stone boxes. Unfortunately, in China we do not yet know of any metal object related,without doubt, to the Chemurchek culture. Kovalev, Erdenebaatar, Tishkin and Grushin found several leaden ear rings and one ring of tin bronze in three excavated Chemurchek stone boxes (Kovalev and Erdenebaatar 2014a; Tishkin et al. 2015). Such lead rings are typical for Elunino culture,which occupied the entire West Altai after 2400–2300 BCE (Tishkin et al. 2015). This culture had developed a tradition of bronze metallurgy with various dopants, primarily tin. Thus, the tradition of bronze metallurgy as early as this time could have penetrated the Mongolian Altai far to the south. In addition, in the Hadat ovoo Chemurchek stone box, Kovalev and Erdenebaatar discovered stone vessels refurbished with the help of copper “patches,” indicating the presence there of metallurgical production (Fig. 2.7) (Kovalev and Erdenebaatar 2014a). In one of the secondary

Chemurchek graves unearthed by Kovalev and Erdenebaatar in Bayan-Ulgi (2400–2220 BCE), a bronze awl was found (Kovalev and Erdenebaatar 2009). Kovalev and Erdenebaatar also discovered a new culture in the territory of Mongolia (Map 2.3), one that begins immediately after Chemurchek – Munkh-Khairkhan culture (Kovalev and Erdenebaatar 2009, 2014b). To date, about 17 mounds of this culture have been excavated in Khovd, Zavkhan, Khovsgol, Bulgan aimag of Mongolia. This culture dates from about 1800 to 1500 BCE, that is, contemporary with the Andronovo culture. Therefore, the Andronovo culture does not extend far into the territory of Mongolia. Three knives without dedicated handles or stems and five awls have been found in the Munkh-Khairkhan culture mounds (Fig. 2.8). All these products are made of tin bronze. (…) Additionally, eight Late Bronze Age burials (c. 1400–1100 BCE) were unearthed in the Bulgan sum of Khovd aimag and belong to another previously unknown culture called Baitag. And in the Gobi Altai, a new group of “Tevsh” sites dating to the Late Bronze Age were defined in Bayankhongor and South Gobi aimags (Miyamoto and Obata 2016: 42–50). From these Tevsh and Baitag sites, we see the expansion of burial goods to include beads of semiprecious stones (carnelian), bronze beads, buttons and rings and even the famous elaborate golden hair ornaments (Tevsh uul;Bogd sum;Uverkhanagia aimag) from the Baitag barrows (Kovalev and Erdenebaatar 2009: Fig. 5; Miyamoto and Obata 2016).

2.1. Chemurchek

About the Chemurchek culture, from A re-analysis of the Qiemu’erqieke (Shamirshak) cemeteries, Xinjiang, China, by Jia and Betts JIES (2010) 38(4):

The major characteristics of Qiemu’erqieke Phase I include:

  1. Burials with two orientations of approximately 20° or 345°.
  2. Rectangular enclosures built using large stone slabs. The size of the enclosure varies from a maximum of 28 x 30 m.*to a minimum of 10.5 x 4.4 m. (Figure 8, Table 2).
  3. *The stone enclosure located near Hayinar is the largest one at approximately 30 x 40 m. based on pacing of the site during a visit by the authors in 2008.

  4. Almost life-sized anthropomorphic stone stelae erected along one side of the stone enclosures (Lin Yun 2008).
  5. Single enclosures tend to contain one or more than one burial, all or some with stone cist coffins.
  6. The cist coffin is usually constructed using five large stone slabs, four for the sides and one on top, leaving bare earth at the base (Zhang Yuzhong 2007). Sometimes the insides of the slabs have simple painted designs (Zhang Yuzhong 2005).
  7. Primary and secondary burials occur in the same grave.
  8. Some decapitated bodies (up to 20) may be associated with the main burial in one cist.
  9. Bodies are commonly placed on the back or side with the legs drawn up.
  10. Grave goods include stone and bronze arrowheads, handmade gray or brown round-bottomed ovoid jars, and small numbers of flat-bottomed jars (Fig. 7).
  11. Clay lamps appear to occur together with roundbottomed jars.
  12. Complex incised decoration on ceramics is common but some vessels are undecorated.
  13. The stone vessels are distinctive for the high quality of manufacture.
  14. Stone moulds indicate relatively sophisticated metallurgical expertise.
  15. Artefacts made from pure copper occur.
  16. Sheep knucklebones (astragali) imply a tradition (as in historical and modern times) of keeping knucklebones for ritual or other purposes. They also indicate the herding of domestic sheep as part of the subsistence economy.
Chemurchek culture ca. 2200-1750 BC. See full culture and ancient DNA maps.

Chemurchek dating

Available evidence suggests that the date range for Qiemu’erqieke Phase I should fall from the later third into the early second millennium BC. There are several reasons to suggest that the time span is around the early second millennium BC. Lin Yun (2008) (…) maintains that the bronze artefacts found in Phase I show a greater sophistication in the level of copper alloy technology than that of the pure copper artefacts common to the Afanasievo tradition. On this basis it might be suggested that the Afanasievo could be considered to be Chalcolithic with a time span across much of the third millennium BC ( Gorsdorf et al. 2004: 86, Fig. 1). Qiemu’erqieke Phase I, however, should more properly be considered as Bronze Age.

Lin Yun also used the bronze arrowhead from burial Ml 7 to narrow down the date of Qiemu’erqieke Phase I. Two arrowheads were found in this burial, one of them leaf shaped with a single barb on the back (Fig. 7:4). A similar arrowhead, together with its casting mould, has been found at the Huoshaogou site of Siba tradition (Li Shuicheng 2005, Sun Shuyun and Han Rufen 1997), in Gansu province, northwest China, dated around 2000-1800 BC (Li Shuicheng and Shui Tao 2000) . This supports a date in the early second millennium BC for the Qiemu’erqieke arrowhead. The painted, round-bottomed jar from the Tianshanbeilu cemetery Qia Weiming, Betts and Wu Xinhua 2008: Fig. 7, bottom left) has been considered as a hybrid between the Upper Yellow River Bronze Age cultures of Siba in northwest China and the steppe tradition of Qiemu’erqieke in west Siberia (Li Shuicheng 1999). If this assumption is correct, the date of Tianshanbeilu, around 2000 BC, can be used as a reference for Qiemu’erqieke Phase I (Jia Weiming, Betts and Wu Xinhua 2008, Lin Yun 2008, Li Shuicheng 1999). Stone arrowheads found in Qiemu’erqieke Phase I also imply that the date is likely to fall within the earlier part of the Bronze Age as no such stone arrowheads have yet been found elsewhere in sites of the Bronze Age in Xinlang dated after the beginning of the second millennium BC.*
*For example Chawuhu and Xiaohe cemeteries (Xinjiang Institute of Archaeology 1999, 2003).

Pottery of Afanasevo and East European traits from the Chemurchek complex. Image modified from Kovalev (2017).

(…) Pottery “oil burners” (goblet-like ceramic vessels, possibly lamps) have been found in three traditions: Afanasievo (Gryaznov and Krizhevskaya 1986:21), Okunevo and Qiemu’erqieke. It is believed that this oil-burner found in Siberia and the Altai is a heritage from the Yamnaya and Catacomb
cultures (Sulimirski 1970: 225, 425; Shishlina 2008:46) in the Caspian steppe further to the west, but does not seem to exist in known Andronovo cultures.
The oil-burner tends to disappear after around 2300 BC during the mid-Okunevo period. It is, however, possible that the tradition continues longer in the Qiemu’erqieke sites.

The construction of the stone enclosures also reveals a close connection between Qiemu’erqieke Phase I and the mid and late Okunevo tradition (Sokolova 2007). Slab built stone enclosures emerged in both the Okunevo and Afanasievo traditions (Gryaznov and Krizhevskaya 1986:15-23, Kovalev 2008, Sokolova 2007, Anthony 2007:310, Koryakova and Epimakhov 2007). In the early Afanasievo the enclosure is circular with no cist coffin (Anthony 2007:310, Gryaznov and Krizhevskaya 1986:20), but in the early stage of the Okunevo square stone enclosures with a single cist burial are dominant. Square or rectangular stone enclosures are a marked feature of Qiemu’erqieke Phase I, suggesting temporal relationships between Qiemu’erqieke Phase I and the Okunevo. In Okunevo chronological group II, possibly with influence from the Anfanasievo, circular stone enclosures appeared in combination with rectangular enclosures within individual cemeteries, referred to by Sokolova (2007: table 2) as hybrid examples. By Okunevo chronological group III, rectangular stone slab enclosures with multi-burials emerged again. This is the dominant form in Qiemu’erqieke Phase I. Okunevo burial traditions changed again to single cist burials in the late stage around chronological group V ( Sokol ova 2007). A specific mortuary rite of decapitated burials exists in both the Qiemu’erqieke and Okunevo traditions (Sokolova 2007, Chen Kwang-tzuu and Hiebert 1995), as does the occasional occurrence of painted designs on the interior of the slabs forming the cists ( e.g., Khavrin 1997: 70, fig. 4; 77: tab. IV.5). Based on these comparisons, the date of Qiemu’erqieke Phase I may well parallel that of the Okunevo from at least chronological group II around 2400 BC (Gorsdorf et al. 2004: fig. 1).

Khuh Udzuuriin I-1 elite barrow (ca. 2470-2190 BC). Modified from Image modified from Kovalev (2014).

In addition to the pottery making tradition, the anthropomorphic stone stelae may also have earlier antecedents. In the Okunevo assemblage there are anthropomorphic stelae that are longer, thinner and more abstract than those of Qiemu’erqieke. There is no indication of such stelae in the Afanasievo tradition (Gryaznov and Krizhevskaya 1986:15-23). However, further to the west, anthropomorphic stone stelae are associated with the Kemi-Oba and Yamnya cultures around the third millennium BC (Telegin and Mallory 1994; Figure 13). Some major characteristics of these stelae such as the icons on the front face of the stelae (Telegin and Mallory 1994:8-9) also appear on stelae found in Qiemu’erqieke Phase I. Recalling the oil burners that may have been inherited from the Yamnya culture and which are found in the Afansievo, Okunevo and Qiemu’erqieke Phase I, it migh t be possible to speculate that Qiemu’erqieke Phase I has its origins even earlier than the first half of the third millennium BC. This idea has also been suggested by Kovalev ( 1999).

Despite the affinities with the Okunevo cultural tradition, Qiemu’erqieke Phase I appears to be a discrete regional variant. The ceramic assemblage shows traits unique to this cluster of sites, while the anthropomorphic stelae are also distinctive markers of this tradition.

Khuh Udzuur anthropomorphic stone stela, oriented toward the south – south-east. Image modified from Kovalev (2014).

3. Bronze Age Xinjiang

I recently reported on this blog the description of Xiaohe and Gumugou cemeteries from interesting Master’s thesis Shifting Memories: Burial Practices and Cultural Interaction in Bronze Age China: A study of the Xiaohe-Gumugou cemeteries in the Tarim Basin, by Yunyun Yang, Uppsala University, Department of Archaeology and Ancient History (2019).

It also offered a full summary of findings from prehistoric sites of Xinjiang related to the arrival of a cultural package from the Altai region, ultimately connected to Afanasievo. Relevant excerpts include the following (emphasis mine):

In Bronze Age Xinjiang, burials were diverse but also show some common features between different geographic sections. The main three mountains, including Kunlun Mountains, Tian Shan (mountains) and Altai Mountains, enclose the Tarim Basin, and the Dzungaria Basin, but leave the eastern part of the Tarim Basin and the south-eastern part of the Dzungaria Basin open (with easy access to the surroundings). The Hami Basin is located at the transitional area, connecting the two basins. Burials are mainly spread along the edge of the mountain ranges.

An assumption of the spreading/expansion routes stone burial construct.

3.1. The Lop Nur region

In the Lop Nur region, the Xiaohe cemetery (2000-1450 BCE) and the Gumugou cemetery (1900-1800 BCE) had many common features shared, and so is the Keliyahe northern cemetery:

  • Cemeteries were located in sandy areas;
  • Rectangular/boat-shaped wooden coffins with monuments of wooden planks or poles;
  • Coffins had no bottoms;
  • The dead were placed lying straight on the back;
  • The dead were commonly buried in single graves.

The Gumugou cemetery contained six special sun-radiating-spokes burial pattern in addition to the normal burials, which were similar to the wooden coffin graves of the Xiaohe cemetery.

NOTE. For more on Xiaohe and Gumugou, see the recent post on Proto-Tocharians. See other papers on the Andronovo horizon for other Early to Middle Bronze Age cultural groups less clearly associated with the Xiaohe horizon, like Hazandu, Xintala, or the Chust culture.

From Shuicheng (2006):

An assemblage of early bronzes had been recovered from northwestern Xinjiang and the periphery of Dzungaria 准噶尔 Basin. It comprises a variety of utilitarian tools and weapons, and a small number of apparels. These artifacts bear the stamps of Andronovo Culture in form, artifact type and decorative pattern. The metallographic analysis on selected artifacts indicates that they comprise mainly of tin-bronzes that contain 2–10% of tin. Moreover, the chemical compositions of these artifacts are similar to that of the Andronovo Culture. Latter date (first half of the 1st millennium BC) artifacts of the assemblage include a small number of arsenic bronzes. In all, during the period between the mid-2nd and mid-1st millennium BC, copper and bronze artifacts coexisted in this region, albeit tin-bronze comprised the majority. The composition of alloy did not show significant change over time. Some colleagues pointed out that the Nulasai 奴拉赛 site at Nileke 尼勒克 County in the Yili 伊犁 River basin of Xinjiang was the pioneer in the use of “sulphuric ore–ice copper–copper”technology. It is also the only early smelting site in Euro-Asia that arsenic ore was added to deliberately produce an alloy

Prehistoric cultures of Xinjiang during the Middle Bronze Age. See full culture and ancient DNA maps.

3.2. The Hami Basin-the Balikun Grassland

From Yang (2019):

The Hami Basin-the Balikun Grassland area is located at the eastern part of Tian Shan. The area is divided in a northern basin and a southern basin by the east-west stretch of the Tian Shan. In the Hami Basin-the Balikun Grassland area, the main type of burials were earth-pit graves in the early Bronze Age, and burials of stone-pit with barrows became more common in the late Bronze Age. The Hami-Tianshan-Beilu cemetery is a representative of the earth-pit graves. The features of the Hami-Tianshan-Beilu cemetery (2000-1500 bce) here were:

  • Rectangular earth pit graves;
  • The dead were often in a hocker position lying on one side;
  • Commonly a single dead in one grave.
The Balikun grassland today (source).

The Hami-Wubu cemetery (earlier than 1000 bce) and the Yanbulake cemetery (1200-600 bce) are representatives of another common earth-pit graves. Common features here were:

  • Rectangular earth pits, with two storeys and/or roofed with wooden boards;
  • The dead was placed in a hocker position lying on one side;
  • Mostly a single dead in one grave.

Later there appeared more stone-pit graves in this area, and the features can be summarized as:

  • Round burial mounds, commonly constructed by stones or a mix of stones and earth;
  • Burial mounds with a sunken top or a normal (dome) top;
  • The diameter of the burial mounds varied between 3 and 25.4 m (but not necessarily limited in this scope);
  • Circular or rectangular stone kerbs;
  • Rectangular stone pits, constructed by earth, or stones, or a mix of earth and stones;
  • Rectangular stone pits contained wooden coffins (represented by the Yiwu Baiqi’er cemetery).
Some representatives of stone burials in the Hami Basin – the Balikun Grassland in the Iron Age (Adapted from: Xinjiang 2011, 29-41). Image modified from Yang (2019).

In the Hami Basin, the Bronze Age cemeteries show common burial features like earth pits and hocker position of the dead. With similar pottery styles in the Hami-Tianshan-Beilu cemetery to those in the Machang and Siba cultures (Xinjiang 2011: 17), it suggests possible cultural influence or people’s migrating from the Hexi Corridor in the east.

In the Balikun Grassland, burials in an earlier time contained mostly earth-pit graves but also a small number of stone-pit graves. The pebbles were imbedded in the floors and the walls of the graves in a rectangular shape, e.g. the Balikun-Nanwan cemetery (1600-1000 bce). In a later time, there appeared huge burial mounds with a sunken top, and with the diameters of the burial mounds varying from 3 to 25.4 m, e.g. the Balikun-Dongheigou cemetery and the Balikun-Heigouliang cemetery. The Yiwu-Bai’erqi and the Yiwu-Kuola cemeteries contained either round stone burial mounds or circular stone kerbs on the ground surface. Considering the three burial elements including burial mounds, stone pits and circular kerbs, the later period cemeteries in the Balikun Grassland were actually similar to cemeteries from the southern edge of the Altai Mountain area.

From Shuicheng (2006):

The Nanwan 南湾 cemetery site at Kuisu 奎苏 Town, Balikun 巴里坤 (1600–1100 BC) also yielded an assemblage of early bronzes. The style of its early phase artifacts is similar to that of the burials distributed in the North Tianshan Route. Some sorts of cultural connection should have existed between the two.

The dates of Yanbulake 焉不拉克 Culture (1300–700 BC) are comparatively late. Its metallurgy was a continuation of the western China tradition. Artifact types include a variety of utilitarian tools, weapons and apparels.

Prehistoric cultures of Xinjiang during the Late Bronze Age. See full culture and ancient DNA maps.

3.3. The Turpan Basin-the middle part of Tian Shan

From Yang (2019):

Turpan Basin is located at the western part of the Hami Basin, and lies at the southern edge of the eastern Tian Shan. In the Turpan Basin-the middle part of Tian Shan area, the main representative of the Bronze Age cemeteries is the Yanghai Nr.1 cemetery. The features here were:

  • Elliptic earth pit graves, commonly covered by round logs on the top;
  • Some graves contained burial beds made of round logs or reeds;
  • The dead were mainly placed lying straight on the back;
  • Mostly a single dead in one grave.

In Iron Age, the stone burials became dominant, but the stone burials varied in different regions of the Turpan Basin-the middle part of Tian Shan area. Graves containing burial mounds, stone pit, and circular stone kerbs are represented by the Shanshan-Ertanggou cemetery, the Tuokexun-Alagou cemetery, the Urumqi-Chaiwobu cemetery and the Urumqi-Yizihu-Sayi cemetery, etc. The stone funeral construction features here are similar to those contemporary cemeteries in the Hami Basin-the Balikun Grassland area.

3.4. The southern edge of the western and middle part of Tian Shan

In the southern edge of the western and middle part of Tian Shan area, the main representatives of the late Bronze Age cemeteries are the Hejing-Chawuhu Nr.4 cemetery (around 1000-500 bce), the Hejing-Xiaoshankou cemetery, the Baicheng-cemetery, etc. The main burial features of the late Bronze Age and the early Iron Age cemeteries (see Fig.12) here were:

  • Burial mounds, constructed by stones or a mix of stones and earth;
  • Irregular circular or rectangular stone kerbs;
  • Stone pit graves in a bell-shape or a rectangular shape;
  • Stone pit graves constructed by imbedding pebbles or stone slabs in walls and floors;
  • The dead were often placed lying on their back with bent legs;
  • The dead were commonly reburied a second time with multiple burials.

From the late Bronze Age to the early Iron Age in this area, the burial traditions tended to be in a more varied way. In the stone burials with stone kerbs, there is a mixture of stone pit and earth pit graves. The burial features of the Iron Age cemeteries in this section were similar to those contemporary both in the Hami Basin-the Balikun Grassland area and in the Turpan Basin-the middle part of Tian Shan area.

From Shuicheng (2006):

The Chawuhu 察吾呼 Culture (1100–500 BC) distributes on the foothills between the middle section of the Tianshan Mountain Ranges and Tarim River. Its bronze assemblage comprises a variety of weapons, utilitarian tools and small apparels. They show no apparent temporal change in form and type through the four cultural phases. In addition, bronzes bear the Chawuhu characteristics were found in Hejing 和静, Baicheng 拜城 and Luntai 轮台 (Bügür). Yet, sites distributed along the Tarim River, such as Heshuo 和硕, Kuga 库车and Aksu 阿克苏, yielded remains of a bronze culture different from that of Chawuhu. Bronzes recovered include double-eared socketed axe, arrowheads, awls, knives, needles and bracelets. Their absolute dates have been estimated to be earlier than that of Chawuhu.

Prehistoric cultures of Xinjiang during the Early Iron Age. See full culture and ancient DNA maps

3.5. The Pamir Plateau

From Yang (2019):

A typical Bronze Age cemetery from the Pamir Plateau area is the Tashenku’ergan-Xiabandi cemetery (around 1000-500 bce). The burial features here were:

  • Mainly inhumations, but also a few cremations;
  • Burial mounds, constructed of stones;
  • Irregular circular or rectangular stone kerbs;
  • Mostly a single dead in one grave;
  • The dead was placed in a hocker position lying on one side.

The adoption of burial customs from the east supports the migration of Afanasievo-related peoples from the Tian Shan up to the Pamir Plateau, strongly influencing the findings of the Xiabandi cemetery, which has been dated from an early Bronze Age phase (ca. 1500-300 BC) to a late date up to ca. 600 BC.

While it is today unclear how far the Afanasievo admixture reached into the western Xinjiang, it seems that the Pamir Plateau remained culturally connected to neighbouring Andronovo-related cultures in pottery and metallurgical innovations, hence their language probably belonged – during most part of the Bronze and Iron Ages – to the Indo-Iranian branch, even though specific dialects might have changed with each new attested group.

In particular, it is possible that the early Andronovo groups related to the Xiaohe Horizon spoke Indo-Aryan or West Iranian dialects, while Saka-related groups replaced them – or an intermediate Tocharian-speaking group – with East Iranian dialects. A close interaction with West Iranian would justify the known ancient borrowings of Tocharian, although they could also be explained by contacts with Chust-related groups farther west. For more on this, see Ged Carling’s work on the different layers of Iranian loans.

Xinjiang BA/IA Summary

From Yang (2019):

In the early Bronze Age, there are distinct regional differences in the burial customs in and surrounding the Tarim Basin. At the southern edge of the Altai Mountains area, the burial customs included stone burial mounds, stone pit graves, circular or rectangular stone kerbs and stone human sculptures; the dead were placed lying straight on the back. In the Hami Basin-the Balikun Grassland area, the burial customs included earth pit graves; the dead were placed in a hocker position lying on one side. In the Turpan Basin-the middle part of Tian Shan area, the burial customs included earth pit graves; the dead were placed lying straight on the back. In the Lop Nur region, the burial customs included wooden coffins buried in sand; the dead were placed lying straight on the back.

But from the late Bronze Age to the early Iron Age, there was a common shift in burial customs from earth pit graves to stone burials in the Hami Basin-the Balikun Grassland area and in the Turpan Basin-the middle part of Tian Shan area. The main features of the stone burials include stone burial mounds, circular or rectangular stone kerbs, and the stone pit graves in the cemeteries. Similar stone burial customs commonly appeared at the southern edge of the western and middle part of Tian Shan area and the Pamir Plateau area in Iron Age. The burial features in most areas are in a mixture of both the earth pit graves and stone pit graves, especially in the Hami Basin-the Balikun Grassland area and the Turpan Basin-the middle part of Tian Shan area.


From Shuicheng (2006):

Historians of metallurgy conducted metallographic analyses on a sample of 234 metal specimens recovered from 16 localities in eastern Xinjiang. They concluded that the metallurgic industry in eastern Xinjiang could be roughly partitioned into three developmental phases. The early phase is represented by the burials distributed in the North Tianshan Route. The majority of the metal assemblage was tin-bronzes; however, copper and arsenic-bronzes maintained considerable proportions. The middle phase is represented by the burials at Yanbulake. During this phase, tin-bronze still maintained the majority; the proportion of arsenic-bronze increased, and some of them were high arsenic-bronzes. The late phase is represented by the burials at Heigouliang 黑沟梁. The composition of lead increased in the bronze alloy in the expense of arsenic. In addition, this phase witnessed the appearance of high tin-bronze that composed up to 16% of tin and the appearance of brass, that is, an alloy of copper and zinc. The bronze alloy consistently contained significant amount of impurities regardless of temporal difference. Casting and forging technologies coexisted throughout the three phases. The early bronzes (2000–500 BC) of eastern Xinjiang, in general, contained arsenic; however, the composition of arsenic was usually under 8%, but a few artifacts contained more than 20% arsenic. In all, arsenic had long been used in the alloy-forming of the early bronzes in eastern Xinjiang. Consequently, arsenic-bronzes were widely found in the prehistoric archaeology of the region. The artifact types, chemical compositions and manufacture techniques of the bronze assemblage of the burials of the North Tianshan Route are similar to those of Siba Culture, indicating that eastern Xinjiang had played a significant role in the East-West interactions.

An assemblage of early bronzes had been recovered from northwestern Xinjiang and the periphery of Dzungaria 准噶尔 Basin. It comprises a variety of utilitarian tools and weapons, and a small number of apparels. These artifacts bear the stamps of Andronovo Culture in form, artifact type and decorative pattern. The metallographic analysis on selected artifacts indicates that they comprise mainly of tin-bronzes that contain 2–10% of tin. Moreover, the chemical compositions of these artifacts are similar to that of the Andronovo Culture. Latter date (first half of the 1st millennium BC) artifacts of the assemblage include a small number of arsenic-bronzes. In all, during the period between the mid-2nd and mid-1st millennium BC, copper and bronze artifacts coexisted in this region, albeit tin-bronze comprised the majority.

Prehistoric cultures of Xinjiang during the Late Iron Age. See full culture and ancient DNA maps.

Tocharians in population genomics

Prehistoric population movements between the Altai and the Tian Shan are difficult to pinpoint, not the least because of the division of these territories among three different countries and their archaeological teams, only recently (more) open to the international scholarship.

The available schematic archaeological picture, where migrations could only be roughly inferred, has been recently updated to a great extent by Ning, Wang et al. (2019), whose genetic analysis of the samples is as thorough as anyone could have asked for, with a level of detail which matches the complex genetic picture of the region by the Iron Age.

As a summary, here is what they described about the samples from Shirenzigou (ca 400-200 BC), corresponding to the Iron Age populations of the Hami Basin-the Balikun Grassland area, and closely related to the preceding Yanbulake Culture:

As shown in Figure S3, the Steppe_MLBA populations including Srubnaya, Andronovo, and Sintashta were shifted toward farming populations compared with Yamnaya groups and the Shirenzigou samples. This observation is consistent with ADMIXTURE analysis that Steppe_MLBA populations have an Anatolian and European farmer-related component that Yamnaya groups and the Shirenzigou individuals do not seem to have. The analysis consistently suggested Yamnaya-related Steppe populations were the better source in modeling the West Eurasian ancestry in Shirenzigou.

Biplot of f3-outgroup tests illustrating the Kostenki14 and Anatolia_N like ancestries in Shirenzigou individuals. Most Shirenzigou individuals were on a cline with Yamnaya and European hunter-gatherer groups, lacking the European farmer ancestry as compared to the Steppe_MLBA populations such as Andronovo, Srubnaya and Sintashta [S1-S5]. Horizontal and vertical bars represent ± 3 standard errors, corresponding to form of outgroup f3 tests on the x axis and y axis respectively.

We continued to use qpAdm to estimate the admixture proportions in the Shirenzigou samples by using different pairs of source populations, such as Yamnaya_Samara, Afanasievo, Srubnaya, Andronovo, BMAC culture (Bustan_BA and Sappali_ Tepe_BA) and Tianshan_Hun as the West Eurasian source and Han, Ulchi, Hezhen, Shamanka_EN as the East Eurasian source. In all cases, Yamnaya, Afanasievo, or Tianshan_Hun always provide the best model fit for the Shirenzigou individuals, while Srubnaya, Andronovo, Bustan_BA and Sappali_Tepe_BA only work in some cases.

Table S2. P values in modelling a two-way (P=rank 1) admixture in Shirenzigou samples using each of the four populations (Bustan_BA, Sappali_Tepe_BA, Andronovo.SG, Srubnaya) together with Han Chinese as two sources [S6], Related to Figure 2. We used the following set of outgroups populations: Dinka, Ust_Ishim, Kostenki14, Onge, Papuan, Australian, Iran_N, EHG, LBK_EN.


In the PCA, ADMIXTURE, outgroup f3 statistics [see Figure S4], as well as f4 statistics (Table S3), we observed the Shirenzigou individuals were closer to the present day Tungusic and Mongolic-speaking populations in northern Asia than to the populations in central and southern China, suggesting the northern populations might contribute more to the Shirenzigou individuals. Based on this, we then modeled Shirenzigou as a three-way admixture of Yamnaya_Samara, Ulchi (or Hezhen) and Han to infer the source from the East Eurasia side that contributed to Shirenzigou. We found the Ulchi or Hezhen and Han-related ancestry had a complicated and unevenly distribution in the Shirenzigou samples. The most Shirenzigou individuals derived the majority of their East Eurasian ancestry from Ulchi or Hezhen-related populations, while the following two individuals M820 and M15-2 have more Han related than Ulchi/ Hezhen-related ancestry

It is unclear whether the Chemurchek population will show a sizeable local contribution from neighbouring groups. The fact that Okunevo shows 20% Yamnaya-related ancestry strongly supports the nature of neighbouring stone-grave-building peoples of the Altai and the northern Tian Shan as mostly Afanasievo-like, and the apparent lack of contributions of Srubna/Andronovo-like ancestry in the early Hami-Balikun stone burial builders also speaks for radical population replacement events reaching the areas south of Tian Shan, at least initially.

While ancestry cannot settle linguistic questions, it seems that nomads of the Gansu and Qinghai grasslands retained an ancestry close to Andronovo, whereas nomads of the Hami Basin-Balikun grasslands and related populations of Xinjiang remained closely related to Afanasievo. This doesn’t preclude that the ancestors of the Yuezhi became acculturated under the influence of peoples from eastern Xinjiang, but all data combined suggest an isolation of both populations – relative to other groups and to each other – and it is therefore more likely that they spoke Indo-Iranian-related languages rather than a language of the Tocharian branch.


In an interesting twist of events, despite the initially reported hg. R1b and Q, Tocharians from Shirenzigou actually show a haplogroup diversity comparable to that attested in other late Iron Age populations: a similar diversity is seen, for example, among Germanic, Baltic, and Balto-Finnic peoples of the Baltic region; among East Germanic or Scythians of the north Pontic region; or among Mediterranean peoples sampled to date. Iron Age peoples show thus a complex sociopolitical setting that overcame the previous patrilineal homogeneity of Bronze Age expansions.

PCA and ADMIXTURE for Shirenzigou Samples. Modified from the original to include in black squares samples related to Yamnaya. Modified from the paper to include labels of modern populations and a dotted lines with the cline formed by Shirenzigou, from (Yamnaya-like) Afanasievo to Central and East Asian-like populations. In red circles, samples with best fit for Andronovo-like ancestry. In green circles, samples with Han-related admixture.

M15-2 (with Han-related ancestry) is of the rare haplogroup Q1a-M120, while the samples with highest Steppe_MLBA-related ancestry are of hg. R1b-PH155, which points to their recent origin among Yuezhi, or to Hun-related populations showing an admixture related to the proto-historic nomads of the Gansu and Qinghai grasslands.

The expansion of Chemurchek-related peoples was probably associated more with hg. Q1a (dubious if it’s a Pre-ISOGG 2017 nomenclature, hence possibly Q1b), a haplogroup that might be found in Khvalynsk as a “significant minority” according to Anthony (2019), and it might also be attested in sampled individuals from Afanasievo in its late phase. This might be, therefore, a case similar to the early expansion of Indo-Europeans with R1b-V1636 lineages through the Volga – North Caucasus region, and of the later expansion with I2a-L699 lineages into the Balkans.

Haplogroup Q1a2-M25 is found in individual X3, whose Steppe ancestry is likely a combination of Afanasievo plus Andronovo-like ancestry heavily admixed with Hezhen/Ulchi-like populations, in line with the expected recent contacts with the neighbouring Xiongnu, Yuezhi, and other population movements affecting eastern Xinjiang.

Sample M4, which packs the most Afanasievo-like ancestry, is of hg. R1a-Z645, which – like sample M8R1 of hg. O – is most likely related to haplogroup resurgence events of local populations, which left the predominant Afanasievo-like admixture brought by builders of stone burials essentially intact, evidenced by the almost 100% of R1a found in the Xiaohe cemetery – and in most of the early Andronovo horizon – and among expanding Kangju and Wusun, as well as by the prevalence of hg. O among sampled East Asian populations.

A question that will only be answered with more samples is how and when the prevalent R1b-L23 and Q1b lineages among Afanasievo-related peoples began to be replaced to reach the high variability seen in Shirenzigou. Given the pastoralist nature of peoples around Tian Shan, the succeeding expansions of Proto-Tocharians, and the late isolation of different Common Tocharian groups, it is more than likely that this variability represents a late and local phenomenon within Xinjiang itself.

Peoples of Xinjiang during Antiquity. See full culture and ancient DNA maps.


Tocharians are one of the main pillars that confirm the Late Proto-Indo-European homeland of the R1b-rich populations of the Don-Volga region. There is already:

Just like the East Bell Beaker expansion from Yamnaya Hungary has confirmed that Corded Ware peoples did not partake in spreading Indo-European languages (spreading Uralic languages instead), data on the expansion of Tocharian speakers from Afanasievo to the Tian Shan was always there; population genomics is merely helping to connect the dots.

In summary, genetic research is supporting the expected linguistic expansions of the Neolithic and Bronze Age step by step, slowly but surely.


North-West Indo-Europeans of Iberian Beaker descent and haplogroup R1b-P312


The recent data on ancient DNA from Iberia published by Olalde et al. (2019) was interesting for many different reasons, but I still have the impression that the authors – and consequently many readers – focused on not-so-relevant information about more recent population movements, or even highlighted the least interesting details related to historical events.

I have already written about the relevance of its findings for the Indo-European question in an initial assessment, then in a more detailed post about its consequences, then about the arrival of Celtic languages with hg. R1b-M167, and later in combination with the latest hydrotoponymic research.

This post is thus a summary of its findings with the help of natural neighbour interpolation maps of the reported Germany_Beaker and France_Beaker ancestry for individual samples. Even though maps are not necessary, visualizing geographically the available data facilitates a direct comprehension of the most relevant information. What I considered key points of the paper are highlighted in bold, and enumerated.

NOTE. To get “more natural” maps, extrapolation for the whole Iberian Peninsula is obtained by interpolation through the use of external data from the British Isles, Central Europe, and Africa. This is obviously not ideal, but – lacking data from the corners of the Iberian Peninsula – this method gives a homogeneous look to all maps. Only data in direct line between labelled samples in each map is truly interpolated for the Iberian Peninsula, while the rest would work e.g. for a wider (and more simplistic) map of European Bronze Age ancestry components.


Iberian Chalcolithic groups and expansion of the Proto-Beaker package. See full map.

The Proto-Beaker package may or may not have expanded into Central Europe with typical Iberia_Chalcolithic ancestry. A priori, it seems a rather cultural diffusion of traits stemming from west Iberia roughly ca. 2800 BC.

Map of Y-DNA haplogroups among Iberia Chalcolithic samples. See full map.

The situation during the Chalcolithic is only relevant for the Indo-European question insofar as it shows a homogeneous Iberia_Chalcolithic-like ancestry with typical Y-chromosome (and mtDNA) haplogroups of the Iberian Neolithic dominating over the whole Peninsula until about 2500 BC. This might represent an original Basque-Iberian community.

Map of mtDNA haplogroups among Iberia Chalcolithic samples. See full map.

Bell Beaker period

Iberian Bell Beaker groups and potential routes of expansion. See full map.

The expansion of the Bell Beaker folk brought about a cultural and genetic change in all Europe, to the point where it has been rightfully considered by Mallory (2013) – the last one among many others before him – the vector of expansion of North-West Indo-European languages. Olalde et al. (2019) proved two main points in this regard, which were already hinted in Olalde et al. (2018):

(1) East Bell Beakers brought hg. R1b-L23 and Yamnaya ancestry to Iberia, ergo the Bell Beaker phenomenon was not a (mere) local development in Iberia, but involved the expansion of peoples tracing their ancestry to the Yamnaya culture who eventually replaced a great part of the local population.

Natural neighbor interpolation of Germany_Beaker ancestry in Iberia during the Bell Beaker period (ca. 2600-2250 BC). See full map.

(2) Classical Bell Beakers have their closest source population in Germany Beakers, and they reject an origin close to Rhine Beakers (i.e. Beakers from the British Isles, the Netherlands, or northern France), ergo the Single Grave culture was not the origin of the Bell Beaker culture, either (see here).

Map of Y-DNA haplogroups among Iberian Bell Beaker samples. See full map.
Map of mtDNA haplogroups among Iberian Bell Beaker samples. See full map.

Early Bronze Age

Iberian Early Bronze Age groups and likely population and culture expansions. See full map.

Interestingly, the European Early Bronze Age in Iberia is still a period of adjustments before reaching the final equilibrium. Unlike the situation in the British Isles, where Bell Beakers brought about a swift population replacement, Iberia shows – like the Nordic Late Neolithic period – centuries of genomic balancing between Indo-European- and non-Indo-European-speaking peoples, as could be suggested by hydrotoponymic research alone.

(3) Palaeo-Indo-European-speaking Old Europeans occupied first the whole Iberian Peninsula, before the potential expansion of one or more non-Indo-European-speaking groups, which confirms the known relative chronology of hydrotoponymic layers of Iberia.

Natural neighbor interpolation of Germany_Beaker ancestry in Iberia during the Early Bronze Age period (ca. 2250-1750 BC). See full map.

This balancing is seen in terms of Germany_Beaker vs. Iberia_Chalcolithic ancestry, but also in terms of Y-chromosome haplogroups, with the most interesting late developments happening in southern Iberia, around the territory where El Argar eventually emerged in radical opposition to the Bell Beaker culture.

Map of Y-DNA haplogroups among Iberia Early Bronze Age samples. See full map.

(4) Bell Beakers and descendants expanded under male-driven migrations, proper of the Indo-European patrilineal tradition, seen in Yamnaya and even earlier in Khvalynsk:

We obtained lower proportions of ancestry related to Germany_Beaker on the X-chromosome than on the autosomes (Table S14), although the Z-score for the differences between the estimates is 2.64, likely due to the large standard error associated to the mixture proportions in the X-chromosome.


Map of mtDNA haplogroups among Iberia Early Bronze Age samples. See full map.

Regarding the PCA, Iberia Bronze Age samples occupy an intermediate cluster between Iberia Chalcolithic and Bell Beakers of steppe ancestry, with Yamnaya-rich samples from the north (Asturias, Burgos) representing the likely source Old European population whose languages survived well into the Roman Iron Age:

PCA of ancient European samples. Marked and labelled are Bronze Age groups and relevant samples. See full image.

Middle Bronze Age

Iberian Middle Bronze Age groups and likely population and culture expansions. See full map.

During the Middle Bronze Age, the equilibrium reached earlier is reversed, with a (likely non-Indo-European-speaking) Argaric sphere of influence expanding to the west and north featuring Iberia Chalcolithic and lesser amount of Germany_Beaker ancestry, present now in the whole Peninsula, although in varying degrees.

Natural neighbor interpolation of Germany_Beaker ancestry in Iberia during the Middle Bronze Age period (ca. 1750-1250 BC). See full map.

All Iberian groups were probably already under a bottleneck of R1b-DF27 lineages, although it is likely that specific subclades differed among regions:

Map of Y-DNA haplogroups among Iberia Middle Bronze Age samples. See full map.
Map of mtDNA haplogroups among Iberia Middle Bronze Age samples. See full map.

Late Bronze Age

Iberian Late Bronze Age groups and likely population and culture expansions. See full map.

The Late Bronze Age represents the arrival of the Urnfield culture, which probably expanded with Celtic-speaking peoples. A Late Bronze Age transect before their genetic impact still shows a prevalent Germany_Beaker-like Steppe ancestry, probably peaking in north/west Iberia:

Natural neighbor interpolation of Germany_Beaker ancestry in Iberia during the Late Bronze Age period (ca. 1250-750 BC). See full map.

(5) Galaico-Lusitanians were descendants of Iberian Beakers of Germany_Beaker ancestry and hg. R1b-M269. Autosomal data of samples I7688 and I7687, of the Final Bronze (end of the reported 1200-700 BC period for the samples), from Gruta do Medronhal (Arrifana, Coimbra, Portugal) confirms this.

In the 1940s, human bones, metallic artifacts (n=37) and non-human bones were discovered in the natural cave of Medronhal (Arrifana, Coimbra). All these findings are currently housed in the Department of Life Sciences of the University of Coimbra and are analyzed by a multidisciplinary team. The artifacts suggest a date at the beginning of the 1st millennium BC, which is confirmed by radiocarbon date of a human fibula: 890–780 cal BCE (2650±40 BP, Beta–223996). This natural cave has several rooms and corridors with two entrances. No information is available about the context of the human remains. Nowadays these remains are housed mixed and correspond to a minimum number of 11 individuals, 5 adults and 6 non-adults.

In particular, sample I7687 shows hg. R1b-M269, with no available quality SNPs, positive or negative, under it (see full report). They represent thus another strong support of the North-West Indo-European expansion with Bell Beakers.

Map of Y-DNA haplogroups among Iberian Late Bronze Age samples. See full map.
Map of mtDNA haplogroups among Iberian Late Bronze Age samples. See full map.

NOTE. To understand how the region around Coimbra was (Proto-)Lusitanian – and not just Old European in general – until the expansion of the Turduli Oppidani, see any recent paper on Bronze Age expansion of warrior stelae, hydrotoponymy, anthroponymy, or theonymy (see e.g. about Spear-vocabulary).

Iron Age

Iberian Pre-Roman Iron Age groups and likely population and culture expansions. See full map.

In a complex period of multiple population movements and language replacements, the temporal transect in Olalde et al. (2019) offers nevertheless relevant clues for the Pre-Roman Iron Age:

(6) The expansion of Celtic languages was associated with the spread of France_Beaker-like ancestry, most likely already with the LBA Urnfield culture, since a Tartessian and a Pre-Iberian samples (both dated ca. 700-500 BC) already show this admixture, in regions which some centuries earlier did not show it. Similarly, a BA sample from Álava ca. 910–840 BC doesn’t show it, and later Celtiberian samples from the same area (ca. 4th c. BC and later) show it, depicting a likely north-east to west/south-west routes of expansion of Celts.

Natural neighbor interpolation of France_Beaker ancestry in Iberia during the Pre-Roman Iron Age period (ca. 750-250 BC). See full map.

(7) The distribution of Germany_Beaker ancestry peaked, by the Iron Age, among Old Europeans from west Iberia, including Galaico-Lusitanians and probably also Astures and Cantabri, in line with what was expected before genetic research:

Natural neighbor interpolation of Germany_Beaker ancestry in Iberia during the Pre-Roman Iron Age period (ca. 750-250 BC). See full map.

A probably more precise picture of the Final Bronze – Early Iron Age transition is obtained by including the Final Bronze samples I2469 from El Sotillo, Álava (ca. 910-875 BC) as Celtic ancestry buffer to the west, and the sample I3315 from Menorca (ca. 904-861 BC), lacking more recent ones from intermediate regions:

Natural neighbor interpolation of Germany_Beaker ancestry in Iberia during the Final Bronze Age – Early Iron Age transition. See full map.
Natural neighbor interpolation of France_Beaker ancestry in Iberia during the Final Bronze Age – Early Iron Age transition. See full map.

In terms of Y-DNA and mtDNA haplogroups, the situation is difficult to evaluate without more samples and more reported subclades:

Map of Y-DNA haplogroups among Iberian Iron Age samples. See full map.
Map of mtDNA haplogroups among Iberian Iron Age samples. See full map.

In the PCA, Proto-Lusitanian samples occupy an intermediate cluster between Iberian Bronze Age and Bronze Age North (see above), including the Final Bronze sample from Álava, while Celtic-speaking peoples (including Pre-Iberians and Iberians of Celtic descent from north-east Iberia) show a similar position – albeit evidently unrelated – due to their more recent admixture between Iberian Bronze Age and Urnfield/Hallstatt from Central Europe:

PCA of ancient European samples. Marked and labelled are Iron Age groups and relevant samples. See full image.

(8) Iberian-speaking peoples in north-east Iberia represent a recent expansion of the language from the south, possibly accompanied by an increase in Iberia_Chalcolithic/Germany_Beaker admixture from east/south-east Iberia.

(9) Modern Basques represent a recent isolation + Y-DNA bottlenecks after the Roman Iron Age population movements, probably from Aquitanians migrating south of the Pyrenees, admixing with local peoples, and later becoming isolated during the Early Middle Ages and thereafter:

[Modern Basques] overlap genetically with Iron Age populations showing substantial levels of Steppe ancestry.

Assuming that France_Beaker ancestry is associated with the Urnfield culture (spreading with Celtic-speaking peoples), Vasconic speakers were possibly represented by some population – most likely from France – whose ancestry is close to Rhine Beakers (see here).

Alternatively, a Vasconic language could have survived in some France/Iberia_Chalcolithic-like population that got isolated north of the Pyrenees close to the Atlantic Façade during the Bronze Age, and who later admixed with Celtic-speaking peoples south of the Pyrenees, such as the Vascones, to the point where their true ancestry got diluted.

In any case, the clear Celtic Steppe-like admixture of modern Basques supports for the time being their recent arrival to Aquitaine before the proto-historical period, which is in line with hydrotoponymic research.


The most interesting aspects to discuss after the publication of Olalde et al. (2019) would have been thus the nature of controversial Palaeohispanic peoples for which there is not much linguistic data, such as:

  • the Astures and the Cantabri, usually considered Pre-Celtic Indo-European (see here);
  • the Vaccaei, usually considered Celtic;
  • the Vettones, traditionally viewed as sharing the same language as Lusitanians due to their apparent shared hydrotoponymic, anthroponymic, and/or theonymic layers, but today mostly viewed as having undergone Celticization and helped the westward expansion of Celtic languages (and archaeologically clearly divided from Old European hostile neighbours to the west by their characteristic verracos);
  • the Pellendones or the Carpetani, who were once considered Pre-Celtic Indo-Europeans, too;
  • the nature of Tartessian as Indo-European, or maybe even as “Celtic”, as defended by Koch;
  • or the potential remote connection of Basque and Iberian languages in a common trunk featuring Iberian/France_Chalcolithic ancestry (also including Palaeo-Sardo).
Pre-Roman Palaeohispanic peoples ca. 300 BC. See full map. Image modified from the version at Wikipedia, a good example of how to disseminate the wrong ideas about Palaeohispanic languages.

Despite these interesting questions still open for discussion, the paper remarked something already known for a long time: that modern Basques had steppe ancestry and Y-DNA proper of the Yamnaya 5,000 years ago, and that Bell Beakers had brought this steppe ancestry and R1b-P312 lineages to Iberia. This common Basque-centric interpretation of Iberian prehistory is the consequence of a 19th-century tradition of obsessively imagining Vasconic-speaking peoples in their medieval territories extrapolated to Cro-Magnons and Atapuerca (no, really), inhabiting undisturbed for millennia a large territory encompassing the whole Iberia and France, “reduced” or “broken” only with the arrival of Celts just before the Roman conquests. A recursive idea of “linguistic autochthony” and “genetic purity” of the peoples of Iberia that has never had any scientific basis.

Similarly, this paper offered the Nth proof already in population genomics that traditional nativist claims for the origin of the Bell Beaker folk in Western Europe were wrong, both southern (nativist Iberian origin) and northern European (nativist Lower Rhine origin). Both options could be easily rejected with phylogeography since 2015, they were then rejected in Olalde et al. and Mathieson et al (2017), then again with the update of many samples in Olalde et al. (2018) and Mathieson et al (2018), and it has most clearly been rejected recently with data from Wang et al. (2018) and its Yamnaya Hungary samples. Findings from Olalde et al. (2019) are just another nail to coffins that should have been well buried by now.

Even David Anthony didn’t have any doubt in his latest model (2017) about the Carpathian Basin origin of North-West Indo-Europeans (see here), and his latest update to the Proto-Indo-European homeland question (2019) shows that he is convinced now about R1b bottlenecks and proper Pre-Yamnaya ancestry stemming from a time well before the Bell Beaker expansion. This won’t be the last setback to supporters of zombie theories: like the hypotheses of an Anatolian, Armenian, or OIT origin of the PIE homeland, other mythical ideas are so entrenched in nationalist and/or nativist tradition that many supporters will no doubt prefer them to die hard, under the most numerous and shameful rejections of endlessly remade reactionary models.


Yamnaya ancestry: mapping the Proto-Indo-European expansions


The latest papers from Ning et al. Cell (2019) and Anthony JIES (2019) have offered some interesting new data, supporting once more what could be inferred since 2015, and what was evident in population genomics since 2017: that Proto-Indo-Europeans expanded under R1b bottlenecks, and that the so-called “Steppe ancestry” referred to two different components, one – Yamnaya or Steppe_EMBA ancestry – expanding with Proto-Indo-Europeans, and the other one – Corded Ware or Steppe_MLBA ancestry – expanding with Uralic speakers.

The following maps are based on formal stats published in the papers and supplementary materials from 2015 until today, mainly on Wang et al. (2018 & 2019), Mathieson et al. (2018) and Olalde et al. (2018), and others like Lazaridis et al. (2016), Lazaridis et al. (2017), Mittnik et al. (2018), Lamnidis et al. (2018), Fernandes et al. (2018), Jeong et al. (2019), Olalde et al. (2019), etc.

NOTE. As in the Corded Ware ancestry maps, the selected reports in this case are centered on the prototypical Yamnaya ancestry vs. other simplified components, so everything else refers to simplistic ancestral components widespread across populations that do not necessarily share any recent connection, much less a language. In fact, most of the time they clearly didn’t. They can be interpreted as “EHG that is not part of the Yamnaya component”, or “CHG that is not part of the Yamnaya component”. They can’t be read as “expanding EHG people/language” or “expanding CHG people/language”, at least no more than maps of “Steppe ancestry” can be read as “expanding Steppe people/language”. Also, remember that I have left the default behaviour for color classification, so that the highest value (i.e. 1, or white colour) could mean anything from 10% to 100% depending on the specific ancestry and period; that’s what the legend is for… But, fere libenter homines id quod volunt credunt.


  1. Neolithic or the formation of Early Indo-European
  2. Eneolithic or the expansion of Middle Proto-Indo-European
  3. Chalcolithic / Early Bronze Age or the expansion of Late Proto-Indo-European
  4. European Early Bronze Age and MLBA or the expansion of Late PIE dialects

1. Neolithic

Anthony (2019) agrees with the most likely explanation of the CHG component found in Yamnaya, as derived from steppe hunter-fishers close to the lower Volga basin. The ultimate origin of this specific CHG-like component that eventually formed part of the Pre-Yamnaya ancestry is not clear, though:

The hunter-fisher camps that first appeared on the lower Volga around 6200 BC could represent the migration northward of un-admixed CHG hunter-fishers from the steppe parts of the southeastern Caucasus, a speculation that awaits confirmation from aDNA.

Natural neighbor interpolation of CHG ancestry among Neolithic populations. See full map.

The typical EHG component that formed part eventually of Pre-Yamnaya ancestry came from the Middle Volga Basin, most likely close to the Samara region, as shown by the sampled Samara hunter-gatherer (ca. 5600-5500 BC):

After 5000 BC domesticated animals appeared in these same sites in the lower Volga, and in new ones, and in grave sacrifices at Khvalynsk and Ekaterinovka. CHG genes and domesticated animals flowed north up the Volga, and EHG genes flowed south into the North Caucasus steppes, and the two components became admixed.

Natural neighbor interpolation of EHG ancestry among Neolithic populations. See full map.

To the west, in the Dnieper-Dniester area, WHG became the dominant ancestry after the Mesolithic, at the expense of EHG, revealing a likely mating network reaching to the north into the Baltic:

Like the Mesolithic and Neolithic populations here, the Eneolithic populations of Dnieper-Donets II type seem to have limited their mating network to the rich, strategic region they occupied, centered on the Rapids. The absence of CHG shows that they did not mate frequently if at all with the people of the Volga steppes (…)

Natural neighbor interpolation of WHG ancestry among Neolithic populations. See full map.

North-West Anatolia Neolithic ancestry, proper of expanding Early European farmers, is found up to border of the Dniester, as Anthony (2007) had predicted.

Natural neighbor interpolation of Anatolia Neolithic ancestry among Neolithic populations. See full map.

2. Eneolithic

From Anthony (2019):

After approximately 4500 BC the Khvalynsk archaeological culture united the lower and middle Volga archaeological sites into one variable archaeological culture that kept domesticated sheep, goats, and cattle (and possibly horses). In my estimation, Khvalynsk might represent the oldest phase of PIE.

(…) this middle Volga mating network extended down to the North Caucasian steppes, where at cemeteries such as Progress-2 and Vonyuchka, dated 4300 BC, the same Khvalynsk-type ancestry appeared, an admixture of CHG and EHG with no Anatolian Farmer ancestry, with steppe-derived Y-chromosome haplogroup R1b. These three individuals in the North Caucasus steppes had higher proportions of CHG, overlapping Yamnaya. Without any doubt, a CHG population that was not admixed with Anatolian Farmers mated with EHG populations in the Volga steppes and in the North Caucasus steppes before 4500 BC. We can refer to this admixture as pre-Yamnaya, because it makes the best currently known genetic ancestor for EHG/CHG R1b Yamnaya genomes.

From Wang et al (2019):

Three individuals from the sites of Progress 2 and Vonyuchka 1 in the North Caucasus piedmont steppe (‘Eneolithic steppe’), which harbour EHG and CHG related ancestry, are genetically very similar to Eneolithic individuals from Khvalynsk II and the Samara region. This extends the cline of dilution of EHG ancestry via CHG-related ancestry to sites immediately north of the Caucasus foothills

Natural neighbor interpolation of Pre-Yamnaya ancestry among Neolithic populations. See full map. This map corresponds roughly to the map of Khvalynsk-Novodanilovka expansion, and in particular to the expansion of horse-head pommel-scepters (read more about Khvalynsk, and specifically about horse symbolism)

NOTE. Unpublished samples from Ekaterinovka have been previously reported as within the R1b-L23 tree. Interestingly, although the Varna outlier is a female, the Balkan outlier from Smyadovo shows two positive SNP calls for hg. R1b-M269. However, its poor coverage makes its most conservative haplogroup prediction R-M343.

The formation of this Pre-Yamnaya ancestry sets this Volga-Caucasus Khvalynsk community apart from the rest of the EHG-like population of eastern Europe.

Natural neighbor interpolation of non-Pre-Yamnaya EHG ancestry among Eneolithic populations. See full map.

Anthony (2019) seems to rely on ADMIXTURE graphics when he writes that the late Sredni Stog sample from Alexandria shows “80% Khvalynsk-type steppe ancestry (CHG&EHG)”. While this seems the most logical conclusion of what might have happened after the Suvorovo-Novodanilovka expansion through the North Pontic steppes (see my post on “Steppe ancestry” step by step), formal stats have not confirmed that.

In fact, analyses published in Wang et al. (2019) rejected that Corded Ware groups are derived from this Pre-Yamnaya ancestry, a reality that had been already hinted in Narasimhan et al. (2018), when Steppe_EMBA showed a poor fit for expanding Srubna-Andronovo populations. Hence the need to consider the whole CHG component of the North Pontic area separately:

Natural neighbor interpolation of non-Pre-Yamnaya CHG ancestry among Eneolithic populations. See full map. You can read more about population movements in the late Sredni Stog and closer to the Proto-Corded Ware period.

NOTE. Fits for WHG + CHG + EHG in Neolithic and Eneolithic populations are taken in part from Mathieson et al. (2019) supplementary materials (download Excel here). Unfortunately, while data on the Ukraine_Eneolithic outlier from Alexandria abounds, I don’t have specific data on the so-called ‘outlier’ from Dereivka compared to the other two analyzed together, so these maps of CHG and EHG expansion are possibly showing a lesser distribution to the west than the real one ca. 4000-3500 BC.

Natural neighbor interpolation of WHG ancestry among Eneolithic populations. See full map.

Anatolia Neolithic ancestry clearly spread to the east into the north Pontic area through a Middle Eneolithic mating network, most likely opened after the Khvalynsk expansion:

Natural neighbor interpolation of Anatolia Neolithic ancestry among Eneolithic populations. See full map.
Natural neighbor interpolation of Iran Chl. ancestry among Eneolithic populations. See full map.

Regarding Y-chromosome haplogroups, Anthony (2019) insists on the evident association of Khvalynsk, Yamnaya, and the spread of Pre-Yamnaya and Yamnaya ancestry with the expansion of elite R1b-L754 (and some I2a2) individuals:

Y-DNA haplogroups in West Eurasia during the Early Eneolithic in the Pontic-Caspian steppes. See full map, and see culture, ADMIXTURE, Y-DNA, and mtDNA maps of the Early Eneolithic and Late Eneolithic.

3. Early Bronze Age

Data from Wang et al. (2019) show that Corded Ware-derived populations do not have good fits for Eneolithic_Steppe-like ancestry, no matter the model. In other words: Corded Ware populations show not only a higher contribution of Anatolia Neolithic ancestry (ca. 20-30% compared to the ca. 2-10% of Yamnaya); they show a different EHG + CHG combination compared to the Pre-Yamnaya one.

Supplementary Table 13. P values of rank=2 and admixture proportions in modelling Steppe ancestry populations as a three-way admixture of Eneolithic steppe Anatolian_Neolithic and WHG using 14 outgroups.
Left populations: Test, Eneolithic_steppe, Anatolian_Neolithic, WHG.
Right populations: Mbuti.DG, Ust_Ishim.DG, Kostenki14, MA1, Han.DG, Papuan.DG, Onge.DG, Villabruna, Vestonice16, ElMiron, Ethiopia_4500BP.SG, Karitiana.DG, Natufian, Iran_Ganj_Dareh_Neolithic.

Yamnaya Kalmykia and Afanasievo show the closest fits to the Eneolithic population of the North Caucasian steppes, rejecting thus sizeable contributions from Anatolia Neolithic and/or WHG, as shown by the SD values. Both probably show then a Pre-Yamnaya ancestry closest to the late Repin population.

Modelling results for the Steppe and Caucasus cluster. Admixture proportions based on (temporally and geographically) distal and proximal models, showing additional AF ancestry in Steppe groups and additional gene flow from the south in some of the Steppe groups as well as the Caucasus groups. See tables above. Modified from Wang et al. (2019). Within a blue square, Yamnaya-related groups; within a cyan square, Corded Ware-related groups. Green background behind best p-values. In red circle, SD of AF/WHG ancestry contribution in Afanasevo and Yamnaya Kalmykia, with ranges that almost include 0%.

EBA maps include data from Wang et al. (2018) supplementary materials, specifically unpublished Yamnaya samples from Hungary that appeared in analysis of the preprint, but which were taken out of the definitive paper. Their location among Yamnaya settlers from Hungary is speculative, although most uncovered kurgans in Hungary are concentrated in the Tisza-Danube interfluve.

Natural neighbor interpolation of Pre-Yamnaya ancestry among Early Bronze Age populations. See full map. This map corresponds roughly with the known expansion of late Repin/Yamnaya settlers.

The Y-chromosome bottleneck of elite males from Proto-Indo-European clans under R1b-L754 and some I2a2 subclades, already visible in the Khvalynsk sampling, became even more noticeable in the subsequent expansion of late Repin/early Yamnaya elites under R1b-L23 and I2a-L699:

Y-DNA haplogroups in West Eurasia during the Yamnaya expansion. See full map and maps of cultures, ADMIXTURE, Y-DNA, and mtDNA of the Early Chalcolithic and Yamnaya Hungary.

Maps of CHG, EHG, Anatolia Neolithic, and probably WHG show the expansion of these components among Corded Ware-related groups in North Eurasia, apart from other cultures close to the Caucasus:

NOTE. For maps with actual formal stats of Corded Ware ancestry from the Early Bronze Age to the modern times, you can read the post Corded Ware ancestry in North Eurasia and the Uralic expansion.

Natural neighbor interpolation of non-Pre-Yamnaya CHG ancestry among Early Bronze Age populations. See full map.
Natural neighbor interpolation of non-Pre-Yamnaya EHG ancestry among Early Bronze Age populations. See full map.
Natural neighbor interpolation of WHG ancestry among Early Bronze Age populations. See full map.
Natural neighbor interpolation of Anatolia Neolithic ancestry among Early Bronze Age populations. See full map.
Natural neighbor interpolation of Iran Chl. ancestry among Early Bronze Age populations. See full map.

4. Middle to Late Bronze Age

The following maps show the most likely distribution of Yamnaya ancestry during the Bell Beaker-, Balkan-, and Sintashta-Potapovka-related expansions.

4.1. Bell Beakers

The amount of Yamnaya ancestry is probably overestimated among populations where Bell Beakers replaced Corded Ware. A map of Yamnaya ancestry among Bell Beakers gets trickier for the following reasons:

  • Expanding Repin peoples of Pre-Yamnaya ancestry must have had admixture through exogamy with late Sredni Stog/Proto-Corded Ware peoples during their expansion into the North Pontic area, and Sredni Stog in turn had probably some Pre-Yamnaya admixture, too (although they don’t appear in the simplistic formal stats above). This is supported by the increase of Anatolia farmer ancestry in more western Yamna samples.
  • Later, Yamnaya admixed through exogamy with Corded Ware-like populations in Central Europe during their expansion. Even samples from the Middle to Upper Danube and around the Lower Rhine will probably show increasing contributions of Steppe_MLBA, at the same time as they show an increasing proportion of EEF-related ancestry.
  • To complicate things further, the late Corded Ware Espersted family (from ca. 2500 BC or later) shows, in turn, what seems like a recent admixture with Yamnaya vanguard groups, with the sample of highest Yamnaya ancestry being the paternal uncle of other individuals (all of hg. R1a-M417), suggesting that there might have been many similar Central European mating networks from the mid-3rd millennium BC on, of (mainly) Yamnaya-like R1b elites displaying a small proportion of CW-like ancestry admixing through exogamy with Corded Ware-like peoples who already had some Yamnaya ancestry.
Natural neighbor interpolation of Yamnaya ancestry among Middle to Late Bronze Age populations (Esperstedt CWC site close to BK_DE, label is hidden by BK_DE_SAN). See full map. You can see how this map correlated with the map of Late Copper Age migrations and Yamanaya into Bell Beaker expansion.

NOTE. Terms like “exogamy”, “male-driven migration”, and “sex bias”, are not only based on the Y-chromosome bottlenecks visible in the different cultural expansions since the Palaeolithic. Despite the scarce sampling available in 2017 for analysis of “Steppe ancestry”-related populations, it appeared to show already a male sex bias in Goldberg et al. (2017), and it has been confirmed for Neolithic and Copper Age population movements in Mathieson et al. (2018) – see Supplementary Table 5. The analysis of male-biased expansion of “Steppe ancestry” in CWC Esperstedt and Bell Beaker Germany is, for the reasons stated above, not very useful to distinguish their mutual influence, though.

Based on data from Olalde et al. (2019), Bell Beakers from Germany are the closest sampled ones to expanding East Bell Beakers, and those close to the Rhine – i.e. French, Dutch, and British Beakers in particular – show a clear excess “Steppe ancestry” due to their exogamy with local Corded Ware groups:

Only one 2-way model fits the ancestry in Iberia_CA_Stp with P-value>0.05: Germany_Beaker + Iberia_CA. Finding a Bell Beaker-related group as a plausible source for the introduction of steppe ancestry into Iberia is consistent with the fact that some of the individuals in the Iberia_CA_Stp group were excavated in Bell Beaker associated contexts. Models with Iberia_CA and other Bell Beaker groups such as France_Beaker (P-value=7.31E-06), Netherlands_Beaker (P-value=1.03E-03) and England_Beaker (P-value=4.86E-02) failed, probably because they have slightly higher proportions of steppe ancestry than the true source population.


The exogamy with Corded Ware-like groups in the Lower Rhine Basin seems at this point undeniable, as is the origin of Bell Beakers around the Middle-Upper Danube Basin from Yamnaya Hungary.

To avoid this excess “Steppe ancestry” showing up in the maps, since Bell Beakers from Germany pack the most Yamnaya ancestry among East Bell Beakers outside Hungary (ca. 51.1% “Steppe ancestry”), I equated this maximum with BK_Scotland_Ach (which shows ca. 61.1% “Steppe ancestry”, highest among western Beakers), and applied a simple rule of three for “Steppe ancestry” in Dutch and British Beakers.

NOTE. Formal stats for “Steppe ancestry” in Bell Beaker groups are available in Olalde et al. (2018) supplementary materials (PDF). I didn’t apply this adjustment to Bk_FR groups because of the R1b Bell Beaker sample from the Champagne/Alsace region reported by Samantha Brunel that will pack more Yamnaya ancestry than any other sampled Beaker to date, hence probably driving the Yamnaya ancestry up in French samples.

The most likely outcome in the following years, when Yamnaya and Corded Ware ancestry are investigated separately, is that Yamnaya ancestry will be much lower the farther away from the Middle and Lower Danube region, similar to the case in Iberia, so the map above probably overestimates this component in most Beakers to the north of the Danube. Even the late Hungarian Beaker samples, who pack the highest Yamnaya ancestry (up to 75%) among Beakers, represent likely a back-migration of Moravian Beakers, and will probably show a contribution of Corded Ware ancestry due to the exogamy with local Moravian groups.

Despite this decreasing admixture as Bell Beakers spread westward, the explosive expansion of Yamnaya R1b male lineages (in words of David Reich) and the radical replacement of local ones – whether derived from Corded Ware or Neolithic groups – shows the true extent of the North-West Indo-European expansion in Europe:

Y-DNA haplogroups in West Eurasia during the Bell Beaker expansion. See full map and see maps of cultures, ADMIXTURE, Y-DNA, and mtDNA of the Late Copper Age and of the Yamnaya-Bell Beaker transition.

4.2. Palaeo-Balkan

There is scarce data on Palaeo-Balkan movements yet, although it is known that:

  1. Yamnaya ancestry appears among Mycenaeans, with the Yamnaya Bulgaria sample being its best current ancestral fit;
  2. the emergence of steppe ancestry and R1b-M269 in the eastern Mediterranean was associated with Ancient Greeks;
  3. Thracians, Albanians, and Armenians also show R1b-M269 subclades and “Steppe ancestry”.

4.3. Sintashta-Potapovka-Filatovka

Interestingly, Potapovka is the only Corded Ware derived culture that shows good fits for Yamnaya ancestry, despite having replaced Poltavka in the region under the same Corded Ware-like (Abashevo) influence as Sintashta.

This proves that there was a period of admixture in the Pre-Proto-Indo-Iranian community between CWC-like Abashevo and Yamnaya-like Catacomb-Poltavka herders in the Sintashta-Potapovka-Filatovka community, probably more easily detectable in this group because of the specific temporal and geographic sampling available.

Supplementary Table 14. P values of rank=3 and admixture proportions in modelling Steppe ancestry populations as a four-way admixture of distal sources EHG, CHG, Anatolian_Neolithic and WHG using 14 outgroups.
Left populations: Steppe cluster, EHG, CHG, WHG, Anatolian_Neolithic
Right populations: Mbuti.DG, Ust_Ishim.DG, Kostenki14, MA1, Han.DG, Papuan.DG, Onge.DG, Villabruna, Vestonice16, ElMiron, Ethiopia_4500BP.SG, Karitiana.DG, Natufian, Iran_Ganj_Dareh_Neolithic.

Srubnaya ancestry shows a best fit with non-Pre-Yamnaya ancestry, i.e. with different CHG + EHG components – possibly because the more western Potapovka (ancestral to Proto-Srubnaya Pokrovka) also showed good fits for it. Srubnaya shows poor fits for Pre-Yamnaya ancestry probably because Corded Ware-like (Abashevo) genetic influence increased during its formation.

On the other hand, more eastern Corded Ware-derived groups like Sintashta and its more direct offshoot Andronovo show poor fits with this model, too, but their fits are still better than those including Pre-Yamnaya ancestry.

Natural neighbor interpolation of non-Pre-Yamnaya EHG ancestry among Middle to Late Bronze Age populations. See full map.
Natural neighbor interpolation of non-Pre-Yamnaya CHG ancestry among Middle to Late Bronze Age populations. See full map.
Natural neighbor interpolation of Anatolia Neolithic ancestry among Middle to Late Bronze Age populations. See full map.
Natural neighbor interpolation of Iran Chl. ancestry among Middle to Late Bronze Age populations. See full map.

NOTE For maps with actual formal stats of Corded Ware ancestry from the Early Bronze Age to the modern times, you should read the post Corded Ware ancestry in North Eurasia and the Uralic expansion instead.

The bottleneck of Proto-Indo-Iranians under R1a-Z93 was not yet complete by the time when the Sintashta-Potapovka-Filatovka community expanded with the Srubna-Andronovo horizon:

Y-DNA haplogroups in West Eurasia during the European Early Bronze Age. See full map and see maps of cultures, ADMIXTURE, Y-DNA, and mtDNA of the Early Bronze Age.

4.4. Afanasevo

At the end of the Afanasevo culture, at least three samples show hg. Q1b (ca. 2900-2500 BC), which seemed to point to a resurgence of local lineages, despite continuity of the prototypical Pre-Yamnaya ancestry. On the other hand, Anthony (2019) makes this cryptic statement:

Yamnaya men were almost exclusively R1b, and pre-Yamnaya Eneolithic Volga-Caspian-Caucasus steppe men were principally R1b, with a significant Q1a minority.

Since the only available samples from the Khvalynsk community are R1b (x3), Q1a(x1), and R1a(x1), it seems strange that Anthony would talk about a “significant minority”, unless Q1a (potentially Q1b in the newer nomenclature) will pop up in some more individuals of those ca. 30 new to be published. Because he also mentions I2a2 as appearing in one elite burial, it seems Q1a (like R1a-M459) will not appear under elite kurgans, although it is still possible that hg. Q1a was involved in the expansion of Afanasevo to the east.

Y-DNA haplogroups in West Eurasia during the Middle Bronze Age. See full map and see maps of cultures, ADMIXTURE, Y-DNA, and mtDNA of the Middle Bronze Age and the Late Bronze Age.

Okunevo, which replaced Afanasevo in the Altai region, shows a majority of hg. Q1b, but also some R1b-M269 samples proper of Afanasevo, suggesting partial genetic continuity.

NOTE. Other sampled Siberian populations clearly show a variety of Q subclades that likely expanded during the Palaeolithic, such as Baikal EBA samples from Ust’Ida and Shamanka with a majority of Q1b, and hg. Q reported from Elunino, Sagsai, Khövsgöl, and also among peoples of the Srubna-Andronovo horizon (the Krasnoyarsk MLBA outlier), and in Karasuk.

From Damgaard et al. Science (2018):

(…) in contrast to the lack of identifiable admixture from Yamnaya and Afanasievo in the CentralSteppe_EMBA, there is an admixture signal of 10 to 20% Yamnaya and Afanasievo in the Okunevo_EMBA samples, consistent with evidence of western steppe influence. This signal is not seen on the X chromosome (qpAdm P value for admixture on X 0.33 compared to 0.02 for autosomes), suggesting a male-derived admixture, also consistent with the fact that 1 of 10 Okunevo_EMBA males carries a R1b1a2a2 Y chromosome related to those found in western pastoralists. In contrast, there is no evidence of western steppe admixture among the more eastern Baikal region region Bronze Age (~2200 to 1800 BCE) samples.

This Yamnaya ancestry has been also recently found to be the best fit for the Iron Age population of Shirenzigou in Xinjiang – where Tocharian languages were attested centuries later – despite the haplogroup diversity acquired during their evolution, likely through an intermediate Chemurchek culture (see a recent discussion on the elusive Proto-Tocharians).

Haplogroup diversity seems to be common in Iron Age populations all over Eurasia, most likely due to the spread of different types of sociopolitical structures where alliances played a more relevant role in the expansion of peoples. A well-known example of this is the spread of Akozino warrior-traders in the whole Baltic region under a partial N1a-VL29-bottleneck associated with the emerging chiefdom-based systems under the influence of expanding steppe nomads.

Y-DNA haplogroups in West Eurasia during the Early Iron Age. See full map and see maps of cultures, ADMIXTURE, Y-DNA, and mtDNA of the Early Iron Age and Late Iron Age.

Surprisingly, then, Proto-Tocharians from Shirenzigou pack up to 74% Yamnaya ancestry, in spite of the 2,000 years that separate them from the demise of the Afanasevo culture. They show more Yamnaya ancestry than any other population by that time, being thus a sort of Late PIE fossils not only in their archaic dialect, but also in their genetic profile:


The recent intrusion of Corded Ware-like ancestry, as well as the variable admixture with Siberian and East Asian populations, both point to the known intense Old Iranian and Old/Middle Chinese contacts. The scarce Proto-Samoyedic and Proto-Turkic loans in Tocharian suggest a rather loose, probably more distant connection with East Uralic and Altaic peoples from the forest-steppe and steppe areas to the north (read more about external influences on Tocharian).

Interestingly, both R1b samples, MO12 and M15-2 – likely of Asian R1b-PH155 branch – show a best fit for Andronovo/Srubna + Hezhen/Ulchi ancestry, suggesting a likely connection with Iranians to the east of Xinjiang, who later expanded as the Wusun and Kangju. How they might have been related to Huns and Xiongnu individuals, who also show this haplogroup, is yet unknown, although Huns also show hg. R1a-Z93 (probably most R1a-Z2124) and Steppe_MLBA ancestry, earlier associated with expanding Iranian peoples of the Srubna-Andronovo horizon.

All in all, it seems that prehistoric movements explained through the lens of genetic research fit perfectly well the linguistic reconstruction of Proto-Indo-European and Proto-Uralic.


Iron Age Tocharians of Yamnaya ancestry from Afanasevo show hg. R1b-M269 and Q1a1

New open access Ancient Genomes Reveal Yamnaya-Related Ancestry and a Potential Source of Indo-European Speakers in Iron Age Tianshan, by Ning et al. Current Biology (2019).

Interesting excerpts (emphasis mine, changes for clarity):

Here, we report the first genome-wide data of 10 ancient individuals from northeastern Xinjiang. They are dated to around 2,200 years ago and were found at the Iron Age Shirenzigou site. We find them to be already genetically admixed between Eastern and Western Eurasians. We also find that the majority of the East Eurasian ancestry in the Shirenzigou individuals is related to northeastern Asian populations, while the West Eurasian ancestry is best presented by ∼20% to 80% Yamnaya-like ancestry. Our data thus suggest a Western Eurasian steppe origin for at least part of the ancient Xinjiang population. Our findings furthermore support a Yamnaya-related origin for the now extinct Tocharian languages in the Tarim Basin, in southern Xinjiang.


The dominant mtDNA lineages of the Shirenzigou people are commonly found in modern and ancient West Eurasian populations, such as U4, U5, and H, while they also have East Eurasian-specific haplogroups A, D4, and G3, preliminarily documenting admixed ancestry from eastern and western Eurasia.

The admixture profile is also shown on the paternal Y chromosome side that 4 out of 6 males in Shirenzigou (Figure S2) belong to the West Eurasian-specific haplogroup R1b (n = 2) and East Eurasian-specific haplogroup Q1a (n = 2), the former is predominant in ancient Yamnaya and nearly 100% in Afanasievo, different from the Middle and Late Bronze Age Steppe groups (Steppe_MLBA) such as Andronovo, [Potapovka], Srubnaya, and Sintashta whose Y chromosomal haplogroup is mainly R1a.



We first carried out principal component analysis (PCA) to assess the genetic affinities of the ancient individuals qualitatively by projecting them onto present-day Eurasian variation (Figure 2). We observed a distinct separation between East and West Eurasians. Our ancient Shirenzigou samples and present-day populations from Central Asia and northwestern China form a genetic cline from East to West in the first PC. The distribution of Shirenzigou samples on the cline is relatively scattered with two major clusters, one being closer to modern-day Uygurs and Kazakhs and the other being closer to recently published ancient Saka and Huns from the Tianshan in Kazakhstan (…).

We applied a formal admixture test using f3 statistics in the form of f3 (Shirenzigou; X, Y) where X and Y are worldwide populations that might be the genetic sources for the Shirenzigou individuals. We observed the most significant signals of admixture in the Shirenzigou samples when using Yamnaya_Samara or Srubnaya as the West Eurasian source and some Northern Asians or Koreans as the East Eurasian source (Table S1). We also plotted the outgroup f3 statistics in the form of f3 (Mbuti; X, Anatolia_Neolithic) and f3 (Mbuti; X, Kostenki14) to visualize the allele sharing between population X and Anatolian farmers. As shown in Figure S3, the Steppe_MLBA populations including Srubnaya, Andronovo, and Sintashta were shifted toward farming populations compared with Yamnaya groups and the Shirenzigou samples. This observation is consistent with ADMIXTURE analysis that Steppe_MLBA populations have an Anatolian and European farmer-related component that Yamnaya groups and the Shirenzigou individuals do not seem to have. The analysis consistently suggested Yamnaya-related Steppe populations were the better source in modeling the West Eurasian ancestry in Shirenzigou.

PCA and ADMIXTURE for Shirenzigou Samples. Modified from the original to include in black squares samples related to Yamnaya.

Genetic Composition of Iron Age Shirenzigou Individuals

We continued to use qpAdm to estimate the admixture proportions in the Shirenzigou samples by using different pairs of source populations, such as Yamnaya_Samara, Afanasievo, Srubnaya, Andronovo, BMAC culture (Bustan_BA and Sappali_Tepe_BA) and Tianshan_Hun as the West Eurasian source and Han, Ulchi, Hezhen, Shamanka_EN as the East Eurasian source. In all cases, Yamnaya, Afanasievo, or Tianshan_Hun always provide the best model fit for the Shirenzigou individuals, while Srubnaya, Andronovo, Bustan_BA and Sappali_Tepe_BA only work in some cases. The Yamnaya_Samara or Afanasievo-related ancestry ranges from ∼20% to 80% in different Shirenzigou individuals, consistent with the scattered distribution on the East-West cline in the PCA


(…) we then modeled Shirenzigou as a three-way admixture of Yamnaya_Samara, Ulchi (or Hezhen) and Han to infer the source from the East Eurasia side that contributed to Shirenzigou. We found the Ulchi or Hezhen and Han-related ancestry had a complicated and unevenly distribution in the Shirenzigou samples. The most Shirenzigou individuals derived the majority of their East Eurasian ancestry from Ulchi or Hezhen-related populations, while the following two individuals M820 and M15-2 have more Han related than Ulchi/Hezhen-related ancestry.

One important question remains, though: how and when did these Proto-Tocharian speakers migrate from the Afanasevo culture in the Altai into the Tarim Basin? The traditional answer, now more likely than ever, is through the Chemurchek culture. See e.g. A re-analysis of the Qiemu’erqieke (Shamirshak) cemeteries, Xinjiang, China, by Jia and Betts JIES (2010) 38(4).

Also, given the apparent lack of (extra farmer ancestry that characterizes) Corded Ware ancestry, if the results were already suspicious before, how likely are now the published R1a(xZ93) and/or radiocarbon dates of the Xiaohe mummies from Li et al. (2010, 2015)? Because, after all, one should have expected in such a late date a generalized admixture with neighbouring Srubna/Andronovo-like populations.


N1c-L392 associated with expanding Turkic lineages in Siberia


Second in popularity for the expansion of haplogroup N1a-L392 (ca. 4400 BC) is, apparently, the association with Turkic, and by extension with Micro-Altaic, after the Uralic link preferred in Europe; at least among certain eastern researchers.

New paper in a recently created journal, by the same main author of the group proposing that Scythians of hg. N1c were Turkic speakers: On the origins of the Sakhas’ paternal lineages: Reconciliation of population genetic / ancient DNA data, archaeological findings and historical narratives, by Tikhonov, Gurkan, Demirdov, and Beyoglu, Siberian Research (2019).

Interesting excerpts:

According to the views of a number of authoritative researchers, the Yakut ethnos was formed in the territory of Yakutia as a result of the mixing of people from the south and the autochthonous population [34].

These three major Sakha paternal lineages may have also arrived in Yakutia at different times and/ or from different places and/or with a difference in several generations instead, or perhaps Y-chromosomal STR mutations may have taken place in situ in Yakutia. Nevertheless, the immediate common ancestor(s) from the Asian Steppe of these three most prevalent Sakha Y-chromosomal STR haplotypes possibly lived during the prominence of the Turkic Khaganates, hence the near-perfect matches observed across a wide range of Eurasian geography, including as far as from Cyprus in the West to Liaoning, China in the East, then Middle Lena in the North and Afghanistan in the South (Table 3 and Figure 5). There may also be haplotypes closely-related to ‘the dominant Elley line’ among Karakalpaks, Uzbeks and Tajiks, however, limitations in the loci coverage for the available dataset (only eight Y-chromosomal STR loci) precludes further conclusions on this matter [25].

17-loci median-joining network analysis of the original/dominant Elley, Unknown and Omogoy Y-chromosomal STR haplotypes with the YHRD matches from outside Yakutia populations.

According to the results presented here, very similar Y-STR haplotypes to that of the original Elley line were found in the west: Afghanistan and northern Cyprus, and in the east: Liaoning Province, China and Ulaanbaator, Northern Mongolia. In the case of the dominant Omogoy line, very closely matching haplotypes differing by a single mutational step were found in the city of Chifen of the Jirin Province, China. The widest range of similar haplotypes was found for the Yakut haplotype Unknown: In Mongolia, China and South Korea. For instance, haplotypes differing by a single step mutation were found in Northern Mongolia (Khalk, Darhad, Uryankhai populations), Ulaanbaator (Khalk) and in the province of Jirin, China (Han population).

14-loci median-joining network analysis for the original/dominant Elley (Ell), Unknown Clan
(Vil), Omogoy (Omo), Eurasian (Eur) and Xiongnu (Xuo) Y-chromosomal STR haplotypes and that for a representative ancient DNA sample (Ch0 or DSQ04) from the Upper Xiajiadian Culture
recovered from the Inner Mongolia Autonomous Region, China.

Notably, Tat-C-bearing Y-chromosomes were also observed in ancient DNA samples from the 2700-3000 years-old Upper Xiajiadian culture in Inner Mongolia, as well as those from the Serteya II site at the Upper Dvina region in Russia and the ‘Devichyi gory’ culture of long barrow burials at the Nevel’sky district of Pskovsky region in Russia. A 14-loci Y-chromosomal STR median-joining network of the most prevalent Sakha haplotypes and a Tat-C-bearing haplotype from one of the ancient DNA samples recovered from the Upper Xiajiadian culture in Inner Mongolia (DSQ04) revealed that the contemporary Sakha haplotype ‘Xuo’ (Table 2, Haplotype ID “Xuo”) classified as that of ‘the Xiongnu clan’ in our current study, was the closest to the ancient Xiongnu haplotype (Figure 6). TMRCA estimate for this 14-loci Y-chromosomal STR network was 4357 ± 1038 years or 2341 ± 1038 BCE, which correlated well with the Upper Xiajiadian culture that was dated to the Late Bronze Age (700-1000 BCE).

Geographical location of ancient samples belonging to major clade N of the Y-chromosome.

NOTE. Also interesting from the paper seems to be the proportion of E1b1b among admixed Russian populations, in a proportion similar to R1a or I2a(xI2a1).

It is tempting to associate the prevalent presence of N1c-L392 in ancient Siberian populations with the expansion of Altaic, by simplistically linking the findings (in chronological order) near Lake Baikal (Damgaard et al. 2018), Upper Xiajiadian (Cui et al. 2013), among Khövsgöl (Jeong et al. 2018), in Huns (Damgaard et al. 2018), and in Mongolic-speaking Avars (Csáky et al. 2019).

However, its finding among Palaeo-Laplandic peoples in the Kola peninsula ca. 1500 BC (Lamnidis et al. 2018) and among Palaeo-Siberian populations near the Yana River (Sikora et al. 2018) ca. AD 1200 should be enough to accept the hypothesis of ancestral waves of expansion of the haplogroup over northern Eurasia, with acculturation and further expansions in the different regions since the Iron Age (see more on its potential expansion waves).

Also, a simple look at the TMRCA and modern distribution was enough to hypothesize long ago the lack of connection of N1c-L392 with Altaic or Uralic peoples. From Ilumäe et al. (2016):

Previous research has shown that Y chromosomes of the Turkic-speaking Yakuts (Sakha) belong overwhelmingly to hg N3 (formerly N1c1). We found that nearly all of the more than 150 genotyped Yakut N3 Y chromosomes belong to the N3a2-M2118 clade, just as in the Turkic-speaking Dolgans and the linguistically distant Tungusic-speaking Evenks and Evens living in Yakutia (Table S2). Hence, the N3a2 patrilineage is a prime example of a male population of broad central Siberian ancestry that is not intrinsic to any linguistically defined group of people. Moreover, the deepest branch of hg N3a2 is represented by a Lebanese and a Chinese sample. This finding agrees with the sequence data from Hallast et al., where one Turkish Y chromosome was also assigned to the same sub-clade. Interestingly, N3a2 was also found in one Bhutan individual who represents a separate sub-lineage in the clade. These findings show that although N3a2 reflects a recent strong founder effect primarily in central Siberia (Yakutia, Sakha), the sub-clade has a much wider distribution area with incidental occurrences in the Near East and South Asia.

Frequency-Distribution Maps of Individual Sub-clades of hg N3a2, by Ilumäe et al. (2016).

The most striking aspect of the phylogeography of hg N is the spread of the N3a3’6-CTS6967 lineages. Considering the three geographically most distant populations in our study—Chukchi, Buryats, and Lithuanians—it is remarkable to find that about half of the Y chromosome pool of each consists of hg N3 and that they share the same sub-clade N3a3’6. The fractionation of N3a3’6 into the four sub-clades that cover such an extraordinarily wide area occurred in the mid-Holocene, about 5.0 kya (95% CI = 4.4–5.7 kya). It is hard to pinpoint the precise region where the split of these lineages occurred. It could have happened somewhere in the middle of their geographic spread around the Urals or further east in West Siberia, where current regional diversity of hg N sub-lineages is the highest (Figure 1B). Yet, it is evident that the spread of the newly arisen sub-clades of N3a3’6 in opposing directions happened very quickly. Today, it unites the East Baltic, East Fennoscandia, Buryatia, Mongolia, and Chukotka-Kamchatka (Beringian) Eurasian regions, which are separated from each other by approximately 5,000–6,700 km by air. N3a3’6 has high frequencies in the patrilineal pools of populations belonging to the Altaic, Uralic, several Indo-European, and Chukotko-Kamchatkan language families. There is no generally agreed, time-resolved linguistic tree that unites these linguistic phyla. Yet, their split is almost certainly at least several millennia older than the rather recent expansion signal of the N3a3’6 sub-clade, suggesting that its spread had little to do with linguistic affinities of men carrying the N3a3’6 lineages.

Frequency-Distribution Maps of Individual Subclade N3a3 / N1a1a1a1a1a-CTS2929/VL29.

It was thus clear long ago that N1c-L392 lineages must have expanded explosively in the 5th millennium through Northern Eurasia, probably from a region to the north of Lake Baikal, and that this expansion – and succeeding ones through Northern Eurasia – may not be associated to any known language group until well into the common era.


Magyar tribes brought R1a-Z645, I2a-L621, and N1a-L392(xB197) lineages to the Carpathian Basin


The Nightmare Week of “N1c=Uralic” proponents (see here) continues, now with preprint Y-chromosome haplogroups from Hun, Avar and conquering Hungarian period nomadic people of the Carpathian Basin, by Neparaczki et al. bioRxiv (2019).


Hun, Avar and conquering Hungarian nomadic groups arrived into the Carpathian Basin from the Eurasian Steppes and significantly influenced its political and ethnical landscape. In order to shed light on the genetic affinity of above groups we have determined Y chromosomal haplogroups and autosomal loci, from 49 individuals, supposed to represent military leaders. Haplogroups from the Hun-age are consistent with Xiongnu ancestry of European Huns. Most of the Avar-age individuals carry east Eurasian Y haplogroups typical for modern north-eastern Siberian and Buryat populations and their autosomal loci indicate mostly unmixed Asian characteristics. In contrast the conquering Hungarians seem to be a recently assembled population incorporating pure European, Asian and admixed components. Their heterogeneous paternal and maternal lineages indicate similar phylogeographic origin of males and females, derived from Central-Inner Asian and European Pontic Steppe sources. Composition of conquering Hungarian paternal lineages is very similar to that of Baskhirs, supporting historical sources that report identity of the two groups.

Interesting excerpts (emphasis mine):

All N-Hg-s identified in the Avars and Conquerors belonged to N1a1a-M178. We have tested 7 subclades of M178; N1a1a2-B187, N1a1a1a2-B211, N1a1a1a1a3-B197, N1a1a1a1a4-M2118, N1a1a1a1a1a-VL29, N1a1a1a1a2-Z1936 and the N1a1a1a1a2a1c1-L1034 subbranch of Z1936. The European subclades VL29 and Z1936 could be excluded in most cases, while the rest of the subclades are prevalent in Siberia 23 from where this Hg dispersed in a counter-clockwise migratory route to Europe (…). All the 5 other Avar samples belonged to N1a1a1a1a3-B197, which is most prevalent in Chukchi, Buryats, Eskimos, Koryaks and appears among Tuvans and Mongols with lower frequency.

First two components of PCA from Hg N1a subbranch distribution in 51 populations including Avars and Conquerors. Colors indicate geographic regions. Three letter codes are given in Supplementary Table S5.

By contrast two Conquerors belonged to N1a1a1a1a4-M2118, the Y lineage of nearly all Yakut males, being also frequent in Evenks, Evens and occurring with lower frequency among Khantys, Mansis and Kazakhs.

Three Conqueror samples belonged to Hg N1a1a1a1a2-Z1936 , the Finno-Permic N1a branch, being most frequent among northeastern European Saami, Finns, Karelians, as well as Komis, Volga Tatars and Bashkirs of the Volga-Ural region.Nevertheless this Hg is also present with lower frequency among Karanogays, Siberian Nenets, Khantys, Mansis, Dolgans, Nganasans, and Siberian Tatars.

The west Eurasian R1a1a1b1a2b-CTS1211 subclade of R1a is most frequent in Eastern Europe especially among Slavic people. This Hg was detected just in the Conqueror group (K2/18, K2/41 and K1/10). Though CTS1211 was not covered in K2/36 but it may also belong to this sub-branch of Z283.

Hg I2a1a2b-L621 was present in 5 Conqueror samples, and a 6th sample form Magyarhomorog (MH/9) most likely also belongs here, as MH/9 is a likely kin of MH/16 (see below). This Hg of European origin is most prominent in the Balkans and Eastern Europe, especially among Slavic speaking groups. It might have been a major lineage of the Cucuteni-Trypillian culture and it was present in the Baden culture of the Chalcolithic Carpathian Basin.

Image modified from the paper, with drawn red square around lineages of likely Ugric origin, and squares around R1a-Z93, R1a-Z283, N1a-Z1936, and N1a-M2004 samples. Y-Hg-s determined from 46 males grouped according to sample age, cemetery and Hg. Hg designations are given according to ISOGG Tree 2019. Grey shading designate distinguished individuals with rich grave goods, color shadings denote geographic origin of Hg-s according to Fig. 1. For samples K3/1 and K3/3 the innermost Hg defining marker U106* was not covered, but had been determined previously.

We identified potential relatives within Conqueror cemeteries but not between them. The uniform paternal lineages of the small Karos3 (19 graves) and Magyarhomorog (17 graves) cemeteries approve patrilinear organization of these communities. The identical I2a1a2b Hg-s of Magyarhomorog individuals appears to be frequent among high-ranking Conquerors, as the most distinguished graves in the Karos2 and 3 cemeteries also belong to this lineage. The Karos2 and Karos3 leaders were brothers with identical mitogenomes 11 and Y-chromosomal STR profiles (Fóthi unpublished). The Sárrétudvari commoner cemetery seems distinct from the others, containing other sorts of European Hg-s. Available Y-chromosomal and mtDNA data from this cemetery suggest that common people of the 10th century rather represented resident population than newcomers. The great diversity of Y Hg-s, mtDNA Hg-s, phenotypes and predicted biogeographic classifications of the Conquerors indicate that they were relatively recently associated from very diverse populations.

Surprising about the Hungarian conquerors – although in line with the historical accounts – is the varied patrilineal origin of clans, including Q1a, G2a2b, I1, E1b1b, R1b, J1, or J2 – some of which (depending on specific lineages) may have appeared earlier in the Carpathian Basin or south-eastern Europe.

However, out of the 27 conqueror elite samples, 17 are of haplogroups most likely related to Ugric populations beyond the Urals: R1a-Z645, I2-L621, and two specific N1a-L392 lineages (see below). In fact, there are three high-ranking conqueror elites of hg. I2-L621 (one of them termed a “leader”, brother to an unpublished leader of Karos3, and all of them possibly family), one of hg. R1a-Z280, one of hg. R1a-Z93 (which should be added to the Árpáds), and one of hg. N1a-Z1936, which gives a good idea of the ruling class among the elite Ugric settlers.

NOTE. The Q1a sample is also likely to be found in the mixed population of the West Siberian forest-steppes, since it was found in Mesolithic-Neolithic samples from eastern Europe to Lake Baikal, and in Bronze Age Siberian groups, although admittedly it may have formed part of an Avar Transtisza group, or even earlier Hunnic or Scythian groups along the steppes. Without precise subclades it’s impossible to know.

The seven chieftains of the Hungarians, detail of Arrival of the Hungarians, from Árpád Feszty’s and his assistants’ vast (1800 m2) cyclorama, painted to celebrate the 1000th anniversary of the Magyar conquest of Hungary, now displayed at the Ópusztaszer National Heritage Park in Hungary. Image from Wikipedia.


I2a-L621 (xS17250) or I2a1b2 in the old nomenclature, is found in 6 early conquerors (including one leader), on a par with R1a and N samples. This haplogroup is found widely distributed in ancient samples, due to its early split (formed ca. 9200 BC, TMRCA ca. 4500 BC) and expansion, probably with Neolithic populations. I can’t seem to find samples of this early haplogroup from the Carpathian Basin, as mentioned in the text, although it wouldn’t be strange, because it appears also in Neolithic Iberia, and in modern populations from western Europe.

Nevertheless, I2a-L621 samples seem to be concentrated mainly in Mesolithic-Neolithic cultures of Fennoscandia, and appeared also in Sikora et al. (2017) in a sample of the High Middle Ages from Sunghir (ca. AD 1100-1200), probably from the Vladimir-Suzdalian Rus’, in a region where clearly tribes of Volga Finns were being assimilated at the time. The reported SNP call by Genetiker is A16681 (see Yfull), deep within I2a-CTS10228. It is possibly also behind a modern Saami from Chalmny Varre (ca. AD 1800) of hg. I2a in Lamnidis et al. (2018).

Lacking precise subclades from Hungarian conquerors this is pure speculation, but modern samples may also point to I2a-CTS10228 (formed ca. 3100 BC, TMRCA ca. 1800 BC) as a Finno-Ugric lineage in common with R1a, which must have expanded to the Urals and beyond with eastern Corded Ware groups or (more likely) succeeding cultures. This is in line with the association of certain I2a lineages with modern Uralic peoples or populations from their historical regions in eastern Europe, and linked thus to the most likely homeland of Uralians in the eastern European forests:

Additional file 6: Table S5. Y chromosome haplogroup frequencies in Eurasia. Modified by me: in bold haplogroup N1c and R1a from Uralic-speaking populations, with those in red showing where R1a is the major haplogroup. Observe that all Uralic subgroups – Finno-Permic, Ugric, and Samoyedic – have some populations with a majority of R1a, and also of I lineages. Data from Tambets et al. (2018).


Regarding the important question of the ethnic makeup of Ugric populations stemming from the Urals, the most interesting (and expected) data is the presence of R1a-Z645 lineages among high-ranking conquerors, in particular four R1a-Z280 subclades proper of Finno-Ugrians.

This proves that, in line with the old split and expansion of R1a-CTS1211 (formed ca. 2600 BC, TMRCA ca. 2400 BC), and its finding in Bronze Age Fennoscandian samples, only some late R1a-Z280 (xZ92) lineages (see Z280 on YFull) may show a clear identification with early acculturated Uralic speakers, with the main early acculturated Balto-Slavic R1a haplogroup remaining R1a-M458.

I recently hypothesized this late connection of Slavs with very specific R1a-Z280 (xZ92) lineages based on analyses of modern populations (like Slovenians), because the connection of ancient Finno-Ugrians with modern Z92 samples was already evident:

(…) subclades of hg. R1a1a1b1a2-Z280 (xR1a1a1b1a2a-Z92) seem to have also been involved in early Slavic expansions, like R1a1a1b1a2b3a-CTS3402 (formed ca. 2200 BC, TMRCA ca. 2200 BC), found among modern West, South, and East Slavic populations and in Fennoscandia, prevalent e.g. among modern Slovenians which points to a northern origin of its expansion (Maisano Delser et al. 2018).

This finding also supports the expected shared R1a-Z280 lineages among ancient Finno-Ugric populations, as predicted from the study of modern Permic and Ugric peoples in Dudás et al. (2019).

Modified image, from Underhill et al. (2015). Spatial frequency distributions of Z282 (green) and Z93 (blue) affiliated haplogroups. Notice the distribution of R1a-Z280 (xZ92), i.e. R1a-M558, compared to the ancient Finno-Ugric distribution.

Furthermore, while we don’t have precise R1a-Z93 lineages to compare with the new Hunnic sample reported, we already know that some archaic R1a-Z2124 subclades stem from the forest-steppe areas of the Cis- and Trans-Urals, and the two newly reported R1a-Z93 Hungarian conqueror elites, like those of the Árpád dynasty, probably belong to them.

There is an obvious lack of continuity in specific paternal lineages among the Hunnic, the Avar, and the Conqueror periods, which makes any simplistic identification of all R1a-Z93 lineages as stemming from Avars, Huns, or the Iron Age Pontic-Caspian steppes clearly flawed. Comparing R1a-Z93 in Hungarian Conquerors with Huns is like comparing them with samples of the Srubna or earlier periods… Similarly, comparing the Hunnic R1b-U106 or the early Avar I1 to later Hungarian samples is not warranted without precise subclades, because they most likely correspond to different Germanic populations: Goths among Huns, then Longobards, then likely peoples descended from Franks and Irish Monks (the latter with R1b-P312).


Second behind R1a subclades are, as expected, N1a-L392 (N1c in the old nomenclature).

Avars are dominated by a specific N1a-L392 subclade, N1a-B197, as we recently discovered in Csáky et al. (2019).

Hungarian conquerors show three N1a-Z1936 subclades, which is known to stem from the northern Ural region, including the Arctic (likely Palaeo-Laplandic peoples) and cross-stamped cultures of the northern Eurasian forests.

Frequency-Distribution Maps of Individual Subclade N3a4 / N1a1a1a1a2-Z1936, probably with the Samic (first) and Fennic (later) expansions into Paleo-Lakelandic and Palaeo-Laplandic territories.

On the other hand, the two N1a-M2118 lineages are more clearly associated with Palaeo-Siberian populations east of the Urals, but became incorporated into the Ugric stock in the Trans-Urals region probably in the same way as N1a-Z1936, by infiltration from (and acculturation of) hunter-gatherers of forest and taiga cultures.

NOTE. You can read more about the infiltration of N1a lineages in the recent post Corded Ware—Uralic (IV): Hg R1a and N in Finno-Ugric and Samoyedic expansions, and in the specific sections for each Uralic group in A Clash of Chiefs.

Frequency-Distribution Maps of Individual Sub-clades of hg N3a2, by Ilumäe et al. (2016).


The picture offered by the paper on Hungarian Conquerors, while in line with historical accounts of multi-ethnic tribes incorporating regional lineages, shows nevertheless patrilineal clans clearly associated with Uralic peoples, in a distribution which could have been easily inferred from ancient Trans-Uralian forest-steppe cultures and modern samples (even regarding I2a-L621).

In spite of this, there is a great deal of discussion in the paper about specific N1a subclades in Hungarian conquerors, while the presence of R1a-Z280 (among early Magyar elites!) is interpreted, as always, as recently acculturated Slavs. This is sadly coupled with the simplistic identification of I2a-L621 as of local origin around the Carpathians.

The introduction of the paper to the history of Hungarians is also weird, for example giving credibility to the mythic accounts of the Árpád dynasty’s origin in Attila, which is in line, I guess, with what the authors intended to support all along, i.e. the association of Magyars with Turks from the Eurasian steppes, which they are apparently willing to achieve by relating them to haplogroup R1a-Z93

The conclusion is thus written to appease modern nation-building myths more than anything else, like many other papers before it:

It is generally accepted that the Hungarian language was brought to the Carpathian Basin by the Conquerors. Uralic speaking populations are characterized by a high frequency of Y-Hg N, which have often been interpreted as a genetic signal of shared ancestry. Indeed, recently a distinct shared ancestry component of likely Siberian origin was identified at the genomic level in these populations, modern Hungarians being a puzzling exception36. The Conqueror elite had a significant proportion of N Hgs, 7% of them carrying N1a1a1a1a4-M2118 and 10% N1a1a1a1a2-Z1936, both of which are present in Ugric speaking Khantys and Mansis. At the same time none of the examined Conquerors belonged to the L1034 subclade of Z1936, while all of the Khanty Z1936 lineages reported in 37 proved to be L1034 which has not been tested in the 23 study. Population genetic data rather position the Conqueror elite among Turkic groups, Bashkirs and Volga Tatars, in agreement with contemporary historical accounts which denominated the Conquerors as “Turks”. This does not exclude the possibility that the Hungarian language could also have been present in the obviously very heterogeneous, probably multiethnic Conqueror tribal alliance.

So, back to square one, and new circular reasoning: If ancient populations from north-eastern Europe believed to represent ancient Finno-Ugrians are of R1a-Z645 lineages, it’s because they were not Finno-Ugric speakers. If ancient and modern populations known to be of Finno-Ugric language show clear connections with R1a-Z645, it’s because they are “multi-ethnic”.

The only stable basis for discussion in genetic papers, apparently, is the own making of geneticists, with their traditional 2000s “R1a=Indo-European” and “N1c=Uralic”, coupled with national beliefs. It does not matter how many predictions based on that have been proven wrong, or how many predictions based on the Corded Ware = Uralic expansion have been proven right.