Arrival of steppe ancestry with R1b-P312 in the Mediterranean: Balearic Islands, Sicily, and Iron Age Sardinia

steppe-balearic-sicily-sardinia

New preprint The Arrival of Steppe and Iranian Related Ancestry in the Islands of the Western Mediterranean by Fernandes, Mittnik, Olalde et al. bioRxiv (2019)

Interesting excerpts (emphasis in bold; modified for clarity):

Balearic Islands: The expansion of Iberian speakers

Mallorca_EBA dates to the earliest period of permanent occupation of the islands at around 2400 BCE. We parsimoniously modeled Mallorca_EBA as deriving 36.9 ± 4.2% of her ancestry from a source related to Yamnaya_Samara; (…). We next used qpAdm to identify “proximal” sources for Mallorca_EBA’s ancestry that are more closely related to this individual in space and time, and found that she can be modeled as a clade with the (small) subset of Iberian Bell Beaker culture associated individuals who carried Steppe-derived ancestry (p=0.442).

Suppl. Materials: The model used was with Bell_Beaker_Iberia_highsteppe, a group of outliers from Iberia buried in a Bell Beaker mortuary context who unlike most individuals from this context in that region had high proportions of Steppe ancestry (p=0.442).

Our estimates of Steppe ancestry in the two later Balearic Islands individuals are lower than the earlier one: 26.3 ± 5.1% for Formentera_MBA and 23.1 ± 3.6% for Menorca_LBA, but the Middle to Late Bronze Age Balearic individuals are not a clade relative to non-Balearic groups. Specifically, we find that f4(Mbuti.DG, X; Formentera_MBA, Menorca_LBA) is positive when X=Iberia_Chalcolithic (Z=2.6) or X=Sardinia_Nuragic_BA (Z=2.7). While it is tempting to interpret the latter statistic as suggesting a genetic link between peoples of the Talaiotic culture of the Balearic islands and the Nuragic culture of Sardinia, the attraction to Iberia_Chalcolithic is just as strong, and the mitochondrial haplogroup U5b1+16189+@16192 in Menorca_LBA is not observed in Sardinia_Nuragic_BA but is observed in multiple Iberia_Chalcolithic individuals. A possible explanation is that both the ancestors of Nuragic Sardinians and the ancestors of Talaiotic people from the Balearic Islands received gene flow from an unsampled Iberian Chalcolithic-related group (perhaps a mainland group affiliated to both) that did not contribute to Formentera_MBA.

This sample, like another one in El Argar, is of hg. R1b-P312. So there you are, the data that connects the Proto-Iberian expansion (replacing IE-speaking Bell Beakers) to the Iberian Chalcolithic population, signaled by the increase in Iberian Chalcolithic ancestry after the arrival of Bell Beakers, most likely connected originally to the Argaric and post-Argaric expansions during the MBA.

balearic-sicily-sardinia-pca
PCA with previously published ancient individuals (non-filled symbols), projected onto variation from present-day populations (gray squares).

Steppe in Sardinia IA: Phocaeans from Italy?

Most Sardinians buried in a Nuragic Bronze Age context possessed uniparental haplogroups found in European hunter-gatherers and early farmers, including Y-haplogroup R1b1a[xR1b1a1a] which is different from the characteristic R1b1a1a2a1a2 spread in association with the Bell Beaker complex. An exception is individual I10553 (1226-1056 calBCE) who carried Y-haplogroup J2b2a, previously observed in a Croatian Middle Bronze Age individual bearing Steppe ancestry, suggesting the possibility of genetic input from groups that arrived from the east after the spread of first farmers. This is consistent with the evidence of material culture exchange between Sardinians and mainland Mediterranean groups, although genome-wide analyses find no significant evidence of Steppe ancestry so the quantitative demographic impact was minimal.

Another interesting data, these (Mesolithic) remnant R1b-V88 lineages closely related to the Italian Peninsula, the most likely region of expansion of these lineages into Africa, in turn possibly connected to the expansion of Proto-Afroasiatic.

We detect definitive evidence of Iranian-related ancestry in an Iron Age Sardinian I10366 (391-209 calBCE) with an estimate of 11.9 ± 3.7.% Iran_Ganj_Dareh_Neolithic related ancestry, while rejecting the model with only Anatolian_Neolithic and WHG at p=0.0066 (Supplementary Table 9). The only model that we can fit for this individual using a pair of populations that are closer in time is as a mixture of Iberia_Chalcolithic (11.9 ± 3.2%) and Mycenaean (88.1 ± 3.2%) (p=0.067). This model fits even when including Nuragic Sardinians in the outgroups of the qpAdm analysis, which is consistent with the hypothesis that this individual had little if any ancestry from earlier Sardinians.

yamnaya-samara
Proportions of ancestry using a distal qpAdm framework on an individual basis (a), and based on qpWave clusters

Sicily EBA: The Lusitanian/Ligurian connection?

(…) While a previously reported Bell Beaker culture-associated individual from Sicily had no evidence of Steppe ancestry, (…) we find evidence of Steppe ancestry in the Early Bronze Age by ~2200 BCE. In distal qpAdm, the outlier Sicily_EBA11443 is parsimoniously modeled as harboring 40.2 ± 3.5% Steppe ancestry, and the outlier Sicily_EBA8561 is parsimoniously modeled as harboring 23.3 ± 3.5% Steppe ancestry. (…) The presence of Steppe ancestry in Early Bronze Age Sicily is also evident in Y chromosome analysis, which reveals that 4 of the 5 Early Bronze Age males had Steppe-associated Y-haplogroup R1b1a1a2a1a2. (Online Table 1). Two of these were Y-haplogroup R1b1a1a2a1a2a1 (Z195) which today is largely restricted to Iberia and has been hypothesized to have originated there 2500-2000 BCE. This evidence of west-to-east gene flow from Iberia is also suggested by qpAdm modeling where the only parsimonious proximate source for the Steppe ancestry we found in the main Sicily_EBA cluster is Iberians.

What’s this? An ancestral connection between Sicel Elymian and Galaico-Lusitanian or Ligurian (based on an origin in NE Iberia)? Impossible to say, especially if the languages of these early settlers were replaced later by non-Indo-European speakers from the eastern Mediterranean, and by Indo-European speakers from the mainland closely related to Proto-Italic during the LBA, but see below.

Regarding the comment on R1b-Z195, it is associated with modern Iberians, as DF27 in general, due to founder effects beyond the Pyrenees. It is a very old subclade, split directly from DF27 roughly at the same time as it split from the parent P312, i.e. it can be found anywhere in Europe, and it almost certainly accompanied the expansion of Celts from Central Europe under the subclade R1b-M167/SRY2627.

The connection is thus strong only because of the qpAdm modeling, since R1b-DF27 and subclade R1b-Z195 are certainly lineages expanded quite early, most likely with Yamna settlers in Hungary and East Bell Beakers.

In this case, if stemming from Iberia, it is most likely of subclade R1b-Z220 – or another Z195 (xM167) lineage – originally associated with the Old European substrate found in topo-hydronymy in Iberia, whose most likely remnants attested during the Iron Age were Lusitanians.

r1b-df27-z195
Left: Modern distribution of R1b-Z195 (YFull estimate 2700 BC); Right: Modern distribution of DF27. Both include later founder effects within Iberia, so the increase in the Basque country and the Crown of Aragon and the decrease in Portugal can safely be ignored. Contour maps of the derived allele frequencies of the SNPs analyzed in Solé-Morata et al. (2017).

We detect Iranian-related ancestry in Sicily by the Middle Bronze Age 1800-1500 BCE, consistent with the directional shift of these individuals toward Mycenaeans in PCA. Specifically, two of the Middle Bronze Age individuals can only be fit with models that in addition to Anatolia_Neolithic and WHG, include Iran_Ganj_Dareh_Neolithic. The most parsimonious model for Sicily_MBA3125 has 18.0 ± 3.6% Iranian-related ancestry (p=0.032 for rejecting the alternative model of Steppe rather than Iranian-related ancestry), and the most parsimonious model for Sicily_MBA has 14.9 ± 3.9% Iranian-related ancestry (p=0.037 for rejecting the alternative model).

The modern southern Italian Caucasus-related signal identified in Raveane et al. (2018) is plausibly related to the same Iranian-related spread of ancestry into Sicily that we observe in the Middle Bronze Age (and possibly the Early Bronze Age).

The non-Indo-European Sicanians and Elymians were possibly then connected to eastern Mediterranean groups before the expansion of the Sea Peoples.

For the Late Bronze Age group of individuals, qpAdm documented Steppe-related ancestry, modeling this group as 80.2 ± 1.8% Anatolia_Neolithic, 5.3 ± 1.6% WHG, and 14.5 ± 2.2% Yamnaya_Samara. Our modeling using sources more closely related in space and time also supports Sicily_LBA having Minoan-related ancestry or being derived from local preceding populations or individuals with ancestries similar to those of Sicily_EBA3123 (p=0.527), Sicily_MBA3124 (p=0.352), and Sicily_MBA3125 (p=0.095).

This increase in Steppe-related ancestry in a western site during the LBA most likely represents either an expansion from the Aegean or – maybe more likely, given the archaeological finds – a regional population similar to Sicily EBA re-emerging or rather being displaced from the eastern part of the island because of a westward movement from nearby Calabria.

Whether this population sampled spoke Indo-European or not at this time is questionable, since the Iron Age accounts show non-IE Elymians in this region.

Actually, Elymians seem to have spoken Indo-European, which fits well with the increase in steppe ancestry.

EDIT (21 MAR): Interesting about a proposed incoming Minoan-like ancestry is the potential origin of the Iran Neolithic-related ancestry that is going to appear in Central Italy during the LBA. This could then be potentially associated with Tyrsenians passing through the area, although the traditional description may be more more compatible with an arrival of Sea Peoples from the Adriatic.

Sad to read this:

This manuscript is dedicated to the memory of Sebastiano Tusa of the Soprintendenza del Mare in Palermo, who would have been an author of this study had he not tragically died in the crash of Ethiopia Airlines flight 302 on March 10.

Related

Aquitanians and Iberians of haplogroup R1b are exactly like Indo-Iranians and Balto-Slavs of haplogroup R1a

eba-indo-iranian-balto-slavs

The final paper on Indo-Iranian peoples, by Narasimhan and Patterson (see preprint), is soon to be published, according to the first author’s Twitter account.

One of the interesting details of the development of Bronze Age Iberian ethnolinguistic landscape was the making of Proto-Iberian and Proto-Basque communities, which we already knew were going to show R1b-P312 lineages, a haplogroup clearly associated during the Bell Beaker period with expanding North-West Indo-Europeans:

From the Bronze Age (~2200–900 BCE), we increase the available dataset from 7 to 60 individuals and show how ancestry from the Pontic-Caspian steppe (Steppe ancestry) appeared throughout Iberia in this period, albeit with less impact in the south. The earliest evidence is in 14 individuals dated to ~2500–2000 BCE who coexisted with local people without Steppe ancestry. These groups lived in close proximity and admixed to form the Bronze Age population after 2000 BCE with ~40% ancestry from incoming groups. Y-chromosome turnover was even more pronounced, as the lineages common in Copper Age Iberia (I2, G2, and H) were almost completely replaced by one lineage, R1b-M269.

iberia-admixture-y-dna
Proportion of ancestry derived from central European Beaker/Bronze Age populations in Iberians from the Middle Neolithic to the Iron Age (table S15). Colors indicate the Y-chromosome haplogroup for each male. Red lines represent period of admixture. Modified from Olalde et al. (2019).

The arrival of East Bell Beakers speaking Indo-European languages involved, nevertheless, the survival of the two non-IE communities isolated from each other – likely stemming from south-western France and south-eastern Iberia – thanks to a long-lasting process of migration and admixture. There are some common misconceptions about ancient languages in Iberia which may have caused some wrong interpretations of the data in the paper and elsewhere:

NOTE. A simple reading of Iberian prehistory would be enough to correct these. Two recent books on this subject are Villar’s Indoeuropeos, iberos, vascos y otros parientes and Vascos, celtas e indoeuropeos. Genes y lenguas.

Iberian languages were spoken at least in the Mediterranean and the south (ca. “1/3 of Iberia“) during the Bronze Age.

Nope, we only know the approximate location of Iberian culture and inscriptions from the Late Iron Age, and they occupy the south-eastern and eastern coastal areas, but before that it is unclear where they were spoken. In fact, it seems evident now that the arrival of Urnfield groups from the north marks the arrival of Celtic-speaking peoples, as we can infer from the increase in Central European admixture, while the expansion of anthropomorphic stelae from the north-west must have marked the expansion of Lusitanian.

Vasconic was spoken in both sides of the Pyrenees, as it was in the Middle Ages.

Wrong. One of the worst mistakes I am seeing in many comments since the paper was published, although admittedly the paper goes around this problem talking about “Modern Basques”. Vasconic toponyms appear south of the Pyrenees only after the Roman conquests, and tribes of the south-western Pyrenees and Cantabrian regions were likely Celtic-speaking peoples. Aquitanians (north of the western Pyrenees) are the only known ancient Vasconic-speaking population in proto-historic times, ergo the arrival of Bell Beakers in Iberia was most likely accompanied by Indo-European languages which were later replaced by Celtic expanding from Central Europe, and Iberian expanding from south-east Iberia, and only later with Latin and Vasconic.

Ligurian is non-Indo-European, and Lusitanian is Celtic-like, so Iberia must have been mostly non-Indo-European-speaking.

The fragmentary material available on Ligurian is enough to show that phonetically it is a NWIE dialect of non-Celtic, non-Italic nature, much like Lusitanian; that is, unless you follow laryngeals up to Celtic or Italic, in which case you can argue anything about this or any other IE language, as people who reconstruct laryngeals for Baltic in the common era do.

EDIT (19 Mar 2019): It was not clear enough from this paragraph, because Ligurian-like languages in NE Iberia is just a hypothesis based on the archaeological connection of the whole southern France Bell Beaker region. My aim was to repeat the idea that Old European topo-hydronymy is older in NE Iberia (as almost anywhere in Iberia) than Iberian toponymy, so the initial hypothesis is that:

  1. a Palaeo-European language (as Villar puts it) expanded into most regions of Iberia in ancient times (he considered at some point the Mesolithic, but that is obviously wrong, as we know now); then
  2. Celts expanded at least to the Ebro River Basin; then
  3. Iberians expanded to the north and replaced these in NE Iberia; and only then
  4. after the Roman invasion, around the start of the Common Era, appear Vasconic toponyms south of the Pyrenees.

Lusitanian obviously does not qualify as Celtic, lacking the most essential traits that define Celticness…Unless you define “(Para-)Celtic” as Pre-Proto-Celtic-like, or anything of the sort to support some Atlantic continuity, in which case you can also argue that Pre-Italic or Pre-Germanic are Celtic, because you would be essentially describing North-West Indo-European

If Basques have R1b, it’s because of a culture of “matrilocality” as opposed to the “patrilocality” of Indo-Europeans

So wrong it hurts my eyes every time I read this. Not only does matrilocality in a regional group have few known effects in genetics, but there are many well-documented cases of population replacement (with either ancestry or Y-DNA haplogroups, or both) without language replacement, without a need to resort to “matrilineality” or “matrilocality” or any other cultural difference in any of these cases.

In fact, it seems quite likely now that isolated ancient peoples north of the Pyrenees will show a gradual replacement of surviving I2a lineages by neighbouring R1b, while early Iberian R1b-DF27 lineages are associated with Lusitanians, and later incoming R1b-DF27 lineages (apart from other haplogroups) are most likely associated with incoming Celts, which must have remained in north-central and central-east European groups.

NOTE. Notice how R1a is fully absent from all known early Indo-European peoples to date, whether Iberian IE, British IE, Italic, or Greek. The absence of R1a in Iberia after the arrival of Celts is even more telling of the origin of expanding Celts in Central Europe.

I haven’t had enough time to add Iberian samples to my spreadsheet, and hence neither to the ASoSaH texts nor maps/PCAs (and I don’t plan to, because it’s more efficient for me to add both, Asian and Iberian samples, at the same time), but luckily Maciamo has summed it up on Eupedia. Or, graphically depicted in the paper for the southeast:

iberia-haplogroups
Y chromosome haplogroup composition of individuals from southeast Iberia during the past 2000 years. The general Iberian Bronze and Iron Age population is included for comparison. Modified from Olalde et al. (2019).

Does this continued influx of Y-DNA haplogroups in Iberia with different cultures represent permanent changes in language? Are, therefore, modern Iberian languages derived from Lusitanian, Sorothaptic/Celtic, Greek, Phoenician, East or West Germanic, Hebrew, Berber, or Arabic languages? Obviously not. Same with Italy (see the recent preprint on modern Italians by Raveane et al. 2018), with France, with Germany, or with Greece.

If that happens in European regions with a known ancient history, why would the recent expansions and bottlenecks of R1b in modern Basques (or N1c around the Baltic, or R1a in Slavs) in the Middle Ages represent an ancestral language surviving into modern times?

Indo-Iranians

If something is clear from Narasimhan, Patterson, et al. (2018), is that we know finally the timing of the introduction and expansion of R1a-Z645 lineages among Indo-Iranians.

We could already propose since 2015 that a slow admixture happened in the steppes, based on archaeological finds, due to settlement elites dominating over common peoples, coupled with the known Uralic linguistic traits of Indo-Iranian (and known Indo-Iranian influence on Finno-Ugric) – as I did in the first version of the Indo-European demic diffusion model.

The new huge sampling of Sintashta – combined with that of Catacomb, Poltavka, Potapovka, Andronovo, and Srubna – shows quite clearly how this long-term admixture process between Uralic peoples and Indo-Iranians happened between forest-steppe CWC (mainly Abashevo) and steppe groups. The situation is not different from that of Iberia ca. 2500-2000 BC; from Narasimhan, Patterson, et al. (2018):

We combined the newly reported data from Kamennyi Ambar 5 with previously reported data from the Sintashta 5 individuals (10). We observed a main cluster of Sintashta individuals that was similar to Srubnaya, Potapovka, and Andronovo in being well modeled as a mixture of Yamnaya-related and Anatolian Neolithic (European agriculturalist-related) ancestry.

Even with such few words referring to one of the most important data in the paper about what happened in the steppes, Wang et al. (2018) help us understand what really happened with this simplistic concept of “steppe ancestry” regarding Yamna vs. Corded Ware differences:

anatolia-neolithic-steppe-eneolithic
Image modified from Wang et al. (2018). Marked are: in red, approximate limit of Anatolia_Neolithic ancestry found in Yamna populations; in blue, Corded Ware-related groups. “Modelling results for the Steppe and Caucasus 1128 cluster. Admixture proportions based on (temporally and geographically) distal and proximal models, showing additional Anatolian farmer-related ancestry in Steppe groups as well as additional gene flow from the south in some of the Steppe groups as well as the Caucasus groups (see also Supplementary Tables 10, 14 and 20).”

As with Iberia (or any prehistoric region), the details of how exactly this language change happened are not evident, but we only need a plausible explanation coupled with archaeology and linguistics. Poltavka, Potapovka, and Sintashta samples – like the few available Iberian ones ca. 2500-2000 BC – offer a good picture of the cohabitation of R1b-L23 (mainly Z2103) and R1a-Z645 (mainly Z93+): a glimpse at the likely presence of R1a-Z93 within settlements – which must have evolved as the dominant elites – in a society where the majority of the population was initially formed by nomad herders (probably most R1b-Z2103), who were usually buried outside of the main settlements.

Will the upcoming Narasimhan, Patterson et al. (2019) deal with this problem of how R1a-M417 replaced R1b-M269, and how the so-called “Steppe_MLBA” (i.e. Corded Ware) ancestry admixed with “Steppe_EMBA” (i.e. Yamnaya) ancestry in the steppes, and which one of their languages survived in the region (that is, the same the Reich Lab has done with Iberia)? Not likely. The ‘genetic wars’ in Iberia deal with haplogroup R1b-P312, and how it was neither ‘native’ nor associated with Basques and non-Indo-European peoples in general. The ‘genetic wars’ in South Asia are concerned with the steppe origin of R1a, to prove that it is not a ‘native’ haplogroup to India, and thus neither are Indo-Aryan languages. To each region a politically correct account of genetic finds, with enough care not to fully dismiss national myths, it seems.

NOTE. Funnily enough, these ‘genetic wars’ are the making of geneticists since the 1990s and 2000s, so we are still in the midst of mostly internal wars caused by what they write. Just as genetic papers of the 2020s will most likely be a reaction to what they are writing right now about “steppe ancestry” and R1a. You won’t find much change to the linguistic reconstruction in this whole period, except for the most multicolored glottochronological proposals…

The first author of the paper has engaged, as far as I could see in Twitter, in dialogue with Hindu nationalists who try to dismiss the arrival of steppe ancestry and R1a into South Asia as inconclusive (to support the potential origin of Sanskrit millennia ago in the Indus Valley Civilization). How can geneticists deal with the real problem here (the original ethnolinguistic group expanding with Corded Ware), when they have to fend off anti-steppists from Europe and Asia? How can they do it, when they themselves are part of the same societies that demand a politically correct presentation of data?

This is how the data on the most likely Indo-Iranian-speaking region should be presented in an ideal world, where – as in the Iberia paper – geneticists would look closely to the Volga-Ural region to discover what happened with Proto-Indo-Iranians from their earliest to their latest stage, instead of constantly looking for sites close to the Indus Valley to demonstrate who knows what about modern Indian culture:

indo-iranian-admixture-similar-iberians
Tentative map of the Late PIE and Indo-Iranian community in the Volga-Ural steppes since the Eneolithic. Proportion of ancestry derived from central European Corded Ware peoples. Colors indicate the Y-chromosome haplogroup for each male. Red lines represent period of admixture. Modified from Olalde et al. (2019).

Now try and tell Hindu nationalists that Sanskrit expanded from an Early Bronze Age steppe community of R1b-rich nomadic herders that spoke Pre-Indo-Iranian, which was dominated and eventually (genetically) mostly replaced by elite Uralic-speaking R1a peoples from the Russian forest, hence the known phonetic (and some morphological) traits that remained. Good luck with the Europhobic shitstorm ahead..

Balto-Slavic

Iberian cultures, already with a majority of R1b lineages, show a clear northward expansion over previously Urnfield-like groups of north-east Iberia and Mediterranean France (which we now know probably represent the migration of Celts from central Europe). Similarly, Eastern Balts already under a majority of R1a lineages expanded likely into the Baltic region at the same time as the outlier from Turlojiškė (ca. 1075 BC), which represents the first obvious contacts of central-east Europe with the Baltic.

Iberia shows a more recent influx of central and eastern Mediterranean peoples, one of which eventually succeeded in imposing their language in Western Europe: Romans were possibly associated mainly with R1b-U152, apart from many other lineages. Proto-Slavs probably expanded later than Celts, too, connected to the disintegration of the Lusatian culture, and they were at some point associated with R1a-M458 and R1a-Z280(xZ92) lineages, apart from others already found in Early Slavs.

pca-balto-slavs-tollense-valley
PCA of central-eastern European groups which may have formed the Balto-Slavic-speaking community derived from Bell Beaker, evident from the position ‘westwards’ of CWC in the PCA, and surrounding cultures. Left: Early Bronze Age. Right: Tollense Valley samples.

This parallel between Iberia and eastern Europe is no coincidence: as Europe entered the Bronze Age, chiefdom-based systems became common, and thus the connection of ancestry or haplogroups with ethnolinguistic groups became weaker.

What happened earlier (and who may represent the Pre-Balto-Slavic community) will be clearer when we have enough eastern European samples, but basically we will be able to depict this admixture of NWIE-speaking BBC-derived peoples with Uralic-speaking CWC-derived groups (since Uralic is known to have strongly influenced Balto-Slavic), similar to the admixture found in Indo-Iranians, more or less like this:

iberian-admixture-balto-slavic
Tentative map of the North-West Indo-European and Balto-Slavic community in central-eastern Europe since the East Bell Beaker expansion. Proportion of ancestry derived from Corded Ware peoples. Colors indicate the Y-chromosome haplogroup for each male. Red lines represent period of admixture. Modified from Olalde et al. (2019).

The Early Scythian period marked a still stronger chiefdom-based system which promoted the creation of alliances and federation-like groups, with an earlier representation of the system expanding from north-eastern Europe around the Baltic Sea, precisely during the spread of Akozino warrior-traders (in turn related to the Scythian influence in the forest-steppes), who are the most likely ancestors of most N1c-V29 lineages among modern Germanic, Balto-Slavic, and Volga-Finnic peoples.

Modern haplogroup+language = ancient ones?

It is not difficult to realize, then, that the complex modern genetic picture in Eastern Europe and around the Urals, and also in South Asia (like that of the Aegean or Anatolia) is similar to the Iron Age / medieval Iberian one, and that following modern R1a as an Indo-European marker just because some modern Indo-European-speaking groups showed it was always a flawed methodology; as flawed as following R1b for ancient Vasconic groups, or N1c for ancient Uralic groups.

Why people would argue that haplogroups mean continuity (e.g. R1b with Basques, N1c with Finns, R1a with Slavs, etc.) may be understood, if one lives still in the 2000s. Just like why one would argue that Corded Ware is Indo-European, because of Gimbutas’ huge influence since the 1960s with her myth of “Kurgan peoples”. Not many denied these haplogroup associations, because there was no reason to do it, and those who did usually aligned with a defense of descriptive archaeology.

However, it is a growing paradox that some people interested in genetics today would now, after the Iberian paper, need to:

  • accept that ancient Iberians and probably Aquitanians (each from different regions, and probably from different “Basque-Iberian dialects” in the Chalcolithic, if both were actually related) show eventually expansions with R1b-L23, the haplogroup most obviously associated with expanding Indo-Europeans;
  • acknowledge that modern Iberians have many different lineages derived from prehistoric or historic peoples (Celts, Phoenicians, Greeks, Romans, Jews, Goths, Berbers, Arabs), which have undergone different bottlenecks, the last ones during the Reconquista, but none of their languages have survived;
  • realize that a similar picture is to be found everywhere in central and western Europe since the first proto-historic records, with language replacement in spite of genetic continuity, such as the British Isles (and R1b-L21 continuity) after the arrival of Celts, Romans, Anglo-Saxons, Vikings, or Normans;
  • but, at the same time, continue blindly asserting that haplogroup R1a + “steppe ancestry” represent some kind of supernatural combination which must show continuity with their modern Indo-Iranian or Balto-Slavic language from time immemorial.
sintashta-y-dna
Replacement of R1b-L23 lineages during the Early Bronze Age in eastern Europe and in the Eurasian steppes: emergence of R1a in previous Yamnaya and Bell Beaker territories. Modified from EBA Y-DNA map.

Behave, pretty please

The ‘conservative’ message espoused by some geneticists and amateur genealogists here is basically as follows:

  • Let’s not rush to new theories that contradict the 2000s, lest some people get offended by granddaddy not being these pure whatever wherever as they believed, and let’s wait some 5, 10, or 20 years, as long as necessary – to see if some corner of the Yamna culture shows R1a, or some region in north-eastern Europe shows N1c, or some Atlantic Chalcolithic sample shows R1b – to challenge our preferred theories, if we actually need to challenge anything at all, because it hurts too much.
  • Just don’t let many of these genetic genealogists or academics of our time be unhappy, pretty please with sugar on top, and let them slowly adapt to reality with more and more pet theories to fit everything together (past theories + present data), so maybe when all of them are gone, within 50 or 70 years, society can smoothly begin to move on and propose something closer to reality, but always as politically correct as possible for the next generations.
  • For starters, let’s discuss now (yet again) that Bell Beakers may not have been Indo-European at all, despite showing (unlike Corded Ware) clearly Yamna male lineages and ancestry, because then Corded Ware and R1a could not have been Indo-European and that’s terrible, so maybe Bell Beakers are too brachycephalic to speak Indo-European or something, or they were stopped by the Fearsome Tisza River, or they are not pure Dutch Single Grave in The South hence not Indo-European, or whatever, and that’s why Iron Age Iberians or Etruscans show non-Indo-European languages. That’s not disrespectful to the history of certain peoples, of course not, but talking about the evident R1a-Uralic connection is, because this is The South, not The North, and respect works differently there.
  • Just don’t talk about how Slavs and Balts enter history more than 1,500 years later than Indo-European peoples in Western and Southern Europe, including Iberia, and assume a heroic continuity of Balts and Slavs as pure R1a ‘steppe-like’ peoples dominating over thousands of kms. in the Baltic, Fennoscandia, eastern Europe, and northern Asia for 5,000 years, with multiple Balto-Slavs-over-Balto-Slavs migrations, because these absolute units of Indo-European peoples were a trip and a half. They are the Asterix and Obelix of white Indo-European prehistory.
  • Perhaps in the meantime we can also invent some new glottochronological dialectal scheme that fits the expansion of Sredni Stog/Corded Ware with (Germano-?)Indo-Slavonic separated earlier than any other Late PIE dialect; and Finno-Volgaic later than any other Uralic dialect, in the Middle Ages, with N1c.
balto-slavic-pca
Genetic structure of the Balto-Slavic populations within a European context according to the three genetic systems, from Kushniarevich et al. (2015). Pure Balto-Slavs from…hmm…yeah this…ancient…region…or people…cluster…Whatever, very very steppe-like peoples, the True Indo-Europeans™, so close to Yamna…almost as close as Finno-Ugrians.

To sum up: Iberia, Italy, France, the British Isles, central Europe, the Balkans, the Aegean, or Anatolia, all these territories can have a complex history of periodic admixture and language replacement everywhere, but some peoples appearing later than all others in the historical record (viz. Basques or Slavs) apparently cannot, because that would be shameful for their national or ethnic myths, and these should be respected.

Ignorance of the own past as a blank canvas to be filled in with stupid ethnolinguistic continuity, turned into something valuable that should not be challenged. Ethnonationalist-like reasoning proper of the 19th century. How can our times be called ‘modern’ when this kind of magical thinking is still prevalent, even among supposedly well-educated people?

Related

Haplogroup R1b-M167/SRY2627 linked to Celts expanding with the Urnfield culture

bronze-age-late-urnfield

As you can see from my interest in the recently published Olalde et al. (2019) Iberia paper, once you accept that East Bell Beakers expanded North-West Indo-European, the most important question becomes how did its known dialects spread to their known historic areas.

We already had a good idea about the expansion of Celts, based on proto-historical accounts, fragmentary languages, and linguistic guesstimates, but the connection of Celtic with either Urnfield or slightly later Hallstatt/La Tène was always blurred, due to the lack of precise data on population movements.

The latest paper on Iberia is interesting for many details, such as:

  • The express dismissal of the newest pet theory based on the simplistic “steppe ancestry = IE”: the obsessive comparisons of Dutch Bell Beakers as the origin of basically anything that moves in Europe.
  • A discrete influx of North African ancestry in certain samples before the Moorish invasion (which was probably mediated by peoples of North African rather than Levantine admixture).
  • The finding of very Mycenaean-like Greek colonies of the 5th century (interestingly, under R1b lineages).
iberia-celts-romans
Modified from section of PCA of ancient samples by Olalde et al. (2019). “IE Iberia” refers to Pre-Celtic Indo-European languages of Iberia, such as Galaico-Lusitanian in the west (see more on Lusitanian), and a potentially Ligurian-related language in the North-East and southern France.

The paper is, however, of particular importance from the perspective of historical linguistics. It confirms that:

  • Celtic-speaking peoples expanded in Iberia likely during the Late Bronze Age – Early Iron Age (probably with the Urnfield culture, before 1000 BC) with North/Central European ancestry.

NOTE. The paper marks what are believed to be the boundaries of non-Indo-European languages during the Iron Age in later times, extrapolating that situation to the past. Mediterranean sites with Iberian traits (ca. 6th century on) were probably non-Indo-European-speaking tribes, but it is unclear what happened in the centuries before their sampling, and there are no clear boundaries. These incoming Celts from central Europe with the Urnfield culture makes it very likely that the Iberian expansion to the north happened later, incorporating thus this central European ancestry in the process. The southern (orientalizing, Tartessian) site of La Angorrilla shows incineration and influence from Phoenician settlers, and their actual language is also far from clear. The other investigated samples, with higher central European contribution, are from Celtiberian sites.

  • The slightly later arrival of (Phoenician, Greek and) Latin-speaking peoples into Iberia is marked by Central/Eastern Mediterranean and North African ancestry.
iberia-migrations-celts-romans
Expansion of different ancestry components in Iberia during Prehistory. Modified from Olalde et al. (2019) to include labels with populations expanding with each component.

While both confirm what was more or less already known about the oldest attested NWIE dialects, and further support the role of East Bell Beakers in expanding North-West Indo-European, the first part is interesting for two main reasons:

  1. Koch’s Celtic from the West hypothesis, which made a recent comeback with a renewed model based on “steppe ancestry”, is once again rejected in population genomics, as expected. At this point I doubt this will mean anything to the supporters of the theory (because you can propose as many “Celtic-over-Celtic” layers as you want), but if you are not obsessed with autochthonous continuity of Celtic languages in the Atlantic area we might begin to judge the most correct dialectal split (and thus classification) among those proposed to date, based on ancestry and haplogroup expansions.
  2. We believed in the 2000s that the expansion of haplogroup R1b-M167 (TMRCA ca. 1100 BC for YTree or 1700 BC for YFull) was coupled with the expansion of Iberians from the Pyrenees, in turn (thus) closely related to Basques. This non-IE presence has been contested with toponymic data in linguistics, and with the testing of many modern samples and the subsequent discovery of the widespread distribution of the subclade in western and northern Europe. Now it has become even more likely (lacking confirmation with aDNA) that this haplogroup expanded with Celts.

NOTE. Regarding R1b SNPs, YTree has more samples (and thus more SNPs) to work with estimates, due to its connection with FTDNA groups, so it is in principle more reliable (although estimates were calculated in 2017). Nevertheless, the methods to estimate the age of the MRCA are different between YTree and YFull.

df27-m167-z262-mcdonald
YTree estimations of TMRCA for R1b-Z262 (left) and R1b-M167 (right).

Why this is important has to do with the realization that Celts must have expanded explosively in all directions during the estimated range for Common Celtic (ca. 1500-1000 BC), and as such R1b-M167 is probably going to be one of the clear Y-DNA markers of the Celtic expansion, when it appears in the ancient DNA record, maybe in new SNP calls from samples of the Olalde et al. (2019) paper, or in future Urnfield/Hallstatt/La Tène papers.

Sister clades derived from R1b-Z262 (TMRCA ca. 1650 BC for YTree, or 2700 for YFull), although sharing a quite old origin, may have taken part in the same communities that expanded R1b-M167, likely from some point in central Europe, possibly as remnants of a previous (Tumulus culture?) central European expansion, as the sample SZ5 from Szólád (R1b-CTS1595) and the distribution of modern samples suggest.

r1b-df27-m167-sry2627
Left: Modern distribution of upstream clade L176.2 (YFull R1b-CTS4188); Right: Modern distribution of M167. Both include later expansions within Iberia (probably with the Crown of Aragon during the Reconquista). Contour maps of the derived allele frequencies of the SNPs analyzed in Solé-Morata et al. (2017).

The Celtic expansion might not have been a mass migration of peoples replacing all male lines of their controlled territories (as was common in the Neolithic and Chalcolithic), because of the Bronze Age dominant chiefdom-based system that relied on alliances, but it is becoming clear that Early Celts are also going to show the expansion of certain successful male lineages.

Oh, and you can say goodbye to the autochthonous “Vasconic = R1b-DF27” (latest heir of the “Vasconic = R1b-P312”) theory, too, if – for some strange reason – you hadn’t already.

EDIT (16 MAR) Just in case the wording is not clear: the fact that this haplogroup most likely expanded with Celts does not mean that its lineages didn’t become eventually incorporated into Iberian cultures and adopted non-IE languages: some of them probably did at some point, in some regions of northern Iberia, and most were certainly later incorporated to the Roman civilization and spoke Latin, then to the medieval kingdoms with their languages, and so on until the present day… Only those eventually associated with Iron Age Aquitanians may have retained their non-IE language, unless those lineages today associated with Basques were incorporated later to the Basque-speaking regions by expanding medieval kingdoms. A complex picture repeated everywhere in Europe: no haplogroup+language continuity in sight, anywhere.

NOTE: This here is currently the most likely interpretation of data based on estimations of mutations; it is not confirmed with ancient samples.

Related

Updates to ASoSaH: new maps, updated PCA, and added newest research papers

steppe-ancestry-cut

The title says it all. I have used some free time to update the series A Song of Sheep and Horses:

I basically added information from the latest papers published, which (luckily enough for me) haven’t been too many, and I have added images to illustrate certain sections.

I have updated the PCAs by including North Caucasus samples from Wang et al. (2018), whose position I could only infer for older versions from previously published PCA graphs.

pca-steppe-eneolithic-early
PCA of ancient and modern Eurasian samples. Early Eneolithic admixture events in the steppe drawn.

I have also added to the supplementary materials the “Tip of the Iceberg” R1b tree by Mike Walsh from the FTDNA R1b group, with permission, because some relevant genetic sections are centered on the evolution of R1b lineages, and the reader can get easily lost with so many subclades.

I have also updated maps, including some of the Y-DNA ones, and managed to finish two new maps I was working on, and I added them to the supplementary materials and to the menu above:

One on Yamna kurgans in Hungary, coupled with contemporaneous sites of Baden-Boleráz or Kostolac cultures:

burials-yamnaya-hungary
Map of attested Yamnaya pit-grave burials in the Hungarian plains; superimposed in shades of blue are common areas covered by floods before the extensive controls imposed in the 19th century; in orange, cumulative thickness of sand, unfavourable loamy sand layer. Marked are settlements/findings of Boleráz (ca. 3500 BC on), Baden (until ca. 2800 BC), Kostolac (precise dates unknown), and Yamna kurgans (from ca. 3100/3000 BC on).

Another one on Steppe ancestry expansion, with a tentative distribution of “steppe ancestry” divided into that of Sredni Stog/Corded Ware origin vs. that of Repin/Yamna origin, a difference that has been known for quite some time already.

It is tentative because there hasn’t been any professional study or amateur attempt to date to differentiate both “steppe ancestries” in Yamna, and especially in Bell Beakers. So much for the call of professional geneticists since 2018 (see here and here) and archaeologists since 2017 (see e.g. here and here) to distinguish fine-scale population structure to be able to follow neighbouring populations which expanded with different archaeological (and thus ethnolinguistic) groups.

steppe-ancestry-corded-ware
Tentative map of fine-scale population structure during steppe-related expansions (ca. 3500–2000 BC), including Repin–Yamna–Bell Beaker/Balkans and Sredni Stog–Corded Ware groups. Data based on published samples and pairwise comparisons tested to date. Notice that the potential admixture of expanding Repin/Early Yamna settlers in the North Pontic area with the late Sredni Stog population (and thus Sredni Stog-related ancestry in Yamna) has been omitted for simplicity purposes, assuming thus a homogeneous Yamna vs. Corded Ware ancestry.

I think both maps are especially important today, given the current Nordicist reactionary trends arguing (yet again) for an origin of Indo-Europeans in The North™, now based on the Fearsome Tisza River hypothesis, on cephalic index values, and a few pairwise comparisons – i.e. an absolutely no-nonsense approach to the Indo-European question (LOL). At least I get to relax and sit this year out just observing how other people bury themselves and their beloved “steppe ancestry=IE” under so many new pet theories…

NOTE. Not that there is anything wrong with a northern origin of North-West Indo-European from a linguistic point of view, as I commented recently – after all, a Corded Ware origin would roughly fit the linguistic guesstimates, unlike the proposed ancestral origins in Anatolia or India. The problem is that, like many other fringe theories, it is today just based on tradition, or (even worse) ethnic, political, or personal desires, and it doesn’t make sense when all findings from disciplines involved in the Indo-European and Uralic questions are combined.

steppe-ancestry-modern-populations
Simple ancestry percentages in modern populations. Recent image by Iain Mathieson 2019 (min. 5.57). A simplistic “Steppe ancestry” defining Indo-European speakers…? Sure.

Within 20 or 30 years, when genetic genealogists (or amateur geneticists, or however you want to call them) ask why we had the opportunity since 2015 to sample as many Hungarian Yamnaya individuals as possible and we didn’t, when it is clear that the number of unscathed kurgans is diminishing every year (from an estimated 4,000 in the 20th century, of the original tens of thousands, to less than 1,500 today) the answer will not be “because this or that archaeologist or linguist was a dilettante or a charlatan‘, as they usually describe academics they dislike.

It will be precisely because the very same genetic genealogists – supposedly interested today in the origin of R1b-L151 and/or genetic marker associated with North-West Indo-Europeans – are obsessed with finding them anywhere else but for Hungary, and prefer to use their money and time to play with a few statistical tools within a biased framework of flawed assumptions and study designs, obtaining absurd results and accepting far-fetched interpretations of them, to be told exactly what they want to hear: be it the Franco-Cantabrian homeland, the Dutch or Moravian Beaker from CWC homeland, the Maykop homeland, or the Moon homeland.

Poetic justice this heritage destruction, whose indirect causes will remain written in Internet archives for everyone to see, as a good lesson for future generations.

A very “Yamnaya-like” East Bell Beaker from France, probably R1b-L151

bell-beaker-expansion

Interesting report by Bernard Sécher on Anthrogenica, about the Ph.D. thesis of Samantha Brunel from Institut Jacques Monod, Paris, Paléogénomique des dynamiques des populations humaines sur le territoire Français entre 7000 et 2000 (2018).

NOTE. You can visit Bernard Sécher’s blog on genetic genealogy.

A summary from user Jool, who was there, translated into English by Sécher (slight changes to translation, and emphasis mine):

They have a good hundred samples from the North, Alsace and the Mediterranean coast, from the Mesolithic to the Iron Age.

There is no major surprise compared to the rest of Europe. On the PCA plot, the Mesolithic are with the WHG, the early Neolithics with the first farmers close to the Anatolians. Then there is a small resurgence of hunter-gatherers that moves the Middle Neolithics a little closer to the WHGs.

From the Bronze Age, they have 5 samples with autosomal DNA, all in Bell Beaker archaeological context, which are very spread on the PCA. A sample very high, close to the Yamnaya, a little above the Corded Ware, two samples right in the Central European Bell Beakers, a fairly low just above the Neolithic package, and one last full in the package. The most salient point was that the Y chromosomes of their 12 Bronze Age samples (all Bell Beakers) are all R1b, whereas there was no R1b in the Neolithic samples.

Finally they have samples of the Iron Age that are collected on the PCA plot close to the Bronze Age samples. They could not determine if there is continuity with the Bronze Age, or a partial replacement by a genetically close population.

PCA-caucasus-yamna
Image modified from Wang et al. (2018). Samples projected in PCA of 84 modern-day West Eurasian populations (open symbols). Previously known clusters have been marked and referenced. Marked and labelled are interesting samples; In red, likely position of late Yamna Hungary / early East Bell Beakers An EHG and a Caucasus ‘clouds’ have been drawn, leaving Pontic-Caspian steppe and derived groups between them. See the original file here. To understand the drawn potential Caucasus Mesolithic cluster, see above the PCA from Lazaridis et al. (2018).

The sample with likely high “steppe ancestry“, clustering closely to Yamna (more than Corded Ware samples) is then probably an early East Bell Beaker individual, probably from Alsace, or maybe close to the Rhine Delta in the north, rather than from the south, since we already have samples from southern France from Olalde et al. (2018) with high Neolithic ancestry, and samples from the Rhine with elevated steppe ancestry, but not that much.

This specific sample, if confirmed as one of those reported as R1b (then likely R1b-L151), as it seems from the wording of the summary, is key because it would finally link Yamna to East Bell Beaker through Yamna Hungary, all of them very “Yamnaya-like”, and therefore R1b-L151 (hence also R1b-L51) directly to the steppe, and not only to the Carpathian Basin (that is, until we have samples from late Repin or West Yamna…)

NOTE. The only alternative explanation for such elevated steppe ancestry would be an admixture between a ‘less Yamnaya-like’ East Bell Beaker + a Central European Corded Ware sample like the Esperstedt outlier + drift, but I don’t think that alternative is the best explanation of its position in the PCA closer to Yamna in any of the infinite parallel universes, so… Also, the sample from Esperstedt is clearly a late outlier likely influenced by Yamna vanguard settlers from Hungary, not the other way round…

Unexpectedly, then, fully Yamnaya-like individuals are found not only in Yamna Hungary ca. 3000-2500 BC, but also among expanding East Bell Beakers later than 2500 BC. This leaves us with unexplained, not-at-all-Yamnaya-like early Corded Ware samples from ca. 2900 BC on. An explanation based on admixture with locals seems unlikely, seeing how Corded Ware peoples continue a north Pontic cluster, being thus different from Yamna and their ancestors since the Neolithic; and how they remained that way for a long time, up to Sintashta, Srubna, Andronovo, and even later samples… A different, non-Indo-European community it is, then.

olalde_pca2
Image modified from Olalde et al. (2018). PCA of 999 Eurasian individuals. Marked is the Espersted Outlier with the approximate position of Yamna Hungary, probably the source of its admixture. Different Bell Beaker clines have been drawn, to represent approximate source of expansions from Central European sources into the different regions. In red, likely zone of Yamna Hungary and reported early East Bell Beaker individual from France.

Let’s wait and see the Ph.D. thesis, when it’s published, and keep observing in the meantime the absurd reactions of denial, anger, bargaining, and depression (stages of grief) among BBC/R1b=Vasconic and CWC/R1a=Indo-European fans, as if they had lost something (?). Maybe one of these reactions is actually the key to changing reality and going back to the 2000s, who knows…

Featured image: initial expansion of the East Bell Beaker Group, by Volker Heyd (2013).

Related

R1a-Z280 lineages in Srubna; and first Palaeo-Balkan R1b-Z2103?

herodotus-world-map

Scythian samples from the North Pontic area are far more complex than what could be seen at first glance. From the new Y-SNP calls we have now thanks to the publications at Molgen (see the spreadsheet) and in Anthrogenica threads, I think this is the basis to work with:

NOTE. I understand that writing a paper requires a lot of work, and probably statistical methods are the main interest of authors, editors, and reviewers. But it is difficult to comprehend how any user of open source tools can instantly offer a more complex assessment of the samples’ Y-SNP calls than professionals working on these samples for months. I think that, by now, it should be clear to everyone that Y-DNA is often as important (sometimes even more) than statistical tools to infer certain population movements, since admixture can change within few generations of male-biased migrations, whereas haplogroups can’t…

Srubna

Srubna-Andronovo samples are as homogeneous as they always were, dominated by R1a-Z645 subclades and CWC-related (steppe_MLBA) ancestry.

The appearance of one (possibly two) R-Z280 lineages in this mixed Srubna-Alakul region of the southern Urals and this early (1880-1690 BC, hence rather Pokrovka-Alakul) points to the admixture of R1a-Z93 and R1a-Z280 already in Abashevo, which also explains the wide distribution of both subclades in the forest zones of Central Asia.

If Abashevo is the cornerstone of the Indo-Iranian / Uralic community, as it seems, the genetic admixture would initially be quite similar, undergoing in the steppes a reduction to haplogroup R1a-Z93 (obviously not complete), at the same time as it expanded to the west with Pokrovka and Srubna, and to the east with Petrovka and Andronovo. To the north, similar reductions will probably be seen following the Seima-Turbino phenomenon.

NOTE. Another R1a-Z280 has been found in the recent sample from Bronze Age Poland (see spreadsheet). As it appears right now in ancient and modern DNA, there seems to be a different distribution between subclades:

  • R1a-Z280 (formed ca. 2900 BC, TMRCA ca. 2600 BC) appears mainly distributed today to the east, in the forest and steppe regions, with the most ‘successful’ expansions possibly related to the spread of Abashevo- and Battle Axe-related cultures (Indo-Iranian and Uralic alike).
  • R1a-M458 (formed ca. 2700, TMRCA ca. 2700 BC) appears mainly distributed to the north, from central Europe to the east – but not in the steppe in aDNA, with the most ‘successful’ expansions to the west.

M458 lineages seem thus to have expanded in the steppe in sizeable numbers only after the Iranian expansions (see a map of modern R1a distributions) i.e. possibly with the expansion of Slavs, which supports the model whereby cultures from central-east Europe (like Trzciniec and Lusatian), accompanied mainly by M458 lineages, were responsible for the expansion of Proto-Balto-Slavic (and later Proto-Slavic).

The finding of haplogroup R1a-Z93, among them one Z2123, is no surprise at this point after other similar Srubna samples. As I said, the early Srubna expansion is most likely responsible for the Szólád Bronze Age sample (ca. 2100-1700 BC), and for the Balkans BA sample (ca. 1750-1625 BC) from Merichleri, due to incursions along the central-east European steppe.

cheek-pieces
Map of decorated bone/antler bridle cheek-pieces and whip handle equivalents. They are often local translations that remained faithful to the originals (from data in Piggott, 1965; Kristiansen & Larsson, 2005; David, 2007). Image from Vandkilde (2014).

Cimmerians

Cimmerian samples from the west show signs of continuity with R1a-Z93 lineages. Nevertheless, the sample of haplogroup Q1a-Y558, together with the ‘Pre-Scythian’ sample of haplogroup N (of the Mezőcsát Culture) in Hungary ca. 980-830 BC, as well as their PCA, seem to depict an origin of these Pre-Scythian peoples in populations related to the eastern Central Asian steppes, too.

NOTE. I will write more on different movements (unrelated to Uralic expansions) from Central and East Asia to the west accompanied by Siberian ancestry and haplogroup N with the post of Ugric-Samoyedic expansions.

Scythians

The Scythian of Z2123 lineage ca. 375-203 BC from the Volga (in Mathieson et al. 2015), together with the sample scy193 from Glinoe (probably also R1a-Z2123), without a date, as well as their common Steppe_MLBA cluster, suggest that Scythians, too, were at first probably quite homogeneous as is common among pastoralist nomads, and came thus from the Central Asian steppes.

The reduction in haplogroup variability among East Iranian peoples seems supported by the three new Late Sarmatian samples of haplogroup R1a-Z2124.

Approximate location of Glinoe and Glinoe Sad (with Starosilya to the south, in Ukrainian territory):

This initial expansion of Scythians does not mean that one can dismiss the western samples as non-Scythians, though, because ‘Scythian’ is a cultural attribution, based on materials. Confirming the diversity among western Scythians, a session at the recent ISBA 8:

Genetic continuity in the western Eurasian Steppe broken not due to Scythian dominance, but rather at the transition to the Chernyakhov culture (Ostrogoths), by Järve et al.

The long-held archaeological view sees the Early Iron Age nomadic Scythians expanding west from their Altai region homeland across the Eurasian Steppe until they reached the Ponto-Caspian region north of the Black and Caspian Seas by around 2,900 BP. However, the migration theory has not found support from ancient DNA evidence, and it is still unclear how much of the Scythian dominance in the Eurasian Steppe was due to movements of people and how much reflected cultural diffusion and elite dominance. We present new whole-genome results of 31 ancient Western and Eastern Scythians as well as samples pre- and postdating them that allow us to set the Scythians in a temporal context by comparing the Western Scythians to samples before and after within the Ponto-Caspian region. We detect no significant contribution of the Scythians to the Early Iron Age Ponto-Caspian gene pool, inferring instead a genetic continuity in the western Eurasian Steppe that persisted from at least 4,800–4,400 cal BP to 2,700–2,100 cal BP (based on our radiocarbon dated samples), i.e. from the Yamnaya through the Scythian period.

(…) Our results (…) support the hypothesis that the Scythian dominance was cultural rather than achieved through population replacement.

Detail of the slide with admixture of Scythian groups in Ukraine:

scythians-admixture

The findings of those 31 samples seem to support what Krzewińska et al. (2018) found in a tiny region of Moldavia-south-western Ukraine (Glinoi, Glinoi Sad, and Starosilya).

The question, then, is as follows: if Scythian dominance was “cultural rather than achieved through population replacement”…Where are the R1b-Z2103 from? One possibility, as I said in the previous post, is that they represent pockets of Iranian R1b lineages in the steppes descended from eastern Yamna, given that this haplogroup appears in modern populations from a wide region surrounding the steppes.

The other possibility, which is what some have proposed since the publication of the paper, is that they are related to Thracians, and thus to Palaeo-Balkan populations. About the previously published Thracian individuals in Sikora et al. (2014):

thracian-samples
Geographic origin of ancient samples and ADMIXTURE results. (A) Map of Europe indicating the discovery sites for each of the ancient samples used in this study. (B) Ancestral population clusters inferred using ADMIXTURE on the HGDP dataset, for k = 6 ancestral clusters. The width of the bars of the ancient samples was increased to aid visualization. https://doi.org/10.1371/journal.pgen.1004353.g001

For the Thracian individuals from Bulgaria, no clear pattern emerges. While P192-1 still shows the highest proportion of Sardinian ancestry, K8 more resembles the HG individuals, with a high fraction of Russian ancestry.

Despite their different geographic origins, both the Swedish farmer gok4 and the Thracian P192-1 closely resemble the Iceman in their relationship with Sardinians, making it unlikely that all three individuals were recent migrants from Sardinia. Furthermore, P192-1 is an Iron Age individual from well after the arrival of the first farmers in Southeastern Europe (more than 2,000 years after the Iceman and gok4), perhaps indicating genetic continuity with the early farmers in this region. The only non-HG individual not following this pattern is K8 from Bulgaria. Interestingly, this individual was excavated from an aristocratic inhumation burial containing rich grave goods, indicating a high social standing, as opposed to the other individual, who was found in a pit.

pca-thracians

The following are excerpts from A Companion to Ancient Thrace (2015), by Valeva, Nankov, and Graninger (emphasis mine):

Thracian settlements from the 6th c. BC on:

(…) urban centers were established in northeastern Thrace, whose development was linked to the growth of road and communication networks along with related economic and distributive functions. The early establishment of markets/emporia along the Danube took place toward the middle of the first millennium BCE (Irimia 2006, 250–253; Stoyanov in press). The abundant data for intensive trade discovered at the Getic village in Satu Nou on the right bank of the Danube provides another example of an emporion that developed along the main artery of communication toward the interior of Thrace (Conovici 2000, 75–76).

Undoubtedly the most prominent manifestation of centralization processes and stratification in the settlement system of Thrace arrives with the emergence of political capitals – the leading urban centers of various Thracian political formations.

getic-thracian
Image from Volf at Vol_Vlad LiveJournal.

Their relationships with Scythians and Greeks

The Scythian presence south of the Danube must be balanced with a Thracian presence north of the river. We have observed Getae there in Alexander’s day, settled and raising grain. For Strabo the coastlands from the Danube delta north as far as the river and Greek city of Tyras were the Desert of the Getae (7.3.14), notable for its poverty and tracklessness beyond the great river. He seems to suggest also that it was here that Lysimachus was taken alive by Dromichaetes, king of the Getae, whose famous homily on poverty and imperialism only makes sense on the steppe beyond the river (7.3.8; cf. Diod. 21.12; further on Getic possessions above the Danube, Paus. 1.9 with Delev 2000, 393, who seems rather too skeptical; on poverty, cf. Ballesteros Pastor 2003). This was the kind of discourse more familiarly found among Scythians, proud and blunt in the strength of their poverty. However, as Herodotus makes clear, simple pastoralism was not the whole story as one advanced round into Scythia. For he observes the agriculture practiced north and west of Olbia. These were the lands of the Alizones and the people he calls the Scythian Ploughmen, not least to distinguish them from the Royal Scythians east of Olbia, in whose outlook, he says, these agriculturalist Scythians were their inferiors, their slaves (Hdt. 4.20). The key point here is that, as we began to see with the Getan grain-fields of Alexander’s day, there was scope for Thracian agriculturalists to maintain their lifestyles if they moved north of the Danube, the steppe notwithstanding. It is true that it is movement in the other direction that tends to catch the eye, but there are indications in the literary tradition and, especially, in the archaeological record that there was also significant movement northward from Thrace across the Danube and the Desert of the Getae beyond it.

Greek literary sources were not much concerned with Thracian migration into Scythia, but we should observe the occasional indications of that process in very different texts and contexts. At the level of myth, it is to be remembered that Amazons were regularly considered to be of Thracian ethnicity from Archaic times onward and so are often depicted in Thracian dress in Greek art (Bothmer 1957; cf. Sparkes 1997): while they are most familiar on the south coast of the Black Sea, east of Sinope, they were also located on the north coast, especially east of the Don (the ancient Tanais). Herodotus reports an origin-story of the Sauromatians there, according to which this people had been created by the union of some Scythian warriors with Amazons captured on the south coast and then washed up on the coast of Scythia (4.110). While the story is unhistorical, it is not without importance. First, it reminds us that passage north from the Danube was not the only way that Thracians, Thracian influence, and Thracian culture might find their way into Scythia. There were many more and less circuitous routes, especially by sea, that could bring Thrace into Scythia. Secondly, the myth offered some ideological basis for the Sauromatian settlement in Thrace that Strabo records, for Sauromatians might claim a Thracian origin through their Amazon forebears. Finally, rather as we saw that Heracles could bring together some of the peoples of the region, we should also observe that Ares, whose earthly home was located in Thrace by a strong Greek and Roman tradition, seems also to have been a deity of special significance and special cult among the Scythians. So much was appropriate, especially from a Classical perspective, in associations between these two peoples, whose fame resided especially in their capacity for war.

skythen
Scythians: cultures and findings (ca. 7th-4th/3rd c. BC). Greek colonies marked with concentric circles.

This broad picture of cultural contact, interaction, and osmosis, beyond simple conflict, provides the context for a range of archaeological discoveries, which – if examined separately – may seem to offer no more than a scatter of peculiarities. Here we must acknowledge especially the pioneering work of Melyukova, who has done most to develop thinking on Thracian–Scythian interaction. As she pointed out, we have a good example of Thracian–Scythian osmosis as early as the mid-seventh century bce at Tsarev Brod in northeastern Bulgaria, where a warrior’s burial combines elements of Scythian and Thracian culture (Melyukova 1965). For, while the manner of his burial and many of the grave goods find parallels in Scythia and not Thrace, there are also goods which would be odd in a Scythian burial and more at home in a Thracian one of this period (notably a Hallstatt vessel, an iron knife, and a gold diadem). Also interesting in this regard are several stone figures found in the Dobrudja which resemble very closely figures of this kind (baby) known from Scythia (Melyukova 1965, 37–38). They range in date from perhaps the sixth to the third centuries bce, and presumably were used there – as in Scythia – to mark the burials of leading Scythians deposited in the area. Is this cultural osmosis? We should probably expect osmosis to occur in tandem with the movement of artefacts, so that only good contexts can really answer such questions from case to case. However, the broad pattern is indicated by a range of factors. Particularly notable in this regard is the observable development of a Thraco-Scythian form of what is more familiar as “Scythian animal style,” a term which – it must be understood – already embraces a range of types as we examine the different examples of the style across the great expanse from Siberia to the western Ukraine. As Melyukova observes, Thrace shows both items made in this style among Scythians and, more numerous and more interesting, a Thracian tendency to adapt that style to local tastes, with observable regional distinctions within Thrace itself. Among the Getae and Odrysians the adaptation seems to have been at its height from the later fifth century to the mid-third century (Melyukova 1965, 38; 1979).

The absence of local animal style in Bulgaria before the fifth century bce confirms that we have cultural influences and osmosis at work here, though that is not to say that Scythian tradition somehow dominated its Thracian counterpart, as has been claimed (pace Melyukova 1965, 39; contrast Kitov 1980 and 1984). Of particular interest here is the horse-gear (forehead-covers, cheek-pieces, bridle fittings, and so on) which is found extensively in Romania and Bulgaria as well as in Scythia, both in hoarded deposits and in burials. This exemplifies the development of a regional animal style, not least in silver and bronze, which problematizes the whole issue of the place(s) of its production. Accordingly, the regular designation as “Thracian” of horse-gear from the rich fourth century Scythian burial of Oguz in the Ukraine becomes at least awkward and questionable (further, Fialko 1995). And let us be clear that this is no minor matter, nor even part of a broader debate about the shared development of toreutics among Thracians and Scythians (e.g., Kitov 1980 and 1984). A finely equipped horse of fine quality was a strong statement and striking display of wealth and the power it implied

(…) while Thracian pottery appears at Olbia, Scythian pottery among Thracians is largely confined to the eastern limits of what should probably be regarded as Getic territory, namely the area close to the west of the Dniester, from the sixth century bce. Rather exceptional then is the Scythian pottery noted at Istros, which has been explained as a consequence of the Scythian pursuit of the withdrawing army of Darius and, possibly, a continued Scythian grip on the southern Danube in its aftermath (Melyukova 1965, 34). The archaeology seems to show us, therefore, that the elite Thracians and Scythians were more open to adaptation and acculturation than were their lesser brethren.

palaeo-balkan-languages
Paleo-Balkan languages in Eastern Europe between 5th and 1st century BC. From Wikipedia.

Conclusion

(…) we see distinct peoples and organizations, for example as Sitalces’ forces line up against the Scythians. Much more striking, however, against that general background, are the various ways in which the two peoples and their elites are seen to interact, connect, and share a cultural interface. We see also in Scyles’ story how the Greek cities on the coast of Thrace and Scythia played a significant role in the workings of relationships between the two peoples. It is not simply that these cities straddled the Danube, but also that they could collaborate – witness the honors for Autocles, ca. 300 bce (SEG 49.1051; Ochotnikov 2006) – and were implicated with the interactions of the much greater non-Greek powers around them. At the same time, we have seen the limited reality of familiar distinctions between settled Thracians and nomadic Scythians and the limited role of the Danube too in dividing Thrace and Scythia. The interactions of the two were not simply matters of dynastic politics and the occasional shared taste for artefacts like horse-gear, but were more profoundly rooted in the economic matrix across the region, so that “Scythian” nomadism might flourish in the Dobrudja and “Thracian-style” agriculture and settlement can be traced from Thrace across the Danube as far as Olbia. All of that offers scant justification for the Greek tendency to run together Thracians and Scythians as much the same phenomenon, not least as irrational, ferocious, and rather vulgar barbarians (e.g., Plato, Rep. 435b), because such notions were the result of ignorance and chauvinism. However, Herodotus did not share those faults to any degree, so that we may take his ready movement from Scythians to Thracians to be an indication of the importance of interaction between the two peoples whom he had encountered not only as slaves in the Aegean world, but as powerful forces in their own lands (e.g., Hdt. 4.74, where Thracian usage is suddenly brought into his account of Scythian hemp). Similarly, Thucydides, who quite without need breaks off his disquisition on the Odrysians to remark upon political disunity among the Scythians (Thuc. 2.97, a favorite theme: cf. Hdt. 4.81; Xen., Cyr. 1.1.4). As we have seen throughout this discussion, there were many reasons why Thracians might turn the thoughts of serious writers to Scythians and vice versa.

It seems, following Sikora et al. (2014), that Thracian ‘common’ populations would have more Anatolian Neolithic ancestry compared to more ‘steppe-like’ samples. But there were important differences even between the two nearby samples published from Bulgaria, which may account for the close interaction between Scythians and Thracians we see in Krzewińska et al. (2018), potentially reflected in the differences between the Central, Southern and the South-Central clusters (possibly related to different periods rather than peoples??).

If these R1b-Z2103 were descended from Thracian elites, this would be the first proof of Palaeo-Balkan populations showing mainly R1b-Z2103, as I expect. Their appearance together with haplogroup I2a2a1b1 (also found in Ukraine Neolithic and in the Yamna outlier from Bulgaria) seem to support this regional continuity, and thus a long-lasting cultural and ethnic border roughly around the Danube, similar to the one found in the northern Caucasus.

However, since these samples are some 2,500 years younger than the Yamna expansion to the south, and they are archaeologically Scythians, it is impossible to say. In any case, it would seem that the main expansion of R1a-Z645 lineages to the south of the Danube – and therefore those found among modern Greeks – was mediated by the Slavic expansions centuries later.

krzewinska-scythians-pca
Modified image from Krzewińska et al. (2018), with added Y-DNA haplogroups to each defined Scythian cluster and Sarmatians. Principal component analysis (PCA) plot visualizing 35 Bronze Age and Iron Age individuals presented in this study and in published ancient individuals in relation to modern reference panel from the Human Origins data set. See image with population references.

On the Northern cluster there is a sample of haplogroup R1b-P312 which, given its position on the PCA (apparently even more ‘modern Celtic’-like than the Hallstatt_Bylany sample from Damgaard et al. 2018), it seems that it could be the product of the previous eastward Hallstatt expansion…although potentially also from a recent one?:

Especially important in the archaeology of this interior is the large settlement at Nemirov in the wooded steppe of the western Ukraine, where there has been considerable excavation. This settlement’s origins evidently owe nothing significant to Greek influence, though the early east Greek pottery there (from ca. 650 bce onward: Vakhtina 2007) and what seems to be a Greek graffito hint at its connections with the Greeks of the coast, especially at Olbia, which lay at the estuary of the River Bug on whose middle course the site was located (Braund 2008). The main interest of the site for the present discussion, however, is its demonstrable participation in the broader Hallstatt culture to its west and south (especially Smirnova 2001). Once we consider Nemirov and the forest steppe in connection with Olbia and the other locations across the forest steppe and coastal zone, together with the less obvious movements across the steppe itself, we have a large picture of multiple connectivities in which Thrace bulks large.

scythian-peoples-balkans
Early Iron Age cultures of the Carpathian basin ca. 7-6th century BC, including steppe-related groups. Ďurkovič et al. (2018).

While the above description of clear-cut R1a-Steppe and R1b-Balkans is attractive (and probably more reliable than admixture found in scattered samples of unclear dates), the true ancient genetic picture is more complicated than that:

  • There is nothing in the material culture of the published western Scythians to distinguish the supposed Thracian elites.
  • We have the sample I0575, an Early Sarmatian from the southern Urals (one of the few available) of haplogroup R1b-Z2106, which supports the presence of R1b-Z2103 lineages among Eastern Iranian-speaking peoples.
  • We also have DA30, a Sarmatian of I2b lineage from the central steppes in Kazakhstan (ca. 47 BC – 24 AD).
  • Other Sarmatian samples of haplogroup R remain undefined.
  • There is R1a-Z93 in a late Sarmatian-Hun sample, which complicates the picture of late pastoralist nomads further.

Therefore, the possibility of hidden pockets of Iranian peoples of R1b-Z2103 (maybe also R1b-P312) lineages remains the best explanation, and should not be discarded simply because of the prevalent haplogroups among modern populations, or because of the different clusters found, or else we risk an obvious circular reasoning: “this sample is not (autosomically or in prevalent haplogroups) like those we already had from the steppe, ergo it is not from this or that steppe culture.” Hopefully, the upcoming paper by Järve et al. will help develop a clearer genetic transect of Iranian populations from the steppes.

All in all, the diversity among western Scythians represents probably one of the earliest difficult cases of acculturation to be studied with ancient DNA (obviously not the only one), since Scythians combine unclear archaeological data with limited and conflicting proto-historical accounts (also difficult to contrast with the wide confidence intervals of radiocarbon dates) with different evolving clusters and haplogroups – especially in border regions with strong and continued interactions of cultures and peoples.

With emerging complex cases like these during the Iron Age, I am happy to see that at least earlier expansions show clearer Y-DNA bottlenecks, or else genetics would only add more data to argue about potential cultural diffusion events, instead of solving questions about proto-language expansions once and for all…

Related

Common pitfalls in human genomics and bioinformatics: ADMIXTURE, PCA, and the ‘Yamnaya’ ancestral component

invasion-from-the-steppe-yamnaya

Good timing for the publication of two interesting papers, that a lot of people should read very carefully:

ADMIXTURE

Open access A tutorial on how not to over-interpret STRUCTURE and ADMIXTURE bar plots, by Daniel J. Lawson, Lucy van Dorp & Daniel Falush, Nature Communications (2018).

Interesting excerpts (emphasis mine):

Experienced researchers, particularly those interested in population structure and historical inference, typically present STRUCTURE results alongside other methods that make different modelling assumptions. These include TreeMix, ADMIXTUREGRAPH, fineSTRUCTURE, GLOBETROTTER, f3 and D statistics, amongst many others. These models can be used both to probe whether assumptions of the model are likely to hold and to validate specific features of the results. Each also comes with its own pitfalls and difficulties of interpretation. It is not obvious that any single approach represents a direct replacement as a data summary tool. Here we build more directly on the results of STRUCTURE/ADMIXTURE by developing a new approach, badMIXTURE, to examine which features of the data are poorly fit by the model. Rather than intending to replace more specific or sophisticated analyses, we hope to encourage their use by making the limitations of the initial analysis clearer.

The default interpretation protocol

Most researchers are cautious but literal in their interpretation of STRUCTURE and ADMIXTURE results, as caricatured in Fig. 1, as it is difficult to interpret the results at all without making several of these assumptions. Here we use simulated and real data to illustrate how following this protocol can lead to inference of false histories, and how badMIXTURE can be used to examine model fit and avoid common pitfalls.

admixture-protocol
A protocol for interpreting admixture estimates, based on the assumption that the model underlying the inference is correct. If these assumptions are not validated, there is substantial danger of over-interpretation. The “Core protocol” describes the assumptions that are made by the admixture model itself (Protocol 1, 3, 4), and inference for estimating K (Protocol 2). The “Algorithm input” protocol describes choices that can further bias results, while the “Interpretation” protocol describes assumptions that can be made in interpreting the output that are not directly supported by model inference

Discussion

STRUCTURE and ADMIXTURE are popular because they give the user a broad-brush view of variation in genetic data, while allowing the possibility of zooming down on details about specific individuals or labelled groups. Unfortunately it is rarely the case that sampled data follows a simple history comprising a differentiation phase followed by a mixture phase, as assumed in an ADMIXTURE model and highlighted by case study 1. Naïve inferences based on this model (the Protocol of Fig. 1) can be misleading if sampling strategy or the inferred value of the number of populations K is inappropriate, or if recent bottlenecks or unobserved ancient structure appear in the data. It is therefore useful when interpreting the results obtained from real data to think of STRUCTURE and ADMIXTURE as algorithms that parsimoniously explain variation between individuals rather than as parametric models of divergence and admixture.

For example, if admixture events or genetic drift affect all members of the sample equally, then there is no variation between individuals for the model to explain. Non-African humans have a few percent Neanderthal ancestry, but this is invisible to STRUCTURE or ADMIXTURE since it does not result in differences in ancestry profiles between individuals. The same reasoning helps to explain why for most data sets—even in species such as humans where mixing is commonplace—each of the K populations is inferred by STRUCTURE/ADMIXTURE to have non-admixed representatives in the sample. If every individual in a group is in fact admixed, then (with some exceptions) the model simply shifts the allele frequencies of the inferred ancestral population to reflect the fraction of admixture that is shared by all individuals.

Several methods have been developed to estimate K, but for real data, the assumption that there is a true value is always incorrect; the question rather being whether the model is a good enough approximation to be practically useful. First, there may be close relatives in the sample which violates model assumptions. Second, there might be “isolation by distance”, meaning that there are no discrete populations at all. Third, population structure may be hierarchical, with subtle subdivisions nested within diverged groups. This kind of structure can be hard for the algorithms to detect and can lead to underestimation of K. Fourth, population structure may be fluid between historical epochs, with multiple events and structures leaving signals in the data. Many users examine the results of multiple K simultaneously but this makes interpretation more complex, especially because it makes it easier for users to find support for preconceptions about the data somewhere in the results.

In practice, the best that can be expected is that the algorithms choose the smallest number of ancestral populations that can explain the most salient variation in the data. Unless the demographic history of the sample is particularly simple, the value of K inferred according to any statistically sensible criterion is likely to be smaller than the number of distinct drift events that have practically impacted the sample. The algorithm uses variation in admixture proportions between individuals to approximately mimic the effect of more than K distinct drift events without estimating ancestral populations corresponding to each one. In other words, an admixture model is almost always “wrong” (Assumption 2 of the Core protocol, Fig. 1) and should not be interpreted without examining whether this lack of fit matters for a given question.

admixture-pitfalls
Three scenarios that give indistinguishable ADMIXTURE results. a Simplified schematic of each simulation scenario. b Inferred ADMIXTURE plots at K= 11. c CHROMOPAINTER inferred painting palettes.

Because STRUCTURE/ADMIXTURE accounts for the most salient variation, results are greatly affected by sample size in common with other methods. Specifically, groups that contain fewer samples or have undergone little population-specific drift of their own are likely to be fit as mixes of multiple drifted groups, rather than assigned to their own ancestral population. Indeed, if an ancient sample is put into a data set of modern individuals, the ancient sample is typically represented as an admixture of the modern populations (e.g., ref. 28,29), which can happen even if the individual sample is older than the split date of the modern populations and thus cannot be admixed.

This paper was already available as a preprint in bioRxiv (first published in 2016) and it is incredible that it needed to wait all this time to be published. I found it weird how reviewers focused on the “tone” of the paper. I think it is great to see files from the peer review process published, but we need to know who these reviewers were, to understand their whiny remarks… A lot of geneticists out there need to develop a thick skin, or else we are going to see more and more delays based on a perceived incorrect tone towards the field, which seems a rather subjective reason to force researchers to correct a paper.

PCA of SNP data

Open access Effective principal components analysis of SNP data, by Gauch, Qian, Piepho, Zhou, & Chen, bioRxiv (2018).

Interesting excerpts:

A potential hindrance to our advice to upgrade from PCA graphs to PCA biplots is that the SNPs are often so numerous that they would obscure the Items if both were graphed together. One way to reduce clutter, which is used in several figures in this article, is to present a biplot in two side-by-side panels, one for Items and one for SNPs. Another stratagem is to focus on a manageable subset of SNPs of particular interest and show only them in a biplot in order to avoid obscuring the Items. A later section on causal exploration by current methods mentions several procedures for identifying particularly relevant SNPs.

One of several data transformations is ordinarily applied to SNP data prior to PCA computations, such as centering by SNPs. These transformations make a huge difference in the appearance of PCA graphs or biplots. A SNPs-by-Items data matrix constitutes a two-way factorial design, so analysis of variance (ANOVA) recognizes three sources of variation: SNP main effects, Item main effects, and SNP-by-Item (S×I) interaction effects. Double-Centered PCA (DC-PCA) removes both main effects in order to focus on the remaining S×I interaction effects. The resulting PCs are called interaction principal components (IPCs), and are denoted by IPC1, IPC2, and so on. By way of preview, a later section on PCA variants argues that DC-PCA is best for SNP data. Surprisingly, our literature survey did not encounter even a single analysis identified as DC-PCA.

The axes in PCA graphs or biplots are often scaled to obtain a convenient shape, but actually the axes should have the same scale for many reasons emphasized recently by Malik and Piepho [3]. However, our literature survey found a correct ratio of 1 in only 10% of the articles, a slightly faulty ratio of the larger scale over the shorter scale within 1.1 in 12%, and a substantially faulty ratio above 2 in 16% with the worst cases being ratios of 31 and 44. Especially when the scale along one PCA axis is stretched by a factor of 2 or more relative to the other axis, the relationships among various points or clusters of points are distorted and easily misinterpreted. Also, 7% of the articles failed to show the scale on one or both PCA axes, which leaves readers with an impressionistic graph that cannot be reproduced without effort. The contemporary literature on PCA of SNP data mostly violates the prohibition against stretching axes.

pca-how-to
DC-PCA biplot for oat data. The gradient in the CA-arranged matrix in Fig 13 is shown here for both lines and SNPs by the color scheme red, pink, black, light green, dark green.

The percentage of variation captured by each PC is often included in the axis labels of PCA graphs or biplots. In general this information is worth including, but there are two qualifications. First, these percentages need to be interpreted relative to the size of the data matrix because large datasets can capture a small percentage and yet still be effective. For example, for a large dataset with over 107,000 SNPs for over 6,000 persons, the first two components capture only 0.3693% and 0.117% of the variation, and yet the PCA graph shows clear structure (Fig 1A in [4]). Contrariwise, a PCA graph could capture a large percentage of the total variation, even 50% or more, but that would not guarantee that it will show evident structure in the data. Second, the interpretation of these percentages depends on exactly how the PCA analysis was conducted, as explained in a later section on PCA variants. Readers cannot meaningfully interpret the percentages of variation captured by PCA axes when authors fail to communicate which variant of PCA was used.

Conclusion

Five simple recommendations for effective PCA analysis of SNP data emerge from this investigation.

  1. Use the SNP coding 1 for the rare or minor allele and 0 for the common or major allele.
  2. Use DC-PCA; for any other PCA variant, examine its augmented ANOVA table.
  3. Report which SNP coding and PCA variant were selected, as required by contemporary standards in science for transparency and reproducibility, so that readers can interpret PCA results properly and reproduce PCA analyses reliably.
  4. Produce PCA biplots of both Items and SNPs, rather than merely PCA graphs of only Items, in order to display the joint structure of Items and SNPs and thereby to facilitate causal explanations. Be aware of the arch distortion when interpreting PCA graphs or biplots.
  5. Produce PCA biplots and graphs that have the same scale on every axis.

I read the referenced paper Biplots: Do Not Stretch Them!, by Malik and Piepho (2018), and even though it is not directly applicable to the most commonly available PCA graphs out there, it is a good reminder of the distorting effects of stretching. So for example quite recently in Krause-Kyora et al. (2018), where you can see Corded Ware and BBC samples from Central Europe clustering with samples from Yamna:

NOTE. This is related to a vertical distorsion (i.e. horizontal stretching), but possibly also to the addition of some distant outlier sample/s.

pca-cwc-yamna-bbc
Principal Component Analysis (PCA) of the human Karsdorf and Sorsum samples together with previously published ancient populations projected on 27 modern day West Eurasian populations (not shown) based on a set of 1.23 million SNPs (Mathieson et al., 2015). https://doi.org/10.7554/eLife.36666.006

The so-called ‘Yamnaya’ ancestry

Every time I read papers like these, I remember commenters who kept swearing that genetics was the ultimate science that would solve anthropological problems, where unscientific archaeology and linguistics could not. Well, it seems that, like radiocarbon analysis, these promising developing methods need still a lot of refinement to achieve something meaningful, and that they mean nothing without traditional linguistics and archaeology… But we already knew that.

Also, if this is happening in most peer-reviewed publications, made by professional geneticists, in journals of high impact factor, you can only wonder how many more errors and misinterpretations can be found in the obscure market of so many amateur geneticists out there. Because amateur geneticist is a commonly used misnomer for people who are not geneticists (since they don’t have the most basic education in genetics), and some of them are not even ‘amateurs’ (because they are selling the outputs of bioinformatic tools)… It’s like calling healers ‘amateur doctors’.

NOTE. While everyone involved in population genetics is interested in knowing the truth, and we all have our confirmation (and other kinds of) biases, for those who get paid to tell people what they want to hear, and who have sold lots of wrong interpretations already, the incentives of ‘being right’ – and thus getting involved in crooked and paranoid behaviour regarding different interpretations – are as strong as the money they can win or loose by promoting themselves and selling more ‘product’.

As a reminder of how badly these wrong interpretations of genetic results – and the influence of the so-called ‘amateurs’ – can reflect on research groups, yet another turn of the screw by the Copenhagen group, in the oral presentations at Languages and migrations in pre-historic Europe (7-12 Aug 2018), organized by the Copenhagen University. The common theme seems to be that Bell Beaker and thus R1b-L23 subclades do represent a direct expansion from Yamna now, as opposed to being derived from Corded Ware migrants, as they supported before.

NOTE. Yes, the “Yamna → Corded Ware → Únětice / Bell Beaker” migration model is still commonplace in the Copenhagen workgroup. Yes, in 2018. Guus Kroonen had already admitted they were wrong, and it was already changed in the graphic representation accompanying a recent interview to Willerslev. However, since there is still no official retraction by anyone, it seems that each member has to reject the previous model in their own way, and at their own pace. I don’t think we can expect anyone at this point to accept responsibility for their wrong statements.

So their lead archaeologist, Kristian Kristiansen, in The Indo-Europeanization of Europé (sic):

kristiansen-migrations
Kristiansen’s (2018) map of Indo-European migrations

I love the newly invented arrows of migration from Yamna to the north to distinguish among dialects attributed by them to CWC groups, and the intensive use of materials from Heyd’s publications in the presentation, which means they understand he was right – except for the fact that they are used to support a completely different theory, radically opposed to those defended in Heyd’s model

Now added to the Copenhagen’s unending proposals of language expansions, some pearls from the oral presentation:

  • Corded Ware north of the Carpathians of R1a lineages developed Germanic;
  • R1b borugh [?] Italo-Celtic;
  • the increase in steppe ancestry on north European Bell Beakers mean that they “were a continuation of the Yamnaya/Corded Ware expansion”;
  • Corded Ware groups [] stopped their expansion and took over the Bell Beaker package before migrating to England” [yep, it literally says that];
  • Italo-Celtic expanded to the UK and Iberia with Bell Beakers [I guess that included Lusitanian in Iberia, but not Messapian in Italy; or the opposite; or nothing like that, who knows];
  • 2nd millennium BC Bronze Age Atlantic trade systems expanded Proto-Celtic [yep, trade systems expanded the language]
  • 1st millennium BC expanded Gaulish with La Tène, including a “Gaulish version of Celtic to Ireland/UK” [hmmm, dat British Gaulish indeed].

You know, because, why the hell not? A logical, stable, consequential, no-nonsense approach to Indo-European migrations, as always.

Also, compare still more invented arrows of migrations, from Mikkel Nørtoft’s Introducing the Homeland Timeline Map, going against Kristiansen’s multiple arrows, and even against the own recent fantasy map series in showing Bell Beakers stem from Yamna instead of CWC (or not, you never truly know what arrows actually mean):

corded-ware-migrations
Nørtoft’s (2018) maps of Indo-European migrations.

I really, really loved that perennial arrow of migration from Volosovo, ca. 4000-800 BC (3000+ years, no less!), representing Uralic?, like that, without specifics – which is like saying, “somebody from the eastern forest zone, somehow, at some time, expanded something that was not Indo-European to Finland, and we couldn’t care less, except for the fact that they were certainly not R1a“.

This and Kristiansen’s arrows are the most comical invented migration routes of 2018; and that is saying something, given the dozens of similar maps that people publish in forums and blogs each week.

NOTE. You can read a more reasonable account of how haplogroup R1b-L51 and how R1-Z645 subclades expanded, and which dialects most likely expanded with them.

We don’t know where these scholars of the Danish workgroup stand at this moment, or if they ever had (or intended to have) a common position – beyond their persistent ideas of Yamnaya™ ancestral component = Indo-European and R1a must be Indo-European – , because each new publication changes some essential aspects without expressly stating so, and makes thus everything still messier.

It’s hard to accept that this is a series of presentations made by professional linguists, archaeologists, and geneticists, as stated by the official website, and still harder to imagine that they collaborate within the same professional workgroup, which includes experienced geneticists and academics.

I propose the following video to close future presentations introducing innovative ideas like those above, to help the audience find the appropriate mood:

Related

Reproductive success among ancient Icelanders stratified by ancestry

iceland-pca

New paper (behind paywall), Ancient genomes from Iceland reveal the making of a human population, by Ebenesersdóttir et al. Science (2018) 360(6392):1028-1032.

Abstract and relevant excerpts (emphasis mine):

Opportunities to directly study the founding of a human population and its subsequent evolutionary history are rare. Using genome sequence data from 27 ancient Icelanders, we demonstrate that they are a combination of Norse, Gaelic, and admixed individuals. We further show that these ancient Icelanders are markedly more similar to their source populations in Scandinavia and the British-Irish Isles than to contemporary Icelanders, who have been shaped by 1100 years of extensive genetic drift. Finally, we report evidence of unequal contributions from the ancient founders to the contemporary Icelandic gene pool. These results provide detailed insights into the making of a human population that has proven extraordinarily useful for the discovery of genotype-phenotype associations.

icelanders
Shared drift of ancient and contemporary Icelanders. (A) Scatterplot of D-statistics reflecting Iceland-specific drift. To aid interpretation, we included values for ancient British-Irish Islanders and a subset of contemporary individuals (who were correspondingly removed from the reference populations).

We estimated the mean Norse ancestry of the settlement population (24 pre-Christians and one early Christian) as 0.566 [95% confidence interval (CI) 0.431–0.702], with a nonsignificant difference betweenmales (0.579) and females (0.521). Applying the same ADMIXTURE analysis to each of the 916 contemporary Icelanders, we obtained a mean Norse ancestry of 0.704 (95% CI 0.699–0.709). Although not statistically significant (t test p = 0.058), this difference is suggestive. A similar difference ofNorse ancestry was observed with a frequency-based weighted least-squares admixture estimator (16), 0.625 [Mean squared error (MSE) = 0.083] versus 0.74 (MSE = 0.0037). Finally, the D-statistic test D(YRI, X; Gaelic, Norse) also revealed a greater affinity between Norse and contemporary Icelanders (0.0004, 95% CI 0.00008–0.00072) than between Norse and ancient Icelanders (−0.0002, 95% CI −0.00056–0.00015). This observation raises the possibility that reproductive success among the earliest Icelanders was stratified by ancestry, as genetic drift alone is unlikely to systematically alter ancestry at thousands of independent loci (fig. S10). We note that many settlers of Gaelic ancestry came to Iceland as slaves, whose survival and freedom to reproduce is likely to have been constrained (17). Some shift in ancestry must also be due to later immigration from Denmark, which maintained colonial control over Iceland from 1380 to 1944 (for example, in 1930 there were 745 Danes out of a total population of 108,629 in Iceland) (18).

icelander-admixture
Shared drift of ancient and contemporary Icelanders. (B) Estimated Norse,
Gaelic, and Icelandic ancestry for ancient Icelanders using ADMIXTURE
in supervised mode.

Five pre-Christian Icelanders (VDP-A5, DAVA9, NNM-A1, SVK-A1 and TGS-A1) fall just outside the space occupied by contemporary Norse in Fig. 3A. That these individuals show a stronger signal of drift shared with contemporary Icelanders is also apparent in the results of ADMIXTURE, run in supervised mode with three contemporary reference populations (Norse, Gaelic, and Icelandic) (Fig. 3B). The correlation between the proportion of Icelandic ancestry from this analysis and PC1 in Fig. 2A is |r| = 0.913.(…)

(…) as the five ancient Icelanders fall well within the cluster of contemporary Scandinavians (Fig. 3C), we conclude that they, or close relatives, likely contributed more to the contemporary Icelandic gene pool than the other pre-Christians. We note that this observation is consistent with the inference that settlers of Norse ancestry had greater reproductive success than those of Gaelic ancestry.

icelanders-y-dna
Haplogroup data, from the paper. Image modified by me, with those close to Gaelic and British/Irish samples (see above Scatterplot of D-statistics and ADMIXTURE data) marked in fluorescent: yellow closer to Gaelic, green less close.

Ancient Icelanders show a clear relation with the typically Norse Y-DNA distribution: I1 / R1a-Z284 / R1b-U106.

  • Among R1a, the picture is uniformly of R1a-Z284 (at least five of the seven reported).
  • There are six samples of I1, with great variation in subclades.
  • Among R1b-L51 subclades (ten samples), there are U106 (at least one sample), L21 (three samples), and another P312 (L238); see above the relationship with those clustering closely with Gaelic samples, marked in fluorescent, which is compatible with Gaelic settlers (predominantly of R1b-L21 lineages) coming to Iceland as slaves.

Probably not much of a surprise, coming from Norse speakers, but they are another relevant reference for comparison with samples of East Germanic tribes, when they appear.

Also, the first reported Klinefelter (XXY) in ancient DNA (sample ID is YGS-B2).

Related: