Happy new year 2019…and enjoy our new books!

song-sheep-horses-header

Sorry for the last weeks of silence, I have been rather busy lately. I am having more projects going on, and (because of that) I also wanted to finish a project I have been working on for many months already.

I have therefore decided to publish a provisional version of the text, in the hope that it will be useful in the following months, when I won’t be able to update it as often as I would like to:

EDIT (20 JAN 2019): For those of you who are more comfortable reading in your native language, I have placed some links to automatic translations by Google Translate. They might work especially well for the texts of A Game of Clans & A Clash of Chiefs.

Don’t forget to check out the maps included in the supplementary materials: I have added Y-DNA, mtDNA, and ADMIXTURE data using GIS software. The PCA graphics are also important to follow the main text.

NOTE. Right now the files are only in my server. I will try to upload them to Academia.edu and Research Gate when I have time, I have uploaded them to Academia.edu and ResearchGate, in case the websites are too slow.

I would have preferred to wait for a thorough revision of the section on archaeology and the linguistic sections on Uralic, but I doubt I will have time when the reviews come, so it was either now or maybe next December…

I say so in the introduction, but it is evident that certain aspects of the book are tentative to say the least: the farther back we go from Late Proto-Indo-European, the less clear are many aspects. Also, linguistically I am not convinced about Eurasiatic or Nostratic, although they do have a certain interest when we try to offer a comprehensive view of the past, including ethnolinguistic identities.

I cannot be an expert in everything, and these books cover a lot. I am bound to publish many corrections as new information appears and more reviews are sent. For example, just days ago (before SNP calls of Wang et al. 2018 were published) some paragraphs implied that AME might have expanded Nostratic from the Middle East. Now it does not seem so, and I changed them just before uploading the text. That’s how tentative certain routes are, and how much all of this may change. And that only if we accept a Nostratic phylum…

NOTE. Since the first book I wrote was the linguistic one, and I have spent the last months updating the archaeology + genetics part, now many of you will probably understand 1) why I am so convinced about certain language relationships and 2) how I used many posts to clarify certain ideas and receive comments. Many posts offer probably a good timeline of what I worked with, and when.

Acknowledgements

I did not add this section to the books, because they are still not ready for print, but I think this is due somewhere now. It is impossible to reference all who have directly or indirectly contributed to this, so this is a list of those I feel have played an important role.

I am indebted to the following people (which does not mean that they share my views, obviously):

First and foremost, to Fernando López-Menchero, for having the patience to review with detail many parts on Indo-European linguistics, knowing that I won’t accept many of his comments anyway. The additional information he offers is invaluable, but I didn’t want to turn this into a huge linguistic encyclopaedia with unending discussions of tiny details of each reconstructed word. I think it is already too big as it is.

I would not have thought about doing this if it were not for the interest of Wekwos (Xavier Delamarre) in publishing a full book about the Indo-European demic diffusion model (in the second half of 2017, I think). It was them who suggested that I extended the content, when all I had done until then was write an essay and draw some maps in my free time between depositing the PhD thesis and defending it.

Sadly, as much as I would like to publish a book with a professional publisher, I don’t think ancient DNA lends itself for the traditional format, so my requests (mainly to have free licenses and being able to review the text at will, as new genetic papers are published) were logically not acceptable. Also, the main aim of all volumes, especially the linguistic one, is the teaching of essentials of Late Proto-Indo-European and related languages, and this objective would be thwarted by selling each volume for $50-70 and only in printed format. I prefer a wider distribution.

At first I didn’t think much of this proposal, because I do not benefit from this kind of publications in my scientific field, but with time my interest in writing a whole, comprehensive book on the subject grew to the point where it was already an ongoing project, probably by the start of 2018.

I would not have been in contact with Wekwos if it were not for user Camulogène Rix at Anthrogenica, so thanks for that and for the interest in this work.

I would not have thought of writing this either if not for the spontaneous support (with an unexpected phone call!) of a professor of the Complutense University of Madrid, Ángel Gómez Moreno, who is interested in this subject – as is his wife, a professor of Classics more closely associated to Indo-European studies, and who helped me with a search for Indo-Europeanists.

EDIT (1 JAN 2019): I remembered that Karin Bojs sent me her book after reading the demic diffusion model. I may have also thought about writing a whole book back then, but mid-2017 is probably too early for the project.

Professor Kortlandt is still to review the text, but he contributed to both previous essays in some very interesting ways, so I hope he can help me improve the parts on Uralic, and maybe alternative accounts of expansion for Balto-Slavic, depending on the time depth that he would consider warranted according to the Temematic hypothesis.

The maps are evidently (for those who are interested in genetics) in part the result of the effort of the late Jean Manco: As you can see from the maps including Y-DNA and mtDNA samples, I have benefitted from her way of organising data and publishing it. Similarly, the work of Iain McDonald in assessing the potential migration routes of R1b and R1a in Europe with the help of detailed maps was behind my idea for the first maps, and consequently behind these, too.

I should thank all people responsible for the release of free datasets to work with, including the Reich and Jena labs, the Veeramah Lab, and also researchers from the Max Planck Institute or the Mainz Palaeogenetics group, who didn’t mind to share with me datasets to work with.

Readers of this blog with interesting comments have also been essential for the improvement of the texts. You can probably see some of your many contributions there. I may not answer many comments, because I am always busy (and sometimes I just don’t have anything interesting to say), but I try to read all of them.

EDIT (1 JAN 2019) I think I should mention at least Chetan, Egg, or Robert George; but then I would leave out old europe, Sgr Ganesh, or Tileman Ehlen; and if I include them I would leave out others…

Users of other sites, like Anthrogenica, whose particular points of view and deep knowledge of some very specific aspects are sometimes very useful. In particular, user Anglesqueville helped me to fix some issues with the merging of datasets to obtain the PCAs and ADMIXTURE, and prepared some individual samples to merge them.

Even without posting anything, Google Analytics keeps sending me messages about increasing user fidelity (returning users), and stats haven’t really changed (which probably means more people are reading old posts), so thank you for that.

I hope you enjoy the books.

Happy new year!

Biparental inheritance of mitochondrial DNA in humans

mtdna-inheritance-paternal

New paper Biparental Inheritance of Mitochondrial DNA in Humans, by Luo et al. PNAS (2018).

Interesting excerpts (emphasis mine):

Abstract

Although there has been considerable debate about whether paternal mitochondrial DNA (mtDNA) transmission may coexist with maternal transmission of mtDNA, it is generally believed that mitochondria and mtDNA are exclusively maternally inherited in humans. Here, we identified three unrelated multigeneration families with a high level of mtDNA heteroplasmy (ranging from 24 to 76%) in a total of 17 individuals. Heteroplasmy of mtDNA was independently examined by high-depth whole mtDNA sequencing analysis in our research laboratory and in two Clinical Laboratory Improvement Amendments and College of American Pathologists-accredited laboratories using multiple approaches. A comprehensive exploration of mtDNA segregation in these families shows biparental mtDNA transmission with an autosomal dominantlike inheritance mode. Our results suggest that, although the central dogma of maternal inheritance of mtDNA remains valid, there are some exceptional cases where paternal mtDNA could be passed to the offspring. Elucidating the molecular mechanism for this unusual mode of inheritance will provide new insights into how mtDNA is passed on from parent tooffspring and may even lead to the development of new avenues for the therapeutic treatment for pathogenic mtDNA transmission.

An example

Compared with Family A, a strikingly similar mtDNA transmission pattern was demonstrated in Families B and C. Taking Family B for illustration, II-3 having 29 heteroplasmic and seven homoplasmic variants should have inherited mtDNA from both his father (I-1, haplogroup of K1b2a) and his mother (I-10, haplogroup of H), who were supposed to possess 34 and nine homoplasmic variants, respectively. II-3 further transmitted his mtDNA that he inherited from I-1 to his son (III-2), who also inherited all of his mother’s mtDNA (II-30, carrying 34 variants and a haplogroup of T2a1a). However, III-2’s sister (III-1) and half-brother (III-5) only inherited the maternal mtDNA. Fresh blood sampling and repeated mtDNA sequencing in a second independent laboratory were also performed to rule out the possibility of sample mix-up for III-2 (III-2, column F-G and H-I). Additionally, these samples were further verified using Pacific Bio single molecular sequencing (see Materials and Methods) and by restriction fragment length polymorphism (RFLP) analysis of Family A, and these results were fully consistent with the previous sequencing.

mtdna-inheritance
Biparental mtDNA inheritance pattern shown in Family B. (A) Pedigree of Family B. The black filled symbols indicate the two family members (II-3 and III-2) showing biparental mtDNA transmission. The IDs of five family members tested by whole mtDNA sequencing analysis have been underlined in the pedigree. (B) Schematic of the mtDNA genotype defined by the homoplasmic and/or heteroplasmic variants aligned from the reference mitochondrial genome. Blue bars represent the genotype of paternally derived mtDNA, whereas purple-red and orange-red bars represent maternally derived mtDNA. Entries labeled (D) represent deduced mtDNA genotypes. (C) Summary of the haplogroup and mtDNA variant numbers in Family B.

A Resurgence of the Paternal Transmission Hypothesis

The results presented in this paper make a robust case for paternal transmission of mtDNA. Here, we report biparental mtDNA inheritance (either directly or indirectly) in 17 members in three multigeneration families. Thirteen of these individuals were identified directly by sequencing of the mitochondrial genome, whereas four could be inferred based on preexisting maternal heteroplasmy caused by biparental inheritance in the previous generation.

To further confirm these remarkable results and to exclude the possibility of sample mix-up and/or contamination, the whole mtDNA sequencing procedure was repeated independently in at least two different laboratories by different laboratory technicians with newly obtained blood samples. All results were reproducible, indicating no artifacts or contamination exist. More importantly, the multiple mtDNA variants that were paternally transmitted differ in both number and position among each of these three families as well as the related haplogroup (R0a1 in Family A, K1b2a in Family B, and K2b1a1a in Family C, respectively), providing two distinct forms of evidence supporting transmission of the paternal mtDNA.

Therefore, we have unequivocally demonstrated the existence of biparental mtDNA inheritance as evidenced by the high number and level of mtDNA heteroplasmy in these three unrelated multigeneration families. Most interestingly, the mixed haplogroups in these samples are very reminiscent of the mixed haplogroups found in the 20 studies that were dismissed by Bandelt et al. as due to contamination or sample mix-up. One is forced to wonder how many other instances of individuals with biparental mtDNA inheritance have been dismissed as technical errors, and whether Schwartz and Vissing’s original discovery has really been given the proper follow-up that it deserves. We suspect that these results will initiate a broader reassessment of the topic.

We propose that the paternal mtDNA transmission in these families should be accompanied by segregation of a mutation in one nuclear gene involved in paternal mitochondrial elimination and that there is a high probability that the gene in question operates through one of the pathways identified above.

If I have to be honest, I was stuck with the paternal transmission hypothesis which we were taught in class long ago. I didn’t know it was controversial or dismissed, I just thought it was really exceptional, and I never thought about learning more on the subject.

This paper proves it may be more complicated than that, especially for population genomics purposes, because biparental mtDNA transmission with autosomal dominant-like inheritance puts a serious barrier to a general, simplistic interpretation of mtDNA.

I don’t think it is a blow to all interpretations based on mtDNA, though, because the traditional interpretation should often work statistically. However, one has to be always very careful when saying “if it’s mtDNA from region X, it’s about female exogamy”, especially when samples are from neighbouring regions and similar periods.

The term “uniparental marker” for mtDNA is obviously misleading and shouldn’t be used, and many research papers and interpretations taking mtDNA as strictly uniparental should be taken with a pinch of salt.

Related

Genetic landscape and past admixture of modern Slovenians

slovenes-snp

Open access Genetic Landscape of Slovenians: Past Admixture and Natural Selection Pattern, by Maisano Delser et al. Front. Genet. (2018).

Interesting excerpts (emphasis mine):

Samples

Overall, 96 samples ranging from Slovenian littoral to Lower Styria were genotyped for 713,599 markers using the OmniExpress 24-V1 BeadChips (Figure 1), genetic data were obtained from Esko et al. (2013). After removing related individuals, 92 samples were left. The Slovenian dataset has been subsequently merged with the Human Origin dataset (Lazaridis et al., 2016) for a total of 2163 individuals.

Y chromosome

First, Y chromosome genetic diversity was assessed. A total of 52 Y chromosomes were analyzed for 195 SNPs. The majority of individuals (25, 48.1%) belong to the haplogroup R1a1a1a (R-M417) while the second major haplogroup is represented by R1b (R-M343) including 15 individuals (28.8%). Twelve samples are assigned to haplogroup I (I M170): five and two samples belong to haplogroup I2a (I L460) and I1 (I M253), respectively, while the remaining five samples did not have enough information to be further assigned.

pca-slovenes
PCA of Slovenian samples with European populations (Slovenian_HO_EU dataset). For details regarding the populations included, see Supplementary Table 1.

PCA

Considering the unbalanced sample size of the Slovenian population compared to the other populations included in the dataset, a subset of 20 Slovenian individuals randomly sampled was used.

All Slovenian samples group together with Hungarians, Czechs, and some Croatians (“Central-Eastern European” cluster) as also suggested by the PCA. All Basque individuals with few French and Spanish cluster together (“Basque” cluster) while a “Northern-European” cluster is made of the majority of French, English, Icelanders, Norwegians, and Orcadians. Five populations contributed to the “Eastern-European” cluster including Belarusians, Estonians, Lithuanians, Mordovians, and Russians. Western and South Europe is split into two cluster: the first (“Western European” cluster) includes all Spanish individuals, few French, and some Italians (North Italy) while the second (“Southern-European” cluster) groups Sicilians, Greeks, some Croatians, Romanians, and some Italians (North Italy).

Admixture Pattern and Migration

admixture-slovenians
Modified image, from the paper (Central-East Europeans marked). Unsupervised admixture analysis of Slovenians. Results for K = 5 are showed as it represents the lowest cross-validation error. Slovenian samples show an admixture pattern similar to the neighboring populations such as Croatians and Hungarians. The major ancestral components are: the blue one which is shared with Lithuanians and Russians, followed by the dark green one that is mostly present in Greek samples and the light blue which characterizes Orcadians and English. For population acronyms see Supplementary Table 1.

All Slovenian individuals share common pattern of genetic ancestry, as revealed by ADMIXTURE analysis. The three major ancestry components are the North East and North West European ones (light blue and dark blue, respectively, Figure 3), followed by a South European one (dark green, Figure 3). Contribution from the Sardinians and Basque are present in negligible amount. The admixture pattern of Slovenians mimics the one suggested by the neighboring Eastern European populations, but it is different from the pattern suggested by North Italian populations even though they are geographically close.

Using ALDER, the most significant admixture event was obtained with Russians and Sardinians as source populations and it happened 135 ± 9.31 generations ago (Z-score = 11.54). (…) When tested for multiple admixture events (MALDER), we obtained evidence for one admixture event 165.391 ± 17.1918 generations ago corresponding to ∼2620 BCE (CI: 3101–2139) considering a generation time of 28 years (Figure 4), with Kalmyk and Sardinians as sources.

We then modeled the Slovenian population as target of admixture of ancient individuals from Haak et al. (2015) while computing the f3(Ancient 1, Ancient 2, Slovenian) statistic. The most significant signal was obtained with Yamnaya and HungaryGamba_EN (Z-score = -10.66), followed by MA1 with LBK_EN (Z-score -9.7) and Yamnaya with Stuttgart (Z-score = -8.6) used as possible source populations (Supplementary Figure 5).

We found a significant signal of admixture by using both pairs as ancient sources. Specifically, for the pair Yamnaya and Hungary_EN the admixture event is dated at 134.38 ± 23.69 generations ago (Z-score = 5.26, p-value of 1.5e-07) while for Yamnaya and LBK_EN at 153.65 ± 22.19 generations ago (Z-score = 6.92, p-value 4.4e-12). Outgroup f3 with Yamnaya put Slovenian population close to Hungarians, Czechs, and English, indicating a similar shared drift between these population with the Steppe populations (Supplementary Figure 6).

admixture-events-slovenes
Admixture events identified with ALDER and MALDER. The gray dots represent significant admixture events detected with ALDER using Slovenians as target, the solid line represents the single admixture event detected using MALDER, dashed lines represent the confidence interval. Only the significant results after multiple testing correction are plotted. For ALDER results see Supplementary Table 5.

Not that any of this would come as a surprise, but:

  • R1a-M458 and some R1a-Z280 (xR1a-Z92) lineages (found among Slovenes) were associated with the Slavic expansion, likely with the Prague-Korchak culture, originally stemming probably from peoples of the Lusatian culture. Other R1a-Z280 lineages remained associated with Uralic peoples, and some became Slavicized only recently.
  • PCA keeps supporting the common cluster of certain West, South, and East Slavs in a “Central-Eastern European” cluster, distinct from the “North-Eastern European” cluster formed by modern Finno-Ugrians, as well as ancient Finno-Ugrians of north-eastern Europe who were only recently Slavicized.
  • Admixture supports the same ancient ‘western’ (a core West+South+East Slavic) cluster, and the admixture event with Yamna + Hungary_EN is logically a proxy for Yamna Hungary being at the core of ancestral Central-East population movements related to Bell Beakers in the mid- to late 3rd millennium.

The theory that East Slavs are at the core of the Slavic expansion makes no sense, in terms of archaeology (see Florin Curta’s dismissal of those recent eastern ‘Slavic’ finds, his commentary on 19th century Pan-Slavic crap, or his book on Slavic migrations), in terms of ancient DNA (the earliest Slavs sampled cluster with modern West Slavs, distant from the steppe cluster, unlike Finno-Ugrians), or in terms of modern DNA.

I don’t know where exactly this impulse for the theory of Russia being the cradle of Slavs comes from today (although there are some obvious political trends to revive 19th c. ideas), but it was always clear for everyone, including Russians, that East Slavs had migrated to the east and north and assimilated indigenous Finno-Ugrians, apart from Turkic-, Iranian-, and Caucasian-speaking peoples to the east. Genetics is only confirming what was clear from other disciplines long ago.

Related

Deep population history of North, Central and South America

human-divergence-americas

Open access Reconstructing the Deep Population History of Central and South America, by Posth et al. Cell (2018).

Abstract:

We report genome-wide ancient DNA from 49 individuals forming four parallel time transects in Belize, Brazil, the Central Andes, and the Southern Cone, each dating to at least ∼9,000 years ago. The common ancestral population radiated rapidly from just one of the two early branches that contributed to Native Americans today. We document two previously unappreciated streams of gene flow between North and South America. One affected the Central Andes by ∼4,200 years ago, while the other explains an affinity between the oldest North American genome associated with the Clovis culture and the oldest Central and South Americans from Chile, Brazil, and Belize. However, this was not the primary source for later South Americans, as the other ancient individuals derive from lineages without specific affinity to the Clovis-associated genome, suggesting a population replacement that began at least 9,000 years ago and was followed by substantial population continuity in multiple regions.

Interesting excerpts:

The D4h3a mtDNA haplogroup has been hypothesized to be a marker for an early expansion into the Americas along the Pacific coast (Perego et al., 2009). However, its presence in two Lapa do Santo individuals and Anzick-1 (Rasmussen et al., 2014) makes this hypothesis unlikely.

The patterns we observe on the Y chromosome also force us to revise our understanding of the origins of present-day variation. Our ancient DNA analysis shows that the Q1a2a1b-CTS1780 haplogroup, which is currently rare, was present in a third of the ancient South Americas. In addition, our observation of the currently extremely rare C2b haplogroup at Lapa do Santo disproves the suggestion that it was introduced after 6,000 BP (Roewer et al., 2013).

(…) Our discovery that the Clovis-associated Anzick-1 genome at ∼12,800 BP shares distinctive ancestry with the oldest Chilean, Brazilian, and Belizean individuals supports the hypothesis that an expansion of people who spread the Clovis culture in North America also affected Central and South America, as expected if the spread of the Fishtail Complex in Central and South America and the Clovis Complex in North America were part of the same phenomenon (direct confirmation would require ancient DNA from a Fishtail-context) (Pearson, 2017). However, the fact that the great majority of ancestry of later South Americans lacks specific affinity to Anzick-1 rules out the hypothesis of a homogeneous founding population. Thus, if Clovis-related expansions were responsible for the peopling of South America, it must have been a complex scenario involving arrival in the Americas of sub-structured lineages with and without specific Anzick-1 affinity, with the one with Anzick-1 affinity making a minimal long-term contribution. While we cannot at present determine when the non-Anzick-1 associated lineages first arrived in South America, we can place an upper bound on the date of the spread to South America of all the lineages represented in our sampled ancient genomes as all are ANC-A and thus must have diversified after the ANC-A/ANC-B split estimated to have occurred ∼17,500–14,600 BP (Moreno-Mayar et al., 2018a).

deep-population-history-americas


New paper (behind paywall) Early human dispersals within the Americas, by Moreno-Mayar et al. Science (2018).

Abstract:

Studies of the peopling of the Americas have focused on the timing and number of initial migrations. Less attention has been paid to the subsequent spread of people within the Americas. We sequenced 15 ancient human genomes spanning Alaska to Patagonia; six are ≥10,000 years old (up to ~18× coverage). All are most closely related to Native Americans, including an Ancient Beringian individual, and two morphologically distinct “Paleoamericans.” We find evidence of rapid dispersal and early diversification, including previously unknown groups, as people moved south. This resulted in multiple independent, geographically uneven migrations, including one that provides clues of a Late Pleistocene Australasian genetic signal, and a later Mesoamerican-related expansion. These led to complex and dynamic population histories from North to South America.

Interesting excerpts:

The Australasian signal is not present in USR1 or Spirit Cave, but only appears in Lagoa Santa. None of these individuals has UPopA/Mesoamerican-related admixture, which ap-parently dampened the Australasian signature in South American groups, such as the Karitiana. These findings suggest the Australasian signal, possibly present in a structured ancestral NA population, was absent in NA prior to the Spirit Cave/Lagoa Santa split. Groups carrying this signal were either already present in South America when the ancestors of Lagoa Santa reached the region, or Australasian-related groups arrived later but before 10.4 ka (the Lagoa Santa 14C age). That this signal has not been previously documented in North America implies that an earlier group possessing it had disappeared, or a later-arriving group passed through North America without leaving any genetic trace. If such a signal is ultimately detected in North America it could help determine when groups bear-ing Australasian ancestry arrived, relative to the divergence of SNA groups.

Although we detect the Australasian signal in one of the Lagoa Santa individuals identified as a “Paleoamerican,” it is absent in other “Paleoamericans” (2, 10), including Spirit Cave with its strong genetic affinities to Lagoa Santa. This indicates the “Paleoamerican” cranial form is not associated with the Australasian genetic signal, as previously suggested (6), or any other specific NA clade (2). The cause of this cranial form, if it is representative of broader population pat-terns, evidently did not result from separate ancestry, but likely multiple factors, including isolation and drift and non-stochastic mechanisms.

australasian
f-statistics–based tests show a rapid dispersal into South America, followed by Mesoamerican-related admixture. Schematic representation of a model for SNA formation. This model represents a reasonable fit to most present-day populations.

Open access The genetic prehistory of the Andean highlands 7000 years BP though European contact, by Lindo et al. Science Advances (2018).

Abstract:

The peopling of the Andean highlands above 2500 m in elevation was a complex process that included cultural, biological, and genetic adaptations. Here, we present a time series of ancient whole genomes from the Andes of Peru, dating back to 7000 calendar years before the present (BP), and compare them to 42 new genome-wide genetic variation datasets from both highland and lowland populations. We infer three significant features: a split between low- and high-elevation populations that occurred between 9200 and 8200 BP; a population collapse after European contact that is significantly more severe in South American lowlanders than in highland populations; and evidence for positive selection at genetic loci related to starch digestion and plausibly pathogen resistance after European contact. We do not find selective sweep signals related to known components of the human hypoxia response, which may suggest more complex modes of genetic adaptation to high altitude.

Related

Corded Ware—Uralic (III): “Siberian ancestry” and Ugric-Samoyedic expansions

siberian-ancestry-tambets

This is the third of four posts on the Corded Ware—Uralic identification. See

An Eastern Uralic group?

Even though proposals of an Eastern Uralic (or Ugro-Samoyedic) group are in the minority – and those who support it tend to search for an origin of Uralic in Central Asia – , there is nothing wrong in supporting this from the point of view of a western homeland, because the eastward migration of both Proto-Ugric and Pre-Samoyedic peoples may have been coupled with each other at an early stage. It’s like Indo-Slavonic: it just doesn’t fit the linguistic data as well as the alternative, i.e. the expansion of Samoyedic first, different from a Finno-Ugric trunk. But, in case you are wondering about this possibility, here is Häkkinen’s (2012) phonological argument:

ugro-samoyedic-uralic

The case of Samoyedic is quite similar to that of Hungarian, although the earliest Palaeo-Siberian contact languages have been lost. There were contacts at least with Tocharian (Kallio 2004), Yukaghir (Rédei 1999) and Turkic (Janhunen 1998). Samoyedic also:

a) has moved far from the related languages and has been exposed to strong foreign influence

b) shares a small number of common words with other branches (from Sammallahti 1988: only 123 ‘Uralic’ words, versus 390 ‘Uralic’ + ‘Finno-Ugric’ words found in other branches than Samoyedic = 31,5 %)

c) derives phonologically from the East Uralic dialect.

The phonological level is taxonomically more reliable, since it lacks the distortion caused by invisible convergence and false divergence at the lexical level. Thus we can conclude that the traditional taxonomic model, according to which Samoyedic was the first branch to split off from the Proto-Uralic unity, is just as incorrect as the view that Hungarian was the first branch to split off.

Seima-Turbino

Late Uralic can be traced back to metallurgical cultures thanks to terms like PU *wäśka ‘copper/bronze’ (borrowed from Proto-Samoyedic *wesä into Tocharian); PU *äsa and *olna/*olni, ‘lead’ or ‘tin’, found in *äsa-wäśka ‘tin-bronze’; and e.g. *weŋći ‘knife’, borrowed into Indo-Iranian (through the stage of vocalization of nasals), appearing later as Proto-Indo-Aryan *wāćī ‘knife, awl, axe’.

It is known that the southern regions of the Abashevo culture developed Proto-Indo-Iranian-speaking Sintashta-Petrovka and Pokrovka (Early Srubna). To the north, however, Abashevo kept its Uralic nature, with continuous contacts allowing for the spread of lexicon – mainly into Finno-Ugric – , and phonetic influence – mainly Uralisms into Proto-Indo-Iranian phonology (read more here).

The northern part of Abashevo (just like the south) was mainly a metallurgical society, with Abashevo metal prospectors found also side by side with Sintashta pioneers in the Zeravshan Valley, near BMAC, in search of metal ores. About the Seima-Turbino phenomenon, from Parpola (2013):

From the Urals to the east, the chain of cultures associated with this network consisted principally of the following: the Abashevo culture (extending from the Upper Don to the Mid- and South Trans-Urals, including the important cemeteries of Sejma and Turbino), the Sintashta culture (in the southeast Urals), the Petrovka culture (in the Tobol-Ishim steppe), the Taskovo-Loginovo cultures (on the Mid- and Lower Tobol and the Mid-Irtysh), the Samus’ culture (on the Upper Ob, with the important cemetery of Rostovka), the Krotovo culture (from the forest steppe of the Mid-Irtysh to the Baraba steppe on the Upper Ob, with the important cemetery of Sopka 2), the Elunino culture (on the Upper Ob just west of the Altai mountains) and the Okunevo culture (on the Mid-Yenissei, in the Minusinsk plain, Khakassia and northern Tuva). The Okunevo culture belongs wholly to the Early Bronze Age (c. 2250–1900 BCE), but most of the other cultures apparently to its latter part, being currently dated to the pre-Andronovo horizon of c. 2100–1800 BCE (cf. Parzinger 2006: 244–312 and 336; Koryakova & Epimakhov 2007: 104–105).

post-eneolithic-steppe-asia
Schematic map of the Middle Bronze Age cultures (steppe and foreststeppe
zone)

The majority of the Sejma-Turbino objects are of the better quality tin-bronze, and while tin is absent in the Urals, the Altai and Sayan mountains are an important source of both copper and tin. Tin is also available in southern Central Asia. Chernykh & Kuz’minykh have accordingly suggested an eastern origin for the Sejma-Turbino network, backing this hypothesis also by the depiction on the Sejma-Turbino knives of mountain sheep and horses characteristic of that area. However, Christian Carpelan has emphasized that the local Afanas’evo and Okunevo metallurgy of the Sayan-Altai area was initially rather primitive, and could not possibly have achieved the advanced and difficult technology of casting socketed spearheads as one piece around a blank. Carpelan points out that the first spearheads of this type appear in the Middle Bronze Age Caucasia c. 2000 BCE, diffusing early on to the Mid-Volga-Kama-southern Urals area, where “it was the experienced Abashevo craftsmen who were able to take up the new techniques and develop and distribute new types of spearheads” (Carpelan & Parpola 2001: 106, cf. 99–106, 110). The animal argument is countered by reference to a dagger from Sejma on the Oka river depicting an elk’s head, with earlier north European prototypes (Carpelan & Parpola 2001: 106–109). Also the metal analysis speaks for the Abashevo origin of the Sejma-Turbino network. Out of 353 artefacts analyzed, 47% were of tin-bronze, 36% of arsenical bronze, and 8.5% of pure copper. Both the arsenical bronze and pure copper are very clearly associated with the Abashevo metallurgy.

seima-turbino-phenomenon-parpola
Find spots of artefacts distributed by the Sejma-Turbino intercultural trader network, and the areas of the most important participating cultures: Abashevo, Sintashta, Petrovka. Based on Chernykh 2007: 77.

The Abashevo metal production was based on the Volga-Kama-Belaya area sandstone ores of pure copper and on the more easterly Urals deposits of arsenical copper (Figure 9). The Abashevo people, expanding from the Don and Mid-Volga to the Urals, first reached the westerly sandstone deposits of pure copper in the Volga and Kama basins, and started developing their metallurgy in this area, before moving on to the eastern side of the Urals to produce harder weapons and tools of arsenical copper. Eventually they moved even further south, to the area richest in copper in the whole Urals region, founding there the very strong and innovative Sintashta culture.

Regarding the most likely expansion of Eastern Uralic peoples:

Nataliya L’vovna Chlenova (1929–2009; cf. Korenyako & Ku’zminykh 2011) published in 1981 a detailed study of the Cherkaskul’ pottery. In her carefully prepared maps of 1981 and 1984 (Figure 10), she plotted Cherkaskul’ monuments not only in Bashkiria and the Trans-Urals, but also in thick concentrations on the Upper Irtysh, Upper Ob and Upper Yenissei, close to the Altai and Sayan mountains, precisely where the best experts suppose the homeland of Proto-Samoyed to be.

cherkaskul-andronovo
Distribution of Srubnaya (Timber Grave, early and late), Andronovo (Alakul’ and Fëdorovo variants) and Cherkaskul’ monuments. After Parpola 1994: 146, fig. 8.15, based on the work of N. L. Chlenova (1984: map facing page 100).

Ugric

The Cherkaskul’ culture was transformed into the genetically related Mezhovka culture (c. 1500–1000 BCE), which occupied approximately the same area from the Mid-Kama and Belaya rivers to the Tobol river in western Siberia (cf. Parzinger 2006: 444–448; Koryakova & Epimakhov 2007: 170–175). The Mezhovka culture was in close contact with the neighbouring and probably Proto-Iranian speaking Alekseevka alias Sargary culture (c. 1500–900 BCE) of northern Kazakhstan (Figure 4 no. 8) that had a Fëdorovo and Cherkaskul’ substratum and a roller pottery superstratum (cf. Parzinger 2006: 443–448; Koryakova & Epimakhov 2007: 161–170). Both the Cherkaskul’ and the Mezhovka cultures are thought to have been Proto-Ugric linguistically, on the basis of the agreement of their area with that of Mansi and Khanty speakers, who moreover in their Fëdorovo-like ornamentation have preserved evidence of continuity in material culture (cf. Chlenova 1984; Koryakova & Epimakhov 2007: 159, 175).

mezhovska-sargary-irmen
Cultures of the Final Bronze Age of the Urals and western Siberia (steppe
and forest-steppe zone).

The Mezhovka culture was succeeded by the genetically related Gamayun culture (c. 1000–700 BCE) (cf. Parzinger 2006: 446; 542–545).

From the Gamayun culture descend Trans-Urals cultures in close contact with Finno-Permic populations of the Cis-Ural region:

  • [Proto-Mansi] Itkul’ culture (c. 700–200 BCE) distributed along the eastern slope of the Ural Mountains (cf. Parzinger 2006: 552–556). Known from its walled forts, it constituted the principal Trans-Uralian centre of metallurgy in the Iron Age, and was in contact with both the Anan’ino and Akhmylovo cultures (the metallurgical centres of the Mid-Volga and Kama-Belaya region) and the neighbouring Gorokhovo culture.
    • [Proto-Hungarian] via the Vorob’evo Group (c. 700–550 BCE) (cf. Parzinger 2006: 546–549), to the Gorokhovo culture (c. 550–400 BCE) of the Trans-Uralian forest steppe (cf. Parzinger 2006: 549–552). For various reasons the local Gorokhovo people started mobile pastoral herding and became part of the multicomponent pastoralist Sargat culture (c. 500 BCE to 300 CE), which in a broader sense comprized all cultural groups between the Tobol and Irtysh rivers, succeeding here the Sargary culture. The Sargat intercommunity was dominated by steppe nomads belonging to the Iranian-speaking Saka confederation, who in the summer migrated northwards to the forest steppe
  • [Proto-Khanty] Late Bronze Age and Early Iron Age cultures related to the Gamayunskoe and Itkul’ cultures that extended up to the Ob: the Nosilovo, Baitovo, Late Irmen’, and Krasnoozero cultures (c. 900–500 BCE). Some were in contact with the Akhmylovo on the Mid-Volga.
sargat-gorokhovo-bolscherechye
Cultural groups of the Iron Age in the forest-steppe zone of western
Siberia. (

Samoyedic

Parpola (2012) connects the expansion of Samoyedic with the Cherkaskul variant of Andronovo. As we know, Andronovo was genetically diverse, which speaks in favour of different groups developing similar material cultures in Central Asia.

Juha Janhunen, author of the etymological dictionary of the Samoyed languages (1977), places the homeland of Proto-Samoyedic in the Minusinsk basin on the Upper Yenissei (cf. Janhunen 2009: 72). Mainly on the basis of Bulghar Turkic loanwords, Janhunen (2007: 224; 2009: 63) dates Proto-Samoyedic to the last centuries BCE. Janhunen thinks that the language of the Tagar culture (c. 800–100 BCE) ought to have been Proto-Samoyedic (cf. Janhunen 1983: 117– 118; 2009: 72; Parzinger 2001: 80 and 2006: 619–631 dates the Tagar culture c. 1000–200 BCE; Svyatko et al. 2009: 256, based on human bone samples, c. 900 BCE to 50 CE). The Tagar culture largely continues the traditions of the Karasuk culture (c. 1400–900 BCE), (…)

chicha-irmen-tagar-baraba-forest-siberian
Map showing the location of Chicha-1.

For the most recent expansions of Samoyedic languages to the north, into Palaeo-Siberian populations, read more about the traditional multilingualism of Siberian populations.

Genetics

Siberian ancestry

The use of a map of “Siberian ancestry” peaking in the arctic to show a supposedly late Uralic population movement (starting in the Iron Age!) seems to be the latest trend in population genomics:

siberian-ancestry-map
Frequency map of the so-called ‘Siberian’ component. From Tambets et al. (2018) (see below for ADMIXTURE in specific populations).

I guess that would make this map of Neolithic farmer ancestry represent an expansion of Indo-European from the south, because Anatolia, Greece, Italy, southern France, and Iberia – where this ancestry peaks in modern populations – are among the oldest territories where Indo-European languages were recorded:

reich-farmer-ancestry
Modern genome-wide data shows that the primary gradient of farmer ancestry in Europe does not flow southeast-to-northwest but instead in an almost perpendicular direction, a result of a major migration of pastoralists from the east that displaced much of the ancestry of the first farmers.

Probably not the right interpretation of this kind of simplistic data about modern populations, though…

The most striking thing about the “Siberian ancestry” white whale is that nobody really knows what it is; just like we did not know what “Yamnaya ancestry” was, until the most recent data is making the picture clearer. Its nature is changing with each new paper, and it can be summed up by “some ancestry we want to find that is common to Uralic-speaking peoples, and should not be CWC-related”. Tambets et al. (2018) explain quite well how they “found it”:

Overall, and specifically at lower values of K, the genetic makeup of Uralic speakers resembles that of their geographic neighbours. The Saami and (a subset of) the Mansi serve as exceptions to that pattern being more similar to geographically more distant populations (Fig. 3a, Additional file 3: S3). However, starting from K = 9, ADMIXTURE identifies a genetic component (k9, magenta in Fig. 3a, Additional file 3: S3), which is predominantly, although not exclusively, found in Uralic speakers. This component is also well visible on K = 10, which has the best cross-validation index among all tests (Additional file 3: S3B). The spatial distribution of this component (Fig. 3b) shows a frequency peak among Ob-Ugric and Samoyed speakers as well as among neighbouring Kets (Fig. 3a). The proportion of k9 decreases rapidly from West Siberia towards east, south and west, constituting on average 40% of the genetic ancestry of FU speakers in Volga-Ural region (VUR) and 20% in their Turkic-speaking neighbours (Bashkirs, Tatars, Chuvashes; Fig. 3a).

siberian-ancestry-modern
Population structure of Uralic-speaking populations inferred from ADMIXTURE analysis on autosomal SNPs in Eurasian context. Individual ancestry estimates for populations of interest for selected number of assumed ancestral populations (K3, K6, K9, K11). Ancestry components discussed in a main text (k2, k3, k5, k6, k9, k11) are indicated and have the same colours throughout. The names of the Uralic-speaking populations are indicated with blue (Finno-Ugric) or orange (Samoyedic). Image from Tambets et al. (2018).

However, this ‘something’ that some people occasionally find in some Uralic populations is also common to other modern and ancient groups, and not so common in some other Uralic peoples. Simply put:

siberian-ancestry-modern-populations
Image modified from Lamnidis et al. (2018). Red line representing maximum “Siberian admixture” in Eastern European hunter-gatherers. In blue, Uralic-speaking groups. “Plot of ADMIXTURE (K=3) results containing West Eurasian populations and the Nganasan. Ancient individuals from this study are represented by thicker bars.”

I already said this in the recent publication of Siberian samples, where a renamed and radiocarbon dated Finnish_IA clearly shows that Late Iron Age Saami (ca. 400 AD) had little “Siberian ancestry”, if any at all, representing the most likely Fennic (and Samic) ancestral components before their expansion into central and northern Finland, where they admixed with circum-polar peoples of asbestos ware cultures.

I will say that again and again, any time they report the so-called “Siberian ancestry” in Uralic samples, no matter how it is defined each time: it does not seem to be that special something people are looking for, but rather (at least in a great part) a quite old ancestral component forming an evident cline with EHG, whose best proximate source are Baikal_EN (and/or Devil’s Gate) at this moment, and thus also East European hunter-gatherers for Western Uralic peoples:

dzudzuana-baikal-en-admixture
Image modified from Lazaridis et al. (2018). In red: samples with Baikal_EN ancestry in speculative estimates. In pink: samples with Baikal_EN ancestry in conservative estimates (probably marking a recent arrival of Baikal_En ancestry, see here). Modeling present-day and ancient West-Eurasians. Mixture proportions computed with qpAdm (Supplementary Information section 4). The proportion of ‘Mbuti’ ancestry represents the total of ‘Deep’ ancestry from lineages that split prior to the split of Ust’Ishim, Tianyuan, and West Eurasians and can include both ‘Basal Eurasian’ and other (e.g., Sub-Saharan African) ancestry. (Left) ‘Conservative’ estimates. Each population 367 cannot be modeled with fewer admixture events than shown. (Right) ‘Speculative’ estimates. The highest number of sources (≤5) with admixture estimates within [0,1] are shown for each population. Some of the admixture proportions are not significantly different from 0 (Supplementary Information section 4).

So either Samara_HG, Karelia_HG, and many other groups from eastern Europe all spoke Uralic according to this ADMIXTURE graphic (and the formation of steppe ancestry in the Volga-Ural region brought the Proto-Indo-European language to the steppes through the CHG/ANE expansion), or a great part of this “Siberian ancestry” found in modern Uralic-speaking populations is not what some people would like to think it is…

Modern populations

PCA clines can be looked for to represent expansions of ancient populations. Most recently, Flegontov et al. (2018) are attempting to do this with Asian populations:

For some Turkic groups in the Urals and the Altai regions and in the Volga basin, a different admixture model fits the data: the same West Eurasian source + Uralic- or Yeniseian-speaking Siberians. Thus, we have revealed an admixture cline between Scythians and the Iranian farmer genetic cluster, and two further clines connecting the former cline to distinct ancestry sources in Siberia. Interestingly, few Wusun-period individuals harbor substantial Uralic/Yeniseian-related Siberian ancestry, in contrast to preceding Scythians and later Turkic groups characterized by the Tungusic/Mongolic-related ancestry. It remains to be elucidated whether this genetic influx reflects contacts with the Xiongnu confederacy. We are currently assembling a collection of samples across the Eurasian steppe for a detailed genetic investigation of the Hunnic confederacies.

jeong-population-clines
Three distinct East/West Eurasian clines across the continent with some interesting linguistic correlates, as earlier reported by Jeong et al. (2018). Alexander M. Kim.

There are potential errors with this approach:

The main one is practical – does a modern cline represent an ancestral language? The answer is: sometimes. It depends on the anthropological context that we have, and especially on the precision of the PCA:

clines-himalayan
Genetic structure of the Himalayan region populations from analyses using unlinked SNPs. (A) PCA of the Himalayan and HGDP-CEPH populations. Each dot represents a sample, coded by region as indicated. The Himalayan region samples lie between the HGDP-CEPH East Asian and South Asian samples on the right-hand side of the plot. From Arciero et al. (2018).

The ‘Europe’, ‘Middle East’, etc. clines of the above PCA do not represent one language, but many. For starters, the PCA includes too many (and modern) populations, its precision is useless for ethnolinguistic groups. Which is the right level? Again, it depends.

The other error is one of detail of the clines drawn (which, in turn, depends on the precision of the PCA). For example, we can draw two paralell lines (or even one line, as in Flegontov et al. above) in one PCA graphic, but we still don’t have the direction of expansion. How do we know if this supposed “Uralic-speaking cline” goes from one region to the other? For that level of detail, we should examine closely modern Uralic-speaking peoples and Circum-Arctic populations:

uralic-cline
Modified from Tambets et al. (2018). Principal component analysis (PCA) and genetic distances of Uralic-speaking populations. a PCA (PC1 vs PC2) of the Uralic-speaking populations

The real ancient Uralic cluster (drawn above in blue) is thus probably from a North-East European source (probably formed by Battle Axe / Fatyanovo-Balanovo / Abashevo) to the east into Siberian populations, and to the north into Laplandic populations (see below also on Mezhovska ancestry for the drawn ‘European cline’, which some may a priori wrongly assume to be quite late).

The fact that the three formed clines point to an admixture of CWC-related populations from North-Eastern Europe, and that variation is greater at the Palaeo-Laplandic and Palaeo-Siberian extremities compared to the CWC-related one, also supports this as the correct interpretation.

However, judging by the two main clines formed, one could be alternatively inclined to interpret that Palaeo-Laplandic and Palaeo-Siberian populations formed a huge ancestral “Uralic” ghost cluster in Siberia (spanning from the Palaeo-Laplandic to the Palaeo-Siberian one), and from there expanded Finno-Samic on one hand, and “Volga-Ugro-Samoyed” on the other. That poses different problems: an obvious linguistic and archaeological one – which I assume a lot of people do not really care about – , and a not-so-obvious genetic one (see below for ancient samples and for the expansion of haplogroup N).

To understand the simplest solution better, one can just have a look at the PCA from Bell Beaker samples in Olalde et al. (2018), which (as Reich has already explained many times) expanded directly from Yamna R1b-L23 lineages:

olalde_pca_clines
Image modified from Olalde et al. (2018). PCA of 999 Eurasian individuals. Marked is the Espersted Outlier with the approximate position of Yamna Hungary, probably the source of its admixture. Different Bell Beaker clines have been drawn, to represent approximate source of expansions from Central European sources into the different regions.

Unlike this PCA with ancient samples, where Bell Beaker clines could be a rough approximation to the real sources for each population, and where a cluster spanning all three depicted Early Bronze Age clusters could give a rough proximate source of European Bell Beakers in Hungary (and where one can even distinguish the Y-DNA bottlenecks in the L23 trunk created by each cline) the PCA of modern Uralic populations is probably not suitable for a good estimate of the ancient situation, which may be found shifted up or down of the drawn “Uralic” cluster along East European groups.

After all, we already know that the Siberian cline shows probably as much an ancient admixture event – from the original Uralic expansion to the east with Corded Ware ancestry – as another more recent one – a westward migration of Siberian ancestry (or even more than one). While we know with more or less exactitude what happened with the Palaeo-Laplandic admixture by expanding Proto-Finno-Samic populations (see here), the Proto-Ugric and Pre-Samoyedic populations formed probably more than one cline during the different ancient migrations through central Asia.

Ancient populations

Apparently, the Corded Ware expansion to the east was not marked by a huge change in ancestry. While the final version of Narasimhan et al. (2018) may show a little more detail about other forest-steppe Seima-Turbino/Andronovo-related migrations (and thus also Eastern Uralic peoples), we have already had enough information for quite some time to get a good idea.

mezhovska-pca
Principal component analysis. PCA of ancient individuals (according colours see legend) projected on modern West Eurasians (grey). Iron Age Scythians are shown in black; CHG, Caucasus hunter-gatherer; LNBA, late Neolithic/Bronze Age; MN, middle Neolithic; EHG, eastern European huntergatherer; LBK_EN, early Neolithic Linearbandkeramik; HG, hunter-gatherer; EBA, early Bronze Age; IA, Iron Age; LBA, late Bronze Age; WHG, western hunter-gatherer.dataset (grey). Iron Age Scythians are shown in black; CHG, Caucasus hunter-gatherer; LNBA, late Neolithic/Bronze Age; MN, middle Neolithic; EHG, eastern European hunter-gatherer; LBK_EN, early Neolithic Linearbandkeramik; HG, hunter-gatherer; EBA, early Bronze Age; IA, Iron Age; LBA, late Bronze Age; WHG, western hunter-gatherer.

Mezhovska‘s position is similar to the later Pre-Scythian and Scythian populations. There are some interesting details: apart from haplogroup R1a-Z280 (CTS1211+), there is one R1b-M269 (PF6494+), probably Z2103, and an outlier (out of three) in a similar position to the recently described central/southern Scythian clusters.

NOTE. The finding of R1b-M269 in the forest-steppe is probably either 1) from an Afanasevo-Okunevo origin, or 2) from an admixture with neighbouring Andronovo-related populations, such as Sargary. A third, maybe less likely option is that this haplogroup admixed with Abashevo directly (as it happened in Sintashta, Potapovka, or Pokrovka) and formed part of early Uralic migrations. In any case, since Mezhovska is a Bronze Age society from the Urals region, its association with R1b-Z2103 – like the association of R1b-Z2103 in Scythian clusters – cannot be attributed to “Thracian peoples”, a link which is (as I already said) too simplistic.

The drawn “European cline” of Hungarians (see above), leading from ‘west-like’ Mansi to Hungarian populations – and hosting also Finnic and Estonian samples – , cannot therefore be attributed simply to late “Slavic/Balkan-like” admixture.

Karasuk – located further to the east – is basically also Corded Ware peoples showing clearly a recent admixture with local ANE / Baikal_EN-like populations. In terms of haplogroups it shows haplogroup Q, R1a-Z2124, and R1a-Z2123, later found among early Hungarians, and present also in ancient Samoyedic populations now acculturated.

The most interesting aspect of both Mezhovska and Karasuk is that they seem to diverge from a point close to Ukraine_Eneolithic, which is the supposed ancestral source of Corded Ware peoples (read more about the formation of “steppe ancestry”). This means that Eastern Uralians derive from a source closer to Middle Dnieper/Abashevo populations, rather than Battle Axe (shifted to Latvian Neolithic), which is more likely the source prevalent in Finno-Permic peoples.

Their initial admixture with (Palaeo-)Siberian populations is thus seen already starting by this time in Mezhovska and especially in Karasuk, but this process (compared to modern populations) is incomplete:

f4-test-karasuk-mezhovska
Visualization of f-statistics results. f4(Test, LBK; Han, Mbuti) values are plotted on x axis and f4(Test, LBK; EHG, Mbuti) values on y axis, positive deviations from zero show deviations from a clade between Test and LBK. A red dashed line is drawn between Yamnaya from Samara and Ami. Iron Age populations that can be modelled as mixtures of Yamnaya and East Eurasians (like the Ami) are arrayed around this line and appear to be distinct from the main North/South European cline (blue) on the left of the x axis.
karasuk-mezhovska-admixture
ADMIXTURE results for ancient populations. Red arrows point to the Iron Age Scythian individuals studied. LBK_EN: Early Neolithic Linearbandkeramik; EHG: Eastern European hunter-gatherer; Motala_HG: hunter-gatherer from Motala (Sweden); WHG: western hunter-gatherer; CHG: Caucasus hunter-gatherer; IA: Iron Age; EBA: Early Bronze Age; LBA: Late Bronze Age.

We know now that Samic peoples expanded during the Late Iron Age into Palaeo-Laplandic populations, admixing with them and creating this modern cline. Finns expanded later to the north (in one of their known genetic bottlenecks), admixing with (and displacing) the Saami in Finland, especially replacing their male lines.

So how did Ugric and Samoyedic peoples admix with Palaeo-Siberian populations further, to obtain their modern cline? The answer is, logically, with East Asian migrations related to forest-steppe populations of Central Asia after the Mezhovska and Karasuk periods, i.e. during the Iron Age and later. Other groups from the forest-steppe in Central Asia show similar East Asian (“Siberian”) admixture. We know this from Narasimhan et al. (2018):

(…) we observe samples from multiple sites dated to 1700-1500 BCE (Maitan, Kairan, Oy_Dzhaylau and Zevakinsikiy) that derive up to ~25% of their ancestry from a source related to present-day East Asians and the remainder from Steppe_MLBA. A similar ancestry profile became widespread in the region by the Late Bronze Age, as documented by our time transect from Zevakinsikiy and samples from many sites dating to 1500-1000 BCE, and was ubiquitous by the Scytho-Sarmatian period in the Iron Age.

We already have some information about these later migrations:

siberian-genetic-component-chronology
Very important observation with implication of population turnover is that pre-Turkic Inner Eurasian populations’ Siberian ancestry appears predominantly “Uralic-Yeniseian” in contrast to later dominance of “Tungusic-Mongolic” sort (which does sporadically occur earlier). Alexander M. Kim

The Ugric-speaking Sargat culture in Western Siberia shows the expected mixture of haplogroups (ca. 500 BC – 500 AD), with 5 samples of hg N and 2 of hg R1a1, in Pilipenko et al. (2017). Although radiocarbon dates and subclades are lacking, N lineages probably spread late, because of the late and gradual admixture of Siberian cultures into the Sargat melting pot.

The Samoyedic-speaking Tagar culture also shows signs of a genetic turnover in Pilipenko et al. (2018):

The observed reduction in the genetic distance between the Middle Tagar population and other Scythian like populations of Southern Siberia(Fig 5; S4 Table), in our opinion, is primarily associated with an increase in the role of East Eurasian mtDNA lineages in the gene pool (up to nearly half of the gene pool) and a substantial increase in the joint frequency of haplogroups C and D (from 8.7% in the Early Tagar series to 37.5% in the Middle Tagar series). These features are characteristic of many ancient and modern populations of Southern Siberia and adjacent regions of Central Asia, including the Pazyryk population of the Altai Mountains.

Before the Iron Age, the Karasuk and Mezhovska population were probably already somehow ‘to the north’ within the ancient Steppe-Altai cline (see image below9 created by expanding Seima-Turbino- and Andronovo-related populations. During the Iron Age, further Siberian contributions with Iranian expansions must have placed Uralians of the Central Asian forest-steppe areas much closer to today’s Palaeo-Siberian cline.

However, the modern genetic picture was probably fully developed only in historic times, when Samoyedic and Ugric languages expanded to the north, only in part admixing further with Palaeo-Siberian-speaking nomads from the Circum-Arctic region (see here for a recent history of Samoyedic Enets), which justifies their more recent radical ‘northern shift’.

east-uralic-clines
Modified image from Jeong et al. (2018), supplementary materials. The first two PCs summarizing the genetic structure within 2,077 Eurasian individuals. The two PCs generally mirror geography. PC1 separates western and eastern Eurasian populations, with many inner Eurasians in the middle. PC2 separates eastern Eurasians along the north-south cline and also separates Europeans from West Asians. Ancient individuals (color-filled shapes), including two Botai individuals, are projected onto PCs calculated from present-day individuals.

This late acquisition of the language by Palaeo-Siberian nomads (without much population replacement) also justifies the wide PCA clusters of very small Siberian populations. See for example in the PCA from Tambets et al. (2018):

uralic-ugric-samoyedic-modern-clines
Approximate Ugric and Samoyedic clines (exluding apparent outliers). Modified from Tambets et al. (2018). Principal component analysis (PCA) and genetic distances of Uralic-speaking populations. a PCA (PC1 vs PC2) of the Uralic-speaking populations

For their relationship with modern Mansi, we have information on Hungarian conqueror populations from Neparáczki et al. (2018):

Moreover, Y, B and N1a1a1a1a Hg-s have not been detected in Finno-Ugric populations [80–84], implying that the east Eurasian component of the Conquerors and Finno-Ugric people are probably not directly related. The same inference can be drawn from phylogenetic data, as only two Mansi samples appeared in our phylogenetic trees on the side branches (S1 Fig, Networks; 1, 4) suggesting that ancestors of the Mansis separated from Asian ancestors of the Conquerors a long time ago. This inference is also supported by genomic Admixture analysis of Siberian and Northeastern European populations [85], which revealed that Mansis received their eastern Siberian genetic component approximately 5–7 thousand years ago from ancestors of modern Even and Evenki people. Most likely the same explanation applies to the Y-chromosome N-Tat marker which originated from China [86,87] and its subclades are now widespread between various language groups of North Asia and Eastern Europe [88].

The genetic picture of Hungarians (their formed cline with Mansi and their haplogroups) may be quite useful for the true admixture found originally in Mansi peoples at the beginning of the Iron Age. By now it is clear even from modern populations that Steppe_MLBA ancestry accompanied the Uralic expansion to the east (roughly approximated in the graphic with Afanasievo_EBA + Bichon_LP EasternHG_M):

siberian-population-expansions
Admixture modelling using qpAdm. Maps showing locations and ancestry proportions of ancient (left) and modern (right) groups. From Sikora et al. (2018).

Continue reading the final post of the series: Corded Ware—Uralic (IV): Haplogroups R1a and N in Finno-Ugric and Samoyedic.

See also

Related

  • The traditional multilingualism of Siberian populations
  • Iron Age bottleneck of the Proto-Fennic population in Estonia
  • Y-DNA haplogroups of Tuvinian tribes show little effect of the Mongol expansion
  • Corded Ware—Uralic (I): Differences and similarities with Yamna
  • Haplogroup R1a and CWC ancestry predominate in Fennic, Ugric, and Samoyedic groups
  • The Iron Age expansion of Southern Siberian groups and ancestry with Scythians
  • Evolution of Steppe, Neolithic, and Siberian ancestry in Eurasia (ISBA 8, 19th Sep)
  • Mitogenomes from Avar nomadic elite show Inner Asian origin
  • On the origin and spread of haplogroup R1a-Z645 from eastern Europe
  • Oldest N1c1a1a-L392 samples and Siberian ancestry in Bronze Age Fennoscandia
  • Consequences of Damgaard et al. 2018 (III): Proto-Finno-Ugric & Proto-Indo-Iranian in the North Caspian region
  • The concept of “Outlier” in Human Ancestry (III): Late Neolithic samples from the Baltic region and origins of the Corded Ware culture
  • Genetic prehistory of the Baltic Sea region and Y-DNA: Corded Ware and R1a-Z645, Bronze Age and N1c
  • More evidence on the recent arrival of haplogroup N and gradual replacement of R1a lineages in North-Eastern Europe
  • Another hint at the role of Corded Ware peoples in spreading Uralic languages into north-eastern Europe, found in mtDNA analysis of the Finnish population
  • New Ukraine Eneolithic sample from late Sredni Stog, near homeland of the Corded Ware culture
  • The traditional multilingualism of Siberian populations

    uralic-languages

    New paper (behind paywall) A case-study in historical sociolinguistics beyond Europe: Reconstructing patterns of multilingualism in a linguistic community in Siberia, by Khanina and Meyerhoff, Journal of Historical Sociolinguistics (2018) 4(2).

    The Nganasans have been eastern neighbours of the Enets for at least several centuries, or even longer, as indicated in Figures 2 and 3.10 They often dwelled on the same grounds and had common households with the Enets. Nganasans and Enets could intermarry (Dolgikh 1962a), while the Nganasans did not marry representatives of any other ethnic groups. As a result, it was not unusual for Enets and Nganasans to live in the same tent and/or to have common relatives. Such close contact must clearly have favoured acquisition of Nganasan by Enets children and of Enets by Nganasan children from an early age.

    The Nenets have been close neighbours of all the Enets groups more recently (Figures 2 and 3). In the seventeenth century, there were only warlike contacts between the Nenets and the Enets, while in the eighteenth century the Nenets started to live on the traditional Enets lands, on the western bank of the Yenisey river, with more peaceful interactions reported. (…) Since then the same situation of intermarriages and common households has been attested for these western Enets neighbours as with the Nganasans (Dolgikh 1962a), and this has also created conditions favouring early acquisition of both languages by children.

    enets-nganasan-early
    The Enets and neighbouring peoples in the middle of the seventeenth century; map by Yuri Koryakov (http://lingvarium.org), adapted from Dolgikh (1960).

    As for the Evenkis and the Selkups, the Enets had regular contact with these peoples (Figures 2 and 3), though they were not their close neighbours: in fact, geographically, the Selkups were not neighbours at all by the end of the nineteenth century. The Evenkis had always been direct south-eastern neighbours (…) Contacts with Selkups could be trade based, or they could simply be occasional encounters on adjacent lands. (…) [With Evenkis] some sporadic contacts were similar in nature to those with the Selkups, however many other contacts were war-like. Traditionally, the Enets considered the Evenkis to have a martial spirit, and the Evenkis were known as being accustomed to stealing Enets women. A number of stories in Dolgikh (1961) concern Evenkis stealing Enets women and Enets men going to Evenki lands to find and return them. It is clear, therefore, that if Evenki or Selkup were acquired by the Enets, this happened later in life, and this acquisition required particular conditions for it, i. e. it was not readily acquired through regular or harmonious contact (as with Nganasan).

    In a pattern similar to the situation with Nganasan, in the second half of the twentieth century most Enets elders could speak Nenets (Vasil’jev 1963; Eugen Helimski p.c., the lead author’s fieldwork experience).

    enets-nganasan-late
    The Enets and neighbouring indigenous peoples: end of the nineteenth century – beginning of the twentieth century; map by Yuri Koryakov (http://lingvarium.org), adapted from
    Bruk (1961).

    At the start of the period studied, in the 1850s, the Enets linguistic community could be characterized as multilingual in the following five languages: Enets, Nganasan, Nenets, Evenki, and Russian (Figure 4). The number of Enets individuals who were able to converse in each of the other four languages differed and generally was a property of the individuals who had regular social contact with speakers of the other four languages. (…) Note that in all cases of interethnic communication there could well be a lack of perfect proficiency in a language for which the multilingualism is ascribed to the Enets community or Enets individuals: as Braunmüller and Ferraresi (2003: 3) put it: “Nobody would ever have expected to know other languages ‘perfectly’ (whatever that may mean in detail). This expectation seems to be a quite modern idea when discussing issues of bilingualism or multilingualism in general”.

    The complex interactions of Siberian populations during the 17th-19th centuries offer a reasonably good picture of the life in the centuries before these accounts, when Samoyedic peoples migrated northwards, and Palaeo-Siberian and Tungusic populations were gradually assimilated into their Uralic culture and language, through intermarriage and close contacts among naturally nomadic populations.

    You can read more about the origin of Nganasans – and other modern Samoyedic-speaking peoples – as Palaeo-Siberian populations (hence probably speaking Palaeo-Siberian languages more or less related to each other) who adopted Samoyedic languages in Wikipedia, which offers a summary of Boris Dolgikh’s On the Origin of the Nganasans (1962). Dolgikh is one of the main sources of information for these Siberian groups, as is reflected in this paper, too.

    samoyedic
    Map of distribution of Samoyedic languages (red) in the XVII century (approximate; hatching) and in the end of XX century (continuous background). Notice late expansion to north and west into the typical territory where Nomadic peoples roamed. Modified from Wikipedia, with the Tuva region labelled (see a recent genetic study on the Tuva region, one of the most likely to be originally Samoyedic-speaking).

    Why some geneticists are using Nganasans – in fact the latest Palaeo-Siberians to learn Samoyedic, already during historic times – as a model for the expansion of Uralic? I have never understood that. Among the many cases of circular reasoning based on modern populations that have been created since the start of population genomics, the use of Nganasans as a model of ‘true Uralians’ is probably the most clearly frontally opposed to what was well known in anthropology before geneticists started this new field.

    If Kallio is right, most “eastern homeland” proposals are due to the interest of Russian nationalism, which is sadly quite likely to be influencing genetic research, too. It’s like letting Hindu nationalists influence publications on steppe-related migrations. As David Reich puts it in his book:

    The tensest twenty-four hours of my scientific career came in October 2008, when my collaborator Nick Patterson and I traveled to Hyderabad to discuss these initial results with Singh and Thangaraj.

    Our meeting on October 28 was challenging. Singh and Thangaraj seemed to be threatening to nix the whole project. Prior to the meeting, we had shown them a summary of our findings, which were that Indians today descend from a mixture of two highly divergent ancestral populations, one being “West Eurasians.” Singh and Thangaraj objected to this formulation because, they argued, it implied that West Eurasian people migrated en masse into India. They correctly pointed out that our data provided no direct evidence for this conclusion. They even reasoned that there could have been a migration in the other direction, of Indians to the Near East and Europe. (…) They also implied that the suggestion of a migration from West Eurasia would be politically explosive. They did not explicitly say this, but it had obvious overtones of the idea that migration from outside India had a transformative effect on the subcontinent.

    If you add the nation-building myths in Eastern Europe (like the Russian Euro-Asian movements) to the now prevalent Indo-European—CWC idea, and a Siberian ancestry peaking in the Arctic, with little demographic or political relevance of modern Uralic-speaking peoples, you have clearly an explosive sociopolitical mix (based on a mythical Pan-Eurasian Indo-Slavonic) in the making…

    euro-asian-empire-dugin
    Russia as the Euro-Asian Empire. Source: A. Dugin (1999), p. 415. From Eberhardt (2018).

    Related

    Iron Age bottleneck of the Proto-Fennic population in Estonia

    tarand-graves-estonia-early-late

    Demographic data and figures derived from Estonian Iron Age graves, by Raili Allmäe, Papers on Anthropology (2018) 27(2).

    Interesting excerpts (emphasis mine):

    Introduction

    Inhumation and cremation burials were both common in Iron Age Estonia; however, the pattern which burials were prevalent has regional as well temporal peculiarities. In Estonia, cremation burials appear in the Late Bronze Age (1100–500 BC), for example, in stone-cist graves and ship graves, although inhumation is still characteristic of the period [28, 18]. Cremation burials have occasionally been found beneath the Late Bronze Age cists and the Early Iron Age (500 BC–450 AD) tarand graves [30, 28]. In south-eastern Estonia, including Setumaa, the tradition to bury cremated human remains in pit graves also appeared in the Bronze Age and lasted during the Pre-Roman period (500 BC–50 AD) and the Roman Iron Age (50–450 AD), and even up to the medieval times [30, 23, 33, 9]. During the Early Iron Age, cremations appear in cairn graves and have occasionally been found in many Pre-Roman early tarand graves where they appear with inhumations [27, 28, 19, 20, 21, 22, 24]. In Roman Iron Age tarand graves, cremation as well inhumation were practiced [28, 36, 37]; however, cremation was the prevailing burial practice during the Roman Iron Age, for example, in tarand graves in south-eastern Estonia [30, 28]. Roman Iron Age (50 AD–450 AD) burial sites have not been found in continental west Estonia [28, 38]). At the beginning of the Middle Iron Age (450–800 AD), burial sites, for example stone graves without a formal structure, like Maidla I, Lihula and Ehmja ‘Varetemägi’, appear in Läänemaa, west Estonia; in these graves cremations as well inhumations have been found [39, 48]. Like underground cremation burial, the stone grave without a formal structure was the most common grave type during the Late Iron Age (800– 1200 AD) in west Estonia [39, 35, 48]. In south-eastern and eastern Estonia, sand barrows with cremation burials appeared at the beginning of the Middle Iron Age. Cremation barrows are attributed to the Culture of Long Barrows and are most numerous in the villages Laossina and Rõsna in northern Setomaa, on the western shore of Lake Peipsi [8, 48]. Apparently during the Iron Age, the practiced burial customs varied in Estonia.

    cist-grave-tarand
    Typical prehistoric Estonian graves. Top: Cist-graves common during the Bronze Age, by Terker (GNU FDL 1.2). Bottom: Tarand graves of the Iron Age, by Marika Mägi (2017)

    Abstract:

    Three Iron Age cremation graves from south-eastern Estonia and four graves including cremations as well inhumations from western Estonia were analysed by osteological and palaeodemographic methods in order to estimate the age and sex composition of burial sites, and to propose some possible demographic figures and models for living communities.

    The crude birth/death rate estimated on the basis of juvenility indices varied between 55.1‰ and 60.0‰ (58.5‰ on average) at Rõsna village in south-eastern Estonia in the Middle Iron Age. The birth/death rates based on juvenility indices for south eastern graves varied to a greater degree. The estimated crude birth/death rate was somewhat lower (38.9‰) at Maidla in the Late Iron Age and extremely high (92.1‰) at Maidla in the Middle Iron Age, which indicates an unsustainable community. High crude birth/death rates are also characteristic of Poanse tarand graves from the Pre-Roman Iron Age – 92.3‰ for the 1st grave and 69.6‰for the 2nd grave. Expectedly, newborn life expectancies are extremely low in both communities – 10.8 years at Poanse I and 14.4 years at Poanse II. Most likely, both Maidla I and Poanse I were unsustainable communities.

    tarand-graves-estonia
    Locations of the investigated Estonian Iron Age graves. Map by R. Allmäe

    According to the main model where the given period of grave usage is 150 years, the burial grounds were most likely exploited by communities of 3–14 people. In most cases, this corresponds to one family or household. In comparison with other graves, Maidla II stone grave in western Estonia and Rõsna-Saare I barrow cemetery in south-eastern Estonia could have been used by a somewhat larger community, which may mean an extended family, a larger household or usage by two nuclear families.

    More papers on the same subject by the author – who participated in the recent Mittnik et al. (2018) paper – include Observations On Estonian Iron Age Cremations (2013), and The demography of Iron Age graves in Estonia (2014).

    Fast life history in Iron Age Estonia

    While the demographic situation in the Gulf of Finland during the Iron Age is not well known – and demography is always difficult to estimate based on burials, especially when cremation is prevalent – , there is a clear genetic bottleneck in Finns, which has been estimated precisely to this period, coincident with the expansion of Proto-Fennic.

    estonian-pca
    PCA of Estonian samples from the Bronze Age, Iron Age and Medieval times. Tambets et al. (2018, upcoming).

    The infiltration of N1c lineages in Estonia – the homeland of Proto-Fennic – happened during the Iron Age – as of yet found in 3 out of 5 sampled Tarand graves – , while the previous period was dominated by 100% R1a and Corded Ware + Baltic HG ancestry. With the Iron Age, a slight shift towards Corded Ware ancestry can be seen, which probably signals the arrival of warrior-traders from the Alanino culture, close to the steppe. They became integrated through alliances and intermarriages in a culture of chiefdoms dominated by hill forts. Their origin in the Mid-Volga area is witnessed by their material culture, such as Tarand-like graves (read here a full account of events).

    This new political structure, reminiscent of the chiefdom system in Sintashta (with a similar fast life history causing a bottleneck of R1a-Z645 lineages), coupled with the expansion of Fennic (and displaced Saamic) peoples to the north, probably caused the spread of N1c-L392 among Finnic peoples. The linguistic influence of these early Iron Age trading movements from the Middle Volga region can be seen in similarities between Fennic and Mordvinic, which proves that the Fenno-Saamic community was by then not only separated linguistically, but also physically (unlike the period of long-term Palaeo-Germanic influence, where loanwords could diffuse from one to the other).

    NOTE. Either this, or the alternative version: an increase in Corded Ware ancestry in Estonia during the Iron Age marks the arrival of the first Fennic speakers ca. 800 BC or later, splitting from Mordvinic? A ‘Mordvin-Fennic’ group in the Volga, of mainly Corded Ware ancestry…?? Which comes in turn from a ‘Volga-Saamic’ population of Siberian ancestry in the Artic region??? And, of course, Palaeo-Germanic widely distributed in North-Eastern Europe with R1a during the Bronze Age! Whichever model you find more logical.

    Related

    The Tungusic Ulchi population probably linked to haplogroup C2b1a

    ulchi-marital

    New paper (behind paywall) Demographic and Genetic Portraits of the Ulchi Population, by Balanovska et al. Russian Journal of Genetics (2018) 54(10):1245–1253.

    Interesting excerpts (emphasis mine):

    Marital structure. The intensity of interethnic marriages puts the existence of the Ulchi population at risk. The colorful ethnic composition of the Ulchi settlements is reflected in the marriage structure [see featured image]. We found that the proportion of single-ethnic marriages of the Ulchi is on average 51%. The greatest number of such marriages takes place in the village of Bulava. Marriages of Ulchi with Russians are in second place. Marriages with indigenous peoples of the Far East, Nanais, Nivkhs, Evenks, and others, are in third place. Thus, almost half of the Ulchi marriages are with representatives of other nationalities. Such a significant level of interethnic mixing makes it possible to talk about intense processes of assimilation of this indigenous people and puts to the forefront the problem of loss of the unique gene pool of the Ulchi.

    Haplogroup C (its branch M48) was genotyped for its five subbranches with markers M86, B470, F13686, B93, and the marker at position 16645386 (GRCh37), which was found by our team for the first time. Variant B93 is rare in the Ulchi, and 14 samples (that is, more than a quarter of the entire gene pool of the Ulchi, Fig. 2) belong to M86 and its subvariants. Therefore, we genotyped STR markers of C-M86 carriers for the Ulchi and neighboring Amur populations and analyzed the relationships of detected haplotypes on the phylogenetic network (Fig. 3, STR haplotypes are available from authors upon request).

    (…) On the network, different clusters are associated with different populations: most Mongols belong to F13686, all Evenks of the Amur River region with this haplogroup form a subcluster within F13686, and part of Upper Nanais is the basis of cluster B470.

    ulchi-y-chromosome
    Frequencies of haplogroups of Y chromosome in the Ulchi population. The nomenclature of haplogroups is given according to [9]. Markers that are not in bold type were not typed, but are ancestral for these nodes.

    An estimate of the age of the entire haplogroup C-F12355 obtained from the data of genome-wide sequencing of seven specimens is 2400 ± 500 years (O.P. Balanovsky, unpublished data). That is, the common ancestor of all the studied representatives of various peoples with this haplogroup lived not so long ago, the first millennium BC. The formation time of cluster F13686 is somewhat later: 1990 ± 600 years.

    (…) obvious traces of the interaction of the gene pool of the Ulchi with neighboring and remote peoples of the Far East and Central Asia in the time range of the last one to three thousand years were revealed. This shows that the results of work [4] on the similarity of the gene pool of the ancient (age of 7500 years) Neolithic genomes of the Amur River region to the Ulchi probably indicate not the uniqueness of the Ulchi, but the fact that this ancient gene pool was preserved in a vast circle of populations of the Far East interwoven with gene flows both with each other and, to a lesser extent, with populations of Central Asia.

    The expansion of C2b1a2a-M86 (among many basal C2-M217 samples) is thus possibly associated with the spread of Tungusic, which puts C2b1a at the root of the Micro-Altaic expansion, with a formation date ca. 12700 BC, TMRCA 12500 BC (and not only Mongolian). This shows that Micro-Altaic is connected with a local population which shows a clear continuity since at least 3500 BC. This, however, tells us little about the origin of the language.

    See also the recent ISBA presentation on the Houtaomuga site, Neolithic transition in Northeast Asia; and also Bronze Age population dynamics and rise of dairy pastoralism in Mongolia, Impact of colonization in north-eastern Siberia

    That leaves the ancestral N lineages found among Far East Asians as Palaeo-Siberian in origin, and their late expansions to the west not particularly linked with any of the known Palaeo-Siberian ethnolinguistic groups, let alone a supposed “Uralo-Altaic” language…

    Related