CHAPTER 3
The Language Question:
Tajik, Uzbek, Mughat In-Group Speech, and the Secret Lexicon
From the proposed monograph:
The Lyuli (Mughat) of Uzbekistan: Language, Ritual, Memory and the Search for South Asian Connections
Dr. Manish Kumar C. Mishra
3.1 Introduction: why language is the decisive archive
Among all the possible forms of evidence concerning Lyuli/Mughat history, language has the greatest potential to move the discussion beyond resemblance, legend, and ethnographic stereotype. Dress can change rapidly; occupations can be shared by unrelated communities; ritual forms can diffuse across religious and regional boundaries; and ethnonyms can be imposed from outside. A structured linguistic system, by contrast, can preserve traces of historical contact and descent in ways that can be tested comparatively.
The linguistic situation of the Lyuli/Mughat is nevertheless unusually difficult. The community lives within a multilingual Central Asian environment in which Tajik and Uzbek are major languages of everyday interaction. At the same time, scholarship has repeatedly referred to a restricted internal repertoire—variously described as Mugat language, secret language, argot, or community speech. These descriptions are not automatically equivalent. The central task of this chapter is therefore conceptual before it is etymological: what exactly is the linguistic object that later chapters will compare with South Asian languages?
A recent Uzbek dissertation explicitly describes the “Mugat language” as the mutual communication language of the Lyuli and discusses its present state and preservation. Earlier comparative linguistic scholarship, however, classifies Jugi/Mugat of the Hissar Valley as Tajik-based, and describes Luli/Multani of the Fergana area as an argot with a Tajik base. The apparent contradiction is productive rather than fatal. It may reflect different meanings of the word language, different communities and regions, or different layers of the same repertoire.
3.2 Four terms that must not be confused
The first distinction is between language, ethnolect, register, and argot. A language is a relatively autonomous linguistic system with grammar and lexicon capable of ordinary communication across domains. An ethnolect is a socially or ethnically associated variety of a wider language. A register is a context-dependent mode of speech used for particular purposes. An argot is a specialised vocabulary or speech practice often associated with occupational, social, or secrecy functions.
The label “secret language” can cover more than one of these realities. A community may speak ordinary Tajik grammar while replacing selected content words with special lexical items. That would differ fundamentally from a complete inherited language whose grammar, pronouns, numerals, verbs, and syntax differ from Tajik. Between these poles are mixed systems: a Tajik grammatical frame with a large non-Tajik lexicon; an ethnolect containing distinctive phonology and morphology; or a repertoire in which ordinary speech and restricted argot alternate according to audience.
Until primary recordings and paradigms are analysed, this monograph will use the neutral expression “Mughat in-group speech” for the restricted repertoire and reserve “Mughat language” for contexts in which the source itself uses that designation.
3.3 The public linguistic ecology: Tajik and Uzbek
The Lyuli/Mughat do not exist outside the linguistic ecology of Uzbekistan and Tajikistan. Tajik, an Iranian language closely related to Persian, has historically played a central role in many Mughat communities, while Uzbek, a Turkic language, is indispensable in much of Uzbekistan’s public life. Russian may also enter the repertoire through education, administration, migration, media, and Soviet/post-Soviet experience.
This multilingualism has methodological consequences. A word found in Mughat speech may have entered from Tajik, Uzbek, Russian, Persianate religious vocabulary, or another regional argot. Even an apparently South Asian-looking word must therefore be tested against all plausible Central Asian sources before it is classified as Indic.
The distinction between language of home, language of neighbourhood, language of school, language of work, and language used when outsiders are present should be documented in future fieldwork. “What language do you speak?” is too crude a question for a community with layered repertoires.
3.4 What Encyclopaedia Iranica tells us
Gernot L. Windfuhr’s survey in Encyclopaedia Iranica provides one of the clearest comparative classifications. For Central Asia it identifies Jugi, with the endonym Mugat, in the Hissar Valley as Tajik-based and relates it to Jogi populations of Afghanistan and Jugi of northern Iran. It separately records Luli of the Fergana area with the endonym Multani and describes their argot as Tajik-based. It also distinguishes Kara-Luli, whose endonym is reported as Hindustani, and an Afghon group with an Indic-based dialect.
The importance of this classification lies in its heterogeneity. The historical umbrella of “Gypsy dialects” in Central Asia includes Tajik-based, Afghan-Persian-based, Persian-based, and Indic-based varieties. Therefore the mere presence of a community within the same social category cannot establish linguistic genealogy.
For the Mughat question, the Iranica classification makes one proposition especially important: the structural base may be Tajik even where the lexicon contains elements of other origins. This is precisely the kind of situation in which lexical stratigraphy becomes necessary.
3.5 The Uzbek dissertation and the expression “Mugat language”
A dissertation associated with the Academy of Sciences of Uzbekistan provides a valuable contemporary perspective. Its English summary states that the scientific findings concern “the present state of the Mugat language which is the mutual communication language of the lyuli and its preservation.” The dissertation also links language with customs, rituals, and cultural heritage.
This formulation should be taken seriously because it reflects recent Uzbek scholarship and may correspond closely to community usage. Yet it should not be used to override the earlier structural classification without examining the underlying data. The expression “mutual communication language” may refer to an in-group ethnolect or argot that functions socially as the community’s own language even if much of its grammar is Tajik-based.
Sociolinguistic identity and genealogical classification answer different questions. A community can legitimately call a repertoire “our language” while a historical linguist describes its grammatical matrix as Tajik. Both statements can be true at the same time.
3.6 Secret speech as a social institution
Restricted speech is not only a linguistic phenomenon; it is a social institution. Its function may include privacy in the presence of outsiders, reinforcement of group solidarity, occupational secrecy, joking, taboo management, protection of commercial information, or the marking of boundaries between insiders and non-members.
This matters for etymology. A vocabulary deliberately designed to be opaque is unusually open to lexical replacement and borrowing. Speakers may select unfamiliar words precisely because outsiders do not understand them. Consequently, the lexicon of an argot can be historically cosmopolitan even when its grammar belongs overwhelmingly to one language.
The social function of Mughat in-group speech must therefore be recorded alongside every lexical item: Who uses it? With whom? In which situations? Do children understand it? Are women and men equally fluent? Does it differ between generations? Is it used in complete sentences or only by inserting special words into Tajik or Uzbek sentences? These questions are essential to deciding whether the object is a language, ethnolect, register, or argot.
3.7 Loterāʾi, Abdoltili, and the wider Persianate argot network
The wider Persianate world contains a long history of specialised argots. Encyclopaedia Iranica’s treatment of Loterāʾi is particularly important because it compares argots of Jugi, Luli, Chistoni, Kavoli, musicians, mendicant dervishes, and other mobile or occupational groups. It defines Abdoltili as the “language of itinerants” associated with Uzbek-speaking artisans and musicians, preachers, and qalandars.
This evidence complicates any attempt to identify every non-Tajik Mughat word as ancestral. Lexical items could circulate across professional and mobile networks without the speakers sharing a single ancestry. The same source nevertheless records explicitly Indic-derived words in related Iranian and Central Asian argots, demonstrating that an Indic layer is not imaginary. The analytical problem is to distinguish inherited material from borrowed argot material.
The existence of an argot network therefore weakens simplistic etymology but strengthens the case for systematic comparison.
3.8 A small verified sample from the comparative argot literature
The Loterāʾi survey gives a useful demonstration of how mixed such vocabularies can be. In the Djougi material of northern Iran—historically compared with Jugi of Tajikistan—it lists forms classified as Indic alongside Arabic and Iranian items. Examples of Indic-classified vocabulary include forms glossed as ‘man’, ‘woman/wife’, ‘iron’, ‘water’, ‘big’, ‘goat’, ‘horse’, and ‘donkey’. The same corpus contains Arabic and Iranian words.
These examples must not be transferred automatically into modern Uzbek Lyuli speech. They belong to a documented comparative argot corpus and are introduced here only to demonstrate the method and the possibility of stratification. The later lexical chapter will include a word only when its community, locality, source, transcription, and gloss can be identified.
This rule is essential: a word attested in Djougi, Loterāʾi, Domari, Romani, or Parya is not a Lyuli/Mughat word merely because the communities have been compared historically.
3.9 The concept of lexical stratigraphy
Lexical stratigraphy treats vocabulary as layers deposited through different periods of contact. For Mughat in-group speech, at least five layers must be tested: Tajik/Persian; Uzbek and other Turkic material; regional professional or religious argot; possible Indic material; and later Russian or modern borrowings. A sixth category—unresolved—must remain available for forms whose origin cannot be established.
The central question is not how many words can be made to resemble Hindi. It is whether the putative Indic forms cluster in historically conservative semantic domains and display regular sound correspondences. Basic verbs, kinship terms, body parts, numerals, pronouns, common animals, and elementary material culture generally carry more genealogical weight than occupational code words that are easily borrowed.
The book will therefore assign both an etymological category and an evidentiary confidence level to every analysed item.
3.10 Grammar is more important than attractive word matches
Popular linguistic comparisons often begin and end with similar-sounding words. Historical linguistics requires more. If Mughat in-group speech possesses Tajik word order, Tajik verb morphology, Tajik case/prepositional structures, and Tajik pronouns, while substituting selected secret nouns, the system is fundamentally different from an Indo-Aryan language such as Parya.
Conversely, if primary material reveals non-Tajik grammatical morphology, inherited pronouns, numeral systems, verb paradigms, or productive suffixes that correspond regularly with Indo-Aryan, the historical implications would be much stronger.
For this reason, future field elicitation must collect sentences and paradigms, not merely vocabulary lists. A hundred isolated words cannot substitute for ten carefully recorded grammatical constructions.
3.11 Parya/Tajuzbeki as the positive control
Parya provides a crucial control because it represents a demonstrably Indo-Aryan language preserved in the Central Asian environment. Bholanath Tiwari’s Tajuzbeki study, together with Oransky’s Parya research, allows the book to observe what genuine Indo-Aryan structural survival looks like after prolonged contact with Tajik and Uzbek.
The comparison should therefore operate at two levels. First, lexical: do Mughat in-group forms correspond systematically with Parya and north-western Indo-Aryan vocabulary? Second, grammatical: does Mughat preserve any structures comparable with Parya that cannot be explained through Tajik?
If the answer is lexical but not grammatical, the likely history may involve borrowing or residual vocabulary after language shift. If both lexicon and grammar align systematically, a much stronger genealogical argument would become possible.
3.12 The India test: Hindi alone is not enough
A serious search for South Asian connections cannot compare Mughat only with Standard Hindi. If the Multoni/Multani clue is historically meaningful, the most relevant comparison may lie farther northwest. Punjabi, Saraiki, Sindhi, Rajasthani, Haryanvi, and related Indo-Aryan varieties must therefore be included, alongside Parya, Romani, and Domari where appropriate.
Geography matters because cognates can reveal subgrouping. A form shared broadly across Indo-Aryan proves less about a precise homeland than a distinctive cluster concentrated in one region. At the same time, Persian loans in Hindi, Punjabi, or Urdu must not be mistaken for Indic inheritance when the same word could have reached Mughat directly through Persian/Tajik.
Every proposed Indian connection must therefore pass a contact-filter test: could this word be explained more economically through Persian, Tajik, Uzbek, Arabic, or regional argot? Only after those alternatives are excluded should an Indic etymology be preferred.
3.13 Proposed lexical database for the monograph
The final book will build a transparent lexical database. Each entry should contain the original Mughat form exactly as recorded, phonetic/transliteration information, English gloss, locality, speaker or source where available, date, grammatical category, example sentence, Tajik comparison, Uzbek comparison, Persian/Dari comparison, Abdoltili or other argot parallels, Parya/Tajuzbeki comparison, South Asian comparisons, proposed etymology, and confidence level.
The database will distinguish ‘probable inherited Indic’, ‘possible Indic’, ‘Iranian’, ‘Turkic’, ‘regional argot’, ‘Russian/modern’, and ‘uncertain’. The category ‘probable inherited Indic’ will require more than phonetic resemblance: semantic fit, plausible sound correspondence, and absence of a more immediate contact-language explanation will all be required.
This format makes the argument reproducible. Readers will be able to see not only the conclusion but the evidence behind each lexical decision.
3.14 Fieldwork protocol: how the language should be recorded
If fieldwork becomes possible in Uzbekistan, the linguistic component should be designed ethically and scientifically. Participation must be voluntary and based on informed consent. Because an in-group repertoire may be intentionally restricted, researchers must not pressure speakers to disclose words they regard as private, sacred, or socially sensitive.
Elicitation should begin with sociolinguistic questions before vocabulary: languages used with parents, spouse, children, neighbours, school, market, officials, and outsiders; self-name for the internal speech; age at which it is learned; and whether younger speakers retain it. Recordings should include natural conversation where consent permits, followed by controlled elicitation of basic vocabulary and grammar.
Regional sampling is essential. Samarkand, Bukhara, Surkhandarya, Fergana, and communities in Tajikistan may not share identical repertoires. Variation itself may preserve migration history.
3.15 Language vitality and intergenerational transmission
The contemporary Uzbek dissertation’s emphasis on preservation raises another major question: is Mughat in-group speech endangered? Language vitality cannot be inferred merely from the existence of older speakers. The decisive issue is intergenerational transmission.
The study should examine whether children understand and actively use the restricted repertoire, whether vocabulary is shrinking, whether Uzbek is replacing Tajik in some regions, and whether schooling, urbanisation, migration, smartphones, and social media are changing the communicative function of secret speech.
A paradox may emerge: a repertoire originally valued because outsiders could not understand it may become less useful as social boundaries change, while simultaneously becoming more important as a symbol of cultural identity. Such a transition from functional secrecy to emblematic heritage would be sociolinguistically significant.
3.16 What the present evidence allows us to conclude
Three conclusions can already be stated with reasonable confidence. First, the Lyuli/Mughat linguistic repertoire is multilingual and cannot be described adequately by a single label. Second, authoritative comparative scholarship describes major Jugi/Mugat and Luli/Multani varieties as Tajik-based, while also situating them within a wider network of specialised argots containing vocabulary from multiple origins. Third, contemporary Uzbek scholarship recognises a socially meaningful “Mugat language” used for mutual communication and treats its preservation as part of Lyuli cultural heritage.
These propositions are compatible if linguistic structure and social identity are separated. A Tajik-based ethnolect or argot can function as an in-group language. The unresolved question is historical: how much of its distinctive vocabulary is inherited from an earlier South Asian language, how much entered through Persianate argot networks, and how much reflects later Central Asian contact?
That question cannot be answered by terminology alone. It requires the corpus.
3.17 Conclusion: from the language question to the lexical test
This chapter has established the methodological foundation for the linguistic core of the book. The phrase “Mughat language” should neither be dismissed nor accepted uncritically as proof of an autonomous genealogical language. Earlier scholarship points to a Tajik grammatical base in important Jugi/Mugat and Luli/Multani varieties; contemporary Uzbek research emphasises the community function and preservation of Mugat speech; and comparative argot scholarship demonstrates that specialised vocabularies in the Persianate world can combine Indic, Iranian, Arabic, Turkic, and other elements.
The strongest India connection will therefore not come from the existence of secrecy itself. It will come, if at all, from a demonstrable pattern inside the lexicon and grammar.
Chapter 4 will undertake that test. It will reconstruct the available Mughat/Lyuli lexical corpus source by source, separate locality-specific data, compare each form with Tajik, Persian, Uzbek, regional argots, Parya/Tajuzbeki, and relevant Indo-Aryan languages, and assign an explicit confidence level to every proposed South Asian correspondence. The aim will not be to prove an Indian origin in advance, but to discover exactly what the linguistic evidence permits us to say.
Table 3.1. Working distinction among the linguistic categories
Table 3.2. Comparative classification relevant to the Mughat question
Figure 3.1. Proposed linguistic stratigraphy
CURRENT PUBLIC REPERTOIRE
Tajik ↔ Uzbek ↔ Russian (context-dependent)
↓ overlapping with ↓
MUGHAT IN-GROUP SPEECH
Possible lexical layers:
Tajik/Persian
+ Uzbek/Turkic
+ Persianate/Central Asian argot (including Abdoltili networks)
+ possible Indic residue
+ later Russian/modern vocabulary
+ unresolved forms
Table 3.3. Template for the Chapter 4 lexical database
References cited in Chapter 3
Windfuhr, Gernot L. 2002. ‘Gypsy ii. Gypsy Dialects’. Encyclopaedia Iranica, XI/4, pp. 415–421.
Windfuhr, Gernot L. 2012. ‘Loterāʾi’. Encyclopaedia Iranica. Survey of Persianate secret languages and argots, including Jugi/Luli-related material and Abdoltili.
Marushiakova, Elena, and Vesselin Popov. 2016. Gypsies of Central Asia and the Caucasus. Cham: Palgrave Macmillan.
Oranskiĭ, I. M. 1961. Studies cited in comparative classifications of Central Asian Jugi/Mugat and related groups.
Oranskiĭ, I. M. 1983. Comparative ethnolinguistic work cited for Mugat/Jugi and Central Asian argot material, especially the corpus discussed in later scholarship.
Tiwari, Bholanath. 1970. Tajuzbeki: Soviet Sangh mein Boli Jane Vali Hindi Boli: Aitihasik aur Tulnatmak Adhyayan tatha Sankshipt Shabdkosh. Delhi: National Publishing House.
Uzbekistan Academy of Sciences-associated dissertation, English summary, 150 pp. Contemporary study describing the Mugat language as the mutual communication language of the Lyuli and discussing its preservation. ZiyoNET digital copy.
Source-critical note
This chapter does not present a fabricated Mughat word list. The few comparative lexical examples discussed in the prose derive from the Encyclopaedia Iranica Loterāʾi survey and are explicitly identified there as Djougi/comparative argot material, not automatically as modern Uzbek Mughat vocabulary. Chapter 4 should proceed only from recoverable primary or clearly attributed lexical corpora. The contemporary Uzbek dissertation is used here for its explicit statement about the social status and preservation of “Mugat language”; its detailed linguistic data must be extracted and checked before being used for word-level etymology.
No comments:
Post a Comment
Share Your Views on this..