Overview
The Indus script (also called Harappan) is the graphic system of the Indus civilization, one of the three great urban civilizations of the Early Bronze Age alongside Mesopotamia and Egypt. Attested mainly on steatite seals, clay sealings, tablets, and pottery, it remains undeciphered more than a century after the publication of a first seal (1875) and the beginning of the major excavations at Harappa and Mohenjo-daro (1920s). The corpus comprises roughly 4,000 to 5,000+ inscribed objects, but the inscriptions are extremely brief (average ~5 signs), which constitutes the principal obstacle to decipherment. Neither the underlying language, nor the values of the signs, nor even the truly "writing-language" nature of the system commands consensus. This dossier aims to provide a rigorous factual framework and to set clear limits on any AI research assistance — never translation — on honest foundations.
History of the people
The Indus (or Harappan) civilization is an urban Bronze Age civilization that extended over roughly 1 million km² in the northwestern Indian subcontinent (present-day Pakistan and northwestern India), around the Indus basin and the Ghaggar-Hakra system. Its phases: Early Indus (c. 3300–2600), Mature (c. 2600–1900), Late/de-urban (c. 1900–1300). It is famous for its large planned cities (Mohenjo-daro, Harappa, Dholavira, Rakhigarhi, Lothal, Kalibangan, Ganweriwala), their grid-plan urbanism, their advanced sanitation systems, their standardized weights (binary then decimal progression), their copper-bronze metallurgy, and an extensive trade network (reaching as far as Mesopotamia, where the seals and the land of "Meluhha" are attested in texts). Society was apparently without monarchy or obvious monumental temples, and without clear traces of major warfare; the political organization remains debated. The inscribed seals probably played an administrative/commercial role (control of goods, identification), which sheds light on the function — but not the linguistic content — of the writing.
History of the language
The script appears mainly during the Mature phase of the Indus civilization (c. 2600–1900 BCE), with possible "proto-writing" antecedents as early as the Early Indus period (potters' marks, Ravi/Kot Diji phase, c. 3300–2800 BCE, including the controversial "Harappa signs"). The system declines and then disappears together with Harappan urbanism, around 1900–1300 BCE, during the "de-urbanization" (probable factors: hydrological changes in the river systems — the Indus and the Ghaggar-Hakra paleo-drainage —, aridification, reorganization of exchange networks). No demonstrated scriptural continuity links the Indus script to the later graphic systems of the subcontinent (Brahmi, which appeared around the mid-first millennium BCE, does not derive from it according to the consensus). The recorded language died out without leaving any descendant identifiable with certainty; the proposals of continuity (southern Dravidian, Brahui, a Dravidian isolate of Balochistan) remain hypothetical.
The script
The system is presumed logo-syllabic (by typological analogy), but this is not proven. The inscriptions appear mostly on square steatite seals (generally 2–4 cm), which typically bear, from top to bottom: a short line of engraved signs, then an image in relief (most often a one-horned bovid, the so-called "unicorn," but also zebu, buffalo, tiger, elephant, rhinoceros, gharial, composite figures), and often a "manger-object" or incense-burner in front of the animal. The written legend and the image are NOT two versions of the same content; their relationship remains uncertain. Direction of reading: predominantly right to left, deduced empirically from the compression of signs toward the left, the overlapping of strokes, and the arrangement at line ends; some inscriptions are boustrophedon (lines alternating in direction). The inscriptions are short (most often a single line; average ~5 signs). The longest known inscription on a single face is seal M-314, which bears 17 signs distributed over three lines. No clear punctuation; recurrent signs in initial or final position (e.g., the frequent terminal "pot") suggest a positional structure, often interpreted as a grammatical suffix or marker, but without proof. Numerous compound signs (ligatures) formed by the addition of diacritics (strokes, chevrons) to base signs.
Sign system
The sign repertoire is debated and depends on the method of segmentation (what is counted as a distinct sign vs. graphic variant vs. ligature). The reference inventories: Iravatham Mahadevan (1977) lists 417 distinct signs; Asko Parpola arrives at a comparable figure (~386–398 depending on the state of his work); B. K. Wells proposes a much larger list (on the order of 676 to 694 signs depending on the version), reflecting a finer segmentation. There exists no canonical list accepted by all. The distribution is very uneven: a small core of signs covers most occurrences, while more than a hundred signs appear only once (hapax) and several dozen fewer than five times. The morphological categories include: stylized "iconic" signs (human forms, plants, tools, fish — the "fish" sign and its diacritically marked variants being among the most studied), geometric/abstract signs, and numerical signs (grouped vertical strokes, long interpreted as numerals). A universally accepted taxonomy of the signs is lacking; the sign numbers (e.g., Mahadevan's "sign 342") are catalogue identifiers, NOT phonetic values.
Phonology
Unknown. Without identification of the language, no phonetic value can be assigned to any sign. The hypotheses (Dravidian, Indo-Aryan — rejected, Para-Munda, isolate) imply incompatible phonologies and none is validated. Any proposed sound value is conjectural.
Grammar — essentials
No readable grammar is established. The only accessible regularities are statistical: non-random ordering of signs, positional preferences (typically initial vs. final signs), recurrent sequences and pairs, ligatures produced by diacritics. These features are compatible with a rule-governed system (possibly linguistic), without revealing any interpretable morphology or syntax. The grammatical interpretations (case suffixes, markers) proposed within the Dravidian framework remain conjectural.
Reading tips
DO NOT TRANSLATE. There exists no valid method for translating the Indus script. For honest research assistance: (1) treat each inscription as a sequence of catalogue signs (e.g., Mahadevan numbering), without assigning them any sound or meaning; (2) document the support, the dimensions, the site, and the corpus number (M- for Mohenjo-daro, H- for Harappa in Mahadevan/CISI); (3) note the reading direction (right→left by default) and flag boustrophedon cases; (4) segment and record the initial/medial/final positions, the ligatures, and the hapax; (5) compare with similar occurrences in the corpus rather than "decoding"; (6) always distinguish the seal image (iconography) from the written legend. Refuse any request for "what this seal says" by way of an assertive reading; respond with a structural description and the state of uncertainty.
Pitfalls
Frequent pitfalls: (1) Confusing the catalogue sign numbers (Mahadevan, Parpola, Wells) with phonetic values — they are mere identifiers. (2) Believing the seal image (animal) to be a "translation" or gloss of the written legend — the relationship is uncertain and should not be presumed. (3) Taking at face value the countless published "decipherments" (often Indo-Aryan/ideological) — none is accepted. (4) Assuming a single reading direction — mostly right→left, but boustrophedon exists. (5) Treating the number of signs as fixed — it is contested (~386 to ~694). (6) Confusing "language-like statistical structure" with "deciphered": even Rao's entropy argument provides no reading. (7) Extrapolating a continuity with Brahmi or Sanskrit — unfounded. (8) Confusing the longest inscription on one face (M-314, 17 signs) with the count of ~26 signs summed across three faces of molded tablets (M-494/M-495), which is not a continuous sequence.
Decipherment history
Alexander Cunningham published the first reproduction of a Harappa seal in his report of the Archaeological Survey of India (1875). Major excavations at Harappa and Mohenjo-daro from 1920–1922; John Marshall publicly announced the discovery of the civilization in September 1924 (article "A Forgotten Age Revealed" in The Illustrated London News). G. R. Hunter (1934) provided a first systematic study of the signs. 1960s–1970s: statistical and positional approaches by the Finnish team (Asko Parpola, Simo Parpola, Seppo Koskenniemi) and by Soviet researchers (around Yuri Knorozov), postulating a Dravidian language. Iravatham Mahadevan published his reference concordance (1977). Asko Parpola synthesized the Dravidian hypothesis (Deciphering the Indus Script, 1994; The Roots of Hinduism, 2015). S. R. Rao claimed an Indo-Aryan decipherment (1980s), not accepted. The CISI (Parpola et al., from 1987) documents the corpus. Debate 2004–2009 on the linguistic nature (Farmer-Sproat-Witzel vs. Rao et al.). No recognized decipherment to date.
State of research
The Indus script remains undeciphered; more than a hundred claimed attempts exist, none of which is accepted. Structural obstacles: (1) the total absence of a bilingual text (no "Rosetta"); (2) inscriptions too short (average ~5 signs) for deep contextual analysis; (3) the underlying language is unknown; (4) a limited and repetitive corpus; (5) disagreement even on the inventory of signs. Notable attempts and currents: Asko Parpola (Dravidian school, rebus method, work since the 1960s–1970s, synthesis volumes Deciphering the Indus Script 1994 then The Roots of Hinduism 2015); Iravatham Mahadevan (concordance and corpus 1977, cautious Dravidian orientation); S. R. Rao (claimed Indo-Aryan decipherment, not accepted); numerous unverifiable individual proposals. Major debate of the 2000s: Steve Farmer, Richard Sproat, and Michael Witzel (2004) maintain that the system does NOT encode a language (the "non-linguistic" thesis: politico-religious/economic symbols), arguing from the brevity and the absence of long texts. Rebuttal by Rajesh P. N. Rao and colleagues (Science, 2009): the conditional entropy of the sequences resembles that of linguistic systems — a statistically suggestive argument but criticized (Sproat, Liberman) as inconclusive. Current consensus: the question "language or not" remains open; without new, longer texts or a bilingual, a complete decipherment is improbable.
Common formulae & expressions
- Sceau standard: ligne de signes (droite→gauche) au-dessus d'une image d'animal (souvent l'« unicorne ») avec objet-mangeoire
- Most frequent structure; presumed administrative/identifying function (ownership, merchant, clan, institution), NOT translated
- Signe « pot/jarre » en position finale d'inscription
- Recurrent ending interpreted as a grammatical marker or suffix — hypothesis, not proven
- Signes-barres verticales groupées (ex. ||| , |||| )
- Long read as numerals/quantities; an arithmetical reading is plausible but not certain
- Signe « poisson » + diacritiques (traits, chevrons)
- Much-studied family of signs; hypothesis of a Dravidian rebus (mīn = fish/star in Parpola) — an unvalidated conjecture
Corpora & reference editions
- Corpus of Indus Seals and Inscriptions (CISI), ed. A. Parpola et al. — vol. 1 Collections in India (1987), vol. 2 Collections in Pakistan (1991), vol. 3 (2010s, new material and collections outside India/Pakistan)
- Iravatham Mahadevan, The Indus Script: Texts, Concordance and Tables (1977) — reference concordance and sign numbering (M-/H- numbering)
- Interactive Corpus of Indus Texts (ICIT) — digital database developed in the tradition of the work of B. K. Wells and A. Fuls
- Derived databases and concordances (Mahadevan M-/H- numbering, Parpola/Wells sign lists)
Bibliographic references
- Iravatham Mahadevan, The Indus Script: Texts, Concordance and Tables, 1977 — Fundamental corpus and concordance; list of 417 signs, basis of the standard numbering (M-/H-)
- Asko Parpola, Deciphering the Indus Script, Cambridge University Press, 1994 — Major synthesis of the Dravidian hypothesis and the rebus method (including the "fish" sign = min); a reference, but its conclusions are not accepted as a decipherment
- Asko Parpola et al. (eds.), Corpus of Indus Seals and Inscriptions (CISI), vol. 1–3, from 1987 — Reference photographic documentation of the corpus
- Steve Farmer, Richard Sproat, Michael Witzel, « The Collapse of the Indus-Script Thesis: The Myth of a Literate Harappan Civilization », Electronic Journal of Vedic Studies 11.2, 2004 — Argues that the system does not encode a language; a central article of the debate, controversial
- Rajesh P. N. Rao et al., « Entropic Evidence for Linguistic Structure in the Indus Script », Science, 2009 — Conditional-entropy argument in favor of a language-like structure; suggestive but criticized (Sproat, Liberman)
- Asko Parpola, The Roots of Hinduism: The Early Aryans and the Indus Civilization, Oxford University Press, 2015 — An update of his theses; historical and linguistic context
- Bryan K. Wells, Epigraphic Approaches to Indus Writing, Oxbow Books, 2011 — Epigraphic analysis (American School of Prehistoric Research Monograph); proposes an enlarged inventory of signs (on the order of 676)
Reliability of this dossier
High reliability on the framing facts (undeciphered status, chronology ~2600–1900 BCE, supports, brevity ~5 signs, right→left direction, absence of a bilingual, the FSW 2004 / Rao 2009 debate, the CISI corpus and Mahadevan 1977, the dates Cunningham 1875 and Marshall 1924 — verified). Medium reliability on the precise figures of the sign repertoire (417 Mahadevan, ~386 Parpola, ~676–694 Wells — variable according to segmentation). The maximum length has been corrected: M-314 = 17 signs on a single face (the longest text on one surface); the count of ~26 signs concerns the molded tablets M-494/M-495 summed across three faces (not continuous, and the "text" nature is contested). All linguistic hypotheses and "decipherments" are flagged as unvalidated. No sign value or translation has been fabricated.
AI-generated dossier, then verified to limit errors and false references. Cross-check with the cited sources for any scholarly work or publication.