Transcription
Y'g welcome to the podcast. So nice to finally meet you and see your face.
Absolutely. Very happy to be here, Abhijit, and thanks for the opportunity.
My pleasure. My absolute pleasure and honor. So, uh, can I address you by your real name?
Yes, I just R it. Okay.
So, uh, so we're going to discuss the work that you have done and you are still doing: the decipherment of the so-called Harappan or Sarasvati-Sindhu script.
Yes. Uh, and this is one of the central matters in the Aryan Invasion/Migration debate. The central thesis of the Aryan Invasion/Migration debate—the central claim—is twofold: firstly, that Hinduism is foreign to India; and secondly, that Sanskrit is foreign to India. The clinching proof, evidence this way or that way, will be the decipherment of the Sarasvati-Sindhu script. And that's what you've done, and you have demonstrated that it is Sanskrit.
Right. Yes. So yeah, let's discuss that. But before that, can we have something about your background? What's your background like, and how did you get into this entire matter?
Right. So, from an early age, during my engineering days—if you're also from a similar background, many attended lectures—we usually get bored, and we do something to pass the time. Some people brought crosswords; I used to bring cryptograms. I see. And it would be good because a crossword is generally about 15-20 minutes; a cryptogram is a good 1 hour. It would fill the lecture. When I was doing the cryptograms, I did not completely know the math and everything behind it, but I had a good sense of how to solve unknown scripts. And eventually, when I ran into the Indus script problem, I started looking at the claimed decipherments of other scholars, and I realized that everything is ad hoc; everything is mostly, "I've looked at the script, I looked at the shape, I think it is an X, I think it is a Y," and that didn't make sense to me because what logic did they use? I didn't get it.
Yeah, it's mostly, "I think it looks like this," and you know, I'm like, "Madan," yeah, like some goddess is telling me how to do it. So it didn't make sense to me, and it—and to the—I think to the general crowd—it always seemed odd that there are so many claims, and not only is it just unfalsifiable, it's also that everyone is claiming a different thing for every symbol.
So, um, I looked at a bunch of decipherments, and then I realized that the problem can be modeled as a cryptogram. And then I, you know, I thought I'll try different languages. In those days, we didn't have—or I just—I didn't find a Tamil dictionary. So in those days, remember that the idea was that Sanskrit arrived in India sometime around 1500 BC to 1200 BC, like Michael Witzel said 1200 BC because 1700 BC there was no early Sanskrit or whatever. So he said 1200. So then, since the Indus Valley is older, it must be something else. So I thought, okay, we'll try Tamil. There's no Tamil downloadable dictionary; other dictionaries existed, but I found, okay, Sanskrit dictionary. So I said, okay, I'll try it. And there were also other scholars who had done work which I was trying to falsify, approve by falsification. And then, so I realized when doing that that this can be modeled as a cryptogram, and I also understood where other people have gone wrong. And after doing a few rounds, I kind of came to the conclusion that this sort of looks like ABA ABA because of the way the regular—so for the audience who may not know, regular expression is it's a way of specifying how the different letters are adjacent to each other.
Okay. Okay. And you can use that specification to extract words that match that specification. You can say the first letter is the same as the third letter, and the second letter is the same as the fifth letter, and there are a total of six letters in this pattern; find me all the words that match it. Okay. So by doing that, I realized that—so that will capture any kind of script except logographic. So if it's syllabic, if it's, you know, upj or anything, you capture it. But so to me, I saw mostly ABAB.
Mhm. So then, after—after a while, after solving about 60-80% of the signs—I kind of thought, okay, I need also a way to prove that what I've done is right, just even to myself before I, you know, go public with this.
Yes. So I was looking around for cryptogram and the math behind it, and I found Shannon's paper. So I'm from an information theory background; I should know about it, but most of us have forgotten, going out of college, all of this—papers, forgot the theory, right? The applications, forgot, or you know, you have in the back of your mind whatever.
Right. That's right. So I rediscovered it, and I saw his proof and his specifications on information theory and how I can mathematically understand, first of myself that I'd done it right, second, I can convince experts, and third, I can convince ordinary people with some level of math education—not expert, but people understand—understand logarithm, factorial.
Yes. High school level math. Yeah. So I was able to do that, and I wrote the—from the paper I got some good criticism. Every round of criticism, I made improvements to the paper. Okay. And eventually, I was able to read all the inscriptions as Sanskrit—when it's grammatically correct, close to Paninian grammar Sanskrit. Okay. And that made me bold enough to announce it.
Okay. Okay. So on Twitter, I was doing, you know, when I started my Twitter account, I was just using it as a kind of database to do my investigations. A lot of the threads go nowhere; like they don't end up in—eventually end up in the paper, but some of the good stuff I recorded—the actual progress. Okay. So that is how I solved it, and um, after that, I had to pro—you know, once you cross a certain threshold, a certain number of characters, then in a non-logographic script, there's always a limit of text. If you read beyond then there's only a single possible solution. So I wanted to calculate that and cross it, which is what I did.
Okay. So when did you begin this process? How many years ago?
Sometime in Co—I think it was either 2020 or 2021. There was nothing to do, nowhere to go, no movies, no—how much Netflix are you going to watch? Exactly afterwards, you—mentally you start to get, you know, um, you need some outlet, you need a creative outlet to occupy your mind. So I was looking for a very hard problem that—'cause the lockdown was going for months, so I thought this could be years, so let's find something I can occupy myself for years, and I came upon this problem. I started working on it, and uh, so yeah, I think 2021 or so I started working on it. In six months I got most of the values.
I see. And then it took me a few months to write the first version of the paper, then after criticism I improved the paper. I think I'm on like the fourth version. When people say, "This is not clear," you know, so I know I need to address that. Many people ask questions that are not, you know, it is from a point of ignorance, but it's a very good question because that is a thing most people would want to know. Exactly. I would incorporate that in my text, and over time the paper became easier and easier to read and understand.
Mhm. And uh, so yeah, I'm, you know, still working on making it better; like the incremental improvements are happening even today. So it's—it's still in the form of a preprint, right?
It's a preprint, and technically it's on Academia.edu, technically still in draft.
Yeah. So let's understand the problem. The problem is that we have this large corpus of—of—of text, yes, ancient text, which would date back roughly two—three—four—three—four—4,300—between 4 and 5,000 years.
4 and 5,000 years. And uh, so we have a whole corpus of text, and it's kind of available online, right?
Yes. Okay. So that's the—that's the subject matter, essentially. That's a data set that you had, and we don't know what the script represents, what language it encodes. There are claims that it is Tamil; there are claims that it is Sanskrit; there's a huge lingering controversy for a very long time.
Yes. And so that's—that's the problem that you have—that you took forth—now, what is the methodology that you use to solve it? What is cryptography, and what's a cryptogram, and how do you use that to solve this problem?
Right, right, right. So the story of cryptography begins—I mean, it's very old, but a nice way to describe it—maybe not 100% accurate, but an easy way to describe it—is that during armed conflict, generals or rulers wanted to send messages to their allies, but if the message is captured in between, then their plans would fail, or you know, they may substitute the message and send the wrong message to the allies and end up, you know, you would end up losing. So Julius Caesar created one of the earliest ciphers, it's called a Caesar cipher. So essentially, he would have a key, which would be like a number five or four or whatever, and he would shift the alphabet by that much. So if it is four, A would become E, B would become F, and so on. And if someone saw the message without knowing that number, they wouldn't be able to read it, right? But the receiver would be able to decipher it. So the Caesar cipher is essentially easily broken because there are only 25 keys.
Yeah. And you could just try all of them, and within 10-25 minutes, uh, you know, even for those days, you could actually crack it if you know the type of cipher. Yeah. And over time, people created more and more complex ciphers, and the cryptogram is one where every letter has a different mapped letter. So this has a large number of keys.
M. So it has actually 26 factorial keys.
26 factorial is a very huge number. You cannot try every combination. It would take you—I mean, nowadays with computers you can, but in those days it would take a very, very long time. And cryptogram was actually the standard even into World War II, of course, in—there were various variations of ciphers, like Vigenère cipher, and some of these that the World War II, you know, Germans had—that Enigma, all that stuff.
Yes. But cryptograms were actually pretty central because you didn't need a machine, and a spy could memorize the key.
Mhm. And he could go into the enemy land, like, you know, dust among them or whatever, and if there's a message he could read it without any specific—without any machines or any kind of—and he could write a reply with just ink and you know, in kind of paper. So cryptograms were the method still in use during World War II. Right. And when—after the—after World War II, the US intelligence wanted to understand what is—what are the limits of various encipherment methods? What is the limit, like how—how well can we be sure that the messages will not be correct? Because what they would find is messages would go through, but no matter how the—how complex the scheme was, some genius somewhere, you know, some German or some Russian would end up cracking on it. So they wanted a theoretical framework to understand how all these encipherment systems work.
M. So they—so that there's a mathematician, Claude Shannon. Shannon who worked on this; it was a classified paper because it was so important. And over time, they—they declassified parts of it bit by bit, and eventually I think after 20 years they got newer, you know, computers and they had different standards; they didn't need cryptograms anymore, so they declassified the whole thing. Okay. So in that paper, he describes the theoretical foundations of cryptography itself, uh, and uh, how to know how much you can encipher without the enemy, uh, you know, deciphering it. And one of the key things he discovered or he presented in the paper was that any cryptogram, any cipher, can be broken if you have sufficiently long text. Okay. And he found—he measured—he made it quantitative—then what is the limit beyond which the text can be deciphered without ambiguity? Okay, like it—it is absolutely—beyond this, and for English text in the Latin alphabet, that came out to be around 30—30 words or characters.
30 characters. Okay. Okay. Using cryptogram. Now there are different uh ways you can compute it, so there's an upper bound. Okay. Uh, there may be for certain messages, uh, you know, you may be able to get it at the lower bound, but this is a measure of an upper bound. It says beyond 30 it's most definitely 100%; there is only one decipherment, and you—you can even extract the key out of it.
Okay. Okay. So what happens if you extract the key? You can not only read this, you can read every message that uses this key.
Correct. Okay. Uh, and this is the theoretical background. The way to decipher it is known much much before this.
Okay. Okay. And there's—there's Mary, Queen of Scots, who was during the time of Elizabeth I.
Yes. And she was imprisoned—I mean, that's the whole history; I don't want to, you know, give a history lesson—but she was imprisoned, and uh, there were, you know, in palaces there's always intrigue; people want to replace one ruler with the other; they kind of collaborate; they try to, you know, so she used to write coded letters, and the code is not just alphabet to alphabet; it's also—they had null values, but the values don't—they don't mean anything; they had for common words like Wednesday, Mr., they had special symbols and so on, so it was very sophisticated even at that time.
Okay. And even those were cracked if they collected sufficient messages. So once Elizabeth's spymaster collected enough messages, he cracked the code. Okay. And what he would do is he would write code pretending to be her collaborators, and he got essentially Mary, Queen of Scots to agree to assassinate Elizabeth I. Now this is a capital punishment.
M. So was she—she ended up getting executed. Yes. So that kind of cracking is fairly simple; it existed for a long time. Cryptograms have existed at least from the 1920s in American newspapers. And my rough calculations based on the number of newspapers that fell in America—of the percentage of newspapers that carried cryptograms over, you know, about 100 years—there have been 100 million cryptograms in America alone that have been solved by so many million people. So cryptogram is a very well-understood thing; it's just that the present generation doesn't read newspapers; they don't know what a cryptogram is.
Right. Uh, so—so that is the bit about cryptography.
M. And so this is a very well-understood—this is literally understood by millions of people. What I—the method I used is not new at all. Okay. So when people say, "Oh, you—what's your method?" It—it is not my method; it's—it's a known method. M. I have just applied it to a different language and a different script, and the websites that do Latin alphabet can also do pretty much any European language, and there are many websites that exist, but I didn't have a website for Sanskrit, so I had to do it by hand.
Okay. That's essentially it. I used regular expressions, and that also is not new. It was first announced on—there's a website called Perl Monks where they use all these techniques. Okay. I think about 20 years ago or 22 years ago, using regular expression software programs was announced. I just essentially did the same thing, and I got the results, and uh, what happens is—and I've described the theoretical—the mathematics behind it also—like why this works, why it's unique, why there's no other possible solution—um, and the people who understand the theory a little bit they should be able to understand it. To those that don't, I will work on making it simpler and simpler, because being able to explain things simply is the true measure of, you know, solving this problem.
Right. So did you already have a knowledge of Sanskrit and Tamil to begin with?
Yeah, I mean, I've taken Sanskrit from 8th to 12th or whatever, obviously forgotten. So everything that I did for this I had to relearn. Right now, the Tamil—I'm actually from Karnataka; I understand Kannada.
Okay. Tamil is related to Kannada. Yes. Some words are different; some endings are different. I can kind of understand most of the words, but—but to solve a cryptogram, you don't need to know a language with fluency; you just need to know the syntax, grammar, and that is something that you can learn as a process of deciphering; you don't need to be a great expert. So, for example, Russians who solve American cryptograms or Germans who solve American cryptograms, they don't need to have a great knowledge of English. Similarly, the Americans solve Russian things; they don't need to, you know, because what you need is you are not actually using your brain power; you're essentially—the dictionary and the set intersection will simply sort out the values and will produce a unique result for you.
Okay. So, uh, but everything I needed to learn, I learned or relearned. I had some background in Sanskrit, yeah, like I said, I had some background in cryptography; I had some background in information theory, M, but the level to which I needed for this paper, I essentially had to re-master it.
Okay. So what was the process for setting up the problem and solving it? First of all, you accumulate all the data, right, which is the corpus of the text.
Yes. Then what do you do with that? How—how do you approach the problem, the solution?
Right. So in—so I'll explain in terms of cryptograms because if there are some people who have done it, you know, it's easier for them. So it's very simple. The set of words that match a pattern,
Mhm. varies with the amount of repetitions you have. Okay. Okay. So, for example, if I have three symbols which are completely different—A, B, C—I'm just going to give them or alpha, beta, gamma. Okay. Okay. Let's assume they're unknown. If I try to find possible matches in the English language, I get about 550 or something like this.
Okay. Okay. But if the first two are the same—if I have alpha, alpha, beta—M, there's only one word that matches it.
Okay. Okay, which is "eye". Okay. M. If only the second two are the same, then I'll get about 10 matches. Okay. Okay. So from 500, I brought it down to 10. M. Now if I find yet another word which has only these symbols, that will go down to like two or three. If I find yet one more, then I'll get a unique result. Okay. So essentially, when you think about it, it's—you have a pattern, M, which when you look it up in the dictionary, you get a set of possible values. Yes. You intersect that with another pattern—like the values of another pattern—you intersect them; your set becomes smaller. M, and you keep doing the intersection till you get only one value, a single value. Right. Now there is a possibility that you get—you try another value, and this becomes a null set. That—that's called a false positive.
False positive, false negative. So you have to do that even after you get the single value; you have to do over and over and over to make sure that value doesn't disappear. So if you continue to do that for many more, and it continues to display the same value, then it's a 99% probability that you have got the right value. Okay. And once you have one value, you can simply use that value to find the other value. So you will start with the symbol that has the highest frequency, Mhm, because you will have the most matches, M, then work yourself down to the rare symbols, the rare symbols, because by then there'll be a lot more matches. So it—it may have only few attestations, but when you do a search, the higher chances it will be next to a symbol you already know. So that's essentially—so it's a very simple process. As I said, millions of people have done it; it's—it's not, you know, anything that I've invented. It's solved to be on—so for this you would need a Sanskrit dictionary; you would need a downloadable dictionary. If you want to use regular expressions, you could still do it manually; it'll just take years.
Yeah. So you need a Sanskrit dictionary online, that download and a Tamil dictionary as well, if you want to match with Tamil as well, if you—any language you want—dictionary.
Yeah, yeah. So you read that, and then you write code for this. So initially I just did the regular expressions because in regular expression—regular expression—so let me—first the—like the information theory description of it—there's something called a residue class. The residue class is—if I have the word, let's say, "EEL," yeah, if I change the symbols to F, F, G, okay, that's the same pattern; it is—it's just the symbols look different, but it's the same pattern. Yeah, if I make it Z, Z, X, mhm, it's the same thing. So the symbols themselves don't have any significance as to what is the set written by the dictionary.
I'm getting it. So those are just dummy variables; the adjacencies and the commonalities of—of the—of the symbols though, that together is called a residue class.
Okay. Okay. That is just a term Shannon used; I'm just using it; I don't want to use a new term. Okay. So once you have a residue class, you can write a regular expression that represents that regular class—that class—um, by just denoting that you're going to capture this particular letter and reference that when it repeats. Okay. So that is called a regular expression, okay, and it's also well-known in the computing world. Anyone who has done almost any language now has regular expression support, and what that—you can—what you can do with that is you can apply it against a dictionary and just say, "Hey, find me all words that match this pattern," yeah, and you will get the results. Okay. So for things that don't have a lot of repetitions, you'll get more, mhm, but things that have repetitions, you get less. Right. And if you have in a corpus—if you have short inscriptions, you can use those, M, and when you have more than one—so even if you don't have repetitions in one inscription, you can take two inscriptions which have the same symbols, and you can do that same intersection. Okay. And you—you will get—by the way, this is all—as I said, it's all well documented; there's a Perl Monks article on it. Okay. I just essentially used that method.
Right. So you initially did this manually?
Initially, I did manually. Yes. And then did you write code for this? So then what I did was—when people said, "Hey, um"—so there was a—there's a Japanese cryptographer who looked at it and said, "This looks good." He doesn't understand Sanskrit, so—but he said, "This looks good," and uh, "I just wish I could, you know, reproduce it, you know, with the uh—like I have the list of regular expressions to reproduce it." So I said, "It's not a big deal; I will write a program," M, and I wanted to write something very simple for two reasons: one is, I want—when—when people read the code, it should be pretty obvious that there's no trickery or anything like that that they have to—so I wanted something—100 lines—just an arbitrary thing. I—I wanted it to fill the screen because I didn't want people to scroll too much or lose interest, uh, and I wanted it to be simple; I wanted it to be simple that people know, "Yes, this is the code," and you run it against the list of regular expressions, and you get all the output. So I was able to—so I wrote all the regular expressions for I think about one-third of the set of alphabets, and when—if you run it, you get all those results. M. So that is open source; that is on my GitHub. GitHub, you'll see. Okay. Something called `script_decipherment`; you can run it yourself, and uh, it actually gets the dictionary; it downloads the dictionary so that I'm not supplying my own dictionary. There are certain changes that are made for accusative and some other endings because the dictionary—um, in Sanskrit—does not have declined forms; it only has the stem. Okay. So, for example, example, "Rama," right, "Rama"—you'll find—you won't find "Ramam," you won't find "Rama," you don't find, you know, all that; you only find "Rama." Okay. So those declined forms, I had to append to the dictionary. Okay. And I've also commented that this is a declined form of which Sanskrit experts can check which is standard. Yeah. So that's essentially—you have a dictionary, you have a script, you have a list of regular expressions, and it outputs it. So yeah, it's mathematically—it's programmatically reproducible.
Okay. Now when it comes to the corpus of the data set of the Indus Valley symbols, yes, the number of symbols that we have or inscriptions that we have is pretty limited, right? It's not a very large corpus that we have, and when it comes to the Sanskrit language, it's an immense language with a very large vocabulary.
Very large vocabulary. So was it a challenge trying to match this with that?
So the—when you say vocabulary is large, the Monier-Williams dictionary—if I remember the dictionary I downloaded had something like 100,000 words; I don't remember the exact.
Okay. 100,000 words; it's a rough number; it could be less or it could be more, but it's not—it's not—it's not like 4,000-5,000. So it's roughly—it's a decent-sized dictionary. Okay. But—but when you say there's a Sanskrit has a large vocabulary, what you mean is that the declined forms, the conjugations, the way to—kantas—which is the way to create out of—roots—all of these can create, um, you know, millions of words, but the dictionary doesn't have that many words; the dictionary has about 100,000.
I see. Okay. So you write the program, and you do the things. So what did you start finding?
So the—so even before I'd—I wrote the program, I was just looking at um some of the seals, and one of the seals had what looked like a Brahmi word—"vada"—V, V. Okay. The W is a Brahmi W; the Da is a Brahmi Da. Now what is "vada"?
Vada can mean "killing," killing. Yeah, it can also mean "killer," killer. Okay. Because that is what's called a Pratipadika. When you add it to a root, it just adds the—and it means the one who does that action. Okay. So—so I—so I saw that, and uh, then I—I wondered if most of the beginnings of this inscriptions—so this particular one had only two signs—two signs—do most of the inscriptions have the starting symbols as the name of a deity? Like "killer" is—sha—right?
Mhm. Do they have that? And this, by the way, is not my idea; this is analysis of the script by other people. Like Rajesh et al. have kind of claimed that in the scripts have like a beginning, a middle, and a terminal; the beginning is two or three letters, and they—they guessed it must be the name of a deity.
I see. So it is not—again, not my contribution, but my validation of—of their research—of the hypothesis. Yes. So I—um—so this—this is the kind of stuff I found—is essentially it starts with the name of a deity in a vocative form. Okay. Means in prayer—for prayer—you have to use what's called a Samhita, okay, and you have to—if—if you're asking for a favor, that verb has to be, you know, it's called a Lata, Lata means—imperative. Okay. Okay. You're asking, "Do X for me," like, you know, "Help me" or something like that. And then there's usually an object; object is—give me a gift, for example; give me water; give me a blessing—that has to be an accusative. Okay. Okay. So that is what—uh—that's essentially what we found, okay, in—while we looked at it, and that's the general pattern. There are many that don't follow this pattern; like they're shorter; there's just the name of the deity, for example; there's just the name of the thing somebody wants—the noun—and so on. Some of them are names of Vedic poets. Like "Vara" is one of the poets. "Vara"—actually, when I first read it, I thought it was "ant," okay, because "ant" is in the dictionary, but then I realized that when I was browsing through the meanings, I found that under the authors of the hymns, one of them is "Vara"—another one—there are many names like that which—for example, "Parna" means "leaf," leaf, yes, right? So I translated it as "leaf," but actually there—it's a name of a king; I think it's also the name of a Vedic poet. So these names exist, uh, and uh, so I—so I started seeing more interesting things, more associations with—so initially I thought it's just—initially my understanding was that the Indus culture is adjacent to Vedic culture—not actually Vedic culture. Okay. And this is not just me; a lot of people claim this because on Vedic iconography and all that, and I also thought, "Yeah, this is like an adjacent culture," and you know, they died out, and the Aryan—where the culture survived and kind of spread. Okay. But based on the names I saw, and eventually when I discovered that Rudra is the precedent deity of the Yajurveda, and I also saw the Puranic iconography, I came to the conclusion that yes, this is—this itself is Vedic culture because the names of the people were pretty unique—if you think about it; some names are repeated, but if you look at all the historical names—you know, Krishna, Vishva—they—most of them don't repeat. So—so—and if I just found "Vara," maybe I would just say it's a name; it's "ant," it doesn't matter, but when I find so many names, I have to assume that these are names of the poets themselves or at least maybe a family name.
Mhm. Okay. So that's what—what you discovered. So by what point had you essentially deciphered all the inscriptions?
So when you say "all," I don't know if you mean all the symbols.
Yeah, yeah. Some symbols—most of the symbols—almost all the symbols—I had done it in six months. So early 2022 when my first paper came out. Now, um, eventually, a couple of them—I think about a handful of symbols—I had to revisit and—and change the values as I read more inscriptions. And—um—then at a certain point, I had to—I had to actually read beyond the unicity distance to be mathematically correct, and I said, "Okay, let me start with the biggest inscription; let's see if I can read it grammatically," Mhm, "and meaningfully," and that is—17 signs—according to, you know, the classical definition, sign is M—but according to me, it's about 20—20. Okay. So I read it grammatically, and uh, I was actually shocked. So I had three big shocks: first, when I discovered that it was Sanskrit.
Okay. Okay. I was like, so—because see everybody has tried—claimed Sanskrit, and so many people have claimed, M, but the thing that I had was—I had the knowledge that these patterns mathematically would not fit in for any other language. Now a lot of people on Twitter and elsewhere have this idea that Sanskrit is this magical language that you can assign symbols, and you can read anything as Sanskrit, and you can touch the grammar and get a meaningful result—context-free—well, it's not context-free—but essentially their imagination of what Sanskrit is—it's not just a language; it's some kind of, you know, some kind of godly thing, okay, that can—bend the rules of mathematics and information theory. So, for example, in their view, they could take the US Constitution, which has a similar size to the Indus corpus—just assign—reassign all the symbols to some random Sanskrit—and they will be able to read it grammatically, and that is not possible mathematically or linguistically. The reason is that would mean certain things—that Sanskrit has an information rate same as the symbol rate, which is impossible. M. It would mean redundancy is zero; this would break down information theory completely, by the way. And it would mean Sanskrit is incompressible. Okay. Okay. So like you have Sanskrit text—you can—if you compress it, you get no compression. And uh, it would also mean that if you take all the words and from each letter you draw an arrow to the next letter and create this what's called a Markov process, that Markov process would be a fully connected graph. Okay. So I know the viewers may not get most of this, but anyone who understands information theory will—will see it's obvious that none of these can be true because A—we know the information rate; we know the redundancy of the language is 0. So—Mhm. And uh—there is—whatever they're claiming is essentially what's called a special pleading claim; you are claiming this language is special without providing any evidence.
Okay. Okay. So you—it is just not mathematically or linguistically possible to do that. Okay. So to answer your question—the—sorry, what was the question? At what point did you—did you essentially decipher all—all the—so yeah—um—so I started—yeah, about six months, I got most of them, then over time I made some improvements, Mhm, and I would say I got most of the symbols within like one year after starting. And you said that you—you were absolutely shocked when you deciphered the longest inscription—70—20 characters—what was the shock that you got?
The shock was that it was grammatically readable. Now grammatical readability of anything of that size is—is very hard—is very hard for—so if you—you may be able to—if you choose and assign values only for those 20, you can get values because a lot of the symbols don
The amount of information the symbols can convey is called redundancy. Most natural languages, including Sanskrit, have a redundancy of 0.7. A specific paper measures Sanskrit's compression, which is a good indicator of redundancy.
Let's consider a language with symbols from 00 to 99—100 combinations. If only 10% have meaning, the information rate is 10%, meaning 90% is redundant. Such a language would have a redundancy of 0.9. Natural language has a redundancy of 0.7; I hypothesize the ideal is 1 - 1/e, approximately 0.62. I have a paper on this. This is due to something called derangement, a mathematical term where symbols are out of place. For example, "cat" becomes "cta"—unintelligible. Redundancy allows interpretation despite mispronunciation or accents. Key space and redundancy map to the number of symbols with equivocation, providing a numerical measurement of language.
This is similar to DNA encoding. DNA has redundancies; only 4% codes for proteins. Sections duplicate proteins; if one is damaged, another can function. This is a good analogy.
Mathematically, I was certain after deciphering the 50 longest inscriptions. Reading the first inscription correctly gave me 90% confidence. That was the turning point.
My initial expectations were that the Indus script would be Tamil or Sanskrit, or perhaps Proto-Indo-Iranian or Indo-Iranian. Indo-Iranian is close to Sanskrit. I expected suffixes and prefixes, but with slight word differences. Avestan and Sanskrit are very close, closer than Avestan is to Vedic Sanskrit. I initially thought I would find suffixes and prefixes, but the word matches were slightly different. I found words that were post-Vedic, and eventually Vedic forms. This indicated Sanskrit. I discovered shortened names, like "ra" for "rudra," to save space. They also flipped icons and cramped letters. Many "rudra" names came from roots meaning "kill" and "roar." The language itself reflects Rudra worship. Initially, I thought Rudra was co-opted, but the iconography and names convinced me it was Harappan. I wondered why they used "Rudra" and not "Indra." "Shakra," a popular name for Indra, is shorter. I figured out Rudra is the deity presiding over the yajnas depicted in Harappan iconography. Many names of Rudra, including nakshatras, were found.
The seals were broken and discarded when obsolete. Later, they used metal (copper and silver). The tradition continued into the Gupta era, using similar formulas and referencing Vedic ideas. My conjecture is that the Indus civilization continued into the Gupta period, the only change being the material of the seals.
The most prolific deity is Rudra, used in prayers or as protection on trade goods. Others include Shakra (Indra), and Maya (Vishnu). A future document will detail deities, meanings, and contexts. There are two types of iconography: organizational and narrative. Organizational icons have inscriptions unrelated to the icon. Narrative scenes, like the Pashupati seal (Ashaman, meaning "honored punisher") or the Laka scene (Shivaratri story), have matching inscriptions.
I wish we had names of kings and dates. Trade data would be less useful. Much of it is religious, including deity names, prayers, references to houses ("asa"), and trader identifiers ("maner," possibly a name of Sha). This information needs to be condensed. The inscriptions show continuity between the Harappan and post-Harappan periods; the civilization didn't disappear.
Place names are partially known: "Daka" (D + Ka, meaning "door" and possibly "port"). "Moha" appears often, possibly "Maha," meaning "union" of city-states. "Hara" is found, possibly related to the Haria river.
Governmental information is inferred. The two blackbucks symbol, widespread across sites, possibly represents nationhood, existing even before 4000 BC. The unicorn, a composite of horse and bull, likely symbolized royalty, similar to Indo-Greek symbols. The four lions/tigers and chakra symbol, dating back to 10,000 BC, is another ancient Indian symbol. Twitter helped synthesize information from various researchers.
The oldest symbols are in Balakot (4000 BC), but unreadable. Abstracted symbols indicate earlier origins. The K site in Kalibangan has a three-sign inscription ("Shani") dating to 580 BC. Indus inscriptions also appear in Tamil Nadu (5th century BC) along with Brahmi. Brahmi evolved from standardized Indus script, with some rotated or simplified signs. Cursive forms and serifs suggest writing on softer materials. Most information was probably on perishable materials.
We know Sanskrit existed by 3000 BC, with continuous civilization. The Indian calendar and astronomical references date the Puranas to the mature Harappan period. The vocabulary includes Vedic, post-Vedic, and some obscure words. No Dravidian words have been found yet (only 40% of inscriptions read). Mesopotamian seals contain Aramaic and Sumerian words.
Sanskrit words, okay? Right. The Indus inscriptions in, uh, uh, in Tamil Nadu are not complete, so I cannot actually read them whether they're Tam or not. Okay, maybe they'll find some more, or I haven't seen all the papers. When they arise, I'll even look at them. But, for example, in the Brahmi inscriptions in Tamil, they use, you know, like, like in Kannada, that that particular retroflex 'sh' which doesn't exist in Tamil. Okay? So that we know that that's a Sanskrit word, and not a Sanskrit borrowing. So when you borrow the 'sh' becomes 's', okay? Like in Tamil, they don't say 'sh', they say 'sa'. Okay, okay. So that's how we know that is actually a Sanskrit word. There were Sanskrit-speaking people there. Okay, so, so yeah, so far I haven't found any Tamil words. My conjecture from that, if we don't find any in the rest of the corpus, is that um, and the collapse people moved South, they took—they essentially Sanskritized and they adopted the Kannada language from there.
So what do you think of the—I mean, it it is said that India has had two parallel civilizations, two cultures: the South Indian culture and the North Indian culture. Do you think that is correct? I mean, I think we've always been an integrated civilization. I can answer that question. So even in the oldest signs of writing and iconography in South India, you find the swastika. The swastika is a sign of the Indian civilization, so you find it there. Even in the oldest writings you see in the script Brahmi, which is a standardized Indus script, you find it there. Mhm. You find Sanskrit there. Mhm. So I don't see any signs of something that is different in terms of culture. You look at the arts and the pantheon and things; it's the same civilization. And it's not just me saying that; it's the, you know, the current government of Tamil Nadu, which is, you know, they have certain political views, they themselves have admitted that the ancient civilization in Tamil Nadu is the same as the Indus Valley. So it is the same civilization. And the, the whole, you know, even before like the 19th century, it wasn't even a thing; people didn't even think that there is a different civilization; everything was one. Yeah, think about it: Ram happens in South India, right? Mahabharata happens in North India, and Madurai is named after Mathura. Wow. Yes, okay. That's interesting. See the 'D' and 'T' are the same in Tamil, so Madurai is named after Mathura. Wow. I had never thought of that. Okay. So, so it is all, you know, it was completely integrated, and before the politics, there was nobody thought of it. I have not seen any discussion of in South India where they thought they were different. Pandya said they were Aryans, Cholas said they were Aryans, who else? Cheras, Vijayanagara thought they were Aryans. My, you know, the Maharajas, Aryans, they would import the Brahmins from the north to do the yagnas, because, you know, then you needed enough of them for certain yagnas, but nobody thought—nobody said we are different. There's no mention of that till the missionaries invented this wonderful idea.
Yeah, that's right, right. So this really is an eye-opener, right? Totally. So, so now going forward, what's your plan with this research? Are you planning to publish this as a paper or as a book? What's up? Yeah, yeah. So there are a few things I want to do. The first thing I want to do is completely transliterate the corpus, yes, um, because it's one thing to say I read 40%; it's a completely different thing if I read 100%. Mhm. Okay. The second thing is, uh, I, you know, as I'm giving talks, I get questions that are challenging to answer. Okay, okay. Not because there are, uh, you know, there's a defect in my decipherment, but because there are things that need to be explained in ways that I'm not explaining clearly yet. Okay. So, for example, you know, I would—I would say—I used to say, you know, beyond a certain length you cannot solve a cipher program in two different ways. Mhm. And if you say it like that as an English sentence, people will say, "Why not? I can solve it," you know? And I used to get all these challenges, you know, "Oh, oh, I'm going to—" Oh, so this is what you did; I can do the same thing in Arabic, I can do the same thing in Chinese. Okay, I said, "Okay, do it." Mhm. And two years, none of the people have come back. But when you—what you want to do is make them understand why not, right? Instead of just saying they challenge me and they're wrong, because that is—that doesn't go anywhere. Yeah, so I figured out that I need to change the argument from qualitative to quantitative. Mhm. And I have to establish the model of the quantitative as a mathematical model and, as you know, in a way that most people would understand. So the other thing I wanted was—I didn't want my paper to be something that only experts can really understand, and everyone has to say, "You know, I Googled it, and Professor X said this is good; Professor Y said this is bad," because what happens in the humanities? Yeah, there will always be people who say this is—this is wrong. Yes, like you could do anything at all, mhm, you'll find people who say this is—that's right, you know? So I wanted the—so this is a hard part. Decipherment, of course, is hard; no one has done it in 100 years. Mhm. It is hard, but I'll tell you what is harder: is to write the paper in a way that ordinary people will understand, and and that is what I want. When ordinary people understand what is decipherment, how it works, and why this is correct—the mathematical basis for it—Mhm, then I think the people will have confidence with their view; they won't be easily bamboozled by experts saying, "Oh, that guy is, you know, he doesn't know what he's doing; his paper is rubbish, this and that," because that—that does happen a lot. You've seen it with, uh, you know, some of our Indic people when they—
Yeah, yeah. So I wanted that to be a rock-solid mathematical foundation. Mhm. And uh, so I had to explain, you know, the various findings and, uh, so one of the things—one of the early criticisms I got from the humanities professors is that, "Why would Harappans write this?" Okay? Some translation. "Why would Harappans write this?" That is called an argument from incredulity. Okay? Because what they're saying is, "If I was a Harappan, I would have written better things," or when—when they see all these multiple symbols having the same value, "Why would they have so many symbols?" Now those are called allographs. Even English has allographs. Mayan, for example, has many allographs for the same symbol, like for 'o' they have 10 symbols. It's very common in all scripts to be inefficient. Okay? There are many inefficiencies in Bronze Age scripts. Mhm. But people, when they look at something, say they have 4,000 years of hindsight behind them, but they don't know that—they don't have that conscious, you know, they don't have the awareness. They—why would they do this? Yeah. So all arguments from incredulity come from a place of ignorance, like they—there's some knowledge they're missing. So the knowledge they're missing for allographs is, for example, that many old scripts have mhm allographs for the same symbol. Then these professors said, "Why would Harappans write this?" I said, "Okay, I've got to look into this. I need a—" because I know it's an argument from incredulity; I can just point out this logical fallacy and stop answering it. But what happens is when ordinary people hear it from experts, yeah, they can come under their spell and say, "Well, Professor so-and-so has said, 'Why would Harappans do that?'" That's right. I would—so I needed, uh, uh, I needed that to be meaningful; I needed the reading to be meaningful. So fortunately, there's a particular inscription that says, uh, you know, essentially, uh, it talks about combing the arrows of Shiva, you know, like the red-bodied one. That's the English translation. And one of the things I had to do because of all these names of Rudra, I had to look for names of Rudra mhm which much matches. So I looked in—you know, I searched Google for "thousand names of Rudra" and so on, and eventually someone told me, "Look at the Shaivagama." Okay? And in the Shaivagama, the first anuvaka itself, the first section itself says, "The red-bodied one whose arrows are combed." So I was like, "This is the same phrase, essentially written more compact, but it's the same phrase." So maybe I should look at all 50 longer ones if there's some similar idea in, like, other Vedic texts. Yeah. So I found—I thought I would find about six or seven; my guess—I found every one of the 50, except this particular one from Shaivagama, which is okay, you know, eventually, is borrowed from—so I found out of that and I said, "If you're saying, 'Why would Harappans write this?' Why—why would Vedic people write the same thing?" Yeah, but why would they write the shortened versions? Why don't they write it word-for-word? Obvious: space. Yeah, they're writing on hard material. But then I found the Gupta seals which write Vedic verses in the compact form. Oh, so the question is, "Why would the Guptas do it?" Well, same reason these guys are doing it. Mhm. So now what happens is when I present it this way, that entire argument from incredulity about the contents of the inscriptions has been answered without a way for them to get out of it, uh, so, and and there is, you know, there was also the question of why there are so many—you know, the reason they ask it is they don't know about scripts. The people who ask these questions, they're allegedly experts in all of these things, mhm, and I'm kind—I was kind of surprised that that every argument they make is essentially a logical fallacy based on things like argument from incredulity, uh, and and, you know, they used to actually—they used to always start with ad hominem. They would, okay, attack the speaker, the writer. Yeah, okay, you know, with the tag line, "They said I don't know what I'm doing."
That's right, right. So fortunately, and because of that, I was anonymous; now they couldn't say anything about me. They don't know anything about me. Yes. So, uh, so I had eliminated that entire thing that way, and so they were left with the argument from incredulity, uh, and there's also, you know, there are other stuff also, other logical fallacies also, but uh, I realized I have to address these things. Okay? So I address most of these things, and uh, some of those things I can address based on materials I found and other inscriptions, and many of the things I can address through mathematics. Mhm. Now, mine is the only mathematical decipherment. Yes. And the things that I used, they're not obscure mathematical papers. Shannon's paper is the most cited mathematical paper in history. Yes, almost 75,000 citations. Citations. Imagine how many people have viewed it. Yes. And you know, even now as we speak, the citations are going to increase. Yeah. And all of information theory recites on that. Mhm. A lot of linguistics, uh, like computational linguistics is based off on that—on—on the micro-process and everything. And the people who understand it understand it very well; people who are adjacent to it, like engineers, can understand it easily enough. Mhm. If I explain to them, which I've done through some of my talks in, um, I and um, you know, even in YouTube talks I've given something because I want ordinary people to have a chance, at least, to understand it. Mhm. So, coming back to the question, what is next? And I'm going to make some incremental information of a paper; I'll try to find a journal that will publish it. Mhm. But like I said, the things like—the script itself, the symbols, what they mean, what's the origin, what are the objects that represent—I think that is going to be so big, it can only be addressed as a book. Mhm. The body of the inscriptions, the names of the deities, all of that—that—that is probably another book. Mhm. And I think it's worth doing one more book on the continuity of symbols and the meaning of symbols, right? So one of the common symbols in Harappan artifacts—not in the inscriptions—is intersecting circles. Mhm. And intersecting circles is actually a symbol of how they form the cardinal directions using Bana's system. Okay? And that's how they used to build the temples and so on. So you see that on a lot of pottery. I see uh, the development of swastikas, like the different kinds of swastikas, the invention or the discovery of fractals. Mhm. There's actually a fractal swastika. I see. Yes, very interesting. So there's a, you know, this Tetris piece with the three horizontal and one thing in the center, with that you can create a swastika, and you can create more of those, and you can create a whole fractal thing out of it. Wow. So they had the knowledge of fractals, they had knowledge of astronomy, uh, they had knowledge of, you know, town planning, all of this stuff, and how that continues into present culture. Right? So, for example, you know, women wear the same kind of jewelry—similar, of course; they know it's a little more—the sculptures are C—but they, you know, the choker, the long thing, T, all of that stuff. I think bangles, bangles, uh, there's one more—in the Sinhalese, right? Sinhalese, they had male and female soldiers, yes, warriors. Mhm. The female shield is different from the male shield. Okay? It's kind of prettier, okay? And it has a symbol on it; it's like a diamond with, like, none. Mhm. That particular pattern is present in Rajasthan, even today. Okay? So that continuity, the meaning of the symbols, mhm, I think that can be another book. Mhm. Uh, it can probably be one book; it would just be 1,200 pages or something like that, you know? So that's an option. I'm going to—you know, there's a lot of demand for it; people want—you know, they keep asking me on uh different things. I think if I could put them in uh one place, it may be useful to people. What—what would it be possible to put on a glossary that would enable just the ordinary person to quickly decipher and and read a—a script inscription?
So I have a website, Induscript.net, okay? And you can—so all the inscriptions are in there, in description of this video, yes. And uh, you can search there in Latin script, in Indus script—if it's deciphered, you can see it in the script; otherwise, you can read it in Latin. You can search for untranslated; you can search for translated; you can search for which roots have been used. Yeah, much—it's a free text search, so mhm you can search for inscriptions that are a particular length. So if you type like capital L6, you'll get all inscriptions of length six. So that site will essentially have all information for Indus seals, and we will essentially add tagging, like, for example, if there's a name of Rudra, we add—add the tag so that you can just search for "name of Rudra" and see—see those inscriptions. Okay? Uh, you can search by symbol. You have a loose match, which is when there are different symbols that are really different variants of the same symbol; it will find you all of those. So, uh, yes, we—that—but yeah, a separate glossary may also be useful for people, for—to, like, do a linear read-through. Yeah, uh, then that's something we can do. Yeah, it would be in, like, book two, for example, right? So I think the biggest takeaway for me is that it's not just Sanskrit that's—that's uh being present for thousands of years. Hinduism also has emerged out of India; it doesn't come from somewhere else, as is alleged. I mean, it's alleged, but if I understand the present academic position, is that Hinduism was formed in India by—I mean, this is obviously falsified by my research, but I'm just stating the—the position is that Indo-Aryan speakers, which is Sanskrit or Proto-Sanskrit, yeah, came to India, and they adopted some Harappan customs mhm and ideas and created Hinduism in India mhm about 900 BC or thereabout, or 1200 BC. 1200 BC. That's—that's the thing. So Hinduism is unique to India, obviously, because there's no Hindu artifacts, inscriptions, any archaeological evidence outside India. So you don't have, you know, uh, you don't have the Vedic altars, like the square, rectangle, and semicircular altars; you don't find them outside. You find them in K-B-B-N. Mhm. You don't find it anywhere outside. You don't find any pictures of Indra, yeah. And—and you know that's the odd thing, because Indra is a rain god; it's a monsoon god, essentially. And um, you know, Afghanistan steppe, they don't see rain; they don't see—they see lightning, like once—one to four times a year; most people would not see lightning in their lives. So it will be an event in their lives, and their lives are not dependent on rain. So what happens is your gods—nature gods—only exist where they have a very strong role separating life and death. Mhm. Volcano goddesses, for example, exist in Hawaii and Japan. Mhm. Why? Because volcanoes can screw everything up. Yep. Uh, you have uh, you know, the deserts have different kinds of gods, uh, you know, like fire and brimstone kind of stuff. Uh, you have like in—the forest places you have the goddess of the hunt, mhm, because if the hunt fails, you know, so—Mhm. Monsoon gods only exist where there's monsoon, mhm, where the rainfall essentially dictates your survival. Mhm. And um, it must have started during the droughts that they—the—the rain god and rain, because you know, we're dying. Yeah, yeah. So it doesn't make sense, like, if you look at all the iconography, you know, the—yeah, the mountains, and the word for um heavens is mountain. Mountain is like stone, *ashman*, right? *Ashman* is the heaven; that means the gods resided in mountains—as above the clouds—which doesn't exist in a lot of places. So it exists in the Himalayas. Himalayas. Yes. Dawn racing over the glowing mountains in the East, so that means mhm are to the east, which is in India. Yeah. In the steppes, the—the—the—the ors are to the west; they wouldn't have a Dawn goddess; they would have a dust goddess, right? Yeah. So if you look at all of these, like every single thing in the ancient—in the text that can only make sense in India. So this is why they said, "Okay, they brought the language; they kind of copy-pasted everything and created this," which—I mean, it's absurd for various reasons. One is the timeline itself, like I told you, the—the Dua star. Yes. You—you can't see that. Okay? That is one. The other is the whole body of your text. Your text—the linguistic changes. So what happens from R to—you have some linguistic changes; you have the—and becoming G in front of, and some of these—and then you know—then it goes back. Uh, in Samskrit, you have this pluta; you have this long words which are very long. So those language changes and all, they don't happen overnight; they don't happen every generation—long-term. So if you add all that—if you add the changes in the vocabulary uh from you know [Music] SAS—then you have the Samskrits; all of this stuff—everything they say that happened before P, like, think SP has mentioned mhm like this much. If I have to squeeze it into 600 years, are those linguistic changes happening every 20–30 years? Like how—how is it—how is it even feasible? So that—that claim itself, even without my decipherment, is uh is just hand-waving, and it's just essentially—look, the European studies essentially cares about Europe, and they have this favorite theory, yeah, and to them India is an edge case that they can hand-wave, saying, "Yeah, it's—they went there, done." They—they don't actually care about you. Like it's—it's about them; it's not about you. You are—uh, thin in their shoe, essentially, like—you're—you're stored in the shoe. Mhm. But anyone with enough—even preliminary Indic knowledge would say, "How is that possible? You fit everything into this 600 years; it just doesn't make sense." Yes. Did they have a linguistic change every two generations? Nonsense, right? I mean, there are other things also, like, you know, um, the—all the words for the Indus technology, there's Sanskrit: brick, *sindhu*, swastika. Brick is *isa*. *Isa*. Yeah, uh, *yava*, barley; all of it is Sanskrit words. Mhm. So did they just tell everybody to, you know, this—be creating 100,000 new words; everybody forget the old words? I mean, it's just—um—so it—it has to be a series of nonsensical things to have happened in India, yeah, in a very short time by a very small number of people for this to have any meaning. Mhm. Right? And—and I ask this, you know, I asked the question: Just play in your mind—like reconstruct whatever you're claiming, and let's see how you can do it. You're in the steppes. Mhm. What was the situation in the steppes? Everybody who was in a group of 20 to 200—average size is 20—see 70–7—okay? And Harappans generally have an animal-to-human ratio of 3:1, so you have about 100 to—you know, 400 animals. You're herding them. Steppes are vast; they're huge. Yes. So you never see another person; you never see the tribe. Mhm. And you see—if when they draw the areas of the cultures, they're all overlapping. Why? Because you don't see them; you cross into others, but you never see—you never meet them. Yeah. And so now you're herding people around you. When you meet someone—some other group—usually there's a battle of the animals, because this guy says, "That's my animal," and he'll say, "It's my animal," little fight. So they generally avoided each other, and if you see each of these cultures, they have—they're completely different in terms of their clothes, their weapons—weapons that traded—so it's kind of similar, uh, but they are essentially—the pottery, everything is different. So they were not a uniform culture; they were all different cultures; they were isolated. Now how do you decide, "Hey, I'm going to organize; I'm going to go across the whole steppes from essentially Western Ukraine, you know, where the forest steppe starts to Mongolia, collect everybody." How you going to do that? Okay? Then let's say you figure out a way. Mhm. Then you say, "Okay, we have to go invade India," right? You have to convince every group of 50 people, okay? And then you say, "Okay, we're going to leave our women behind," mhm "and we're going to leave all the animals behind," because remember there's no evidence of steppe animals ever coming into India. That's right. Okay? There's no evidence of steppe clothes coming into India, so they leave their clothes behind also, and uh, there's no evidence of steppe transportation—chariot, cart, nothing. Yeah. So they walk naked, basically, only the men, mhm, and they cross the Indus Kush. Mhm. Anyone seen the map of the Himalayas? It's—so many square—million kilometers of maze—so mountain passes are not apparent; people get lost in them. Y—even as far as the 19th century or 20th century, when they wanted to create highways in the United States, they had to get Native Americans to tell them, "Where are the passes?" Yeah, to create the highways. You just can't walk into it. So they didn't have GPS, nothing. So these guys somehow walked in, and uh—I know—so even if males walked in—I know 5,000 males—well, I guess the babies and the old people wouldn't be able to fight. Yeah, it's about 2,000 men walked in, and uh, there are 10,000 sites, right? 1,000 sites, but uh, including forests and all of this, there must be quite a lot of villages to—to address. So they went everywhere, and there's no sign of battle, right? No battles at all. No. So they have to tell everyone—convince them—convince them, "Hey, uh, what you're doing is cool. Mhm. Swastika's cool; *sindhu*'s cool; we'll continue all that, but we want to rename all this stuff to our language." Yeah. Now how do you even tell them if you don't know their language, right? How—how you going to—imagine you go to a place where that—it's a different language; you know nothing; you don't know how to speak to them; how are you going to convince them, "Hey, from tomorrow, these 100,000 things you have—navigation boards, farming equipment, crops, administration, weights and measures, scales, mhm uh, royal seals, geography, um, names of rivers—literally 100, especially 100,000 things. Mhm. You're going to tell everyone, 'Hey, I thought of 100,000 new different names. Yeah, I want you to completely forget your old names; like whatever you've written, I want it erased; I want you to never mention these words—old words—again; I want you to use these new words,' and um, you know, I'm naked and a bunch of guys, we're going around this, you know, 10,000 sites, and uh, somehow they did it in like a flash." Because if you're going—if the effort to do this takes like 5–10 years, by the time they reach the 10,000th village, that's 100,000 years. It's a—yeah. So if they had to do that whole thing in—I don't know how long it took them, they claim 100 years—that means they were doing every village in one-tenth of a year, like in every month they were converting 100,000 words, yeah, and convincing everyone to switch the language. So mathematically, it's—it's impossible. Whatever they're talking about—already before my decipherment—that's right. Yeah. So the people who claim this, they're just throwing out terms; they're not thinking the thing through; how it—how it could have happened. Uh, they're just making claims because that is the way to address an ed—is—Mhm. And the problem is they can do it in Europe because Europe doesn't have—other than Greece—doesn't have old attested language. Yes. Yeah. Even in Greece, most of the attestation is from 600–700 BC, from their Homer and all that. They have some from Mycenaean, and they can, you know—most of Europe they can claim anything they want. In India, they cannot claim it because the Indian texts are much more voluminous, and all the European texts—just a huge amount of text—thousands and thousands of years of tradition. So it is—you have to kind of hand-wave it and say, "Yeah, this is what happened; don't worry about it." So and in the beginning when I was a student in the eighth grade or whatever, I—I said, "Yeah, those guys can't give the language." We all thought the same. Yeah. But as I learned more about the Indian tradition, the texts, the volumes of the texts, yeah, the astronomy in the texts, the linguistics of the various, you know, Vedas and all that, every day the time—from my understanding, the time required to create this keeps expanding. Okay? First, I was of the opinion this could be done in a few hundred years—thousand years—4,000 years. I mean, it's just—because that is so much—so voluminous, isn't it? Yeah, yeah. So it's—you know, for me, it's hard to take any of this like—even seriously; I can't even treat it like a serious claim. That's right. So what you have demonstrated is that Sanskrit is the oldest Indo-European language. I mean, they claim that the Harappan or something—some—some extinct language is the oldest Indic language, but clearly now—at least—yeah—
H. But if Sanskrit goes back to at least 4,000 BC, when—that's clearly the oldest Indo-European language. So uh, so to be precise, there's no such thing as the oldest language because language already existed. Just oldest—oldest you can say Sanskrit is the oldest attested attested Indo-Aryan language. Yes, we have uh from, you know, uh definitely from 3,000 BC, uh, but possibly 4,000 BC because the symbols are indeed the same. What's your view about this Indo-European language family, that Proto-Indo-European structure? Yeah, that's—what's your view about that? You think it makes sense? Is it absurd? No, no, it's not absurd. It's—um—so you have to understand the background of this and who did this and what they thought and how it, you know, so it's a little involved. Uh, let me try to give a compact version of this. So essentially, uh, you know, it was pretty obvious that Greek, Latin, and Sanskrit—you see—Greek and Latin were learned by educated people in Europe. That's right. For—because it's just classical; they have a lot of literature and so on. So educated people knew about these languages. Mhm. And when the British came to India, the earliest East India Company people were actually very curious; they were very intelligent; they were actually, uh, you know, the top 1% of—of the—the population of Britain. Yes. Of course, the later people were, you know, different, but the earliest people were, you know, they were curious; they learned a whole bunch of things, and one of the things they learned was Sanskrit. Yes. And they realized that these three languages are very close in many different things. Yes. And then it was proposed that they are somehow interconnected. So now what is this interconnection? Yeah, how are they interconnected? And how—what's the reason for it? Mhm. So eventually, some people came up with the hypothesis that they were originally one language and to split into different languages, and uh, a lot of this work was done by the Germans. They actually called this Indo-Germanic; it was Indo-German before it was Indo-European, right? Okay. And uh, the idea that it's connected as a tree, like starting from a root, branching, you know, few branches and then sub-branches and so on, actually comes from the biblical idea of the Tower of Babel. Mhm. In the Tower of Babel story, humans tried to build a tower to reach the heavens, and God just messes up their language, and you know, they—yeah. So that is essentially the basis. And what they did was they presumed that the language is connected as a tree, mhm, and they started to reconstruct that tree. And this is called an axiomatic model. Mhm. So an axiomatic model is like a—it's an assumption, mhm, and you create uh these rules that are called axioms, and based on that—in that axiomatic system—things will be consistent. So the issue with an axiomatic model is that you have to understand that it has nothing to do with reality outside that axiomatic world. Yeah. So it is useful, for example—so in the old days, they—they had to discuss Sanskrit words, and you can see this even today in the Sanskrit dictionaries; they would have a Greek cognate, and they would have a Latin cognate. That is because the old dictionaries were written before this whole Indo-European uh tree was developed to this level. Okay? So if you see in Monier-Williams, you'll see Greek and Latin, and sometimes other languages also. And so they would say—when—when talking about a word, they would say, uh, "This word is present in Greek as this word, present in Latin as this word." Okay? Now, the purpose of this Indo-European uh language tree is that you could navigate up to Proto-Indo-European and navigate it down to the other side. So now you can describe the relationship through a series of what's called regular changes. Okay? So in terms of that, it's useful. When the tree was—when the language—when the Indo-European tree, or what's called Proto-Indo-European, that whole tree was created, mhm, it was mentioned—or it was—people were cautioned that this is a model; it is not temporal, mhm, which means that it is not a language that some people spoke at some point in time in some place. Mhm. It was very clearly mentioned uh, and as an axiomatic model, I think it's fine; it's useful to see how obscure branches are connected to each other and how the changes occur and so on, but you should not—conflict with the empirical model. An empirical model means that is based on actual data, actual attestations. And the parts of the tree that are attested are empirical; the parts that are not attested—the—the ones—words with the star—are axiomatic. The axiomatic part, you can treat it like any tree in uh graph theory. Mhm. You can rotate the tree; you can make something else the root; you just have to invert the relationship—invert the—invert the transformations. So I can—if everything was axiomatic, I can pick any node and make it the root. I can—I can pick Modern English, or I can pick Ionianics, which is the hip-hop in a—a—pap language; I can make that the root, mhm, and you know, I just have to invert the transformation. But if your uh lower branches are attested—you know, Middle English or NS—all this—even those—I can insert other languages. Yeah. So think about