Showing posts with label indo-aryan. Show all posts
Showing posts with label indo-aryan. Show all posts

Monday, 18 June 2012

"Oh, no, maadarcho-": On subtitling vulgarities in Hindi films

John McWhorter, in his recent New Republic article, "Gosh, Golly, Gee: Mitt Romney's verbal stylings", discusses what he appears to view as Romney's over-sanitised style of public speaking as a marker of inauthenticity. Lucy Ferriss, in her post "Jeepers!" on the Lingua Franca site, is somewhat sceptical of this argument.

However, what struck me in Ferriss's post was the following paragraph, because it touches on something I was pondering a couple of days ago.
It is certainly true, as McWhorter observes, that public discourse has grown more casual and that examples of “taking the name of the Lord in vain” are not so proscribed as they once were. I only became aware of my own habitual use, not only of various expletives involving Judeo-Christian names for the deity, but of designated euphemisms, when I was in Pakistan recently. I would start to say, “Jesus, it’s hot,” and realize that my hosts’ theological frame of reference was somewhat different. Soon I began censoring not only “God” and “Christ,” but also “jeez,” “criminy,” “omigod,” and “lordy.” It was surprisingly easy to do, and as my speech changed, I also noticed no swearing (at least in English) on the part of my interlocutors, who did use other American slang freely.
 An oddly persistent feature of Hindi-language film English subtitling is the bowdlerisation of cursing. A particularly amusing instance of this occurs in the film Murder 2, a somewhat gruesome thriller. The main character, a hard-boiled ex-cop, is verbally abusing another character, and calls him मादरचोद (mādarchod).* Now mādarchod means "one who has sexual relations with his mother" and thus has a readily available and obvious English gloss. However, in the English subtitles mādarchod is rendered as "scoundrel". The disparity between the original and the translation afforded me a good chuckle (my wife simply ignores the subtitles, so wondered why I started laughing).

What is even more amusing is that this bowdlerised subtitling extends to subtitling English as well. So, in the same film, when the hero disgustedly says "Fuck." in sotto voce, the subtitles tell us that he said "Oh, no!".

I wonder if there is a certain subset of South Asians (who can speak English, and reside somewhere in South Asia, as opposed to abroad) who are uncomfortable with cursing in English (even if they do so in other languages) - this subset would seem to include everyone who provides English subtitles for Hindi films.

Postscript: "Taking the name of the lord in vain" doesn't translate well very into a Hindu setting. Hindi speakers will exclaim हे भगवान! (he bhagwān) "Oh, lord!" in times of crisis (or mock-crisis), and likewise will say "Oh, lord!" or "Oh, god!" in English in the same fashion. But these are all what I would call vocative uses, supplications to divine powers for assistance (and I would think "lordy" would fit into this category too). I can't think of Hindi language curses which parallel zounds (< "by god's wounds"). Hindi swearing usually involves some sort of reference to sex or sexual organs, usually involving someone else's mother or sister --- बहिनचोद (bahinchod) "one who has sexual relations with his sister" being in fact a bit more typical than मादरचोद (mādarchod).

 * मादरचोद (mādarchod) is interesting from the standpoint that mādar is a borrowing from Persian but is infrequent outside of this compound. That may well not be accidental --- its use in other contexts may be "blocked" by association with mādarchod.

Tuesday, 20 September 2011

Lizards, Walls, Dragons: on an apparently undocumented Nepali lexeme (भित्ति)

I have not posted in some time due to dissertating, searching for (and thankfully finding) a job, and subsequently moving. Here's a short posting on a Nepali word which I heard from my wife which I can't find in any Nepali dictionary.

When we moved into our new house, we discovered that there were a number of house-lizards already resident (and, less amusingly, quite a few German roaches), which our cat has really enjoyed hunting down. I remembered having such lizards in our house in India, and immediately I saw them remarked to my wife "देखो! छिपकली है!" (Look! There's a lizard!"), using the Hindi word for "lizard", छिपकली [chipkalī]. My wife replied, "in Nepali we call them 'bhitti' (भित्ति)."


I'd never heard this word before, and was curious. I checked Turner's A comparative and etymological dictionary of the Nepali language as well as his mammoth four-volume A comparative Dictionary of the Indo-Aryan languages. Neither mentions bhitti or anything like it. I also checked a number of Hindi dictionaries, none of which turned up anything. Except for Platts' A dictionary of Urdu, classical Hindi, and English, which has an entry for भित्तिका bhittikā:
S بهتکا भित्तिका bhittikā, s.f. Wall (=bhīt, q.v.); small house lizard.
This isn't quite bhitti, but it's close. I had already supposed (and my wife had already suggested) that bhitti was connected with the word for "wall" (in Nepali, भित्तो bhitto or भित्ता bhittā), given that they're often found on walls. So bhitti is something like "wall-(related) creature". [Turner does have an entry for bhitti, but he gives the meaning "wall".] 

Platts' entry indicates a Sanskrit origin, and indeed  bhittikā looks awfully Sanskritic, with the "diminutive" -(i)ka suffix, which is not really always diminutive, but rather can also attach to words with no change in meaning. But here perhaps a diminutive based on "wall" makes sense. 

Interesting, the Sanskrit word for "wall, panel, partition", bhittí, comes from a root √bhid- "to split", which is very dear to my heart (part of the Proto-Indo-European dragon mythology). 

So there's a "new" Nepali word:  bhitti "house-lizard", which doesn't seem to have been recorded before. It may be dialectal (i.e. I'm not sure that Kathmandu Nepali speakers would use it), and that's perhaps why it wasn't previously recorded. In any case, I think it's a cool word, given that it does sort of connect lizards and dragons, indirectly.

[Incidentally, Platts suggests that Hindi छिपकली [chipkalī] derives from the root chip- "to hide", which is what I always assumed (going back to an early Indo-Aryan *chapp- "press, cover, hide". Turner, on the other hand, derives it from Sanskrit शेप्या śepyā which means "tail" (and "penis", but I think "tail" is what is relevant here). The (potential) Nepali cognate of Hindi छिपकली [chipkalī] is छेपारो chepāro, though the latter might be more plausibly derived from  Sanskrit शेप्या śepyā "tail", especially as छेपारो chepāro seems to refer to outdoor lizards (while माङ्सुलि māṅsuli is used for house lizards).]

Friday, 8 July 2011

Some ponderings on Google's research on inter-language linking (Bengali <-> Swahili, Nepali <-> Marathi)

On the Google Research Blog, the latest post (by ) concerns inter-language linking, i.e. looking at webpages' off-site links which go to a page in another language. From the post:
Most web pages link to other pages on the same web site, and the few off-site links they have are almost always to other pages in the same language. It's as if each language has its own web which is loosely linked to the webs of other languages. However, there are a small but significant number of off-site links between languages. These give tantalizing hints of the world beyond the virtual.
I'm particularly interested in the data on Indian language webpages' inter-language linking, especially as there are some perplexing findings. But let's start with some findings which aren't really that surprising.

One of the features measured is the degree to which webpages in a particular language are "introverted" or "extroverted", where more "introverted" webpage languages have fewer inter-language off-site links. The data are summarised here:























Webpage languages which are higher (on the y-axis) are more introverted; webpage languages which are further to the right (on the x-axis) represent languages with a greater number of total webpages.

First, a word about the apparently high degree of English-language webpage "extroversion". The relatively high percentage of English-language websites which link to non-English websites is unlikely to represent a high percentage of native English speakers who are linking to non-English websites. Rather, this would seem to simply reflect English's status as a/the world language, so that even sites whose audience may largely consist of non-native English speakers may choose to create English-language websites simply in order to have a larger audience. And I suspect the "extroverted" English-language webpages are of that type: English is the language chosen for this type of website due to its ability to reach a more "universal" audience, but the site itself may have "local" interests, reflected by its linking to non-English language websites.

But it's the Indian languages that I really want to talk about. Given the large number of Hindi speakers, one might at first be surprised at the relatively small number of Hindi language sites (compared to say Japanese). This, I think, is easily explained by the status of English in India, especially amongst people who would be more likely to create and use Internet sites. In another words, many native Hindi speakers would choose to create English- rather Hindi-language webpages. The high degree of insularity ("introversion") of Hindi-language webpages in terms of inter-language linkage is likely not unconnected. In the context of modern India, choosing to create a Hindi- rather than English-language website is already a more "insular" choice, given the widespread use of English in India itself. Those website content creator who choose Hindi medium over English medium are likely to have more "insular" interests, and thus would not be as likely to link to non-Hindi sites (and even less likely to link to non-Indian language sites).

So, thus far, there isn't really anything terribly surprisingly about these findings. But when we look at the particular inter-language link connections which are strongest, especially in the case of Indian languages, there are some weird data:
















[The arrows indicate directionality of linkage; red connections are stronger than green connections.] As point out:
Surprising links include those from Hindi to Ukrainian, Kurdish to Swedish, Swahili to Tagalog and Bengali, and Esperanto to Polish.
I would add that the Swahili-Bengali and Swahili-Tagalog links are not only strong (red), but also bidirectional (e.g. Swahili pages are linking to Bengali pages, and Bengali pages to Swahili pages). It is hard to think of convincing explanations for the connections between Swahili and Bengali (or Swahili and Tagalog). One possibility comes to mind, which is that, in terms of total Internet representation, the number of pages in Bengali, Swahili, and Tagalog is relatively small. Here the Google researchers' webpage selection criteria is presumably relevant:
The particular choice of pages in our corpus here reflects decisions about what is `important'. For example, in a language with few pages every page is considered important, while for languages with more pages some selection method is required, based on pagerank for example.
This means that for languages with a smaller Internet population individuals could have a greater effect on the particular inter-language linkages than is the case for languages with larger Internet populations.  And thus perhaps the existence of a few creators of Bengali webpage content who happen to live in central eastern Africa could be responsible for some these unexpected inter-language linkages. I would be curious to what sort of Bengali sites link to Swahili sites (and vice-versa) to see if this is a plausible idea.

There is something which worries me about these data though: look at the linkages between the Indo-Aryan languages (Punjabi, Gujarati, Marathi, Bengali, Nepali, Hindi). Punjabi, Gujarati, Marathi, Bengali, and Nepali all have strong bidirectional links with Hindi, which is to be expected given Hindi's status as a Indian lingua franca. Notice however that other than being linked with Hindi, none of the other Indo-Aryan languages are inter-linked with each other: except for Nepali and Marathi.

In India,there are large Nepali communities in West Bengal and other eastern parts of India.Marathi is spoken in Maharashtra in the far western part of India. I would be unsurprised if there were strong Marathi-Gujarati inter-language linkages (since these two languages are spoken in the neighbouring states), or if there were a strong inter-language linkage between Nepali and Bengali. But a Nepali-Marathi link doesn't make sense, at least in absence of other intra-Indo-Aryan linkages.

There is one property which I can think of which does link Nepali and Marathi, namely the fact that they both are written in Devanagari script (also used for Hindi). Gujarati, Punjabi, and Bengali, on the other hand, are each written in their own scripts (distinct from Devanagari). So I wonder if there is any possibility that the script is creating "false hits" when the off-site link connections for Nepali and Marathi are being computed. 

That also makes me worry about the other surprising inter-language linkages, such as Bengali-Swahili, Swahili-Tagalog. Not, obviously, that these languages share a common script, but whether some of the apparent connections are artefacts of the algorithm, whether due to use of a common script or some other factor. If they're not simply artefacts, then it certainly would be interesting to find out why, for instance, Bengali-language and Swahili-language webpages are linking to each other.

Wednesday, 23 March 2011

Linguistics Behind the Wicket (LBW) #1: Shahid Afridi and Free Love Friday

In belated celebration of the breaking of Australia's 34-match unbeaten run in World Cup matches by Pakistan, I offer the first in what I plan to be a recurring series of cricket-related linguistic investigations. I'm dubbing this series LBW ("Linguistics Behind (the) Wicket").

Shahid Afridi after the 2011 World Cup Pakistani victory over Australia
Shahid Afridi during the Pakistani World Cup 2011 match with Australia

This first investigation is a study in onomastics, taking as its subject the name of the skipper of the Pakistan team: Shahid Afridi (Urdu: شاہد آفریدی). To find out the connection between Afridi and free, "love", and Friday, read on!

[A brief word about the sources of Hindi/Urdu words: alongside of the native Indo-Aryan vocabulary (inherited, ultimately, from a vernacular cousin of Sanskrit), both the Hindi and Urdu varieties of Hindi/Urdu employ a large number of Persian and Arabic words (as a result of the Mughal invasion of India).]

Shahid (Hindi: शहीद; Urdu: شاہد) is an Hindi/Urdu word of Perso-Arabic origins, meaning "martyr" (religious or political). It derives ultimately from an Arabic root شہد, which Platts[1] glosses as meaning "to give testimony". Not being a semiticist, I cannot offer any further interesting discussion.

It is rather the name Afridi (Hindi: आफ़्रीदी; Urdu: آفریدی) which is of more interest for me. Jokingly, I have sometimes referred to Afridi as "Afriti", since his aggressive cricketing (Afridi holds the record (37 deliveries) for fastest century in one-day cricket) and mercurial temperament is suggestive of an Arabian Afreet (an angry sort of djinn): Arabic ʻIfrīt عفريت, pl. ʻAfārīt عفاريت. [The origin of this word is rather opaque to me: Platts[1] derives it from an Arabic root عفر meaning "to roll in the dust"; the Wikipedia article suggests that it comes from عفرت (`afrt) meaning "the evil"; the translation of the Qur'anic passage, Sura An-Naml (27:39-40) seems gloss it as "strong one". Maybe semiticists could enlighten me here?]

However, Āfrīdī, in fact, has no connection with Arabic "Afreet". Rather, it is a word of Iranian origin, which, being the name of a certain Pathan tribe, is thus presumably indicative of Shahid Afridi's ancestral origins.
afridi soldiers
Some afridis in the Khyber Rifles

In terms of its etymology, the word āfrīdī can be derived from the Persian word آفريده āfrīda, which means "creature" (noun) or "created" (adjective). (The āfrīdīs are thus perhaps "the created people".)

Āfrīda itself can be derived as the past/perfect participial form of the Avestan root frī- "love" combined with the prefix ā- (theoretically contributing a sense of "near, towards", but sometimes resulting in idiosyncratic meanings). Avestan āfrīda would corresponds to Sanskrit āprīta, both meaning "gladdened, joyous" etc.

The semantic change from Avestan "gladdened, joyous" to Persian "created" is intriguing. The earlier meaning of "joy" still seems to be present in Persian (and Hindi/Urdu) āfrīn/āfirīn, which can be used to mean "bravo! well done!" (though it too can have the "create" sense, at least in the compound jahān-āfirīn "creator of the world").

The root underlying both Sanskrit āprīta and Avestan āfrīda is Proto-Indo-Iranian *prī-, which itself can be traced back to the Proto-Indo-European root *prī- whose most basic sense is "to love".

The PIE root *prī- (see Watkins[2]) is also the source of English free (from Old English frēo, derived from the verb frēon "to love, to set free"), friend (from Old English frēond "friend, lover"), and Friday (from Old English Frīgedæge "Frigg's day", where Frigg, the name of the Scandinavian goddess of love, Odin's wife, derives from Proto-Germanic *frijjō "beloved, wife"); as well as Old English frioðu "peace", which sadly has no direct reflexes in modern English.

In fact, PIE *prī- underlies not only the Persian tribal name Afridi, but also a variety of Germanic-derived names (see Watkins[2]), including:
  1. Siegfried, from Old High German Sigi-frith "victorious peace"
  2. Godfrey, from Old High German Goda-frid "peace of god"
  3. Frederick, from French Frédéric, itself a borrowing of Old High German Fridu-rīh "peaceful ruler"
  4. Geoffrey, from Old French Geoffroi from mediaeval Latin Gaufridus, itself a borrowing from Germanic *Gawja-frithu- "(having a) peaceful region"
Thus perhaps Geoffrey Boycott can mention his "prī-" connection with Shahid Afridi if he ever needs some filler material when commentating a Pakistan match...

So, this concludes the first LBW. I'm open to suggestions for other cricketers or cricket terminology to etymologise for future episodes.

Bibliography:
[1]Platts, John T. 1884. A dictionary of Urdū, classical Hindī, and English. London: W. H. Allen & Co., 1884. (Reprinted, New Delhi: Munshiram Manoharlal, 2000.) [online]
[2]Watkins, Calvert. 2000. The American Heritage dictionary of Indo-European roots. Boston: Houghton Mifflin, 2nd edn.
[3]McGregor, R.S. 1993. The Oxford Hindi-English dictionary. Oxford: Oxford University Press. (Indian edition: New Delhi: Oxford University Press, 1994.)

Tuesday, 8 March 2011

Indian voices from 1913-1929: Gramophone Recordings from the Linguistic Survey of India

George Grierson pioneered the vast Linguistic Survey of India in 1894, an immensely useful resource for anyone working on languages of the Indian subcontinent. A set of recordings were also made as part of the survey, which were recently uncovered in the British Library. These recordings are now freely available from the University of Chicago's Digital South Asian Library at http://dsal.uchicago.edu/lsi/



In order that the languages might be more easily compared (and because "it contains the three personal pronouns, most of the cases found in the declension of nouns, and the present, past, and future tenses of the verb"), Grierson chose to use translations of the Biblical "Parable of the prodigal son", and many of the recordings are of speakers reciting this parable in their native language.

Here is the recording of the "Parable of the prodigal son":
  In Hindi (one of the major languages of India)
  In Khasi (a Mon-Khmer language spoken in Shillong, Meghayala, [the former capital of Assam])

OPEN Magazine has a great article about these recordings, their rediscovery and content, available here: Voices from Colonial India

It's well worth a read, but here are a few highlights. For instance:

Some of the Sanskrit recording took a bit of doings. Background: strict followers of the Vedic/Hindu tradition are supposed to safeguard the Vedas from the ears of those who are not dvijas ("twice-born", those who wear the sacred thread). This prohibition was taken seriously by some authorities, for instance, in the Gautama Dharma Sutra we find:
अथ हास्य वेदमुपशृणवतस्त्रपुजतुभ्यांश्रोत्रप्रतिपूरणमुदाहरणे जिह्वाच्छेदो धारणेशरीरभेदः
"Now if he [a Shudra = a non-dvija/untouchable] listens intentionally to (a recitation of) the Veda, his ears shall be filled with (molten) tin or lac.   [Gautama Dharma Sutra 12.4]
From the OPEN Magazine article:
...All of this, of course, could not have been accomplished without some Brahminical drama. The scholar Ganganath Jha, who was approached for the Sanskrit reading, was scandalised to learn that a mlechha [Sanskrit for "barbarian", "foreign devil", and thus by definition a non-dvija] would be privy to his chaste Sanskrit. A demand was made for a certifiably Brahmin gramophone operator. The Raj, almost as unbending as Brahmins, refused. A compromise was reached: Jha sat in a room and spoke into a large horn-like object that projected his voice into another room where the operator sat. Communication between the two was by means of a complicated system of switches to ensure that the operator didn’t physically hear the Sanskrit. And that was enough to assuage the Brahmin guilt about speaking Sanskrit into a device that held the power to broadcast it to the world...
Jha's recording must have been of some Vedic text, because I am unaware of any general prohibition against speaking Sanskrit in the presence of non-dvijas. Sadly, I cannot find this recording on the University of Chicago's Digital South Asian Library site (they do have a general entry for Jha here: https://coral.uchicago.edu:8443/display/lasa/Ganganath+Jha+Ken.+Sanskrit+Vidyapith+%28Allahabad%2C+India%29).

[Brahminical rationalisations can be both amusing and creative: My advisor, who is a (German) Sanskrit scholar, once told me about one spoken Sanskrit conference he attended (where, I believe, he was the only non-Brahmin/non-Indian) at which there was one attendee who was a bit unhappy with the presence of a non-Brahmin, and was careful not to let my advisor's shadow touch him... Other attendees came up with rationalisations: German sounds a bit like Sharma, a Brahmin surname, and so they theorised that Germans are perhaps "long-lost" Brahmins, and therefore my advisor's presence could be a acceptable.]

Another interesting bit from the OPEN Magazine article:
Many of the speakers chose to sing or recite poems or limericks. Particularly lingering is the voice of Hassaina of Delhi who has clips in the Ahirwati and Mewati languages. Who was this girl who sang with such sang-froid of love and waiting on 26 April 1920?  Nothing is known of her. She survives only as a voice.
Here is Hassaina's song: http://dsal.uchicago.edu/lsi/6838AK

[27 May 2011: Nepali is actually represented too, hidden under "Khaskura", including both the parable of the prodigal son translation, and a delightful song sung by a Shillongwala Nepali, Babu Dhan.]

Wednesday, 8 September 2010

English "like" can, like, function like Sanskrit "इति"

In a recent blog post, "How Old is Parasite 'Like'?", Oxford Etymologist Anatoly Liberman explores the history of (modern) English like when used as a type of filler/discourse-marker. However, modern English like has another function, which I think is often conflated with filler like (presumably because it commonly occurs in the speech of people who also use filler like): namely, as a sort of quotative marker (also noted by commenters Mike Gibson and Charles Wells).
She was like OMG! And then I was like wow!
The above sentences might be "translated" as:
She said, "Oh my god!" And then I said, "Wow!"
Or (since there seems to be some ambiguity):
She said, "Oh my god!" And then I thought, "Wow!"
This use as a sort of quotation mark is a separate function from its pragmatic discourse-marking use (which Liberman focusses on) in examples like:
You wanna, like, go see a movie?
Which might be uttered by a teenage boy asking a girl out on a date , where like can either act as a "hedge", foreseeing the possibility of rejection ("....it's ok if you don't want to"), or to allow for the possibility of other activities ("...or get some ice cream").

Both uses of like are stigmatised; again, the stigmatisation of "quotative" like is probably via guilt-by-association with the discourse-marking/filler like. I'll admit that, like Liberman, I find both uses rather aesthetically displeasing (which doesn't mean that I never use them---they are, as Liberman suggests, somewhat viral). But the "quotative" like is interesting. Though Liberman remarks that:
Particularly disconcerting is the fact that the analogs of like swamped other languages at roughly the same time or a few decades later. Germans have begun to say quasi in every sentence. Swedes say liksom, and Russians say kak by; both mean “as though.” In this function quasi, liksom, and kak by are recent. The influence of American like is out of the question, especially in Russian. So why, and why now? Delving into the depths of Indo-European and Proto-Germanic requires courage and perspicuity. But here we are facing a phenomenon of no great antiquity and are as puzzled as though we were trying to decipher a cuneiform inscription.
Interestingly, however, the "quotative" function of English like has a couple of parallels in Sanskrit. One is these---the one most closely resembling English "quotative" like, at least in its frequency---is the Sanskrit particle iti (इति).

Both Sanskrit iti and English like can occur in the following contexts:
A. When quoting words actually utttered, alongside a verb of speaking:
(Skt-1) kathitam avalokitayā "madanodyānam gato mādhava" iti
"Avalokita had told me that Madhava was gone to the grove of Kama." [Mālatīmādhava I, p. 11; cited from Speijer[1]:§493a]

(Eng-1) "She said like 'I want to go too'."

B. Expressing the contents of one's thought:
(Skt-2) manyate pāpakam kṛtvā "na kaścid vetti mām" iti
"After committing some sins, one thinks 'nobody knows me'." [Mahabharata 1.74.29; cited from Speijer[1]:§493b]

(Eng-2) “And I thought like 'wow, this is for me'.” [OED, 2nd Supplement[2]; 1970, no earlier citations]

C. More general setting forth of motives, emotions, judgements etc.:
(Skt-3) vyāghro mānuṣam khādati iti lokāpavādaḥ
"'The tiger eats the man' is slanderous gossip." [Hitopadesha 10; cited from Speijer[1]:§493c]

(Eng-3) "I was like 'wow'!"
There are obvious differences between English quotative like and Sanskrit iti, including the fact that English quotative like precedes the "quotation", while Sanskrit iti follows it (in conformity with the general left-branching nature of Sanskrit syntax).

Further, Sanskrit iti doesn't have any of the other functions or meanings associated with English like. English like derives ultimately from Proto-Germanic *lîko- "body, form, appearance", while Sanskrit iti is built from the pronominal stem i-. In fact, iti still has pronominal uses, even in Classical Sanskrit, as in the following example.
(Skt-4) tebhyas pratijnāya nalaḥ kariṣya iti
"Nala promised them he would do thus." [Nala 3,1; cited from Speijer[1]:§492]
Amusingly, I find that (pretending that a parallel development has taken place in English) replacing "quotative" like with thus actually seems grammatical to me---though wholly unidiomatic, e.g.:
(Eng-4) "I was thus: 'Wow!'"
(Somehow I imagine that if thus had been recruited as a quotative in English rather than like, the use of a quotative marker wouldn't be so stigmatised, since there would be no association with filler like and, moreover, thus is largely used in formal registers of English.)

However, there is another element in Sanskrit which---though not as frequently used in this function as iti---actually is more similar to English quotative like in its syntax and semantics: yathā. Yathā is, properly speaking, a relative pronoun and is often part of relative-correlative constructions of the form yathā X...tathā Y "As X...., so Y". However, it can occur without correlative tathā, and in fact can have the meaning "like", as in the following example:
(Skt-5) mansyante mām yathā npam
"They will consider me like a king." [Mahabharata 4.2.5; cited from Speijer[1]:§470a]
Yathā can also function as a sort of quotative, but---unlike iti and like like---it precedes rather than follows the quoted discourse:
(Skt-6) viditam eva yathā "vayam malayaketau kimcitkālāntaram uṣitāḥ".
"It is certainly known (to you) that I stayed for some time with Malayaketu." [Mudrarakshasa VII; cited from Speijer[1]:§494]
(Or, maybe: "You certainly know, like, 'I stayed for some time with Malayaketu'.")
(Yathā and iti (since they occupy different syntactic positions) can also co-occur.)

So there is at least one antique parallel for the development of modern English like as a quotative marker.

Returning to the more commonly used iti, the following Sanskrit example---occurring when one of the heroes of the Mahabharata has performed an act of generosity so great that even the gods are impressed---I think is a great parallel for examples like "I was like, 'Wow!'":

(Skt-7) tato 'ntarikṣe vāg āsīt "sādhu sādhv" iti
"Then a voice in the sky was like 'Wow! Wow!'" [Mahabharata 14.91.15]
This line might be more usually translated as "then a voice in the sky said 'Bravo! Bravo!'", but there is actually no verb of speaking: āsīt means "was".

References:
[1] Speijer, J.S. 1886.
Sanskrit syntax. Leiden: E.J. Brill. [reprinted, Delhi: Motilal Banarsidass, 1973.]
[2] The Oxford English Dictionary, September 2009 rev. ed.


Wednesday, 9 December 2009

Nepali, Nez Perce, and Na'vi: On alien-language in Cameron's Avatar, with remarks on etymology and "Universal Grammar"

In order to lend authenticity to his film Avatar, James Cameron had the language of the alien Na'vi people designed by linguist Paul Frommer, as reported by Benjamin Zimmer[1] in his 4 December article "Skxawng!" (in his New York Times column "On Language"). Cameron apparently choose Frommer partly on the basis of his co-authored textbook Looking at Languages[2], where one of the exercises involves deciphering Klingon word order [spoiler: it's object-verb-subject] (Klingon is another linguist-designed language).


An interview with Frommer is available at the Unidentified Sound Object blog, in which Frommer reports on some interesting features of the language he developed. The language of the Na'vi (who look sort of like blue cat-people, see above) involves some typologically-unusual linguistic features, including: the presence of ejectives in the phonological inventory, specifically [k'], [t'], and [p'] (click to hear what these sound like), and---more interesting to syntacticians and morphologists---a tripartite system of case marking.

Like ejectives, tripartite case-marking is present but rare in human languages, found in the Australian languages Wangkumara and Kala Lagaw Ya (though these two languages are apparently unrelated) as well as in the Amerindian language Nez Percé spoken in the northwest of the USA (on which see further Cash Cash[3]). The tripartite case-marking system involves differences in morphological case-marking on (a) agents of transitive verbs [agentive/ergative case], (b) objects of transitive verbs [objective/accusative case], and (c) agents of intransitive verbs [absolutive/nominative case].

However, there are languages which are much less exotic (at least to me) that could also be seen as employing tripartite case-marking, including many Indo-Aryan languages such as Hindi and Nepali:-- see examples below (where nom=nominative/absolutive case; acc=accusative/objective case; erg=ergative/agentive case).
Hindi:
(1) लड़का कल आया
laṛkā-[Ø] kal āyā
boy-nom yesterday came
"The boy came yesterday."

(2) लड़के ने लड़की को देखा
laṛke-ne laṛkī-ko dekhā
boy-erg girl-acc saw
"The boy saw the girl."

Nepali:
(3) केटा हिजो आयो
keṭā-[Ø] hijo āyo
boy-nom yesterday came
"The boy came yesterday."

(4) केटाले केटीलाई हेर्यो
keṭā-le keṭī-lāī heryo
boy-erg girl-acc saw
"The boy saw the girl."
[Though the "accusative" case-marker (Hindi ko, Nepali lāī) in Indo-Aryan is not straightforwardly a marker of objects of transitive verbs, rather it tends to occur particularly on objects which are animate and/or specific--see Bhatia[4].]

The fact that the Na'vi language shares this feature with Indo-Aryan perhaps makes all the more appropriate that the name of Cameron's film is also Indo-Aryan. Avatar, from Sanskrit अवतार (avatāra), is usually translated into English as "incarnation", used to refer to gods assuming human bodies (e.g. the god Vishnu becoming Krishna). It also has an extended use in the world of cyberspace, where it refers to the graphic representation of a user or his alter ego. The sense in Cameron's film, I take it, actually draws on both of these meanings, as some of the human characters control Na'vi-appearing bodies.

Interestingly, in the Mahabharata, one of the two major Indian epic poems, where avatars are a cental concept, the term avatāra is actually never employed (Sutton[5]:156-7); however the concept is frequently alluded to (Biardeau[6]:1621n2, Hiltebeitel[7]:109n56) by usages of the verb avatr̥̄-, which literally means something like "stepping down" (prefix ava- "down, off" + √tr̥̄ "to cross over"). In fact, the verb avatr̥̄- is conventionally used in the Mahabharata to refer to people "stepping down from their chariots" (Hiltebeitel[7]:232).

Returning to Na'vi, in his Unidentified Sound Object interview, Frommer remarks that:
As I mentioned, there’s nothing in Na’vi that couldn’t be found in some human language—and that’s important, since humans have learned to speak it.
I found this idea that the Na'vi language is learnable by humans rather intriguing, since part of the Chomskian notion of (natural human) language is that it relies on biocognitive structures which are unique (at least on Earth) to humans (i.e. not present in any other Terran creatures). Would/could language as developed in an extraterrestrial species rely on biocognitive structures which would be equivalent to those underlying human language?

This reminds me of a story that Prof. Peter Lasersohn told in one of his semantics courses; paraphrased (as well as I can remember it):
Logicians and philosophers had long treated human language not being expressable in terms of formal logic. Richard Montague famously developed a system of formal semantics for language (Montague[8,9,10]); in one of the earlier accounts he states: "There is in my opinion no important theoretical difference between natural languages and the artificial languages of logicians; indeed, I consider it possible to comprehend the syntax and semantics of both kinds of language within a single natural and mathematically precise theory. On this point I differ from a number of philosophers, but agree, I believe, with Chomsky and his associates" (Montague[8]).
However, the Chomskian notion of "Universal Grammar" involves an abstract (but biocognitively instantiated) system which underlies all human language but is unique to humans. Montague, on the other hand, used "Universal Grammar" in the sense of a formal syntax and semantics which would be truly "universal", that is, applicable to any language, human or otherwise.
When Barbara Partee (a semanticist who was instrumental in popularising Montague-Grammar among generative linguists) explained Chomsky's sense of "Universal Grammar" to Montague, he was perplexed, remarking that he did not understand why linguists would adopt a human-only conception of "Universal Grammar" which would thus automatically disqualify them from being the ones the world would to turn to---in the event of humans making contact with aliens---for the decryption of extraterrestrial language.
References:
[1]Zimmer, Benjamin. 2009. "Skxawng!" On Language, New York Times, 4 December 2009.
[2]Frommer, Paul R. & Finegan, Edward. 2004. Looking at languages: A workbook in elementary linguistics. Boston: Wadsworth, 3rd edn.
[3]Cash Cash, Phillip. 2004. "Nez Perce verb morphology". Ms., University of Arizona, Tucson.
[4]Bhatia, Archna. 2008. "Animacy, specificity and overt object case marking in Hindi". Ms., University of Illinois, Urbana-Champaign.
[5]Sutton, Nicholas. 2000. Religious doctrines in the Mahābhārata. Delhi: Motilal Banarsidass.
[6]Biardeau, Madeleine. 1999. Le Rāmāyaṇa de Vālmīki. Paris: Gallimard.
[7]Hiltebeitel, Alf. 2001. Rethinking the Mahābhārata: A reader’s guide to the education of the dharma king. New Delhi: Oxford University Press [Indian edition].
[8]Montague, Richard. 1970a. “Universal grammar”. Theoria 36: 373-398.
[9]Montague, Richard. 1970b. “English as a formal language”. In Bruno Visentini et al. (ed.), Linguaggi nella società e nella tecnica. Milan: Edizioni di Comunità, 188-221.
[10]Montague, Richard. 1973. “The proper treatment of quantification in ordinary English”. In K.J.J. Hintikka, J.M.E. Moravcsik, & P. Suppes (eds.), Approaches to natural language. Dordrecht: Reidel, 221-242.


Monday, 30 November 2009

Nother post on nor

A recent Language Log post discusses the use of nor in the sentence:
"The snow fell nor did it cease to fall."
Since this topic touches on disjunction and ultimately on wh-words (interrogative pronouns), central issues in my dissertation, I can (almost) justify taking the time to investigate some of the antecedents of McCarthy's use of nor. Nor in the above sentence, as Mark Liberman observes, conforms to the sense in the OED's[1] entry 5a. for nor:
5. And — not; neither. In later use normally with inversion of subject and verb.

a. Following an affirmative clause, or in continuing narration. Obs. (chiefly poet. in later use).
Liberman offers discussion of modern and (late) early modern English examples in the aforementioned Language Log post; I shall concentrate on earlier examples, such as:
[1423] Guildhall Let.-bk. in R. W. Chambers & M. Daunt Bk. London Eng. (1931) 114 He shalle wirke..without fraude..nor he shall nat entermete of sekenes, sore, or hurte..vnknowynge to hym in eny maner.

[1492-3] in T. Pape Medieval Newcastle-under-Lyme (1928) 180 The aforesaid William shall delyuer all evedence and writings that belonges to the lands in the Newcastle, nor hurt nor truble the aforesaid John Leighton.

[1523] LD. BERNERS tr. J. Froissart Cronycles I. cxxxv. 162, I greatly desyre to se the kynge my maister, nor I wyll lye but one nyght in a place, tyll I come there.
These are the three earliest examples the OED gives for sense 5a. It is interesting to note that the first two appear to come from legal documents.

The OED suggests that nor derives from earlier nother1 (a contracted form of Old English nōhwæðer "neither", on which more presently), for which its earliest example means "neither of two preceding things or persons":
[eOE] KING ÆLFRED tr. Gregory Pastoral Care (Hatton) li. 399 Ne fornime incer noðer oðer ofer will butan geðafunge.
"Let neither of you deprive the other without consent."
As Mitchell[2]:§§1847-51 observes, OE nohwæðer/noðer cannot always be interpreted as a pronoun, as in:
[Blickling Homilies[3]:45.14]...þæt hi þonne ne mihtan nawþer ne him sylfum ne þære heorde þe hi ær Gode healdan sceoldan, nænige gode beon.
"[For the good teacher has said that, when the priest or bishop was led into eternal perdition,] that they could not be any good, neither for himself nor for the flock which they previously should have kept for God.",
where it plays a similar role to modern English neither.

This nother1 is not to be confused with a nother [sic] development which led to a form nother: namely the reanalysis of another/an other as a nother2, for which we find early examples such as:
[c1390] MS Vernon Homilies in Archiv f. das Studium der Neueren Sprachen (1877) 57 280 He wolde him say his onswere on a noþer day.
And, of course, this nother2 is frequent in the collocation (that's) a whole nother story, but it may be found outside of this formula, as in:
[1977] C. MCFADDEN Serial (1978) xxviii. 62/2 I'm in a whole nother space.
[1993] Wired Dec. 18/3 A new direction and a new name seem inevitable. But ‘tekkies?’ It seems too much like ‘Trekkies’, which invokes a whole 'nother set of connotations.
Interestingly, there is also a dialectal English (apparently particularly in Southwest England, if the prominence of Zomerzet zs in the last two examples is any indication) development which the OED suggests represents convergence between nother1 and nother2:-- neither nother, originally "neither one nor another", and thence "no other":
[?a1425 (c1380)] CHAUCER tr. Boethius De Consol. Philos. V. met. iii. 52 Who so that sekith sothnesse, he nis in neyther nother habit, for he not nat al, ne he ne hath nat al foryeten.
[1533] T. MORE Apologye 180 There are fewe or none good in neyther nother parte.
[1640] R. BROME Sparagus Garden IV. v, No sir, we come with no zick intendment on neither nother zide.
[1888] F. T. ELWORTHY W. Somerset Word-Bk. 523 There idn nother-nother lemon vor to be had in the town, nit vor love nor money, zo Mr. Baker zess.
Neither nother is also prominent in West Indian English, with the sense "no other":
[1957] F. A. COLLYMORE Notes for Gloss. Barbadian Dial. (ed. 2) 59, I ain't got neither-nother sixpence.
[1975] T. CALLENDER It so Happen 99 He never going look at neithernother girl again.
Nother1 also appears with the OED's nor sense 5a., as far back as the Old English of Beowulf:
Swā wē þǣr inne andlangne dæg
nīode nāman oð ðæt niht becwōm
ōðer tō yldum; Þā wæs eft hraðe
gearo gyrnwræce Grendeles mōdor
sīðode sorhfull; sunu dēað fornam,
wīghete Wedra; wīf unhӯre
hyre bearn gewræc; beorn ācwealde
ellenlīce; þǣr wæs Æschere
frōdan fyrnwitan feorh ūðgenge.
Nōðer hӯ hine ne mōston syððan mergen cwōm
dēaðwērigne Denia lēode
bronde forbærnan nē on bǣl hladan
lēofne mannan; hīo þæt līc ætbær
fēondes fæðme under firgenstrēam;
þæt wæs Hrōðgāre hrēowa tornost
þāra þe lēodfruman lange begēate.
Beowulf ll.2115-30
"We were happy therein all day long,
and enjoyed ourselves, until another
night descended on man. Then suddenly
Grendel's mother, ready to revenge her sorrow,
journeyed, sorrowful--- death had taken her son,
the war-hate of the Wederas [=Beowulf]. The ghastly woman
avenged her child, slew a warrior
boldly. Thus from Ashhere,
the wise counsellor, life departed.
Nor could the Danish people, when morning came,
cremate the dead one in the fire,
could not lay on the funeral pyre
the body of the beloved man: she had carried off the corpse,
held in fiend's embrace, beneath the mountain-stream.
That was for Hrothgar the most bitter grief
which had long befallen the ruler of the people."
Thus this use of nother1/nor (in the OED sense 5a. for nor) appears to have a long history in English. [Additional note: Nōðer here does not seem to mean "neither", in the sense "neither...nor", despite the present of in the sentence (see this comment on Languagelog), since on bǣl hladan "lay/load (his body) on the pyre" is really just a variation of bronde forbærnan "cremate in the fire" --- these aren't two different funerary options that the Danes have. But see this post for what is perhaps a clearer example from The Fortunes of Men.]

Old English nōhwæðer, originally a pronoun meaning "neither of two persons or things", from which nother1 (OE nōðer) derives, is itself etymologically-interesting. Nōhwæðer is morphologically composed of ne "not" + ā/ō "always" + hwæðer "whether". Without ne we find āhwæðer (with contracted forms āwðer, ōwðer, āðer), with essentially the sense of modern English either:
[KING ALFRED, Trans. of Orosius, 290.21] Þa oferhogode he þæt he him aðer dyde, oþþe wyrnde, oþþe tigþade...
"Then he scorned to do either, forbid it or grant it..."
More interesting is the original sense of hwæðer (ancestor of modern English whether): "which of two", as illustrated by Beowulf's speech to his men before his fight with the dragon:
'Gebīde gē on beorge byrnum werede
secgas on searwum hwæðer sēl mæge
æfter wælrǣse wunde gedӯgan
uncer twēga; nis þæt ēower sīð
nē gemet mannes nefne mīn ānes.'
Beowulf, ll.2529-33
"'Wait you here in the barrow, wearing mailcoats,
warriors in armour, (and see) which of the two can better,
during the slaughter-race, survive wounds,
of the two of us; this is not your adventure,
nor in the power of any man, save mine alone.'"
The predominant modern use of whether for introducing indirect yes/no questions was originally only one of its many functions, which including introducing alternative questions (note that, like other wh-words, in matrix questions it triggers verb-raising to the second-position):
[c1000] Ags. Gosp. Matt. xxi. 25 Hwæðer wæs iohannes fulluht, þe of heofonum, þe of mannum?
"Was John's baptism from heaven or from man?"
[1595] SHAKES. John I. i. 134 Whether hadst thou rather be a Faulconbridge,..Or the reputed sonne of Cordelion?
[1713] BERKELEY Hylas & Phil. I. (1725) 5 Whether does Doubting consist in embracing the Affirmative or Negative Side of a Question?
[a1822] SHELLEY Ion Pr. Wks. 1888 II. 115 Whether do you demonstrate these things better in Homer or Hesiod?
As well introducing as indirect alternative questions:
[c1000] ÆLFRIC Hom. II. 120 Eft ða Gregorius befran, hwæðer þæs landes folc cristen wære ðe hæðen.
"Then Gregorius asked whether the people of the land were christian or heathen."
[1610] SHAKES. Temp. V. i. 123 Whether this be, Or be not, I'le not sweare.
[1849] MACAULAY Hist. Eng. iv. I. 464 His neighbours might well doubt whether it were more dangerous to be at war or at peace with him.
The modern function as introducing indirect yes/no questions is attested early as well:
[c1000] Ags. Gosp. Matt. xxvi. 25 Cwyst þu, lareow, hwæðer ic hyt si?
"Do you say, teacher, whether it is I?"
[1470-85] MALORY Arthur VII. xx. 244 He mette with a poure man..& asked hym whether he mette not with a knyghte.
All of these functions can be seen to derive from the original sense "which of two". The morphological formation of hwæðer is curious however: it derives from Proto-Germanic *χwaþaraz (with cognates in other Germanic languages, e.g. Old Frisian hwedder, Old Saxon hweðar, Old High German hwedar, Old Norse hvaðarr (> Swedish hvar), Gothic hwaþar), which itself can be traced to PIE *kwo- "what, who etc." + the comparative suffix *-tero-.

What is curious is the use of the comparative suffix: all of these forms would literally be something like "what-er" ("more what")! (Though of course, since hwæðer etc. are used to inquire about "which of two", the comparative suffix, which compare two things, does make a certain amount of sense.)

Proto-Germanic *χwaþaraz has cognates in other old Indo-European languages, e.g. Greek πότερος, and Sanskrit katará-, the latter is found for example in the Rgvedic hymn on "Heaven and Earth":
katarā́ pū́rvā katarā́parāyóḥ
RV 1.185,1a
"Which of the two is earlier, which of the two is later?"
In Sanskrit, the interrogative pronoun can also combine with the superlative suffix (PIE *-temo-), to mean "which amongst many", as in the following Rgvedic passage praising Varuṇa:
kásya nūnáṁ katamásyāmŕ̥tānām
RV 1.24,1a
"Who now is he? Which among the many immortals?" [Lit. "Whichest of the immortals?"]
We're now of course a long way from the snow fell nor did it cease to fall, but following nor back along the path to nother "neither of two", and then off on the side path of its component morpheme hwæðer (mod. Engl. whether) "which of two", originally "what-er"(!), seemed an interesting enough detour.

References:
[1]The Oxford English Dictionary, September 2009 rev. ed.
[2]Mitchell, Bruce. 1985. Old English syntax. 2 vols. Oxford: Clarendon.
[3]Morris, Rev. R. 1880. The Blickling Homilies of the tenth century. London: Early English Text Society.
[4]Graßmann, Hermann. 1873. Wörterbuch zum Rig-Veda. Wiesbaden: Harrassowitz.