Monday, October 25, 2010

Broken Conlangs

So, on a sort  of dare, I was starting to learn a new natural language, something I haven't done -- if you don't count my annual two-day frustration with written Chinese -- in a long time.  Lesson one gave me an overview: the language was SOV and NA, VP was typically a participle and one of a small number of verbs that did the tense, etc. work, noun agglutinative and on the ergative model.  Then came the first of those working verbs, izan, 'to be (permanent or inherent relationship --Sp ser, more or less)(with personal pronouns): ni naiz, hi haiz (the familiar second singular, nice pattern emerging), zu zaude (common 2nd sing. -- that is often hauled in from somewhere else -- no problem),  hura da (whoa, Nelly!  where did that come from?), gu gara, zuek zarete (no pl. familiar), halek dira.  Oh, goody, a whole bunch of puzzles to figure out.

And then it struck me why so many conlangs are quickly boring: they are so damned regular!  Most are isolating, so they have no conjugations or declensions, or, if they do, there is just one and totally regular across the vocabulary. The fun, if there is any, is in trying to say strange things with a limited vocabulary -- and some don't even limit the vocabulary. Once you see the program and how it's implemented, there is little left to attract a person, intrinsically. Of course, a person may have extrinsic reasons to go into a language more deeply which may offset its boringness.

But, the reply is, regularity makes a language easy to learn; you can carry over what you learn for one word or construction to another. Maybe.  But maybe not.  Most people say the hardest part of learning a language is learning the vocabulary and the range of meanings of its words (which never match up with English exactly).  The rest is just drill and exceptions.  And the exceptions are easy because they are exceptions and so stick in your mind, like unusual thing in every area.  And they give some spice (though, admittedly, some pain) to the language.

I don't at this point know what to make of this realization.  I'm not about to advocate more irregularities in conlangs, though they would be a Pooh-bah  for languages in concultures that purport to have histories (and many such conlangs are irregular in various ways), I'm not going to downgrade a conlang for being to regular (although I may for its being boring).  I guess I'll just throw this out and hope it has some effect on the next generation of conlangs.

Tuesday, October 19, 2010

Articles in Logjam

(This is my opinionated memory of how things went and where they ended up.  I am open to factual corrections and will listen to other (wrongheaded) opinions.)

Logjam (Loglan up to about 1984 and Lojban) has basically three articles: la, le and lo (there are others out there, some of which arose from the concerns that follow, but none of which have much practical significance in either language).

The easiest is 'la'.  It turns whatever follows it (assuming solutions to RHE problems) into a name, a purely designating expression, one that refers to one "thing", however that plays out.  If what follows 'la' is a name form (ends in a consonant), this is obvious; if what follows is a legitimate Logjam expression (or close to it or whatever), then whatever content it may appear to have is to be disregarded: the referent of the name is not subject to the expression's supposed message applying to it.  As a historical example, Afraid-of-horses was not afraid of horses or much else, on the contrary, his enemies were afraid of even just his horses (the English loses some of the point of the original Lakota).  But even the latter point is not relevant to the applicability of the name; his son had the same name but was of a very different character.  The rules for 'la' has scarcely changed throughout the history of Logjam (though, early on, everything had to be whipped into name form).

The second article, the one often associated with "the", is 'le.'  This derives ultimately from the definite description operator in formal logic: ixFx, "the one and only thing that is F".  This expression, once introduced into formal logic (Peano, I think, but Frege had something like it , too), caused a problem (Russell to raise and "solve"): what if there is no "one and only thing that is F" and the expression therefore does not designate anything by virtue of its internal structure?  What if there are no Fs, or more than one?  Russell solved the problem by claiming that the expression wasn't there at all, but was, when appearing as an argument to a predicate, GixFx, part of a short hand for a much longer expression, which asserted that there is exactly one F and it is G.  So, if there isn't such a thing, for whatever reason, the simple sentence is, when expanded, simply false, as false as it would be if there were such a thing but it wasn't G.  Others took other views, mainly insisting that the expression was a real expression, not a shorthand and really referred to what it appeared to do when that reference was legitimate.  But when reference failed? Some said that sentences involving such expressions were simply not truth valued or had truth values other than True and False (logicians will take any excuse to go to a non-standard logic, which are much more fun than the usual one).  Others said that the expression always referred to an existing object -- the unique F if there was one, the null object, ix x=/=x if not. Since the identity of the null object is in the semantics and might be anything at all, the syntactic problem was solved.  But some went further in positing that the expression was simply referring, without any necessary connection to the property mentioned in it.  To be sure, if there are Fs it might be a good idea usually to have it refer to one of them, but even that was not a requirement.  That screwed up some of the standard theorems, of course, but had advantages in other areas.  Logjam 'le' comes out of this last position, with modifications: 1) the identity of the referent was not put off to the semantics but had to be established (at least in the speaker's mind) at the time of use, and 2) the referent could be plural.  The second modification raised problems within logic (that were eventually assuaged by plural quantification or a better understanding of L-sets) and the corresponding problems in Logjam, which were generally solved either by ignoring them (which turned out to be the right way to go) or by generating a number of half-hearted compromises, most of which ended up in new articles (for sets, for example, and masses, and mushes, and Lord knows what all else), and slopped over into 'lo'.


Probably, 'lo' started life closely parallel to 'le', differing in that it was nonspecific (or indefinite or whatever) and that it was "veridical", that is, the referent of 'lo broda' had to be broda.  The second feature follows  from the first, since, if you don't have a particular thing in mind, the only way to get to what you want is through what it is -- unlike the 'le' case, where, since you have a particular in mind, anything that gets a person to that thing is OK (though calling broda broda is generally a good way of doing that, but not always, as witness then do of The Crying Game or an early Mike Hammer story).  Somehow this notion did not catch on, or rather, other notions came into play.  Without rehearsing the arguments involved, 'lo broda' was variously explained as:
A.all the broda, taken disjunctively: that is, if even a single lion lives in Africa then lions live in Africa (the stock example).  This, of course, raised the question of how to say all -- or some group of -- broda did something conjunctively, working together (the stock case here is boys carrying a piano). So
B. a mass of broda acting conjunctively.  But then how say they acted disjunctively?  Eventually, this generated a new article. But then
C. the substance of which broda are made (stock example is "There's dog on the windshield, the tires and all down the left side" after hitting  a large dog with a car).  This arose from the notion of a mass as a product of a blender, which might also act in broda shaped units under certain conditions, which proved hard to define.Whence
D. brodahood, which is indefinitely related to the sense of 'broda' (and comes somehow from Quine's remark about the meaning of 'gavagai').  This solves the problem of there being no broda, for brodahood is presumably always around.  On the other hand, it  (like the others above) make predicates mean something other than they mean or involve some other way to get from a warehouse of broda to the individual broda that  predicates apply to.  This was done by hiding the mechanism as much as possible, using default quantifiers and default expressions after the quantifiers, too.  So, in 'lo broda cu brode' either 'brode' is not a property of broda but rather of the set/mass/goo/hood of broda (one which the set has, presumably, by virtue of individual broda having some other property not mentioned) or 'lo broda' does not really refer to the presented abstract object, but to some unspecified member/participant/chunk/avatar of that object.  Not very transparent for a logical language.


Lojban inherited this muddle, went through most of the same discussion again and finally settled on basically choice A, with bits of the rest creeping in as needed.  So, 'lo broda cu brode'  probably really means "some member of  lo broda is brode"  Or "members", of course, but that raises another issue, namely that what you get out of a set this way are not individuals, but a subset, which, being a set, presents the same problem anew.  Enter xorxes with the reasonable suggestion that 'lo' be just like 'le' except indefinite (or ...) and hence veridical.  We are still dealing with sets though, so the problem above arise still -- and using the empty set (which is always available) when there are not broda has some undesirable consequences.  So there emerged (again) Mr. Broda, who combined somehow (hours spent trying to figure out how) the best of the indefinite set with the best of -hood, to cover all eventualities.  And then (ta-DA!) in the midst of all this, xorxes found a book about plural quantification and plural reference, where reference and assignment were not a function  from a referring expression to an object, but rather a relation between an expression and objects (one or many).  Of course, such a change from standard logic was anathema to tired old logicians who had lived with Cantor sets for years; it looked not to be a consistent, coherent logical system at all.  But then it turned out to be exactly Lesniewski's mereology (or Quine-Goodman-Leonard's part-whole logic), a known -- but previously unintelligible -- system with all its metalinguistic credentials in place (and now intelligible).  So the problem brought on by always needing a single referent and that being something like a set disappeared.  The referent could be several things and, because of the way the logic works, these could individually or collectively do various things, without hidden mechanisms (collective action is the norm, independent action is indicated by external quantifiers).  


So, 'lo broda' refers broda, one or several or many, not specified beyond what is given in the sentences in which the expression occurs.  These broda take on properties individually or collectively as indicated by predicates which can take individuals separately or collectively as arguments.  All the various problems which arose around 'lo' (and occasionally 'le', since the same problems occur there) disappear.  


Except when there is no broda. Plural reference or Lesniewski sets don't have an empty set to fall back on.  Sometimes, bits of Mr. Broda have emerged to deal with this, but the basic solution seems to have been in defining the appropriate universe of discourse (what things there are available to be referred to in a given context).  When last we went round on this, xorxes seemed to be holding a view on this that led to some bizarre (to me, with my ingrained habits) consequences.  But there is another approach which avoids these consequences but still yields all the right results.  




    

Friday, September 4, 2009

C-sets and L-sets. Draft

C-sets. or Cantorian sets, are the usual sets of set theory. They were first (for practical purposes) described by Cantor in his work with infinity, which gave them a slightly shady reputation, enhanced by the discovery of an array of paradoxes in the earliest attempts to axiomatize the theory.  Eventually,number of axiomatizations that avoided the known paradoxes in various ways were devised.  All the various axiomatizations are incomplete, of course, since arithmetic is and can be defined within them. They also differ in various ways at their outer edges (is the next largest cardinal after Aleph-null, that of the set of natural numbers, R, that of the set of reals?). But they agree in the basic area we are concerned with, namely, simple set, their subsets, and simple operations on these. Briefly, a set is different from its members, in particular, {a} (the set whose only member is a) is different from a (the thing in it). Further, from a given set, {abc}, say. we can get other sets using only the members of the original set (its subsets): in this case {a}, {b}, {c}, {ab}, {ac}, {bc}, and, for completeness, the original set {abc} and the empty set{}, which has no members. The order in which the members are listed is not significant, but the fact that these members are in different sets will be: {{a}{b}} is different from {ab}, of course, and {{ab}{ac}} is different from {abc}, and so on. A set, once constructed, can now behave as an element in further sets, as just applied, and so, all of the subsets can be gathered together into a single set (called the power set because its size for a set with n members is 2 to the power n). In this way, quite large sets can be built from relatively small beginnings (the natural numbers, in one approach, are built up from the empty set -- about as small as you can get -- by taking its power set, {{}}, joining the two into a new set, {{}{{}}}, then combining this set with its members to form a set with three elements, and so on forever).

L-sets were developed first by Lesniewski and independently by Leonard and Goodman (and Quine).. They are mathematically more untidy that C-sets (there is no empty set, so a set with n members has only 2^n -1 subsets) and also less useful mathematically (it is hard to get bigger sets). But they prove to be more useful in representing many situations in language (and thus expanding logic a bit). For L-sets, {a} is the same as a, further, {{a}{ab}} collapses to {ab}, {{ab}{c}} = {{a}{bc}} = {{a}{b}{c}} = {abc}, and so on. That is, a set is always a set of its ultimate components, even if we talk about intermediate subsets.

In set theories, of course, sets are individuals (the values for variables, the referents of constants, and so on) and so, they have properties and enter relations. But, in the theory, these attributes tend to be either dummies or set-theoretic ones. Little is said about sets and ordinary properties and relations , "carries a piano", say, or "wins a race." But, if we have any intuitions about sets, they don't seem to be the sorts of thing that carry anything or even enter races, even though their members might be.

And yet, something formally very like sets do these thing in everyday language: "The boys carried the piano" need not mean that each carried it by himself; it might mean that they carried it together, acting formally as a set. (Each of the boys participated in carrying the piano, though each's exact role is unspecified.) And, of course, a team (looks like a set in a general way -- a number of things conceived of as together) can win a race even if only one member runs (the other participate by being on the same team, I suppose -- that is a characteristic of teams among sets). And it turns out in many cases to simplify thing quite a bit to take plural references as to sets, rather than somehow to each of the several individuals we take as making up the set. We can disambiguate if we need to, but often it does not matter (as the race case suggests). So long as the piano gets carried, we don't care how the boys do it. If we really want to specify that each carried it alone part of the way, then we can say so:'"each of the boys carried the piano."

While we could do this sort of thing with either C-sets or L-sets, L-sets seem the more natural. Bunches of things such as we have in mind don't seem to grow into bigger bunches by subdividing and recombining, Further, if we are mistaken about the referent being plural, the same pattern applies since {a}=a, the singleton set reduces to its member. Even the lack of an empty set, so damaging to the mathematical uses, proves useful here, since a reference to a something having a property which nothing in fact has automatically renders a sentence false for L-sets, but makes a reference to the null-set in C-sets and the null set has many properties, mostly irrelevant to whatever we were talking about or relevant but holding only through logical tricks -- neither desirable situations.

It should be noted that some people raise objections to using L-sets to treat languages. The objection runs that doing so compels the language to implicitly recognize the existence of such things as L-sets, since they are necessarily in the range of quantifiers in the language. While I am not sure why this is a problem, especially in a language, like English, which regularly interchanges L-set words like "bunch" and "group" with simple plurals, I give way to others' quest for ontological purity and say, that we ought not think of L-sets as something different (or over and above) the members considered together. This is, of course, very easy to do, given the transparency of L-sets. It requires some changes in the way we describe the logic of the situation, but only minor ones. In particular, we need to allow that quantifiers may take a number of instances simultaneously and together and that we have a way of pulling out particular instances from these. That being done, we can deal with plurals as really being about several things rather than one thing of which the several are members, which does seem more natural. The point is that the logics of the two approaches are exactly the same and even the two metalanguages, while not the same, are directly translatable the one into the other and so totally congruent. As an old-timer raised on set theory, I stick to the locution I am most comfortable with, but I try always to include the other reading as well.






Zipf's Wall -draft

Zipf's Law is a more precise formulation of the obvious (once you think about it) proposition that common words tend to be short and rarely used words long. The full version does the math and gives correlations between frequency and length (relative to the norm for the language). This is all, of course, descriptive of how natural languages in fact work, not a prescription of how they should work. But the underlying logic of language as a human instrument gives it some projective power.

In creating a language, then, one wants to keep this pattern in mind. In particular, one does not want to set the language up in such a way that common topics of conversation will inevitably involve words that are overlong for their commonality. This is not a problem for a language like Esperanto, which can borrow freely from the languages around it (with a few -- often ignored -- restrictions), nor even for Lojban, which makes enforced restrictions on borrowings but ones that cost only a syllable or two.

It is a problem, however, for languages with a fixed base of concepts and no means to add new items. Depending upon what the base is and the structure of the language, the problem of too long words can arise sooner or later. In this context, "word" should probably be "phrase," since many languages do not have a means to construct new words with fixed meanings (combinations of the basic concepts) but must do the work stringing words together. In any case, there will come a time when repeating a certain referring expression come to be felt to be too onerous and the cry goes up for a replacement. Aside from simply giving in to the plea and adding a new concept to the base pile, here are a few strategies to meet this issue.

1. Simplify the definition. If the concept dog is represented by something like "furry beast that we have around the house for protection and to play with and take hunting," we can surely trim this back to something like "beast that ...." with only one or two things in the gap. This definition is, of course, purely accidental, i.e., gives neither necessary nor sufficient conditions for being a dog, but the strict definition is going to be either the Linnean binomial or the biological description behind it, and that is likely too long also -- aside from likely not fitting into the language's patterns. So definitions are likely to be contextual and in that context finding an appropriate phrase is simplified. We need, perhaps, only to distinguish dogs from cats and so any short thing that does that will do.

2. Choose your base wisely. This is sorta ex post facto, but presumably you can go back and revise before too many people get too committed to the original. NSM offers a short list of concepts which are said to occur in all languages and with which all others can be defined. The definition process is complex, however, and does not lend itself to simple expression construction. Still, any starting point should be sure to cover those concepts. Swadesh's list of concepts you can be pretty sure to find expressed readily in any language is about four times the size of the NSM list. It is meant primarily as an entryway into a new language: you can be sure there are words for these, and once you get them they will enable you to ask about other things as well -- but there is no claim that everything else can be defined in terms of these. Basic English starts with a list about four times the length of Swadesh's but does claim to be able to say anything using only those words. But, as some examples will show, the problem of Zipf's wall is simply ignored (and so people don't use BE much). BE does also claim that its words are among the most commonly used in English and so provides a further guide: even though the list is too long for direct use (probably), if you have it covered with appropriately short words you are well on your way.

3. Invent a slang. This is a bit of a cheat, but one that natural languages use all the time. The version that is likely to occur to a language creator is apocopation, dropping out stuff. If "Geheimnis Staat Polizei" is something you say a lot, lop it down to "Gestapo." Everybody knows what you mean and you are not really introducing a new word, just saying the old one faster. American go for acronyms (initial letters of the underlying words), but our former enemies seem to prefer slightly larger chunks of the original (see above and "Ogpu" and "Sudoku") -- which is often more informative (a number of CYAs have rather conflicting agendas and "confusing" two of them can make for really bad jokes). The fullest form of this sort is the word-forming rules in Loglan and Lojban, although is not quite what is intended here, since the results are new words and even have definitions which could not be derived readily from the sources.

Of course, there are other forms that slang can take. One is frozen metaphor, where a word that does not mean what you want but can be connected with it in some poetic (or not so) context comes to stand for what you want (there is a nice rhetoricians' name for this but I can't remember it now; I'll ask my resident English major). Assuming that the picked up word is incongruous enough in the context where it is used for what you want, this will work -- perhaps with a little training (both sheep and cotton have been compared to clouds, so one might use "cloud" for either or both of them -- you don't pick sheep and you don't herd cotton and you don't do either to clouds (ah, but you do make cloth from both sheep and cotton, though still not clouds)).

Or the reference might be indirect, through an accidental or a causal intermediary. Rhyming slang is a good example of this: since "Bees and honey stand for money" so does "bees" alone. Along the causal line we have all the American terms for money, listing the essential it buys: "dough," "bread," "bacon," and so on.

4. Expand the meaning of some basic terms. This might be viewed as a case of the last sort, but it comes into play at a different level, as an official part of the language. Again, natural languages provide models, as when "times" went from several points or periods in time to multiplication.

Actual created fixed-morpheme languages use some combination of all of these dodges, as they must if they are to achieve their goal of saying all that needs to be said, but with limited resources. Critics from the outside tend to take these moves as a proof that the program of such languages cannot be accomplished, rather than noting the ingenuity of language (and language creators, of course) to deal with situations as they come to the fore.


Friday, August 28, 2009

Just noodlin'

I got to thinking the other day about what kind of language I would like to create if I were to go into the active phase of this game. I came up with two features, one phonologic and the other everything else.

For phonology, I would like a language without vowels. Any continuant can be a syllabic peak, so that need not come from within the triangle i-u-a. And many language use some of these continuant consonants (especially lmnr, but Chinese us s-sh-sr) occasionally or in paralinguistic utterances (Pfft! Psst!) . These usages often get disguised with added vowels in the orthography, but in the language I have in mind there would be no vowels to begin with, so no temptation to add them (unless the habit is so strong that one nominal vowel would be used throughout).

Maybe some implosives and clicks, too?

For everything else, I think of Whorf's occasional almost intelligible formulations of SWH and of his idea of what the world is really like and what language would bring us to that perception and wonder how to build such a language. He worked with Hopi and Menominee, in which (I gather) most sentences reduce to complex verb forms, subject and object (as we SAE speakers say) being incorporated somehow. I have to assume, to get close to the ideal BLW was after, that the incorporated bits were also verb forms and that the notion of a verb here ceases to have its distinctive value (from nouns and adjective and ...) and means a reference to an action, motion, stasis, etc. in the flow. But (even after looking at bits of Hopi and Menominee) my SAE mind cannot visualize how to do this (and maybe go beyond what happens there). Maybe I should go read a few thousand pages of Li Er and Whitehead and Hartshorne.

I think these ideas arose for me out of the languages I have played with and the stuff I have read and taught over the years. toki pona claims a Daoist inspiration and has very fluid grammatical categories (though not syntactic slots), Loglan/Lojban started as a test for SWH, albeit not a very appropriate one, I think. And years of reading Daoist and Mahayana Buddhist literature has put me in the midst of a landscape of constant flow -- or at least instant ontic replacement.

I don't suppose I would derive any SW effects from this language, because I would have to get those effects in order to construct it correctly. And that might be an interesting thing to try, if I ever figure out how to begin. Some suggestions are quantum mechanics, ordinary mechanics, and hydrodynamics, all of which start out with things v things (except maybe the first -- and Lord knows what it starts with).

Well, I can mess with the phonology alone anyhow.

Monday, August 17, 2009

aUI -- the language of space

aUI (capitalization significant) was created in the 1950s by John Weilgart, an Austrian-born psychiatrist working in Iowa. (We can discount the story that he learned it in an instant from a little green man when he was a child of 5 on internal evidence alone: the precise fit with the English alphabet - including some pushing to make the fit -- the frequent coincidences of aUI and English or German words, the rigorous SAE grammar, and so on for quite a while). He publishedaUI The Language of Space first in 1961 and continued working on it until at least 1979, when he published the 4th revised and expanded edition, with the further subtitle Pentecostal Logos of Love & Peace. The basic language changed little over the years, the new books added new ways to approach the topic, new stories (apparently autobiographical, but probably not -- see above), and new commendation from various scientific and "scientific" sources. There seems to have been an occasional flurry of interest in aUI: a now defunct list, a commercial site (also defunct) for Cosmic Communication Co. (run by a daughter?) and a recent Facebook family with a handful of active members.

Purpose

aUI is a philosophical language, i.e., words bear their meaning on their faces: related concepts have related spellings and the spelling defines (at least delimits) the concept named. Thus, aUI aims to clarify our concepts and thus avoid misleading speech and propaganda. This, if universally adopted, would end misunderstandings at all levels, leading to peace, love and maybe the parousia. So, it is intended also as an auxlang -- indeed, as a replacement language -- for the world (indeed, for the cosmos). In the meantime, it is useful in logotherapy, helping confused people (at whatever degree) to get clear their thoughts by expressing them analytically in this language.

Tools

Weilgart uses the English alphabet (and, covertly, parts of the German and Scandanavian) to the fullest: all the letters, plus all the capitalized vowels and both sets of vowels underlined. Each of these has a sound and a meaning (and the sound or its manner of production is somehow iconic of its meaning). Words are then constructed by concatenating sounds to show interaction of meanings. All of this totally a priori. He also provides a native aUI alphabet in which each sound is provided with an iconic symbol for its meaning, perhaps less a priori, since some human conventions clearly play a part.

Specs
Phonology

The alphabet has it usual English values except that
the vowels have the Italian values, lowercase are shorter and generally lower than upper case
underlined vowels of both sorts are nasalized
j is ezh
c is esh
q is umlaut o (Mach den Mund rund und sag 'e')
y is umlaut u (ditto but 'i') between consonants or spaces; before vowels it is y.
underlined (and usually capitalized) Y is nasalized
Orthographically, the use of capital L and capital Q are encouraged (to prevent confusion with 1 and I on the one hand, g on the other). Otherwise capitalas are used only meaningfully with vowels and with borrowed names.

Consecutive vowels do not diphthongize but are pronounced separately, though without a marked break.

Separate words do have a marked break between them, since run together they might form a single word (though one related to what is intended).

Stress accent (which is also higher pitched) falls on the nasalized vowel, if there is one; on the long (capital) vowel, if there is one but no nasal: and, in the remaining cases (neither nasal nor long, or two or more of the dominant type) , on the penult or as near as possible while meeting the dominance requirement. Some words, when the stem of a verb, keep their original accent even if syllables are inserted after.

Morphology

Each sound is also a morpheme with an assigned meaning, an idea it conveys. Words are formed by first defining the idea you wish to convey, arranging already defined ideas as modifier and modified, then concatenating the two words in that order. This starts with letter-letter pairs, but such pairs and their extentions often become fixed words, which then act as a unit in later definitions. So the structure is basically right grouping but with possible left grouped chunks (which may be right grouping internally, of course). As a matter of etiquette, when introducing a new term in writing, the left groups should be marked off with dashes, since the straightforward form might be capable of many interpretations (though context and familiarity with the standing vocabulary do allow for fairly rapid understanding). The morph y, not, etc., is particularly like to form tight left groups. the polar opposite of the group it modifies (cf. Eo. mal-) (in the native spelling, the bar which is the symbol for y extends over the whole modified block, so is more clear than either the spoken or the English-written form).

aUI words are generally concrete nouns originally. Any of them can be made into an adjective by suffixing m, quality. From these in turn, abstract nouns can be made by adding U, thought, mind, etc., and then, from these, words for concepts by adding z, part, etc. Adjectives may serve as adverbs or become official adverbs by suffixing Q (O umlaut, remember), condition, manner, etc. In all of these, the original noun remains as a left unit within the right grouping.
Neither gender nor number is required, but, if wanted, plurals are formed by inserting (or suffixing) n, number, after the last vowel. Gender goes unmarked but can be part of a word in the course of things, with the (nonfinal) components v, male, or yv, female, occurring (pronouns use these ro modify words for the right sort of thing: u, human, os, animal, living thing, io, plant, light-life, Es, thing, material object; the resulting words are also the basic words for male and female of the given type thing.)

In a similar way, nouns can be verbed by add v, activity, etc., and then become available for a number of modifications, basically the handy appendix to your Latin textbook:
Imperative: insert r, good, positive, etc., before v (before yv in passive verb forms.
Passive: as just noted, insert y before v, creating the opposite of activity.
Past tense inserts pA, before-time, before v
Future tense inserts tA, toward-time, before v. These can be repeated and mixed to get past perfect and future perfect and the (unused) converses.
Participles end in Am, time quality, added after the v for the present, dropping the v in the past tense to give -pAm for the past. and shifting the A around the v in the future tense to give -tvAm.
Causative verbs are formed by inserting (or prefixing) v before the first vowel (or, if that is modified by y, just before that y).
Optative mood: - -Or-, feeling-good, like to, before v.
Subjunctive (contrary-to-fact) suffix -yEc, not materially existing.

There are a few other frequently used patterns, but they all, like these, are simply applications of the general pattern for constructing words.

Syntax


Weilgart says "There is no special grammar, All elements of meaning and their combinations still mean what they say. The rule is we talk 'as clear as we must, as short as we can.'" This seems to mean that, if it works in English, it works in aUI, subject to the following overriding fact.

aUI is a rigorously SVO language, with AN modification structure (as in word construction) and prepositions in lieu of cases -- except direct object is positional, right after the verb. Prepositional phrases tend to come at the end, after the object. Relative clauses are not inverted, nor are questions, the relative or interrogative word comes at its natural place in the order. If the relative word (starting with x, relative) is buried too deep in the structure -- object of a complex verb, say, or a preposition, a warning marker, xQ, relative condition, may be placed at the head of the clause. If there is no question word in the question (so yes-no questions), hI, question sound, is placed at the end. There is no distinction between restrictive and nonrestrictive relative clauses, although Weilgart does seem to use commas in the American way in writing.

Indirect discourse is a regular sentence introduced by Uf, think this, in the (usually) object place in the expressing sentence. Indirect questions seem to be just questions set off with a colon. Direct quotations are merely enclosed in quotation marks (sometimes with a colon before as well), which have no spoken matches.

Conjuncted expressions occupying various slots do not require different conjunctions (it works in English, ...). "Or" has both simple gaf, room this (don't ask), between alternatives and an "either ... or" form with gaf before each alternative. But "if," Qg, qualified inside, does not appear to have a matching "then," for the temporal yfA, not this time, clearly won't work (we might try something like fQ, this condition, if something like the right sense of "under" or "on" were available). The contrary-to-fact sense of conditionals is carried by the verbs.

Predicate adjectives occupy the V slot preceded by c, is, exists; predicate nouns do so with various extensions of j, identity, or z, part of.

Semantics

Starting with only thirty some concepts to define everything means that the initial concepts must be very broad indeed and that the means of combining them various. The first point means that one concept in aUI may seem like a random mix of several concepts, distinct in English. Presumably (I haven't done thorough research here) the definitions in aUI of those English concepts will help to see the unified nature of the aUI concepts. Similarly, since there is only one way of showing relationships, the exact nature of the relation may be obscure. Happily, Weilgart provides a number of discussions of particular cases as a guide.

The basic pattern is, of course, differentia and genus: picking out one subconcept (or subset of things) by indicating how it (they) differ from the rest. Thus, from s, thing (the notion seems to involve boundaries setting off from others), by adding a, space, we get as, place, a delimited bit of space. Similarly, As, time thing, instant, and Us, thought thing, (individual) thought. Another way of breaking down concepts is descent to instances: fa, this space, here (though fas would work as well). An interesting case of this approach is fu, this human, I, me, a nicely oriental touch (without the kowtow), saving a lot of concept space. Sometimes, however, the relation is part of the compound itself: gaf, or, instead of this, in place this.

Reading Weilgart's translation dictionaries and, more interestingly, his encyclopedia of compounds and his explanation of the "100 basic" ones, results in one of two responses,usually: "Oh, now I see it" or "But why doesn't it mean...," often simultaneously; the third, normal second speaker, response is "But wouldn't ... be better for that?" In the end, one has to say that one constructs what is right for one's concept, which may not be the same as someone else's though they use the same English word. Weilgart tells (whether a report or a thought experiment is unclear) of a group of children learning aUI and being asked for the aUI word for "love." The results are all over the place, but each has a plausible claim to be right; indeed, all are right -- for what the speaker means by "love" at that time. So aUI can convey simply shades of meaning that would be difficult or impossible to convey in English, say.

Discussed Problems

I have no evidence of an active aUI groups working over the material in the book. Clearly, a few people have done some things with the language, but they have left few records. Outside observers, however, have pointed out a number of things, mainly having to do with presuppositions (prejudices?) embedded in the language as presented. The other comments have been about the essential weirdness of the language as a human language (some evidence that it did come from little green men, perhaps, or just the result of being a consistent philosophical language). One instance of this is the lack of special status for the personal pronouns (in so far as there are such, separate from generic words). We saw an example of this in the first person case, fu, this person, but it carries through the rest: fnu, we, bu, with-person, and the plural bnu (the person with you in the conversation). The other is the deviation from the almost universals of human languages, the -m- in words for mother, for example (sometimes lost in later sound shifts, to be sure). ytLu holds little hope for this: not-toward round person = from-woman, the woman you come from.

This shows the earlier mentioned problem, presuppositions. Why should a round person, Lu, be equated with a woman? (My IHOP experiences show men at least on a par with women in this area). Actually, this case is a part of the evidence that other people have worked on this language, for the earlier -- and still legitmate -- word for woman was yvu, not-active human, passive human, and the stem yv is still the normal qualifier for females in other species. And examples like this (though not so egregious) can be found on every page. On the other hand, "purely by chance," some things come out familiar, like bru, together-good human, friend.

The definitions also presuppose a certain state of science, more or less the current one as popularly understood. So, for example, elements are named by their atomic number suffixed to Ez, matter part, element, so oxygen. atomic number 8, is Ez8 [the numerals are the symbols for the nasalized vowels, in the order aeiuo (so we go from alpha to omega) AEIUO. nasalized O is written 0, of course, but does not stand for zero, only the place holder in decimal notation; the real zero is nasalized Y, also written with 0, but never in strings: nasal O is always preceded by at least one other number, a multiplier on 10]. While this is not likely to change or to be different on another planet enough like ours to have recognizable science, the color terms are less sure. These are formed by prefixing numbers to i, light, the numbers corresponding to the position of the color on the spectrum, going up from red =1. Quite aside from questions about other color ranges (less than this or more, or shifted) the list is strange in that it has green as 3 and violet as 5, but not orange, which is 12i, red-yellow light (another type of conection, mixtures or going together -- but how distinguished from twelve?).

Comments

Though I have studied aUI off and on for 30 years (an awful winter in Iowa for a start), I have never lived in it or even learned a bit of it, so I cannot comment on how the language works as lived. But viewed as an object, it is an interesting specimen.

It is, first of all, a remarkably complete philosophical language. With some (mainly early and remarkable) exceptions, philangers have been so concerned to get the right set of basic concepts and to put them together in just the right way that they have never gotten beyond a vocabulary and even that often only writable, not speakable. aUI is speakable (OK, so bnu may not flow easily, but it does come) and has enough padding (though we don't call it that) to buffer the worst possibilities that might arise -- and enough various ways of combining to avoid most of these possilities any way. It has a clear phonology attached to its concepts and a plausible (or at least mnemonically useful) explanation for the association of each sound with its concept. It has the framework of a grammar, which is not always "do what you do in English (or Latin)," though it does fall back on that from time to time.

The definitions/word construction in aUI are usually interesting, sometimes because they are horribly wrong, but often because they insightful or at least thought-provoking. Many of them might serve as guides -- either directly or because of the process used -- for constructions in other semantic prime languages.

I suppose that Weilgart's aim would not be fulfilled by this language. While people might be able to express what they mean more precisely and thus understand what they are saying to each other, this would not bring an end to misunderstanding and propaganda and war. People lie, and those in positions to generate propaganda and war more so than most. So, while what I say may be clear as a bell, it may not be what I think. Not everyone can be trusted to put yr, bad, (Lojban mal-) in front of every derogatory word they use. But among trustworthy people, clarity is nice and -- assuming the trustworthiness -- imagine how many disasters of one sort or another would have been avoided, if, instead of "I love you," each had spelled out what they meant.

Monday, August 10, 2009

NSM -- Natural Semantic Metalanguage

This is a linguistic approach founded by Anna Wierzbicka in the 1970s and carried on by her and her followers, mainly in Australia. The central idea is that there are a small number of concepts (about 70 -- 63 the last time I counted) and sentence types which are realized in simple forms in all languages and which, taken together in a given language, allow one to define all the words in that language in that language. Both the concepts and the sentence types have been arrived at empirically, cutting items out of an original intuitive list as ways to define them were found, adding to the list when as yet unsolved problems arose. The word list and sentence type list might then be taken as an empirical minimum for language (though this is not the point of NSM).

For constructing a language, however, this is probably not the best guideline (assuming you want to start with semantic primes or even just the smallest possible vocabulary) . For one thing, it is designed to be used in defining other terms, not in conversation or narrative exposition, so, while it does contain soome words you would need immediately (for I and you, for example), it lacks others (day, for example, or path).

For another thing, the definitions NSM provides are not (or rarely are) simple isosemic phrases of the cat = domestic felid sort. They tend rather to require an imaginative journey. Think of a situation, specified in appropriate detail, and then the word to be defined will be the appropriate thing to say: a broke b is appropriate to say when 1) a did something to b, 2) because of this something happened to b at the same time, 3) it happened in one moment, and 4) because of this afterwords b was not one thing any more. While this looks about right, it is not clear that it can be collapsed into a replacive definition and so that it could be use for the most common sort of introduction of new terms into a conlang. It is also not clear that this process will really work in more complex cases; the imagined situation may call up for the native speaker some other notion than the one sought -- or may not even apply to his experiences: the definition of green, for example, requires imagining a field of grass or other green plant, which may not be in some people's experience at all.

For all that, the theory is a viable one in linguistics, often ably attacked and often ably defended (and occasionally not so much of either). For more detailed information about the genral theory and its detailed applications -- and its controversies -- see the bibliography at
www.une.edu.au/bcss/linguistics/nsm/.