Showing posts with label methodology. Show all posts
Showing posts with label methodology. Show all posts

Saturday, July 28, 2018

Scripts, Golems and Ideal Types

In an attempt to bring together the viewpoints of the digital humanities on the one side and the knowledge representation community on the other, it is instructive to ask what the commonalities and differences are between scripts (Schank & Abelson, 1977), Golems and Ideal Types (Weber, 1912).

Scripts

Scripts are supposed to structure habitually reoccurring events. They have to do with how people internalize appropriate behaviors and how to negotiate situations. The classical case is eating at a restaurant. There are actors for capturing the participants; there are roles that these actors play in the various scenes that structure the event, and there are rules of continuity, consistency and temporal ordering that hold between the scenes and the events. Non-actors or matériel is present as well and has its own rules of continuity between the scenes. Not all scenes are required, and some scenes can be repeated multiple times. 

  • At the restaurant, the patrons, as well as the serving staff, are actors.
  • The other patrons belong to a backdrop unless people are sharing tables in a full cafe or similar.
  • The menu and the tableware and the condiments form part of the matériel that is utilized in the individual scenes.
  • On a bad day, the scene of an order being rejected due to lack of ingredients might repeat several times. That scene might be elaborate if the waiter has to check each time with the chef if there is any clam chowder, bass, salmon, lobster or shortcake left.
  • The action relations are keyed to the scene of specific types which structure the roles; for example, the paying of the check is an economic transaction or a purchase, with the patron in the role of the buyer, the restaurant as represented by the waiter in the role of the seller. At the same time, it is a praising if the waiter receives a tip or a criticizing if the tip is small or withheld altogether. 
  • Different matériel may require different sub-scenes structures; for cash, there is often change brought back; for credit card transaction, a receipt has to be signed to complete the transaction. If the payment is a gift card or a voucher, none of these might be necessary, unless the payment is partial. 
All in all, a pattern matching grammar for events that talks about expectations, down to the {n:m} notation of repetitions, and what one is going to find.

Golems

While the statements about scripts, roles, actors, matériel and similar cache out as rules with universal or existential quantification, the process of authoring these statements is sometimes not as intuitive as one would like.

Taking a leaf from prototype-based programming, Golems sketch a situation using generic terms that stand in for the involved types. If a script is more like a narrative or a "what I did last summer" essay, the Golem is more like an explanatory drawing in an encyclopedia.


(Image from the Public Domain Pictures Net)

Given such a setup, one can easily speak about the sternum being connected to the second (true) rib "through the intervention of the costal cartilage anterior." The naturalness of the authoring comes from bracketing the need to be specific about quantifiers, which are implied. Furthermore, there is a completeness to the golem that can be exploited to note that any connection not enumerated is therefore implicitly absent, akin to a completely enumerated collection type.

On the whole, while the script appears to be a more flexible form of representation, the Golem is a more concise form of representation that has a clearer affinity to natural language use.

Ideal Types

The notion of ideal types was introduced by Max Weber in his essay on the objectivity in the social sciences. Weber pointed out that expectations and recognition are guided in the social sciences by superimposing the features and substructures discovered on exemplars into a coherent role that subsumed all but over-described any of them. The canonical example has been the Medieval city, with its burghers and guilds and churches and fountains and market squares, defensive installations and city government, rights and duties and privileges and so on and so forth. And though no specific Medieval city ever covered all of these aspects, at least not at the same point in time, the ideal type guides discovery of information and serves as a backdrop to explanations and narrations.



In terms of knowledge representation, the ideal type is of course a very different beast altogether than either the script or the golem representation and some thoughts on how that difference cashes out is going to be the main concern 

Sunday, August 13, 2017

What about all these Stories? Rereading Schank

In the aftermath of the civil disturbances of Charlottesville, Michael Eric Dyson wrote in the August 12th, 2017 edition of the New York Times:
Such an ungainly assembly of white supremacists rides herd on political memory. Their resentment of the removal of public symbols of the Confederate past — the genesis of this weekend’s rally — is fueled by revisionist history. They fancy themselves the victims of the so-called politically correct assault on American democracy, a false narrative that helped propel Mr. Trump to victory. Each feeds on the same demented lies about race and justice that corrupt true democracy and erode real liberty.
Dyson then distinguishes between political memories, between revisionist history and other history (proper? academic? he has no name for it, but cites Du Bois), between true and false narratives.

Tools of Story Telling

It is in this context relevant to look at the infrastructure that Roger Schank has identified for story telling (bracketing Schank's occupation with intelligence, however). Because of his particular slant to story analysis, Schank provides a repertoire of narratological tools and distinctions that differ from those in the literary toolkit. (This cuts both ways, Morson argues in the Foreword, missing for example the notion of Genre in Schank's musings, p.xxff.)

Stories are received and generated structures that people label in an attempt to index them, so that they can retrieve them and match them for comparison purposes. People are assisted in story manipulation by story skeletons, which give a basic direction to a narrative, and by the hierarchy of scripts, plans, goals and themes that Schank had identified in his classical work with Robert Abelson (1977). Because stories can be retold in a variety of ways and for various purposes, Schank posits that they are remembered in a different format, the gist, of which the individual retelling is a production. Stories vary by culture and subculture---Schank specifically discusses the French restaurant on the one hand and the teenager Rock fan on the other.

In much of the exposition of the book, and especially in the discussion of the story production and generation, Schank brackets the question of truth. Schank writes as if the intentionality of the stories would make the question of whether the stories are true or not irrelevant or even impossible to answer.

Diplomatic Incidents

Schank uses a diplomatic conflict, the shooting down of an Iranian airliner by a US Navy warship, the Vincennes, on July 3, 1988, to illustrate the way that the same event is used to make hay for the divergent political positions.
If we construct our own version of truth by reliance upon skeleton stories, two people can know exactly the same facts but construct a story that relays those facts in very different ways. (p.152)
Or even more patronizingly:
The real problem in using skeletons whis way is that the storytellers usually believe what they themselves are saying. Authors construct their own reality by finding the events that fit the skeleton convenient for them to believe. They enter a storytelling situation wanting to tell a certain kind of story and only then worrying about whether || the facts fit onto the bones of the skeleton that they have previously chosen. (p.154f)
Schank then goes on to label the skeletal stories that backed the reactions of the various government as the US using the understandable tragedy (p.152f), Great Britain the justifiability of self-defense (p.153),  Libya insolence and state terrorism (p.154),  and Bahrain justifiable bad effects of war on the aggressor and moral courage (p.155).

So, we can summarize as a first impression that, while Schank will admit that some facts do not fit onto the bones of a skeleton, he is willing to talk about people's own realities and how the skeleton stories are what make up people's notions of truth.
... no matter what happens next, all the viewers of the play [that plays out in the international diplomatic incident, RCK] will retell the story according to the skeletons they have already selected; i.e. they will probably not be moved to reinterpret any new event in terms of some skeleton that they do not already have in mind. (p.158)
So even though the stories dominate to the point of where they shape the memory of the events and the way that future facts will attach to past narrations, there is still something that Schank can call coherence.
One of the oddities of story-based understanding is that people have difficulty making decisions if they know they will have trouble constructing a coherent story to explain their decision. (p.159)
So justificatory stories do have coherence requirements, otherwise they will not properly support the decisions through explanations. Schank goes on to demonstrate this with divorce stories (pp.160ff), which however need not concern us here.

Stories that were Challenged

While there is much plausibility to the fact that political agents have apriori decided how participants will respond to events in the arena that is international diplomacy---theorizing, as Sherlock Holmes would warn us, in advance of the facts (p.159), matters become more complicated when Schank looks at stories that were strongly challenged by agents able to do something about disagreement.

Schank discusses this problem in two examples, in the claimed rape of Tawana Brawley by white supremacists (pp.208ff) and the disqualification of Canadian runner Ben Johnson (pp.210ff) for allegations of doping. However, the larger context is the problem of the cultural and subcultural specificity of stories, and how to acquire that skill (pp.189-218). As a result, the analysis of these stories that were challenged (and ultimately rejected) by authoritative agents come across as situations of choosing bad raw materials for story telling on part of Brawley and Johnson.

Tawana Brawley

... Tawana Brawley ran away for four days, and she needed to tell a story. For her own reasons, she decided not to tell the true story of where she had been, but to invent one instead. (p.209)
It is easy to agree with Schank here, though one has to emphasize that he distinguishes between "true story" and "invend(ed) one" explicitly.

Schank then continues:
When we invent a story in order to mislead people, we try to figure out the story that they want to hear, and we tell it. Children frequently tell a he made me do it story or an it wasn't me that did it story when they are caught having done something wrong. And this kind of story is what Tawana Brawley told too. (p.209)
Again, we are on the same page as Shank here. But then Schank takes a departure that is interesting for the problem of reliability.
Her [Tawana Brawley's, RCK] problem was selecting a believable story. She failed to assess how many listeners would hear her story and failed to understand that what each of them wanted to hear was quite different. (p.209)
Even with the discussion of enculturation into story cultures and subcultures and the difficulties for teenagers to acquire these skills going before (p.208), the turn here is unexpected. The problem according to Schank at this point is not that too many listeners would eventually exhaust Tawana's range of fictional supports, but that they story expectations would eventually exceed her capabilities in some other way.

Schank argues that Brawley picked two basic skeletons, both shockingly realistic in 1980s America ( and as Dyson might remind us, even in 2017), namely young girl is kidnapped and raped as well as young black is victim of racial attack (p.209). Crucially for Schank's argument, the first one comes from the general American culture, and the second one from Brawley's subculture (p.209).
While either of these standard stories is bad enough, the combination of them produced something we can only assume Tawana had not counted on---the match of two types of stories || sought by the news media and black activists. (pp.209f)
Schank then discusses these two groups in turn:
The media looks for horrifying stories involving assaults on especially innocent people. Consequently, Tawana's story matched a skeleton story that news people are always looking to report. (p.210)
Her [i.e. Tawana Brawley's, RCK] story also matched a skeleton story that black activists are always on the lookout for: innocent blacks as easy victims of white officials. She had added that her rapists were state police officers and other officials of her area. The factor, of course, had it been true, should well have caused alarm on the part of the activists. (p.210)
Schank is entirely plausible in identifying the activation patterns for various groups' involvement, and the in arguing that Brawley did not understand how her chosen skeletons meshed with these activation skeletons. But even here, Schank is willing to admit that there are true stories of innocent blacks as victims---"had it been true" (p.210)---that activists should be alarmed by.
Tawana's mistake, apart from fabricating events in the first place, was to invent stories without understanding the standard nature of the stories that others look for. (p.210)
There are a polite and a critical reading to this sentence. The admission, and as an aside only, that "fabricating events in the first place" was a mistake either is Schank's backdoor to admit that the events matter (this is the polite reading) or an indication of his lack of appreciation for how much events matter (the critical reading), i.e. that Brawley's failure to grasp "the standard nature of stories" (p.210) pale to the point of irrelevance in comparison with the need to fabricate facts that could hold up to the scrutiny of multiple organizations.

It is thus with much less agreement than previously, or with a feeling of the emphasis lying on the wrong part, that we head into the description of the role of the standard nature of stories.
Thus, when doctors are handed a rape case or a case of unconsciousness from beating and deprivation, they have certain tests they perform to aid the victim. They also use skeleton stories to understand such cases, and they seek to fit the details of any new story into the familiar story. (p.210)
This  is just an adapted restating of Schank's general epistemology that understanding in humans is story-based.
Similarly, the police know a story about rape, and in the case of Tawana Brawley, they tried to fit the details of her story into theirs. (p.210)
As a result of casting the problem in this way, Schank posits that the police became suspicious because of the mismatch between Tawana's story and their expected story.
For both the doctors and the police, Tawan's story did not agree with the standard stories about rape, and this disparity caused them to question the truth of Tawana's claims. (p.210)
Notice that the whole problem of truth---now however applied to claims, not to facts or events or stories---is admitted and raised by Schank. Yet his conclusion is cast in terms of story skeletons and of subcultures:
Not surprisingly, since she is young and rather unsophisticated, Tawana Brawley failed to understand how to make up a story that matched the ones that the people who were listening to her expected to hear. She did not know the stories of the other subcultures. (p.210)
And almost wistfully, Schank concludes:
Traveling across cultures, it seems, requires the help of a translator. (p.210)
Perhaps a stronger stance that Schank could have taken was to recur on his prior notion of scripts. Tawana failed, because she did not really understand the rape script and the deprivation script, and was therefore unable to tell stories that matched well enough with the script that the doctors and the police officers had for these incidents.

Furthermore, as the New York Times article that Schank quotes (pp.208f) suggests, while the doctors and the police officers took their departure from their questioning the validity of the story, they did not condemn her on that account but used it to find actual evidence that contradicted Tawana's narrative. Notice the use of the word "evidence" in the cited article:
The conclusion that Miss Brawley fabricated her story is supported by ... evidence that she ran away and spent much of the next four days at her former apartment, evidence that she concocted the condition in which she was found, and evidence that she tried to mislead the police, doctors and others about what had happened to her.  --New York Times, September 27, 1988, cited in: Schank 1990, (p.209)
In fact, the article, a two-page spread as Schank admits (p.208), was called "Evidence points to Deceit by Brawley".

Even granted all the usual caveats of the mutual interaction between evidence and narrative, it seems hard to support the stance that for Miss Brawley, the story skeletons did her in. At best they tipped the police off to problems. And the failure in the collection of confirming evidence, which would have been required equally for a court of law, would have done her in just as much as the success in finding fabricated evidence.

Canadian Runner Ben Johnson

Schank of course believes that he has proven that Brawley became entrapped by being unfamiliar with the story expectations. Thus he can use the case of Ben Johnson, a Canadian runner disqualified from the Olympics for doping, as "another case of story misunderstanding" (p.210). 
The story is interesting in this context because the press assumed that everyone who used steroids also knows how to avoid getting caught. (p.211)
Schank's quotes from the NY Times actually do not bear this statement out. It is Dr Voy, the chief medical officer of the US Olympic Committee is the one that is quoted as commenting on the masking know-how in the athletes community.

Schank then goes on to enumerate the stories that Dr Voy, the physician and official, has---to wit: the miscalculate the dose, the screwy system, and the reckless gamble story, the panic to insure victory story and the masking the drug story (p.211)---as well as the stories that Johnson was familiar with---to wit: spiked my drink with drugs and bad lab test (p.211).

The first thing to notice is that Voy's and Johnson's narrative intentions are at odds. None of Dr Voy's stories would have gotten Johnson a pass for testing positive. Indeed, the article gives no examples of stories that Dr Voy would have considered legitimate for a positive testing candidate to continue on as a medal bearer. So Johnson could not have used any of these stories. In fact, all we learn is that the stories Johnson alleged were not part of Dr Voy's repertoire of acceptable excuses; that set may have been empty. If there was no story for Johnson to choose, then Schank's comment, that "in essence, the only thinking we have here is the selection of stories" (p.211), is puzzling in the extreme.

Summary

Though Schank at times sounds as if he is flirting with both sides of the relativism divide, this effect is produced partially by sloppy wording and partially by not distinguishing the following points correctly:
  1. Even if skeletal stories are hard to disprove because of their intentionality, which functions as a political stance does, their are notions of quality of match that can lead to not selecting them.
  2. Stories include factual information and can be challenged much more readily along that dimension. Admittedly, facts may themselves be the outcome of stories (a topic that Schank unfortunately bypasses), and facts have some of the same subcultural aspects that Schank identified for the stories (e.g. water-logged dinosaurs supporting a Biblical flood narrative) as to acceptability and validity.
  3. Though it may not be possible to show that some stories are true, it is sufficient to be able to show that some stories are false for historiography to have the ability to stem the flood of politically-charged revisionist history. Tawana Brawley's story for sure and possibly Ben Johnson's story as well (Schank's source material is too terse) are stories that are false, due to the evidence found, not the skeletons chosen.
It is possible that Schank's side-stepping of the question of truth is due to his general goal of moving AI work from theorem proving and expert systems (p.xlii) toward story generation / understanding / summarization and explanation.

If we take a step back and look at the larger sources of information available for the examples that Schank uses, such as the Admiral Fogarty report from August 1988 or the investigative research journalism by the New York Times reporters Robert McFadden et al, Outrage: The Story Behind the Tawana Brawley Hoax, from August 1990,  we get the suspicion that Schank is used in the small forms, such as the newspaper article, not the hundred-plus pages slugfest of details that Fogarty or McFadden and their collaborators provided with.

Postscriptum

In an amusing twist, the McFadden's book Outrage, published contemporaneously with Schank's book in 1990, possibly even contradicts the sequence of events claimed by Schank. 
Later, Dr [Alice, RCK] Pena [the emergency room physician, cf. p20, RCK] had her put her feet into stirrups for an examination of the vaginal, rectal and pelvic areas. The examination revealed no cuts, dried blood, bruises, swelling, deep redness, or other indications of injury. There were no signs of trauma to the mouth …. Indeed, the girl’s [Tawana Brawley’s, RCK] teeth were surprisingly clean and her mouth did not even have a bad odor. Dr Pena decided to forego the use of a rape-detection kit, at least for a while. There weren’t enough signs to warrant it …. (p.21)
At this point, with Tawana still giving signs of semi-consciousness (p.17), and had hardly launched into a story of the type that Schank is looking for. Already the initial interaction with the paramedics and the emergency room physician was gathering evidence against the narrative that Tawana was planning or had been told to use.

Bibliographic Record

Roger C. Schank, Tell Me A Story: Narrative and Intelligence,  Evanston, IL (Northwestern University Press), 1995 (= reprint of the 1990 edition published in New York (Scribner), 1990; with a foreword of Gary Saul Morson, from 1995). [Google Book Selections]

Monday, July 31, 2017

Some notes on Corpus Linguistics and their Criticism

My friend Paige Morgan, Digital Humanities Librarian at the University of Miami, recently suggested Laurence Anthony's tool AntConc for corpus linguistic analysis that is well supported and straightforward to accomplish thanks to online tutorials and YouTube videos, some of them even by the maintainer himself.

As so often, from that seed one can spiral out into similar corpus-investigating tools, such as the online Voyant Toolkit, which makes statistical properties of texts visible and accepts cut-and-pasted text for quick analysis. I am not clear yet how to use some of tools, such as Cirrus-display or the term-berry, but my example was rather short and had little in terms of repeating words.

Because Corpus Linguistics has been going on for a while now, there are a couple of good articles or even books on how to construct corpora on Amazon, mostly targeting linguists however, with a few exceptions.

There has also been criticism in the community, distinguishing corpus linguistics from discourse analysis. The recent exercise that I went through for the purposes of this blog, on Ernest Gellner's book Plow, Sword and Book, effectively plays off the discourse analytic stance against the corpus linguistic one, as I now suspect. The most recent book I was able to obtain that splits the difference comes from the analysis of academic writing.

There have been more principled attacks against corpus linguistics, for example from Noam Chomsky, for example discussed here, which see in corpus linguistics the mistaken assumption that data will eventually induce itself into a theory. (That sounds familiar ....)

There are also people who simply try to explain how corpus linguistics came about and developed, e.g. Karen Fort at INIST.fr. Fort reminds her readers of the incident reported in Hill 1962, when Chomsky overplayed his hand and denying that perform could be used with mass-nouns in English, citing himself as a native speaker as an authority. The British national Corpus revealed however that perform magic is indeed just such a construction. (I would argue that Chomsky fell into a recognition/recall trap here, overestimating the latter based on his excellence at the former. Still---bad for him.)

A more detailed review with 20th century references from the University of Lancaster can be found here. It cites British linguists like H. Widdowson, who in 2000 wrote an article entitled The Limitations of Linguistics Applied in Applied Linguistics, 21/1:3-25.
For, obviously enough, the computer can only cope with the material products of what people do when they use language. It can only analyse the textual traces of the processes whereby meaning is achieved: it cannot account for the complex interplay of linguistic and contextual factors whereby discourse is enacted. ... In reference to Hymes' components of communicative competence (Hymes 1972), we can say that corpus analysis deals with the textually attested, but not with the encoded possible, or the contextually appropriate.  (no page number provided in extract)
Though Widdowson is talking about language learning, he makes a point that bears repeating in a larger context:
... the textual product that is subjected to quantitative analysis is itself a static abstraction. The texts which are collected in a corpus have a reflected reality: they are only real because of the presupposed reality of the discourses of which they are a trace.  (no page number provided in extract)
That seems to me to be the entry point for the importance of discourse analysis. In fact, Widdowson later summarizes his point as
... corpus linguistics provides us with the description of text, not discourse. (no page number provided in extract)
The document than provides a rejoined by M Stubbs,  from 2001, entitled Texts, corpora and problems of interpretation: A response to Widdowson, published in Applied Linguistics 22/2. pp.149-172.
Corpus linguistics therefore investigates relations between frequency and typicality, and instance and norm. It aims at a theory of the typical, on the grounds that this has to be the basis of interpreting what is attested but unusual. (no page number provided in extract)
And more fully later on:
Frequency is not necessarily the same as interpretative significance: an occurrence might be significant in a text precisely because it is rare in a corpus. But unexpectedness is recognizable only against the norm. (no page number provided in extract)
Stubbs notes that this insight is especially important to note conventionality:
... [A] major finding of corpus linguistics is that pragmatic meanings, including evaluative connotations, are more frequently conventionally encoded than is often realized (Kay 1995; Moon 1998; Channell 2000). (no page number provided in extract)
Concepts of convention and norm raise problems in the not infrequent cases when interpretations diverge. (no page number provided in extract)
Stubbs cites the case of cronies in corpus-based dictionaries, but emphasizes that the analysis is made possible by the empirical aspect of corpus studies.

In a 2016 paper in Dialogic Pedagogy by Richards and Pilcher (which works with a somewhat static distinction of objective and subjective language systems) quote Ädel (2010), p.48, who worries about "the inevitable focus on surface forms in corpus work" as well as "the risk of focusing exclusively on the word and the phrase level when using computer-assisted methods" p.49.
It it was, in contrast [to the previously stated, RCK] accepted that the usage and the meaning of the language was creative [IS2], individual [IS1], and only represented the inert hardened crust of the language [IS4] then the linguist would be unable to analyse it isolated from the context in which it was used. (A128)
Except, what choice do the historians have? Voloshinov (or Mikhail Bakhtin, however the debate around Morris 1984 comes out, cf. A123) criticized the departure of linguistics from the antiquarian concerns, where
the ancient written monument [is considered] ... the ultimate realium" (Voloshinov, 1973, p.73, cited in Richards & Pilcher, A123).
Pilcher and Richards cite Bakhtin in observing:
Fundamental to the meaning of language in such a view ... are dialogue and context. The importance of dialogue (Bakhtin 1981, 1986) means that language consists of a stream of unfinished utterances that is continually evolving and is never completed. (A129)
Context is with Bakhtin referred to as a linked chain of previous utterances (A129). Pilcher and Richards mostly focus on spoken language here (e.g. the contribution of intonation to interpreting the word `well`), but some of their concerns are true in larger situations also. Citing Fecho, they write
... to expect that just because you and I are using the same term or phrase that we have a consensus understanding of its meanings is to deny that context and experience have anything to do with our understandings (Fecho, 2011, p.19; cited A130)
Corpus linguistics then generates frequency lists of
... decontextualized signifiers, which in turn are only evidence of past thoughts. (A130)
There was finally also a paper by Nelya Koteyko trying to distinguish the different forms of discourse used in science and their applicability to corpus linguistics, but I did not quite catch the main drift.

Sunday, July 30, 2017

The problem of the range of the Discourse

During a discussion of couples' interactions in the New Yorker, the author reminded the reader that the range of a discourse in presidential politics and the presidential White House extends to the previous occupants and their actions as well.
On Tuesday, after Melania [Trump] appeared again to reject the President [Donald Trump], this time on the tarmac in Rome with a slick “down low, too slow” move, Pete Souza, President Obama’s official photographer, posted a photo to his Instagram account of Barack and Michelle tenderly holding hands in Selma, Alabama, a gesture that needed no interpretation.
This is an example of the kind of interaction that is difficult to track or detect without establishing the precise discourse that the item belongs to. Here models of layers of discourse that need to be attended to are crucial.

Eventually, the Washington Post made it clear at a description level, by linking to these (and other) clips and photos, providing the interpretation for those that had missed the discourse contributions. So the hope of large scale ingesting of documents for interpretation is that discourse contributions that are clever in the way that Souza's was will eventually have the kind schoolmaster who spells out what the others suspected. (In some sense, the historians often end up in that role.)

Of course, not every hand-holding couple posted that day is a commentary on the Trumps' situation, but most likely, the George W. Bushes' holding hands would have been, within a specific window of time, of course.

Appendix


Monday, June 6, 2016

The power role of historiography

As Justin E.H. Smith put it in his ruminations on what it means to compare Hitler and Trump:
History may be rooted in storytelling, but we can summon it to be something more — the arbiter of truth against lies told in pursuit of power.

Saturday, June 4, 2016

Bronze Age Britain and the "Brexit"

In an article from the New Yorker, describing the latest discoveries at Must Farm quarry near Peterborough in England, Charlotte Higgins reflects on the openness of the evidence toward the larger cultural narratives.
At a moment when Britain is anxiously debating its relationship to the Continent, the kind of story one chooses to tell about the nation’s past matters a great deal. According to recent polling, the area around Peterborough is the second most skeptical of E.U. membership in Britain. One archeologist at Must Farm told me that he and his colleagues sometimes joke about the contradictory narratives that could plausibly describe the former life there. “You could probably write either a Euroskeptic or a Europhile account of this site,” he said. One researcher might see an indigenous community that was closely tied to the landscape of East Anglia; another might see an outward-looking people linked to the Continent by its waterways. Archeology is always an encounter between a fixed past and a shifting present; we bring to it our fantasies, prejudices, and predilections—this year different from last year, next year different again.

Saturday, April 30, 2016

anecdotally approximating the needs of digital humanity work

(The following is akin to the prologue of a future paper in the digital humanities, a motivator for my research and work.)

Biographically speaking, I am the embodiment of a digital humanities researchers, as I aspire to spend about half of my time in the digital world as a software engineer, and half of my time in the humanities, mostly working on problems of socio-economic history and its impact on religious thought. Moving in both of these worlds makes the disconnect between their tools and capabilities noticeable in a way that I find instructive for the efforts of tooling the digital humanities.

Assume I open my email inbox in the morning and find an email from my fellow software engineer Pace, who has difficulties with a piece of software that I am responsible for. For historical reasons, software engineers have termed these problems "bugs", and the process of bringing this issue to my attention is termed a "bug report". In her bug report, Pace will describe what she did, what the expected outcome was, and what happened instead. Software engineers have a standard method for dealing with these issue: we boil the "bug" down to a small program that verifies that the expected inputs produce the expected outcomes---we call that a "unit test"---and then tweak and modify the existing software until the unit test passes, that is, the program no longer exhibits the erroneous behavior. In doing so, good software projects draw upon their existing unit tests---long-running projects will have tens of thousands of these. Passing all unit tests ensures that eliminating one bug did not introduce any other issues: Primum, non nocere.

Assume that in the afternoon, when I check my other inbox, I find an email from fellow digital humanities researcher Paige. Paige just found a problem with one of the transcribed sources that I shared with her from the archival work that I did for my dissertation. The problem may be very small. Perhaps we can even sort out what the difficulty was and rectify it---the source may now be digitized and accessible on the web. But I have no quick and safe way to incorporate that correction into my dissertation, though it exists as a set of LaTeX files that are themselves easy to change. I have no representation of the argument that my dissertation is making. I have no way of verifying that incorporating Paige's correction is local and will not affect other claims that I make.

Of course, that situation is not very different from the one that I found myself in while writing the dissertation to begin with. It seems statistically implausible that there are no flawed steps in an argument spanning a dozen chapters, a couple hundred of pages and consisting of tens of thousands of words. No manager would hire a developer who wrote a program of that size and had checked its correctness only in their head.

Don't get me wrong: I love lemmatized POS-tagged corpora of major writers as much as the next guy does. But the main product of the humanities is arguments, and crucially counter-arguments, narratives that go against what the public already believes to be the case. Should we not devote the same care to them as software developers toward their code?


Wednesday, April 20, 2016

On estimation of confounders

I came to this paper on controlling for statistical confounders via Scott Alexander's Slate Star Codex website, which covers statistical topics.

Westfall and Yarkoni remind us that we use proxies to estimate the true latent constructs, e.g. a "specific survey item asking about respondents’ income bracket" to estimate "socioeconomic status". Thus, statistical arguments cannot easily control for the potentially confounding latent construct, but only for a measure of that construct, which has its own error brackets.

At the root of the problem lies the insight countervening conventional wisdom that 
The relationship between n and Type 1 error may be less obvious: all else equal, as sample size increases, error rates also increase.
The problem is exacerbated by the way reliability interacts with error rates, giving a non-monotonic relationship.
In the middle [of the reliability range, RCK], however, there exists a territory where effects are large enough to afford detection, but reliability is too low to prevent misattribution, leading to particularly high Type 1 error rates.
Furthermore, this problem is independent of the statistical analysis approach used (frequentist vs Bayesian, parameter estimation) as long as reliability is not explicitly accounted for.

Indeed, somehow getting a grip on the reliability is the core of the problem. But that is easier said than done. Westfall and Yarkoni point out that
econometric studies attempting to control for SES [i.e. socioeconomic status, cf. above, RCK] hardly ever estimate or report the reliability of the actual survey item(s) used to operationalize the SES construct ....
This is why Westfall and Yarkoni recommend structural equation modeling (SEM) in order to avoid the trap altogether. Lacking any estimations for the reliability, the researchers can then at least plot their results across a range of estimates to see how sensitive their findings are on reliability.

Bibliographic Record:
Westfall J, Yarkoni T (2016) Statistically Controlling for Confounding Constructs Is Harder than You Think. PLoS ONE 11(3): e0152719. doi:10.1371/journal.pone.0152719

Sunday, April 17, 2016

Nelson Goodman on Ampliative Inference

By ampliative inference is central to many historiographical arguments, such as induction by enumeration, analogical reasoning, inference to the best explanation, and other forms that expand on what is available in the premise.

However, it is in these contexts precisely that projectibility can be lost. In my current understanding this has to do with discontinuities in the underlying space, as the bleen and grue examples show that Nelson Goodman introduced in the 1940s and revisited in his 1984 paper New Riddle of Induction in Fact, Fiction and Forecast. Thus, while
  1. All observed emeralds are green.
  2. Thus, all emeralds are green.
is valid, the form
  1. All observed emeralds are grue.
  2. Thus, all emeralds are grue.
fails, the fact that the definition of "grue" contains the notion of "observed before the year X" in its definition means that the property of "grue" is insufficiently general for projectability. 

Chris Swoyer, who wrote the SEP article used here, argues that the grue problem is "just an instance of the ubiquitous underdetermination of hypotheses by finite bodies of data". And such underdetermination is a problem for inductive inference, unless the induction is controlled in some fashion, i.e. some hypothesis are rejected from the inductive step. Restriction to projectable relations is one way to tackle the problem.

In his 3rd Lecture, Goodman looks at the way that justification and inference are related, and states a co-dependency that is virtuously circular:
A rule is amended if it yields an inference we are unwilling to accept; an inference is rejected if it violates a rule we are unwilling to amend. (p.67)
Thus:
The process of justification is the delicate one of making mutual adjustments between rules and accepted inferences; and in the agreement achieved lies the only justification needed for either. (p.67)
Similarly for inductive inference.
An inductive inference, too, is justified by conformity to general rules, and a general rule by conformity to accepted inductive inferences. Predictions are justified if they conform to valid canons of induction; and the canons are valid fi they accurately codify accepted inductive practice. (p.67)
Applying this to inductive work, which is neglected as compared to deductive inference, is what Goodman calls the Constructive Task of Confirmation Theory (p.68). However, just as with defining any term, we tweak the usage and the definition to make sure that the definition picks out only the instances of usage an no others (p.69).

Goodman's discussion of the black and the raven things starting on pp.70ff is dependent on Carl Gustav Hempel's Studies in the Logic of Confirmation, from: Mind, v.54 n.213 (January 1945), pp.1-26, where Hempel (p.9ff) pulls apart Nicod's criterion for confirmation or invalidation, as Nicod termed it. The basic point is that the basic structure of a rule in its logical form as disjunct no longer separates the assumptions of the hypothesis from the predictions of the hypothesis. This has the effect that logically equivalent hypotheses (R => B and !B => !R) are supported or neutral with respect to the same observations (e.g. the observation !R & !B supports the second but not the first hypothesis). Clearly independence of formulation is an important desideratum (Hempel 1945, p.12).

[Note: I could not find an easy simplification rule that allows to rewrite S1 in the multi-variable case (Hempel 1945, p.13 Fn) into S2, i.e. ~R(x,y) => R(x,y) & ~R(y,x) into R(x,y), but the truth tables show that this is so.]
Taking a leaf from Carl Gustav Hempel, Goodman points out that the only substitutions in the hypothesis step can be the variables for the objects, not the relations (p.71), or put differently:
... the hypothesis says of all things what the evidence statement says of one thing .... (p.71)
The central idea for an improved definition is that, within certain limitations, what is asserted to be true for the narrow universe of the evidence statements is confirmed for the whole universe of discourse. (p.73) 
 (to be continued)

the possibility of digitally supporting argumentation in the humanities

As noted in my previous post,  while the decision process regarding the quality of arguments in the digital humanities is complicated and usually involves mainly sins of omission, the basic model of an argument being open to confirmation makes for a sensible launching point to bring DH support to bear.

It is important to understand that at the level of generality under discussion here, there is no principal bias toward formal models of argumentation. By this I mean that there is no bias toward argumentation whose rules of inference have already been formalized (such as first-order predicate calculus or decision logic or similar). However, it seems to me that the interest in the ability of others to follow our arguments implies that the rules of inference I use in my argument could at least in principle be formalized. If such a "mechanical" or "symbolic" notion of inference is admitted, then Turing has good news for us: such a process of validation of an argument is a Turing-computable function and we can bring effective digital support to bear on the problem.

On the one hand it seems clear that, within the schools of the various disciplines of the humanities, rules for what are valid and proper forms of argumentation (and what are not) have developed and are taught successfully to the new generation. (This is the direction that Thomas Kuhn's arguments about the paradigms of scientific practice are hinting toward.) On the other hand, most humanities scholars would be hard-pressed to successfully formalize their underlying rules of inference. They may even have difficulties to press the forms of argumentation of their school of thinking correctly into another formalized system of inference---such as first-order predicate logic or decision logic.

The situation appears perhaps even more difficult when faced with heuristics, by which I mean argument constructs that are rule-like but are mutually exclusive and are known to be only valid in some situations. Perhaps best-known in the humanities are the rules for cross-document textual interpretation when reconstructing an Urtext, such as lectio difficilior or lectio brevior.  These heuristics come into conflict if the shorter text eliminated difficulties.

But even if we are admonished not to apply these heuristics mechanically, the fact that their weighing against each other could be taught suggests that they can be explained and in the limit applied mechanically by those wishing to understand the process of our weighing, that is our argument.

Sunday, April 10, 2016

digital support for argumentation in the humanities

Perhaps it is especially obvious to someone who has just published his first book that monographs are large complex theoretical undertakings that could benefit from more digital support than word processing, map drawing, or Google and Amazon Kindle books (as searchable sources) provide---however helpful all of these things are!

At the same time, these analogous tools give an appreciation for the limitations of the support one should expect realistically: for example, it took about two weeks of low-level LaTeX tweaking to get my monograph just right, with judicious page stretching and negative-space fiddling and, in the end, we were stuck splicing a page-break into a generated file (the output from the indexer) in order to get the look "just right" (TM by Goldilocks)---a computer-science no-no if there ever was one!

The basic idea behind argumentation in the humanities is simple enough: I would want other people, given the data that I used, and applying to these the analysis that I applied, to come to the same conclusion that I arrived at. In the limit, that process should be effectively mechanical.

The practice of actual scholarly argumentation is of course vastly more complicated, where people often challenge arguments, by

  • challenging the data selection, meaning that there is other data that should have been used
  • challenge the data use, meaning that other features of the data set should have been used 
  • challenging the models (scripts, in the language of Schank and Abelson) that are used to interpret people's actions 
  • etc
The implied criticism is always that these differences in strategies would have led to differences in the argument outcomes (everything else would be supporting evidence). This means that an argument that can be followed and agreed to under its premises is only a small part of the scientific process in the humanities. 

While this is true,  it also stands to reason that the infrastructure developed for digitally supporting the construction and maintenance of humanities arguments and narratives are a good launching pad for digitally supporting the comparison of different arguments and narratives in the humanities. 
So, while effort expended in the digital support of humanities argumentation will not solve "the problem", it is an interesting incremental improvement of supporting ever more of the scientific work of the humanities digitally.

Thursday, May 28, 2015

O'Boyle et al on the Chrysalis Effect

A fascinating study of questionable research practices that requires access to SAGE publications to read, however.

Richard Klein's Many Labs replication effort

Richard Klein and many many collaborators performed a multi-lab replication study of some thirteen key social psychology claims.

Chris Chambers on the First Anniversary of Registered Reports'

One year after the introduction of Registered Reports at Cortex magazine (Elsevier) in May of 2013, Chris Chambers summarizes the lessons learned.

Chambers gives not only a useful process chart for the new strategy, but also pokes fun at the hypothetico-deductive method by inserting the cheats that the false incentive structure produces.

Chambers successfully reused that chart in his APS 2014 presentation.