Friday, April 29, 2016

digital humanities on Twitter

Courtesy of Doron Goldfarb a list of good Twitter handles for digital humanities, including a list by Martin Grandjean, who seems to be a key player in the Swiss DH scene.

https://twitter.com/ADHOrg
https://twitter.com/eadh_org
https://twitter.com/dhnow
https://twitter.com/DH_UniWien

http://www.martingrandjean.ch/digital-humanities-on-twitter/
https://twitter.com/GrandjeanMartin/lists/digital-humanities

applying TEI modeling to financial information

The MEDEA project (an acronym for Modelling semantically Enriched Digital Edition of Accounts) was pointed out to me by Paige Morgan (University of Miami, Florida) as an effort that she is involved in. It would be fun to take the accounts from the steam-boat "Maid of Iowa" and transcribe them into their TEI format, once that encoding format has been fully defined.

Paige also pointed me to the Guelph workshop on Digital Humanities where she is teaching for four days from May 9-12, 2016. Here's wishing her all the best.

digital humanity tools

I looked at WordHoard today, a tool spearheaded by Martin Mueller of NWU, one of the grandmasters in the TEI text-encoding initiative (he was also behind the MONK project).

I was struck by some of the cool aspects of WordHoard, such as the stemming and POS infrastructure and the Bean scripting window, as well as some of the weirdness (MySQL-backed hibernate objects that require full rebuild for every XML edit?).

Some of the linguistic services available from MONK are incorporated in the MorphAdorner that is still available on BitBucket (with cool online examples hosted at NWU). There is also a collaboration with the Lucene-based BlackLab index.

The TEI community seems to be big fans of the Oxygen XML editor, but I can handle a lot of pain for 240US$, to be quite honest about it.

I also noticed that the next TEI conference is in Vienna, in September, and that they have a call for papers out now. Maybe I can think of something cool to write on.

Wednesday, April 20, 2016

On estimation of confounders

I came to this paper on controlling for statistical confounders via Scott Alexander's Slate Star Codex website, which covers statistical topics.

Westfall and Yarkoni remind us that we use proxies to estimate the true latent constructs, e.g. a "specific survey item asking about respondents’ income bracket" to estimate "socioeconomic status". Thus, statistical arguments cannot easily control for the potentially confounding latent construct, but only for a measure of that construct, which has its own error brackets.

At the root of the problem lies the insight countervening conventional wisdom that 
The relationship between n and Type 1 error may be less obvious: all else equal, as sample size increases, error rates also increase.
The problem is exacerbated by the way reliability interacts with error rates, giving a non-monotonic relationship.
In the middle [of the reliability range, RCK], however, there exists a territory where effects are large enough to afford detection, but reliability is too low to prevent misattribution, leading to particularly high Type 1 error rates.
Furthermore, this problem is independent of the statistical analysis approach used (frequentist vs Bayesian, parameter estimation) as long as reliability is not explicitly accounted for.

Indeed, somehow getting a grip on the reliability is the core of the problem. But that is easier said than done. Westfall and Yarkoni point out that
econometric studies attempting to control for SES [i.e. socioeconomic status, cf. above, RCK] hardly ever estimate or report the reliability of the actual survey item(s) used to operationalize the SES construct ....
This is why Westfall and Yarkoni recommend structural equation modeling (SEM) in order to avoid the trap altogether. Lacking any estimations for the reliability, the researchers can then at least plot their results across a range of estimates to see how sensitive their findings are on reliability.

Bibliographic Record:
Westfall J, Yarkoni T (2016) Statistically Controlling for Confounding Constructs Is Harder than You Think. PLoS ONE 11(3): e0152719. doi:10.1371/journal.pone.0152719

Sunday, April 17, 2016

Nelson Goodman on Ampliative Inference

By ampliative inference is central to many historiographical arguments, such as induction by enumeration, analogical reasoning, inference to the best explanation, and other forms that expand on what is available in the premise.

However, it is in these contexts precisely that projectibility can be lost. In my current understanding this has to do with discontinuities in the underlying space, as the bleen and grue examples show that Nelson Goodman introduced in the 1940s and revisited in his 1984 paper New Riddle of Induction in Fact, Fiction and Forecast. Thus, while
  1. All observed emeralds are green.
  2. Thus, all emeralds are green.
is valid, the form
  1. All observed emeralds are grue.
  2. Thus, all emeralds are grue.
fails, the fact that the definition of "grue" contains the notion of "observed before the year X" in its definition means that the property of "grue" is insufficiently general for projectability. 

Chris Swoyer, who wrote the SEP article used here, argues that the grue problem is "just an instance of the ubiquitous underdetermination of hypotheses by finite bodies of data". And such underdetermination is a problem for inductive inference, unless the induction is controlled in some fashion, i.e. some hypothesis are rejected from the inductive step. Restriction to projectable relations is one way to tackle the problem.

In his 3rd Lecture, Goodman looks at the way that justification and inference are related, and states a co-dependency that is virtuously circular:
A rule is amended if it yields an inference we are unwilling to accept; an inference is rejected if it violates a rule we are unwilling to amend. (p.67)
Thus:
The process of justification is the delicate one of making mutual adjustments between rules and accepted inferences; and in the agreement achieved lies the only justification needed for either. (p.67)
Similarly for inductive inference.
An inductive inference, too, is justified by conformity to general rules, and a general rule by conformity to accepted inductive inferences. Predictions are justified if they conform to valid canons of induction; and the canons are valid fi they accurately codify accepted inductive practice. (p.67)
Applying this to inductive work, which is neglected as compared to deductive inference, is what Goodman calls the Constructive Task of Confirmation Theory (p.68). However, just as with defining any term, we tweak the usage and the definition to make sure that the definition picks out only the instances of usage an no others (p.69).

Goodman's discussion of the black and the raven things starting on pp.70ff is dependent on Carl Gustav Hempel's Studies in the Logic of Confirmation, from: Mind, v.54 n.213 (January 1945), pp.1-26, where Hempel (p.9ff) pulls apart Nicod's criterion for confirmation or invalidation, as Nicod termed it. The basic point is that the basic structure of a rule in its logical form as disjunct no longer separates the assumptions of the hypothesis from the predictions of the hypothesis. This has the effect that logically equivalent hypotheses (R => B and !B => !R) are supported or neutral with respect to the same observations (e.g. the observation !R & !B supports the second but not the first hypothesis). Clearly independence of formulation is an important desideratum (Hempel 1945, p.12).

[Note: I could not find an easy simplification rule that allows to rewrite S1 in the multi-variable case (Hempel 1945, p.13 Fn) into S2, i.e. ~R(x,y) => R(x,y) & ~R(y,x) into R(x,y), but the truth tables show that this is so.]
Taking a leaf from Carl Gustav Hempel, Goodman points out that the only substitutions in the hypothesis step can be the variables for the objects, not the relations (p.71), or put differently:
... the hypothesis says of all things what the evidence statement says of one thing .... (p.71)
The central idea for an improved definition is that, within certain limitations, what is asserted to be true for the narrow universe of the evidence statements is confirmed for the whole universe of discourse. (p.73) 
 (to be continued)