A. Burchardt, S. Pado, D. Spohr, A. Frank, and U. Heid. Constructing Integrated Corpus and Lexicon Models for Multi-Layer Annotations in OWL DL. Linguistic Issues in Language Technology 1, 1-33. Journal page.


We present a general approach to formally modelling corpora with multi-layered annotation in a typed logical representation language, OWL DL. By defining abstractions over the corpus data, we can generalise from a large set of individual corpus annotations, thereby inducing a lexicon model.

The resulting combined corpus and lexicon model can be interpreted as a graph structure that offers flexible querying functionality beyond current XML-based query languages. Its powerful methods for characterising and checking consistency can be used for incremental model refinement. In addition, the formalisation in a graph-based structure offers the means of defining flexible lexicon views over the corpus data. These views can be tailored for linguistic inspection or to define clean interfaces with other linguistic resources.

We illustrate our approach by applying it to the syntactically and semantically annotated SALSA/TIGER corpus, a collection of German newspaper text.



@Article{burchardt08:constructing,
  author = 	 {Aljoscha Burchardt and Sebastian Pad\'o and 
                  Dennis Spohr and Anette Frank and Ulrich Heid},
  title = 	 {Constructing Integrated Corpus and Lexicon Models 
                  for Multi-Layer Annotations in {OWL DL}},
  journal = 	 {Linguistic Issues in Language Technology},
  year = 	 2008,
  volume =	 1,
  pages =        {1--33}
}