Discourse Models for Collaboratively Edited Corpora (2008)
BibTeX
@MISC{Chen08discoursemodels,
author = {Erdong Chen and Regina Barzilay},
title = {Discourse Models for Collaboratively Edited Corpora},
year = {2008}
}
OpenURL
Abstract
This thesis focuses on computational discourse models for collaboratively edited corpora. Due to the exponential growth rate and significant stylistic and content variations of collaboratively edited corpora, models based on professionally edited texts are incapable of processing the new data effectively. For these methods to succeed, one challenge is to preserve the local coherence as well as global consistence. We explore two corpus-based methods for processing collaboratively edited corpora, which effectively model and optimize the consistence of user generated text. The first method addresses the task of inserting new information into existing texts. In particular, we wish to determine the best location in a text for a given piece of new information. We present an online ranking model which exploits this hierarchical structure – representationally in its features and algorithmically in its learning procedure. When tested on a corpus of Wikipedia articles, our hierarchically informed model predicts the correct insertion paragraph more accurately than baseline methods. The second method concerns inducing a common structure across multiple articles in similar domains to







