The main motivation of this study is based on the current demands of people in an information society, who faced with the information overload. Although a large number of news aggregator websites could help people easily get news articles relevant to the same event, they still contain much redundant information and has no structure of information.
We intend to build a hierarchical structure of tiles that reflects the information discussed in such a set of news articles. This structure could help the reader quickly get an overview of the event and locate the interesting parts. Based on this motivation, we plan some future directions of this study as follows.
Research One of our future works is to pay more attention to remaining issues addressed throughout this thesis. In Chapter 4, we have improved the performance of the text segmentation task with supportive knowledge. Although the experimental results overcome the current state-of-the-art result, our model has no mechanism for deter-mining the number of segments automatically. Therefore, we have to investigate a theoretical and practical analysis to propose a criterion to determine the number of segments.
We also have to do more works on title generation to improve the readability, fluency of the generated title. We will also investigate the way to determine the length of generated title automatically.
In the segment combination task, a challenge is how to build a tree of segments that reflects the hierarchical structure of information. We will investigate on some generative model such as hLDA [9] to deal with this obstacle.
Application We plan to implement a module that can retrieve a set of documents as the input and produce the hierarchical structure of topic-information. That system contains three modules corresponding to three tasks: text segmentation, segment combination, and title generation, respectively. With that system, we can easily implement our improvements and verify them on the real data. Furthermore, it can be easily integrated into the readily news aggregator to provide an option to the users to quickly access needed information. We hope this application is a useful and attractive part of a news aggregator or newspaper website.
Appendix A
Cohesion in English
Halliday and Hassan [41] describetexture as a property possessed by a text, but which an arbitrary combination of sentences does not have. Readers can frequently tell whether or not a series of sentences exhibits texture. In the following example, the sentences in (a) do exhibit it, while those in (b) do not [41].
(a) Wash and core six cooking apples. Put them into a fire proof dish.
(b) Wash and core six cooking apples. The prices of computers drop regularly.
Cohesion is one of the elements of a discourse which contributes to its texture. Cohe-sion is present when an element in a text is best interpreted in light of a previous (or less frequently, following) element of the same text. Halliday and Hasan identify five cohesive relations which contribute texture to a document. The details are summarized by Reynar [87] as follows.
Reference are like pointers. Rather than repeat a phrase in the text, a writer or speaker may use a pointer to the entity selected by a phrase instead. Halliday and Hasan distinguish two main types of reference. Exophoric references are to entities in the world of the discourse and endophoric references are to portions of the text itself.
The word “he” in (a) is an exophoric reference and so is an endophoric reference in (b).
(a) John likes apples, but he loves pears.
(b) For he’s a jolly good fellow. And so say all of us.
Substitution and reference are similar, but differ in that substitution occurs prior to semantic interpretation while reference occurs after interpretation. That is a substi-tute acts merely as a pointer to a region of text which refers to an entity in the world of the discourse, while a reference refers directly to an entity without the mediation of the original referring phrase. In the following example, “does” substitutes for the phrase “like apples”: Do you like apples? Everybody does.
Ellipsis is similar to substitution. It can be viewed as substitution by a zero. In the following example, “bought” has been replaced by a null phrase in “Mary some flowers”: John bought some chocolates and Mary some flowers.
Conjunction is more difficult to define than the previous three relations. It holds be-tween elements of a text when they are ordered temporally, one causes the other, when they describe a contrast or when one elaborates on the other. Examples from [41] will demonstrate these relations. Each of the sentences (a) through (d) should be read immediately following the first sentence in the following example.
For the whole day he climbed up the mountainside, almost without stopping.
(a) Then, as dusk fell, he sat down to rest. (Temporal order) (b) So by night time the valley was far below him. (Causation)
(c) Yet he was hardly aware of being tired. (Contrast) (d) And in all this time he met no one. (Elaboration)
Lexical Cohesion holds between two tokens in a text which are either of the same type or are semantically related in a particular way. There are five semantic relations that constitute lexical cohesion.
1. Reiteration with identity of referenceoccurs when a particular entity previously referred to in a discourse is referred to again.
(a) John saw a dog.
(b) The dog was a retriever.
In above example, (a) refers to a particular dog and (b) refers to the same dog again.
2. Reiteration without identity of reference occurs when reference is made to the entire class to which an entity previously referred to in a discourse belongs.
(a) John saw a small retriever.
(b) Retrievers are usually large.
In above example, (a) refers to one particular member of the set of dogs iden-tified as retrievers while (b) refers to the entire class of retrievers.
3. Reiteration by means of superordinate occurs when reference is made to a su-perclass of the class to which a previously mentioned entity belongs.
(a) John saw the retriever.
(b) Dogs are his favorite animals.
In above example, (a) refers to a retriever, which is a type of dog, while (b) refers to dogs in general.
4. A systematic semantic relation holds when a word, or group of words, has a clearly definable relationship with a previously used word or phrase. For example, both could refer to members of the same set.
(a) John likes retrievers.
(b) He doesn’t like collies.
In above example, (a) refers to retrievers and (b) mentions collies, both of which are subsets of the species of dogs. In this case the relationship can be classified as membership in a particular class.
5. A nonsystematic semantic relation holds between two words or phrases in a discourse when they pertain to a particular theme or topic, but the nature of their relationship is difficult to specify. Recognizing this category in a compu-tational system would be more difficult than recognizing the other categories.
(a) John spent the afternoon studying in his dormitory room.
(b) He loves attending college.
A semantic connection exists between the word “dormitory” in (a) and “col-lege” in (b), but it is hard to classify and unlikely that all such relations, or even the preponderance of them, could be found in a knowledge source in the way that many synonymy relations can be identified using a thesaurus.
Halliday and Hasan’s categories overlap to some degree. For example, it can be difficult to distinguish instances of substitution from endophoric reference. Substitution is subtly different in that it relates words of the text, is not a semantic relation and requires the substituted phrase to have the same role as the phrase it substitutes for. This is not the case with reference. Nonetheless, Halliday and Hasan acknowledge that there are instances where more than one category applies equally well.
Halliday and Hasan explain that texts frequently exhibit varying degrees of cohesion in different sections. Obviously, the start of a text cannot be cohesive with preceding sections, nor can the end exhibit cohesion with later sections. In the middle of a text, however, the quantity of cohesion can vary greatly. Some authors, Halliday and Hasan suggest, prefer to alternate between high and low degrees of cohesion.
Texture—which is more frequently calledcoherence—and cohesion are often confused, but differ significantly. Cohesion relates elements of a text and can generally be identified out of context. Texture, however, is a property that applies to an entire text. It is more difficult to define, but can be recognized upon reading a text in its entirety.
Appendix B
Tools and Datasets
This appendix briefly describes the datasets and tools that have been used for conducting experiments in this study.