Generalized ensemble model for document ranking in information retrieval

Yanshan Wang; In Chan Choi; Hongfang Liu

doi:10.2298/CSIS160229042W

Generalized ensemble model for document ranking in information retrieval

Yanshan Wang, In Chan Choi, Hongfang Liu

Digital Health Sciences

Research output: Contribution to journal › Article › peer-review

1 Scopus citations

Abstract

A generalized ensemble model (gEnM) for document ranking is proposed in this paper. The gEnM linearly combines the document retrieval models and tries to retrieve relevant documents at high positions. In order to obtain the optimal linear combination of multiple document retrieval models or rankers, an optimization program is formulated by directly maximizing the mean average precision. Both supervised and unsupervised learning algorithms are presented to solve this program. For the supervised scheme, two approaches are considered based on the data setting, namely batch and online setting. In the batch setting, we propose a revised Newton’s algorithm, gEnM.BAT, by approximating the derivative and Hessian matrix. In the online setting, we advocate a stochastic gradient descent (SGD) based algorithm—gEnM.ON. As for the unsupervised scheme, an unsupervised ensemble model (UnsEnM) by iteratively co-learning from each constituent ranker is presented. Experimental study on benchmark data sets verifies the effectiveness of the proposed algorithms. Therefore, with appropriate algorithms, the gEnM is a viable option in diverse practical information retrieval applications.

Original language	English (US)
Pages (from-to)	123-151
Number of pages	29
Journal	Computer Science and Information Systems
Volume	14
Issue number	1
DOIs	https://doi.org/10.2298/CSIS160229042W
State	Published - Jan 2017

Keywords

Document ranking
Ensemble model
Information retrieval
Mean average precision
Optimization

ASJC Scopus subject areas

General Computer Science

Access to Document

10.2298/CSIS160229042W

Cite this

@article{c4f0a19d21b54240bad923e5b66af1a3,

title = "Generalized ensemble model for document ranking in information retrieval",

abstract = "A generalized ensemble model (gEnM) for document ranking is proposed in this paper. The gEnM linearly combines the document retrieval models and tries to retrieve relevant documents at high positions. In order to obtain the optimal linear combination of multiple document retrieval models or rankers, an optimization program is formulated by directly maximizing the mean average precision. Both supervised and unsupervised learning algorithms are presented to solve this program. For the supervised scheme, two approaches are considered based on the data setting, namely batch and online setting. In the batch setting, we propose a revised Newton{\textquoteright}s algorithm, gEnM.BAT, by approximating the derivative and Hessian matrix. In the online setting, we advocate a stochastic gradient descent (SGD) based algorithm—gEnM.ON. As for the unsupervised scheme, an unsupervised ensemble model (UnsEnM) by iteratively co-learning from each constituent ranker is presented. Experimental study on benchmark data sets verifies the effectiveness of the proposed algorithms. Therefore, with appropriate algorithms, the gEnM is a viable option in diverse practical information retrieval applications.",

keywords = "Document ranking, Ensemble model, Information retrieval, Mean average precision, Optimization",

author = "Yanshan Wang and Choi, {In Chan} and Hongfang Liu",

year = "2017",

month = jan,

doi = "10.2298/CSIS160229042W",

language = "English (US)",

volume = "14",

pages = "123--151",

journal = "Computer Science and Information Systems",

issn = "1820-0214",

publisher = "ComSIS Consortium",

number = "1",

}

TY - JOUR

T1 - Generalized ensemble model for document ranking in information retrieval

AU - Wang, Yanshan

AU - Choi, In Chan

AU - Liu, Hongfang

PY - 2017/1

Y1 - 2017/1

N2 - A generalized ensemble model (gEnM) for document ranking is proposed in this paper. The gEnM linearly combines the document retrieval models and tries to retrieve relevant documents at high positions. In order to obtain the optimal linear combination of multiple document retrieval models or rankers, an optimization program is formulated by directly maximizing the mean average precision. Both supervised and unsupervised learning algorithms are presented to solve this program. For the supervised scheme, two approaches are considered based on the data setting, namely batch and online setting. In the batch setting, we propose a revised Newton’s algorithm, gEnM.BAT, by approximating the derivative and Hessian matrix. In the online setting, we advocate a stochastic gradient descent (SGD) based algorithm—gEnM.ON. As for the unsupervised scheme, an unsupervised ensemble model (UnsEnM) by iteratively co-learning from each constituent ranker is presented. Experimental study on benchmark data sets verifies the effectiveness of the proposed algorithms. Therefore, with appropriate algorithms, the gEnM is a viable option in diverse practical information retrieval applications.

AB - A generalized ensemble model (gEnM) for document ranking is proposed in this paper. The gEnM linearly combines the document retrieval models and tries to retrieve relevant documents at high positions. In order to obtain the optimal linear combination of multiple document retrieval models or rankers, an optimization program is formulated by directly maximizing the mean average precision. Both supervised and unsupervised learning algorithms are presented to solve this program. For the supervised scheme, two approaches are considered based on the data setting, namely batch and online setting. In the batch setting, we propose a revised Newton’s algorithm, gEnM.BAT, by approximating the derivative and Hessian matrix. In the online setting, we advocate a stochastic gradient descent (SGD) based algorithm—gEnM.ON. As for the unsupervised scheme, an unsupervised ensemble model (UnsEnM) by iteratively co-learning from each constituent ranker is presented. Experimental study on benchmark data sets verifies the effectiveness of the proposed algorithms. Therefore, with appropriate algorithms, the gEnM is a viable option in diverse practical information retrieval applications.

KW - Document ranking

KW - Ensemble model

KW - Information retrieval

KW - Mean average precision

KW - Optimization

UR - http://www.scopus.com/inward/record.url?scp=85011634648&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85011634648&partnerID=8YFLogxK

U2 - 10.2298/CSIS160229042W

DO - 10.2298/CSIS160229042W

M3 - Article

AN - SCOPUS:85011634648

SN - 1820-0214

VL - 14

SP - 123

EP - 151

JO - Computer Science and Information Systems

JF - Computer Science and Information Systems

IS - 1

ER -

Generalized ensemble model for document ranking in information retrieval

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this