On multilabel classification and ranking with bandit feedback

Articolo

Data di Pubblicazione:

2014

Abstract:

We present a novel multilabel/ranking algorithm working in partial information settings. The algorithm is based on 2nd-order descent methods, and relies on upper-confidence bounds to trade-off exploration and exploitation. We analyze this algorithm in a partial adversarial setting, where covariates can be adversarial, but multilabel probabilities are ruled by (generalized) linear models. We show O(T1/2 log T) regret bounds, which improve in several ways on the existing results. We test the effectiveness of our upper-confidence scheme by contrasting against full-information baselines on diverse real-world multilabel data sets, often obtaining comparable performance.

Tipologia CRIS:

Articolo su Rivista

Keywords:

Contextual bandits; Generalized linear; Online learning; Ranking; Regret bounds; Structured prediction

Elenco autori:

Gentile, Claudio; Orabona, F.

Link alla scheda completa:

https://irinsubria.uninsubria.it/handle/11383/1959521

Pubblicato in:

JOURNAL OF MACHINE LEARNING RESEARCH

Journal