Document Hashing by Exploiting Noisy Neighborhood Information with Fault-Tolerant Mutual-Information-Preserving VAE
摘要
Semantic hashing is an effective technique for large-scale information retrieval. Currently, some methods have suggested learning high-quality binary hash codes of documents by leveraging both document contents and neighborhood information. However, it is found that erroneous connections often exist in the provided neighborhood information, but were never taken into account in these models. To alleviate their negative impacts on hash code learning, we first build a basic generative model to simultaneously model the document content and neighborhood. Then, we show that the basic generative model can be placed under a more general framework, dubbed mutual-information (MI) preserving variational auto-encoder (VAE). Capitalizing on this connection, a new hashing method that can tolerate the noisy characteristic of the neighborhood information is further developed by proposing a novel fault-tolerant lower bound for MI. Extensive experiments are conducted on six real-world datasets, and significant performance gains are observed over current state-of-the-art models.