<p>We propose a quantum circuit for word embeddings, serving as a quantum counterpart to the classical Word2vec model. While there are some quantum-inspired models in the literature, no quantum circuits have been proposed for this purpose. The proposed quantum circuit, which we call Q-word2vec, consists of two components: an encoding part that maps data into a Hilbert space and a decoding part. These components are implemented using parameterized unitaries and entangling unitaries, which introduce quantum entanglement into the circuit. Given word-context pairs as training data, Q-word2vec is trained using a hybrid quantum-classical framework. The embedding data are then obtained by measuring the output states of the encoding part. A loss function that incorporates the correlation coefficients between the distance vectors of the embedding data and those of the labels in the training data is used to enhance the performance of Q-word2vec. Furthermore, a heuristic formula is proposed for estimating the circuit length. Synthetic data and two small corpora are used to compare the performance of Q-word2vec with that of Word2vec. Although the performance of Q-word2vec is slightly inferior to that of Word2vec, we believe that leveraging the Hilbert space makes Q-word2vec a promising model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Quantum circuits for word embeddings

  • Hiroshi Ohno

摘要

We propose a quantum circuit for word embeddings, serving as a quantum counterpart to the classical Word2vec model. While there are some quantum-inspired models in the literature, no quantum circuits have been proposed for this purpose. The proposed quantum circuit, which we call Q-word2vec, consists of two components: an encoding part that maps data into a Hilbert space and a decoding part. These components are implemented using parameterized unitaries and entangling unitaries, which introduce quantum entanglement into the circuit. Given word-context pairs as training data, Q-word2vec is trained using a hybrid quantum-classical framework. The embedding data are then obtained by measuring the output states of the encoding part. A loss function that incorporates the correlation coefficients between the distance vectors of the embedding data and those of the labels in the training data is used to enhance the performance of Q-word2vec. Furthermore, a heuristic formula is proposed for estimating the circuit length. Synthetic data and two small corpora are used to compare the performance of Q-word2vec with that of Word2vec. Although the performance of Q-word2vec is slightly inferior to that of Word2vec, we believe that leveraging the Hilbert space makes Q-word2vec a promising model.