{"id":622,"date":"2022-08-29T15:10:45","date_gmt":"2022-08-29T15:10:45","guid":{"rendered":"https:\/\/blogs.kcl.ac.uk\/kclip\/?p=622"},"modified":"2022-08-29T15:16:10","modified_gmt":"2022-08-29T15:16:10","slug":"the-born-supremacy-in-learning-how-to-learn","status":"publish","type":"post","link":"https:\/\/blogs.kcl.ac.uk\/kclip\/2022\/08\/29\/the-born-supremacy-in-learning-how-to-learn\/","title":{"rendered":"The Born Supremacy in Learning How to Learn"},"content":{"rendered":"<p style=\"text-align: left\">Whilst the true impact of quantum computers is anybody&#8217;s guess, there seems to be some consensus on the advantages offered by near-term devices in modeling more complex probability distributions. These distributions can be used to model complex particle interactions, e.g., in quantum chemistry, or, as we will see next, to train principled machine learning models &#8211; in this case, <strong>binary Bayesian neural networks<\/strong> &#8211; and enable fast adaptation to new learning tasks from few training examples.<\/p>\n<div id=\"attachment_631\" style=\"width: 310px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-631\" class=\"wp-image-631 size-medium\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/sys_mod-300x175.png\" alt=\"\" width=\"300\" height=\"175\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/sys_mod-300x175.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/sys_mod-768x447.png 768w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/sys_mod-676x394.png 676w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/sys_mod.png 831w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><p id=\"caption-attachment-631\" class=\"wp-caption-text\">Fig. 1. (left) A binary Bayesian neural network, i.e., a neural network with stochastic binary weights, is trained to carry out a learning task. (right) The probability distribution of the binary weights of the neural network is modelled by a Born machine, i.e., by a parametric quantum circuit (PQC), leveraging the PQC&#8217;s capacity to model complex distributions [1].<\/p><\/div>\n<h2 style=\"text-align: left\">Setting<\/h2>\n<p style=\"text-align: left\">In our latest work, accepted for presentation at the IEEE MLSP, we are interested in training <strong>Bayesian binary neural networks<\/strong>, i.e., classical neural networks with stochastic binary weights, in a sample-efficient manner by means of meta-learning, as illustrated in Fig. 1. The key idea of this work is to model the distribution of the binary weights <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-633\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bthetasolo.png\" alt=\"\" width=\"15\" height=\"15\" \/> via a <strong>Born machine<\/strong>, i.e., via a probabilistic parametric quantum circuit (PQC), due to the capacity of PQCs to efficiently implement complex probability distributions [1]-[4]. We propose a novel method that integrates <strong>meta-learning<\/strong> with the gradient-based optimization of quantum Born machines [3], with the aim of speeding up adaptation to new learning tasks from few examples.<\/p>\n<h2 style=\"text-align: left\"><\/h2>\n<h2 style=\"text-align: left\">Born Machines<\/h2>\n<p style=\"text-align: left\">A Born machine produces random binary strings\u00a0 <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-640\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/btheta-300x62.png\" alt=\"\" width=\"87\" height=\"18\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/btheta-300x62.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/btheta.png 368w\" sizes=\"auto, (max-width: 87px) 100vw, 87px\" \/> , where\u00a0 <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-639\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bn.png\" alt=\"\" width=\"67\" height=\"21\" \/>\u00a0 denotes the total number of model parameters, by measuring the output of a PQC\u00a0 <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-638\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/buphi.png\" alt=\"\" width=\"28\" height=\"20\" \/> defined by parameters\u00a0 <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-637\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bphi.png\" alt=\"\" width=\"18\" height=\"18\" \/>.<\/p>\n<div id=\"attachment_630\" style=\"width: 310px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-630\" class=\"wp-image-630 size-medium\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/circuit-300x128.png\" alt=\"\" width=\"300\" height=\"128\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/circuit-300x128.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/circuit-676x287.png 676w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/circuit.png 708w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><p id=\"caption-attachment-630\" class=\"wp-caption-text\">Fig. 2. Hardware-efficient ansatz for a Born machine. All qubits are initialized in the ground state. The rotations are parametrized by the entries of the variational vector.<\/p><\/div>\n<p style=\"text-align: left\">As illustrated in Fig. 2, the PQC takes the initial state\u00a0 \u00a0<img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-636\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bpsi-300x68.png\" alt=\"\" width=\"106\" height=\"24\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bpsi-300x68.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bpsi.png 380w\" sizes=\"auto, (max-width: 106px) 100vw, 106px\" \/> of <em>n<\/em> qubits as an input, and operates on it via a sequence of unitary gates described by a unitary matrix\u00a0 \u00a0<img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-638\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/buphi.png\" alt=\"\" width=\"30\" height=\"21\" \/>. This operation outputs the final quantum state<\/p>\n<p style=\"text-align: left\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-635 aligncenter\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bpsitau-300x45.png\" alt=\"\" width=\"200\" height=\"30\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bpsitau-300x45.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bpsitau.png 504w\" sizes=\"auto, (max-width: 200px) 100vw, 200px\" \/><\/p>\n<p style=\"text-align: left\">which is measured in the computational basis to produce a random binary string <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-640\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/btheta-300x62.png\" alt=\"\" width=\"97\" height=\"20\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/btheta-300x62.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/btheta.png 368w\" sizes=\"auto, (max-width: 97px) 100vw, 97px\" \/>. Note that each basis vector of the computational basis corresponds to one of all the possible <em>2^n<\/em> patterns of model parameters\u00a0 <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-633\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bthetasolo.png\" alt=\"\" width=\"19\" height=\"20\" \/>.<\/p>\n<p style=\"text-align: left\">The PQC can be implemented using a <strong>hardware-efficient ansatz<\/strong> [2], in which a layer of one-qubit unitary gates, parametrized by vector <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-637\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bphi.png\" alt=\"\" width=\"24\" height=\"23\" \/>, is followed by a layer of fixed, entangling, two-qubit gates. This pattern can be repeated any number of times, building a progressively deeper circuit. Another option is using the <strong>mean-field ansatz<\/strong> that does not use entangling gates, and only relies on one-qubit gates.<\/p>\n<p style=\"text-align: left\">By Born&#8217;s rule (hence the name of the circuit), the probability distribution of the output model parameter vector <img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-633\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bthetasolo.png\" alt=\"\" width=\"19\" height=\"20\" \/> is given by<\/p>\n<p style=\"text-align: left\"><img loading=\"lazy\" decoding=\"async\" class=\"size-medium wp-image-634 aligncenter\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bbr-300x36.png\" alt=\"\" width=\"300\" height=\"36\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bbr-300x36.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bbr-676x81.png 676w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/bbr.png 714w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/p>\n<p style=\"text-align: left\">Importantly, Born machines only provide samples, while the actual distribution above can only be estimated by averaging multiple measurements of the PQC&#8217;s outputs. Therefore, Born machines model <strong>implicit distributions,<\/strong> and only define a stochastic procedure that directly generates samples.<\/p>\n<h2 style=\"text-align: left\">Some Results<\/h2>\n<p style=\"text-align: left\">Fig. 3 illustrate the results in terms of the prediction root mean squared error (RMSE) as a function of the number of meta-training iterations. By comparison with conventional per-task learning, the figure illustrates the capacity of both joint learning and meta-learning to transfer knowledge from the meta-training to the meta-test task, with hardware-efficient (HE) and mean-field (MF) quantum meta-learning clearly outperforming joint learning. For example, HE meta-learning requires around <em>150<\/em> meta-training iterations to achieve the same RMSE ideal per-task training, whilst joint-learning requires more than <em>200<\/em> to achieve comparable performance. The HE ansatz performs best, due to the use of entangling unitaries; however, the MF ansatz approaches the minimal RMSE after <em>230<\/em> iterations. The classical solution based on MF Bernoulli does not achieve lower RMSE than the quantum-aided meta-learning schemes, even with joint learning.<\/p>\n<div id=\"attachment_629\" style=\"width: 310px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-629\" class=\"wp-image-629 size-medium\" src=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/iter_RMSE_v5-300x233.png\" alt=\"\" width=\"300\" height=\"233\" srcset=\"https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/iter_RMSE_v5-300x233.png 300w, https:\/\/blogs.kcl.ac.uk\/kclip\/files\/2022\/08\/iter_RMSE_v5.png 546w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><p id=\"caption-attachment-629\" class=\"wp-caption-text\">Fig. 3. Average RMSE for a new, meta-test, task as a function of the number of meta-training iterations. The results are averaged over <em>5<\/em> independent trials.<\/p><\/div>\n<p style=\"text-align: left\">Please see the paper for a more detailed exposition, available <a href=\"https:\/\/arxiv.org\/pdf\/2203.17089.pdf\">here<\/a>.<\/p>\n<h2 style=\"text-align: left\"><span dir=\"ltr\" role=\"presentation\">References<\/span><\/h2>\n<p style=\"text-align: left\"><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">[1] Arute, F., Arya, K., Babbush, R., Bacon, D., Bardin, J.C., Barends, R., Biswas, <\/span><span dir=\"ltr\" role=\"presentation\">R., Boixo, S., Brandao, F.G., Buell, D.A., et al.: Quantum supremacy using a pro<\/span><span dir=\"ltr\" role=\"presentation\">grammable superconducting processor. Nature<\/span> <span dir=\"ltr\" role=\"presentation\">574<\/span><span dir=\"ltr\" role=\"presentation\">(7779), 505\u2013510 (2019)<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">[2] Kandala, A., Mezzacapo, A., Temme, K., Takita, M., Brink, M., Chow, J.M., Gam<\/span><span dir=\"ltr\" role=\"presentation\">betta, J.M.: Hardware-efficient variational quantum eigensolver for small molecules <\/span><span dir=\"ltr\" role=\"presentation\">and quantum magnets. Nature<\/span> <span dir=\"ltr\" role=\"presentation\">549<\/span><span dir=\"ltr\" role=\"presentation\">(7671), 242\u2013246 (2017)<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">[3] Liu, J.G., Wang, L.: Differentiable learning of quantum circuit Born machines. Phys<\/span><span dir=\"ltr\" role=\"presentation\">ical Review A<\/span> <span dir=\"ltr\" role=\"presentation\">98<\/span><span dir=\"ltr\" role=\"presentation\">(6), 062324 (2018)<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">[4] Sweke, R., Seifert, J.P., Hangleiter, D., Eisert, J.: On the quantum versus classical <\/span><span dir=\"ltr\" role=\"presentation\">learnability of discrete distributions. Quantum<\/span> <span dir=\"ltr\" role=\"presentation\">5<\/span><span dir=\"ltr\" role=\"presentation\">,<\/span> <span dir=\"ltr\" role=\"presentation\">417 (2021)<\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Whilst the true impact of quantum computers is anybody&#8217;s guess, there seems to be some consensus on the advantages offered by near-term devices in modeling more complex probability distributions. These distributions can be used to model complex particle interactions, e.g., in quantum chemistry, or, as we will see next, to train principled machine learning models [&hellip;]<\/p>\n","protected":false},"author":1042,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-622","post","type-post","status-publish","format-standard","hentry","category-uncategorized","post-preview"],"_links":{"self":[{"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/posts\/622","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/users\/1042"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/comments?post=622"}],"version-history":[{"count":10,"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/posts\/622\/revisions"}],"predecessor-version":[{"id":647,"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/posts\/622\/revisions\/647"}],"wp:attachment":[{"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/media?parent=622"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/categories?post=622"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.kcl.ac.uk\/kclip\/wp-json\/wp\/v2\/tags?post=622"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}