{"id":9580,"date":"2023-07-10T14:40:53","date_gmt":"2023-07-10T12:40:53","guid":{"rendered":"https:\/\/www.dm.unipi.it\/eventi\/on-the-influence-of-stochastic-rounding-bias-in-implementing-gradient-descent-with-applications-in-low-precision-training-lu-xia-eindhoven-university-of-technology\/"},"modified":"2023-07-10T14:40:53","modified_gmt":"2023-07-10T12:40:53","slug":"on-the-influence-of-stochastic-rounding-bias-in-implementing-gradient-descent-with-applications-in-low-precision-training-lu-xia-eindhoven-university-of-technology","status":"publish","type":"unipievents","link":"https:\/\/www.dm.unipi.it\/en\/eventi\/on-the-influence-of-stochastic-rounding-bias-in-implementing-gradient-descent-with-applications-in-low-precision-training-lu-xia-eindhoven-university-of-technology\/","title":{"rendered":"On the influence of stochastic rounding bias in implementing gradient descent with applications in low-precision training &#8211; Lu Xia (Eindhoven University of Technology)"},"content":{"rendered":"<h4>Venue<\/h4>\n<p>Dipartimento di Matematica, Aula Magna.<\/p>\n<h4 class='mt-4'>Abstract<\/h4>\n<p>In the context of low-precision computation for the training of neural networks with the<br \/>gradient descent method (GD), the occurrence of deterministic rounding errors often leads<br \/>to stagnation or adversely affects the convergence of the optimizers. The employ-<br \/>ment of unbiased stochastic rounding (SR) may partially capture gradient updates that<br \/>are lower than the minimum rounding precision, with a certain probability. We<br \/>provide a theoretical elucidation for the stagnation observed in GD when training neural<br \/>networks with low-precision computation. We analyze the impact of floating-point round-<br \/>off errors on the convergence behavior of GD with a particular focus on convex problems.<br \/>Two biased stochastic rounding methods, signed-SR$_\\varepsilon$ and SR$_\\varepsilon$, are proposed, which have<br \/>been demonstrated to eliminate the stagnation of GD and to result in significantly faster<br \/>convergence than SR in low-precision floating-point computation.<br \/>We validate our theoretical analysis by training a binary logistic regression model on<br \/>the Cifar10 database and a 4-layer fully-connected neural network model on the MNIST<br \/>database, utilizing a 16-bit floating-point representation and various rounding techniques.<br \/>The experiments demonstrate that signed-SR$_\\varepsilon$ and SR$_\\varepsilon$ may achieve higher classification<br \/>accuracy than rounding to the nearest (RN) and SR, with the same number of training<br \/>epochs. It is shown that a faster convergence may be obtained by the new rounding<br \/>methods with 16-bit floating-point representation than by RN with 32-bit floating-point<br \/>representation.<\/p>\n<p class='mt-4'>Further information is available on the <a href=\"https:\/\/events.dm.unipi.it\/event\/201\/\">event page<\/a> on the Indico platform.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>In the context of low-precision computation for the training of neural networks with thegradient descent method (GD), the occurrence of deterministic rounding errors often leadsto stagnation or adversely affects the convergence of the optimizers.&hellip;<\/p>\n<p><a class=\"btn btn-dark btn-sm unipi-read-more-link\" href=\"https:\/\/www.dm.unipi.it\/en\/eventi\/on-the-influence-of-stochastic-rounding-bias-in-implementing-gradient-descent-with-applications-in-low-precision-training-lu-xia-eindhoven-university-of-technology\/\">Read More&#8230;<\/a><\/p>\n","protected":false},"author":6,"featured_media":0,"template":"","tags":[],"unipievents_taxonomy":[],"class_list":["post-9580","unipievents","type-unipievents","status-publish","hentry"],"acf":[],"unipievents_startdate":1689688800,"unipievents_enddate":1689692400,"unipievents_place":"Dipartimento di Matematica, Aula Magna.","unipievents_externalid":201,"jetpack_sharing_enabled":true,"publishpress_future_workflow_manual_trigger":{"enabledWorkflows":[]},"_links":{"self":[{"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/unipievents\/9580","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/unipievents"}],"about":[{"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/types\/unipievents"}],"author":[{"embeddable":true,"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/users\/6"}],"version-history":[{"count":0,"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/unipievents\/9580\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/media?parent=9580"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/tags?post=9580"},{"taxonomy":"unipievents_taxonomy","embeddable":true,"href":"https:\/\/www.dm.unipi.it\/en\/wp-json\/wp\/v2\/unipievents_taxonomy?post=9580"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}