Changes

Jump to: navigation, search

Richard Sutton

530 bytes added, 14:52, 12 April 2021
no edit summary
'''Richard Stuart Sutton''',<br/>
an American computer scientist and [[Artificial Intelligence|AI]]-researcher. Since 2003, Richard S. Sutton is a professor in the Department of Computing Science <ref>[http://www.cs.ualberta.ca/ Home | Department of Computing Science]</ref> at the [[University of Alberta]] and is principal investigator of the [[Reinforcement Learning]] and [[Artificial Intelligence]] (RLAI) <ref>[http://rlai.cs.ualberta.ca/RLAI/ualberta.html Reinforcement Learning and Artificial Intelligence (RLAI)]</ref> group. Rich's research interests center on the [[Learning|learning]] problems facing a decision-maker interacting with its environment, which he sees as central to artificial intelligence. He is the author of the original paper on [[Temporal Difference Learning]] <ref>[[Richard Sutton]] ('''1988'''). ''Learning to Predict by the Methods of Temporal Differences''. [https://en.wikipedia.org/wiki/Machine_Learning_(journal) Machine Learning], Vol. 3, No. 1</ref> and, with [[Andrew Barto]], of the textbook ''Reinforcement Learning: An Introduction'' <ref>[http://incompleteideas.net/book/the-book.html Reinforcement Learning: An Introduction] ebook by Richard Sutton and [[Andrew Barto]]</ref> . He is also interested in animal learning psychology, in [https://en.wikipedia.org/wiki/Connectionism connectionist] networks, and generally in systems that continually improve their representations and models of the world <ref>[http://incompleteideas.net/BriefBio.html Brief Biography for Richard Sutton]</ref> .
=Selected Publications=
<ref>[http://dblp.uni-trier.de/pers/hd/s/Sutton:Richard_S= dblp: Richard S. Sutton]</ref><ref>[http://ilk.uvt.nl/icga/journal/docs/References.pdf [ICGA Journal#RefDB|ICGA Reference Database] (pdf)]</ref>
==1978==
* [[Richard Sutton]] ('''1978'''). ''Single channel theory: A neuronal theory of learning''. Brain Theory Newsletter 3, No. 3/4, pp. 72-75. [http://www.cs.ualberta.ca/%7Esutton/papers/sutton-78-BTN.pdf pdf]
==1980 ...==
* [[Richard Sutton]], [[Andrew Barto]] ('''1981'''). ''Toward a modern theory of adaptive networks: Expectation and prediction''. Psychological Review, Vol. 88, pp. 135-170. [http://www.cs.ualberta.ca/%7Esutton/papers/sutton-barto-81-PsychRev.pdf pdf]
* [[Richard Sutton]] ('''1984'''). ''[http://scholarworks.umass.edu/dissertations/AAI8410337/ Temporal Credit Assignment in Reinforcement Learning]''. Ph.D. dissertation, [https://en.wikipedia.org/wiki/University_of_Massachusetts University of Massachusetts]
* [[Richard Sutton]] ('''1988'''). ''Learning to Predict by the Methods of Temporal Differences''. [https://en.wikipedia.org/wiki/Machine_Learning_%28journal%29 Machine Learning], Vol. 3, No. 1, [https://webdocs.cs.ualberta.ca/~sutton/papers/sutton-88-with-erratum.pdf pdf]
==1990 ...==
* [[Richard Sutton]], [[Andrew Barto]] ('''1990'''). ''Time-Derivative Models of Pavlovian Reinforcement''. in [http://node.realityspline.net/ari/work/neuro/people/showpeople.php?person=faculty/mgabriel.php Michael Gabriel], [http://people.umass.edu/jwmoore/people.htm#JWMoore John Moore] (eds.) ('''1990'''). ''Learning and Computational Neuroscience: Foundations of Adaptive Networks''. [https://en.wikipedia.org/wiki/MIT_Press MIT Press], [https://webdocs.cs.ualberta.ca/~sutton/papers/sutton-barto-90.pdf pdf]* [[Doina Precup]], [[Richard Sutton]] ('''1997'''). ''Multi-time Models for Temporally Abstract Planning''. [httphttps://dblp.uni-trier.de/db/conf/nips/nips1997.html#PrecupS97 NIPS 1997], [https://papers.nips.cc/paper/1997/file/a9be4c2a4041cadbf9d61ae16dd1389e-Paper.pdf pdf]
* [[Richard Sutton]], [[Andrew Barto]] ('''1998'''). ''[http://incompleteideas.net/book/the-book.html Reinforcement Learning: An Introduction]''. [https://en.wikipedia.org/wiki/MIT_Press MIT Press]
* [[Richard Sutton]], [[Doina Precup]], [[Mathematician#SSingh|Satinder Singh]] ('''1999'''). ''Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning''. [https://en.wikipedia.org/wiki/Artificial_Intelligence_(journal) Artificial Intelligence], Vol. 112, [https://people.cs.umass.edu/~barto/courses/cs687/Sutton-Precup-Singh-AIJ99.pdf pdf]
==2000 ...==
* [[Michael L. Littman]], [[Richard Sutton]], [[Mathematician#SSingh|Satinder Singh]] ('''2001'''). ''Predictive Representations of State''. [http://dblp.uni-trier.de/db/conf/nips/nips2001.html#LittmanSS01 NIPS 2001], [http://web.eecs.umich.edu/~baveja/Papers/psr.pdf pdf]
* [[Richard Sutton]], [http://dblp.uni-trier.de/pers/hd/t/Tanner:Brian Brian Tanner] ('''2005'''). ''Temporal-Difference Networks''. Advances in Neural Information Processing Systems 17, pages 1377-1384. [http://www.cs.ualberta.ca/%7Esutton/papers/sutton-tanner-04.pdf pdf]* [[David Silver]], [[Richard Sutton]], [[Martin Müller]] ('''2007'''). ''Reinforcement learning of local shape in the game of [[Go]]''. In Twentieth International Joint Conference on Artificial Intelligence ([[Conferences#IJCAI2007|20th IJCAI), pages 1053-1058]], [https://en.wikipedia.org/wiki/Hyderabad,_India Hyderabad, India]. [http://webdocs.cs.ualberta.ca/%7Emmueller~mmueller/ps/silver-ijcai2007.pdf pdf]* [[Richard Sutton]], [[Csaba Szepesvári]], [[Hamid Reza Maei]] ('''2008'''). ''A Convergent O(n) Algorithm for Off-policy Temporal-difference Learning with Linear Function Approximation''. [https://dblp.uni-trier.de/db/conf/nips/nips2008.html#SuttonSM08 NIPS 2008], available as [httphttps://wwwproceedings.sztakineurips.hucc/paper/%7Eszcsaba2008/papersfile/gtdnips08e0c641195b27425bb056ac56f8953d24-Paper.pdf pdf] (draft)
* [[David Silver]], [[Richard Sutton]], [[Martin Müller]] ('''2008'''). ''Sample-Based Learning and Search with Permanent and Transient Memories''. In Proceedings of the 25th International Conference on Machine Learning, [http://icml2008.cs.helsinki.fi/papers/564.pdf pdf]
* [[Maria Cutumisu]], [[Michael Bowling]], [[Duane Szafron]], [[Richard Sutton]] ('''2008'''). ''Agent Learning using Action-Dependent Learning Rates in Computer Role-Playing Games''. [https://www.aaai.org/Library/AIIDE/aiide08contents.php Proceedings of the Fourth Artificial Intelligence and Interactive Digital Entertainment Conference], [https://webdocs.cs.ualberta.ca/~duane/publications/pdf/2008aiide.pdf pdf]
* [[Hamid Reza Maei]], [[Csaba Szepesvári]], [[Shalabh Bhatnagar]], [[Doina Precup]], [[David Silver]], [[Richard Sutton]] ('''2009'''). ''Convergent Temporal-Difference Learning with Arbitrary Smooth Function Approximation.'' Accepted in Advances in Neural Information Processing Systems 22, Vancouver, BC. December 2009[https://dblp.uni-trier. MIT Pressde/db/conf/nips/nips2009. html#MaeiSBPSS09 NIPS 2009], [httphttps://bookspapers.nips.cc/paperspaper/files2009/nips22file/NIPS2009_11213a15c7d0bbe60300a39f76f8a5ba6896-Paper.pdf pdf]* [[Richard Sutton]], [[Hamid Reza Maei]], [[Doina Precup]], [[Shalabh Bhatnagar]], [[David Silver]], [[Csaba Szepesvári]], [[Eric Wiewiora]]. ('''2009'''). ''[https://dl.acm.org/doi/10.1145/1553374.1553501 Fast Gradient-Descent Methods for Temporal-Difference Learning with Linear Function Approximation]''. In Proceedings of the 26th International Conference on Machine Learning (ICML-09). [httphttps://wwwdblp.sztakiuni-trier.hude/db/~szcsabaconf/papersicml/GTD-ICML09icml2009.pdf pdfhtml#SuttonMPBSSW09 ICML 2009]
==2010==
* [[Hamid Reza Maei]], [[Csaba Szepesvári]], [[Shalabh Bhatnagar]], [[Richard Sutton]] ('''2010'''). ''Toward Off-Policy Learning Control with Function Approximation''. [httphttps://wwwdblp.incompleteideasuni-trier.netde/db/conf/suttonicml/publicationsicml2010.html#GQ MaeiSBS10 ICML 2010], [https://icml.cc/Conferences/2010/papers/627.pdf pdf]* [[Hamid Reza Maei]], [[Richard Sutton]] ('''2010'''). ''[https://www.researchgate.net/publication/215990384_GQlambda_A_general_gradient_algorithm_for_temporal-difference_prediction_learning_with_eligibility_traces GQ(λ): A general gradient algorithm for temporal-difference prediction learning with eligibility traces]''. In Proceedings of the Third Conference on Artificial General Intelligence[https://agi-conf.org/2010/ AGI 2010]
* [[David Silver]], [[Richard Sutton]], [[Martin Müller|Martin Mueller]] ('''2013'''). ''Temporal-Difference Search in Computer Go''. Proceedings of the [http://icaps13.icaps-conference.org/technical-program/workshop-program/planning-and-learning/ ICAPS-13 Workshop on Planning and Learning], [http://webdocs.cs.ualberta.ca/~sutton/papers/SSM-ICAPS-13.pdf pdf]
* [[Huizhen Yu]], [[A. Rupam Mahmood]], [[Richard Sutton]] ('''2017'''). ''On Generalized Bellman Equations and Temporal-Difference Learning''. Canadian Conference on AI 2017, [https://arxiv.org/abs/1704.04463 arXiv:1704.04463]
* [http://videolectures.net/icml09_sutton_fgdm/ Fast Gradient-Descent Methods for Temporal-Difference Learning with Linear Function Approximation], [https://en.wikipedia.org/wiki/VideoLectures.net videolecture] by Richard Sutton, June 2009
* [https://deepmind.com/blog/deepmind-office-canada-edmonton/ DeepMind expands to Canada with new research office in Edmonton, Alberta] by [[Demis Hassabis]], [[DeepMind]], July 5, 2017
* [https://en.chessbase.com/post/standing-on-the-shoulders-of-giants Standing on the shoulders of giants] by [[Albert Silver]], [[ChessBase|ChessBase News]], September 18, 2019
=References=
<references />
 
'''[[People|Up one level]]'''
[[Category:Researcher|Sutton]]

Navigation menu