Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach

Jurenka, Irina; Kunesch, Markus; McKee, Kevin R.; Gillick, Daniel; Zhu, Shaojian; Wiltberger, Sara; Phal, Shubham Milind; Hermann, Katherine; Kasenberg, Daniel; Bhoopchand, Avishkar; Anand, Ankit; Pîslar, Miruna; Chan, Stephanie; Wang, Lisa; She, Jennifer; Mahmoudieh, Parsa; Rysbek, Aliya; Ko, Wei-Jen; Huber, Andrea; Wiltshire, Brett; Elidan, Gal; Rabin, Roni; Rubinovitz, Jasmin; Pitaru, Amit; McAllister, Mac; Wilkowski, Julia; Choi, David; Engelberg, Roee; Hackmon, Lidan; Levin, Adva; Griffin, Rachel; Sears, Michael; Bar, Filip; Mesar, Mia; Jabbour, Mana; Chaudhry, Arslan; Cohan, James; Thiagarajan, Sridhar; Levine, Nir; Brown, Ben; Gorur, Dilan; Grant, Svetlana; Hashimshoni, Rachel; Weidinger, Laura; Hu, Jieru; Chen, Dawn; Dolecki, Kuba; Akbulut, Canfer; Bileschi, Maxwell; Culp, Laura; Dong, Wen-Xin; Marchal, Nahema; Van Deman, Kelsie; Misra, Hema Bajaj; Duah, Michael; Ambar, Moran; Caciularu, Avi; Lefdal, Sandra; Summerfield, Chris; An, James; Kamienny, Pierre-Alexandre; Mohdi, Abhinit; Strinopoulous, Theofilos; Hale, Annie; Anderson, Wayne; Cobo, Luis C.; Efron, Niv; Ananda, Muktha; Mohamed, Shakir; Heymans, Maureen; Ghahramani, Zoubin; Matias, Yossi; Gomes, Ben; Ibrahim, Lila

Computer Science > Computers and Society

arXiv:2407.12687 (cs)

[Submitted on 21 May 2024 (v1), last revised 19 Jul 2024 (this version, v2)]

Title:Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach

Authors:Irina Jurenka, Markus Kunesch, Kevin R. McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, Ankit Anand, Miruna Pîslar, Stephanie Chan, Lisa Wang, Jennifer She, Parsa Mahmoudieh, Aliya Rysbek, Wei-Jen Ko, Andrea Huber, Brett Wiltshire, Gal Elidan, Roni Rabin, Jasmin Rubinovitz, Amit Pitaru, Mac McAllister, Julia Wilkowski, David Choi, Roee Engelberg, Lidan Hackmon, Adva Levin, Rachel Griffin, Michael Sears, Filip Bar, Mia Mesar, Mana Jabbour, Arslan Chaudhry, James Cohan, Sridhar Thiagarajan, Nir Levine, Ben Brown, Dilan Gorur, Svetlana Grant, Rachel Hashimshoni, Laura Weidinger, Jieru Hu, Dawn Chen, Kuba Dolecki, Canfer Akbulut, Maxwell Bileschi, Laura Culp, Wen-Xin Dong, Nahema Marchal, Kelsie Van Deman, Hema Bajaj Misra, Michael Duah, Moran Ambar, Avi Caciularu, Sandra Lefdal, Chris Summerfield, James An, Pierre-Alexandre Kamienny, Abhinit Mohdi, Theofilos Strinopoulous, Annie Hale, Wayne Anderson, Luis C. Cobo, Niv Efron, Muktha Ananda, Shakir Mohamed, Maureen Heymans, Zoubin Ghahramani, Yossi Matias, Ben Gomes, Lila Ibrahim

View PDF HTML (experimental)

Abstract:A major challenge facing the world is the provision of equitable and universal access to quality education. Recent advances in generative AI (gen AI) have created excitement about the potential of new technologies to offer a personal tutor for every learner and a teaching assistant for every teacher. The full extent of this dream, however, has not yet materialised. We argue that this is primarily due to the difficulties with verbalising pedagogical intuitions into gen AI prompts and the lack of good evaluation practices, reinforced by the challenges in defining excellent pedagogy. Here we present our work collaborating with learners and educators to translate high level principles from learning science into a pragmatic set of seven diverse educational benchmarks, spanning quantitative, qualitative, automatic and human evaluations; and to develop a new set of fine-tuning datasets to improve the pedagogical capabilities of Gemini, introducing LearnLM-Tutor. Our evaluations show that LearnLM-Tutor is consistently preferred over a prompt tuned Gemini by educators and learners on a number of pedagogical dimensions. We hope that this work can serve as a first step towards developing a comprehensive educational evaluation framework, and that this can enable rapid progress within the AI and EdTech communities towards maximising the positive impact of gen AI in education.

Subjects:	Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2407.12687 [cs.CY]
	(or arXiv:2407.12687v2 [cs.CY] for this version)
	https://doi.org/10.48550/arXiv.2407.12687

Submission history

From: Irina Jurenka [view email]
[v1] Tue, 21 May 2024 19:27:59 UTC (5,681 KB)
[v2] Fri, 19 Jul 2024 14:03:41 UTC (5,681 KB)

Computer Science > Computers and Society

Title:Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computers and Society

Title:Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators