SapientML: Synthesizing Machine Learning Pipelines by Learning from Human-Written Solutions

Saha, Ripon K.; Ura, Akira; Mahajan, Sonal; Zhu, Chenguang; Li, Linyi; Hu, Yang; Yoshida, Hiroaki; Khurshid, Sarfraz; Prasad, Mukul R.

doi:10.1145/3510003.3510226

Computer Science > Machine Learning

arXiv:2202.10451 (cs)

[Submitted on 18 Feb 2022 (v1), last revised 19 Apr 2022 (this version, v2)]

Title:SapientML: Synthesizing Machine Learning Pipelines by Learning from Human-Written Solutions

Authors:Ripon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu, Linyi Li, Yang Hu, Hiroaki Yoshida, Sarfraz Khurshid, Mukul R. Prasad

View PDF

Abstract:Automatic machine learning, or AutoML, holds the promise of truly democratizing the use of machine learning (ML), by substantially automating the work of data scientists. However, the huge combinatorial search space of candidate pipelines means that current AutoML techniques, generate sub-optimal pipelines, or none at all, especially on large, complex datasets. In this work we propose an AutoML technique SapientML, that can learn from a corpus of existing datasets and their human-written pipelines, and efficiently generate a high-quality pipeline for a predictive task on a new dataset. To combat the search space explosion of AutoML, SapientML employs a novel divide-and-conquer strategy realized as a three-stage program synthesis approach, that reasons on successively smaller search spaces. The first stage uses a machine-learned model to predict a set of plausible ML components to constitute a pipeline. In the second stage, this is then refined into a small pool of viable concrete pipelines using syntactic constraints derived from the corpus and the machine-learned model. Dynamically evaluating these few pipelines, in the third stage, provides the best solution. We instantiate SapientML as part of a fully automated tool-chain that creates a cleaned, labeled learning corpus by mining Kaggle, learns from it, and uses the learned models to then synthesize pipelines for new predictive tasks. We have created a training corpus of 1094 pipelines spanning 170 datasets, and evaluated SapientML on a set of 41 benchmark datasets, including 10 new, large, real-world datasets from Kaggle, and against 3 state-of-the-art AutoML tools and 2 baselines. Our evaluation shows that SapientML produces the best or comparable accuracy on 27 of the benchmarks while the second best tool fails to even produce a pipeline on 9 of the instances.

Comments:	Accepted to the Technical Track of ICSE 2022
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
Cite as:	arXiv:2202.10451 [cs.LG]
	(or arXiv:2202.10451v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2202.10451
Related DOI:	https://doi.org/10.1145/3510003.3510226

Submission history

From: Ripon Saha [view email]
[v1] Fri, 18 Feb 2022 20:45:47 UTC (1,560 KB)
[v2] Tue, 19 Apr 2022 21:23:57 UTC (1,560 KB)

Computer Science > Machine Learning

Title:SapientML: Synthesizing Machine Learning Pipelines by Learning from Human-Written Solutions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:SapientML: Synthesizing Machine Learning Pipelines by Learning from Human-Written Solutions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators