文献详情 >M3-20M: A Large-Scale Multi-Mo... 收藏

arXiv

M3-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery

作者：Guo, Siyuan Wang, Lexuan Jin, Chang Wang, Jinxian Peng, Han Shi, Huayang Li, Wengen Guan, Jihong Zhou, Shuigeng

作者机构：Department of Computer Science and Technology Tongji University No. 4800 Cao’an Road Shanghai201804 China Shanghai Key Lab of Intelligent Information Processing School of Computer Science Fudan University 2005 Songhu Road Shanghai200438 China

出版物：《arXiv》 (arXiv)

年卷期：2024年

核心收录：

主　　题：Web crawler

摘要：This paper introduces M3-20M, a large-scale Multi-Modal Molecule dataset that contains over 20 million molecules, with the data mainly being integrated from existing databases and partially generated by large language models. Designed to support AI-driven drug design and discovery, M3-20M is 71 times more in the number of molecules than the largest existing dataset, providing an unprecedented scale that can highly benefit the training or fine-tuning of models, including large language models for drug design and discovery tasks. This dataset integrates one-dimensional SMILES, two-dimensional molecular graphs, three-dimensional molecular structures, physicochemical properties, and textual descriptions collected through web crawling and generated using GPT-3.5, offering a comprehensive view of each molecule. To demonstrate the power of M3-20M in drug design and discovery, we conduct extensive experiments on two key tasks: molecule generation and molecular property prediction, using large language models including GLM4, GPT-3.5, GPT-4, and Llama3-8b. Our experimental results show that M3-20M can significantly boost model performance in both tasks. Specifically, it enables the models to generate more diverse and valid molecular structures and achieve higher property prediction accuracy than existing single-modal datasets, which validates the value and potential of M3-20M in supporting AI-driven drug design and discovery. The dataset is available at https://***/bz99bz/M-3. Copyright © 2024, The Authors. All rights reserved.

本地馆藏 | 借阅须知 | 我要预约

已订购，未入库

sda

目录详情 | 试阅读 |

读者评论与其他读者分享你的观点

学校读者

用户名:未登录

我的评分

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

时间限定

文献类型

馆藏选择

核心期刊

语言

文献类型

帮助

文字说明：

检索规则说明：

检索范例：

分类表

所选分类

看过本文的还看了

相关文献

该作者的其他文献

CADAL相关文献

M3-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery

读者评论与其他读者分享你的观点

请选择收藏分类：

建议与咨询 留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

时间限定

文献类型

馆藏选择

核心期刊

语言

文献类型

帮助

文字说明：

检索规则说明：

检索范例：

分类表

所选分类

看过本文的还看了

相关文献

该作者的其他文献

CADAL相关文献

M3-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery

读者评论 与其他读者分享你的观点

请选择收藏分类： 新增自定义分类 确定 取消

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

读者评论与其他读者分享你的观点

请选择收藏分类：