Add GraphTokenizer implementation - #260
Open
baihchou8787 wants to merge 6 commits into
Open
baihchou8787 wants to merge 6 commits into
baihchou8787 wants to merge 6 commits into
Conversation
Author
|
已补充 GraphTokenizer 改进(commit
验证: |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
概述
新增 GraphTokenizer 算法实现,包括图序列化、Graph-BPE 分词、GraphBERT/GraphGTE 模型、训练示例与使用文档。
新增
QM9、OGBG-molhiv和Peptides-struct三个分子数据集的数据加载器,支持数据下载、预处理及官方训练/验证/测试划分。最新提交(989b48a)
GraphGTE为随机初始化,GTE 结果须在官方 checkpoint 加载完成后重新验证。测试
此前在
TL_BACKEND=torch环境下运行相关测试:本次提交已通过
compileall与git diff --check;当前环境未安装pytest,因此未重复运行聚焦 pytest。