mindspore/docs/api/api_python/dataset_text/mindspore.dataset.text.GloV...

25 lines
1.1 KiB
ReStructuredText
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

mindspore.dataset.text.GloVe
=============================
.. py:class:: mindspore.dataset.text.GloVe
用于将tokens映射到向量的GloVe对象。
.. py:method:: from_file(file_path, max_vectors=None)
从文件构建CharNGram向量。
参数:
- **file_path** (str) - 包含向量的文件路径。预训练向量集的格式必须是 `glove.6B.*.txt`
- **max_vectors** (int可选) - 用于限制加载的预训练向量的数量。
大多数预训练的向量集是按词频降序排序的。因此,如果内存不能存放整个向量集,或者由于其他原因不需要,
可以传递 `max_vectors` 限制加载数量。默认值None无限制。
返回:
GloVe 根据文件构建的GloVe向量。
异常:
- **RuntimeError** - `file_path` 参数所指向的文件非法或者包含的数据异常。
- **ValueError** - `max_vectors` 参数值错误。
- **TypeError** - `max_vectors` 参数不是整数类型。