mindspore/docs/api/api_python/dataset_text/mindspore.dataset.text.Fast...

25 lines
1.1 KiB
ReStructuredText
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

mindspore.dataset.text.FastText
================================
.. py:class:: mindspore.dataset.text.FastText
用于将tokens映射到向量的FastText对象。
.. py:method:: from_file(file_path, max_vectors=None)
从文件构建FastText向量。
参数:
- **file_path** (str) - 包含向量的文件路径。预训练向量集的文件后缀必须是 `*.vec`
- **max_vectors** (int可选) - 用于限制加载的预训练向量的数量。
大多数预训练的向量集是按词频降序排序的。因此,如果内存不能存放整个向量集,或者由于其他原因不需要,
可以传递 `max_vectors` 限制加载数量。默认值None无限制。
返回:
FastText 根据文件构建的FastText向量。
异常:
- **RuntimeError** - `file_path` 参数所指向的文件非法或者包含的数据异常。
- **ValueError** - `max_vectors` 参数值错误。
- **TypeError** - `max_vectors` 参数不是整数类型。