mindspore/docs/api/api_python/dataset_text/mindspore.dataset.text.Fast...

20 lines
852 B
ReStructuredText
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

mindspore.dataset.text.FastText
================================
.. py:class:: mindspore.dataset.text.FastText
用于将tokens映射到矢量的FastText对象。
.. py:method:: from_file(file_path, max_vectors=None)
从文件构建FastText向量。
参数:
- **file_path** (str) - 包含向量的文件的路径。预训练向量集的文件后缀必须是 `*.vec`
- **max_vectors** (int可选) - 用于限制加载的预训练向量的数量。
大多数预训练的向量集是按词频降序排序的。因此,在如果内存不能存放整个向量集,或者由于其他原因不需要,
可以传递 `max_vectors` 限制加载数量。默认值None无限制。
返回:
FastText 根据文件构建的FastText向量。