Accelerate fastText-style word vector training with GPU efficiency.
Project details
floret-torch focuses on training fastText-style word vectors on GPUs, delivering significant speed improvements compared to CPU training. This open-source solution seamlessly integrates with spaCy, enabling enhanced embeddings for extensive corpora and rare words through its optimized training approach. Ideal for those looking to boost performance while maintaining quality.
floret-torch is an advanced tool for training fastText-style word vectors on GPU, significantly increasing training speed—up to 20× faster than CPU-based floret models—without compromising quality. This repository builds on the foundation of fastText by leveraging the power of GPUs, making it ideal for scenarios where CPU training becomes a bottleneck, such as when handling large corpora or using limited computing resources like a laptop.
The performance of floret-torch has been benchmarked using the text8 dataset, showcasing impressive results:
| Epochs | Model | Hardware | Train Time | WS353 |
|---|---|---|---|---|
| 1 | CPU floret | i7-7500U, 4 threads | 188.1s | 0.2875 |
| 1 | floret-torch | GeForce 940MX (2016 laptop GPU) | 178.5s | 0.3582 |
| 3 | CPU floret | i7-7500U, 4 threads | 593.1s | 0.4738 |
| 3 | floret-torch | Colab T4 | ~30s | 0.4740 |
At 3 epochs, floret-torch achieves comparable quality to CPU floret while drastically reducing training time—demonstrating an unmatched efficiency advantage for users.
To train the model, follow these steps:
python -m floret_torch.train \
--input data/text8 --output out/vectors \
--model cbow --dim 300 --minn 4 --maxn 5 --hashCount 2 --bucket 50000 \
--epoch 3 --lr 0.05 --batch 8192 --dtype fp32 --device cuda
python -m spacy init vectors en out/vectors.floret out/pipeline --mode floret
Inspired by Amazon’s SageMaker BlazingText service, floret-torch democratizes access to high-speed word vector training on local GPU resources, making it open-source and more accessible for developers and researchers.
floret-torch offers a robust solution for training fastText-style word vectors on GPU, enhancing productivity and efficiency for those working with large datasets. With its compatibility with spaCy, it ensures that trained models can be seamlessly integrated into NLP applications.
Comments
0Start the conversation
Share the first comment.