metaai·lightalo unofficial · independent
Universe / Open Source Infra / fastText
Open Source Infra · 2016

fastText

Language detection for 176 languages in under a megabyte

open source archived

Latest: v0.9.2 final line; repository archived March 19, 2024

Library for fast word representations and text classification built on subword n-grams, from FAIR in 2016. It shipped pre-trained word vectors for 157 languages and a famously tiny language-identification model, making industrial-strength NLP possible on a laptop CPU. The repo was archived in March 2024 — mission accomplished.

Why it matters

Made industrial-strength text classification and word embeddings run on a laptop CPU using subword n-grams, and shipped pretrained vectors for 157 languages plus a famously tiny language-identification model that still runs inside data pipelines everywhere. Archived in March 2024 after its ideas became standard practice.

Facts

Try it yourself

Sources

More in Open Source Infra

PyTorchfaissProphetPyTorch Domain LibrariesParlAINevergrad

Read the Open Source Infra story on the sky →

✦ Open on the map Explore Open Source Infra Quiz me