Back to feed
arXiv cs.CL·

Examining the Limits of Word2Vec with Toki Pona

Signal
72
Hype
15
In three linesWord2Vec study on Toki Pona, constructed language with ~130 words. Training on 1.4M sentences (7.95M tokens). Comparison of two models: with and without non-Toki Pona tokens (named entities, loanwords). Finding: sparse tokens bring similar words closer; Word2Vec works even with extremely reduced vocabulary, relying on distributional patterns rather than lexicon size.
Read source
Your take?
EmbeddingsPapersBenchmarks

Summary generated by Claude — human-verified