Back to feed
Reddit r/LocalLLaMA·

A benchmark for tiny LLMs based on a real world problem: natural language file search (using monkeSearch)

Signal
72
Hype
25
In three linesBenchmark for small LLMs (<3B parameters) evaluating natural language parsing into structured JSON for file search. 9 models tested (Gemma-3 270M to DeepSeek R1 Distill 1.5B) on 80 queries covering file types, temporal context, and specificity. Results: 0.8B–1.5B models significantly outperform sub-0.5B.
Read source
Your take?
BenchmarksOpen sourceCode generationTools

Summary generated by Claude — human-verified