Retour au feed
Hacker News (AI)·

GateGPT: 56k tokens per second Transformer (KV cache) on FPGA at 80 MHz

Signal
65
Hype
25
En 3 lignesGateGPT atteint 56k tokens/sec sur FPGA à 80 MHz en optimisant le cache KV des Transformers. Démonstration d'accélération matérielle pour l'inférence.
Lire la source
Ton avis ?
InfrastructureBenchmarks

Résumé généré par Claude — vérifié par l'humain