Page 118 sur 192

ToutHaut signalRécent
7679 articles
Reddit r/LocalLLaMA·

2X tk/s (from 19.4 -> 38.1 tk/s on 1 x MI50) Playing with a hypothesis like speculative decoding.. but instead of an additional side model, exploiting that I can run multiple computations side-by-side AS IF I had Qwen3.6-27B loaded twice in memory - small quants don't use all the available compute.

Optimisation de décodage spéculatif : doublement du débit (19.4 → 38.1 tk/s sur MI50) en exécutant plusieurs calculs parallèles avec le même modèle quantifié Q8, exploitant que les petites quantifications n'utilisent que 25% de la puissance de calcul disponible.

Open source
SIG
65
HYP
25
Reddit r/LocalLLaMA·

I bundled a fully local LLM inside my Unity game. No internet, no cloud, no API key. The conversation is the gameplay.

Un développeur a intégré un LLM local dans son jeu Unity 'Simulation Simulator', sans internet ni API. Le gameplay repose sur des conversations naturelles avec une IA qui génère 5 fins différentes selon les interactions. Le jeu explore la théorie de la simulation et la philosophie via un dialogue organique et non-scriptés.

LlamaGénération de codeAgents IA
SIG
65
HYP
45
The Decoder·

Microsoft tightens rules for conflict zones after investigation into Israel's military use of Azure

Microsoft conclut son enquête sur l'utilisation militaire d'Azure par Israël et met en place de nouvelles vérifications des droits humains. Le rapport ne précise pas le contenu réel des données militaires examinées et omet les départs de personnel chez Microsoft Israël. Enjeux : infrastructure cloud, surveillance de masse et sélection de cibles IA à Gaza.

RégulationSécurité IAAlignement
SIG
65
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> google /</span> skills

Google publie Skills, une collection d'outils et de compétences pour agents IA compatibles avec les produits Google. Le projet fournit des intégrations et des capacités étendues pour les systèmes multi-agents.

Agents IAMulti-agentsDeepMind
SIG
65
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> Andyyyy64 /</span> whichllm

Outil open-source pour identifier le meilleur LLM local selon votre matériel. Classement basé sur benchmarks réels et à jour, pas sur le nombre de paramètres. Une commande pour tester instantanément.

Open sourceBenchmarksOutils
SIG
65
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> xerrors /</span> Yuxi

Plateforme multi-tenant Agent Harness intégrant base de connaissances LightRAG et graphes de connaissances. Stack : LangChain + Vue + FastAPI, support DeepAgents, MinerU PDF, Neo4j, MCP.

Agents IARAGMCP
SIG
65
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> Andyyyy64 /</span> whichllm

Outil CLI pour identifier le meilleur LLM local sur son matériel. Classement basé sur benchmarks réels et actualisés, pas sur le nombre de paramètres. Exécution instantanée en une commande.

Open sourceOutilsBenchmarks
SIG
65
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> google /</span> skills

Google publie Skills, une collection d'outils et de ressources pour développer des agents IA compatibles avec les produits Google. Le projet fournit des compétences réutilisables et des intégrations pour construire des systèmes multi-agents.

Agents IAMulti-agentsDeepMind
SIG
65
HYP
25
Reddit r/LocalLLaMA·

llama-server router: a model pinned to one GPU still grabs a CUDA context on every card, so it OOMs when my others are full. Am I missing a flag or is this just how it is?

Utilisateur signale que llama-server en mode router alloue un contexte CUDA sur toutes les cartes GPU même quand un modèle est épinglé à une seule. Gemma 4B sur RTX 5060 Ti réserve ~256 MiB sur chaque 3090 et ~120 sur 4060 Ti, causant des OOM quand les 3090s sont saturées. Le problème vient de l'héritage de l'env du router par les enfants sans `CUDA_VISIBLE_DEVICES` par modèle.

LlamaInfrastructureOpen source
SIG
65
HYP
15
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> aaif-goose /</span> goose

Goose est un agent IA open-source extensible qui dépasse les simples suggestions de code : il installe, exécute, édite et teste avec n'importe quel LLM.

Agents IAGénération de codeOpen source
SIG
65
HYP
35
Reddit r/MachineLearning·

Two independent ML/CV researchers (M.Eng, ex-research-institute in Europe) looking for an arXiv cs.CV endorser for a nearly finished paper. Happy to share the full draft, repo, or talk collaboration [D]

Deux chercheurs indépendants (M.Eng, ex-instituts européens) cherchent un endorser arXiv cs.CV pour leur papier Locate-SAM2. Ils proposent un pipeline training-free connectant LocateAnything-3B (NVIDIA) à SAM 2.1 (Meta) via un adaptateur léger. Sur RefCOCO val : 0.772 mIoU vs 0.717 pour Grounding DINO Base. Code et papier disponibles.

VisionOpen sourcePapers
SIG
65
HYP
25