Page 38 of 192

AllHigh signalRecent
7679 articles
arXiv cs.CL·

Learning What to Learn: Stage-Specific Data Sets for SFT-then-RL in Small Language Model Reasoning

SFT-then-RL training framework for small language models: SFT acquires not-yet-mastered reasoning skills, RL consolidates them. Bridge mechanism transforms raw reasoning traces into learnable supervision. Critique Fine-Tuning converts zero-reward failures into diagnostic supervision. Consistent improvements across five reasoning benchmarks.

Fine-tuningReinforcement learningReasoning
SIG
75
HYP
15
arXiv cs.CL·

GlossAssist -- A Tool to Simplify Corpus Creation and Study the Effect of NLP Models in Low-Resource Documentation Settings

GlossAssist is an automated glossing tool for linguistic documentation built on CWoMP (Contrastive Word-Morpheme Pre-training). It integrates active learning: each annotator correction enriches a mutable lexicon of morpheme representations without model retraining. The interface enables field linguists to incorporate expertise directly into model behavior.

RAGFine-tuningEvals
SIG
75
HYP
15
arXiv cs.AI·

AgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement Learning

AgentJet is a distributed framework for reinforcement learning of LLM agents. Its decoupled architecture separates server nodes (GPU optimization) from client nodes (agent execution). It supports heterogeneous multi-model RL, multi-task cocktail training, fault tolerance, and live code iteration. A context tracking module with timeline merging accelerates training 1.5-10x.

AI AgentsMulti-agentReinforcement learning
SIG
75
HYP
25
arXiv cs.AI·

R-APS: Compositional Reasoning and In-Context Meta-Learning for Constrained Design via Reflective Adversarial Pareto Search

R-APS improves LLM reliability in agentic settings via reasoning-mode decomposition. Tested on planar mechanism synthesis, it delivers robustness certificates 3.5× tighter than baselines, 46% faster iterations-to-first-admission, and 2.1× Chamfer-distance reduction. No fine-tuning required; operates via structured protocol on frozen LLM.

AI AgentsReasoningRobotics
SIG
75
HYP
25
arXiv cs.AI·

SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models

SMAC-Talk extends StarCraft Multi-Agent Challenge with natural language communication to evaluate LLM-based agents in cooperative multi-agent environments. Open-source benchmark testing decentralized control, partial observability and long-horizon decision-making, including scenarios with deceptive communicators. Evaluation on Qwen3.5 models.

Multi-agentAI AgentsBenchmarks
SIG
75
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> aquasecurity /</span> trivy

Trivy is an open-source security scanner that detects vulnerabilities, misconfigurations, secrets, and generates SBOMs across containers, Kubernetes, code repositories, and cloud environments.

Open sourceAI safetyInfrastructure
SIG
75
HYP
15
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> lyogavin /</span> airllm

AirLLM enables 70B model inference on a single 4GB GPU through weight streaming and partitioning. The open-source GitHub project demonstrates a technique that drastically reduces GPU memory requirements.

Open sourceInfrastructureLlama
SIG
75
HYP
35
arXiv cs.LG·

Human-in-the-Loop Contextual Bandits for Short-Term Rental Dynamic Pricing: Structural Equivalence of Historical Warm-Up and Approval-Gated Live Learning

HITL-GB framework for short-term rental dynamic pricing: a contextual bandit algorithm generates price recommendations that a human can accept, modify, or reject. Authors show historical data is structurally equivalent to on-policy warm-up, reducing cold-start from ~150 to ~30 episodes. Validated on 1,461 real nights (April 2022–2026).

AI AgentsReinforcement learningBenchmarks
SIG
75
HYP
15
arXiv cs.AI·

DELTAMEM: Incremental Experience Memory for LLM Agents via Residual Trees

DeltaMem organizes LLM agent experience memory into two residual trees: one stores goal-conditioned tasks as reusable skills, another stores scene-level environment knowledge. Each tree uses root nodes for generalized base experiences and delta nodes for variations, eliminating redundancy. An autonomous consolidation mechanism distills high-frequency paths into new root nodes.

AI AgentsReasoningPapers
SIG
75
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> chopratejas /</span> headroom

Headroom compresses tool outputs, logs, files, and RAG chunks before sending to LLM. Reduces token consumption by 60-95% without quality loss. Available as library, proxy, and MCP server.

RAGMCPTools
SIG
75
HYP
25