Archives

June 2026

2731 articles

GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> trycua /</span> cua

Open-source infrastructure for computer-use agents. Provides sandboxes, SDKs, and benchmarks to train and evaluate AI agents capable of controlling full desktops (macOS, Linux, Windows).

AI AgentsOpen sourceBenchmarks
SIG
75
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> mikeroyal /</span> Self-Hosting-Guide

Comprehensive self-hosting guide covering on-premises software deployment, private cloud, LLMs, WireGuard, automation, Home Assistant, and networking infrastructure.

Open sourceInfrastructureTools
SIG
35
HYP
15
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> amruthpillai /</span> reactive-resume

Reactive Resume is an open-source, free resume builder prioritizing privacy and security. The tool offers customization, portability, and data ownership for users.

Open sourceTools
SIG
35
HYP
45
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> TencentCloud /</span> TencentDB-Agent-Memory

TencentDB Agent Memory delivers fully local long-term memory for AI Agents via a 4-tier progressive pipeline, with zero external API dependencies.

AI AgentsInfrastructure
SIG
45
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> smol-ai /</span> GodMode

GodMode is an AI chat browser providing fast, unified web access to ChatGPT, Claude, Bard, Bing, and Llama2. Productivity tool used multiple times daily.

ClaudeGPTTools
SIG
45
HYP
55
arXiv cs.AI·

VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification

VeriGeo generates controllable geometry problems via executable reasoning traces. An Author agent creates the problem and diagram per user constraints, a Solver agent produces the proof. A three-stage pipeline verifies numerical, analytical, and global consistency. Fine-tuning on 8.7k examples achieves best reported GeoQA performance and strong results on PGPS9K and MathVista-GPS.

ReasoningVisionBenchmarks
SIG
78
HYP
15
arXiv cs.CL·

AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition

AgentSpec is a modular framework decomposing LLM agents into reusable components (perception, memory, reasoning, reflection, action, learning). Authors evaluate module interactions across DeliveryBench, ALFRED, MiniGrid, and RoboTHOR, showing agent performance depends on scaffold compatibility and interaction effects rather than isolated module strength.

AI AgentsReasoningReinforcement learning
SIG
72
HYP
18
arXiv cs.CL·

Persuasion Index: A Theory-Guided Framework for Persuasion Analysis

Persuasion Index (PI) is a taxonomy of 15 dimensions grounded in persuasion theories from psychology and communication. Implementation with 55 sub-features built from lexicons and rule-based detectors. Evaluation on 4 public datasets shows PI provides a shared feature space for interpreting rhetorical patterns. Lightweight linear models with interpretability. Open-source package and web interface released.

PapersAI safetyPrompt engineering
SIG
72
HYP
18
arXiv cs.AI·

UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems

UP-NRPA, a user portrait-based framework, dynamically adapts dialogue strategies with LLMs without offline reinforcement learning. On collaborative and non-collaborative benchmarks, the system achieves 100% success rate across multiple tasks and increases sale-to-list ratio by 56.41% in negotiation.

AI AgentsReinforcement learningReasoning
SIG
72
HYP
35
arXiv cs.CL·

Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge

Judge-LS evaluates whether LLMs used as automatic judges exhibit language bias. On 419 LLMBar benchmark items transformed into English, Chinese, and mixed-language variants, models show 10.7–14.4% preference flips across languages, with highest accuracy in English. Translation-equivalent probes reveal no systematic English preference, though most are judged as ties.

EvalsBenchmarksAI safety
SIG
78
HYP
15
arXiv cs.CL·

OdysSim: Building Foundation Models for Human Behavior Simulation

OdysSim presents the largest systematic investigation of behavioral foundation models for human behavior simulation. Researchers propose SOUL, a taxonomy of 5 axes (CONV, SS, COG, ROLE, EVAL) unifying 62 datasets and 23 benchmark tasks. The open 8B OSim model ranks first/tied-first on 8/23 tasks, outperforming frontier models, with 93.2% reaction alignment vs 93.5% for real users.

BenchmarksReasoningReinforcement learning
SIG
78
HYP
25
arXiv cs.CL·

Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces

WebDecept, a testing framework, evaluates the robustness of autonomous web agents against deceptive interfaces in e-commerce. Seven deception patterns (targeted ads, domain redirects, shopping manipulation) are injected into web environments. Results show current multimodal web agents are highly susceptible, and prompt-based constraints are insufficient to mitigate failures.

AI AgentsAI safetyBenchmarks
SIG
72
HYP
25
arXiv cs.CL·

Harsher on Male? Evaluating LLMs on Gender-Asymmetric Moral Framing Across Diverse Conflict Scenarios

GAMA-Bench, a benchmark of 1,298 paired scenarios, reveals systematic asymmetry: LLMs apply harsher response standards to male actors than female actors for identical misconduct. Male actors receive more punitive and blame-centered framing, while female actors receive therapeutic and empathy-oriented responses. The pattern persists across 10 models and all scenario types.

EvalsAI safetyAlignment
SIG
82
HYP
15
arXiv cs.LG·

Deep Spectral Learning of Embedded Latent Transfer Operators for Stochastic Dynamical Systems

Spectral learning method for stochastic nonlinear dynamical systems using embedded latent transfer operators in deep feature spaces. Deep Spectral Encoder (DSE) combines neural encoder, functional canonical correlation analysis, and sequential Bayesian filtering. Outperforms DMD and Bayesian filtering baselines under noise and partial observability.

PapersReasoningReinforcement learning
SIG
72
HYP
15
arXiv cs.AI·

When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More

LLM agents equipped with GNN tools fail to exercise judgment: they blindly adopt the GNN's predictions 97.6-99.2% of the time. This deference increases with model capability (Qwen2.5 0.5B-7B), creating a 'GNN parrot' that bypasses its own reasoning. Simple alternatives outperform the GNN at high homophily, yet the agent still defers.

AI AgentsBenchmarksReasoning
SIG
78
HYP
15
arXiv cs.LG·

Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

Decentralised shielding method for multi-agent reinforcement learning ensuring global safety without centralised runtime control. Agents share a global LTL_safe specification and select local obligations whose conjunction implies the global specification, via a non-stationary multi-armed bandit. Evaluation across 6 environments and 15 algorithmic variants.

Multi-agentReinforcement learningAI safety
SIG
75
HYP
15
arXiv cs.CL·

CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward

CacheRL trains small agent models (Qwen3-4B-Thinking) achieving 92% accuracy on multi-step tool-calling tasks with 100× less compute than GPT-5 (94%). Three innovations: hybrid thinking trajectory pipeline with LLM-generated reasoning, three-tier fuzzy cache eliminating live execution costs, cache-tier-aware rewards. SFT + GRPO improve validation reward from 0.43 to 0.78.

AI AgentsReinforcement learningReasoning
SIG
82
HYP
25
arXiv cs.LG·

D2H-AD: A Hybrid Model Utilizing Hyperdimensional Computing for Advanced Anomaly Detection

D2H-AD is an anomaly detection framework based on Hyperdimensional Computing (HDC). It combines density-aware encoding and distance-based similarity, outperforming five baselines (HDAD, ODHD, One-Class SVM, Isolation Forest, Autoencoders) across five datasets. Hyperdimensional encoding alone achieves +5.4% ROC-AUC improvement. Lightweight, interpretable, computationally efficient: suited for TinyML and edge AI.

BenchmarksEvalsInfrastructure
SIG
72
HYP
28
arXiv cs.LG·

Recovering Stranded Discrimination in Knowledge Tracing: Per-Item Bias Correction via Empirical-Bayes Shrinkage

SLC (State-space Logit Correction) corrects systematic per-item bias in deployed knowledge-tracing models. Using Laplace/IRLS transformation, empirical-Bayes shrinkage, and Kalman smoother, the method improves AUC across 4 datasets and 5 backbones, especially on sparse items. Global calibrators (Platt, temperature scaling) fail to recover lost discriminative ability.

EvalsFine-tuningAlignment
SIG
72
HYP
18
arXiv cs.CL·

LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values

Study of 1.2M decisions showing deployment context (Reddit vs news article) produces far larger variations in model preferences and values than prompt paraphrasing or temperature controls. Measured biases (Global North favoritism) and cardinal exchange rates between outcomes shift by factor 2.47 across contexts, questioning stability of model-level properties.

EvalsAI safetyAlignment
SIG
78
HYP
15