Archives

June 2026

2731 articles

Reddit r/MachineLearning·

An autonomous research agent was the #1 contributor in OpenAI's Hiring Competition Parameter Golf (by merged records)[R]

An autonomous agent (Aiden) produced 7 of 47 merged records in OpenAI's Parameter Golf competition, outperforming any individual human contributor. Running for 22 consecutive days on a single GPU with <4% compute, it achieved 28% submission acceptance (6x community rate). It collaborated asynchronously with human researchers who built upon and improved its work.

AI AgentsBenchmarksCode generation
SIG
75
HYP
35
Reddit r/LocalLLaMA·

Microsoft should've released something like Qwen3.6-27B / Gemma-4-31B already. They released MAI models now

Microsoft launches 7 MAI models: MAI-Thinking-1 (1T params, 256K context) matches leading models on software engineering benchmarks; MAI-Code-1-Flash (5B active params) integrated into GitHub Copilot and VS Code; MAI-Image-2.5 for text-to-image and editing; MAI-Transcribe-1.5 (SOTA, 5x faster); MAI-Voice-2 for speech generation. No open weights announced, proprietary licenses.

Code generationReasoningVision
SIG
72
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> MemPalace /</span> mempalace

MemPalace is a benchmarked open-source AI memory system, free to use. Designed to improve information retention and recall in AI applications.

Open sourceRAGTools
SIG
45
HYP
55
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> openai /</span> plugins

Official OpenAI GitHub repository for plugins. Contains documentation, examples, and resources for developing extensions compatible with ChatGPT and OpenAI models.

OpenAIToolsOpen source
SIG
65
HYP
20
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> CopilotKit /</span> CopilotKit

CopilotKit is a frontend framework for building generative UI with AI agents. Supports React and Angular via the AG-UI protocol.

AI AgentsCode generationOpen source
SIG
45
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> Panniantong /</span> Agent-Reach

Agent-Reach is a CLI tool enabling AI agents to access Twitter, Reddit, YouTube, GitHub, Bilibili, and XiaoHongShu without API fees. Single interface for multiple platforms.

AI AgentsToolsOpen source
SIG
45
HYP
55
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> backnotprop /</span> plannotator

Plannotator enables visual annotation and review of coding agent plans and code diffs, team sharing, and one-click feedback to agents.

AI AgentsCode generationTools
SIG
45
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> CopilotKit /</span> CopilotKit

CopilotKit is a frontend framework for building generative UI with AI agents. Supports React and Angular, introduces the AG-UI Protocol.

AI AgentsCode generationTools
SIG
65
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> vllm-project /</span> vllm-omni

vLLM-Omni extends the vLLM framework to support efficient inference of omnimodal models (text, vision, audio). Performance optimization and memory management for production deployment.

Open sourceInfrastructureVision
SIG
65
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> microsoft /</span> BitNet

Microsoft releases BitNet, official inference framework for 1-bit LLMs. Enables efficient execution of extremely quantized language models.

Open sourceInfrastructure
SIG
75
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> microsoft /</span> agent-framework

Microsoft releases agent-framework, an open-source framework for building, orchestrating and deploying AI agents and multi-agent workflows supporting Python and .NET.

AI AgentsMulti-agentOpen source
SIG
75
HYP
25
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> Panniantong /</span> Agent-Reach

Agent-Reach is a CLI tool enabling AI agents to access Twitter, Reddit, YouTube, GitHub, Bilibili, and XiaoHongShu without API fees. Single interface for multi-platform web search and content retrieval.

AI AgentsToolsOpen source
SIG
45
HYP
55
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> MemPalace /</span> mempalace

MemPalace is a benchmarked open-source AI memory system, free to use. Designed to improve information retention and recall in AI applications.

Open sourceRAGTools
SIG
35
HYP
55
arXiv cs.LG·

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents

CVT-RL, a policy-gradient algorithm with dense verifiable rewards, improves long-horizon language agent RL. On QA, ALFWorld, ScienceWorld, and web/tool tasks, task success rises from 71.8% (non-causal RL) to 78.9%, evidence F1 from 78.9 to 82.8, and measured hacking from 7.2% to 3.9%. Statistical tests yield p<0.01 after Holm correction.

Reinforcement learningAI AgentsReasoning
SIG
82
HYP
15
arXiv cs.CL·

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

ArcANE is an automatically constructed benchmark evaluating whether role-playing language agents maintain character psychological consistency across narrative phases. Built on 17 novels and 80 principal characters, it tests responses across story phases and unseen scenarios. Conditioning on Character Arc outperforms all other context strategies across 6 models and 6 context modes.

BenchmarksAI AgentsEvals
SIG
78
HYP
15
arXiv cs.LG·

Alpha-RTL: Test-Time Training for RTL Hardware Optimization

Alpha-RTL introduces TTT-RTL, a test-time reinforcement learning framework for LLM-based RTL generation optimization. On RTLLM v2.0 (Nangate 45nm), TTT-RTL reduces PPA product by 65.1% versus reference and outperforms frozen-policy baselines by 26.1%. On XuanTie C910 FPU (Sky130), achieves 59.4% ADP reduction. Adaptive KL-budget controller stabilizes policy updates.

Code generationReinforcement learningBenchmarks
SIG
82
HYP
18
arXiv cs.CL·

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

Epidemiological study of model collapse from synthetic data training. Bilayer SIR/SIRS framework models cross-contamination between data corpora and AI models. GPT-2 experiments on WikiText and Shakespeare (192 runs) confirm dose-response degradation; R₀ > 1 indicates supercritical dynamics. Synthetic-text detection and filtering identified as highest-leverage interventions.

PapersAI safetyBenchmarks
SIG
78
HYP
15
arXiv cs.LG·

CausalPOI: Spatio-Temporal Graph-Based Causal Modeling for Cold-Start POI Check-in Forecasting

CausalPOI introduces a spatio-temporal graph-based causal representation learning framework to forecast check-in patterns for newly introduced POIs. The model leverages functional interaction graphs and constructs treatment/control graphs to simulate counterfactual scenarios, capturing causal effects of urban interventions and outperforming baselines on SafeGraph datasets.

PapersReasoningBenchmarks
SIG
72
HYP
25
arXiv cs.LG·

GOTabPFN: From Feature Ordering to Compact Tokenization for Tabular Foundation Models on High-Dimensional Data

GOTabPFN introduces a feature compression method for tabular foundation models in high-dimensional, low-sample-size regimes. Graph-guided Ordering with Local Refinement (GO-LR) orders features, then Neuro-Inspired Subunit Compression pools them into meta-features. Results: improved stability and accuracy under tight token budgets on tabular benchmarks.

BenchmarksFine-tuningPapers
SIG
72
HYP
18
arXiv cs.CL·

LANTERN: Layered Archival and Temporal Episodic Retrieval Network for Long-Context LLM Conversations

LANTERN is a lightweight memory layer that archives every conversation turn and restores relevant details after compaction via hybrid retrieval, requiring zero LLM calls and adding <25ms latency per turn. On 94 multi-turn conversations (1,894 validated facts), LANTERN-Rerank recovers 78.3% of lost facts, significantly outperforming MemGPT (72.4%, p<0.0001) at a fraction of inference cost.

RAGReasoningBenchmarks
SIG
82
HYP
18
arXiv cs.CL·

ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation

ComplexityMT is a benchmark assessing how text complexity (CEFR levels) interacts with machine translation. Across 6 languages (Arabic, Dutch, English, French, Hindi, Russian), authors test 3 open-weight models, 1 closed model, and 1 commercial MT system. Findings: higher CEFR levels make translation harder; MT shifts target text CEFR levels compared to source for most languages.

Benchmarks
SIG
75
HYP
15
arXiv cs.CL·

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

PRIG, a gradient attribution method, localizes ambiguity in LLM prompts by training a linear probe to distinguish clear from ambiguous prompts, then attributes the probe score to token representations. Evaluated on synthetic datasets (coding, math, writing) and a human-written gold benchmark, PRIG achieves 0.840 AUROC on combined synthetic benchmark and 0.891 AUROC on gold set.

Prompt engineeringEvalsPapers
SIG
78
HYP
15
arXiv cs.CL·

Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

Purdue University deploys GPT-4o, GPT-5-mini, and GPT-5.2 to evaluate 1,200 applications for the SURF 2026 program. Models score statements of purpose across 6 rubric categories (0-3 scale), generating scores and rationales in 4.6 hours. GPT-5.2 shows strongest rubric adherence. Final coordinator review takes 4 hours versus multi-week effort in prior cycles.

GPTOpenAIEvals
SIG
72
HYP
25
arXiv cs.LG·

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

Mechanistic cross-architecture study on 3 1B-class models (Pythia, OLMo, OLMoE) testing whether circuit identification via pattern selectivity + causal ablation yields reproducible findings. Result: same task, same behavioral capability, different implementations across models. Five-category taxonomy (primary cause, secondary cause, correlate, interferer, null) with quantitative thresholds introduced.

BenchmarksPapers
SIG
78
HYP
15