Archives

June 2026

2731 articles

arXiv cs.CL·

HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule

HKJudge is the first sentence-level expert-annotated legal discourse corpus. It contains ~290k sentences and ~6.5M tokens from Hong Kong criminal judgments across all court levels, annotated by legal linguistics experts. Two benchmark tasks: rhetorical role classification (26 categories) and legal element extraction. Evaluation on BERT models, open-source and commercial LLMs.

BenchmarksPapersFine-tuning
SIG
82
HYP
15
arXiv cs.CL·

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

Study on LLM overgeneralization beyond training data. Authors propose the Piggyback Hypothesis: chat-template tokens propagate finetuned behaviors to out-of-distribution domains. They introduce Token-Regularized Finetuning (TReFT) to mitigate emergent misalignment, achieving 33.5% more reduction than data interleaving on Llama-3.1-8B legal domain finetuning.

Fine-tuningAlignmentAI safety
SIG
78
HYP
25
arXiv cs.LG·

Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws

Study of scaling laws for language model pretraining in data-constrained regime. Authors propose MIR (masked-input regularization), an auxiliary next-token prediction loss on randomly masked inputs, and SoftQ, a scaling law coupling model and data size under repeated data. MIR improves validation loss on 72M–1.4B models and equals ~1.3× more unique training data.

Fine-tuningBenchmarks
SIG
78
HYP
15
arXiv cs.LG·

Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring

PROBE, a multi-stage pipeline, diagnoses 802.11 packet captures by combining deterministic PCAP-to-text normalization, multi-model ensembles, and evidence-grounded reliability scoring (without LLM self-assessment). On 87 enterprise Wi-Fi captures, achieves F1=0.957 vs 0.871 expert baseline, eliminates LLM hallucinations and uncalibrated confidence scores.

ReasoningEvalsBenchmarks
SIG
78
HYP
15
arXiv cs.LG·

TALAN: Task-Aligned Latent Adaptation Networks for Targeted Post-Training of Large Language Models

TALAN (Task-Aligned Latent Adaptation Networks) combines a low-rank adapter with a sequence-conditioned latent side path inserted into the transformer's residual stream. Tested on four Qwen3 backbones and four STEM/code benchmarks, TALAN improves LoRA (+1.41 points) and DoRA (+1.85 points) baselines with <1% additional parameters and 1.01-1.02x inference overhead.

Fine-tuningReasoningCode generation
SIG
78
HYP
15
arXiv cs.LG·

The Geography of Algorithmic Judgment: LLM Intermediaries, Place Identity, and Racial Steering in Housing Search

Behavioral audit of 7 LLMs (open-weight and closed-source) across 4 US cities reveals racial steering emerges from interaction between user identity, stated preferences, and the model's learned spatial representations. Steering is not uniform: preference-conditioned testing often amplifies bias. Results do not generalize across local markets.

AI safetyAlignmentEvals
SIG
78
HYP
25
Reddit r/LocalLLaMA·

llama-server router: a model pinned to one GPU still grabs a CUDA context on every card, so it OOMs when my others are full. Am I missing a flag or is this just how it is?

User reports llama-server router mode allocates CUDA context on all GPUs even when a model is pinned to one. Gemma 4B on RTX 5060 Ti reserves ~256 MiB per 3090 and ~120 on 4060 Ti, causing OOM when 3090s are saturated. Issue stems from child processes inheriting router env without per-model `CUDA_VISIBLE_DEVICES` support.

LlamaInfrastructureOpen source
SIG
65
HYP
15
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> ggml-org /</span> llama.cpp

llama.cpp is a C/C++ project for LLM inference. Available on GitHub, it enables efficient execution of language models in C/C++.

LlamaCode generationOpen source
SIG
75
HYP
15
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> RyanCodrai /</span> turbovec

TurboVec is a vector index built on TurboQuant, written in Rust with Python bindings. Optimized for high-performance vector search.

Vector searchEmbeddingsOpen source
SIG
45
HYP
15
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> aaif-goose /</span> goose

Goose is an open-source, extensible AI agent that goes beyond code suggestions—it installs, executes, edits, and tests with any LLM.

AI AgentsCode generationOpen source
SIG
65
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> cline /</span> cline

Cline is an autonomous coding agent available as SDK, IDE extension, or CLI assistant. Open-source tool for automating development tasks.

AI AgentsCode generationOpen source
SIG
45
HYP
35
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> luongnv89 /</span> claude-howto

Visual, example-driven guide to Claude Code from basic concepts to advanced agents, with copy-paste templates for immediate implementation.

Claude CodeAI AgentsPrompt engineering
SIG
45
HYP
55
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> RyanCodrai /</span> turbovec

TurboVec is a vector index built on TurboQuant, written in Rust with Python bindings. Optimized for high-performance vector search.

Vector searchEmbeddingsOpen source
SIG
45
HYP
15
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> ashishpatel26 /</span> 500-AI-Agents-Projects

Curated collection of 500 AI agent projects across healthcare, finance, education, retail and more. Practical use cases with links to open-source implementations.

AI AgentsOpen source
SIG
35
HYP
45
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> SudoHopeX /</span> KaliGPT

KaliGPT is a multi-model agentic AI (Gemini, ChatGPT, Ollama, OpenRouter) fine-tuned for ethical hackers and offensive security students. Streamlines penetration testing workflows.

AI AgentsFine-tuningGemini
SIG
45
HYP
55
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> AstrBotDevs /</span> AstrBot

AstrBot is an AI agent framework integrating multiple IM platforms, LLMs, and plugins. Open-source alternative to OpenClaw for building AI assistants.

AI AgentsOpen sourceTools
SIG
35
HYP
45
Reddit r/MachineLearning·

Two independent ML/CV researchers (M.Eng, ex-research-institute in Europe) looking for an arXiv cs.CV endorser for a nearly finished paper. Happy to share the full draft, repo, or talk collaboration [D]

Two independent researchers (M.Eng, ex-European research institutes) seeking arXiv cs.CV endorser for Locate-SAM2 paper. Training-free pipeline connecting NVIDIA's LocateAnything-3B to Meta's SAM 2.1 via lightweight adapter. RefCOCO val: 0.772 mIoU vs 0.717 for Grounding DINO Base. Code and paper available.

VisionOpen sourcePapers
SIG
65
HYP
25
Reddit r/MachineLearning·

Got told my open-source model experiments are too scattered. I'm organizing a journal to provide clarity before structuring the first git release. Is this readable for ML folks who aren’t in mech interp? Open to ANY feedback [D]

Mechanistic interpretability experiment on Qwen3.5-35B-A3B: a routed expert (E114, layer 14) correlates with first-person self-examination register during generation. Author documents results before git release, using W/S/Q decomposition of MoE routing.

QwenOpen source
SIG
45
HYP
25