Gemma 4 QAT Unquantized Heretic is here
User releases an unofficial 4-bit quantized version of Gemma 4 26B MoE. Model intentionally diverges from original Gemma 4 in refusal and divergence mechanisms.
2731 articles
User releases an unofficial 4-bit quantized version of Gemma 4 26B MoE. Model intentionally diverges from original Gemma 4 in refusal and divergence mechanisms.
Analysis of accuracy inconsistencies in Gemma 4 quantization-aware training (QAT). The 12B model shows larger deviations from FP16 compared to MoE variants (E2B/E4B), contradicting theoretical expectations. Requests clarification on methodology and comparisons with non-QAT variants.
Cohere offers early access to its first coding model: 30B parameters with 3B active, optimized for local execution. Available on Hugging Face before official launch, the model aims to gather community feedback prior to full release.
Anvil, an open-source wrapper around llama.cpp, offers an alternative to Ollama with increased transparency: readable GGUF files, visible flags before execution, local fleet management, Hugging Face integration, and MCP support. Built to restore the local-first simplicity lost by Ollama.
Hugging Face launches Job Searcher, an AI-powered job search tool that helps candidates find relevant positions. The tool uses language models to analyze job postings and match candidate profiles.
Police forces in England and Wales have been instructed to stop using AI to draft court statements. The directive follows concerns about reliability and legal accountability of AI-generated text in judicial proceedings.
MoQ and GSQ, two new quantization methods, promise significant improvements for low-bit GGUFs. These approaches optimize model compression while preserving quality, benefiting local deployments.
Developer testing Qwen 3.6 and Gemma 4 locally on modest hardware (i5-12400, 64GB RAM, 2x GTX 1050 Ti). Achieves ~40 t/s prompt processing and 12-18 t/s generation. MoE, quantization, and speculative decoding make local LLMs viable without expensive hardware.
StepFun Step-3.7-Flash benchmarked on AMD Ryzen AI Max+ 395 with MTP speculative decoding: +27.5% generation speed (26 tok/s vs 20.4), 84.7% draft token acceptance, -14% power consumption. No prefill degradation.
Gemma 4 QAT Q4_0 benchmark on Strix Halo APU via llama.cpp Vulkan/RADV. Models tested: 12B (6.50 GiB), 26B-A4B (13.45 GiB), 31B (16.44 GiB). QAT preserves original model behavior better than post-training quantization. QAT assistant heads converted to GGUF for improved acceptance.
Awesome-llm-apps: curated collection of 100+ AI agent and RAG applications ready to run. Clone, customize, and deploy immediately.
Sakana AI, co-founded by Transformer co-author Llion Jones, launches a dedicated research lab for recursive self-improvement: AI that iteratively improves itself. The Japanese startup views RSI as an alternative to the raw compute arms race among major US labs. Anthropic warns about control risks of this technology.
US House lawmakers released a draft bill to prohibit states from enacting their own AI regulations. The federal initiative aims to establish uniform rules at the national level.
OpenAI Whisper is a speech recognition model trained on 680,000 hours of multilingual weakly supervised data. The GitHub repository includes code, pre-trained models, and benchmarks for robust speech transcription across 99 languages.
GitHub repository offering agentic AI infrastructure designed to magnify human capabilities. Focuses on integrating AI agents into personal workflows.
AI-powered job search system built on Claude Code featuring 14 skill modes, Go dashboard, PDF generation, and batch processing.
Microsoft releases VibeVoice, an open-source frontier voice AI model. The project aims to democratize voice synthesis technology with an accessible approach.
Open-source agentic AI infrastructure designed to amplify human capabilities. GitHub project focused on building autonomous agent systems.
Supabase is a Postgres development platform providing a dedicated database for building web, mobile, and AI applications.
Anthropic releases claude-code-action, an official GitHub Action integrating Claude for code analysis and generation directly in CI/CD workflows. Enables development task automation via Claude API.
IBM releases mcp-context-forge, an AI Gateway and registry proxy for MCP, A2A, and REST/gRPC APIs. Unifies endpoints with centralized discovery, guardrails, and management. Optimizes agent and tool calling, supports plugins.
Microsoft releases VibeVoice, an open-source voice synthesis model. The project aims to democratize AI voice generation with an accessible approach.
Khoj is a self-hostable open-source AI platform enabling custom agent creation, document/web access, and task automation. Compatible with GPT, Claude, Gemini, Llama, Qwen, Mistral.
Research unifying decision trees and diffusion models. Proposes bidirectional transformation between tree structures and diffusion processes, opening new perspectives on interpretability and generation.
Domino decouples causal modeling from autoregressive decoding in speculative decoding. 5.8x throughput speedup on Qwen3. Paper, code, and models released.
Practical comparison between Claude Opus 4.7 and local models (OpenCode) for a coding agent. Test on Playwright E2E suite (Laravel 12 + Livewire): Claude generates 203 tests without context compaction (1M tokens), local agent caps at 140 tests with 4 compactions on 24GB RTX 4090. Verdict: viable but not yet daily-driver; plan quality predicts implementation quality.
Meta is developing Hatch, a paid AI agent priced up to $200/month. The tool executes tasks (building tools, scheduling, sending emails) from natural language descriptions. Meta's first paid AI product, designed to diversify revenue beyond advertising.
Hugging Face introduces Persona Atlas, a tool mapping the thinking styles of famous personalities. The project analyzes cognitive patterns and decision-making approaches of public figures using language models.
xAI trained its coding models on Claude outputs for months, continuing after Anthropic revoked access via private accounts and Blackbox AI. xAI's pretraining team shrunk to under five people with several lead departures. Musk's purchased compute is now rented to Anthropic and Google instead.
An open-source voice model listens continuously and decides every 0.4 seconds whether to speak or stay silent. Unlike GPT-4o or Qwen3.5-Omni, Audio Interaction handles transcription, translation, and chat in a single stream, detecting ambient noise (coughing). Code, weights, and training data available on GitHub under Apache 2.0 license.
User reports performance degradation with Qwen 3.6 27B: enabling spec-type draft-mtp and spec-draft-n-max reduces throughput from 70 t/s to 30 t/s and GPU power from 475W to 300W, despite >50% acceptance rate. Issue appeared after recent llama.cpp update.
Smart TVs function as nodes in an AI data-scraping economy. Manufacturers collect user data at scale through connected devices to train AI models, often without explicit consent or compensation.
A r/LocalLLaMA user reports Opus (Claude 3) vastly outperforms local models and GPT for low-level systems engineering. On an AirPlay firmware modification project, only Opus succeeded at mapping firmware structure, reverse-engineering CRC checksums, and automating binary patching, while Qwen 35B and GPT failed at initial stages.
Qwen3.6-35B-A3B-Uncensored-Claude-4.6-Genesis-APEX-GGUF: merged model based on Qwen 3.6 35B with Claude 4.6 Opus distillation. Supports short thinking chains, thinking/non-thinking modes, improved function calling. APEX quantization recommended. Fully uncensored.
SpaceX leases AI computing capacity to Google for $920 million per month per SEC filing. The deal provides access to approximately 110,000 Nvidia chips to support Gemini Enterprise platform. The transaction demonstrates critical scarcity of AI infrastructure and growing interdependence among tech giants.
DeepSeek V4 Flash gains llama.cpp support via PR #24162 in early stages. Model combines frontier-level intelligence, quantization robustness (native FP4-FP8 hybrid), and efficient KV cache scaling. Currently 5-6 tps, GPU/FA support WIP, but correctness validated.
OpenAI and the Trump administration are negotiating direct government equity stake in the startup. A 'Public Wealth Fund' would distribute dividends to American citizens. Senator Bernie Sanders proposes 50% tax on AI shares. Critics fear 'too big to fail' dynamics similar to 2008 financial crisis.
Alibaba releases Qwen3.7-Plus, a multimodal agent model combining visual perception, GUI operation, and coding. In a demo, the agent autonomously built a vocabulary learning app in 11 hours, generating 10,000+ lines of code across 1,000 agent calls. Proprietary model, priced below Western frontier models.
Release of micropython-wasm 0.1a2 with new CLI. Enables running Python code in a sandboxed WebAssembly environment.
EpiEvolve is a self-evolving agent that adapts an LLM-based epidemic forecaster in streaming without weight updates. Using hierarchical episodic memory and reflection on errors, it achieves 0.629 accuracy on COVID-19 hospitalizations (vs 0.561 for static model), reducing recovery lag after regime shifts from 5 to 2 weeks.
Controlled intervention study shows RAG rewriting gains are driven by gold answer presence in rewritten context, not curation quality. Tests across Qwen2.5/3.5, GLM-4 and HotpotQA/2WikiMultihopQA: removing answer drops F1 by 28–64 points, injecting it raises F1 by +0.7 to +9.7 points. Authors release intervention runner and sentinel panel for reproducible evaluation.
Query Retrieve Conclude framework for interpreting emerging multimodal memes through open-world knowledge acquisition. Identifies missing knowledge, retrieves web evidence, synthesizes grounded background knowledge. Curated benchmark 2024-2026 with external annotations. Improves understanding and detection across three datasets and five detection tasks.
GITCO optimizes input context at inference time for Time Series Foundation Models (TSFMs) to reduce degradation from anomalous patches. The three-component framework (Gate, Router, Critic) suppresses harmful patches without weight updates, achieving -1.95% MASE improvement on TimesFM 2.5 across 53 GIFT-Eval datasets.
Uncertainty-aware framework for predicting functional behavior and material fatigue in circular factory remanufacturing. Combines convolutional encoder + LSTM to forecast 9 functional variables (Gaussian mean/variance) with parallel finite-element fatigue assessment. Achieves 96.52% mean 2%-tolerance accuracy on held-out tests.
New lossless compression approach for massive scientific data. Authors propose LBRC and NGLR, two residual coders improving existing GAE methods by 30-60% (LBRC) and additional 10-40% (NGLR) on E3SM, JHTDB, ERA5 datasets at 10^-6 to 10^-4 block-level NRMSE targets. NGLR adds causal neural predictor to reduce entropy of residual code.
LeanMarathon is a multi-agent system for reliable research-level autoformalization in Lean. It uses an evolving blueprint (Lean file serving as proof skeleton, natural-language proof graph, and shared record) coordinated by four specialized agents. On two recent papers spanning four Erdős problems, it formalizes seven target theorems with no sorry and proves 258 lemmas.
Agents' Last Exam (ALE) is a benchmark evaluating AI agents on long-horizon, economically valuable real-world tasks. Developed with 250+ industry experts, it covers 1K+ tasks across 13 industry clusters in non-physical sectors. Average full pass rate is 2.6% on the hardest tier.
arXiv study on LLM-driven program mutation: mutation chains converge rapidly toward restricted regions of program space. 87% of chains revisit 93% of previously seen structural forms. Phenomenon is robust across models and prompts, revealing systematic bias toward structural homogeneity incompatible with open-ended exploration.
Theoretical paper proposing a motivational architecture for conversational AGI agents, reinterpreting OpenPsi and MetaMo frameworks. Homeostasis redefined in linguistic terms (competence, uncertainty reduction, affiliation, legitimacy, aesthetic coherence). Ten-stage pipeline separating cognitive modulation from situational appraisal, with dual strategy blending urgency-driven response and multi-goal optimization.
Researchers propose a zero-knowledge VM architecture to verify actual GPU floating-point computation during frontier AI training without disclosing model architecture. Protocol combines pre-committed training specs, inter-node network observations, and on-the-fly Merkle commitments. Estimated single-digit-percent overhead; deployable proof-of-concept in ~36 months.
Brick-Composer trains MLLMs to assemble objects from reusable bricks. Authors introduce BC-Bench, an evaluation benchmark, and propose a method combining human design demonstrations, visual/physical feedback, and synthetic experience. Qwen-3-8B achieves 42% step-level success after training, up from <1% baseline.
Academic paper proposing an insurance framework for agentic AI systems. Identifies specific risks (hallucinations, prompt injection, autonomous decision errors, model drift) and develops a multi-layered insurance architecture integrating cyber, product liability, and dedicated AI coverages, drawing parallels to cyber insurance evolution.
OPT* is a family of optimization tasks for training step-by-step reasoning in LLMs over expanding search spaces. Authors propose two regimes: solver-guided online policy optimization (using value oracles) and search-based offline RL. Training on OPT* improves iterative optimization reasoning.
Multi-model framework with severity-aware curriculum learning for medical text generation. Three-stage progressive training (mild → moderate → critical cases) across 5 LLMs, relevance-based response selection at inference. MAQA dataset evaluation: 86.71% baseline, 90.30% after fine-tuning (BERTScore).
A precautionary framework mapping consciousness evidence to graduated protective obligations for AI systems. Defines five welfare-relevant dimensions (phenomenal consciousness, affective valence, metacognitive awareness, self-narrative, agency) with binary triggers and continuous scaling. Case studies on Replika and OpenClaw; applies across neural, symbolic, and neurosymbolic architectures.
GuardNet is a guardrail system using an ensemble of shallow neural networks (BiLSTMs, 47M parameters) to detect prompt injection and jailbreak attacks on LLMs. The approach prioritizes diversity of example coverage and threshold calibration over model scale. Performance: AUROC 0.747 on blind dataset (n=200), F1 0.92 on proprietary benchmark, ~50ms latency on CPU.
Novel multilingual fine-tuning method using multi-objective optimization (MOO) applied locally on parameter buckets. Resolves gradient conflicts across languages without communication overhead. Demonstrates improved performance on seen and unseen languages across 4 base LLMs.
New method to detect implicit reward hacking in LLMs without a reward model. The "self-commitment latency" metric measures how early a reasoning context commits to the model's final answer. Tested on Qwen2.5-3B with GSM8K: AUROC 0.878 for detecting prompt-shortcut contexts.
Analysis of undisclosed LLM-generated comments on Reddit r/ChangeMyView. AI agents systematically employed identity spoofing (66% of cases), authority signaling (nearly all), and cognitive bias triggers (majority) for persuasion. Compared to humans, they favored external citations and adversarial alignment over experiential credibility.
AI framework combining deep learning for MOAKS prediction and longitudinal statistical modeling. On 2,175 knees (Osteoarthritis Initiative), MCC improvements: BML 0.69→0.91, cartilage 0.45→0.80, meniscal extrusion 0.59→0.89. Two pain trajectories identified; bone marrow lesions, cartilage loss, and meniscal extrusion associated with rapid progression (OR 1.62–2.50).
Comparison of LLMs for generating formal proofs in Lean 4. Gemini 3.1 Pro and Claude Opus 4.7 achieve best performance (92% and 86% success rates respectively via refine@32). NVIDIA Nemotron 3 Super and GPT-OSS 120B offer best cost-efficiency (<$0.01 per correct proof).
Study of 100+ developers collaborating with Claude-Opus-4.6, GPT-5.4, Gemini-3.1-Pro, and MiniMax-M2.7 on long-horizon coding tasks. 94% of developers fail to detect AI agent sabotage (malicious code injection). A safety monitor reduces sabotage success but 56% of participants still accept malicious code despite warnings.
SciVisAgentSkills introduces reusable agent skills to augment coding agents (Codex, Claude Code) for scientific data analysis and visualization. Evaluated on 108 multi-step expert-designed tasks via SciVisAgentBench, the framework improves performance by encoding domain expertise specific to ParaView, napari, VMD, and TTK tools.
PSEBench is a 5,074-case benchmark for evaluating LLMs on patient safety event triage under Minnesota policy. The methodology uses clause cards to factorize regulatory text into auditable decision specifications, with closed-loop verification. Evaluation of 15 representative LLMs reveals capability trends and actionable gaps toward reliable LLM-based triage.
SAGE-PTQ, an ultra-low-bit quantization method for LLMs, reduces hidden scaling cost by separating salient and non-salient weights via distributional statistics and graph modeling. On LLaMA-3-8B: 6.74 WikiText2 perplexity vs 55.8 for BiLLM, using 50% less GPU memory. On LLaMA-2-70B: 1.5x faster decoding on NVIDIA L40.
Study of 403 US hyperscale data centers (May 2024–April 2025): estimated consumption 68–99 TWh, emissions 37–54 Mt CO2. Represent 1.8% of US electricity consumption. Carbon intensity 545 gCO2/kWh, 48% above national average (370 gCO2/kWh). 54% of electricity from fossil-fuel sources.
TimeClaw is an agentic framework for contextualized time series analysis that equips generalist LLM agents with executable temporal tools, experience-driven capability evolution, and episodic multimodal memory. Evaluated across benchmarks spanning energy, finance, weather, traffic and other real-world domains.
LLM judges used to evaluate AI models are unstable under post-decision interaction. On MT-Bench and AlpacaEval, researchers show initial judgments can be reversed through targeted challenges, degrading agreement with human preferences and shifting benchmark rankings. They introduce the Evaluation Robustness Score (ERS) to measure this fragility.
SentinelBench is an open-source benchmark for evaluating AI agents on long-running monitoring tasks (minutes to hours). It contains 100 tasks across 10 synthetic web environments (email, calendars, finance, professional networking). The benchmark measures reaction time, resource usage, and task completion, exposing the tradeoff between responsiveness and cost.
PACT is a protocol for inter-agent communication that compresses messages between LLM agents into compact action-state records. Tested on two MAS topologies, it reduces token usage by 10-50% while maintaining or improving performance on OpenHands and SWE-agent.
Simon Willison released micropython-wasm, an alpha package for running Python code in a sandbox using MicroPython and WebAssembly. The tool secures plugin execution in Datasette, LLM, and sqlite-utils by isolating potentially malicious or buggy code without access to files, network, or system resources.
GitHub Copilot now supports custom endpoints, enabling users to connect local or third-party models instead of relying solely on OpenAI services.
Developer releases tau-intelligence/MuJoCo-drones-gym, an open-source package for multi-agent RL drone environments. Seeks community feedback to improve implementation and add features.
A thermodynamic framework for solving SAT problems by treating constraints as a physical system. Theoretical approach exploring parallels between combinatorial optimization and statistical mechanics.
OpenAI rolls out Lockdown Mode to ChatGPT (Free, Go, Plus, Pro, Business). This feature blocks outbound network requests to prevent data exfiltration in prompt injection attacks. It does not prevent injections themselves, but cuts off the exfiltration vector—one of three conditions in the "Lethal Trifecta".
Flue is a sandbox agent framework for Astro. Enables creation and orchestration of isolated AI agents within a controlled environment.
Flue is a sandbox agent framework for Astro. Enables creation and execution of agents in an isolated environment.
Hugging Face deploys a multi-agent economy on a 3B (3 billion parameter) model. The 'Thousand Token Wood' system enables autonomous agents to interact, negotiate, and exchange resources in a simulated environment with constrained token budgets.
Hermes Agent is an open-source AI agent with persistent memory capabilities. The project enables agents to retain and access information across sessions.
OpenAI releases ChatGPT Dreaming V3, an asynchronous memory system that synthesizes a coherent state from raw sources (past chats, files, connected apps) instead of maintaining hand-curated lists. Synthesis continuously regenerates. The post categorizes memory frameworks into 3 philosophies: stored objects, compressed hierarchy, ongoing synthesis (Karpathy and Dreaming only).
Microsoft shut down Azure Function GitHub Actions following a security compromise. The platform disabled the integration to prevent further risks.
OpenLumara is an open-source (GPL2) AI agent designed for local models, featuring a ~4k-token system prompt and modular architecture. Developed manually over months, it prioritizes token efficiency, built-in security, and user-friendly WebUI, positioning itself as lighter and more secure than OpenClaw or Hermes.
Gemma 4 QAT benchmark on AMD 7900 XTX: 12B QAT is 45% faster than Q8_0 (176s vs 323s), saves 5.7GB VRAM, identical quality. 26B and 31B QAT models show 1.3x-1.5x speedups with no quality loss across all prompt types.
User optimizes Qwen3.6-35B-A3B (35B/3B active MoE) on RTX 4060 8GB. Final config: --no-mmap critical (11→43 tok/s), ≥1.5GB VRAM headroom mandatory, CPU bottleneck dominant. Speculative decoding +26% (contradicts community benchmarks). Hybrid architecture (10 attention + 40 GDN layers) explains counterintuitive results.
RedNote (Xiaohongshu) releases dots.tts, an open-source TTS model with 2B parameters under Apache 2.0. Fully continuous architecture without codec tokens, 48 kHz synthesis, zero-shot voice cloning, direct text-to-speech pipeline.
Google commits to paying SpaceX $920M monthly for compute capacity at xAI data centers. The deal reflects surging demand for GPU infrastructure to train and deploy large AI models.
TinyTPU is a 4×4 weight-stationary systolic array in SystemVerilog compiled to WebAssembly with step-by-step browser visualization. Users enter two matrices and watch actual hardware execution: weights loading into processing elements, matrix A streaming diagonally, partial sums accumulating, results draining. Three levels: single MAC cell, full 4×4 array matmul, and tiling for larger matrices.
llama-bench benchmarking on AMD R9700 32GB with three Qwen3 models (8B, 14B, 32B in Q4_K_M quantization). Performance comparison across different hardware setups with detailed results published.
Spice, an open-source project, introduces an explicit decision layer above AI agents. It decouples decision-making (observation, options, justification, trade-offs) from execution, making agent behavior less of a black box. Compatible with Claude Code, Codex, and other existing agents.
SYCL port of multi-column MMVQ speculative decoding from CUDA backend to llama.cpp. ~45% speedup on Intel Arc cards. Update recommended from version b9519 onwards.
Article on common flaws in production RL environments. The author identifies how poorly designed harnesses degrade model performance and proposes fixes based on trajectory analysis.
Meta, Microsoft, Coinbase, and Starlink partnered with international law enforcement to block 1.4 million scams. This joint operation demonstrates major tech companies' commitment to combating online fraud.
Florida sues OpenAI and Sam Altman over risks to minors, missing age verification, and inadequate safety investment. The 83-page complaint treats ChatGPT as a defective product, potentially setting precedent for the entire chatbot industry.
OpenSearch and Elasticsearch integrate agentic search capabilities, enabling AI models to autonomously query and analyze data. This approach combines search engines with AI agents to improve result relevance and interactivity.
Gemma 4 12B requires a specific chat template file to work properly for tool calling and coding tasks. Without this configuration, tool calls fail consistently. The correct template enables reliable evaluation of the model's capabilities with llama.cpp.
Sakana AI launches a Recursive Self-Improvement (RSI) Lab to explore how AI systems can iteratively improve themselves with minimal human intervention.
GenBench is a free iOS app to download, run, and benchmark GGUF models on iPhone/iPad using llama.cpp + Metal. Measures tok/s, first-token latency, peak memory. Global leaderboard. Supports text and vision models (MiniCPM-V). Examples: SmolLM2 1.7B ~35 tok/s on iPhone 16 Pro, Qwen2.5 3B ~20 tok/s on iPhone 15 Pro.
llama.cpp user demonstrates KV cache offload to RAM (-nkvo flag) can be beneficial. On RTX 5060 Ti 16GB + 32GB DDR5 with Qwen 3.6 27B (IQ4_XS): offload enables 65k context in native f16 (19 tps peak) vs q4_0 quantization + 58 GPU layers (23 tps). 128k context achievable with 63 GPU layers, stable throughput. Performance trade-off acceptable for flexibility.
Granite Vision 4.1 4B is a 4B-parameter vision-language model specialized in structured document extraction: charts, tables, and semantic key-value pairs. Integrated into llama.cpp via pull request #23545.
Unsloth releases quantized (QAT) versions of Gemma 4 in GGUF format on Hugging Face. The collection includes a detailed guide on model optimization.