Topic

#GPT

GPT (Generative Pre-trained Transformer) is a family of language models trained on large text corpora to generate, summarize, or translate natural language content. OpenAI's GPT-4 is the most widely known instance, powering products such as ChatGPT.

40Articles
11Sources
66Avg. signal
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> smol-ai /</span> GodMode

GodMode is an AI chat browser providing fast, unified web access to ChatGPT, Claude, Bard, Bing, and Llama2. Productivity tool used multiple times daily.

ClaudeGPTTools
SIG
45
HYP
00
arXiv cs.CL·

Shopping Reasoning Bench: An Expert-Authored Benchmark for Multi-Turn Conversational Shopping Assistants

Shopping Reasoning Bench: expert-authored benchmark of 525 missions (232 single-turn, 293 multi-turn) with 10,863 importance-weighted binary rubrics for evaluating conversational shopping assistants. Evaluation of 9 models (GPT, Claude, Gemini): pass rates 57–77%, performance degrades 4–18 points across conversation turns, 13–29 point gap between required and optional criteria.

BenchmarksGPTClaude
SIG
75
HYP
00
Reddit r/MachineLearning·

Routing LLMs by task verifiability: a small experiment (n=120, 3 models) inspired by Karpathy's framework [D]

Experiment on 120 tasks testing whether weaker models match frontier models on high-verifiability tasks (Karpathy framework). Claude Sonnet 4.6, GPT 5.5, Mistral 3 8B compared. Code/structured extraction: narrower gaps with retry (Mistral 87%→95% code). Multi-hop reasoning: real capability gap (Sonnet 78%, Mistral 51%). Creative summarization: expected advantage for stronger models.

ClaudeGPTMistral
SIG
45
HYP
00
arXiv cs.AI·

Moonshine: An Autonomous Mathematical Research Agent Centered on Conjecture Generation

Moonshine is an autonomous agent generating mathematical conjectures by extracting structure from classical problems and formulating significant conjectures. Applied to the Jacobian conjecture, it transfers the logic to affine-ridge sigmoid networks, formulating the Neural Jacobian Conjecture (NJC). GPT-5.5-pro and DeepSeek-V4-pro obtained complete proofs for N=n+1.

AI AgentsReasoningPapers
SIG
72
HYP
00
Reddit r/MachineLearning·

LLM Relational Intelligence: A 4-Month Research Experiment on Multi-Model Behavioral Alignment with Human Communication [R]

4-month experiment testing whether context windows can be engineered so frontier models (GPT, Claude, Gemini, Grok) interact indistinguishably from human-to-human interaction. Gemini demonstrates highest relational intelligence. Author treats context window as behavioral environment rather than query interface, using modeling, accountability, humor, and social correction.

Prompt engineeringGPTClaude
SIG
35
HYP
00
arXiv cs.CL·

Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses

Evaluation study of LLMs (GPT-5.1, GPT-5 mini, Claude Sonnet 4.5 + Thinking, DeepSeek-V3.1) on their ability to generate multiple responses to the same scientific query while varying language complexity. On 98 queries, Claude Sonnet 4.5 maintains consistent complexity only 46% of the time. Evaluation framework based on formative study with 16 participants.

EvalsClaudeGPT
SIG
72
HYP
00
arXiv cs.CL·

Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

Purdue University deploys GPT-4o, GPT-5-mini, and GPT-5.2 to evaluate 1,200 applications for the SURF 2026 program. Models score statements of purpose across 6 rubric categories (0-3 scale), generating scores and rationales in 4.6 hours. GPT-5.2 shows strongest rubric adherence. Final coordinator review takes 4 hours versus multi-week effort in prior cycles.

GPTOpenAIEvals
SIG
72
HYP
00
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> 0x4m4 /</span> hexstrike-ai

HexStrike AI MCP Agents is an MCP server enabling AI agents (Claude, GPT, Copilot) to autonomously run 150+ cybersecurity tools for automated pentesting, vulnerability discovery, and security research.

MCPAI AgentsClaude
SIG
65
HYP
00