Topic

#Embeddings

Embeddings are numerical vector representations of text, images, or audio that capture their semantic meaning. For example, OpenAI's text-embedding-3-small model converts sentences into vectors used for search or similarity tasks.

40Articles
7Sources
69Avg. signal
arXiv cs.CL·

MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval

MCompassRAG improves RAG systems by using topic-level metadata as a semantic compass for paragraph-level retrieval. The method enriches chunk representations with topic signals in the same embedding space and trains a lightweight retriever via LLM-teacher distillation. Across six benchmarks, it gains 8.24% in information efficiency with 5× lower latency than efficient RAG baselines.

RAGEmbeddingsBenchmarks
SIG
78
HYP
00
arXiv cs.AI·

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

Multimodal fusion framework for time-to-event prediction (PE mortality, CVD outcomes) aligning CT and longitudinal EHR representations using foundation models. Four strategies tested (late fusion, contrastive alignment, cross-attention, co-attention) on 3,099–2,951 patients. Contrastive fusion improves concordance index by 1.5–5.4% vs unimodal baselines.

BenchmarksEmbeddingsVision
SIG
72
HYP
00
arXiv cs.AI·

A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis

Unified framework integrating PPO, time-series prediction, in-context learning, game theory, and cross-modal sentiment analysis for financial systems. Results: +23.7% portfolio optimization, -31.2% high-frequency trading error, +18.9% recommendation accuracy, +27.4% Nash convergence, +15.6% sentiment analysis.

Reinforcement learningBenchmarksReasoning
SIG
45
HYP
00
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> chroma-core /</span> chroma

Chroma is a vector search infrastructure for AI applications. The trending GitHub project provides storage and querying tools for embeddings to support RAG and language model-based systems.

Vector searchEmbeddingsRAG
SIG
65
HYP
00
Reddit r/LocalLLaMA·

Used local Ollama (gemma4:e4b + nomic-embed-text) to bulk-generate AI summaries for 4300 arXiv papers and push them to a remote Cloudflare DB — pipeline walkthrough

Developer built ArxivExplorer, semantic arXiv search engine with AI-generated summaries. Local pipeline uses Ollama: gemma4:e4b (8B) for structured JSON summaries, nomic-embed-text (137M) for 768-dim embeddings. 4300 papers processed, ~95% first-pass success rate, storage via Cloudflare D1/Vectorize. REST API 100× faster than wrangler.

RAGEmbeddingsOpen source
SIG
75
HYP
00
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> RyanCodrai /</span> turbovec

TurboVec is a vector index built on TurboQuant, written in Rust with Python bindings. Optimized for high-performance vector search.

Vector searchEmbeddingsOpen source
SIG
45
HYP
00
GitHub Trending·

<svg aria-hidden="true" data-component="Octicon" height="16" viewBox="0 0 16 16" version="1.1" width="16" data-view-component="true" class="octicon octicon-repo mr-1 tmp-mr-1 color-fg-muted"> <path d="M2 2.5A2.5 2.5 0 0 1 4.5 0h8.75a.75.75 0 0 1 .75.75v12.5a.75.75 0 0 1-.75.75h-2.5a.75.75 0 0 1 0-1.5h1.75v-2h-8a1 1 0 0 0-.714 1.7.75.75 0 1 1-1.072 1.05A2.495 2.495 0 0 1 2 11.5Zm10.5-1h-8a1 1 0 0 0-1 1v6.708A2.486 2.486 0 0 1 4.5 9h8ZM5 12.25a.25.25 0 0 1 .25-.25h3.5a.25.25 0 0 1 .25.25v3.25a.25.25 0 0 1-.4.2l-1.45-1.087a.249.249 0 0 0-.3 0L5.4 15.7a.25.25 0 0 1-.4-.2Z"></path> </svg> <span data-view-component="true" class="text-normal"> RyanCodrai /</span> turbovec

TurboVec is a vector index built on TurboQuant, written in Rust with Python bindings. Optimized for high-performance vector search.

Vector searchEmbeddingsOpen source
SIG
45
HYP
00
arXiv cs.LG·

Training-Free Lexical-Dense Fusion for Conversational-Memory Retrieval

Training-free lexical-dense fusion study for long-term conversational memory retrieval. Score-level fusion of late-interaction dense + BM25 improves Hit@1 by +8.8 to +17.2 points across six encoders (Hit@1 0.752 with e5-large-v2). Web search cross-encoder reranker degrades results (-6.9 pp). Analysis shows division of labor: dense excels on multi-hop/temporal questions, BM25 on adversarial ones.

RAGEmbeddingsBenchmarks
SIG
78
HYP
00
Embeddings — AI news · Signal IA