Back to feed
arXiv cs.AI·

Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

Signal
72
Hype
15
In three linesHSCHG method for open-vocabulary audio-visual event localization. Constructs hierarchical heterogeneous graph in Euclidean space with segment and video-level nodes, applies selective cross-modal fusion with dual-threshold filtering, and projects representations into hyperbolic space with hierarchical entailment regularization. Outperforms existing methods on OV-AVEL benchmark.
Read source
Your take?
VisionVoiceBenchmarksPapers

Summary generated by Claude — human-verified