Back to feed
arXiv cs.CL·

Are you speaking my languages? On spoken language adherence in multimodal LLMs

Signal
72
Hype
18
In three linesLLM-based ASR systems often misidentify output languages in multilingual contexts. Authors propose three mitigation strategies: zero-shot prompting, supervised fine-tuning, and Chain-of-Thought reasoning to improve language adherence while preserving code-switching flexibility and ASR performance.
Read source
Your take?
VoicePrompt engineeringFine-tuningReasoningPapers

Summary generated by Claude — human-verified