Back to feed
arXiv cs.AI·

Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs

Signal
78
Hype
25
In three linesZEDD (Zero-Shot Embedding Drift Detection) detects prompt injections by measuring semantic shifts in embedding space between benign and suspect inputs. Without model internals access or retraining, the method achieves >93% accuracy on Llama 3, Qwen 2, Mistral with <3% false positive rate.
Read source
Your take?
AI safetyEmbeddingsPrompt engineeringBenchmarks

Summary generated by Claude — human-verified