Researchers have introduced Visual Contrastive Self-Distillation (VCSD), a new technique to improve vision-language models without external teachers or privileged information.
VCSD leverages the contrast between predictions on original and content-erased images to provide a more informative self-distillation signal for model training.
Tested on Qwen3 models, VCSD consistently outperformed traditional methods, suggesting a simpler and more effective approach to enhancing visual grounding in AI systems.