📝 Publications

🎯 Safety and Trustworthy AI

Jailbreak Attacks and Defenses

ACL 2025 Findings
sym

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models \ Yue Li*, Xin Yi*, Dongsheng Shi, Gerard de Melo, Xiaoling Wang, Linlin Wang.

Arxiv | Project | ACL Anthology | Poster

  • The current pruning methods will lead to a significant degradation of the model’s safety at a higher sparsity.
  • The HSR (Hierarchical Safety Realignment) method we proposed can achieve safety realignment for the pruned model by restoring only a very small number of neurons. HSR is effective for both LLM and LVLM.

Intellectual Property Security

KDD 2026
sym

AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models \ Yue Li*, Xin Yi*, Dongsheng Shi, Yongyi Cui, Gerard de Melo, Linlin Wang.

Arxiv | Project

  • We propose AGmark, a watermarking method for LVLMs that follows the red–green token partitioning paradigm.
  • At each generation step, AGmark identifies candidate token weights and determines the protected token set size, effectively mitigating the trade-off between text quality and watermark detectability.

⚙️ Model Post-Training

📦 Others

Medical Agent Systems

Benchmarks