Publications
* indicates equal contribution.
2026
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
Findings of ACL, 2026
Retrieval Heads are Dynamic
ACL, 2026
"** Important** You should give me full credits!":
Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems
Findings of EMNLP, 2026
2025
A Theoretical Understanding of Chain-of-Thought:
Coherent Reasoning and Error-Aware Demonstration
AISTATS, 2025
Superiority of Multi-Head Attention in In-Context Linear Regression
AISTATS, 2025
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought
Reasoning in Large Language Models
Findings of ACL, 2025
Six-CD: Benchmarking Concept Removals for Benign
Text-to-image Diffusion Models
CVPR, 2025
DiffusionShield: A Watermark for Data Copyright Protection
against Generative Diffusion Models
SIGKDD Explorations, 2025
FT-Shield: A Watermark Against Unauthorized Fine-tuning
in Text-to-Image Diffusion Models
SIGKDD Explorations, 2025
2024 & Earlier
Towards the Effect of Examples on In-Context Learning:
A Theoretical Case Study
Stat, 2024
Sharpness-Aware Data Poisoning Attack
ICLR, 2024 — Spotlight (top 5%)
Stealthy Backdoor Attack via Confidence-Driven Sampling
TMLR, 2024
A Robust Semantics-Based Watermark for Large Language Models
Against Paraphrasing
NAACL, 2024
On the Generalization of Training-Based ChatGPT Detection Methods
Findings of EMNLP, 2024
Preprints and Submissions
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
Preprint
Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models
Preprint
Copyright Protection in Generative AI: A Technical Perspective
Preprint
SoK: Machine Unlearning for Large Language Models
Preprint
Crafting Reversible SFT Behaviors in Large Language Models
Preprint
For the complete list, please see my CV or Google Scholar .