About me
I am a posdoc research scientist at Argonne National Lab. I have completed PhD in Computer Science at Virginia Tech, where I focus on developing robust, trustworthy multimodal and vision-language models that can withstand adversarial manipulation, backdooring, and misinformation attacks. My research spans adversarial robustness, multimodal alignment, secure fine-tuning, and safety evaluation in foundation models, with the broader goal of ensuring reliability and ethical deployment of AI at scale.
My recent work proposes defense strategies against multimodal poisoning attacks, designs universal perturbations for multi-image tasks, and builds fine-grained inconsistency detection systems capable of analyzing visual-textual contradictions in multimedia content. I also explore vulnerabilities and probing techniques in emerging web agents, as well as imperceptible adversarial signals for audio-vision systems.
I interned at GE Healthcare (2024), where I helped build one of the largest medical image-text datasets (2M pairs) across CT, MR, and ultrasound modalities, and contributed to designing a state-of-the-art multimodal medical foundation model. In 2025, I joined Futurewei Technologies as a research intern to develop self-evolving RAG architectures and reference-free hallucination mitigation frameworks for enterprise applications.
I am actively looking for full time applied scientist, research scientist, research engineer
π Recent Highlights
August 2026: Our paper on multimodal multidocument event extraction and coreference resolution benchmark accepted in EMNLP 2026 Main
April 2026: Our paper on audio attack for trimodel model accepted in ACL 2026 Main
Feb 2026: Our paper on model immunization accepted in CVPR 2026 as Highlight
Dec 2025: Our paper on Universal Adversarial Perturbations for Multi-Image Tasks accepted in AAAI 2026
May 2025: Started internship on Futurewei Research
June 2024: Attented CVPR 2024 in Seattle, WA.
May 2024: Started my internship at GE Healthcare
Feb 2024: Semantic Shield β Fine-grained knowledge alignment to defend VLMs against backdoor and poisoning attacks accepted in CVPR 2024
August 2024: M3D β Multimodal, multidocument inconsistency detection for disinformation and knowledge verification accepted in EMNLP 2024
September 2024: JourneyBench β A large-scale benchmark for evaluating VLMs on generated images accepted in NeurIPS 2024
