About me

I am a posdoc research scientist at Argonne National Lab. I have completed PhD in Computer Science at Virginia Tech, where I focus on developing robust, trustworthy multimodal and vision-language models that can withstand adversarial manipulation, backdooring, and misinformation attacks. My research spans adversarial robustness, multimodal alignment, secure fine-tuning, and safety evaluation in foundation models, with the broader goal of ensuring reliability and ethical deployment of AI at scale.

My recent work proposes defense strategies against multimodal poisoning attacks, designs universal perturbations for multi-image tasks, and builds fine-grained inconsistency detection systems capable of analyzing visual-textual contradictions in multimedia content. I also explore vulnerabilities and probing techniques in emerging web agents, as well as imperceptible adversarial signals for audio-vision systems.

I interned at GE Healthcare (2024), where I helped build one of the largest medical image-text datasets (2M pairs) across CT, MR, and ultrasound modalities, and contributed to designing a state-of-the-art multimodal medical foundation model. In 2025, I joined Futurewei Technologies as a research intern to develop self-evolving RAG architectures and reference-free hallucination mitigation frameworks for enterprise applications.

I am actively looking for full time applied scientist, research scientist, research engineer

πŸš€ Recent Highlights

August 2026: Our paper on multimodal multidocument event extraction and coreference resolution benchmark accepted in EMNLP 2026 Main

April 2026: Our paper on audio attack for trimodel model accepted in ACL 2026 Main

Feb 2026: Our paper on model immunization accepted in CVPR 2026 as Highlight

Dec 2025: Our paper on Universal Adversarial Perturbations for Multi-Image Tasks accepted in AAAI 2026

May 2025: Started internship on Futurewei Research

June 2024: Attented CVPR 2024 in Seattle, WA.

May 2024: Started my internship at GE Healthcare

Feb 2024: Semantic Shield β€” Fine-grained knowledge alignment to defend VLMs against backdoor and poisoning attacks accepted in CVPR 2024

August 2024: M3D β€” Multimodal, multidocument inconsistency detection for disinformation and knowledge verification accepted in EMNLP 2024

September 2024: JourneyBench β€” A large-scale benchmark for evaluating VLMs on generated images accepted in NeurIPS 2024