Jiachen Ma (马嘉晨)

Fudan University & Shanghai AI Lab

mjc.jpg

I am a first-year Ph.D. student at Fudan University and Shanghai Artificial Intelligence Laboratory (PJLab), where I am privileged to be advised by Prof. Bowen Zhou and Prof. Chao Yang.

Prior to my doctoral studies, I earned my M.S. from Zhejiang University in 2025.

My research is driven by a fundamental commitment to Trustworthy Machine Learning. I am dedicated to exploring the boundaries of AI safety and ensuring that the next generation of intelligence remains beneficial and controllable.

🔍 Research Interests

I am broadly interested in the intersection of safety and general intelligence. Currently, I am dedicated to the following directions:

  • Value Alignment: Designing mechanisms that ensure large-scale models internalize human values and safety constraints fundamentally.
  • Autonomous Agents: Investigating the safety alignment and red teaming of agents in complex, long-horizon decision-making environments.
  • Explainability of Model Safety: Uncovering the internal mechanisms and representation-level dynamics that govern model (mis)behaviors to build more transparent safety frameworks.

“My vision is to make Safe AI — ensuring safety is not just a constraint, but a foundational property of intelligence.”

news

Jun 01, 2026 🎉 Reflector has been accepted by ICML 2026! Meanwhile, ReSaM is now under submission to NeurIPS 2027.
Sep 01, 2025 🚀 Excited to share that I have officially started my Ph.D. journey at Fudan University and Shanghai AI Lab! Working under the supervision of Prof. Bowen Zhou and Prof. Chao Yang, I will be dedicating my research to the vision: “Make Safe AI” — focusing on Value Alignment and the interpretability of model safety.
Mar 10, 2025 🎓 Officially graduated with a Master’s degree in Software Engineering from Zhejiang University. Honored to be recognized as an Outstanding Graduate Student (Top 5%). Gratitude to my advisors and colleagues at the Bigdata Lab!
Jan 10, 2025 📢 New Publication: Our work “Jailbreaking Prompt Attack (JPA)”, a controllable adversarial attack framework against Diffusion Models, has been accepted by NAACL 2025.
Jan 15, 2016 A simple inline announcement with Markdown emoji! :sparkles: :smile:

selected publications

  1. ICML’26
    reflect.jpg
    REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
    Jiachen Ma, Jiawen Zhang, Xiangtian Li, and 3 more authors
    ICML, 2026
  2. Arxiv
    resam.jpg
    ReSaM: Representation-Level Safety Margin Alignment for Vision-Language Models
    Jiachen Ma, Jiawen Zhang, Chaochao Lu, and 1 more author
    Arxiv, 2026
  3. NAACL’25
    jpa.jpg
    Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
    Jiachen Ma, Yijiang Li, Zhiqing Xiao, and 4 more authors
    In NAACL, 2025