Jiachen Ma (马嘉晨)
I am a first-year Ph.D. student at Fudan University and Shanghai Artificial Intelligence Laboratory (PJLab), where I am privileged to be advised by Prof. Bowen Zhou and Prof. Chao Yang.
Prior to my doctoral studies, I earned my M.S. from Zhejiang University in 2025.
My research is driven by a fundamental commitment to Trustworthy Machine Learning. I am dedicated to exploring the boundaries of AI safety and ensuring that the next generation of intelligence remains beneficial and controllable.
🔍 Research Interests
I am broadly interested in the intersection of safety and general intelligence. Currently, I am dedicated to the following directions:
- Value Alignment: Designing mechanisms that ensure large-scale models internalize human values and safety constraints fundamentally.
- Autonomous Agents: Investigating the safety alignment and red teaming of agents in complex, long-horizon decision-making environments.
- Explainability of Model Safety: Uncovering the internal mechanisms and representation-level dynamics that govern model (mis)behaviors to build more transparent safety frameworks.
“My vision is to make Safe AI — ensuring safety is not just a constraint, but a foundational property of intelligence.”
news
| Jun 01, 2026 | 🎉 Reflector has been accepted by ICML 2026! Meanwhile, ReSaM is now under submission to NeurIPS 2027. |
|---|---|
| Sep 01, 2025 | 🚀 Excited to share that I have officially started my Ph.D. journey at Fudan University and Shanghai AI Lab! Working under the supervision of Prof. Bowen Zhou and Prof. Chao Yang, I will be dedicating my research to the vision: “Make Safe AI” — focusing on Value Alignment and the interpretability of model safety. |
| Mar 10, 2025 | 🎓 Officially graduated with a Master’s degree in Software Engineering from Zhejiang University. Honored to be recognized as an Outstanding Graduate Student (Top 5%). Gratitude to my advisors and colleagues at the Bigdata Lab! |
| Jan 10, 2025 | 📢 New Publication: Our work “Jailbreaking Prompt Attack (JPA)”, a controllable adversarial attack framework against Diffusion Models, has been accepted by NAACL 2025. |
| Jan 15, 2016 | A simple inline announcement with Markdown emoji! |
selected publications
- NAACL’25
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion ModelsIn NAACL, 2025