About
I am a researcher at Alibaba Group, applying
RL / Multi-Agent RL / Agentic RL to e-commerce marketing and pricing.
My current focus is on RL-driven methods that advance LLM reasoning, planning, and long-horizon decision-making.
Previously at Xiamen University, I worked on
cooperative MARL and neural activation dynamics in Deep RL.
Exploration > exploitation, in life too.
Research Interests
- Reinforcement Learning — policy optimization, value factorization, RL for LLMs
- Multi-Agent RL — cooperation, credit assignment, parameter sharing
- AI Agent & Agentic RL — planning under uncertainty, role-aware orchestration, long-horizon reasoning
- AI for E-Commerce — dynamic pricing, value alignment, marketing decision automation
News
-
Jul 2026
Invited to serve as a Reviewer for AAAI 2027.
-
Jun 2026
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing
is selected as Oral at KDD 2026 ADS Track. 🎉
-
Jun 2026
Invited to serve as a Reviewer for TMLR.
-
May 2026
Recognized as a Gold Reviewer for ICML 2026.
-
May 2026
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing
accepted to KDD 2026 ADS Track.
-
Mar 2026
Invited to serve as a Reviewer for NeurIPS 2026.
-
Mar 2026
AIGP accepted to ICLR 2026 AIMS Workshop.
-
Feb 2026
RASO: Role-Aware Shared Reflection for Multi-Agent Orchestration in E-Commerce Long-Horizon Planning
accepted to ICAPS 2026.
-
Feb 2026
Invited to serve as a Reviewer for ICLR 2026 AIMS Workshop.
-
Jan 2026
SkillPrice: Semantic Skill Hierarchical Reinforcement Learning for Interpretable E-commerce Dynamic Price Recommendation
accepted to DASFAA 2026.
-
Dec 2025
Invited to serve as a Reviewer for ICML 2026.
-
Sep 2025
PlanU: Large Language Model Decision Making through Planning under Uncertainty
accepted to NeurIPS 2025.
-
Sep 2025
Invited to serve as a Reviewer for ICLR 2026.
Selected Publications
Author names in bold indicate myself. Full list on Google Scholar.
Chennan Ma, Yanning Zhang, Siqi Hong, Xiuchong Wang, Fei Xiao, Keping Yang
KDD 2026 ADS · Oral
An LLM-Agent framework that aligns short-term pricing with
long-term platform and user value in e-commerce marketing.
Paper
Chennan Ma, Yanning Zhang, Siqi Hong, Xiuchong Wang, Fei Xiao, Keping Yang, Bo Zheng
ICLR 2026 Workshop · AIMS
Workshop version presenting the long-term value alignment
formulation for LLM-driven e-commerce pricing.
Paper
Ziwei Deng, Mian Deng, Chengjing Liang, Zeming Gao, Chennan Ma, Chenxing Lin, Haipeng Zhang, Songzhu Mei, Siqi Shen, Cheng Wang
NeurIPS 2025
A planning-under-uncertainty framework that boosts LLM reasoning
via explicit belief modeling and search over uncertain outcomes.
Paper
Haoyuan Qin, Zhengzhu Liu, Chenxing Lin, Chennan Ma, Songzhu Mei, Siqi Shen, Cheng Wang
ICML 2025
Diagnoses and repairs futile neurons in parameter-sharing MARL networks,
improving cooperation and sample efficiency.
Paper ·
Code
Haoyuan Qin, Chennan Ma, Mian Deng, Zhengzhu Liu, Songzhu Mei, Xinwang Liu, Cheng Wang, Siqi Shen
NeurIPS 2024
Uncovers dormant neurons in MARL value factorization and proposes
a reactivation strategy that restores network capacity and performance.
Paper ·
Code
Siqi Shen, Chennan Ma, Chao Li, Weiquan Liu, Yongquan Fu, Songzhu Mei, Xinwang Liu, Cheng Wang
NeurIPS 2023
A risk-sensitive value factorization framework for cooperative MARL,
enabling reliable multi-agent decisions under uncertainty.
Paper ·
Code ·
Slides
→ See all publications on Google Scholar
Experience & Education
-
2025 – Present
Researcher, Alibaba Group · Taobao & Tmall Group
RL / MARL / Agentic RL for e-commerce marketing and pricing.
-
2022 – 2025
M.Eng. in Computer Science, Xiamen University · ASC Lab (spAtial Sensing and Computing)
Advisor: Associate Prof. Siqi Shen. Cooperative MARL & neural activation dynamics in Deep RL.
-
2018 – 2022
B.Eng. in Software Engineering, Jilin University
Awards & Honors
- 2026 KDD 2026 ADS Track — Oral Presentation
- 2026 ICML 2026 Gold Reviewer
Academic Service
Reviewer for NeurIPS, ICML, ICLR,
AAAI, and TMLR.